mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs repeatable evaluation contracts. > - Evaluation code must stay separate from provider launch and production orchestration. > - Offline fixtures need stable compatibility, scoring, traceability, and report rules. > - Published Runner consumers need only the supported evaluation contract surface. > - This pull request adds offline evaluation tooling and a workspace-private matrix kernel. > - The benefit is deterministic evaluation without credentials or paid provider calls. ## Linked Issues or Issue Description Refs #11297 This pull request extracts the offline evaluation unit from the earlier aggregate Runner work. ## What Changed - Add a workspace-private, provider-neutral evaluation matrix kernel. - Add the public `@paperclipai/paperclip-runner/evals` compatibility and native execution contracts. - Add fail-closed runnerd artifact and protocol compatibility checks. - Add deterministic workflow catalogs, scoring, traceability, and report generation. - Add sanitized Codex, OpenCode, and ACPX fixtures. - Add package-boundary and clean-consumer checks. - Add the eval package manifest to the Docker dependency stage. - Add the generated protocol fixture digest without changing the lockfile. ## Verification GitHub Actions must run: - Runner TypeScript and Rust type checks. - Runner unit and protocol tests. - Evaluation kernel tests. - Workflow traceability checks. - Clean-consumer and package-boundary checks. - Repository test, type-check, build, policy, and security gates. No local test command was run. The repository owner requested GitHub-only verification. ## Risks This is a large greenfield review surface with 51 files. The code does not launch a live provider or load credentials. Package and protocol drift fail closed. The workspace lockfile remains under the existing CI-owned process. ## Model Used OpenAI Codex with the GPT-5 agent model. The work used high reasoning, repository inspection, tool use, and parallel code review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
48 lines
2.3 KiB
Markdown
48 lines
2.3 KiB
Markdown
# Runner SDK and lab documentation
|
|
|
|
These documents cover the package-local browser SDK, React components,
|
|
standalone demos, scenario explorer, issue-thread surface, and live-session
|
|
development tools. They do not enable a production provider or change
|
|
Paperclip's runtime selection.
|
|
|
|
## Tutorials
|
|
|
|
- [Local runner](tutorials/local-runner.md)
|
|
- [Live console protocol server](tutorials/live-console-protocol-server.md)
|
|
- [Live console](tutorials/live-console.md)
|
|
- [SDK console](tutorials/sdk-console.md)
|
|
- [Standalone adapter](tutorials/standalone-thin-paperclip-adapter.md)
|
|
- [Scenario explorer](tutorials/capability-scenario-explorer.md)
|
|
- [Scenario chat](tutorials/scenario-chat.md)
|
|
- [Issue-thread lab](tutorials/capability-issue-thread.md)
|
|
- [Clean-room Codex chat](tutorials/capability-clean-room-chat.md)
|
|
|
|
## Reference
|
|
|
|
- [Runner eval scoring slice](capability-eval-slice.md)
|
|
- [Deterministic runner workflow evals](runner-workflow-evals.md)
|
|
- [Evals integration contract](evals-integration.md)
|
|
- [Local runner and supervision](local-runner.md)
|
|
- [Durable transport and recovery](durable-recovery.md)
|
|
- [Live console protocol server](live-console-protocol-server.md)
|
|
- [Live console](live-console.md)
|
|
- [Browser and React SDK](sdk.md)
|
|
- [Standalone adapter](standalone-thin-paperclip-adapter.md)
|
|
- [Scenario explorer](capability-scenario-explorer.md)
|
|
- [Scenario chat](scenario-chat.md)
|
|
- [Mock control plane](capability-mock-control-plane-port.md)
|
|
- [Live runnerd/Codex loop](capability-live-runnerd-codex.md)
|
|
- [Issue-thread UI](capability-issue-thread-ui.md)
|
|
- [Clean-room chat](capability-clean-room-chat.md)
|
|
- [Issue-thread UX contract](design/capability-issue-thread-ux-contract.md)
|
|
- [Scenario explorer UX](design/capability-scenario-explorer-ux.md)
|
|
- [Scenario chat UX](design/scenario-chat-ux.md)
|
|
- [Live-console interaction map](design/live-console-interaction-map.md)
|
|
- [Live-console component decisions](design/live-console-component-decisions.md)
|
|
- [SDK component decisions](design/sdk-component-decisions.md)
|
|
- [Standalone adapter boundary](design/standalone-thin-paperclip-adapter.md)
|
|
- [UI library compatibility note](research/2026-08-07-ui-library-compatibility.md)
|
|
|
|
The SDK and labs are additive. Existing direct adapters continue to use their
|
|
existing composer, transcript, interaction, and finalization paths.
|