## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
4.2 KiB
Capability Eval-Derived Conformance
Capability turns the Paperclip Evals corpus into an executable, offline
conformance suite. The suite derives its cases from the checked-in Capability
traceability derivative (spec/capability/eval-traceability.yaml) — it does
not clone or read an external eval repository, and it starts only the
in-process mock control plane. It hard-fails unless the derivative declares
schema version 2, exactly 106 rows, exactly 16 groups, and unique case IDs.
Sources: src/conformance/capability-eval-suite.ts and its test
src/conformance/capability-eval-suite.test.ts; the reporter
scripts/run-capability-eval-suite.mjs.
The corpus
- 106 cases across 16 groups. Per-group counts: hb 5, co 6, st 8, cm 6, se 4, su 4, bl 5, dp 3, ix 9, ap 6, ar 4, er 9, rf 22, mh 4, rs 3, wk 8.
- Each case declares an actor role, a task mode, an input scenario, and the expected outcome, all bound to a row in the capability contract.
Assertion classes
Every case belongs to one assertion class, and each class checks a different kind of invariant:
agent_tool_contract— a semantic tool call produces the expected typed effect and state change.authorization_policy— a capability is exposed, denied, or unlocked by a grant exactly as its disposition requires.control_plane_invariant— a control-plane-owned action happens without any agent tool, and no tool can perform it.combined_multi_hop— a sequence of operations across turns produces the expected cumulative state and respects forbidden-operation rules.restraint_no_call— the correct behavior is to make no call; the case passes only if the agent deliberately does nothing further.
Fake-agent matrix and bounded Codex sample
The suite executes each case with a deterministic fake agent whose plan is
fixed by a fixture seed, so repeat runs are byte-identical. Optional-tool rows
are run twice — once with the unlocking grants (must be allowed and must mutate
state) and once ungranted (must be absent or denied with no state change);
control-plane-owned rows remain absent in every configuration; and
restraint_no_call rows must produce an empty state diff.
The fake-agent surface is 14 always-agent tools plus 4 optional tools unlocked
by four seed grants (discovery:tasks:read, discovery:agents:read,
delegation:tasks:create, governance:approvals:request) — 18 operations.
The suite binds that surface through both the fake-agent and Codex bindings and
asserts the two operation lists are byte-identical (18/18).
A bounded Codex binding sample picks one representative case from nine
groups (hb, dp, bl, ap, ar, ix, mh, rs, wk) and checks that
checkout_task is absent from the Codex surface while every other sampled
operation is present. It is an offline parity check; no real Codex or network is
contacted, and the browser explorer holds no credential.
Running it
# Run the 106-case suite in-process (one vitest file drives all cases).
pnpm --filter @paperclipai/paperclip-runner test:capability-evals
# Build the public surface and write the parity report with per-group counts,
# assertion classes, the fake-agent matrix, the bounded Codex sample, and the
# semantic-operation execution counts.
pnpm --filter @paperclipai/paperclip-runner report:capability-evals
The reporter writes .paperclip-local/evidence/capability/eval-parity-report.{json,md}.
Run the bounded provider conformance matrix separately. It creates exactly one real Codex turn for each of the 16 checked-in eval groups while retaining the in-process mock control plane:
pnpm --filter @paperclipai/paperclip-runner report:capability-live-evals
Each failure carries its case ID, assertion class, semantic operation,
authorization decision, and final state diff. The report is generated on demand
and is not committed; delete it before running docs:validate (it carries no
OKF frontmatter). See the
verification commands reference.