## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
5.0 KiB
Local runner: Run the Local Runner and Fake Harness
What this phase is
Local runner is a local live-run path. A TypeScript mock core starts a Rust runner. The Rust runner starts a scripted fake harness. The processes use JSON lines over stdio.
What this phase proves
This phase proves that a native session can run from start to finish without Paperclip, a network provider, or a real model. It proves cleanup, request resolution, interruption, bounded logs, one terminal result, and live/replay parity.
Prerequisites
- Node.js 24.11 or newer
- pnpm 9 or newer
- a stable Rust toolchain with
cargo - a Chromium-compatible browser
Start from a clean repository checkout and run every command from the repository root. Install the package dependencies and browser:
pnpm install --filter @paperclipai/paperclip-runner --lockfile=false --ignore-scripts --dev
pnpm --filter @paperclipai/paperclip-runner exec playwright install chromium
On a minimal Linux host, install the browser libraries too:
pnpm --filter @paperclipai/paperclip-runner exec playwright install-deps chromium
playwright install-deps requires root. On a Debian or Ubuntu host where root
is unavailable, use the package's rootless verification command instead. It
downloads the required browser-library packages into a user-owned cache,
extracts them without installing system packages, and scopes LD_LIBRARY_PATH
to the verification process:
pnpm --filter @paperclipai/paperclip-runner verify:rootless
1. Run the package verification path
pnpm --filter @paperclipai/paperclip-runner verify
Expected: Rust, TypeScript, protocol, boundary, documentation, and browser checks pass. The final Local runner line reports one terminal event and a successful semantic result.
2. Run the happy path in the CLI
pnpm --filter @paperclipai/paperclip-runner trace:local-runner -- --scenario happy-path --quiet
Expected summary facts:
eventCountis19;terminalCountis1;semanticResultisdone;- harness and runner exit codes are
0.
Capture the live trace as a validated replay fixture:
# Recorded evidence generation is deferred from this release.
The command writes
packages/paperclip-runner/.paperclip-local/evidence/local-runner-happy-path.json.
3. Run the scripted control flows
pnpm --filter @paperclipai/paperclip-runner trace:local-runner -- --scenario permission-input --quiet
pnpm --filter @paperclipai/paperclip-runner trace:local-runner -- --scenario interrupted --quiet
pnpm --filter @paperclipai/paperclip-runner trace:local-runner -- --scenario error --quiet
pnpm --filter @paperclipai/paperclip-runner trace:local-runner -- --scenario duplicate-terminal --quiet
pnpm --filter @paperclipai/paperclip-runner trace:local-runner -- --scenario happy-path --duplicate-turn-command --quiet
Expected:
- permission and input finish with
done; - interruption exits the harness with
130and reportsyielded; - scripted error exits the harness with
7and reportsyielded; - duplicate terminal still reports
terminalCount: 1; - duplicate turn command does not start a second turn.
4. Use the live browser
Start the package-local server:
pnpm --filter @paperclipai/paperclip-runner browser:dev --host 127.0.0.1 --port 4179
Open http://127.0.0.1:4179, then follow these steps:
- Keep
Happy pathselected and choose Start local run. - Confirm the terminal badge says
succeeded. - Confirm
Harness process exitis0andSemantic resultisdone. - Confirm
Live and replay reducer outputsaysMatch. - Select
Permission and input, start the run, and choose Allow. - Enter a trace name and choose Send input. Confirm the run completes. Each interactive request remains open for five minutes. If it expires, the page reports that the run finished and asks you to start a new run.
- Select
Interruption, start the run, and choose Interrupt turn when the button becomes active. - Confirm the terminal badge says
cancelled, the semantic result saysyielded, and the timeline contains onerun.terminalevent. - Switch to Static replay and confirm the Replay fixture path still works.
Stop Vite with Ctrl+C.
The browser regression suite writes its temporary screenshots under the
ignored packages/paperclip-runner/test-results/ directory. It never rewrites
the committed protocol fixtures.
5. Inspect the fixtures
What this proves
- the Rust runner owns and cleans the fake-harness process group;
- the fake driver emits deterministic live events and requests;
- logs are bounded;
- process exit and semantic result stay separate;
- duplicate commands and terminal proposals do not repeat effects;
- every browser event passes the Replay validator and reducer;
- replay after completion produces the same final snapshot.
Durable recovery must not begin until the Local runner human checkpoint is accepted.