## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
5.7 KiB
Capability live runnerd and Codex loop
Reference lab, not the production sandbox topology. This API remains the runner-lab/session implementation used by UI and recovery tests. Production sandbox execution and live protocol evals use the Rust-owned bridge described in
../../../doc/plans/2026-08-20-single-daemon-runner-tool-bridge.md: external control plane → PRP → one Rustpaperclip-runnerd→ provider. The TypeScript dispatcher below does not run in the sandbox.
Capability binds the provider-neutral semantic catalog to a real package-local
paperclip-runnerd process and a real Codex app-server session. Paperclip data
remains deterministic mock state behind ControlPlanePort; no request reaches
the Paperclip API.
Process and authority boundary
The process chain is:
CapabilityLiveSession -> paperclip-runnerd -> codex app-server
-> CapabilitySemanticDispatcher -> ControlPlanePort mock
paperclip-runnerd owns the Codex child and proxies newline-delimited JSON-RPC
over stdio. The transport starts a dedicated Unix process group. Normal close,
stop/reset cleanup, and fatal protocol errors terminate that group with a
bounded TERM/KILL sequence so a Codex child is not abandoned.
Only the allowlisted Codex host environment is copied. PAPERCLIP_* variables,
provider credentials other than Codex's server-side home, and credentialed
proxy URLs are not passed to runnerd or Codex. Model-issued commands still use
the separate skillless, network-disabled workspace permission profile.
Stable session API
Production workers import the live entrypoint and bind a durable store to the attempt's immutable authority tuple:
import {
CapabilityLiveSessionService,
DurableCapabilityLiveSessionStore,
} from "@paperclipai/paperclip-runner/live";
const binding = { sessionId, runId, companyId, actorId, taskId };
const store = new DurableCapabilityLiveSessionStore({ directory, binding });
const service = new CapabilityLiveSessionService({ store });
const session = await service.create({ ...binding, attemptId, workingDirectory });
const turn = await session.sendMessage("Read the mock task and report progress.");
await session.interrupt(); // only when a turn is active
await service.stop(session.id);
After a worker terminates, a new worker must construct a new service and use the
production resume entrypoint. It must not call create again:
const resumed = await new CapabilityLiveSessionService({ store }).resume({
sessionId,
attemptId: resumeAttemptId,
resumeOf: killedAttemptId,
});
await resumed.reconcileActiveTurn();
resume first commits the killed attempt as terminated and the successor as
running. Only then does it start runnerd, read the checkpointed provider
thread, and resume that exact thread. A missing or corrupt checkpoint, authority
binding mismatch, attempt-lineage mismatch, or provider thread/session drift
fails closed. The killed attempt and its usage remain immutable.
The handoff surface for later tracks is:
create(input)starts runnerd, Codex, one mock run, and one dynamic-tool thread.sendMessage(text)supports repeated turns on the same provider thread.pendingInteractions()andresolveInteraction(input)preserve typed human interactions and return their results to that same thread.reconnect(sessionId)closes the old process group, starts a fresh runnerd and Codex app-server, then reads and resumes the persisted provider thread.restore(sessionId)recreates mock state, transcript, authority, authorization records, pending interactions, and the provider thread from a stored snapshot.resume({ sessionId, attemptId, resumeOf })is the cross-worker production path and records distinct linked attempts before provider recovery.recordUsage(receipt)durably commits an attempt-bound, exactly-once provider response receipt before the caller acknowledges that response. Reusing a receipt with different contents fails closed.reconcileActiveTurn()interrupts and records the terminal fact for a turn that was active in the checkpoint when its worker terminated.interrupt(reason)cancels only the active turn and retains session authority.service.stop(sessionId)clears authority and reaps the process group;resetalso deletes the old snapshot and restores the original clean mock seed under new run/session authority.
DurableCapabilityLiveSessionStore writes a checksummed, revisioned checkpoint
with atomic rename plus file and directory fsync. It persists provider identity,
mock state and semantic idempotency receipts, attempt lineage, active and
terminal turn facts, and the usage ledger. The included in-memory store remains
limited to tests and single-process consumers.
Every snapshot includes bounded transcript/evidence, serialized mock state,
semantic authorization records, runner/Codex PIDs and exit state, and explicit
network evidence. Tool calls are admitted only when their thread and turn match
the active Codex turn. Their typed CapabilitySemanticToolResult is serialized into
the app-server response, allowing Codex to use the resulting state revision in
its next response.
Verification
Run the deterministic contract suite:
pnpm --filter @paperclipai/paperclip-runner test:scenarios
Run a real runnerd and Codex app-server smoke:
pnpm --filter @paperclipai/paperclip-runner trace:live-runner -- --json
The smoke requires an authenticated local Codex installation. It checks a real semantic tool mutation, typed-result response, same-thread second turn, process ownership/cleanup, cleared authority, and zero Paperclip network/child-env exposure.