Files
PaperClipAI/packages/paperclip-runner/docs/capability-execution-modes.md
T
Dotta 560e7e48b5 feat(runner): add SDK and developer tooling (#12608)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner package already provides the production protocol and
execution spine.
> - Contributors still need stable SDK surfaces, deterministic test
tools, and local inspection tools.
> - Those surfaces share generated contracts and must change as one
package boundary.
> - This pull request adds the package-local SDK, labs, examples, and
drift checks.
> - The benefit is a reviewable developer platform that does not change
application execution selection.

## Linked Issues or Issue Description

**Subsystem affected**

`packages/paperclip-runner` — runner SDK, conformance tools, and
developer tooling.

**Problem or motivation**

The production runner spine is present, but package consumers cannot
build deterministic integrations, inspect sessions, or verify
provider-neutral behavior through supported surfaces.

**Proposed solution**

Add browser, React, standalone, live-session, scenario, conformance, and
evaluation surfaces. Add generated contract inventories and
package-local verification scripts. Keep production application routing
unchanged.

**Alternatives considered**

We considered splitting each generated catalog, SDK surface, and demo
into separate pull requests. Those changes share exports, fixtures, and
drift gates. Splitting them would create intermediate package states
that do not build.

**Roadmap alignment**

No overlapping item appears in `ROADMAP.md`. This work extends the
runner package that is already on `master`.

## What Changed

- Add browser, React, standalone, live-session, and issue-thread SDK
surfaces.
- Add deterministic mock control-plane, scenario, conformance, replay,
and evaluation tools.
- Add bounded Codex, OpenCode, and ACPX development transports and
fixtures.
- Keep deferred managed-provider execution fail-closed. Persisted
compatibility data remains readable.
- Add generated capability inventories with their source files and drift
checks.
- Add examples, package documentation, browser checks, and
clean-consumer checks.
- Preserve the reviewed protocol bounds, replay compatibility aliases,
process environment isolation, and semantic redaction limits.
- Update the ACPX package patch that the existing workspace patch
registry already tracks.
- Do not change `pnpm-lock.yaml`, repository workflows, server runtime
selection, or the application UI.

## Verification

GitHub Actions is the verification authority for this pull request. The
repository CI, package TypeScript and Rust checks, package tests,
generated-output drift checks, browser checks, security scans, and
Greptile review must pass on the exact head.

Local test suites were not run because this series uses parallel GitHub
Actions for verification.

## Risks

This is a large greenfield package change. The main risks are public
export drift, generated-output drift, and optional React consumer
compatibility. Package boundary checks, clean-consumer checks, and
browser tests cover those risks. Production adapter selection and server
execution are outside this pull request.

## Stack

1. **This PR:** runner SDK and developer tooling.
2. [Codex production server
integration](https://github.com/paperclipai/paperclip/pull/12616).
3. [Provider-neutral task-thread
UI](https://github.com/paperclipai/paperclip/pull/12617).

## Model Used

OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the feature request
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-31 21:33:11 -05:00

4.5 KiB

Capability execution modes and identity

Capability runs the same semantic tool catalog, authorization engine, and mock ControlPlanePort under two agent modes. Whichever mode is active, the surface always names three separate actors so a reader never mistakes one for another.

Three actors, always separate

Actor What it is in Capability What it is not
Real Codex A real Codex app-server session driving the turn loop. The agent has no Paperclip skill; it sees only the semantic tools the scenario exposes. Not a scripted stand-in, and not a Paperclip-aware agent.
Real runnerd The real package-local paperclip-runnerd binary. It owns the Codex child process group and proxies newline-delimited JSON-RPC over stdio. Not an in-process fake and not the production Paperclip runtime.
Mock Paperclip The deterministic in-process ControlPlanePort adapter. It holds every issue, comment, document, interaction, approval, and audit record as mock state. Not the Paperclip control plane, database, or API. No request leaves the process for a Paperclip service.

The live surface renders a Real Codex, Real runnerd, and Mock Paperclip marker at all times. Mock records carry an MCK- identifier prefix so a mock issue is never confused with a real PAP- issue.

Two agent modes

Scripted (deterministic) — fake mode

The scripted driver replays a recorded conversation over the same mock core. It needs no provider credential, no runnerd process, and no network after pnpm install. It is the mode used for:

  • the 106-case conformance suite and its bounded parity report;
  • the byte-stable screenshot matrix;
  • replay of a recorded session.

Because it is fully deterministic, two runs of the same route produce byte-identical output. A scripted (mode=fake) artifact cannot satisfy a live acceptance criterion — it proves the contract, not a real provider turn.

Codex (bounded) — live mode

Live mode starts a real paperclip-runnerd process and a real Codex app-server session. The package server — never the browser — owns the Codex process, the session, the tool loop, and provider authentication. The browser receives no provider, runner, mock-control-plane, or real-control-plane credential.

Live mode requires a locally authenticated Codex installation. When that relay is not present, the UI keeps Codex disabled with a named reason and scripted mode stays available. Live mode is the only mode eligible for the final Revision 3/4 acceptance criteria that require a real provider turn.

Mode-independent controls

These behave the same in either mode because they operate on the mock core and the runner session, not on the agent:

  • Reset restores the original clean mock seed under new run/session authority and rotates or clears the old session authority. It does not affect another browser session's state.
  • Stop cancels only the active turn and reaps the runnerd process group; the session and its transcript survive. Cleanup evidence shows no abandoned child process and no active session afterward.
  • Replay reproduces a recorded canonical timeline. It is always scripted and always labelled as fake-derived, even inside a session that also ran a live turn.
  • Refresh / reconnect restores the same durable session, pending interaction, transcript, and mock state. Reconnect starts a fresh runnerd and Codex app-server and resumes the persisted provider thread; the session does not restart because the browser reconnected.

Two entry points, one live mode

Live mode is reachable from both primary surfaces:

  • the preset scenario explorer, with mode=live on a chosen scenario;
  • the clean-room chat at #/chat, which has no other mode. It starts blank on a freshly minted mock tenant, pins mode=live in the route, and reports a failure to start real Codex rather than falling back to a fixture or a recording.

Eligibility summary

Evidence Eligible for
Scripted / mode=fake run 106-case conformance, replay determinism, screenshot matrix, contract proofs
Live Codex run on the final build Revision 3/4 final acceptance criteria that require a real provider turn
Revision 2 preview (build da0d32d74a, historical URL) Historical comparison only — not eligible for final acceptance

See the live runnerd and Codex loop reference for the session API and verification commands behind each row above.