Files
PaperClipAI/packages/paperclip-runner/spec/prp-v1-expressiveness-audit.md
T
Dotta 560e7e48b5 feat(runner): add SDK and developer tooling (#12608)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner package already provides the production protocol and
execution spine.
> - Contributors still need stable SDK surfaces, deterministic test
tools, and local inspection tools.
> - Those surfaces share generated contracts and must change as one
package boundary.
> - This pull request adds the package-local SDK, labs, examples, and
drift checks.
> - The benefit is a reviewable developer platform that does not change
application execution selection.

## Linked Issues or Issue Description

**Subsystem affected**

`packages/paperclip-runner` — runner SDK, conformance tools, and
developer tooling.

**Problem or motivation**

The production runner spine is present, but package consumers cannot
build deterministic integrations, inspect sessions, or verify
provider-neutral behavior through supported surfaces.

**Proposed solution**

Add browser, React, standalone, live-session, scenario, conformance, and
evaluation surfaces. Add generated contract inventories and
package-local verification scripts. Keep production application routing
unchanged.

**Alternatives considered**

We considered splitting each generated catalog, SDK surface, and demo
into separate pull requests. Those changes share exports, fixtures, and
drift gates. Splitting them would create intermediate package states
that do not build.

**Roadmap alignment**

No overlapping item appears in `ROADMAP.md`. This work extends the
runner package that is already on `master`.

## What Changed

- Add browser, React, standalone, live-session, and issue-thread SDK
surfaces.
- Add deterministic mock control-plane, scenario, conformance, replay,
and evaluation tools.
- Add bounded Codex, OpenCode, and ACPX development transports and
fixtures.
- Keep deferred managed-provider execution fail-closed. Persisted
compatibility data remains readable.
- Add generated capability inventories with their source files and drift
checks.
- Add examples, package documentation, browser checks, and
clean-consumer checks.
- Preserve the reviewed protocol bounds, replay compatibility aliases,
process environment isolation, and semantic redaction limits.
- Update the ACPX package patch that the existing workspace patch
registry already tracks.
- Do not change `pnpm-lock.yaml`, repository workflows, server runtime
selection, or the application UI.

## Verification

GitHub Actions is the verification authority for this pull request. The
repository CI, package TypeScript and Rust checks, package tests,
generated-output drift checks, browser checks, security scans, and
Greptile review must pass on the exact head.

Local test suites were not run because this series uses parallel GitHub
Actions for verification.

## Risks

This is a large greenfield package change. The main risks are public
export drift, generated-output drift, and optional React consumer
compatibility. Package boundary checks, clean-consumer checks, and
browser tests cover those risks. Production adapter selection and server
execution are outside this pull request.

## Stack

1. **This PR:** runner SDK and developer tooling.
2. [Codex production server
integration](https://github.com/paperclipai/paperclip/pull/12616).
3. [Provider-neutral task-thread
UI](https://github.com/paperclipai/paperclip/pull/12617).

## Model Used

OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the feature request
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-31 21:33:11 -05:00

6.1 KiB

PRP v1 expressiveness audit

Verdict

PRP v1 is sufficient for runner lifecycle, replay integrity, negotiated runner features, runtime/issue-thread request routing, structured work disposition, terminal causality, and provider-neutral semantic-control-plane receipts. The optional semantic_tool and terminal.stopReason envelopes now type tool invocation, authorization denial, redaction, optimistic concurrency, mutation receipts, artifacts, governed targets, causal continuations, and budget stops. They remain receipts for control-plane-owned operations, not a provider tool catalog or permission to move control-plane policy into the runner.

prp-v1-expressiveness-crosswalk.json is the machine-checked source for this verdict. Its Vitest gate rejects an unclassified required fact, a fact justified only by a permissive field, an unbounded control-plane-local fact, or an accepted additive change without required fixtures.

Evidence boundary

This audit compares the canonical capability authority and its traceability inputs, the real-surface ledger, the deterministic mock, live Codex/runnerd transport, and PRP schemas/replay fixtures. The wire contract is not the tool catalog: capability placement controls exposure, while PRP carries execution and causality. generic_api_request remains test-only and cannot satisfy any coverage row.

What v1 proves directly

  • capabilities.schema.json negotiates driver identity, session reuse, steering, interruption, resume, runtime requests, structured results, typed events, and explicit unsupported features.
  • Events have stable source instance/event ids, source ordering, run/session/ turn/item correlation, priority, and an explicit v1 schema version. Replay rejects unsupported required event versions, requires byte-equivalent duplicate source events, records source gaps, and is side-effect free.
  • Commands have controller ordering and preconditions; runtime and issue-thread requests have typed kinds/statuses; results require completion claims, verification, artifacts, and blocker/yield continuations where applicable.
  • run.terminal is a typed terminal envelope with terminal state, work assessment, and issue-status decision references. Existing golden replay, duplicate-event, unknown-optional-fields, and unsupported-required-version fixtures demonstrate the compatibility boundary.
  • capabilities.semanticTools advertises versioned, provider-neutral operation availability and redaction rules. Paired mcp_app.tool_input and mcp_app.tool_result events carry one safe semantic_tool envelope per call.
  • terminal.stopReason carries a safe budget/cost aggregate and the decision receipt that stopped the run. Eval trace_completeness reads these PRP wire receipts when available while retaining the existing live-evidence fallback.

Landed additive v1 envelope

The correction is one optional, versioned semantic_tool envelope for mcp_app.tool_input / mcp_app.tool_result, plus an optional terminal.stopReason envelope. It defines:

  • operation id, call id, correlation ids, idempotency key, outcome, stable code, retryability, and audit/operation receipt id;
  • admitted input/output references or digests, redaction disposition, and only safe identifiers; never raw credentials, hidden-company identifiers, or secret payloads;
  • authorization boundary (company, actor, active_task, grant, governed_action, lock, or revision), current revision where safe, and a deterministic conflict/duplicate result;
  • artifact/work-product references linked to the semantic receipt and terminal result; immutable interaction/approval document or decision targets; and budget stop reason/aggregate/decision id.

This remains additive because existing v1 consumers can ignore the optional typed envelopes. The reducer intentionally does not project them into session state; they supply inspectable trace evidence only. Unknown optional fields are accepted, while an unknown required envelope version fails closed.

Classification and fixture plan

The crosswalk assigns every required fact to exactly one of: direct v1, compositional v1 with a listed invariant, control-plane-local, additive v1, missing fixture/docs, or breaking v2 work. Direct and compositional support is never inferred merely from additionalProperties: true.

The additive change graduated to direct support with these fixtures and golden projections:

  1. semantic-tool-artifact-happy-path.json — artifact and work-product receipt;
  2. semantic-tool-denial-redaction.json — denied/redacted with no fallback;
  3. semantic-tool-conflict-duplicate-retry.json — stale conflict and exact retry;
  4. semantic-tool-governance-wake-monitor.json — immutable governed targets and continuation chain;
  5. budget-cost-stop-reason.json — typed terminal budget/cost stop;
  6. semantic-tool-unknown-optional-envelope.json — optional fields accepted with unchanged projection;
  7. semantic-tool-unsupported-required-version.json — required v2 envelope rejected fail-closed.

The six accepted fixtures replay to TypeScript/Rust parity summaries and golden snapshots. Shared high-risk semantic vectors compare normalized receipt and state-diff observations across deterministic mock and production bindings.

Explicit exclusions

Checkout/release, task selection, wake routing, budget enforcement, audit and run persistence/replay, assignment, and monitor management are control-plane actions, not runner-wire semantic operations. Cross-company discovery, broad audit access, destructive document lifecycle, and company administration are breaking v2/product-governance work. The runner must report typed unavailable or denied outcomes rather than tunnel those operations through generic payloads.

Review checklist

  • Golden replay, duplicate/retry, and unknown-version fixtures accompany every accepted protocol schema/event addition.
  • A semantic receipt is typed and safe before any surface is marked supported.
  • Control-plane-local decisions stay out of the semantic tool wire.
  • Eval trace scoring consumes wire receipts without changing projection semantics.