Files
PaperClipAI/packages/paperclip-runner/docs/protocol-compatibility.md
T
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00

15 KiB

PRP Compatibility and Versioning Policy

Authority

The JSON Schema files in protocol/schemas/ are the language-neutral source of truth for the Replay and Local runner executable contract. The generated TypeScript schema module is checked against those files before every TypeScript typecheck. Rust consumes the same fixtures and must produce the same golden parity summaries.

The reviewable architecture and trust boundaries are defined in architecture.md; durable transport and recovery behavior is defined in durable-recovery.md. Local runner reuses the protocol contract for local live events. It adds package-local stdio and stream envelopes, but it does not add durable transport, persistence, or production control-plane behavior.

Version fields

Field Replay support Compatibility rule
protocolVersion 1 Required. Negotiate the highest overlapping version; no overlap fails closed.
fixtureVersion 1 Required by the conformance corpus. Unknown values fail closed.
event.schemaVersion 1 Required on every event. Unknown values fail closed before reduction.
capabilities.semanticTools.schemaVersion 1 Optional advertisement. When present, an unknown required version fails closed.
payload.semantic_tool.schemaVersion 1 Optional on paired semantic tool input/result events. When present, an unknown required version fails closed.
terminal.stopReason.schemaVersion 1 Optional budget/cost receipt. When present, an unknown required version fails closed.
Typed schema discriminators *.v1 Required. Unknown required schema identities fail JSON Schema validation.

Wire protocol versions and fixture-corpus versions are independent. A fixture format can evolve without changing PRP, and a future PRP version can be represented only after the consumer advertises support for it.

Forward compatibility

  • Unknown object properties are accepted and preserved by validation. Reducers ignore fields they do not understand until a later schema version gives those fields defined behavior.
  • Unknown required versions, schema discriminators, enum values, and required fields fail closed. A consumer must never guess at their semantics.
  • Scripted fixtures bind every event to the fixture run/session, require contiguous controller command order, exactly one unique proposed result, and exactly one unique terminal event.
  • The top-level fixture result must equal the run.result.proposed payload after canonical key ordering. Repeated sourceEventId deliveries must be byte-equivalent after the same normalization.

The forward-compatibility fixture proves that optional fields survive validation without changing the v1 snapshot. The unsupported-version fixture proves that a required v2 protocol cannot be replayed by this consumer.

Within-turn checklist snapshots

plan.updated / paperclip.plan.updated.v1 is a complete, ordered snapshot of the provider's checklist for one active turn. It is not a Paperclip Plan document and must never be inferred from assistant prose, Codex proposed-plan items, or generic TodoWrite output. Every replacement uses the provider turn ID as planId; PRP sourceSeq, not an optional provider revision, determines snapshot order. An empty step array clears the checklist, and complete is true only when a non-empty snapshot contains only completed steps. The legacy document-coupling fields are always syncStatus: "not_applicable" and documentRevision: null.

Qualified adapter profile Checklist support
Direct Codex App Server turn/plan/updated
ACPX Codex Structured ACP plan entries
ACPX Claude Structured ACP plan entries
ACPX Pi Unavailable; no production-qualified profile is exposed
OpenCode Unsupported until it exposes a structured plan event

Codex turn/diff/updated is normalized separately as the latest same-turn workspace.change.updated snapshot. Together these two independent event families can drive a turn-status UI without changing the public PRP family or adding a control-plane endpoint.

Provider-neutral semantic receipts

  • capabilities.semanticTools advertises stable operation IDs, availability, required claims, and redaction disposition without naming a provider API.
  • mcp_app.tool_input and mcp_app.tool_result may carry paired semantic_tool envelopes. Correlation IDs must match the containing event; operation ID and idempotency key must match across the pair.
  • Content is represented by a canonical SHA-256 digest plus allowlisted typed references. Raw credentials, provider payloads, and hidden identifiers do not belong on the wire.
  • Result receipts distinguish success, denial, conflict, exact duplicate, unavailable, and failure. They can name the authorization boundary, safe revision, artifact/work-product refs, immutable governed targets, and bounded wake/monitor causality.
  • terminal.stopReason records budget/cost kind, stable code, retryability, limit class, safe aggregate, and decision receipt.

These fields are trace evidence only. The v1 reducer ignores semantic_tool payloads, so adding or extending the optional envelope has no projection effect. Eval trace_completeness treats PRP wire receipts as authoritative when present and retains the pre-existing scalar fallback for live evidence that has not yet emitted them.

Provider-neutral structured input

Harness-initiated forms cross PRP as paperclip.runtime_request.v2 with requestKind: "runtime", type: "input", and an embedded paperclip.question_set.v1. A submission is always { "action": "submit", "response": paperclip.question_response.v1 }. Codex answer objects, OpenCode answer arrays, and ACP typed content exist only inside their adapters; origin may retain the provider method and adapter name for diagnostics but never provider response data.

Question and option order is significant, while answers are keyed by stable question IDs and selections reference stable option IDs. The canonical modes are text, single_select, and multi_select. Text validation is repeated at the untrusted server edge and again against the persisted question set before a provider receives the translated response.

Every harness adapter must add a paperclip.question_adapter_fixture.v1 fixture proving its native request normalizes to the canonical shape and its canonical response can be translated back. The shared fixture format deliberately contains both native and canonical objects so adding a provider does not change PRP or the UI contract.

V1 runtime requests and their legacy resolutions remain accepted during the migration. A request that contains no structured form stays on the legacy path; once a provider supplies a form, malformed or unsupported fields fail closed instead of silently degrading. ACPX sidecars advertise only form elicitation and use sidecar protocol v2 runtime.input_requested / input.resolve frames.

The live lifecycle pauses and resumes the same provider turn. If the provider process is lost first, Paperclip emits one non-replayable runtime_request.expired fact and materializes an idempotent durable ask_user_questions interaction using the identical question set. Explicit cancellation and already-resolved requests never create that fallback.

Native execution permission compatibility

paperclip.native-execution-input.v4 pins the effective harness permission policy in the closed provider configuration: approvalPolicy for Codex and permissionMode for OpenCode and ACPX. The pinned value participates in provider-session identity, so an incompatible idle or recovered session is replaced on the next execution. An active turn is never mutated in place.

paperclip.native-execution-input.v5 is the current input format. It retains the v4 permission pinning and adds optional completion-source references; paperclip.native-model-envelope.v3 is the corresponding explicit model projection. Readers continue to accept persisted v1-v4 inputs, and v5 readers must preserve the historical behavior of inputs that do not carry the new optional fields. Missing Codex and OpenCode policy fields retain their historical effective behavior. Legacy ACPX permissionPolicy: "interactive" is interpreted as approve-reads, while new v4 and v5 ACPX executions default to approve-all at the server boundary.

When a healthy provider session is resumed, its persisted v4 or v5 input format is retained even if the newly built input uses the other format. This avoids rotating an active session for a presentation-only schema change. A safe rollback from v5 to v4 removes only the optional completion-source references; it retains the task, contract, provider, workspace, and permission fields. Format changes still go through the normal provider-session identity checks, and an active turn is never mutated in place.

See Adding a harness for the permission catalog, isolation rules, and provider conformance requirements.

Seven conformance fixtures cover artifact success, redacted denial without fallback, stale conflict plus duplicate retry, governed target and continuation causality, budget/cost stop, unknown optional fields, and rejection of an unknown required version. The six accepted fixtures have shared TypeScript and Rust golden parity summaries.

Replay semantics

  • Events are applied in fixture order and ordered independently by (sourceKind, sourceInstanceId, sourceSeq).
  • A repeated source event ID has no second projection effect.
  • A forward source-sequence gap is recorded explicitly; the reducer never invents a missing event.
  • An event at or behind the committed source cursor is ignored and recorded as out of order.
  • Replaying an already-applied batch leaves the snapshot unchanged.

The CLI and browser import the same replayReplayFixtureText function, so validation, compatibility errors, and final snapshots cannot drift between the two surfaces.

Local envelope rules

  • Mock-core commands use paperclip.prp.command.v1 over stdin JSONL.
  • Runner output uses paperclip.runner.stream.v1 over stdout JSONL.
  • Fake-harness commands use paperclip.fake_harness.command.v1.
  • Fake-harness output uses paperclip.fake_harness.message.v1.
  • An equivalent repeated commandId returns a duplicate receipt and has no second driver effect. Reuse with different data is rejected.
  • A new command must use the next contiguous controllerSeq.
  • Harness logs are bounded diagnostic data. They are not canonical PRP events.
  • run.result.proposed, harness.exited, and run.terminal are separate facts and appear in that order when a semantic result exists.
  • The live browser rejects an event with an invalid schema, run ID, or session ID before it reaches the reducer.

These envelopes are local Local runner implementation contracts.

Durable wire rules

  • The runner opens loopback ws:// or hostname-verified wss://, or accepts a preview-proxy connection on its fixed listener, and completes the PRP v1 authenticated handshake before any command result or event.
  • A one-use bootstrap bearer capability returns a short-lived connection lease in welcome. Later connections use that lease. Neither raw capability is durable state.
  • welcome.payload.connectionLeaseRenewalVersion: 1 opts into authenticated lease_renew / lease_renewed control frames. Renewal extends the persisted expiry on the same live authority without restarting provider work. Identity, protocol, and revocation epoch remain fixed; expired or revoked leases cannot renew. See durable recovery for retry and warm-handoff rules. Peers lacking this capability retain their original lease expiry.
  • hello.resume reports the last processed controller sequence, next source sequence, cumulative ACK cursor, and current unacknowledged range.
  • welcome selects the one overlapping protocol version, returns the core's cumulative ACK cursor, and carries at most one durable pending command.
  • An event is durable before send. Event IDs and source sequences stay stable across replay and process restart.
  • An ACK is cumulative. The runner rejects a cursor behind its durable ACK or beyond its produced source cursor.
  • An equal repeated command ID and canonical digest returns its stored result. Reuse with different bytes fails closed and cannot repeat an effect.
  • Frames are bounded at 1 MiB and upgrade headers at 16 KiB. Unknown or invalid required protocol data fails closed; malformed JSON is a bounded diagnostic.

Runnerd build-metadata contract v2 advertises the exact transport inventory: dial_ws_loopback, dial_wss, and listen_ws. Plaintext dial destinations must resolve entirely to loopback. Public dial targets require TLS trust and hostname validation; a private CA bundle augments the platform roots and must be a bounded, private, regular file. Listener mode is fixed to port 43127 and a single run-bound path. All modes retain the same message/frame bounds and PRP authentication.

These are package-local Durable recovery and transport rules. Control-plane admission and deployment policy remain separately reviewed work.

Change policy

  1. Change JSON Schema first.
  2. Regenerate the TypeScript schema module.
  3. Add or revise a shared fixture and its golden snapshot/summary.
  4. Prove TypeScript and Rust parity.
  5. Update this policy and the normative spike specification when behavior changes.

Breaking changes require a new required version. Additive optional fields may remain in v1 only when old consumers can safely ignore them.

Package-level compatibility

PRP is one independently versioned component of the runner bundle. Catalog, runner-client, control-plane-adapter, testkit, and eval-corpus compatibility is declared by PAPERCLIP_RUNNER_COMPATIBILITY and checked before execution by assertPaperclipRunnerCompatibility. A mismatch fails with paperclip_runner_incompatible and stable per-issue codes; a provider-specific tool error is not a compatibility negotiation mechanism.

See ADR 0001 for the component rules and clean-consumer packaging gate.

Evals integration negotiation

The packed ./evals entry point adds a stricter execution preflight for the App/Evals join. assertPaperclipRunnerEvalCompatibility requires simultaneous agreement on package semver, runnerd build metadata, a common PRP version, semantic catalog version and SHA-256 digest, harness-driver contract and required capabilities, and the native-execution version. It reports paperclip_runner_eval_incompatible with expected/received values for every mismatch and must run before launching a provider.

runnerd itself reports paperclip-runner/runnerd-build-metadata/v1 from --build-metadata. The consumer passes its path and expected content digest to resolvePaperclipRunnerdArtifact; implicit PATH or source-tree discovery is not part of the contract. Native attempt output is paperclip-runner/native-execution/v1, whose parser accepts unknown additive fields but rejects unknown required versions and inconsistent terminal, semantic-denial, usage, or transcript facts. Full fields and the deterministic gate are exercised by the package-local deterministic conformance suite.