Files
PaperClipAI/packages/paperclip-runner/docs/protocol-compatibility.md
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00

282 lines
15 KiB
Markdown

# PRP Compatibility and Versioning Policy
## Authority
The JSON Schema files in [`protocol/schemas/`](../protocol/schemas/) are the
language-neutral source of truth for the Replay and Local runner executable contract. The
generated TypeScript schema module is checked against those files before every
TypeScript typecheck. Rust consumes the same fixtures and must produce the same
golden parity summaries.
The reviewable architecture and trust boundaries are defined in
[`architecture.md`](./architecture.md); durable transport and recovery behavior
is defined in [`durable-recovery.md`](./durable-recovery.md). Local runner reuses
the protocol contract for local live events. It adds package-local stdio and
stream envelopes, but it does not add durable transport, persistence, or
production control-plane behavior.
## Version fields
| Field | Replay support | Compatibility rule |
|---|---:|---|
| `protocolVersion` | `1` | Required. Negotiate the highest overlapping version; no overlap fails closed. |
| `fixtureVersion` | `1` | Required by the conformance corpus. Unknown values fail closed. |
| `event.schemaVersion` | `1` | Required on every event. Unknown values fail closed before reduction. |
| `capabilities.semanticTools.schemaVersion` | `1` | Optional advertisement. When present, an unknown required version fails closed. |
| `payload.semantic_tool.schemaVersion` | `1` | Optional on paired semantic tool input/result events. When present, an unknown required version fails closed. |
| `terminal.stopReason.schemaVersion` | `1` | Optional budget/cost receipt. When present, an unknown required version fails closed. |
| Typed `schema` discriminators | `*.v1` | Required. Unknown required schema identities fail JSON Schema validation. |
Wire protocol versions and fixture-corpus versions are independent. A fixture
format can evolve without changing PRP, and a future PRP version can be
represented only after the consumer advertises support for it.
## Forward compatibility
- Unknown object properties are accepted and preserved by validation. Reducers
ignore fields they do not understand until a later schema version gives those
fields defined behavior.
- Unknown required versions, schema discriminators, enum values, and required
fields fail closed. A consumer must never guess at their semantics.
- Scripted fixtures bind every event to the fixture run/session, require
contiguous controller command order, exactly one unique proposed result, and
exactly one unique terminal event.
- The top-level fixture result must equal the `run.result.proposed` payload after
canonical key ordering. Repeated `sourceEventId` deliveries must be
byte-equivalent after the same normalization.
The forward-compatibility fixture proves that optional fields survive validation
without changing the v1 snapshot. The unsupported-version fixture proves that a
required v2 protocol cannot be replayed by this consumer.
## Within-turn checklist snapshots
`plan.updated` / `paperclip.plan.updated.v1` is a complete, ordered snapshot of
the provider's checklist for one active turn. It is not a Paperclip Plan
document and must never be inferred from assistant prose, Codex proposed-plan
items, or generic TodoWrite output. Every replacement uses the provider turn ID
as `planId`; PRP `sourceSeq`, not an optional provider revision, determines
snapshot order. An empty step array clears the checklist, and `complete` is true
only when a non-empty snapshot contains only `completed` steps. The legacy
document-coupling fields are always `syncStatus: "not_applicable"` and
`documentRevision: null`.
| Qualified adapter profile | Checklist support |
|---|---|
| Direct Codex App Server | `turn/plan/updated` |
| ACPX Codex | Structured ACP `plan` entries |
| ACPX Claude | Structured ACP `plan` entries |
| ACPX Pi | Unavailable; no production-qualified profile is exposed |
| OpenCode | Unsupported until it exposes a structured plan event |
Codex `turn/diff/updated` is normalized separately as the latest same-turn
`workspace.change.updated` snapshot. Together these two independent event
families can drive a turn-status UI without changing the public PRP family or
adding a control-plane endpoint.
## Provider-neutral semantic receipts
- `capabilities.semanticTools` advertises stable operation IDs, availability,
required claims, and redaction disposition without naming a provider API.
- `mcp_app.tool_input` and `mcp_app.tool_result` may carry paired
`semantic_tool` envelopes. Correlation IDs must match the containing event;
operation ID and idempotency key must match across the pair.
- Content is represented by a canonical SHA-256 digest plus allowlisted typed
references. Raw credentials, provider payloads, and hidden identifiers do not
belong on the wire.
- Result receipts distinguish success, denial, conflict, exact duplicate,
unavailable, and failure. They can name the authorization boundary, safe
revision, artifact/work-product refs, immutable governed targets, and bounded
wake/monitor causality.
- `terminal.stopReason` records budget/cost kind, stable code, retryability,
limit class, safe aggregate, and decision receipt.
These fields are trace evidence only. The v1 reducer ignores `semantic_tool`
payloads, so adding or extending the optional envelope has no projection
effect. Eval `trace_completeness` treats PRP wire receipts as authoritative when
present and retains the pre-existing scalar fallback for live evidence that has
not yet emitted them.
## Provider-neutral structured input
Harness-initiated forms cross PRP as `paperclip.runtime_request.v2` with
`requestKind: "runtime"`, `type: "input"`, and an embedded
`paperclip.question_set.v1`. A submission is always
`{ "action": "submit", "response": paperclip.question_response.v1 }`.
Codex answer objects, OpenCode answer arrays, and ACP typed content exist only
inside their adapters; `origin` may retain the provider method and adapter name
for diagnostics but never provider response data.
Question and option order is significant, while answers are keyed by stable
question IDs and selections reference stable option IDs. The canonical modes
are `text`, `single_select`, and `multi_select`. Text validation is repeated at
the untrusted server edge and again against the persisted question set before a
provider receives the translated response.
Every harness adapter must add a
`paperclip.question_adapter_fixture.v1` fixture proving its native request
normalizes to the canonical shape and its canonical response can be translated
back. The shared fixture format deliberately contains both native and canonical
objects so adding a provider does not change PRP or the UI contract.
V1 runtime requests and their legacy resolutions remain accepted during the
migration. A request that contains no structured form stays on the legacy path;
once a provider supplies a form, malformed or unsupported fields fail closed
instead of silently degrading. ACPX sidecars advertise only form elicitation
and use sidecar protocol v2 `runtime.input_requested` / `input.resolve` frames.
The live lifecycle pauses and resumes the same provider turn. If the provider
process is lost first, Paperclip emits one non-replayable
`runtime_request.expired` fact and materializes an idempotent durable
`ask_user_questions` interaction using the identical question set. Explicit
cancellation and already-resolved requests never create that fallback.
## Native execution permission compatibility
`paperclip.native-execution-input.v4` pins the effective harness permission
policy in the closed provider configuration: `approvalPolicy` for Codex and
`permissionMode` for OpenCode and ACPX. The pinned value participates in
provider-session identity, so an incompatible idle or recovered session is
replaced on the next execution. An active turn is never mutated in place.
`paperclip.native-execution-input.v5` is the current input format. It retains
the v4 permission pinning and adds optional completion-source references;
`paperclip.native-model-envelope.v3` is the corresponding explicit model
projection. Readers continue to accept persisted v1-v4 inputs, and v5 readers
must preserve the historical behavior of inputs that do not carry the new
optional fields. Missing Codex and OpenCode policy fields retain their
historical effective behavior. Legacy ACPX `permissionPolicy: "interactive"`
is interpreted as `approve-reads`, while new v4 and v5 ACPX executions default
to `approve-all` at the server boundary.
When a healthy provider session is resumed, its persisted v4 or v5 input format
is retained even if the newly built input uses the other format. This avoids
rotating an active session for a presentation-only schema change. A safe
rollback from v5 to v4 removes only the optional completion-source references;
it retains the task, contract, provider, workspace, and permission fields.
Format changes still go through the normal provider-session identity checks,
and an active turn is never mutated in place.
See [Adding a harness](adding-a-harness.md) for the permission catalog,
isolation rules, and provider conformance requirements.
Seven conformance fixtures cover artifact success, redacted denial without
fallback, stale conflict plus duplicate retry, governed target and continuation
causality, budget/cost stop, unknown optional fields, and rejection of an
unknown required version. The six accepted fixtures have shared TypeScript and
Rust golden parity summaries.
## Replay semantics
- Events are applied in fixture order and ordered independently by
`(sourceKind, sourceInstanceId, sourceSeq)`.
- A repeated source event ID has no second projection effect.
- A forward source-sequence gap is recorded explicitly; the reducer never
invents a missing event.
- An event at or behind the committed source cursor is ignored and recorded as
out of order.
- Replaying an already-applied batch leaves the snapshot unchanged.
The CLI and browser import the same `replayReplayFixtureText` function, so
validation, compatibility errors, and final snapshots cannot drift between the
two surfaces.
## Local envelope rules
- Mock-core commands use `paperclip.prp.command.v1` over stdin JSONL.
- Runner output uses `paperclip.runner.stream.v1` over stdout JSONL.
- Fake-harness commands use `paperclip.fake_harness.command.v1`.
- Fake-harness output uses `paperclip.fake_harness.message.v1`.
- An equivalent repeated `commandId` returns a duplicate receipt and has no
second driver effect. Reuse with different data is rejected.
- A new command must use the next contiguous `controllerSeq`.
- Harness logs are bounded diagnostic data. They are not canonical PRP events.
- `run.result.proposed`, `harness.exited`, and `run.terminal` are separate facts
and appear in that order when a semantic result exists.
- The live browser rejects an event with an invalid schema, run ID, or session
ID before it reaches the reducer.
These envelopes are local Local runner implementation contracts.
## Durable wire rules
- The runner opens loopback `ws://` or hostname-verified `wss://`, or accepts a
preview-proxy connection on its fixed listener, and completes the PRP v1
authenticated handshake before any command result or event.
- A one-use bootstrap bearer capability returns a short-lived connection lease
in `welcome`. Later connections use that lease. Neither raw capability is
durable state.
- `welcome.payload.connectionLeaseRenewalVersion: 1` opts into authenticated
`lease_renew` / `lease_renewed` control frames. Renewal extends the persisted
expiry on the same live authority without restarting provider work. Identity,
protocol, and revocation epoch remain fixed; expired or revoked leases cannot
renew. See [durable recovery](durable-recovery.md#execution-duration-and-operation-deadlines)
for retry and warm-handoff rules. Peers lacking this capability retain their
original lease expiry.
- `hello.resume` reports the last processed controller sequence, next source
sequence, cumulative ACK cursor, and current unacknowledged range.
- `welcome` selects the one overlapping protocol version, returns the core's
cumulative ACK cursor, and carries at most one durable pending command.
- An event is durable before send. Event IDs and source sequences stay stable
across replay and process restart.
- An ACK is cumulative. The runner rejects a cursor behind its durable ACK or
beyond its produced source cursor.
- An equal repeated command ID and canonical digest returns its stored result.
Reuse with different bytes fails closed and cannot repeat an effect.
- Frames are bounded at 1 MiB and upgrade headers at 16 KiB. Unknown or invalid
required protocol data fails closed; malformed JSON is a bounded diagnostic.
Runnerd build-metadata contract v2 advertises the exact transport inventory:
`dial_ws_loopback`, `dial_wss`, and `listen_ws`. Plaintext dial destinations
must resolve entirely to loopback. Public dial targets require TLS trust and
hostname validation; a private CA bundle augments the platform roots and must
be a bounded, private, regular file. Listener mode is fixed to port 43127 and a
single run-bound path. All modes retain the same message/frame bounds and PRP
authentication.
These are package-local Durable recovery and transport rules. Control-plane
admission and deployment policy remain separately reviewed work.
## Change policy
1. Change JSON Schema first.
2. Regenerate the TypeScript schema module.
3. Add or revise a shared fixture and its golden snapshot/summary.
4. Prove TypeScript and Rust parity.
5. Update this policy and the normative spike specification when behavior
changes.
Breaking changes require a new required version. Additive optional fields may
remain in v1 only when old consumers can safely ignore them.
## Package-level compatibility
PRP is one independently versioned component of the runner bundle. Catalog,
runner-client, control-plane-adapter, testkit, and eval-corpus compatibility is
declared by `PAPERCLIP_RUNNER_COMPATIBILITY` and checked before execution by
`assertPaperclipRunnerCompatibility`. A mismatch fails with
`paperclip_runner_incompatible` and stable per-issue codes; a provider-specific
tool error is not a compatibility negotiation mechanism.
See [ADR 0001](adr/0001-runner-testing-eval-package-boundaries.md) for the
component rules and clean-consumer packaging gate.
## Evals integration negotiation
The packed `./evals` entry point adds a stricter execution preflight for the
App/Evals join. `assertPaperclipRunnerEvalCompatibility` requires simultaneous
agreement on package semver, runnerd build metadata, a common PRP version,
semantic catalog version and SHA-256 digest, harness-driver contract and
required capabilities, and the native-execution version. It reports
`paperclip_runner_eval_incompatible` with expected/received values for every
mismatch and must run before launching a provider.
runnerd itself reports `paperclip-runner/runnerd-build-metadata/v1` from
`--build-metadata`. The consumer passes its path and expected content digest to
`resolvePaperclipRunnerdArtifact`; implicit PATH or source-tree discovery is
not part of the contract. Native attempt output is
`paperclip-runner/native-execution/v1`, whose parser accepts unknown additive
fields but rejects unknown required versions and inconsistent terminal,
semantic-denial, usage, or transcript facts. Full fields and the deterministic
gate are exercised by the package-local deterministic conformance suite.