<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task descriptions, comments, continuation data, skills, and execution rules enter several agent adapters. > - The same source can be rendered by more than one automatic input carrier. > - Failed resumes can also rebuild input from stale or compact context. > - This pull request gives each Paperclip-owned source one delivery owner and preserves the required transport boundaries. > - It adds deterministic adapter, interaction, runner, and browser tests for these boundaries. > - The benefit is more predictable context delivery with explicit evidence for later live qualification. ## Linked Issues or Issue Description Related: #13144 removes a duplicate environment payload and bounds wake lists. Related: #11360 addresses Hermes resume behavior. This pull request preserves compatible active-session formats while repairing context ownership and stale question creation. **What happened?** Task descriptions and comments could enter more than one automatic context block. Native transports could wrap a complete model input in a second task envelope. Some legacy and gateway adapters could omit the owned assignment on ordinary tasks or rebuild a failed resume with stale compact context. A continuation could also request a question after newer human comments had arrived. **Expected behavior** Each task or comment source has one automatic model-facing owner. Distinct comment IDs and repeated wording remain distinct. Fresh fallback attempts rebuild the required full context. A question request is rejected when newer queued human direction makes it stale. Harness access policy remains owned by execution configuration. **Steps to reproduce** 1. Build a task with a description and current comments. 2. Capture the actual adapter or runner input. 3. Compare source ownership and task-envelope nesting. 4. Queue a human comment before a continuation requests a question. 5. Trigger a failed resume and inspect the fresh retry input. 6. Run the focused adapter, interaction, runner, and browser checks. ## What Changed - Add shared prompt-section selection at the provider-attempt boundary. - Deliver owned assignment context through native, legacy CLI, ACP, gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and Hermes paths. - Rebuild full or compact context after resume recovery changes the attempt. Add native and Claude ACP tests of actual recovery requests. - Preserve custom templates, loaded instruction files, execution policies, and older active-session formats. - Record continuation source metadata and reject stale question creation under the issue-row lock. - Add explicit Product E2E context-integrity profiles, prerequisite gates, credential-isolation checks, and report fixtures. - Bypass service-worker forwarding for same-origin Vite development modules. A real Chromium test fails with resource exhaustion before the repair and passes after it. Production asset caching keeps its existing policy. - Add browser diagnostics and service-worker module-loading regressions. - Add an explicit zero-retry eval option. The default retry behavior remains unchanged. Each campaign records its effective policy. - Remove the model-facing working-directory sentence from four prompt builders. Existing workspace, sandbox, permission, and custom-template configuration remains unchanged. - Align the everyday workflow assertion with the current 47-entry catalog. Compared with current upstream master, the branch carries the context-ownership implementation and its tests, the explicit context-integrity catalog and evidence harness, and the focused browser regression checks. ## Verification **Merge assessment:** focused regression evidence supports merge. This is not full completion of the original broad qualification matrix. The maintainer has authorized merge after fresh verification of the master integration. - Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`. All 14 conflicts are resolved. Cancellation checks, workspace finalization, native Grok support, and both sets of tests are retained. - Current-head Greptile: **5/5**, with no blocking findings. The review names this exact commit. All **59 reported checks are terminal: 55 successful, 4 skipped, zero pending or failing**. This includes the full root general and serialized suites, separate runner checks, typecheck, build, canary, browser E2E, Docker, and security checks. The successful legacy security status is included in that total. - After integration: workspace typecheck and full build passed. Separate runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust tests, and 39 preparation checks**. Other passing checks include 621 Product E2E harness units, 376 focused shared/adapter tests, 160 real-database/API tests, 86 Hermes tests, 18 browser-support checks, and Product E2E typechecking. The complete root suite passed in CI. The duplicate local monolithic root run was stopped after that CI result; it is not counted as a completed local pass. - New native recovery coverage retains full assignment, completion contract, and explicit skill selection after safe replacement, for old and prepared input formats. Full native session test file: **136/136 passed**. - New Claude ACP coverage captures actual fresh, resumed, and missing-session fallback requests. It verifies one assignment copy, comment order, identical text under distinct comment IDs, and full fallback context. Full file: **33/33 passed**. Both affected TypeScript checks passed. - Existing deterministic tests cover source revisions, approval and trust boundaries, completion validation, custom templates, compatible sessions, standalone driver wrapping, and maintained adapter transport requests. - Provider-free browser support: **17/17 passed** after the master merge. Service-worker unit tests: **33/33 passed**. The module-overload regression failed before the repair and passed after it in real Chromium. ### Fresh live comparisons The new batch ran exactly four Product E2E attempts. **All four passed on the first attempt; no retries.** Each has six terminal matchers plus the existing browser lifecycle and invariant checks. | Exact case ID | Control | Candidate | |---|---|---| | `core-compatibility.runner-codex.local.plan-revise-accept` | Passed | Passed | | `local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume` | Passed | Passed | The plan case checks a revised canonical plan and revision-bound approval before completion. The question case restarts the server before submitting the answer, then verifies the continuation completes. Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical frozen definitions and provider versions: Codex `0.156.0` with `gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with `claude-sonnet-5`. The September 24 head added master browser recovery and test-only changes. The September 28 head also integrates newer master changes, including cancellation, workspace finalization, and native Grok. These are frozen-source live results, not exact-head live runs. The candidate received one description copy where the control initially received three. The submitted initial plan envelopes were 7,969 versus 19,097 characters. Question envelopes were 7,592 versus 18,919. These are structural measurements, not whole-provider token or dollar savings. ### Earlier evidence and failed attempts - The preceding fresh batch has four effective passing pairs: OpenCode comment continuation and assigned skill, native Codex comment continuation, and native Claude comment continuation. It retains **11 attempts: eight passed and three failed**. - Original failures remain recorded: missing local PostgreSQL library links before task creation; host-sleep cleanup after task/page checks passed; and a Claude **control** session-open rejection before a model turn. Setup was repaired identically on both worktrees. The permitted unchanged infrastructure retries passed. The underlying Claude provider startup error was not retained and remains unknown. - Older R2 retains **17 passes and one failure** across 18 attempts, including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode blank-page failure led to the service-worker repair. R2 is historical evidence: master changed the native fixed prompt and removed duplicate wake environment data afterward. - The September 24 CI run initially failed one unrelated preview readiness test (`ECONNREFUSED` on its local fixture). Its test and production code match master. Isolated local verification passed **28 tests, 3 skipped**. One unchanged CI retry passed the full shard: **831 passed, 1 skipped**, including all **31 preview-exposure tests**. The aggregate CI gate passed afterward. The precise startup cause remains unknown; a port race is a hypothesis, not a proved cause. ### Limits The original wider profile/workflow matrix, repeated trials, and remote Daytona qualification are incomplete. These results support a focused merge recommendation, not statistical equivalence or universal harness qualification. Some usage receipts are missing in both variants, so no token or dollar savings are claimed. The $500 ceiling was preserved using conservative allowances; failed attempts and unknown charges remain in the ledger. Reproduce the focused additions with `pnpm exec vitest run packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/native-session-runtime.test.ts`. Full checks use `pnpm -r typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner checks. Paid evals require the frozen definitions, profiles, and credentials; do not use `--all` as a substitute for the selected cases. ## Risks - Context placement changes can affect model behavior. Deterministic checks cover the selected paths, but live qualification remains incomplete. - The stale-question guard can reject a request when queued human comments arrived during the run. This is intended. - New stored inputs and model envelopes retain compatibility readers for older active sessions. - Custom templates may intentionally repeat content. - Removing a model-facing working-directory sentence does not change filesystem, command, sandbox, or permission configuration. - The worker bypass applies only to same-origin development module paths. Cache-policy tests preserve private-response handling and production asset caching. Mounted HTTP fixture changes remain test-only. - This PR does not claim measured token savings or statistical equivalence across every harness. ## Model Used OpenAI Codex, exact model gpt-6-astra, with repository tools and code execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The serving context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR using the required issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused local checks and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect these changes - [x] I have considered and documented risks above - [x] All current-head Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups for the current head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
15 KiB
PRP Compatibility and Versioning Policy
Authority
The JSON Schema files in protocol/schemas/ are the
language-neutral source of truth for the Replay and Local runner executable contract. The
generated TypeScript schema module is checked against those files before every
TypeScript typecheck. Rust consumes the same fixtures and must produce the same
golden parity summaries.
The reviewable architecture and trust boundaries are defined in
architecture.md; durable transport and recovery behavior
is defined in durable-recovery.md. Local runner reuses
the protocol contract for local live events. It adds package-local stdio and
stream envelopes, but it does not add durable transport, persistence, or
production control-plane behavior.
Version fields
| Field | Replay support | Compatibility rule |
|---|---|---|
protocolVersion |
1 |
Required. Negotiate the highest overlapping version; no overlap fails closed. |
fixtureVersion |
1 |
Required by the conformance corpus. Unknown values fail closed. |
event.schemaVersion |
1 |
Required on every event. Unknown values fail closed before reduction. |
capabilities.semanticTools.schemaVersion |
1 |
Optional advertisement. When present, an unknown required version fails closed. |
payload.semantic_tool.schemaVersion |
1 |
Optional on paired semantic tool input/result events. When present, an unknown required version fails closed. |
terminal.stopReason.schemaVersion |
1 |
Optional budget/cost receipt. When present, an unknown required version fails closed. |
Typed schema discriminators |
*.v1 |
Required. Unknown required schema identities fail JSON Schema validation. |
Wire protocol versions and fixture-corpus versions are independent. A fixture format can evolve without changing PRP, and a future PRP version can be represented only after the consumer advertises support for it.
Forward compatibility
- Unknown object properties are accepted and preserved by validation. Reducers ignore fields they do not understand until a later schema version gives those fields defined behavior.
- Unknown required versions, schema discriminators, enum values, and required fields fail closed. A consumer must never guess at their semantics.
- Scripted fixtures bind every event to the fixture run/session, require contiguous controller command order, exactly one unique proposed result, and exactly one unique terminal event.
- The top-level fixture result must equal the
run.result.proposedpayload after canonical key ordering. RepeatedsourceEventIddeliveries must be byte-equivalent after the same normalization.
The forward-compatibility fixture proves that optional fields survive validation without changing the v1 snapshot. The unsupported-version fixture proves that a required v2 protocol cannot be replayed by this consumer.
Within-turn checklist snapshots
plan.updated / paperclip.plan.updated.v1 is a complete, ordered snapshot of
the provider's checklist for one active turn. It is not a Paperclip Plan
document and must never be inferred from assistant prose, Codex proposed-plan
items, or generic TodoWrite output. Every replacement uses the provider turn ID
as planId; PRP sourceSeq, not an optional provider revision, determines
snapshot order. An empty step array clears the checklist, and complete is true
only when a non-empty snapshot contains only completed steps. The legacy
document-coupling fields are always syncStatus: "not_applicable" and
documentRevision: null.
| Qualified adapter profile | Checklist support |
|---|---|
| Direct Codex App Server | turn/plan/updated |
| ACPX Codex | Structured ACP plan entries |
| ACPX Claude | Structured ACP plan entries |
| ACPX Pi | Unavailable; no production-qualified profile is exposed |
| OpenCode | Unsupported until it exposes a structured plan event |
Codex turn/diff/updated is normalized separately as the latest same-turn
workspace.change.updated snapshot. Together these two independent event
families can drive a turn-status UI without changing the public PRP family or
adding a control-plane endpoint.
Provider-neutral semantic receipts
capabilities.semanticToolsadvertises stable operation IDs, availability, required claims, and redaction disposition without naming a provider API.mcp_app.tool_inputandmcp_app.tool_resultmay carry pairedsemantic_toolenvelopes. Correlation IDs must match the containing event; operation ID and idempotency key must match across the pair.- Content is represented by a canonical SHA-256 digest plus allowlisted typed references. Raw credentials, provider payloads, and hidden identifiers do not belong on the wire.
- Result receipts distinguish success, denial, conflict, exact duplicate, unavailable, and failure. They can name the authorization boundary, safe revision, artifact/work-product refs, immutable governed targets, and bounded wake/monitor causality.
terminal.stopReasonrecords budget/cost kind, stable code, retryability, limit class, safe aggregate, and decision receipt.
These fields are trace evidence only. The v1 reducer ignores semantic_tool
payloads, so adding or extending the optional envelope has no projection
effect. Eval trace_completeness treats PRP wire receipts as authoritative when
present and retains the pre-existing scalar fallback for live evidence that has
not yet emitted them.
Provider-neutral structured input
Harness-initiated forms cross PRP as paperclip.runtime_request.v2 with
requestKind: "runtime", type: "input", and an embedded
paperclip.question_set.v1. A submission is always
{ "action": "submit", "response": paperclip.question_response.v1 }.
Codex answer objects, OpenCode answer arrays, and ACP typed content exist only
inside their adapters; origin may retain the provider method and adapter name
for diagnostics but never provider response data.
Question and option order is significant, while answers are keyed by stable
question IDs and selections reference stable option IDs. The canonical modes
are text, single_select, and multi_select. Text validation is repeated at
the untrusted server edge and again against the persisted question set before a
provider receives the translated response.
Every harness adapter must add a
paperclip.question_adapter_fixture.v1 fixture proving its native request
normalizes to the canonical shape and its canonical response can be translated
back. The shared fixture format deliberately contains both native and canonical
objects so adding a provider does not change PRP or the UI contract.
V1 runtime requests and their legacy resolutions remain accepted during the
migration. A request that contains no structured form stays on the legacy path;
once a provider supplies a form, malformed or unsupported fields fail closed
instead of silently degrading. ACPX sidecars advertise only form elicitation
and use sidecar protocol v2 runtime.input_requested / input.resolve frames.
The live lifecycle pauses and resumes the same provider turn. If the provider
process is lost first, Paperclip emits one non-replayable
runtime_request.expired fact and materializes an idempotent durable
ask_user_questions interaction using the identical question set. Explicit
cancellation and already-resolved requests never create that fallback.
Native execution permission compatibility
paperclip.native-execution-input.v4 pins the effective harness permission
policy in the closed provider configuration: approvalPolicy for Codex and
permissionMode for OpenCode and ACPX. The pinned value participates in
provider-session identity, so an incompatible idle or recovered session is
replaced on the next execution. An active turn is never mutated in place.
paperclip.native-execution-input.v5 is the current input format. It retains
the v4 permission pinning and adds optional completion-source references;
paperclip.native-model-envelope.v3 is the corresponding explicit model
projection. Readers continue to accept persisted v1-v4 inputs, and v5 readers
must preserve the historical behavior of inputs that do not carry the new
optional fields. Missing Codex and OpenCode policy fields retain their
historical effective behavior. Legacy ACPX permissionPolicy: "interactive"
is interpreted as approve-reads, while new v4 and v5 ACPX executions default
to approve-all at the server boundary.
When a healthy provider session is resumed, its persisted v4 or v5 input format is retained even if the newly built input uses the other format. This avoids rotating an active session for a presentation-only schema change. A safe rollback from v5 to v4 removes only the optional completion-source references; it retains the task, contract, provider, workspace, and permission fields. Format changes still go through the normal provider-session identity checks, and an active turn is never mutated in place.
See Adding a harness for the permission catalog, isolation rules, and provider conformance requirements.
Seven conformance fixtures cover artifact success, redacted denial without fallback, stale conflict plus duplicate retry, governed target and continuation causality, budget/cost stop, unknown optional fields, and rejection of an unknown required version. The six accepted fixtures have shared TypeScript and Rust golden parity summaries.
Replay semantics
- Events are applied in fixture order and ordered independently by
(sourceKind, sourceInstanceId, sourceSeq). - A repeated source event ID has no second projection effect.
- A forward source-sequence gap is recorded explicitly; the reducer never invents a missing event.
- An event at or behind the committed source cursor is ignored and recorded as out of order.
- Replaying an already-applied batch leaves the snapshot unchanged.
The CLI and browser import the same replayReplayFixtureText function, so
validation, compatibility errors, and final snapshots cannot drift between the
two surfaces.
Local envelope rules
- Mock-core commands use
paperclip.prp.command.v1over stdin JSONL. - Runner output uses
paperclip.runner.stream.v1over stdout JSONL. - Fake-harness commands use
paperclip.fake_harness.command.v1. - Fake-harness output uses
paperclip.fake_harness.message.v1. - An equivalent repeated
commandIdreturns a duplicate receipt and has no second driver effect. Reuse with different data is rejected. - A new command must use the next contiguous
controllerSeq. - Harness logs are bounded diagnostic data. They are not canonical PRP events.
run.result.proposed,harness.exited, andrun.terminalare separate facts and appear in that order when a semantic result exists.- The live browser rejects an event with an invalid schema, run ID, or session ID before it reaches the reducer.
These envelopes are local Local runner implementation contracts.
Durable wire rules
- The runner opens loopback
ws://or hostname-verifiedwss://, or accepts a preview-proxy connection on its fixed listener, and completes the PRP v1 authenticated handshake before any command result or event. - A one-use bootstrap bearer capability returns a short-lived connection lease
in
welcome. Later connections use that lease. Neither raw capability is durable state. welcome.payload.connectionLeaseRenewalVersion: 1opts into authenticatedlease_renew/lease_renewedcontrol frames. Renewal extends the persisted expiry on the same live authority without restarting provider work. Identity, protocol, and revocation epoch remain fixed; expired or revoked leases cannot renew. See durable recovery for retry and warm-handoff rules. Peers lacking this capability retain their original lease expiry.hello.resumereports the last processed controller sequence, next source sequence, cumulative ACK cursor, and current unacknowledged range.welcomeselects the one overlapping protocol version, returns the core's cumulative ACK cursor, and carries at most one durable pending command.- An event is durable before send. Event IDs and source sequences stay stable across replay and process restart.
- An ACK is cumulative. The runner rejects a cursor behind its durable ACK or beyond its produced source cursor.
- An equal repeated command ID and canonical digest returns its stored result. Reuse with different bytes fails closed and cannot repeat an effect.
- Frames are bounded at 1 MiB and upgrade headers at 16 KiB. Unknown or invalid required protocol data fails closed; malformed JSON is a bounded diagnostic.
Runnerd build-metadata contract v2 advertises the exact transport inventory:
dial_ws_loopback, dial_wss, and listen_ws. Plaintext dial destinations
must resolve entirely to loopback. Public dial targets require TLS trust and
hostname validation; a private CA bundle augments the platform roots and must
be a bounded, private, regular file. Listener mode is fixed to port 43127 and a
single run-bound path. All modes retain the same message/frame bounds and PRP
authentication.
These are package-local Durable recovery and transport rules. Control-plane admission and deployment policy remain separately reviewed work.
Change policy
- Change JSON Schema first.
- Regenerate the TypeScript schema module.
- Add or revise a shared fixture and its golden snapshot/summary.
- Prove TypeScript and Rust parity.
- Update this policy and the normative spike specification when behavior changes.
Breaking changes require a new required version. Additive optional fields may remain in v1 only when old consumers can safely ignore them.
Package-level compatibility
PRP is one independently versioned component of the runner bundle. Catalog,
runner-client, control-plane-adapter, testkit, and eval-corpus compatibility is
declared by PAPERCLIP_RUNNER_COMPATIBILITY and checked before execution by
assertPaperclipRunnerCompatibility. A mismatch fails with
paperclip_runner_incompatible and stable per-issue codes; a provider-specific
tool error is not a compatibility negotiation mechanism.
See ADR 0001 for the component rules and clean-consumer packaging gate.
Evals integration negotiation
The packed ./evals entry point adds a stricter execution preflight for the
App/Evals join. assertPaperclipRunnerEvalCompatibility requires simultaneous
agreement on package semver, runnerd build metadata, a common PRP version,
semantic catalog version and SHA-256 digest, harness-driver contract and
required capabilities, and the native-execution version. It reports
paperclip_runner_eval_incompatible with expected/received values for every
mismatch and must run before launching a provider.
runnerd itself reports paperclip-runner/runnerd-build-metadata/v1 from
--build-metadata. The consumer passes its path and expected content digest to
resolvePaperclipRunnerdArtifact; implicit PATH or source-tree discovery is
not part of the contract. Native attempt output is
paperclip-runner/native-execution/v1, whose parser accepts unknown additive
fields but rejects unknown required versions and inconsistent terminal,
semantic-denial, usage, or transcript facts. Full fields and the deterministic
gate are exercised by the package-local deterministic conformance suite.