## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
73 KiB
Cursor ACP capability inventory
Current source (2026-10-06): Cursor profile v14 combines the model-selection correction with mainline ACP resource-not-found handling. The bundled ACPX runtime forwards unlisted model IDs unchanged, verifies model selection responses, and reapplies a selected model when a loaded session reports another model. Its 14 package tests exercise startup and reconnect through a real ACPX manager with offline ACP fixtures. Cursor activity projection now belongs to its adapter. The release attestation binds the updated ACPX patch; older profile identities cannot be reused for new launches. Committed plan waits remain readable. The live evidence below belongs to profile v11; it does not certify v14 artifacts.
Historical release candidate (2026-10-04): Cursor profile v11 passed local and Daytona qualification.
It ports snapshot 22c78242a4e0c2369fecf0c2dc4e7600fbad6706 onto mainline
dd868ed125cd709506dd9b29fca640a44d580501, preserving newer recovery and owned
warm agent-file handoff. The readiness checklist
records all ten local and all ten remote gates, including cleanup. The final
public installed-package smoke is recorded separately there. The historical
results below retain their own identities and limits.
The public installation path is paperclipai runtime setup cursor. It explicitly
downloads the pinned CLI 2026.09.26-dd393fe, applies the source-owned patch, and
verifies the complete runtime closure. npm installation does not download Cursor.
Run setup as the OS user that runs Paperclip. Local assets live in that account's
~/.paperclip/runtimes/cursor/<platform>-<arch>/<closure-sha256>, independently of
npm directory ownership and isolated provider HOME/XDG settings. Existing image
assets remain authoritative; an invalid packaged runtime cannot fall back to the
user cache. Release assembly and runtime loading reject stale Cursor profile or
closure identities even when the supplied manifest has been rehashed.
Ordinary provider packs include those assets, and Daytona image preparation uses
the same distribution. Supported targets are macOS ARM64/x64 and Linux x64.
Configure a company secret and bind it in the agent environment as
CURSOR_API_KEY or CURSOR_AUTH_TOKEN. Select the model explicitly; native ACP
must acknowledge that exact model before execution. Missing assets, credentials,
model access, and entitlement failures are surfaced without selecting a substitute.
acpxSessionMode supports Agent (default), Plan, and Ask and is part of recovery
identity. Acceptance of a native plan succeeds the planning run while leaving the
task open awaiting explicit direction; acceptance alone never starts implementation.
Paperclip semantic questions are the supported question path. The native AskQuestion handler is defensive compatibility code and is not certified for this release. Per-run dollar usage is unavailable. Partial counters remain diagnostic observations with unknown semantics; they do not establish measured spend or an enforceable per-run dollar bound. Qualification uses the separately authorized account cap. Image-input delivery, detailed native diffs, and deeper child transcripts remain follow-ups.
Historical source candidate (2026-10-01): Cursor profile v10 is unqualified.
Native tool IDs containing C0, C1 or DEL control characters now use the same
bounded hash in permission details, passive evidence and the sidecar tool event.
Blank permission IDs are rejected because ACPX drops blank tool-event identity.
Other admitted IDs stay unchanged at this boundary; distinct IDs stay distinct.
Canonical tool execution then applies the existing Rust opaque-ID conversion.
For example, tool/1 remains the native permission/evidence key while its
execution ID is the deterministic opaque hash. Lifecycle readers explicitly
convert between these keys; they must not compare them directly. Both-order
bridge tests cover this distinction, including 161-character IDs. The original ACP request and option identity
remain intact for response delivery. The declaration binds the identity helper,
permission adapter and evidence projector. Retained v9 sessions are incompatible.
The native distribution and its usage limits are unchanged. Deterministic tests
cover both permission/tool arrival orders and an unresolved real ACPX callback;
fresh local and Daytona qualification is still required.
Historical source candidate (2026-09-30): Cursor profile v9 is unqualified.
Its declaration binds paperclip-cursor-usage-v4, all three newly verified native
closures and the shared ACPX patch that persists bounded diagnostic observations.
Every native counter receipt remains partial with unverified semantics; it cannot
satisfy token or dollar accounting. The exact model remains unpriced until a
verified rate is available. Prior v7 source, builds and paid observations below
are historical and do not qualify v9. Old v7/v8 sessions must be reopened.
The v9 collector marks missing, invalid or reused child run ordinals as
child_run_attribution_unverified. The offline proof executes the pinned
vendor child creation and reuse methods on all three platforms; child counter
aggregation semantics remain unverified.
Historical qualification checkpoint (2026-09-30): Cursor profile v9 remains unqualified. Draft PR #14724 at 3adee6f3fc652d5f5ee40b2061a16cea0c7341f5 passes Apex 5/5 and all 51 check runs plus one context; one unchanged-head failed-job retry is recorded. It fixes future eval-notice preservation but does not regrade the retained strict protocol-accounting failure or establish native USD. The verified fixed $25 account-cycle cap resets October 28 and applies to account on-demand fees, not per-cell costs. The separate 15-test helper proposal is unintegrated and grants no launch authority. Profile v9 binds usage-v4 and three native closures; counters remain partial with unverified semantics. Local and Daytona qualification remain pending.
A separate native-question attempt ended without a captured
cursor/ask_question callback, and the model reported that AskQuestion was
unavailable. Public upstream review
attributed the reported symptom to server-side ACP session identification,
without identifying a missing initialize capability or a verified fix release.
Native AskQuestion therefore remains unverified for the pinned model/mode; this
does not establish universal absence or support in every mode. Semantic question
continuation is separate and does not qualify the native callback. Evidence is
retained at runtime-9-8-10-preparation/cursor-question-public-review/root-findings.json
(SHA-256 2fc34c1942e85ab58577b7cfa8ccdeafe520cc7c539aa67a7eef5a474848ad73).
Historical v7 checkpoint (2026-09-30): Cursor profile v7 remains unqualified. The corrected local native-plan case on controller/Product source 3d21d375de2b6249d9f5322bf001ce0ce0e052aa and native runtime 5b8e4454ef0bf12d0bb068c2e41d8c9df9356a1c passed all 26 checks. Reject, feedback, revision-bound acceptance, browser reconnect, exact delivery and unchanged workspace bytes passed. The single planning run succeeded while the task remained In Progress with an explicit next-message wait; no implementation or task completion is claimed. Cleanup and full end integrity passed. Result SHA-256: 685609f01291db84f40d890ef562473b95065e86907bd964b656ad732f6dd97a. The earlier controller40064 attempt remains failed. Controller fix c58a8881f uses the production schema/policy/contract hash envelope; native runtime/profile bytes are unchanged. That exact Cursor PR head has green CI and Greptile 5/5. All three platform builds exist, but the remaining current-profile local/Daytona cases are still required. Native AskQuestion availability remains unverified. Cursor on-demand usage remains $0 under the approved fixed $25 account cap; native per-run USD is unknown. See the comparative capability report for exact source roles, evidence and remaining gates.
Evidence updated: 2026-09-29. Candidate: 2026.09.26-dd393fe. This profile is not live-qualified. One authorized task-context prompt on the initial free account returned an upgrade requirement, invoked no semantic tools, and supplied no usage receipt. Its cost is unknown; the account dashboard remained unchanged at coarse precision. The user then selected another account, whose exact-model local Runner Eval passed task-context and history tools with clean terminal settlement. Its dashboard attributed 56K rounded tokens to included usage and no incremental charge; ACP still supplied no token or dollar receipt. The local Product hello subsequently passed through the real browser, server, database, native runner, and authenticated completion tool. A second local Product case passed file creation, editing, command validation, an independent exact-byte matcher, and visible workspace artifact presentation. The merged-source local Product question and revision-bound Plan approval also passed browser response and semantic-tool continuation. A serialized follow-up passed pending semantic-question recovery across a server restart. A verified native ACP probe also denied a shell write before any observed side effect. Native Cursor questions/plans and their recovery, durable permission presentation, interruption, and Daytona execution remain required qualification gates.
The earlier 2026-09-29 source used profile v5, with native session-mode admission and the instruction correction described below. All prior paid results (including v4) and packaged builds remain historical and cannot qualify v5. The earlier mode-selection limitation below describes the pre-v5 implementation.
The reference is the runner's Codex app-server integration and its closed thread-item inventory in src/provider-events.ts. Cursor ACP and its private extension methods are the transport; the legacy Cursor adapter is unchanged.
Evidence and packaging
- Cursor's ACP contract describes stdio JSON-RPC, sessions, permissions and five extensions. The pinned binary reveals additional capabilities and one material documentation discrepancy below.
packages/paperclip-runner/cursor-distributions.jsonpins vendor archive and complete extracted execution-closure SHA256 digests for macOS ARM64/x64 and Linux x64. All three archives were downloaded and independently inventoried (446 files on macOS, 454 on Linux). A real macOS ARM64 installation passed the materializer's full closure check.- The archive is a complete runtime: bundled
node,index.js, dynamically loaded numbered JavaScript chunks, native addons, workers, and assets. Pinning only the launcher or one executable does not pin its execution dependencies. The materializer rejects links/special files and overwrites, and verifies every file against the pinned closure digest. test/fixtures/cursor-acp/credential-free-probe.jsonrecords an actual macOS ARM64 initialization and nonbillable method probes.initializereports ACP v1, session load/list, HTTP/SSE MCP, image prompts, and opt-in subagent events. Audio and embedded context are false.session/forkandsession/resumereturn method-not-found. Creating/listing a session requires authentication.- Static evidence comes from
5672.index.js(ACP implementation),8096.index.js(ACP SDK protocol schemas),190.index.js(hook paths) andindex.js(CLI/configuration), all bound by the archive and closure pins. SDK schema support alone is not counted as implemented harness capability.
Capability matrix against Codex app-server
This historical field inventory distinguishes implemented transport behavior from its original qualification gaps. The current release proof and exclusions are in the readiness report linked above; native AskQuestion remains uncertified.
| Capability | Cursor harness/ACP evidence | Paperclip treatment | Remaining gap or qualification |
|---|---|---|---|
| Streaming assistant/reasoning | agent_message_chunk, agent_thought_chunk in live and replay presenters |
Existing ACP streaming path | Authenticated event capture required |
| Tool lifecycle/input/output | tool_call and tool_call_update, including raw input/output, locations and content |
Canonical lifecycle, bounded output, input-changed indicator and one relative target; Cursor extensions add semantic activity | Raw input body, structured content/diffs/media and additional locations are not projected. See the exact field audit below |
| Permissions | Native session/request_permission offers allow_once, allow_always, and reject_once |
Shared durable approval mapping preserves supported decisions; indefinite grant is not exposed without proven scope | A live raw ACP write denial preserved request ID 0 and original reject-once, with 98 absent-marker samples through cleanup. This does not qualify durable approval presentation or reconnect. The original grader failure and separate offline correlation assessment are retained |
| Human questions | cursor/ask_question, native single/multiple selections |
normalizeCursorQuestionRequest: all questions required; opaque UI IDs round-trip native IDs; forged/duplicate/missing selections rejected |
Active-turn adapter maps canonical answers, decline to skipped, and cancellation; durable delivery still requires end-to-end evidence |
| Proposed plans and decisions | cursor/create_plan with full Markdown, overview, todos, named phases and project flag |
normalizeCursorPlanRequest: complete presentation; accept/reject/cancel; revision-derived question ID; rejection feedback |
Foundation question descriptions allow 100,000 chars; this boundary also enforces the 196 KiB UTF-8 durable payload cap. Oversize rejects; it is never truncated. No generated planUri, since no uploaded plan artifact is fabricated |
| Execution checklist | Standard ACP plan and cursor/update_todos (merge) |
Per-turn reducer merges by native ID and emits complete snapshots with increasing revisions | Native cancelled status is visible as “Cancelled” with canonical blocked status, since canonical plan schema has no cancelled enum |
| Child activity | Optional subagent_spawned, subagent_state_update; metadata has tool-call/agent/model identity |
Stateful child normalizer emits canonical delegation and retains role/task across state deltas; active-turn adapter validates parent, known child spawn, and immutable tool/agent origin | Requires client _meta.subagents: true. ACPX SDK 1.4.0 lacks these two update variants. A closed opt-in pre-parser intercept binds the parent and declared descendants; child text/reasoning/plan/tool summaries are attributed to their child, never flattened into the parent. Nested media and raw tool details remain partial |
| Completed subagent tasks | cursor/task: description, prompt, subtype, model, ID, duration |
Canonical delegation with role, model, task and duration in activity summary | Dedicated numeric duration field absent from canonical delegation schema |
| Subagent questions | Pinned ACP subagent handler explicitly rejects interactive questions | No unsupported affordance advertised | Confirmed harness omission; requires upstream support |
| Generated images | cursor/generate_image: description, path and reference-image paths |
Canonical generated/viewed artifact events; description notice; realpath/regular-file containment checks; registered:false |
Paths are references, not automatic uploads. Revalidate before artifact ingestion. No image bytes invented |
| Files and diffs | Tool content includes diff blocks with path, oldText, newText; native edit output and locations |
Tool lifecycle plus the runner’s separate workspace-change evidence | Provider-reported before/after blocks are preserved by ACPX but not consumed by the canonical mapper. Additional and absolute locations are dropped; structured provider-diff projection needs implementation and live proof |
| True active-turn steering | No session/steer; handlePrompt cancels its existing pending prompt before starting another |
Advertise steering unsupported | Concurrent prompt is interruption/replacement, not Codex-style in-place steering. Do not mislabel it |
| Queued follow-ups | No harness queue method discovered | Control-plane scheduling can send another prompt after settlement | No native queue-management events; do not send concurrently to emulate queueing |
| Cancellation | session/cancel calls the active prompt cancellation callback |
Existing interruption path | Process/child cleanup and exact terminal settlement require live qualification |
| Turn settlement | Prompt implementation drains background subagent completion and awaits child publishers | Existing prompt result terminal boundary; explicit owned process cleanup | Native write denial settled without side effects. Background shell/child error settlement remains unverified. Credential-free stdin EOF did not exit within five seconds; do not advertise clean EOF |
| Session continuity | session/new, session/load; load replays historical messages, reasoning, tools, images and children |
Existing session load/recovery boundary with identity fencing | Duplicate replay suppression and provider-death recovery need authenticated proof |
| Session discovery | session/list implemented, cwd filter absolute; pagination cursor rejected |
Discovered and reported | Currently unused: runner recovers only its exact recorded session; arbitrary history browsing needs company-scoped discovery API |
| Fork/resume methods | session/fork and session/resume absent in agent implementation and return method-not-found |
Unsupported | Session load is available; do not infer fork from SDK schema |
| Modes | agent, plan, ask; session/set_config_option returns exact mode; native mode updates are observable |
v5 admits an explicit selected mode, binds recovery identity, and gates every prompt against native acknowledgement | Default Agent; mode is independent of permission policy. Plan acceptance does not switch to Agent. Native callback availability and Plan-mode completion remain unqualified |
| Model selection | session/set_model, config option model, model-parameter options, config_option_update; cursor/list_available_models extension |
Exact requested model must be selected/verified by shared admission | Authenticated listing and exact Luna variant echo passed. Native parameterized picker metadata and separate parameter changes also passed; this richer configuration path is not exposed by the runner. See the authenticated discovery evidence below |
| Available commands | available_commands_update, including skills/slash commands |
ACPX persists command metadata; runner canonical mapper deliberately emits no display event | Command picker and bounded command-discovery surface are not implemented |
| Session metadata | session_info_update title after automatic naming |
Capability documented | Runner owns normalized session identity; provider title is currently not surfaced |
| Prompt media | Image=true, audio=false, embeddedContext=false in observed initialize | The shared runner turn contract accepts text only; it does not forward image attachment blocks to ACP | P1: implement bounded, validated attachment-to-ACP conversion with capability checks, then qualify the exact model. Audio/embedded context are explicitly unsupported by this advertisement |
| MCP and semantic tools | Session MCP injection supports stdio/HTTP/SSE; user/project config also loaded | Runner-owned authenticated MCP bridge only; the owned distribution patch never initializes ambient MCP | Exact tool discovery and company/token boundary remain live tests |
| Usage/cost | Vendor ACP emits only prompt stopReason; native TurnEndedUpdate separately exposes optional input/output/cache/reasoning counters |
The observation-only patch below preserves bounded native counters as partial diagnostic metadata; it emits no standard usage or dollar receipt | Native counter aggregation semantics and exact-model pricing remain unverified; accounting stays unavailable |
| Reviews, compaction, hooks, memory | Native mode changes, hooks present internally; no dedicated ACP counterparts to Codex context/hook/memory lifecycle found | No fabricated event families | Distinguish confirmed absence of ACP event mapping from unverified native internals |
| Goal controls / lineage | No ACP goal or fork lineage protocol found | Unsupported | No equivalent of Codex native goals/thread lineage |
Wire and isolation findings
The observation-only usage candidate adds _meta.paperclipCursorUsage to a
prompt response. Its paperclip.cursor.native-usage.v1 envelope contains a fresh
opaque prompt ID and separate native invocation/run observations. It allows at
most 64 observations, 64 invocation identities and 16 KiB of JSON. Only observed,
safe nonnegative numeric input/output/cache-read/cache-write/reasoning counters
are retained; absent or invalid fields are omitted. Every envelope remains
completeness: "partial" with native_counter_semantics_unverified, even when
all fields are present. There are no standard ACP usage fields, counter sums,
estimated dollars or billed dollars. A reused child ID has no native callback
generation, so its counters are omitted with child_run_attribution_unverified.
The invocationId is a generated per-prompt ordinal. nativeRun is likewise a
generated invocation ordinal for parents, and the vendor's child run ordinal
for children; neither is advertised as a native request ID. Finalization closes
all tracked invocations, marks any missing terminal observation and enforces the
15,744-byte emission limit after all reasons are present, trimming observations
explicitly. The declared 16 KiB limit includes a 640-byte reserve for the bounded
ACPX request/message identity wrapper persisted alongside the envelope.
History and callbacks arriving after their invocation/prompt closes cannot add
observations to another prompt. Cancelled prompts return their own partial
envelope after the existing cancellation handling; they do not claim complete
cancelled-work accounting. A thrown JSON-RPC error returns no prompt envelope,
so these observations are not retained on that error path. This is a new source candidate, not a change to
the frozen v7 runtime or a regrading of historical evidence.
The paperclip-cursor-usage-v4 candidate retains the vendor version and archive
pins. Private full-tree materialization verifies 446 files per macOS target and
454 files on Linux; only the ACP implementation chunk differs from each vendor
tree. Its patched closures are:
| Platform | Patched closure SHA-256 |
|---|---|
| macOS ARM64 | 257424bd48e35412091c6adfc61e4648e836757ec1d240d890bba81a24918c30 |
| macOS x64 | 6f28c799c5afdc64fbdff8a2157f565f17ae7615efe014ac389d63bf70cf2be2 |
| Linux x64 | eadb8bb8ffb0450455b15b88c9b230307a9e149958a0c452d16dd157b4633d74 |
usage-v4-runtime-patch-offline-proof.json retains fresh isolation/instruction
checks, and native-usage-v4-offline-proof.json records the usage checks. Earlier
proof files remain historical. These are offline source proofs; native emission
semantics, accounting completeness and live qualification remain unproven.
The pinned-source offline harness executes the patched prompt method, native invocation expression, child construction/reuse and child-update methods on all three platform chunks with transport/session dependency doubles. It proves routing and bounds, not actual backend emissions, retry/subagent inclusion, cache/reasoning overlap or billing. Shared metadata persistence and new distribution/profile identities are separate integration requirements. Until semantics and the exact selected model's pricing are verified, these observations grant no accounting coverage. The vendor package declares no license grant; its bundled license notices describe third-party components. Cursor's terms and the applicable agreement remain release-review evidence for modified distributions; this source work does not establish redistribution permission.
The documentation calls todo/task/image methods notifications. The pinned presenter actually calls connection.extMethod(...).catch(...): these messages can carry JSON-RPC IDs. They need immediate acknowledgement plus activity normalization. They must never become human input requests or block awaiting an unnecessary decision.
Question and plan requests lack a session ID. The shared hook binds them to the retained connection, active execution/session/turn and originating request ID; it rejects stale/duplicate answers, cancels on turn termination, persists requests before exposing them, and awaits exact JSON-RPC pipe-write delivery before recording a resolved interaction. This proves transport handoff, not an additional provider application-level acknowledgement. The provider module does not substitute a best-guess session.
Fixed launch arguments are --disable-project-configs --disable-auto-update acp, using the verified bundled Node and absolute verified index.js. Private HOME/XDG roots are required together with CURSOR_CONFIG_DIR, CURSOR_DATA_DIR, disabled compilation caching, AGENT_CLI_CREDENTIAL_STORE=memory, and NO_OPEN_BROWSER=1. Only explicitly bound CURSOR_API_KEY or CURSOR_AUTH_TOKEN may enter. The controller creates a provider/session-scoped credential-name binding from the explicit task environment, ignoring inherited or caller-supplied markers. Rust forwards it through the closed sidecar environment boundary. Before host admission, the sidecar rejects missing, stale, wrong-provider or unbound credentials and removes the marker before provider launch. Tests cover this full boundary and preserve legacy credential behavior. Do not use --force, --trust, or --approve-mcps to paper over governance.
--disable-project-configs only suppresses .cursor/cli.json. It does not suppress project MCP, Cursor/Claude hooks, or installed plugins. assertCursorWorkspacePolicy refuses ambient project execution config from the workspace through its nearest Git root, including symlinks and unreadable configuration. Ordinary Claude settings with neither hooks nor plugins remain admissible. Enterprise system hooks are checked too. The host repeats admission checks before native launch. The vendor admission check alone cannot prevent a concurrent file mutation. The current paperclip-cursor-usage-v4 distribution patch removes local hook loading, asynchronous team hook synchronization, ambient MCP loader initialization, and ambient-client merging from the ACP source. The retained admission check remains defense in depth. This closes those discovery paths independently of filesystem polling; authenticated validation on the patched final runtime remains required.
Artifact paths are never automatically read or uploaded. References require a present regular file under the physical workspace, reject traversal and symbolic links, and retain registered:false. Private/outside paths are omitted and produce a visible warning. This is metadata validation, not durable file ownership or a replacement for artifact registration checks.
Verification and open work
Passed in the provider worktree:
- 17 Cursor Vitest cases: questions, multi-selection, native ID round-trip, invalid answers, exact plan revisions, outcomes, todo delta merging, delegation metadata, path containment/symlink refusal, child identity retention, private launch defaults and project config gates; active-turn canonical response mapping and immediate acknowledgements; stale/foreign/unknown child identity rejection.
- 5 materializer Node tests: platform pins, entire-tree tampering, linked paths, archive tampering before extraction, and overwrite refusal.
- Targeted strict TypeScript compilation for both provider modules and their imported contracts.
- Credential-free wire initialization and auth/method probes, plus real macOS ARM64 materialization and archive/closure inventories for the other two platforms.
The generic native-distribution lease now snapshots the pinned closure, launches Cursor through its verified bundled Node with a closed-module guard, and uses the existing lifetime guardian. A real nonbillable Cursor initialization passed through that lease. Its six focused tests passed, including argument fencing, external module rejection, and native child ownership/exit proof. The existing snapshot suite also passed. The foundation’s later master integration refreshes the common Claude/Codex pins; the candidate pack below was built against that synchronized graph.
Priority follow-ups before advertising support: (1) per-attempt cost attribution despite missing native usage receipts; (2) durable restrictive-approval and interruption proof; (3) remaining local and Daytona files/artifacts, questions/plans, recovery and settlement proof; (4) parent tool content/diff projection and richer nested child presentation; (5) authenticated trace audit for additional unconsumed native fields. Shared receipt/recovery coverage now proves pipe delivery, provider-loss expiry, exact lifecycle IDs and no approval replay.
Remaining rich capabilities intentionally unused: session discovery, generic mode/config-option UI, provider session titles, command/skill picker metadata, full nested child transcripts beyond bounded delegation summaries, native numeric task duration, and optional planUri. They are listed above with reasons; neither “unsupported” nor “complete” should be inferred for unverified live behaviors. The final combined report must update this inventory with integration evidence and include any further events observed during authenticated qualification.
Shared display-channel verification additionally covers exact rich-event schema validation, terminal/semantic/source injection rejection, active-turn fences, complete Unicode plan preservation through durable outbox persistence, secret redaction disclosure, and oversized UTF-8 rejection. The 4 KiB diagnostic preview cap remains active for diagnostics, never full approval documents.
Child transcript boundary and remaining fields
Pinned source creates a separate session-update presenter for every child and can use another child as the immediate parent. Enabling _meta.subagents therefore requires handling both lifecycle and ordinary child session updates. The Cursor-specific ACPX compatibility path tracks declared descendants per stream and active prompt, forwards them only under the retained root parent authority, and suppresses unknown child sessions. It leaves the SDK's standard parent update schema closed. Known child updates never become parent assistant messages. Old-stream, foreign-parent, pre-turn, settled, aborted, oversized, unknown-child and duplicate-spawn cases are covered.
Child assistant/reasoning text, plan steps and tool lifecycle are displayed in each child's activity summary. The 4,000-character activity surface retains a clearly marked tail when it fills; this is a summary, not a complete nested transcript. Child tool raw input/output, locations, diffs, media content and other child update fields are not rendered by this summary surface and produce explicit partial-detail notices. Adding canonical child transcript/tool surfaces is the next priority for these meaningful exposed fields. Native capabilities:{} currently carries no populated child capability fields; _meta.cursor.agentId and toolCallId are retained internally for immutable origin checks. Nested parent identity is represented in activity metadata because canonical delegation has no parent-child relation field. Native disconnected now yields a failed child with a visible explanation, never a successful completion.
Three real ACP stream tests prove child lifecycle/transcript preservation before SDK parsing, nested parent attribution, standard parent parsing, and old-stream/turn/connection fences. Shared extension, reply-delivery and Pi receipt package cases pass against the resulting patched package; the current totals are recorded below. Three native factory cases verify all platform pins, exact profiles, and rejection of unsupported distributions; native assets resolve only from the runner-owned package authority.
Integrated pack and offline launch evidence
The full build-provider-pack.mjs --candidate-providers=cursor path passed on source 25fb1b5b317e52a8ad50208d7681a1ee34bd939c (foundation includes upstream master 18e8c121d). It used standalone Node, pnpm 9.15.4, a fresh production deployment and the pinned vendor archive. The build checked portable Node relocation, fresh ACPX import and complete Cursor materialization. The retained manifest is packages/paperclip-runner/test/fixtures/cursor-acp/provider-pack-darwin-arm64.json; it contains only relative paths and digests.
- Pack digest:
sha256:38bc6dd39c13b0f5478a1018e26deb61b0333ab0242f3d857e3c79d3a34b7682. - Cursor declaration:
sha256:c91aa592ec867071ec4457b9b7230399ce42ad918787ceed99b4b651ea5607e2. - macOS ARM64 closure:
sha256:77394184a89b0e7384971181da19c3c82d83a399b04e7944b3f94fe3d7e62b25. - Resolved dependency lock:
sha256:9eea60187c6c808efb4d936a234a8d8000beac45253091abd4c5348132bd1408, also the reviewed Daytona build digest for this branch. This digest was resolved from the clean tracked lock with the exact Docker resolution-only command. A second resolution from the clean tracked lock produced the identical digest; the frozen install applied the Cursor/Grok combined patch successfully. The image consumes the reviewed resolved-lock artifact without registry resolution. Image creation rejects a different digest before execution.
provider-pack-offline-proof.json records a real launch through the generic profile installation registry from the pack’s verified native closure and retained snapshot, using private directories and no credentials. ACP v1 initialize succeeds. The correct cursor/list_available_models, session/list and session/new methods return authentication-required; session/fork and session/resume return method-not-found. The process produced no stderr, was explicitly terminated after these probes, and its lease closed. The earlier discovery fixture’s _cursor/list_available_models probe used an incorrect method prefix; this later probe corrects that uncertainty without changing the original evidence.
A separate credential-free macOS probe on the same final pack initialized with JSON-RPC request ID 0, then closed stdin at 20:56:57.776 UTC. The native process did not exit within five seconds and produced no stderr. The probe then sent the owned process group SIGTERM and closed its native lease. provider-pack-offline-proof.json retains eofSettled:false. This is a native lifecycle limitation: no clean-EOF guarantee is claimed. Production cleanup uses explicit process ownership and termination; Linux image qualification must separately record EOF behavior and verify bounded explicit cleanup. This probe sent no credentials or model prompts.
At the prior foundation checkpoint 7721662f2, the TypeScript/verified-sidecar build and all 177 targeted cases passed: Cursor normalization/policy/artifacts, credential bindings, the durable launcher's actual process specification, candidate backend admission, and sidecar behavior. The actual process-launch regressions verify explicit candidate credentials and their marker survive the final closed allowlist; ambient credentials and caller-supplied markers cannot select themselves. The prior source f09a9a3f8 passed all 47 CI jobs after one retry for runner shutdown and a wall-clock rate-limit case, plus Greptile 5/5; it also passed 29 installed ACPX/materializer Node cases and 6 Daytona image-content cases. These prior results remain historical, distinct from the final source proof.
The current pack/probe binds foundation f063fbf2b, including candidate admission, biller attribution, explicit unavailable usage, actual-spawn credential binding, and preservation of usage receipts on terminal provider failure. The verified TypeScript build, 56 Cursor/sidecar/eval-provider tests, 67 shared live/eval tests, and 25 mixed-usage contract tests passed during these qualification fixes. Earlier pack evidence remains in Git history. Its explicit offline factory model sentinel is not an authenticated model selection; no ACP prompt was sent. Daytona content identity includes the candidate selection, Cursor materializer and distribution manifest. This is local macOS ARM64 packaging/initialization evidence only. It does not qualify Linux execution, a Daytona image build/run, authenticated model work, permissions, or dollar accounting. No screenshot is claimed for these headless tests. That offline probe submitted no prompts and incurred no model usage. Subsequent authenticated attempts are recorded separately below.
The final pack passed 31 installed-package/materializer contract tests, 82 Cursor/integrity profile tests (six platform-specific skips), and 12 candidate/image-content tests. TypeScript and verified-sidecar compilation passed. Its generic installation registry probe initialized the pinned native executable without credentials or prompts, then shut down cleanly. A fresh release daemon was built from source 78360fa55adff11da9f200b228467eca782b81e1, Rust tree 17d327094be5759d80ef5bd02e15c59e78548ea2 (identical to foundation f063fbf2b), SHA256 e6a9fb5170b76a49b8411834b3706e8edf8f1a1ae85ad13b55368158aa7f67a0. This source includes bounded snapshot copying and public Grok packaging. Earlier paid attempts retain their original sources, package identities, and daemon digests; refreshing packaging does not rerun or extend their qualification.
The locally built Linux x64 image also passed credential-free native initialization through the generic installation registry and immutable lease, with Docker networking disabled. Image ID is sha256:5fa7951d1d6dd99305555fe00a5baf2bf5b834d16f021053737301975b93986f, built from provider source 25fb1b5b317e52a8ad50208d7681a1ee34bd939c. The Linux pack digest is sha256:5fab4bc1f0633561b7d7e858a535ebd23faeccedb5f86044df071a3206d949ba; the runner daemon SHA256 is 5eb3031aed118112caf0b80ea67a6cd01283098e2b4e526ef02d1416d125a1fb. The pinned Cursor closure is sha256:bf04ef8ea6a63191067a6b3e15139aa2c793d8419daab246aa3f0a012e1bbcd9. ACP initialize request ID 0 returned protocol version 1. Linux also reported eofSettled:false; bounded explicit process-group termination and lease cleanup passed. linux-image-offline-proof.json retains the identities and limits. This local container proof sent no credentials or model prompts and is not a paid Daytona run or authenticated Linux qualification.
Authenticated discovery evidence
test/fixtures/cursor-acp/authenticated-capability-probe.json records the user-authorized follow-up on 2026-09-28. The pinned verified distribution returned 41 model entries, created a session, accepted gpt-5.6-luna[context=272k,reasoning=medium,fast=false] through both session/set_model and session/set_config_option, and returned that exact value. No model prompt was sent. These facts establish authentication and model selection, not inference entitlement or full runner qualification.
The native picker has two materially different interfaces. Ordinary ACP advertises fixed variant IDs. It rejected grok-4.7[context=256k,reasoning_effort=low,fast=false], although each individual parameter value is advertised by the separate catalog. Advertising client _meta.parameterizedModelPicker:true enables base model grok-4.7 and separate session/set_config_option calls; context 256k, reasoning_effort:low, and fast:false all succeeded and were echoed. The runner does not yet negotiate or project this parameterized configuration path. Follow-up P2: bind the model plus complete verified parameter map to profile/recovery identity and expose supported options without silently choosing a default. In particular, this account's Grok and Composer catalog defaults enable the more expensive Fast option.
Credential setup uses Cursor's documented browser login, with an isolated task-owned home and the pinned file credential store. Only the access token was explicitly saved in the operator's private secret file; credentials and refresh tokens are absent from retained public fixtures. The official usage guide distinguishes included usage from on-demand charges and warns that spend-limit enforcement is delayed. The first prompt therefore requires a separately reserved budget and attribution from the account usage records; missing ACP usage remains unknown, never zero-cost inference.
Authenticated qualification attempts
test/fixtures/cursor-acp/authenticated-qualification-attempts.json retains sanitized per-attempt source, package, daemon, model, timing, cleanup, and cost evidence. The first canonical get-task-context attempt used pack source e5159fb67ad28bf6a691bfa64ffddb332d96f8be, exact Luna selection, one turn, a 60-second limit, a $0.50 declared envelope inside a $2 reservation, and zero retries. This is a Runner Eval against the mock control plane, not Product E2E or restrictive-permission qualification.
The provider returned only an upgrade requirement and then normal completion. Source inspection of pinned 5672.index.js confirms that actionable authentication/entitlement errors are rendered as ordinary assistant text and lose their typed failure before ACP settlement. No text-matching terminal heuristic is used. This was a native-wrapper gap in that unmodified candidate. The owned distribution patch described below preserves those typed errors as failed ACP RPC responses; the historical attempt is unchanged. The semantic oracle must still fail a response that did not invoke the required tool.
The replacement-account attempt passed all four canonical get-task-context checks at source 01959b8a602683f13706807983f02c3cba9d36a0, using newly rebuilt daemon source f77ae83aeb6d802c73f1c955760017ae272eba29. Both get_task_context and get_task_history returned successful authenticated semantic results. It took 29.291 seconds and exited cleanly. The Usage page showed a new gpt-5.6-luna-medium row at 19:28:55 UTC with 56K tokens, marked Included. Incremental cash was zero for that row; metered dollars remain unavailable. A conservative $0.07 list-price bound assumes at most 57K rounded tokens at the maximum listed $1.20/M rate and is not an invoice amount. There was no retry.
The local Product hello first stopped at agent creation with HTTP 422 before any provider prompt. Foundation 26dde3cfe fixed exact host-owned candidate/model admission at the public agent API. The separately authorized replacement passed at source edf538e61e712dddb6b4d59045c3dcfd445686c7, from 19:41:56.376 to 19:42:38.761 UTC. All six independent matchers passed: one exact completion marker, task Done, run succeeded, native mode, and local execution. The retained run confirms the 120-second configured deadline, no timeout, committed finalization, and clean cleanup. The final screenshot was inspected and shows the Cursor assignee, tool activity, one completion marker, and Done status. Native usage is null. The account Usage page attributes a new 19:42:25 UTC Luna row with 41.7K tokens, Included in Pro Plus, and zero incremental cash. Metered dollars remain unavailable; a conservative $0.06 list-price bound uses at most 42K rounded tokens at $1.20/M and is not a billing receipt. This proves the local completion-tool product path, not native question/plan, permission, file, recovery, or Daytona qualification. Both attempts remain in the sanitized evidence fixture.
The local file-edit Product case passed at source fe132224c2b30a8d9ce7b46cea38b8760af233fc, from 19:44:43.582 to 19:45:24.406 UTC. All seven matchers passed, including an independent exact-byte read of the 24-byte final file. A native shell command created then rewrote the file and asserted the expected bytes; tool lifecycle and output were retained. The inspected browser screenshot shows the file, a downloadable workspace artifact card, edited-file activity, the completion marker, and Done. Cleanup passed. This is workspace artifact discovery and presentation evidence, not provider-reported diff projection. The account Usage page attributes a 19:44:59 UTC Luna row with 84.1K tokens, Included in Pro Plus, and zero incremental cash. Metered dollars remain unknown; the conservative list-price bound is $0.102, not an invoice. The trace also confirms a meaningful remaining field gap: native command rawOutput.exitCode:0 survives in serialized output, while canonical exitCode remains null; extracting only validated command results is a P2 follow-up.
The merged-source local question Product case passed at a7e01a0cec397dd5048f5d5b5825658dc6e91450, from 20:04:44.147 to 20:05:52.691 UTC. Its two expected provider runs submitted at 20:05:00.738 and 20:05:30.525 UTC, with no retry. The real request_human_input semantic tool produced a pending single-choice question, the browser selected Cobalt through the public response endpoint, the answered interaction woke the assignee, and the task completed with one final marker. Both pending and final screenshots were inspected; all six terminal matchers, lifecycle invariants, and cleanup passed. Bounds were 120 seconds per turn, 300 seconds overall, and a $2 reservation. Native usage is unavailable. The dashboard attributes Included Luna rows at 20:05:01 UTC (35.5K tokens) and 20:05:35 UTC (42K tokens), with zero incremental cash; metered dollars remain unavailable. A conservative $0.10 list-price allowance is not an invoice. This is semantic interaction and continuation evidence, explicitly not qualification of native cursor/ask_question, server restart recovery, or plan decisions.
The local Product Plan case also passed on the same frozen source a7e01a0cec: 20:09:56.429–20:10:49.114 UTC, two expected turns, six passing terminal matchers and no invariant failures. The inspected UI displayed the complete two-step Plan at revision 1. The pending confirmation's target.revisionId exactly equaled the displayed document's latestRevisionId; it was accepted before continuation and successful completion. Both account Usage rows (48.5K and 43.7K) were Included with zero incremental cash; native metered dollars remain unavailable. This covers semantic write_document plus request_human_input confirmation, not native cursor/create_plan.
The first subsequent server-restart attempt failed; the later authorized retry is recorded below. Its first launcher attempt failed before any provider call because the task's temporary Unix socket path exceeded the macOS limit; the zero-prompt failure remains recorded. An explicitly authorized infrastructure replacement used a short canonical private path and produced the required pending question. During the requested restart, embedded PostgreSQL failed semget because the host semaphore limit was exhausted. No answer or continuation occurred. The original result remains failed (transient_infrastructure), including failed API cleanup verification; a separate process inspection found no surviving task-owned server, PostgreSQL, runner, or provider process. The dashboard attributes one 20:13:38 UTC Luna row with 23.7K tokens, Included in Pro Plus, zero incremental cash, and no continuation row; native metered dollars remain unavailable. Exact submission time was not recoverable from the post-failure API snapshot, so attribution uses the retained case window. No automatic retry was made.
The explicitly authorized serialized restart retry passed on source d35b83074a018537f5475568d7410e3b1d676789, using the final pack and matching release daemon, from 20:42:13.120 to 20:43:10.985 UTC. All six matchers passed, no invariant failed, and cleanup passed; a separate process inspection confirmed no remaining task-owned processes. The real pending semantic question survived the server restart, was answered Cobalt through the browser, and produced one successful continuation and Done. Both screenshots were inspected. This establishes semantic interaction persistence and wake-assignee recovery, not restoration of a blocked native cursor/ask_question request. The dashboard attributes two Luna rows at 20:42:26 UTC (23.8K tokens) and 20:42:57 UTC (41.7K), both Included with zero incremental cash; the displayed current-account total is 458.1K across eleven requests. Native metered dollars remain unknown. Original launcher and PostgreSQL failures remain retained.
A separate native AskQuestion sidecar probe on final pack source 25fb1b5b317e52a8ad50208d7681a1ee34bd939c sent one prompt at 20:39:47.990 UTC, with no Paperclip semantic tools. It requested single and multiple selection through Cursor's native question tool. The provider instead replied “The native AskQuestion tool is unavailable in this session,” emitted zero native input requests, and ended normally. The qualification oracle correctly failed; session close passed. The dashboard attributes a 20:39:48 UTC row with 17.4K tokens as Included, with zero incremental cash and unknown native metered dollars. This is observed default-mode tool unavailability, not proof of absence in every Cursor mode. The pinned source implements session/set_mode for agent/plan/ask and native question/plan handlers; the runner currently has no explicit mode-selection command. Retained native initialization exposes modes and model settings but no native tool inventory or question/plan availability bit. The bundled ask_question_all_modes:true default does not prove the server-selected tool catalog. No native Plan prompt was attempted, and no native interaction pass is claimed.
The restrictive native denial probe used the final verified pack and exact Luna model, with one prompt submitted at 20:52:07.462 UTC. Cursor emitted numeric request ID 0 for the shell write. Its offered IDs were allow-once, allow-always, and reject-once; the client selected the unchanged reject-once and acknowledged the pipe write at 20:52:10.919 UTC. The provider ended at 20:52:11.980 UTC. The forbidden marker was absent in all 98 independent samples, including before launch, at the terminal, five seconds later, and after process-group and lease cleanup at 20:52:16.996 UTC. This establishes the native harness denial behavior, not sidecar normalization, Rust durability, or Product approval UI.
The original probe grader failed because it expected rawInput on the permission frame. That frame omits the arguments, while the preceding same-session tool_call carries the exact command under the identical native toolCallId. The original failed result remains immutable. A separate offline assessment correlates that exact command, session, tool identity, actual response and delivery acknowledgment with the original marker samples; it passes without another model call. Tests reject missing binding, a different tool ID, a different command, wrong kind, or a title alone. native-denial-proof.json retains this distinction and sanitized wire frames. The account Usage page attributes 24.4K tokens at 20:52:07 UTC, Included in Pro Plus, bringing the displayed total to 482.5K across twelve requests with zero on-demand charge. Native metered usage remains unavailable. The conservative $0.03 list-price allowance is not an invoice.
The first attempt also exposed the shared live-eval layer's rejection of missing usage. The foundation now retains explicit unavailable-usage records for diagnostic candidates, leaves the numeric receipt ledger empty, preserves the response for grading, and keeps unknown aggregate costs unknown across mixed turns. This fix does not retroactively turn the retained failed attempt into a pass.
Cursor's Grok Bot billing documentation says linking SuperGrok grants Grok Bot usage while leaving the Cursor plan unchanged. Cursor's pricing page lists Composer for free Hobby and frontier models for paid plans. These documents and model-list presence do not establish entitlement for a particular CLI prompt. The operator selected a replacement Pro+ account; its exact Luna configuration was echoed successfully, with no automatic model substitution.
Exact remaining field audit
This source audit follows pinned native update → ACPX runtime → canonicalProviderEventsFromAcpxRuntimeEvent → Paperclip activity. It supplements the higher-level capability matrix and does not claim an authenticated trace was captured.
| Native field/interface | Preserved boundary | Current projection / reason for partial use | Follow-up |
|---|---|---|---|
| Prompt image blocks | Pinned native initialize advertises image input | Shared AcpxRuntimeTurnInput has only text; codex-runtime-adapter.ts passes that text to ACPX, so no image blocks reach the wrapper from a runner turn |
P1: add validated, bounded attachment inputs through the runner/ACPX contract, then prove exact-model behavior; native advertisement alone is not implementation evidence |
Parent tool rawInput |
ACPX forwards the body | Canonical tool schema exposes only inputUpdated; full native arguments are not displayed |
P1: bounded/redacted tool-argument presentation with company and path checks |
Parent tool content[] diff path, oldText, newText |
ACPX forwards structured content; native permission requests also use it for proposed changes | Canonical tool mapper uses rawOutput, so completed provider-reported diff blocks have no dedicated output surface; separate workspace evidence is not equivalent |
P1: validate physical workspace locations and expose attributed diff artifacts |
Parent tool content[] images/resource blocks |
ACPX forwards blocks and produces a summary | Structured media is unused by canonical tool output; safe registration/location semantics must precede ingestion | P1: validated artifact reference projection with explicit provenance |
Parent tool locations[] |
ACPX forwards bounded locations | Only first relative location becomes target; absolute paths and additional locations are discarded |
P1: workspace-aware multi-file attribution; never strip arbitrary absolute prefixes |
Parent tool rawOutput |
ACPX forwards raw value | Canonical output is bounded serialized text, not lossless structured data; truncation is marked by output metadata | P2: typed tool-result detail viewer when useful |
| Plan step priority/native extras | ACPX plan entries | Display uses body/status and fixed ordinal step IDs; priority and unknown metadata have no canonical fields | P2: add explicit bounded fields if native traces show product value |
| Commands, config options, session title | ACPX persists/normalizes available_commands_update, config_option_update, session_info_update |
Canonical display mapper explicitly skips these tags; exact model admission remains separate | P2: company-scoped command/config/title surfaces |
| Child tool raw arguments/results, locations, diffs and media | Cursor child intercept retains validated update for its adapter | Child activity keeps bounded attributed summaries, with visible partial-detail notices; there is no nested transcript/tool detail surface | P1: canonical child transcript and tool/artifact relation model |
Child _meta.cursor.agentId / tool origin / parent ID |
Retained for immutable-origin and descendant checks | Internal identity guards; parent relationship is textual metadata because canonical delegation lacks an edge field | P2: typed delegation lineage |
Task numeric durationMs, optional accepted planUri |
Native task/plan extensions | Duration is visible in summary; no numeric delegation field. No plan URI is returned because no plan artifact is registered | P2: typed duration; only return an actually registered plan URI |
| Session list, model parameters, generic mode control | Confirmed native ACP methods | No operator UI or cross-session discovery API in this integration; recovery loads only the runner-owned exact session | P2: explicit company-scoped controls; authenticated model catalog first |
True steering, native queueing, fork/resume, native usage receipts and interactive child questions are absent from this pinned ACP implementation as described above. Their absence is distinct from the implemented-but-unused fields in this table. Remote team hooks, provider-native artifact traces, interruption settlement and authenticated Linux/Daytona behavior remain unverified, not confirmed absent.
Initial isolation patch, profile v3 (2026-09-29)
The initial profile v3 materialized paperclip-cursor-isolation-v1 on top of the
unchanged vendor 2026.09.26-dd393fe archives. The materializer first verifies
the archive and vendor execution closure, requires exact single-occurrence
source anchors in an input-digest-pinned ACP chunk, applies the owned patch,
then verifies a separately pinned patched execution closure. Both identities
remain in the distribution manifest. Unknown, changed, already-patched, or
ambiguously matched source fails closed. The patch does not change the legacy
Cursor adapter or the vendor interactive CLI.
The ACP shared-services initializer no longer initializes or loads its ambient MCP loader. Each ACP session starts with an empty client lease and adds only the explicit session MCP definitions, preserving native last-definition-wins behavior. It never borrows an earlier or ambient lease, including when the session MCP list is empty or invalid. Local hook config is an empty snapshot, and the remote team hook fetch/install/update path cannot start. Newly created project config during a session therefore cannot enter either removed discovery path. This is source isolation, not an operating-system sandbox: explicitly allowed native shell commands still have their granted filesystem/network powers.
Typed ActionRequiredError instances now produce a failed JSON-RPC response
with data schema paperclip.cursor.provider-error.v1, kind action_required,
and the closed action enum login, upgrade, payment, config, or unknown.
Login and native ConnectError with the Unauthenticated code use ACP error
-32000; other action requirements use -32603. Provider-private error detail
is not copied into these responses. Ordinary assistant text is never examined to
infer an entitlement failure. Generic provider exceptions outside these typed
branches retain the vendor behavior and remain an upstream audit item.
qualify-cursor-runtime-patch.mjs executes expressions extracted from each actual
input-digest-verified platform chunk. Its poisoned ambient loaders prove no
ambient access while owned MCP, empty sessions, and duplicate definitions work;
it also executes all typed error branches and ignores entitlement-shaped ordinary
errors. The entire patched chunks compile for macOS ARM64/x64 and Linux x64.
runtime-patch-offline-proof.json records these checks with zero provider calls.
All three authentic archives separately passed materialization and post-patch
closure verification. Seven materializer tests cover drift, ambiguous anchors,
closure tampering, links, archive tampering, identity consistency, and overwrites.
Both patched macOS binaries also returned ACP v1 from credential-free real-process
initialization and exited after explicit SIGTERM; x64 ran through the ARM host
compatibility layer. runtime-patch-initialization.json retains this distinction.
These are offline checks, not authenticated macOS x64 or Linux/Daytona qualification.
Prior Product/Runner results above measured the earlier unpatched runtime. The profile must retain pending qualification until the rebuilt final foundation runtime passes the remaining native interaction, durable approval, cancellation, recovery, accounting, and authenticated target-platform gates.
Frozen profile v3 build
The retained build record
identifies exact runtime source 526081b9d6c905de5e2da680e3a37c620899b380,
profile v3, the ARM64 portable package and matching release daemon, both lock
hashes, the committed ACPX patch hash, and credential-free launch results.
Dependencies were installed into a fresh private store with copied files to
avoid borrowing mutable package bytes from another worktree. The initial frozen
install failed because the committed lock's patch configuration differed; that
failure is retained. A private derived staging lock then passed frozen install.
The repository lock was not changed.
The package's real immutable lease passed ACP initialization, required authentication for model/session discovery and creation, rejected unsupported fork/resume methods, and closed with no stderr. The TypeScript and release daemon builds passed. Twenty-six installed-package contracts passed. Another 184 control tests passed, including two Cursor-specific tests using the real installed ACP SDK stream: sessionless native single/multi-selection and complete revision-bound plan replies preserve request ID zero, await response pipe delivery, reject stale plan revisions, and suppress late answers after cancellation. The surrounding suite covers mocked admission, permission identity, provider-loss expiry, recovery fences, and cleanup ownership. These are deterministic tests, not live native interaction, durable browser approval, or provider-death qualification.
The new cursor-runtime-patch.mjs is present in the Docker build context. It is
also a required Daytona content-identity input: changing patch source must change
the image identity even when another declared input is unchanged. Publishing or
qualification must use an image whose content hash includes that script.
Profile v4: instruction delivery and admission (2026-09-29)
The pinned native ACP implementation ignores _meta.systemPrompt on session
creation, and ACPX does not send it on load. Native workspace rules do not read
Paperclip's registered AGENT_HOME. The hidden --system-prompt option is an
internal-team-gated interactive CLI feature and is not used by this integration.
Those facts invalidate any assumption that prior transport success alone proved
canonical governance or custom entry instructions reached Cursor.
The owned paperclip-cursor-instructions-v2 patch preserves the v3 isolation and
typed-error changes and adds the composed instructions through Cursor's native
LocalResourceProvider.additionalRules path. The sandbox supplies an ephemeral,
controller-owned PAPERCLIP_CURSOR_INSTRUCTIONS payload. Ambient environment
cannot select or replace it, and it is not persisted in ACPX's launch environment
record. The payload binds exact UTF-8 content (maximum 32 KiB, no NUL) to SHA-256;
it includes the already-composed governance, custom entry content and registered
agent-home context, without discovering another workspace instruction file.
Explicit empty content has an empty rules list and the SHA-256 of empty bytes.
Native session creation and load both use this process-owned snapshot and return
_meta.paperclipCursorInstructions with schema, exact digest, and byte length.
The additional rule uses the native global-rule representation and retains
native ordering: additional rules precede the ordinarily discovered project
rules in RequestContext.rules. It is not prepended to the user's message and
does not replace Cursor's built-in system prompt. The synthetic rule identifier
is paperclip://runtime/instructions; no generated instruction file is exposed
to workspace discovery.
A synchronous ACPX protocolGuardFactory checks each connection independently.
It correlates the acknowledgement with the exact new/load request and session,
rejects missing or mismatched acknowledgements before SDK dispatch, and blocks
outbound prompts until that session is admitted. Provider death and automatic
reload reset authority. The adapter also asserts admission immediately after
ensureSession. A best-effort observer cannot authorize a prompt. The shared
patch retains upstream terminal-failure diagnostics, rich interaction callbacks,
response-delivery receipts, and Cursor child-event handling.
The v4 declaration
binds profile digest sha256:b1440d559ebc4eef5c7a582f1c81fc153270cfbafa1731a8ee76d83713bdf61b
to all platform closures and ACPX patch SHA-256
64180c194ab841d6f4bba7e6eca5ead621c6222c9a8057e4bb2bb6f007d8bb3c.
Vendor archive and vendor closure hashes are unchanged. All three pinned archives
passed fresh materialization and the new patched closure check:
| Platform | Patched closure SHA-256 |
|---|---|
| macOS ARM64 | 912a37edb67fd7737809c24049dce2db20b1330f98b15a6e43809f0e0943cb6c |
| macOS x64 | 322f54c7bd533b202a311ce8e2c8916bfe0202790ef8a40fbb36bb89cbd506cb |
| Linux x64 | a375fc771dfe86b9b24532f95956ec4b02c5f4f24d80f5db06565160eb3b672c |
qualify-cursor-instructions.mjs, called by the existing patch qualifier, executes
all three digest-verified native resource constructors and bounded payload logic.
On ARM64 it additionally executes the exact patched new/load methods and native
cache/RequestContext construction with explicit dependency doubles, proving the
composed bytes and native ordering at that boundary. This is an offline source
execution proof, not a claim that an authenticated model obeyed the instructions.
The retained proof
and runtime-patch-offline-proof.json record the limits.
Nine new deterministic tests cover exact UTF-8 and empty instructions, ambient injection exclusion, recovery refresh, request/session/connection authority, actual SDK new/load admission, and actual runtime automatic reload after provider death. Valid acknowledgements permit the prompt; missing or wrong digests fail before prompt bytes. Another 113 adapter/sandbox/identity tests, 25 fresh isolated ACPX contracts, seven materializer cases, and the runner TypeScript compile pass. No paid calls or billing mutations were made. A new final-source provider pack and daemon, authenticated instruction semantics, and the remaining live gates are still required before qualification.
Cursor profile v5: admitted native session modes (2026-09-29)
The explicit acpxSessionMode configuration selects agent (the default),
plan, or ask for Cursor. Historical v5 qualification snapshots used
cursorMode; those saved results retain their original identities. The current
production port uses the generic provider/sidecar mode identifier, separate from ACPX's persistent/oneshot session lifecycle. Shared
TypeScript/Rust transport and recovery code treat it as an opaque bounded string.
Each provider adapter owns supported values, defaults, native translation and
acknowledgement; currently only the Cursor adapter qualifies configurable modes.
Mode is included in the immutable session key and recovery identity;
a missing or changed mode cannot reopen an existing v5 session.
Admission checks the pinned native session/new or session/load response's
mode configuration and mode state. If necessary, it sets the selected mode using
session/set_config_option and requires its exact echoed configuration before
admitting a prompt. A bare session/set_mode response or asynchronous
current_mode_update alone is not acknowledgement. Every new native connection
gets a fresh guard: a mismatched automatic reload blocks prompt delivery, and an
unsolicited active mode change fails closed. Explicit startup/reopen admission
can reassert the selected mode within the existing admission deadline and renews
the consumed launch lease after each temporary control connection.
The pinned CLI maps Agent to native default, Plan to plan, and Ask to
search, and carries that metadata into the native UserMessage mode. The
inspectable mode-v5-offline-proof.json executes the pinned native methods with
storage, transport, and enum-converter doubles. Real ACPX fixture tests separately
prove cold-manager load/control ordering and automatic reconnect rejection before
prompt bytes. Neither is paid model/tool qualification.
Native mode does not change Paperclip permission policy, company governance, or filesystem isolation. The native descriptions say Plan is read-only planning and Ask has no edits or command execution; these descriptions are not an OS sandbox claim. Native ACP CreatePlan acceptance does not itself mutate session mode in the pinned handler. An approved plan therefore stays in the selected Plan mode; there is no implicit promotion to Agent mode. Actual native question/plan callback availability, Plan-mode governed completion, and accepted/revised artifact flows remain pending Product qualification. Existing v4 paid results and macOS x64 Rosetta build evidence remain historical and do not qualify v5.
The v5 declaration
binds digest sha256:aa8c0b2b84786982bcd06b7634bf95be6f2bc42bb8def48a8751dc50cf739b6b.
Profile v5 retains the native distribution closures and instruction patch. Its
canonical declaration adds the admitted mode policy and binds ACPX patch
79aad2d688b03362e8cfcbf7a08f78a8383869f6882a9c9ed66f1d18efb94f2b, which preserves
identified empty rich-input chunks at the native parser boundary.
Accepted native plans wait for explicit continuation (2026-09-30)
With Cursor CLI 2026.09.26-dd393fe, accepting native CreatePlan can finish the
planning turn normally without a Paperclip semantic completion result. The
controller recognizes this boundary only from the admitted Plan-mode run,
company/task scope, exact native plan request and accepted revision, acknowledged
human response delivery, and normal terminal event. An explicit semantic result
still takes precedence.
The planning run succeeds while the task stays in_progress. Its durable comment
says: “Plan accepted. This task is waiting for your next message. This run used
Plan mode; no implementation or task completion is claimed.” Acceptance neither
changes the mode nor schedules implementation or an automatic follow-up. A user
can send a new task message to continue; that message follows ordinary admission
with the settings selected for the new run.
The committed wait survives controller restart, agent pause, changes to settings for future runs, and later profile-catalog revisions. Those changes are not permission to start task work. New waits must match the current qualified profile; recovery retains an older committed wait only while its original admission, contract, native request/answer/terminal proof, accepted-result identity, assignment, and applied task decision remain unchanged. A new user message or superseding task decision/run ends that wait's authority. Altering the original proof does not receive the historical-profile exception.
Deterministic tests cover real result persistence and transactional finalization, stale delivery rejection, restart/recovery without automatic wakes, catalog and future-setting changes, ordinary user continuation, and semantic-finish priority. The Product fixture preserves its native decision and workspace no-effect checks and now expects a successful planning run with an unfinished task and visible next-message guidance. Paid requalification of this controller behavior is still pending; the retained earlier Product failure is not reclassified as a pass. Runner, sidecar, provider distribution, profile, and image bytes are unchanged by this controller settlement change.
Cursor7 native plan lifecycle binding
The retained ca702 Cursor6 local plan attempt delivered both native decisions,
including the revised plan's acceptance, but failed controller settlement. The
provider emitted the accepted CreatePlan tool's activity after the callback
request had been created. The earlier proof treated that activity as unrelated
work. Workspace snapshots remained unchanged; that failed attempt is preserved.
Cursor7 binds the native callback's parent tool identity to the canonical
runtime_request item identity through both the sidecar and direct driver. The
controller requires exactly one successful lifecycle for that same tool, in the
same session and turn, completing after the accepted answer and before the normal
turn terminal. It rejects unrelated activity during this interval, failed or
missing lifecycle records, and altered durable proof. Tool names and plan titles
are not identity evidence. The committed receipt includes the correlated rows'
digest. Existing exact Cursor6 committed waits retain their historical contract;
Cursor6 cannot create a new wait or reopen under Cursor7 admission.
Focused deterministic checks cover the actual observed event order, identity normalization boundaries, cancellation/expiry reconstruction, unrelated tools, and persisted historical waits and proof tampering. Current runtime/sidecar/pack builds are complete; Cursor7 paid qualification remains blocked by the failed settlement case in the checkpoint above. This changes the runtime projection and profile contract, while the pinned Cursor native executable and distribution patch stay unchanged.
Runner Eval diagnostic retention
The live-session recorder admits the bounded cursor_native_usage_observed
notice only for the current Cursor session and turn. It validates the fixed
provenance/reasons and generated observation identities, rejects unknown or
oversized fields, and preserves parent/child observations through durable eval
checkpoint reload. These unsummed, partial counters never enter the usage ledger
or establish native USD. Deterministic transport/store tests cover this capture
path; paid qualification remains pending. Earlier artifacts produced by the old
recorder cannot establish whether these native counters were emitted and must
retain their original accounting failure.