Commit Graph
10 Commits
Author SHA1 Message Date
DottaandPaperclip a6306ba606 feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native Runner keeps provider sessions under company authority,
approvals, budgets and durable recovery.
> - Cursor work was spread across candidate branches. The published
branch lacked later plan, permission and cleanup fixes.
> - Production also needs public installation and matching runtime
assets for local and Daytona execution.
> - This pull request consolidates Cursor onto current mainline recovery
behavior and completes that installation path.
> - The installed v11 release passed focused local and Daytona
qualification after the generic mode and lifecycle cleanup. The later
model-selection correction and current mainline merge produce v14
artifacts that need matching release qualification.
> - Cursor admission is enabled in source; publish only an artifact
combination with matching qualification. Native AskQuestion and complete
per-run dollar accounting remain excluded.

## Linked Issues or Issue Description

Refs: #14435, #14631, #14669, #14699, #14724.

This completes the Cursor implementation by @cryppadotta from combined
source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer
mainline recovery, completion and warm-directory behavior. Pi and
Copilot remain gated.

## What Changed

- Generate named Rust and TypeScript ACPX release profiles from one
manifest. Share runtime pins with packaging and server verification.
Preserve vendor runtime versions; bind the updated ACPX patch to Cursor
profile v14 and reject stale generated declarations at build/typecheck.
- Remove ACPX model allowlists, including the former Codex and Pi
restrictions and the duplicate developer test-drive gate. Send any
explicit model ID unchanged to its provider and verify the effective
selection before prompting. The bundled ACPX package forwards unlisted
IDs, rejects mismatched acknowledgements, and restores the exact
selection after session load. It does not expand Cursor model aliases.
Provider rejection, mismatch, or missing model controls fails without a
fallback. Model examples live in evaluation fixtures, outside runtime
declarations.

- Add pinned Cursor execution, contained instructions, exact model
verification and Agent/Plan/Ask modes.
- Carry an opaque generic `mode` identifier in shared native execution,
sidecar, Rust and recovery contracts. The provider adapter owns
supported modes, defaults, native translation and acknowledgement.
- Keep native RPC recognition, accepted-plan interpretation and
permission evidence behind provider adapters. Shared settlement and
recovery verify normalized facts and their committed evidence.
- Replace the Cursor-only warm-attachment branch with a runner-owned
capability. Only Cursor opts into it. Move profile compatibility and
optional usage parsing into provider metadata and adapters.
- Write generic plan-wait receipts. Read exact historical Cursor
receipts through a separate compatibility decoder. Reject mixed formats
and preserve existing authority checks.
- Carry native plans, semantic questions, todos, child activity,
permission identities and partial usage diagnostics through the Runner.
- Preserve durable response delivery, cancellation, warm ownership and
process retirement.
- Finish accepted planning runs successfully. Keep their tasks open for
explicit direction. Acceptance does not start implementation.
- Ship `paperclipai runtime setup cursor` and its provisioner through
the public package. npm installation does not download Cursor. Setup
uses the OS account's closure-keyed cache so system-wide npm packages
can remain read-only. Run it as the Paperclip service account.
- Include Cursor in normal provider packs and Daytona images for macOS
ARM64/x64 and Linux x64.
- Reject stale release packs by source revision and current ACPX/Cursor
pins before assembly writes files. Verify current Cursor
version/profile/closure again at runtime.
- Ship all three daemon targets and the expected Linux image-pack
identity. A macOS controller uses its packaged Linux daemon for Daytona.
Image mismatches fail before provider launch.
- Use the vendored Runner boundary for installed readiness probes.
Verify the actual installed Cursor probe.
- Verify compiled public Daytona plugins and their release versions in
installed smokes.
- Record exact artifacts, the acceptance matrix, retained failures,
supported capabilities and rollback behavior in the [readiness
report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md).

## Verification

- Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges
mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor
plan and cancellation guards alongside mainline historical-question
filtering. The evaluation catalog includes both Cursor and expanded
adapter accounting cases (683 total). Recursive typecheck, full build,
696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck
passed. Current-head CI passed: 56 successful checks, one neutral and
four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535).
The fresh Base Greptile review is 5/5 on this exact head, with 304 files
reviewed, zero new comments and zero unresolved threads. The user
authorized overriding the CODEOWNER review gate after checks passed; no
failing checks are overridden. Prior results below retain their own head
identities.
- Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the
post-merge Apex finding. Automatic-review and new-evidence
reconciliation preserve pending child results and recheck delivery under
the status lock before completing. Account repair now excludes unrelated
secret consumers and requires the failed agent's identity. Regression
coverage includes the commit race, delivery statuses,
current-run/current-intent exclusions, repeated reconciliation, both
database reconciliation paths, and credential consumer boundaries. All
184 affected tests, server typecheck and server build passed.
Current-head Base Greptile review is 5/5, with 304 files reviewed, zero
new comments and zero unresolved threads. Current-head CI passed: 56
successful checks, one neutral and four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724).
This Base review is distinct from the earlier Apex review.
- Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves
both accepted-plan waits and pending-child-completion checks, current
provider selectors, task-creation response identities, and mainline ACPX
missing-file handling. The combined patch is bound to Cursor profile
v14; historical records keep their original identities.
- Merge head `5957c257a` passed recursive typecheck, full build, 43
installed ACPX/package contracts, 107 provider UI and plan/recovery
tests, 593 database-backed lifecycle tests, 49 profile/native contract
tests, 45 Product E2E fixture tests, fixture typecheck, token gates,
three provider-free browser task-creation cases, and Runner
conformance/replay checks. Its complete CI passed (55 successful checks,
one neutral and four skipped), while Apex returned 2/5 with a
child-delivery finding addressed below.
- The local full-suite attempt again failed the unchanged Git streaming
test (360-second timeout) and was stopped. The concurrent local Rust
attempt failed four unchanged Codex process/deadline tests; all four
passed serially without code changes in 7.29 seconds after removing the
competing test load. These failed commands are retained and are not
reported as full-suite passes; the fresh Linux CI runs are tracked
separately.
- The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned
Apex 5/5 with zero comments after fixing all three findings: per-user
install cache, stale release-pack rejection, and public Linux smoke
account/home handling. Its real built installer passed from read-only
public packages on macOS ARM64 and Linux x64. All 137 release-registry
checks and 64 ACPX package contracts passed. That review does not cover
this mainline reconciliation.
- Prior `beadd3654` passed the full CI matrix; its one unchanged chat
test failure and successful single retry remain in the [CI
history](https://github.com/paperclipai/paperclip/actions/runs/37521449327).
Historical results below remain attributed to their original builds.

- Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes
provider-specific model choices from generic offline ACPX tests. The
fake sidecar preserves the model and session identity selected at open
through suspension. Affected verification passed: 106 Rust tests and 73
TypeScript tests. This commit changes test code only; the
production-code checks below retain their recorded identities. Its CI
and Greptile review later passed; those results belong to that
historical head.
- Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`:
252 focused Runner tests passed (six platform skips), covering all six
ACPX agents, native model acknowledgement, rejected selections,
installation integrity and recovery identity. The merged branch passed
recursive typecheck, full build, token gates, server admission (19
tests), and the Product E2E catalog (45 tests). The acceptance catalog
passed all four tests. The full Rust suite passed: 643 tests, 2 ignored.
It verifies sidecar acknowledgement of unlisted models and rejection of
model mismatches. The final commits only update Rust tests; production
sources match the verified build at
`65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls
were made.
- The merge preserves both Cursor and the new mainline public-MCP
fixture cases. Auto-merge remains disabled; the latest follow-up status
is recorded above. The local `pnpm test:run` attempt hit the unchanged
Git streaming test's 300-second timeout and was interrupted before
merging mainline. The broad Runner attempt found obsolete single-model
assertions plus three macOS fixture-path failures caused by a
`/private/tmp` override. The assertions are corrected; affected
TypeScript checks passed with the standard macOS temporary directory,
and the complete Rust suite passed. Neither interrupted command is a
full-suite pass.
- Earlier declaration-cleanup head `6f4a5e9e2` passed recursive
typecheck, build, Rust and focused tests. Its CI later exposed a test
expecting duplicated Grok digest literals. The current source fixes that
assertion to compare launcher bytes with the shared manifest. Historical
successes and failed attempts are retained; no new live provider
qualification is claimed.
- Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed
complete CI (56 successful checks, one neutral, four skipped) and
Greptile 5/5. [Historical complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305).
Those results are not claimed for the cleanup head.
- Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`.
Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The
declaration cleanup preserves release pins and does not relabel that
tested artifact as a build of the new source. Mainline through
`e34abee670` was reconciled while preserving accepted-plan waits,
provider-capacity handling, and both Cursor and public-MCP fixtures.
- Clean normal installation, explicit Cursor setup and daemon resolution
passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm
lifecycle hooks ran without silently downloading Cursor.
- Historical v11 live matrix: **18/18 passed with cleanup** (nine local,
nine Daytona) after the generic mode and lifecycle cleanup. The campaign
has 23 attempts; all five failures and their diagnoses remain recorded.
Exact case identities, hashes and limits are in the readiness report.
All provider calls are real, use the explicit Luna model and
company-bound credentials, and run without qualification or
runtime-asset overrides.
- The immutable Daytona image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`.
The public Daytona plugin is installed independently and its version is
checked.
- Recursive typecheck, full build, token gates and Runner
contract/conformance/replay checks passed on the frozen application. Its
complete Linux CI suite passed. The duplicate local full-suite command
was incomplete after timing failures; affected repeats passed, but that
command is not reported as a clean pass.
- Qualification fixtures passed typecheck, 1,675 Vitest tests (one
skip), 128 Node checks, three provider-free browser tests, and 150
focused lifecycle tests after the final diagnostic correction. The
affected legacy Cursor command file also passed all five tests after
removing its shorter 10-second override; it now inherits the suite’s
standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no
unresolved review threads. CI results above are recorded separately from
historical build results.

## Risks

- Cursor v14 includes the updated ACPX dependency patch and release
identity. The v11 live matrix and image below remain historical
evidence. They do not certify new v14 package/image artifacts.

- ACPX accepts models beyond the qualification fixtures. Availability
and entitlement depend on the provider. Successful configuration is not
a claim of live qualification for every model.
- Shared mode is an opaque identifier. Provider adapters own its
meaning. Incompatible historical sessions remain fenced; exact committed
plan waits and task history remain inspectable.
- Native AskQuestion is excluded. Paperclip semantic questions are
supported. Authoritative per-run dollar accounting is unavailable;
partial counters remain diagnostics and unknown cost is not zero.
- Image input, detailed native diffs, deeper child transcripts and
native plan-file export remain follow-ups.
- macOS x64 has clean-install and daemon-startup proof under Rosetta,
not a separate live campaign on Intel hardware.
- Release only the tested package/image combination. Merging this PR
does not publish npm packages or deploy that image. Later builds need
their own release verification. Rollback disables new Cursor admission
while preserving records and recovery inspection.
- A model can fail an exact instruction: one cancelled-plan attempt
returned the wrong summary marker despite correct cancellation. The
unchanged repeat passed; both results remain in the report.

> ROADMAP.md was checked. This completes existing native Runner/Cursor
work; it does not add an independent core feature proposal.

## Model Used

OpenAI Codex, GPT-6. The exact serving variant and context window are
not exposed in this session. The agent used reasoning, repository
inspection, code execution, protocol tests and browser-backed Product
E2E tools. Cursor acceptance uses the explicit
`gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is
the evaluated provider model.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — affected suites passed;
full CI and the retained local failed attempts are recorded separately
above.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green — 56 successful checks, one
neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`;
zero new comments and no unresolved threads. The earlier Apex finding
remains fixed.
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:48:15 -05:00
6c36c07a4f feat(adapters): add GPT-6.1 Sol and refresh shared coding harness pins (#14942)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents run through coding-agent adapters and the native runner. Both
use the same installed provider CLIs, model catalogs, and reasoning
controls.
> - OpenAI released GPT-6.1 Sol (`gpt-6.1-sol`) in Codex. Anthropic
released Claude Sonnet 5.5. The static Codex, Bedrock, and OpenCode
catalogs do not list these IDs.
> - The shared provider pack pins Codex 0.156.0 and OpenCode 1.18.32.
The evaluation image pins older Grok, Gemini, Kimi, Cursor, and GitHub
CLI releases. Codex 0.156.0 has no bundled metadata for GPT-6.1 Sol.
> - A model entry without a current harness, or a harness pin without
its runner integrity checks, fails at run time.
> - This pull request adds the verified model IDs and moves the harness
pins, executable digests, controller checks, and image pins together.
> - The benefit is that operators can select the current models, and the
native and local adapters share one current CLI installation.

## Linked Issues or Issue Description

Refs #13829 and #13838 (the September 22, 2026 model and harness
refresh). Related pull requests: #14993 (merged October 5, 2026,
superseding #14816) added the direct Claude Sonnet 5.5 entry and
refreshed the Claude runtime to Agent SDK 0.3.286 / Claude Code 2.1.286.
This pull request does not change the Claude runtime or the direct
Claude model list; it keeps the #14993 pins and adds only the Bedrock
Sonnet 5.5 ID. After #14993 merged, this branch was rebased onto
`master` (October 5, 2026). The six overlapping pin regions
(`docker/daytona-runner/Dockerfile`, `docker/daytona-runner/README.md`,
`package.json`, `pnpm-workspace.yaml`,
`packages/adapters/claude-local/src/index.test.ts`,
`packages/paperclip-runner/src/backends/native-backend-factory.test.ts`)
were resolved by keeping this pull request's Codex 0.160.0 and OpenCode
1.18.34 pins next to #14993's Claude 0.3.286 / 2.1.286 pins, taking the
union of the Sonnet 5.5 model IDs in the Claude test, and merging both
README paragraphs. The Sonnet 5.5 effort and CLI-gate lines in the
Claude adapter were identical in both pull requests and merged without a
diff. #14917 and #14918 reordered the Claude and Codex model lists
earlier; the new entries sit where those ordering rules put them.

Sources checked on 2026-10-02:

- [OpenAI Codex models](https://learn.chatgpt.com/docs/models): GPT-6.1
Sol uses `gpt-6.1-sol`, supports reasoning efforts from Light to Ultra,
and has Standard and Fast modes at launch. The page also records that
`gpt-5.4` and `gpt-5.4-mini` retired from Codex with ChatGPT sign-in on
August 31, 2026, and that `gpt-5.5` retires on October 14, 2026. Neither
retirement applies to the OpenAI API.
- [Codex CLI releases](https://github.com/openai/codex/releases) 0.157.0
through 0.160.0. The bundled model metadata in the 0.160.0 Linux binary
contains `gpt-6.1-sol`.
- [Claude Sonnet
5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview):
Bedrock ID `anthropic.claude-sonnet-5-5`, released September 28, 2026.
- [OpenCode releases](https://github.com/anomalyco/opencode/releases)
1.18.33 and 1.18.34 (fixes only). The OpenCode model registry lists both
added provider-qualified IDs.
- npm `latest` tags for `@xai-official/grok` 1.0.46,
`@google/gemini-cli` 0.62.0, and `@moonshot-ai/kimi-code` 2.1.1.
[xAI](https://docs.x.ai/docs/models),
[Google](https://ai.google.dev/gemini-api/docs/models), and
[Kimi](https://www.kimi.com/code/docs/en/kimi-code/models.html) list no
newer coding models.
- Cursor CLI 2026.10.01-e373342 is the version the official installer
resolves. The pinned digest is the SHA-256 of the versioned Linux x64
archive.
- [GitHub CLI 2.102.0](https://github.com/cli/cli/releases/tag/v2.102.0)
(security fixes). The pinned digest matches the release `checksums.txt`.

## What Changed

- Codex adapter: add `gpt-6.1-sol` to the model list, the Fast mode
list, and the Ultra effort set. It is the first entry: #14918 orders the
list newest version first, and its description notes the ChatGPT app
lists GPT-6.1 Sol first. Update the adapter documentation text.
- Claude adapter: add `us.anthropic.claude-sonnet-5-5` (Bedrock Sonnet
5.5) to the Bedrock catalog in the newest-Sonnet slot after Opus 5.5;
`us.anthropic.claude-sonnet-5` moves into the older-Sonnet group,
matching what `sortClaudeModels` from #14917 produces at runtime. Any
Sonnet 5.5 ID (direct or Bedrock-qualified) now gets the documented
`xhigh` and `max` efforts and requires Claude Code 2.1.284 or later on
the CLI lane (the Claude Code changelog entry for 2.1.284 adds
`claude-sonnet-5-5`). These two lines are identical to the ones #14993
merged, so the branch carries no diff for them.
- OpenCode adapter: add `openai/gpt-6.1-sol` and
`anthropic/claude-sonnet-5-5` to the static fallback catalog.
- Codex runtime pin 0.156.0 → 0.160.0 in the root and workspace
overrides, the runner package, the Codex ACP package patch, the
qualified ACPX profiles, the Linux x64 executable digest, the Rust
provider backend and its tests, the provider-pack manifest pins, the
remote controller pins, the sandbox npm install spec, and the opt-in
qualification scripts.
- Remote Codex compatibility window: upper bound 0.157.0 → 0.161.0. The
minimum stays at 0.149.0.
- OpenCode runtime pin 1.18.32 → 1.18.34 in the runner package, the
materialization script, the server and Rust qualified versions, the eval
and live-session labels, fixtures, and the configuration label.
- Evaluation image (`docker/daytona-runner/Dockerfile`): Grok CLI
1.0.46, Gemini CLI 0.62.0, Kimi Code 2.1.1, Cursor CLI
2026.10.01-e373342 with its digest, GitHub CLI 2.102.0 with its digest,
Codex and OpenCode version probes, and the refreshed lockfile digest.
The Claude Code 2.1.286 probe comes from #14993 and is unchanged here.
- `pnpm-lock.yaml` is not part of this pull request. The repository's
pull request gate rejects lockfile edits, and the refresh bot
regenerates the lockfile on master (the same flow #13838 used). The
Dockerfile `PAPERCLIP_RUNNER_LOCK_SHA256` default is the digest of the
lockfile that `pnpm install --resolution-only --ignore-scripts
--no-frozen-lockfile` (the refresh workflow's command) produces for the
combined pins on the rebased branch (`e1856797…`); that lockfile differs
from master only in the `@openai/codex` 0.160.0 platform packages, the
`@anthropic-ai/claude-agent-sdk` 0.3.286 override that #14993 introduced
(the open refresh-bot pull request #14872 carries that part),
`opencode-ai` 1.18.34 with its Linux x64 baseline, and the `codex-acp`
patch hash.
- Documentation: runner README, runner compatibility doc, environment
variable example, and a new `doc/adapter-model-audit-2026-10-02.md` with
sources and deferred items.
- Tests: Codex adapter catalog, server adapter models, Codex
compatibility window, native session executor pins, runner package
contract, OpenCode materialization, and UI effort options.

Unchanged on purpose: Claude Agent SDK 0.3.286 / Claude Code 2.1.286
(already on `master` from #14993), ACP bridges (`acpx` 0.13.1,
`claude-agent-acp` 0.73.0, `codex-acp` 1.6.2; newer upstream releases
need a separate qualification), the native Grok runtime 1.0.13, Pi
0.84.2 / 0.87.1 (the Pi 1.0 runner stack covers it), and Hermes 0.19.0
(current). `gpt-5.4` and `gpt-5.4-mini` stay in the picker because the
OpenAI API still serves them.

## Verification

Run on Linux x64 with Node 25.9.0 and pnpm 9.15.4 after `pnpm install
--no-frozen-lockfile` (the refreshed lockfile stays local; see above).
The results below were re-run on the rebased head (October 5, 2026) for
the suites the conflict resolution touches; the other rows are from the
original run and are covered by CI on every push:

- Rebased head: `packages/adapters/codex-local` 482 passed;
`packages/adapters/claude-local` 340 passed, 4 failed (`execute.remote`,
`test.probe`, `execute.acp-fallback`, `acp` spawn/env-hardening cases
that fail identically on unchanged `master` in this host environment);
`server` adapter-models + codex-runtime-compatibility +
native-session-executor + adapter-registry 607 passed, 1 failed (the
same adapter-registry override-pause case as before, also failing on
`master` here); `packages/paperclip-runner` native-backend-factory +
qualified-profiles 36 passed; `ui` codex-reasoning-effort +
config-fields + model-utils 19 passed. Rust, full typecheck, build, and
the Docker image are left to CI as before.

- `vitest run` in `packages/adapters/codex-local`: 13 passed. `vitest
run` in `packages/adapters/claude-local` (whole package, including the
new Sonnet 5.5 gate and effort tests): see the latest CI run and the
comment below. `vitest run` in `packages/adapters/opencode-local`: 48
passed, 1 failed (`runtime-config.test.ts` reads the host
`PAPERCLIP_OPENCODE_PROVIDERS` variable; it fails the same way on the
unchanged base).
- `vitest run src/__tests__/adapter-models.test.ts
src/services/native-runtime/codex-runtime-compatibility.test.ts
src/__tests__/adapter-registry.test.ts` in `server`: 84 passed, 1 failed
(`adapter-registry.test.ts` override pause test; it fails the same way
on the unchanged base).
- `vitest run` in `ui` for `codex-reasoning-effort`,
`agent-setup-fields`, `config-fields`, and `ComposerRunSettingsPicker`:
25 passed.
- `node --test test/acpx-codex-package-contract.test.mjs
scripts/materialize-opencode-binary.test.mjs
scripts/runner-protocol-eval-campaign.test.mjs` in
`packages/paperclip-runner`: 23 passed. The package contract test
verifies the installed Codex ACP executable digest and the 0.160.0 patch
pin.
- `vitest run src/drivers/acpx src/backends src/drivers/opencode
src/live/live-session.test.ts` in `packages/paperclip-runner`: 626
passed, 5 failed, 1 skipped. The 5 failures
(`installation-integrity.test.ts` `/proc/self/fd` module loading and one
OpenCode answer-selection test) also fail on the unchanged base under
Node 25; Linux CI runs Node 24.
- `pnpm run test:opencode:qualification` in `packages/paperclip-runner`
against the installed OpenCode 1.18.34 executable: passed.
- `codex --version` from the installed pack prints `codex-cli 0.160.0`.
The Linux x64 executable digest `12eb3e81…652aad` was computed from the
`@openai/codex@0.160.0-linux-x64` archive after checking its registry
`dist.integrity`.
- `pnpm check:token-gates`: all gates clean.
- `pnpm run typecheck:typescript` in `packages/paperclip-runner`:
passed. Package typechecks ran one at a time; see the comment below for
the server and UI results.

Not run here, and needed from CI:

- Rust tests and `pnpm -r typecheck` / `pnpm build` for the server (no
`cargo` in this environment; the server typecheck prepares the runner
vendor build).
- The Docker evaluation image build and the real-binary Codex startup
and session-resume probes (no Docker; the probes need the compiled
`paperclip-runnerd`). The trusted CI runner workflow covers them.
- Authenticated inference with any new model. This change is metadata
and startup validation only.

## Risks

- Codex 0.160.0 changes the bundled model catalog and app-server
behaviour (authoritative provider catalogs, incremental running-turn
tracking). The patched `codex-acp` 1.6.2 bridge is unchanged and
declares `^0.148.0`; it worked with 0.156.0 under the same override. If
CI probes show a protocol change, the pin can return to 0.156.0 by
reverting this pull request.
- The compatibility window upper bound moves to `<0.161.0`. Remote
images with Codex 0.157 to 0.160 become accepted. Older images stay
accepted down to 0.149.0.
- Until the refresh bot lands the regenerated lockfile on master, the
Dockerfile lockfile digest default does not match the committed
lockfile. The trusted CI workflow computes the digest from its own
resolution at build time, so this affects only a local build that passes
no digest.
- Existing saved model selections and effort settings are not changed.
Agents on `gpt-5.4` or `gpt-5.5` with ChatGPT sign-in need a model
change before the OpenAI retirement dates; that is documented, not
enforced.
- Rollout order: deploy the controller and runner from this change
before promoting a sandbox image that carries these pins. Older
controllers reject the new provider-pack pins.

## Model Used

- Claude Fable 5.1 (Anthropic, model ID `claude-fable-5-1`), 1M context
window, adaptive thinking, tool use. The model ran as a Paperclip agent
through the Claude Code harness, performed the web research, edited the
code, and ran the tests listed above.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Bender (Fable) <noreply@paperclip.ing>
Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 13:59:08 -07:00
cab4263dc9 feat(claude-local): add Sonnet 5.5 and refresh the qualified Claude runtime (#14993)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Claude local adapter lists models for users with a Claude
subscription.
> - The list needs Claude Sonnet 5.5 and its supported effort levels.
> - Sonnet 5.5 needs Claude Code 2.1.284 or later on both execution
paths.
> - The qualified ACP runtime previously used Claude Code 2.1.280.
> - This change adds Sonnet 5.5 and pins Agent SDK 0.3.286, which
includes Claude Code 2.1.286.
> - Users can select the model and run it with a qualified runtime.

## Linked Issues or Issue Description

Refs #3936. This replaces #14816 because the maintainer integration
cannot write to the contributor fork.

Thank you to @SkilLab-Tech for the model support, runtime refresh,
tests, and platform digest verification. This branch preserves both
original commits: `8d2f3261af61a2ac1120e51e8a8618732ace543b` and
`e68d1d002a3ed745f016fb11c50ac5a3c5a9ff8d`.

Related work:

- #14917 added Claude model ordering. This branch includes that merged
change and resolves its conflicts with #14816.
- #14942 updates the other models and harnesses. It remains separate.
Its matching Sonnet effort and CLI-gate changes are identical. Both PRs
merge with master. The second PR will need a rebase after the first
merges because adjacent runtime-pin and test edits conflict.
- #14954 is another Sonnet 5.5 change. It overlaps with the model
additions but does not include the qualified runtime refresh.
- #14039 makes the per-task effort picker model-aware. #3937 is also
related to effort selection.

The original author checked the [Claude Code
changelog](https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md)
and [effort
documentation](https://platform.claude.com/docs/en/build-with-claude/effort)
on 2026-10-01.

## What Changed

- Add the direct `claude-sonnet-5-5` model and Low, Medium, High,
X-High, and Max effort levels.
- Require Claude Code 2.1.284 or later for that model on the CLI path.
- Put Sonnet 5.5 after Opus 5.5 in the current-model group. Keep Sonnet
5 in the older-model group.
- Retain the Sonnet 5.5 assertions and the model-order assertions in the
server tests.
- Pin Agent SDK 0.3.286 and Claude Code 2.1.286 across overrides,
integrity digests, qualified profiles, Rust provider pins, and the
Daytona version check.
- Update the related adapter and runtime documentation.

## Verification

Local verification uses the resolved source tree and pnpm 9.15.4. Model
tests passed on Node 25.9.0. Runtime integrity tests use CI's Node
24.21.0.

- Five focused Claude test files pass: 62 tests. They cover model
defaults, model ordering, CLI gates, and remote execution probes.
- Server model-list and UI setup tests pass: 30 tests.
- The Claude adapter typecheck passes.
- Runner integrity and qualification tests pass on Node 24.21.0: 78
tests. Four descriptor-loader tests fail on Node 25.9.0; all four pass
on the CI version.
- The runner package contract passes: 10 tests.
- Full local typecheck stopped with exit 137 in the database package
under the container's 4 GB memory limit. The production build reached
the runner Rust build, then stopped because `cargo` is absent.
- The full stable local Vitest run was stopped after all current-head CI
test shards passed. It did not complete locally. The 180 focused tests
listed above passed.
- `git diff --check` passes. The branch changes 20 files against master.
It has no lockfile or workflow changes.
- The original author verified all three platform digests against
registry integrity and ran the Linux executable. Its version was
`2.1.286 (Claude Code)`. See #14816 for that evidence.
- Greptile reviewed head `6db3d3f1` and gave 5/5 with zero comments.
Both Superagent scans and Commitperclip pass. All current-head CI jobs
pass, including build, typecheck, Rust, test shards, browser tests, and
the canary dry run.

## Risks

- The controller and provider pack must use matching runtime pins.
Deploy them together.
- CI owns `pnpm-lock.yaml`. The master lockfile refresh must resolve the
SDK override. Refresh the Daytona lock digest with that lockfile.
- Images built with Claude Code older than 2.1.284 need a rebuild before
the CLI path can use Sonnet 5.5.
- The runtime remains at SDK 0.3.286. This PR does not take the later
0.3.287 patch.
- A live Sonnet 5.5 session and a Daytona image build are not part of
the local verification.

## Model Used

- Original work: Anthropic Claude Code, `claude-sonnet-5-5`. Review:
`claude-opus-5-5`. The author reported `xhigh` effort, tool use, and
code execution. The original context window was not reported.
- Merge repair and PR preparation: OpenAI Codex, based on GPT-6, with
tool use and code execution. The runtime does not expose the exact model
identifier or context window in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

## Squash Attribution

Keep these trailers in the squash commit to preserve the original author
and AI attribution:

```text
Co-Authored-By: Claude Code (Ivan) <SkilLab-Tech@users.noreply.github.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
```

---------

Co-authored-by: Claude Code (Ivan) <ivan@skillab.com.br>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 11:03:20 -07:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandPaperclip 18e8c121d9 fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path

> - Paperclip manages agents through a shared native runner.
> - Built-in harness support should ship with Paperclip's public
distribution.
> - Grok already speaks ACP; it does not require a new public bridge
package.
> - Sandbox provisioning owns the native executable and its pinned
version.
> - The runner must verify that prerequisite without downloading it
during npm installation.
> - This change separates built-in launcher identity from external
runtime identity.
> - Clean npm installation and live staging checks verify the
distribution boundary.

## Linked Issues or Issue Description

Refs #13882, #13973, #13977, #13979.

This follow-up now targets master after #13882 was squash-merged. It
replaces the private `@paperclipai/grok-acp` workspace package with
runner-owned assets. Current master is included so the branch also
contains the merged scheduler, complete-event capture, and durable
cleanup fixes.

## What Changed

- Ship Grok launcher and qualification metadata inside the runner's
compiled output and the public server's vendored runner tree.
- Remove the separate Grok npm package and all package-manager install
hooks for this runtime.
- Require the checksum-verified Grok Build 1.0.13 binary at
`/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution
environment. Provision it explicitly in the Daytona image and CI setup.
- Keep native binaries outside the provider pack. Bind the built-in
launcher into the pack manifest.
- Preserve executable leases, descriptor-backed startup, credential
fences, permissions, and exact ACP model admission.
- Use `builtin:grok-acp` and `native:grok` as profile identities.
Historical package-profile sessions fail closed on resume rather than
being silently reinterpreted.
- Resolve built-in assets from the authenticated sidecar location,
including public server npm layouts. Keep the controller path out of
provider environments.
- Add clean npm tarball installation verification to the existing
trusted canary CI job and the admitted manual EC2 verification path. It
stages a unified release version and runs npm lifecycle scripts, then
verifies missing-prerequisite rejection and admission after separate
provisioning without credentials or inference.
- Include the controller-owned provider pack in stamped Cloud images.
Unstamped local images omit the pack and remain usable; remote ACPX
requires full source provenance.
- Correct CLI approval-page metadata for an already authenticated Cloud
board user; approval authorization remains unchanged.
- Honor explicit native-runner enablement in the Cloud agent picker and
direct setup page, keeping the flag disabled by default.
- Allow selecting the execution environment before connecting
credentials. Include Grok in the existing authenticated hello-probe
flow, targeting its pinned native prerequisite for runner setup.
- Recover an existing subscription sign-in conflict through an explicit
cancel-and-retry action, serialized after cancellation succeeds.
- Preserve the selected ACPX harness before normalizing config fields,
so new Grok agents use the Grok default model.
- Keep the credential-free Cloud provider pack root-owned and readable
after runtime UID remapping; verify manifest and referenced asset access
under an unrelated unprivileged UID during image builds.
- Archive prior failover backups alongside explicitly replaced harness
state, preserving evidence while preventing stale backups from blocking
a fresh replacement.
- Update Daytona image content inputs and contract tests for the
built-in assets and explicit provisioner.
- Document and regression-test the shared `approve-all` default for Grok
setup, saved configuration, and native execution. Explicitly saved
restrictions remain unchanged.

## Verification

Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05`
incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the
base PR was squash-merged. All 12 conflicts came from incoming files
identical to the tested pre-squash base. The final tree exactly matches
a three-way merge using that original base, preserving built-in Grok
distribution and removal of the obsolete private package. All 252
focused runner/UI tests, six npm-isolation tests, and token gates pass.
Fresh exact-head Greptile review is 5/5 with no outstanding findings;
security scans and EC2 native compilation pass. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([run
36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)).
The repository owner explicitly authorized bypassing code-owner approval
after all checks passed; no CI checks or repository protection settings
are bypassed or changed. The only remaining PR was removed from the
completed stack metadata to permit native auto-merge.

Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)).
The initial attempt lost two EC2 runners to shutdown signals and stalled
a third shard during dependency preparation; all three passed the
same-commit failed-job-only retry. Trunk code-owner requirements remain
enforced. The review summary’s non-blocking saved-asset offset
classification note concerns code already merged in #14301; those
runtime files are identical to master and outside this stack’s diff.
Historical live evidence below retains its original source revisions.
[Final public npm
verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764)
passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public
packages, an executed offline lifecycle sentinel, unchanged consumer
lock, built-in launcher, missing-prerequisite rejection, and verified
separately provisioned binary/command lease. Provisioning and cleanup
require no host privilege elevation; only the positive probe mounts the
temporary native binary read-only. The verifier is unchanged by the
final master merge. All six isolation tests and an offline npm smoke
test pass. The prior head had 56 green CI checks and a 5/5 review after
two unchanged tests timed out and passed a failed-job-only retry ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)).
All 56 recovery-display/lineage tests pass; re-review cleared the
already-covered missed-retry concern. Earlier EC2 failures remain
retained: [npm lockfile
rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203),
[missing compiler in the slim
image](https://github.com/paperclipai/paperclip/actions/runs/36440210984),
and the aggregate 15-minute test timeouts in those broad runs. Both
broad attempts passed typecheck, token gates, Product E2E type/unit
checks and build. The focused EC2 lane preserves the existing
trusted-actor and immutable-source gates.


Earlier documentation/test checkpoint
`ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior
unchanged. 154 focused tests pass across configuration building, native
provider resolution, permission policy, credentials, UI configuration,
and new-agent setup (including both Grok auth modes); token gates pass.
All fresh CI is green for this head: 56 successful checks/statuses and
two intentional skips ([run
36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)).
Greptile is 5/5 with no new findings. Grok already inherits the shared
`approve-all` default, so unattended setup requires no manual permission
change.

Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final
staging continuation failure before provider startup: explicit
replacement archived the old harness but left its failover backups
active, which caused `runner_harness_state_mismatch`. The regression
fails before the fix and passes after it; all eight adjacent
recovery-safety cases also pass. Old backups remain inspectable inside
the continuity archive. All fresh CI is green at this head ([run
36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)),
with a 5/5 review. One unrelated Cursor test timed out in the initial
server shard; the same-commit failed-job rerun passed, and both attempts
are retained. Staging deployment is confirmed healthy on this revision.
The controller image is
`ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`.
The final browser-created staging task passed on this exact revision
with API authentication: context read → structured human question →
controller restart → answer submission → same native provider session
resumed → document saved → task Done. The two turns took approximately
119s and 77s. The actual write receipt was applied, and the saved
document has exactly one revision containing the selected answer and
requested marker. Usage and cost were not reported. [Controller image
build](https://github.com/paperclipai/paperclip/actions/runs/36360889243).

- Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`:
all CI green (53 successful checks/statuses, two intentional skips),
including repository typecheck/build/tests, native Runner tests, browser
shards, and canary installation checks. [CI run
36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529).
Greptile is 5/5 with no unresolved findings.
- Focused checks cover Grok credentials, executable admission, launcher
assets, provider-pack paths/permissions, workflow contracts, setup
defaults, CLI authorization, and subscription conflict recovery. All 39
protocol definitions validate. Final integration checks pass 124
catalog/evidence/cache tests and nine project-form tests; token gates
pass. Some local dependency checks could not load the stale installed
dependency tree; the corresponding fresh EC2 checks pass.
- Clean public npm installation passed on EC2 at
`8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run
36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)):
17 unified-version packages, lifecycle scripts enabled, built-in
launcher present, no separate Grok package or npm-downloaded binary,
missing prerequisite rejected, separately provisioned native executable
and command lease verified. No credentials or inference were used.
Subsequent changes preserve this npm asset layout.
- The immutable Daytona prerequisite image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`,
built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous
Cloud controller image was
`ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`,
built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded
by the latest image above. Its EC2 build verified provider-pack access
under an unrelated unprivileged UID.
- Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed
full Grok onboarding with the correct `grok-4.7` model, saved credential
delivery, and pinned Daytona execution. A browser-created task read
context and asked the structured human question. After a controller
restart, answering the persisted question resumed the same native
provider session, saved the requested document, and completed the task.
Actual tool outcomes and durable state agree: one question and one
document revision. The two successful turns took 42.7s and 63.1s; usage
and cost were not reported.
- Restricted policy returned the expected `approval_required` outcome.
Functional staging tests explicitly selected `approve-all`; controller
authorization and governed approvals remain enforced. Temporary board
CLI access was revoked and verified rejected (HTTP 401), and the
disposable onboarding agent was paused.

Failures remain retained: the pre-fix continuation failure (its task
remains blocked; the passing final task is fresh), the original Cloud
provider-pack permission failure, the expected restricted-policy denial,
the superseded npm staging failure, and an earlier monolithic CI
infrastructure timeout. Browser CI exposed a project alias/form race;
the final stack uses master's stronger draft-preservation fix and all
browser shards pass. Historical full subscription/API protocol and
Product rosters retain their original source revisions and do not
qualify this packaging revision. No local Docker or Rust build was used.

## Risks

The branch includes master’s draft-preservation fix for project URL
aliases. It keeps the same project’s edit form mounted and clears prior
data when the project or company changes.

Custom sandboxes and local execution hosts must provision the pinned
binary before Grok starts. Missing, changed, unsupported-platform, and
symlinked executables fail admission. The new builtin profile cannot
resume sessions created with the former private-package profile.
Existing Claude/Codex npm bridge profiles retain their package pins.
Grok restricted modes preserve the selected policy but cannot
automatically admit Paperclip calls: ACP permission metadata does not
independently bind tool authority, so those calls stop with
`approval_required`. New Grok configurations default to `approve-all`,
including API configurations that omit the mode. Existing explicitly
restricted configurations remain restricted; controller authorization
and governed approvals remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:54:44 -05:00
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00
DottaandPaperclip 60c7c9cd1a fix(runner-e2e): pass verified lock digest to Daytona image build (#13876)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Product E2E campaigns test the native runner in local and Daytona
environments.
> - Each campaign resolves one target lockfile and verifies its
downloaded artifact.
> - The Daytona image job did not pass that artifact digest to Docker.
> - Docker used an older default digest and stopped before any selected
task ran.
> - This pull request passes and validates the campaign digest at the
image build boundary.
> - The image keeps its checksum check and frozen package installation.

## Linked Issues or Issue Description

**What happened?**

The merged-master [qualification
campaign](https://github.com/paperclipai/paperclip/actions/runs/35863582409)
stopped in the Daytona image build. The resolved target lock digest was
`e0c928a494f90ddad3c00791e83f09315ee8c82df2a0418a809dcf93649a8ab3`.
Docker used its default digest,
`57b298aceebc48bb94ea0593347348256475da7b2fddb77025d9e57cc8759420`. The
checksum check rejected the mismatch. All 12 selected model cells were
skipped. This follows the image provenance work in #13814.

**Expected behavior**

The image build must check the same lockfile artifact that the campaign
restored and verified. A changed lockfile must still fail the checksum
check.

**Steps to reproduce**

1. Start a Product E2E campaign with a Daytona cell on master
`7944ed3d976d1a7cc26a2d0cee51f227f3542084`.
2. Resolve a target lockfile whose digest differs from the Dockerfile
default.
3. Observe the provider-pack image stage reject the lockfile before
model execution.

**Paperclip version or commit**

`7944ed3d976d1a7cc26a2d0cee51f227f3542084`.

**Deployment mode**

GitHub Actions Product E2E campaign with a Daytona image build.

**Install method**

Built from source with the campaign lockfile artifact.

**Agent adapter(s) involved**

Native Codex and ACPX Claude cells were selected. No model cell ran in
this failed campaign.

**Database mode**

Not involved. The failure occurs during image creation.

**Access context**

The authorized default-branch paid workflow. The build receives no
provider credentials.

## What Changed

- Read the image checksum from the existing target-lock job output.
- Require a 64-character lowercase hexadecimal digest before image
inspection or build.
- Pass the digest as the existing Docker build argument.
- Add regression checks and document the campaign checksum handoff.

## Verification

- The Daytona image regression fails with the original workflow and
passes with the fix.
- All six Daytona image contract tests pass.
- All 450 Product E2E unit tests pass.
- Product E2E typecheck passes.
- Actionlint passes for the changed workflow.
- A context-shaped resolution probe preserves the downloaded lockfile
bytes and digest.
- All latest-head CI checks passed on
`67d41fd9439b2a9a809ddb05765f8617585072c5`
([run](https://github.com/paperclipai/paperclip/actions/runs/35865739359)).
- Greptile gave 5/5 on this head; its test-scoping comment is addressed
and resolved.
- A hosted Daytona image rebuild and the three remote qualification
cells remain pending after merge.

## Risks

The campaign digest comes from the existing trusted target-lock job. The
restored artifact checks, Docker checksum check, frozen install, content
identity, image signing, and verification remain in place. The
standalone Docker default remains available. This change does not alter
task behavior, prompts, credentials, or dependency versions.

## Model Used

OpenAI `gpt-6-astra` through Codex performed diagnosis and review with
code execution tools. OpenAI `gpt-5.6-luna` assisted with investigation,
implementation, and verification. Context window limits are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 08:45:52 -05:00
Devin FoleyandPaperclip be6f49a425 feat(runner): refresh shared coding harness runtimes (#13838)
## Thinking Path

> - Paperclip runs agents through local adapters and the native runner.
> - Both paths must use the same installed provider CLI.
> - New models require current harness releases.
> - The runner still pins Codex 0.153.4, Claude SDK 0.3.263, and
OpenCode 1.18.29.
> - Changing the image alone would fail the runner's exact version and
executable checks.
> - This pull request updates those dependencies, integrity checks,
controller checks, and image pins together.
> - Shared installations can then run the current models without a
task-time download.

## Linked Issues or Issue Description

Refs #13829, which updates model choices and reasoning controls.
Searches found no open PR that updates these runtime pins.

**Current behavior**

The shared provider pack ships old CLIs. Claude Code 2.1.263 cannot run
Opus 5.5, which requires 2.1.280. Remote controllers reject provider
packs whose versions differ from their declared pins.

**Proposed behavior**

Use Codex 0.156.0, Claude Agent SDK 0.3.280 / Claude Code 2.1.280, and
OpenCode 1.18.32 throughout the runner. Keep the reviewed ACP bridge
patches and one shared CLI installation per provider.

**Reason and benefit**

Current harnesses support the new model IDs while preserving executable
verification and remote provider-pack compatibility checks.

## What Changed

- Update dependency overrides, the Codex ACP package patch, runtime
profiles, and remote controller pins.
- Verify the new Claude Linux x64 and macOS arm64/x64 executables and
Codex Linux x64 executable against integrity-verified npm archives.
- Refresh OpenCode version checks, fixtures, and the runner
configuration label.
- Refresh the eval image's Grok, Gemini, Kimi, Cursor, and GitHub CLI
pins and archive hashes. Hermes remains current at 0.19.0.
- Refresh the build-time lock digest from clean pnpm 9.15.4 resolution.
Leave lockfile commits to repository automation.
- Document model compatibility and the separation between CLI runtimes
and patched ACP bridges.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- Rust workspace release tests passed.
- Package/patch and OpenCode binary-materialization contract tests: 11
passed.
- Real Codex 0.156.0 startup-ownership and paginated session-resume
probes passed with isolated synthetic homes and no model turn.
- Codex app-server `thread/start` preserved `gpt-6-sol` and
`gpt-6-luna`; no `turn/start` was sent. An unauthenticated built-in
catalog does not include those account-served entries.
- Installed Claude integrity probes passed for `claude-opus-5-5` and
`claude-fable-5-1`.
- `pnpm --filter @paperclipai/paperclip-runner
test:opencode:qualification` passed with the actual OpenCode 1.18.32
executable under Node 24 and Node 25. The loopback provider exercise
covers health/version, session creation/read/delete, SSE, and a
completed async prompt.
- `pnpm check:token-gates` passed.
- The targeted runner suite passed 130 tests. Three macOS failures in
snapshot module lookup and OpenCode final-message selection also
reproduce on the unchanged base; Linux CI will provide the platform
check.
- [Final Linux
CI](https://github.com/paperclipai/paperclip/actions/runs/35798076399):
all gates passed. Four jobs needed one retry after their CI workers
received shutdown signals. The PR has 55 successful checks, two skipped
checks, Greptile 5/5, and no unresolved review threads.
- Changed runner configuration UI tests: 5 passed.
- Full macOS `pnpm test:run` reached 13,094 passing server tests, 84
skipped, and 18 failures before the wrapper stopped. Failures involved
skill-cache publication permissions, missing bundled connector skills in
the worktree, and a conversation-reset timing case. The 10 cache
permission failures reproduce on the unchanged base; both
conversation-reset cases passed on a targeted retry. The wrapper did not
reach its later workspace/serialized groups locally; Linux CI covers
those groups.
- The local Docker daemon did not respond, so no local Docker build was
run. No billable model requests were made.

## Risks

- Deploy the matching controller and provider pack together. Older
controllers enforce their previous exact pins.
- Current upstream CLIs can change behavior. Existing protocol tests and
isolated real Codex probes cover the integration boundaries;
authenticated model inference is not part of these checks.
- ACP bridge package versions and executable digests stay unchanged
because their executable bytes are unchanged. Only the underlying
CLI/SDK dependencies move.
- No schema migration. Revert the runtime and image pins together to
roll back.

## Model Used

OpenAI GPT-6 via Codex, with repository tools, code execution, and web
research. The exact serving model ID and context window were not exposed
by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass for the changed surfaces
and real-executable probes; full macOS-suite limitations are listed
above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 17:02:29 -07:00
DottaandPaperclip f6a211479f fix: share current CLI runtimes across sandbox adapters (#12994)
## Thinking Path

- Paperclip Runner needs its runtime preinstalled for fast sandbox
startup.
- Native and local adapters should launch one current CLI installation
per provider.
- An older global copy can shadow that installation, and exact native
compatibility pins must match it.
- Update the qualified releases and binary digests, expose shared CLI
entrypoints from the provider pack, and prefer the image-owned bin
directory.
- Keep dependency installation in the image build; task startup only
discovers, links, and verifies artifacts.

## Linked Issues or Issue Description

**What happened?**
Remote native startup rejected a stale global Codex, while CLI-only
images lacked runnerd entirely.

**Expected behavior**
An image-baked runtime starts without uploading binaries or installing
packages. All adapters share the same current provider CLI.

**Steps to reproduce**
Start a native remote task with the old global Codex and the updated
runtime available only under `/opt/paperclip-runner/bin`.

**Paperclip version or commit**
Discovery behavior at `54a99d884`.

**Deployment mode**
Docker with a remote sandbox.

## What Changed

- Prefer `/opt/paperclip-runner/bin`, then the user's local bin
directory, then PATH. Existing metadata and version validation remains
in force.
- Qualify Codex 0.153.4, OpenCode 1.18.29, and Claude SDK 0.3.263 / CLI
2.1.263. Update binary digests, TypeScript/Rust checks, registry
defaults, and the displayed OpenCode version together.
- Share Codex and Claude's native executable with the ACP bridges
through exact dependency overrides. Preserve the separately qualified
ACP bridge implementations and their security patches.
- Expose shared provider-pack CLI launchers; fail the pack build if
Codex ACP resolves a separate Codex installation. Update the eval
image's other agent CLIs to current stable releases and remove duplicate
global provider installs.
- Document the single-current-CLI policy in source comments and
development guidance. Latest stable releases are resolved at
review/build preparation and pinned; task startup never auto-updates.

## Verification

- Native-session and adapter-registry suites: 158 tests passed.
- Provider suites: 88 tests passed, 7 Linux-only checks skipped on
macOS. One existing macOS temporary-path alias assertion passed when
rerun with canonical `TMPDIR=/private/tmp`.
- Package-contract and OpenCode materialization tests: 11 passed.
- Full typecheck, build, and token gates passed. Rust
native-provider/recovery tests: 19 passed.
- Broad local suite: 5,974 passed, 23 failed, 41 skipped. Failures are
in unchanged macOS workspace/path/port and connection suites; focused
runtime tests pass. All latest-head Linux PR checks passed, including
the full test shards, typecheck, build, runner verification, browser
suites, and canary dry run.
- The standalone fleet image built with one current provider CLI each
and passed native Codex/Claude binary-integrity checks. A disposable
Daytona sandbox reported ready in 798 ms; its baked runner completed an
API-key `gpt-5.6-luna` turn in 2,430 ms and returned the expected marker
with a usage receipt. No runtime artifacts were uploaded or installed.
- The normal shared `codex exec` entrypoint also completed an API-key
`gpt-5.6-luna` turn in 2,321 ms.
- Both image builds verify the complete generated lockfile against a
reviewed SHA-256 before package installation or lifecycle execution.
Root lockfile changes remain CI-owned. Merge and rollout remain on hold
for operator review.

## Risks

- Updating provider CLIs changes their behavior for all adapters;
version probes and live native smoke testing are required before image
promotion.
- The image-owned directory takes precedence. Its entries must launch
the same shared CLI as the global PATH, not a private older/newer copy.
- Application qualification pins and the deployed image must move
together. No startup fallback installation is added.
- No schema or authentication-policy changes.

## Model Used

OpenAI GPT-6 (Codex). The session does not expose a more specific model
ID or context-window size. Used reasoning, repository inspection, code
execution, and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-07 10:09:29 -05:00
Dotta 5716fe907e test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner subsystem executes agent work across local and managed
provider backends.
> - The lower pull requests restore the task runtime, provider backends,
and managed-provider control plane.
> - The restored system needs repeatable full-stack checks before it can
ship safely.
> - Paid live checks also need clear access, cost, and secret controls.
> - This pull request adds acceptance, live evaluation, chaos, and
release gates for the restored runner stack.
> - The benefit is measurable runner parity with safer release
decisions.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting. This change covers runner tests, release workflows,
server contracts, and evaluation tools.

**Problem or motivation**

The runner stack did not have one complete acceptance surface for native
Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could
miss provider drift, task-view regressions, cost-policy errors, and
destructive cleanup errors.

**Proposed solution**

Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid
workflows. Add live evaluation, chaos, cost-limit, redaction, and
release contract checks. Add AWS AgentCore infrastructure and guarded
provisioning tools. Keep the native runner experimental flag off by
default.

**Alternatives considered**

We considered manual smoke tests only. They do not give repeatable
evidence and they do not protect release branches. We also considered
one large pull request. The stacked pull requests keep each review below
the Greptile file limit.

**Roadmap alignment**

This work supports the shipped Cloud / Sandbox agents milestone and the
shipped Agent evals & feedback milestone in `ROADMAP.md`.

Related stack:

- #12699 adds managed provider backends and lifecycle support.
- #12691 adds qualified OpenCode and ACPX provider backends.
- #12685 restores task runtime rendering and steering.

## What Changed

- Add the runner full-stack harness with 57 catalog cells and 60 unit
tests.
- Add a Daytona runner image with digest-pinned base images and
base-aware image-content checks.
- Add guarded live evaluation and chaos workflows with a fixed
40-execution matrix; live and full-stack paid schedules now run only on
Sundays or by manual dispatch.
- Add in-flight reported-usage cost stops, post-turn cost caps,
exact-threshold failure classification, secret redaction, retry
classification, and actor authorization.
- Reattach stream and hard-budget listeners before restart-recovery
continuations so restored paid sessions cannot bypass in-flight
interruption.
- Preserve OpenCode usage and cost across tool-loop messages and turns
while exposing an explicit current-run delta to durable accounting.
- Keep PNG/WebM evidence in access-controlled artifacts only, reject
SVG, and publish only pruned inert structured per-attempt evidence.
- Add AWS AgentCore infrastructure, provisioning checks, and smoke
tools; reject unsafe model identifiers, require exact stack ownership
markers, and make failed-stack replacement explicit.
- Add evaluation-session contracts and capability reports.
- Add release workflow checks for immutable action pins, frozen
dependency installs, exact weekly cron shape, paid-run guards,
provider-secret isolation, and chaos test paths.
- Reauthorize the original and triggering numeric actor IDs as the first
step of every provider-secret job, including partial reruns, before
checkout or provider access.
- Give each full-stack matrix cell only its matching provider
credential, expose Daytona only to Daytona cells, and disable shared
dependency caches anywhere paid credentials or OIDC write access are
present.
- Protect the legacy manual E2E workflow with the same default-branch,
allowlist, environment, and per-job authorization boundary.
- Rotate live-eval candidates by week and retain 120 days of compatible
history so the seven-week trend window remains viable.
- Restore the root runner-acceptance commands and reconcile reported
snapshots,
raw receipts, and terminal usage without double counting or losing late
usage.
- Mark ACPX token deltas exact only when every budget field is present,
keep
cumulative cost/request authority separate, reject non-USD cost
labeling,
  and include thought tokens in output-token budgets.
- Keep `enableNativeRunner` off by default. The acceptance harness
enables it only in its isolated test instance.

## Verification

Passed locally:

- `pnpm --filter @paperclipai/paperclip-runner typecheck`
- `pnpm test:runner-acceptance:typecheck`
- `pnpm test:runner-acceptance` (19 tests)
- focused OpenCode proxy, driver, runnerd transport, live-session, and
turn-stream tests (106 tests)
- `pnpm --filter @paperclipai/paperclip-runner exec vitest run
src/live/clean-room-server.test.ts` (22 tests)
- `pnpm test:e2e:runner:typecheck`
- `pnpm test:e2e:runner:unit` (62 tests)
- `node --test scripts/__tests__/release-verify-workflow.test.mjs`
- `pnpm --filter @paperclipai/paperclip-runner
test:runner-workflow-evals` (22 tests)
- `pnpm -r typecheck`
- `pnpm build`
- `node --test
packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs`
(6 tests)
- `git diff --check`
- `cargo test --manifest-path
packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core
--lib --locked` (161 tests)
- focused ACPX provider-event tests (10 tests)
- The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged.

I did not run paid live provider jobs or provision AWS resources. Those
checks need credentials and can create cost.

## Risks

The paid workflows can create provider cost. They require an allowlisted
original and triggering actor, the protected `runner-e2e-paid`
environment, explicit opt-in variables, and cost limits. The four
provider credentials exist only in that master-only environment, which
requires allowlisted reviewer approval and disables administrator
bypass; repository and organization Actions scopes contain no copies.

Provider usage arrives after a billable request, so the live guard
cannot prevent one request from crossing a threshold. It interrupts
immediately on the first reported threshold hit and permits no
continuation.

Visual evidence can contain secrets rendered as pixels. PNG/WebM remain
only in access-controlled workflow artifacts; SVG and per-attempt XML
are excluded, and S3/Pages receive a pruned structured dashboard.

The AWS scripts can create cloud resources. They use explicit commands,
least-privilege roles, KMS encryption, saved nonsecret metadata, and
explicit teardown.

This pull request does not enable the experimental native runner for
existing instances.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex with GPT-5. The model used extended reasoning, tool use,
code execution, and parallel subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-02 08:55:08 -05:00