mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
codex/plugin-task-execution
41
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
941a3fa991 |
Pin Copilot native dependencies with the maintained lockfile refresh (#15572)
Pin the three optional GitHub Copilot 1.0.88 native packages for Runner and server using the maintained lockfile workflow. Synchronize the package contract and bound initial render readiness in the deliberately throttled browser fixture. Current-head CI and focused checks pass. Co-Authored-By: Dotta <cryppadotta@users.noreply.github.com> Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a6306ba606 |
feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
854af7df19 |
build(deps-dev): bump vitest from 4.1.11 to 5.0.3 (#12969)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.11 to 5.0.3. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v5.0.3</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Isolate <code>result.status</code> between <code>repeats</code> runs - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11218">vitest-dev/vitest#11218</a> <a href="https://github.com/vitest-dev/vitest/commit/5dbebe9e3"><!-- raw HTML omitted -->(5dbeb)<!-- raw HTML omitted --></a></li> <li>Don't print an interceptor warning in browser mode - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11377">vitest-dev/vitest#11377</a> <a href="https://github.com/vitest-dev/vitest/commit/15cc006aa"><!-- raw HTML omitted -->(15cc0)<!-- raw HTML omitted --></a></li> <li>Don't retry when <code>test.fails</code> expectedly failed - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11219">vitest-dev/vitest#11219</a> <a href="https://github.com/vitest-dev/vitest/commit/b24585f08"><!-- raw HTML omitted -->(b2458)<!-- raw HTML omitted --></a></li> <li>Scope cache key generators to projects - by <a href="https://github.com/ecoyoung"><code>@ecoyoung</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11281">vitest-dev/vitest#11281</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11301">vitest-dev/vitest#11301</a> <a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1d"><!-- raw HTML omitted -->(92ba7)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Delay server <code>listen</code> until tests start running - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11366">vitest-dev/vitest#11366</a> <a href="https://github.com/vitest-dev/vitest/commit/7d8ed3e9b"><!-- raw HTML omitted -->(7d8ed)<!-- raw HTML omitted --></a></li> <li>Check mock path boundaries - by <a href="https://github.com/saryn17"><code>@saryn17</code></a>, <strong>Ryosei Sato</strong> and <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11361">vitest-dev/vitest#11361</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11362">vitest-dev/vitest#11362</a> <a href="https://github.com/vitest-dev/vitest/commit/1c3888bce"><!-- raw HTML omitted -->(1c388)<!-- raw HTML omitted --></a></li> <li>Keep config of browser-consumed environments - by <a href="https://github.com/kasperpeulen"><code>@kasperpeulen</code></a> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11378">vitest-dev/vitest#11378</a> <a href="https://github.com/vitest-dev/vitest/commit/aafc0996f"><!-- raw HTML omitted -->(aafc0)<!-- raw HTML omitted --></a></li> <li>Ignore page crash while cancelling - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11386">vitest-dev/vitest#11386</a> <a href="https://github.com/vitest-dev/vitest/commit/7c36748fa"><!-- raw HTML omitted -->(7c367)<!-- raw HTML omitted --></a></li> <li><code>toMatchScreenshot</code> uses wrong reference on retried tests - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11393">vitest-dev/vitest#11393</a> <a href="https://github.com/vitest-dev/vitest/commit/c22aba992"><!-- raw HTML omitted -->(c22ab)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>cache</strong>: <ul> <li>Revalidate imports of cached modules - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11381">vitest-dev/vitest#11381</a> <a href="https://github.com/vitest-dev/vitest/commit/38f98855f"><!-- raw HTML omitted -->(38f98)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>deps</strong>: <ul> <li>Pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into <code>ERR_PNPM_TRUST_DOWNGRADE</code> - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11403">vitest-dev/vitest#11403</a> <a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977"><!-- raw HTML omitted -->(f6c9a)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Pass current equality testers to <code>expect.extend</code> asymmetric matchers - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11401">vitest-dev/vitest#11401</a> <a href="https://github.com/vitest-dev/vitest/commit/3e794a96b"><!-- raw HTML omitted -->(3e794)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Support Blob on jsdom 30.1 - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11379">vitest-dev/vitest#11379</a> <a href="https://github.com/vitest-dev/vitest/commit/6c49b7197"><!-- raw HTML omitted -->(6c49b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>pool</strong>: <ul> <li>Preserve unique pool ids when <code>groupOrder</code> is set - by <a href="https://github.com/mtorp"><code>@mtorp</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11392">vitest-dev/vitest#11392</a> <a href="https://github.com/vitest-dev/vitest/commit/50312ebb4"><!-- raw HTML omitted -->(50312)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>ui</strong>: <ul> <li>Split-pane handle overlapping iframe - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11221">vitest-dev/vitest#11221</a> <a href="https://github.com/vitest-dev/vitest/commit/f91db0dfd"><!-- raw HTML omitted -->(f91db)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vitest</strong>: <ul> <li>Remove root temp dir on close - by <a href="https://github.com/abhinav-phi"><code>@abhinav-phi</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11248">vitest-dev/vitest#11248</a> <a href="https://github.com/vitest-dev/vitest/commit/7c7119cf7"><!-- raw HTML omitted -->(7c711)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vm</strong>: <ul> <li>Do not optimize deps from index.html - by <a href="https://github.com/ezefernandezyf"><code>@ezefernandezyf</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11329">vitest-dev/vitest#11329</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11360">vitest-dev/vitest#11360</a> <a href="https://github.com/vitest-dev/vitest/commit/caf2887de"><!-- raw HTML omitted -->(caf28)<!-- raw HTML omitted --></a></li> <li>Don't reuse scripts across vite environments - by <a href="https://github.com/MO2k4"><code>@MO2k4</code></a>, <strong>Martin Oehlert</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11395">vitest-dev/vitest#11395</a> <a href="https://github.com/vitest-dev/vitest/commit/346d3896b"><!-- raw HTML omitted -->(346d3)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v5.0.2...v5.0.3">View changes on GitHub</a></h5> <h2>v5.0.2</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Bind <code>process</code> in case global is overwritten - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11343">vitest-dev/vitest#11343</a> <a href="https://github.com/vitest-dev/vitest/commit/0b79231ad"><!-- raw HTML omitted -->(0b792)<!-- raw HTML omitted --></a></li> <li><strong>detect-async-leaks</strong>: <ul> <li>Ignore <code>process.stdio</code> handles - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11333">vitest-dev/vitest#11333</a> <a href="https://github.com/vitest-dev/vitest/commit/0fd6b9790"><!-- raw HTML omitted -->(0fd6b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Fix <code>toMatchObject</code> with asymmetric matchers - by <a href="https://github.com/ShreeBohara"><code>@ShreeBohara</code></a>, <strong>Claude Opus 5</strong>, <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-5)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11100">vitest-dev/vitest#11100</a> <a href="https://github.com/vitest-dev/vitest/commit/42523289e"><!-- raw HTML omitted -->(42523)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Fix <code>Request</code> with <code>Blob</code> body on jsdom 28+ - by <a href="https://github.com/harshit-d3v"><code>@harshit-d3v</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11295">vitest-dev/vitest#11295</a> <a href="https://github.com/vitest-dev/vitest/commit/d1c3ecc93"><!-- raw HTML omitted -->(d1c3e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporter</strong>: <ul> <li><code>agent</code> to respect <code>--silent</code> - by <a href="https://github.com/Raj4478"><code>@Raj4478</code></a> and <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11271">vitest-dev/vitest#11271</a> <a href="https://github.com/vitest-dev/vitest/commit/5b95efb6d"><!-- raw HTML omitted -->(5b95e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporters</strong>: <ul> <li>Handle concurrent <code>createReport</code> calls - by <a href="https://github.com/7rulnik"><code>@7rulnik</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11278">vitest-dev/vitest#11278</a> <a href="https://github.com/vitest-dev/vitest/commit/e8e556ff7"><!-- raw HTML omitted -->(e8e55)<!-- raw HTML omitted --></a></li> <li><code>hanging-process</code> to use ESM entrypoint - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11316">vitest-dev/vitest#11316</a> <a href="https://github.com/vitest-dev/vitest/commit/4e91e5668"><!-- raw HTML omitted -->(4e91e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>spy</strong>: <ul> <li>Fix stack overflow when spying <code>Set.prototype.add</code> - by <a href="https://github.com/fengmk2"><code>@fengmk2</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11299">vitest-dev/vitest#11299</a> <a href="https://github.com/vitest-dev/vitest/commit/a0a939653"><!-- raw HTML omitted -->(a0a93)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/33cadea62e8763c455c7fca38d9ab1dda87c5f75"><code>33cadea</code></a> chore: release v5.0.3 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11409">#11409</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/346d3896b65c3c907174447a035807342799f346"><code>346d389</code></a> fix(vm): don't reuse scripts across vite environments (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11395">#11395</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977ad3363f796a737834572e54c6ad5c18"><code>f6c9a49</code></a> fix(deps): pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into `...</li> <li><a href="https://github.com/vitest-dev/vitest/commit/062c75d8b63519211d951d8293ea81b5a9e3c124"><code>062c75d</code></a> chore: fix standalone docs build, update exports maps (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11394">#11394</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/caf2887dee8987a60118d53933f6e9cabd6b3e2a"><code>caf2887</code></a> fix(vm): do not optimize deps from index.html (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11329">#11329</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11360">#11360</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/50312ebb4eca6a98f6d0b2b61d5d9d38cbbabcef"><code>50312eb</code></a> fix(pool): preserve unique pool ids when <code>groupOrder</code> is set (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11392">#11392</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/7c36748fad1eae9687312f2f7ceadce6ec88b5df"><code>7c36748</code></a> fix(browser): ignore page crash while cancelling (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11386">#11386</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1df16a4fa5bbee3f198c582fbd56689d8"><code>92ba7fc</code></a> fix: scope cache key generators to projects (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11281">#11281</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11301">#11301</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/38f98855fa9cd7fd376afb84094eba0fda256a74"><code>38f9885</code></a> fix(cache): revalidate imports of cached modules (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11381">#11381</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/b24585f08f2ea267746a2d6ca0e43edcbb29726f"><code>b24585f</code></a> fix: don't retry when <code>test.fails</code> expectedly failed (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11219">#11219</a>)</li> <li>Additional commits viewable in <a href="https://github.com/vitest-dev/vitest/commits/v5.0.3/packages/vitest">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
6c36c07a4f |
feat(adapters): add GPT-6.1 Sol and refresh shared coding harness pins (#14942)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents run through coding-agent adapters and the native runner. Both use the same installed provider CLIs, model catalogs, and reasoning controls. > - OpenAI released GPT-6.1 Sol (`gpt-6.1-sol`) in Codex. Anthropic released Claude Sonnet 5.5. The static Codex, Bedrock, and OpenCode catalogs do not list these IDs. > - The shared provider pack pins Codex 0.156.0 and OpenCode 1.18.32. The evaluation image pins older Grok, Gemini, Kimi, Cursor, and GitHub CLI releases. Codex 0.156.0 has no bundled metadata for GPT-6.1 Sol. > - A model entry without a current harness, or a harness pin without its runner integrity checks, fails at run time. > - This pull request adds the verified model IDs and moves the harness pins, executable digests, controller checks, and image pins together. > - The benefit is that operators can select the current models, and the native and local adapters share one current CLI installation. ## Linked Issues or Issue Description Refs #13829 and #13838 (the September 22, 2026 model and harness refresh). Related pull requests: #14993 (merged October 5, 2026, superseding #14816) added the direct Claude Sonnet 5.5 entry and refreshed the Claude runtime to Agent SDK 0.3.286 / Claude Code 2.1.286. This pull request does not change the Claude runtime or the direct Claude model list; it keeps the #14993 pins and adds only the Bedrock Sonnet 5.5 ID. After #14993 merged, this branch was rebased onto `master` (October 5, 2026). The six overlapping pin regions (`docker/daytona-runner/Dockerfile`, `docker/daytona-runner/README.md`, `package.json`, `pnpm-workspace.yaml`, `packages/adapters/claude-local/src/index.test.ts`, `packages/paperclip-runner/src/backends/native-backend-factory.test.ts`) were resolved by keeping this pull request's Codex 0.160.0 and OpenCode 1.18.34 pins next to #14993's Claude 0.3.286 / 2.1.286 pins, taking the union of the Sonnet 5.5 model IDs in the Claude test, and merging both README paragraphs. The Sonnet 5.5 effort and CLI-gate lines in the Claude adapter were identical in both pull requests and merged without a diff. #14917 and #14918 reordered the Claude and Codex model lists earlier; the new entries sit where those ordering rules put them. Sources checked on 2026-10-02: - [OpenAI Codex models](https://learn.chatgpt.com/docs/models): GPT-6.1 Sol uses `gpt-6.1-sol`, supports reasoning efforts from Light to Ultra, and has Standard and Fast modes at launch. The page also records that `gpt-5.4` and `gpt-5.4-mini` retired from Codex with ChatGPT sign-in on August 31, 2026, and that `gpt-5.5` retires on October 14, 2026. Neither retirement applies to the OpenAI API. - [Codex CLI releases](https://github.com/openai/codex/releases) 0.157.0 through 0.160.0. The bundled model metadata in the 0.160.0 Linux binary contains `gpt-6.1-sol`. - [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview): Bedrock ID `anthropic.claude-sonnet-5-5`, released September 28, 2026. - [OpenCode releases](https://github.com/anomalyco/opencode/releases) 1.18.33 and 1.18.34 (fixes only). The OpenCode model registry lists both added provider-qualified IDs. - npm `latest` tags for `@xai-official/grok` 1.0.46, `@google/gemini-cli` 0.62.0, and `@moonshot-ai/kimi-code` 2.1.1. [xAI](https://docs.x.ai/docs/models), [Google](https://ai.google.dev/gemini-api/docs/models), and [Kimi](https://www.kimi.com/code/docs/en/kimi-code/models.html) list no newer coding models. - Cursor CLI 2026.10.01-e373342 is the version the official installer resolves. The pinned digest is the SHA-256 of the versioned Linux x64 archive. - [GitHub CLI 2.102.0](https://github.com/cli/cli/releases/tag/v2.102.0) (security fixes). The pinned digest matches the release `checksums.txt`. ## What Changed - Codex adapter: add `gpt-6.1-sol` to the model list, the Fast mode list, and the Ultra effort set. It is the first entry: #14918 orders the list newest version first, and its description notes the ChatGPT app lists GPT-6.1 Sol first. Update the adapter documentation text. - Claude adapter: add `us.anthropic.claude-sonnet-5-5` (Bedrock Sonnet 5.5) to the Bedrock catalog in the newest-Sonnet slot after Opus 5.5; `us.anthropic.claude-sonnet-5` moves into the older-Sonnet group, matching what `sortClaudeModels` from #14917 produces at runtime. Any Sonnet 5.5 ID (direct or Bedrock-qualified) now gets the documented `xhigh` and `max` efforts and requires Claude Code 2.1.284 or later on the CLI lane (the Claude Code changelog entry for 2.1.284 adds `claude-sonnet-5-5`). These two lines are identical to the ones #14993 merged, so the branch carries no diff for them. - OpenCode adapter: add `openai/gpt-6.1-sol` and `anthropic/claude-sonnet-5-5` to the static fallback catalog. - Codex runtime pin 0.156.0 → 0.160.0 in the root and workspace overrides, the runner package, the Codex ACP package patch, the qualified ACPX profiles, the Linux x64 executable digest, the Rust provider backend and its tests, the provider-pack manifest pins, the remote controller pins, the sandbox npm install spec, and the opt-in qualification scripts. - Remote Codex compatibility window: upper bound 0.157.0 → 0.161.0. The minimum stays at 0.149.0. - OpenCode runtime pin 1.18.32 → 1.18.34 in the runner package, the materialization script, the server and Rust qualified versions, the eval and live-session labels, fixtures, and the configuration label. - Evaluation image (`docker/daytona-runner/Dockerfile`): Grok CLI 1.0.46, Gemini CLI 0.62.0, Kimi Code 2.1.1, Cursor CLI 2026.10.01-e373342 with its digest, GitHub CLI 2.102.0 with its digest, Codex and OpenCode version probes, and the refreshed lockfile digest. The Claude Code 2.1.286 probe comes from #14993 and is unchanged here. - `pnpm-lock.yaml` is not part of this pull request. The repository's pull request gate rejects lockfile edits, and the refresh bot regenerates the lockfile on master (the same flow #13838 used). The Dockerfile `PAPERCLIP_RUNNER_LOCK_SHA256` default is the digest of the lockfile that `pnpm install --resolution-only --ignore-scripts --no-frozen-lockfile` (the refresh workflow's command) produces for the combined pins on the rebased branch (`e1856797…`); that lockfile differs from master only in the `@openai/codex` 0.160.0 platform packages, the `@anthropic-ai/claude-agent-sdk` 0.3.286 override that #14993 introduced (the open refresh-bot pull request #14872 carries that part), `opencode-ai` 1.18.34 with its Linux x64 baseline, and the `codex-acp` patch hash. - Documentation: runner README, runner compatibility doc, environment variable example, and a new `doc/adapter-model-audit-2026-10-02.md` with sources and deferred items. - Tests: Codex adapter catalog, server adapter models, Codex compatibility window, native session executor pins, runner package contract, OpenCode materialization, and UI effort options. Unchanged on purpose: Claude Agent SDK 0.3.286 / Claude Code 2.1.286 (already on `master` from #14993), ACP bridges (`acpx` 0.13.1, `claude-agent-acp` 0.73.0, `codex-acp` 1.6.2; newer upstream releases need a separate qualification), the native Grok runtime 1.0.13, Pi 0.84.2 / 0.87.1 (the Pi 1.0 runner stack covers it), and Hermes 0.19.0 (current). `gpt-5.4` and `gpt-5.4-mini` stay in the picker because the OpenAI API still serves them. ## Verification Run on Linux x64 with Node 25.9.0 and pnpm 9.15.4 after `pnpm install --no-frozen-lockfile` (the refreshed lockfile stays local; see above). The results below were re-run on the rebased head (October 5, 2026) for the suites the conflict resolution touches; the other rows are from the original run and are covered by CI on every push: - Rebased head: `packages/adapters/codex-local` 482 passed; `packages/adapters/claude-local` 340 passed, 4 failed (`execute.remote`, `test.probe`, `execute.acp-fallback`, `acp` spawn/env-hardening cases that fail identically on unchanged `master` in this host environment); `server` adapter-models + codex-runtime-compatibility + native-session-executor + adapter-registry 607 passed, 1 failed (the same adapter-registry override-pause case as before, also failing on `master` here); `packages/paperclip-runner` native-backend-factory + qualified-profiles 36 passed; `ui` codex-reasoning-effort + config-fields + model-utils 19 passed. Rust, full typecheck, build, and the Docker image are left to CI as before. - `vitest run` in `packages/adapters/codex-local`: 13 passed. `vitest run` in `packages/adapters/claude-local` (whole package, including the new Sonnet 5.5 gate and effort tests): see the latest CI run and the comment below. `vitest run` in `packages/adapters/opencode-local`: 48 passed, 1 failed (`runtime-config.test.ts` reads the host `PAPERCLIP_OPENCODE_PROVIDERS` variable; it fails the same way on the unchanged base). - `vitest run src/__tests__/adapter-models.test.ts src/services/native-runtime/codex-runtime-compatibility.test.ts src/__tests__/adapter-registry.test.ts` in `server`: 84 passed, 1 failed (`adapter-registry.test.ts` override pause test; it fails the same way on the unchanged base). - `vitest run` in `ui` for `codex-reasoning-effort`, `agent-setup-fields`, `config-fields`, and `ComposerRunSettingsPicker`: 25 passed. - `node --test test/acpx-codex-package-contract.test.mjs scripts/materialize-opencode-binary.test.mjs scripts/runner-protocol-eval-campaign.test.mjs` in `packages/paperclip-runner`: 23 passed. The package contract test verifies the installed Codex ACP executable digest and the 0.160.0 patch pin. - `vitest run src/drivers/acpx src/backends src/drivers/opencode src/live/live-session.test.ts` in `packages/paperclip-runner`: 626 passed, 5 failed, 1 skipped. The 5 failures (`installation-integrity.test.ts` `/proc/self/fd` module loading and one OpenCode answer-selection test) also fail on the unchanged base under Node 25; Linux CI runs Node 24. - `pnpm run test:opencode:qualification` in `packages/paperclip-runner` against the installed OpenCode 1.18.34 executable: passed. - `codex --version` from the installed pack prints `codex-cli 0.160.0`. The Linux x64 executable digest `12eb3e81…652aad` was computed from the `@openai/codex@0.160.0-linux-x64` archive after checking its registry `dist.integrity`. - `pnpm check:token-gates`: all gates clean. - `pnpm run typecheck:typescript` in `packages/paperclip-runner`: passed. Package typechecks ran one at a time; see the comment below for the server and UI results. Not run here, and needed from CI: - Rust tests and `pnpm -r typecheck` / `pnpm build` for the server (no `cargo` in this environment; the server typecheck prepares the runner vendor build). - The Docker evaluation image build and the real-binary Codex startup and session-resume probes (no Docker; the probes need the compiled `paperclip-runnerd`). The trusted CI runner workflow covers them. - Authenticated inference with any new model. This change is metadata and startup validation only. ## Risks - Codex 0.160.0 changes the bundled model catalog and app-server behaviour (authoritative provider catalogs, incremental running-turn tracking). The patched `codex-acp` 1.6.2 bridge is unchanged and declares `^0.148.0`; it worked with 0.156.0 under the same override. If CI probes show a protocol change, the pin can return to 0.156.0 by reverting this pull request. - The compatibility window upper bound moves to `<0.161.0`. Remote images with Codex 0.157 to 0.160 become accepted. Older images stay accepted down to 0.149.0. - Until the refresh bot lands the regenerated lockfile on master, the Dockerfile lockfile digest default does not match the committed lockfile. The trusted CI workflow computes the digest from its own resolution at build time, so this affects only a local build that passes no digest. - Existing saved model selections and effort settings are not changed. Agents on `gpt-5.4` or `gpt-5.5` with ChatGPT sign-in need a model change before the OpenAI retirement dates; that is documented, not enforced. - Rollout order: deploy the controller and runner from this change before promoting a sandbox image that carries these pins. Older controllers reject the new provider-pack pins. ## Model Used - Claude Fable 5.1 (Anthropic, model ID `claude-fable-5-1`), 1M context window, adaptive thinking, tool use. The model ran as a Paperclip agent through the Claude Code harness, performed the web research, edited the code, and ran the tests listed above. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Bender (Fable) <noreply@paperclip.ing> Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
cf8ad63c80 |
build(deps-dev): bump tsx from 4.23.12 to 4.23.15 (#12965)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.23.12 to 4.23.15. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/privatenumber/tsx/releases">tsx's releases</a>.</em></p> <blockquote> <h2>v4.23.15</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.14...v4.23.15">4.23.15</a> (2026-09-20)</h2> <h3>Bug Fixes</h3> <ul> <li>exclude bare builtins from namespace inheritance (<a href="https://github.com/privatenumber/tsx/commit/38e158857e50bca311be7c232a5057e5a2e5347a">38e1588</a>)</li> <li>expose require.cache and require.extensions to tsImport CommonJS modules (<a href="https://github.com/privatenumber/tsx/commit/2da34075afaed43e2b7fd0aca5fbebaaf337ff3a">2da3407</a>)</li> <li>make namespaced register() overloads portable for declaration emit (<a href="https://github.com/privatenumber/tsx/commit/562c434a5c8695e74327bbeb51cfeb9b86fc7e15">562c434</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.15"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.14</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.13...v4.23.14">4.23.14</a> (2026-09-20)</h2> <h3>Bug Fixes</h3> <ul> <li>restore the CJS bridge namespace for Node 24 require(esm) under tsImport() (<a href="https://redirect.github.com/privatenumber/tsx/issues/802">#802</a>) (<a href="https://github.com/privatenumber/tsx/commit/6e5236b065738d3687a06396d064774cfede390f">6e5236b</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.14"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.13</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.12...v4.23.13">4.23.13</a> (2026-08-30)</h2> <h3>Bug Fixes</h3> <ul> <li><strong>cache:</strong> bound shared transform cache memory (<a href="https://redirect.github.com/privatenumber/tsx/issues/835">#835</a>) (<a href="https://github.com/privatenumber/tsx/commit/28e1f12d04cd2afe1db17f8555b14fe5fb567c6e">28e1f12</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.13"><code>npm package (@latest dist-tag)</code></a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/privatenumber/tsx/commit/ca66105a17a2a4c6503fe3a12b5b9ec408286011"><code>ca66105</code></a> test: fix drive-less file URLs in ESM resolver fixtures</li> <li><a href="https://github.com/privatenumber/tsx/commit/2da34075afaed43e2b7fd0aca5fbebaaf337ff3a"><code>2da3407</code></a> fix: expose require.cache and require.extensions to tsImport CommonJS modules</li> <li><a href="https://github.com/privatenumber/tsx/commit/38e158857e50bca311be7c232a5057e5a2e5347a"><code>38e1588</code></a> fix: exclude bare builtins from namespace inheritance</li> <li><a href="https://github.com/privatenumber/tsx/commit/562c434a5c8695e74327bbeb51cfeb9b86fc7e15"><code>562c434</code></a> fix: make namespaced register() overloads portable for declaration emit</li> <li><a href="https://github.com/privatenumber/tsx/commit/edfb1f05a3f40b879a41a03a0801c2abd3a3ecf9"><code>edfb1f0</code></a> build: upgrade pkgroll and externalize CJS loader reference</li> <li><a href="https://github.com/privatenumber/tsx/commit/70e78284837c859f09b96cd10cd71d007aa4b795"><code>70e7828</code></a> test: upgrade tinyspy for disposable API</li> <li><a href="https://github.com/privatenumber/tsx/commit/9ed2022dfa9ea1be9511fe6abcde8110c25055a7"><code>9ed2022</code></a> ci: avoid duplicate release notifications</li> <li><a href="https://github.com/privatenumber/tsx/commit/872e77ffc5e96ca5c4727e74c0694debcb26219b"><code>872e77f</code></a> refactor: use disposables for cleanup</li> <li><a href="https://github.com/privatenumber/tsx/commit/6e5236b065738d3687a06396d064774cfede390f"><code>6e5236b</code></a> fix: restore the CJS bridge namespace for Node 24 require(esm) under tsImport...</li> <li><a href="https://github.com/privatenumber/tsx/commit/28e1f12d04cd2afe1db17f8555b14fe5fb567c6e"><code>28e1f12</code></a> fix(cache): bound shared transform cache memory (<a href="https://redirect.github.com/privatenumber/tsx/issues/835">#835</a>)</li> <li>See full diff in <a href="https://github.com/privatenumber/tsx/compare/v4.23.12...v4.23.15">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
18e8c121d9 |
fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path > - Paperclip manages agents through a shared native runner. > - Built-in harness support should ship with Paperclip's public distribution. > - Grok already speaks ACP; it does not require a new public bridge package. > - Sandbox provisioning owns the native executable and its pinned version. > - The runner must verify that prerequisite without downloading it during npm installation. > - This change separates built-in launcher identity from external runtime identity. > - Clean npm installation and live staging checks verify the distribution boundary. ## Linked Issues or Issue Description Refs #13882, #13973, #13977, #13979. This follow-up now targets master after #13882 was squash-merged. It replaces the private `@paperclipai/grok-acp` workspace package with runner-owned assets. Current master is included so the branch also contains the merged scheduler, complete-event capture, and durable cleanup fixes. ## What Changed - Ship Grok launcher and qualification metadata inside the runner's compiled output and the public server's vendored runner tree. - Remove the separate Grok npm package and all package-manager install hooks for this runtime. - Require the checksum-verified Grok Build 1.0.13 binary at `/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution environment. Provision it explicitly in the Daytona image and CI setup. - Keep native binaries outside the provider pack. Bind the built-in launcher into the pack manifest. - Preserve executable leases, descriptor-backed startup, credential fences, permissions, and exact ACP model admission. - Use `builtin:grok-acp` and `native:grok` as profile identities. Historical package-profile sessions fail closed on resume rather than being silently reinterpreted. - Resolve built-in assets from the authenticated sidecar location, including public server npm layouts. Keep the controller path out of provider environments. - Add clean npm tarball installation verification to the existing trusted canary CI job and the admitted manual EC2 verification path. It stages a unified release version and runs npm lifecycle scripts, then verifies missing-prerequisite rejection and admission after separate provisioning without credentials or inference. - Include the controller-owned provider pack in stamped Cloud images. Unstamped local images omit the pack and remain usable; remote ACPX requires full source provenance. - Correct CLI approval-page metadata for an already authenticated Cloud board user; approval authorization remains unchanged. - Honor explicit native-runner enablement in the Cloud agent picker and direct setup page, keeping the flag disabled by default. - Allow selecting the execution environment before connecting credentials. Include Grok in the existing authenticated hello-probe flow, targeting its pinned native prerequisite for runner setup. - Recover an existing subscription sign-in conflict through an explicit cancel-and-retry action, serialized after cancellation succeeds. - Preserve the selected ACPX harness before normalizing config fields, so new Grok agents use the Grok default model. - Keep the credential-free Cloud provider pack root-owned and readable after runtime UID remapping; verify manifest and referenced asset access under an unrelated unprivileged UID during image builds. - Archive prior failover backups alongside explicitly replaced harness state, preserving evidence while preventing stale backups from blocking a fresh replacement. - Update Daytona image content inputs and contract tests for the built-in assets and explicit provisioner. - Document and regression-test the shared `approve-all` default for Grok setup, saved configuration, and native execution. Explicitly saved restrictions remain unchanged. ## Verification Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05` incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the base PR was squash-merged. All 12 conflicts came from incoming files identical to the tested pre-squash base. The final tree exactly matches a three-way merge using that original base, preserving built-in Grok distribution and removal of the obsolete private package. All 252 focused runner/UI tests, six npm-isolation tests, and token gates pass. Fresh exact-head Greptile review is 5/5 with no outstanding findings; security scans and EC2 native compilation pass. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([run 36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)). The repository owner explicitly authorized bypassing code-owner approval after all checks passed; no CI checks or repository protection settings are bypassed or changed. The only remaining PR was removed from the completed stack metadata to permit native auto-merge. Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)). The initial attempt lost two EC2 runners to shutdown signals and stalled a third shard during dependency preparation; all three passed the same-commit failed-job-only retry. Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. [Final public npm verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764) passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public packages, an executed offline lifecycle sentinel, unchanged consumer lock, built-in launcher, missing-prerequisite rejection, and verified separately provisioned binary/command lease. Provisioning and cleanup require no host privilege elevation; only the positive probe mounts the temporary native binary read-only. The verifier is unchanged by the final master merge. All six isolation tests and an offline npm smoke test pass. The prior head had 56 green CI checks and a 5/5 review after two unchanged tests timed out and passed a failed-job-only retry ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)). All 56 recovery-display/lineage tests pass; re-review cleared the already-covered missed-retry concern. Earlier EC2 failures remain retained: [npm lockfile rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203), [missing compiler in the slim image](https://github.com/paperclipai/paperclip/actions/runs/36440210984), and the aggregate 15-minute test timeouts in those broad runs. Both broad attempts passed typecheck, token gates, Product E2E type/unit checks and build. The focused EC2 lane preserves the existing trusted-actor and immutable-source gates. Earlier documentation/test checkpoint `ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior unchanged. 154 focused tests pass across configuration building, native provider resolution, permission policy, credentials, UI configuration, and new-agent setup (including both Grok auth modes); token gates pass. All fresh CI is green for this head: 56 successful checks/statuses and two intentional skips ([run 36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)). Greptile is 5/5 with no new findings. Grok already inherits the shared `approve-all` default, so unattended setup requires no manual permission change. Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final staging continuation failure before provider startup: explicit replacement archived the old harness but left its failover backups active, which caused `runner_harness_state_mismatch`. The regression fails before the fix and passes after it; all eight adjacent recovery-safety cases also pass. Old backups remain inspectable inside the continuity archive. All fresh CI is green at this head ([run 36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)), with a 5/5 review. One unrelated Cursor test timed out in the initial server shard; the same-commit failed-job rerun passed, and both attempts are retained. Staging deployment is confirmed healthy on this revision. The controller image is `ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`. The final browser-created staging task passed on this exact revision with API authentication: context read → structured human question → controller restart → answer submission → same native provider session resumed → document saved → task Done. The two turns took approximately 119s and 77s. The actual write receipt was applied, and the saved document has exactly one revision containing the selected answer and requested marker. Usage and cost were not reported. [Controller image build](https://github.com/paperclipai/paperclip/actions/runs/36360889243). - Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`: all CI green (53 successful checks/statuses, two intentional skips), including repository typecheck/build/tests, native Runner tests, browser shards, and canary installation checks. [CI run 36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529). Greptile is 5/5 with no unresolved findings. - Focused checks cover Grok credentials, executable admission, launcher assets, provider-pack paths/permissions, workflow contracts, setup defaults, CLI authorization, and subscription conflict recovery. All 39 protocol definitions validate. Final integration checks pass 124 catalog/evidence/cache tests and nine project-form tests; token gates pass. Some local dependency checks could not load the stale installed dependency tree; the corresponding fresh EC2 checks pass. - Clean public npm installation passed on EC2 at `8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run 36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)): 17 unified-version packages, lifecycle scripts enabled, built-in launcher present, no separate Grok package or npm-downloaded binary, missing prerequisite rejected, separately provisioned native executable and command lease verified. No credentials or inference were used. Subsequent changes preserve this npm asset layout. - The immutable Daytona prerequisite image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`, built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous Cloud controller image was `ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`, built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded by the latest image above. Its EC2 build verified provider-pack access under an unrelated unprivileged UID. - Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed full Grok onboarding with the correct `grok-4.7` model, saved credential delivery, and pinned Daytona execution. A browser-created task read context and asked the structured human question. After a controller restart, answering the persisted question resumed the same native provider session, saved the requested document, and completed the task. Actual tool outcomes and durable state agree: one question and one document revision. The two successful turns took 42.7s and 63.1s; usage and cost were not reported. - Restricted policy returned the expected `approval_required` outcome. Functional staging tests explicitly selected `approve-all`; controller authorization and governed approvals remain enforced. Temporary board CLI access was revoked and verified rejected (HTTP 401), and the disposable onboarding agent was paused. Failures remain retained: the pre-fix continuation failure (its task remains blocked; the passing final task is fresh), the original Cloud provider-pack permission failure, the expected restricted-policy denial, the superseded npm staging failure, and an earlier monolithic CI infrastructure timeout. Browser CI exposed a project alias/form race; the final stack uses master's stronger draft-preservation fix and all browser shards pass. Historical full subscription/API protocol and Product rosters retain their original source revisions and do not qualify this packaging revision. No local Docker or Rust build was used. ## Risks The branch includes master’s draft-preservation fix for project URL aliases. It keeps the same project’s edit form mounted and clears prior data when the project or company changes. Custom sandboxes and local execution hosts must provision the pinned binary before Grok starts. Missing, changed, unsupported-platform, and symlinked executables fail admission. The new builtin profile cannot resume sessions created with the former private-package profile. Existing Claude/Codex npm bridge profiles retain their package pins. Grok restricted modes preserve the selected policy but cannot automatically admit Paperclip calls: ACP permission metadata does not independently bind tool authority, so those calls stop with `approval_required`. New Grok configurations default to `approve-all`, including API configurations that omit the mode. Existing explicitly restricted configurations remain restricted; controller authorization and governed approvals remain enforced. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1a394bd30 |
feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
be6f49a425 |
feat(runner): refresh shared coding harness runtimes (#13838)
## Thinking Path > - Paperclip runs agents through local adapters and the native runner. > - Both paths must use the same installed provider CLI. > - New models require current harness releases. > - The runner still pins Codex 0.153.4, Claude SDK 0.3.263, and OpenCode 1.18.29. > - Changing the image alone would fail the runner's exact version and executable checks. > - This pull request updates those dependencies, integrity checks, controller checks, and image pins together. > - Shared installations can then run the current models without a task-time download. ## Linked Issues or Issue Description Refs #13829, which updates model choices and reasoning controls. Searches found no open PR that updates these runtime pins. **Current behavior** The shared provider pack ships old CLIs. Claude Code 2.1.263 cannot run Opus 5.5, which requires 2.1.280. Remote controllers reject provider packs whose versions differ from their declared pins. **Proposed behavior** Use Codex 0.156.0, Claude Agent SDK 0.3.280 / Claude Code 2.1.280, and OpenCode 1.18.32 throughout the runner. Keep the reviewed ACP bridge patches and one shared CLI installation per provider. **Reason and benefit** Current harnesses support the new model IDs while preserving executable verification and remote provider-pack compatibility checks. ## What Changed - Update dependency overrides, the Codex ACP package patch, runtime profiles, and remote controller pins. - Verify the new Claude Linux x64 and macOS arm64/x64 executables and Codex Linux x64 executable against integrity-verified npm archives. - Refresh OpenCode version checks, fixtures, and the runner configuration label. - Refresh the eval image's Grok, Gemini, Kimi, Cursor, and GitHub CLI pins and archive hashes. Hermes remains current at 0.19.0. - Refresh the build-time lock digest from clean pnpm 9.15.4 resolution. Leave lockfile commits to repository automation. - Document model compatibility and the separation between CLI runtimes and patched ACP bridges. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - Rust workspace release tests passed. - Package/patch and OpenCode binary-materialization contract tests: 11 passed. - Real Codex 0.156.0 startup-ownership and paginated session-resume probes passed with isolated synthetic homes and no model turn. - Codex app-server `thread/start` preserved `gpt-6-sol` and `gpt-6-luna`; no `turn/start` was sent. An unauthenticated built-in catalog does not include those account-served entries. - Installed Claude integrity probes passed for `claude-opus-5-5` and `claude-fable-5-1`. - `pnpm --filter @paperclipai/paperclip-runner test:opencode:qualification` passed with the actual OpenCode 1.18.32 executable under Node 24 and Node 25. The loopback provider exercise covers health/version, session creation/read/delete, SSE, and a completed async prompt. - `pnpm check:token-gates` passed. - The targeted runner suite passed 130 tests. Three macOS failures in snapshot module lookup and OpenCode final-message selection also reproduce on the unchanged base; Linux CI will provide the platform check. - [Final Linux CI](https://github.com/paperclipai/paperclip/actions/runs/35798076399): all gates passed. Four jobs needed one retry after their CI workers received shutdown signals. The PR has 55 successful checks, two skipped checks, Greptile 5/5, and no unresolved review threads. - Changed runner configuration UI tests: 5 passed. - Full macOS `pnpm test:run` reached 13,094 passing server tests, 84 skipped, and 18 failures before the wrapper stopped. Failures involved skill-cache publication permissions, missing bundled connector skills in the worktree, and a conversation-reset timing case. The 10 cache permission failures reproduce on the unchanged base; both conversation-reset cases passed on a targeted retry. The wrapper did not reach its later workspace/serialized groups locally; Linux CI covers those groups. - The local Docker daemon did not respond, so no local Docker build was run. No billable model requests were made. ## Risks - Deploy the matching controller and provider pack together. Older controllers enforce their previous exact pins. - Current upstream CLIs can change behavior. Existing protocol tests and isolated real Codex probes cover the integration boundaries; authenticated model inference is not part of these checks. - ACP bridge package versions and executable digests stay unchanged because their executable bytes are unchanged. Only the underlying CLI/SDK dependencies move. - No schema migration. Revert the runtime and image pins together to roll back. ## Model Used OpenAI GPT-6 via Codex, with repository tools, code execution, and web research. The exact serving model ID and context window were not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass for the changed surfaces and real-executable probes; full macOS-suite limitations are listed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
352153b5ed |
fix: run the cargo-building native-runner CI suite in the Rust-cached vitest lane (#13586)
## Thinking Path Post-merge of #13557, the slowest check on the freshest fully-green PR run ([35246999382](https://github.com/paperclipai/paperclip/actions/runs/35246999382)) was `ci / General tests (server (1/12))` at **339s**. The cause is one suite: `server/src/services/native-runtime/native-codex-runner.integration.test.ts` runs 1 test in **277s of a 291s vitest step (95%)** because its `beforeAll` cargo-builds the Runner release binaries, and the general-server shards carry no Rust cache — every PR run cold-compiles the full third-party crate graph. The other 19 suites in that shard finish in under 70ms each. The obvious fix (a dedicated Rust-cached matrix lane) requires editing workflow files, which the available GitHub App credentials cannot push (`workflows` permission). But the `Verify Paperclip Runner` lanes **already restore the shared `release-runner-v1` Rust cache read-only**, and their commands are `pnpm --filter @paperclipai/paperclip-runner <package script>` — so the suite can move into a Rust-cached lane purely through script changes. ## What Changed - `scripts/run-vitest-stable.mjs`: new `general-server-native-runner` group carrying exactly that suite. Under the PR workflow (`GITHUB_WORKFLOW == "PR"`, inherited from `pr.yml` by the reusable `pr-trusted.yml`) the without-chat server shards exclude it and rebalance to ~211s of tests each. Every other caller — local runs, `release-verify.yml` under the Release / Cloud readiness workflows — keeps the suite in the shards, so a renamed or unknown workflow degrades to today's slower-but-covered behavior instead of dropping coverage. - `packages/paperclip-runner`: `test:typescript:vitest` now routes through `scripts/run-pr-vitest-lane.mjs` — the identical `ensure:eval-build-deps && build:rust && vitest run` chain (shard flags passed through), plus the native-runner group on the **final PR shard only** (`--shard=N/M` with `N == M`, i.e. today's `vitest 2/2`, the 122s lane). With the restored cache the suite's cargo build becomes an incremental rebuild. - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: guards pin the whole contract — PR 12-shard coverage (shards + chat + native-runner = full server group exactly), Release/local 10-shard runs keep the suite, `pr.yml` is named `PR`, the vitest lanes partition with exactly one final shard, the package-script wiring, and the wrapper's shard/workflow gating via its `--dry-run` plan output. No workflow files change. `.github/workflows/*` are untouched. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`: **36/36 pass** locally on this branch (includes the new coverage, wiring, and wrapper-gating guards). Both files run in CI's `Test general-server shard partition` / `Test release verify workflow wiring` steps. - Wrapper `--dry-run` plan matrix verified for all six shard/workflow combinations plus malformed-shard rejection (pinned as a guard test). - The executing proof is this PR's own CI: `ci / Verify Paperclip Runner (vitest 2/2)` must go green while running the native-runner suite (its log will show the `general-server-native-runner` group after the package vitest shard), and the 12 `ci / General tests (server (x/12))` shards must go green without it. ## Risks - The exclusion keys on `GITHUB_WORKFLOW == "PR"`. Failure mode of a rename is safe (suite falls back into the server shards, slower but covered) and the guard test on `pr.yml`'s name makes it loud. - `vitest 2/2` grows from ~122s to an expected ~210–260s — still well under the ~306s `vitest 1/2` and ~326s e2e shards, and inside the 20-minute lane timeout. If the cache misses (key drift), the lane pays a cold compile like the server shard does today; a miss is slow, never wrong. - Double-run/coverage-loss combinations are enumerated in the wrapper header and pinned by tests: each caller runs the suite exactly once. ## Model Used Claude (Bender agent, Paperclip) — Fable 5. --- Expected savings once merged: the 339s `server (1/12)` check drops to ~265s-equivalent shard levels (~211s of tests), the slowest `ci /` check becomes the ~326s e2e shard (~13–33s off PR wall time), and every PR run stops paying ~4.5 min of billed cold Rust compile. For the merger (squash): please keep the trailer below in the squash body to preserve authorship. `Co-Authored-By: Bender (Fable) <Paperclip-Paperclip@users.noreply.github.com>` ## Related PRs Searched the GitHub PR list for prior work on this surface — related groundwork, none duplicate this change: - #13457 — restored master's Rust dependency cache on the PR runner lane (the read-only cache this PR relies on) - #13500 — made that cache key image-toolchain-independent so GitHub-hosted PR runners actually hit it - #13521 — rebalanced PR shards and split the Verify Paperclip Runner lanes this PR extends - #13557 — previous health-check iteration (split the runnerd transport suite); this PR targets the next slowest check ## Checklist - [x] I have searched GitHub for duplicate or related PRs and linked them above Co-authored-by: Bender (Fable) <Paperclip-Paperclip@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d08abcba15 |
ci: cut PR wall clock from ~16 to ~6 minutes (#13521)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every pull request runs the Trusted PR CI workflow before merge > - The test suites roughly tripled in six weeks, and shard balance did not keep up, so PR runs crept from ~4 to ~17 minutes > - Slow CI delays every merge and every contributor > - This pull request rebalances the shards from fresh measurements, splits the largest test files, reuses the Rust build cache in three more jobs, and takes the policy job off the critical path > - The benefit is a PR wall clock near 6 minutes with the same coverage ## Linked Issues or Issue Description **What existing behavior does this improve?** PR CI wall clock. A typical green run took 16-17 minutes. Two months ago it took about 4 minutes. **Subsystem affected** The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the shard-duration manifests, the vitest shard runner scripts, the `paperclip-runner` package scripts, and the dry-run branch of `release.sh`. **Current behavior** The shard-duration manifests were stale. The general-server manifest had durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29 specs. Stale median weights made shard steps range 417s-806s (server) and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release build. Every test lane waited ~60s for the policy job before it could start. **Proposed behavior** All lanes finish in a narrow ~200-290s band. The manifests carry fresh measured durations for every suite. The three largest test files are split so no single file caps a shard. The Rust cache restore runs in every job that builds the Runner binary. Test lanes start as soon as the gate resolves. **Reason and benefit** Merges stop waiting on CI. The projected wall clock is ~6 minutes for the same test coverage. ## What Changed - Rebuild `scripts/general-server-shard-durations.json` (646 suites) and `scripts/e2e-shard-durations.json` (all specs) from per-suite completion timestamps in runs 35036001734 and 35024948947. - Move the PR server lane to the release-verify shape: `general-server-without-chat` across twelve duration-balanced shards, plus the chat integration suite split by collected test location across three dedicated lanes. - Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and `-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions` and `-projects` specs. Each pair shares fixtures through a `.shared.ts` module. Playwright collects the same test sets (39 and 20 tests). - Raise e2e shards to eight and serialized shards to nine. - Run the runner package's `check:all` as four matrix lanes: `check:static`, `check:runner`, and two native vitest `--shard` halves. The union is exactly `check:all`. - Add the read-only Rust cache restore (toolchain pin, `save-if: false`) to the Canary Dry Run, Build, and Typecheck jobs. - Make release.sh preview publish payloads concurrently in batches of eight during `--dry-run`. The real publish path stays strictly serial. - Drop the policy-job lockfile artifact chain. Each lane installs with `--frozen-lockfile` and falls back to an inline `--resolution-only` regeneration. The policy job stays a required check through the `verify` and `e2e` aggregates. - Update the shard-count mirrors and workflow assertions in the partition and gate tests. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/e2e-shard.test.mjs` — 30 pass. - `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass. - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass. - `playwright test --list` collects 39 tests across the chat-adapters split and 20 across the agent-chat split, equal to the original files. - A local vitest collection of the chat suite partitions 995 tests into 498/497 line shards. - Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s x8, serialized ~216s x9. ## Risks - The split spec files reorder tests relative to the original files. Every describe seeds its own company, so the specs stay independent; a hidden cross-describe dependency would surface as a deterministic failure in one shard. - The inline lockfile fallback changes install behavior for manifest-changing and stacked PRs. The policy job still validates resolution as a required check. - `release.sh` changes are confined to the `--dry-run` preview branch. The publish loop is untouched. `bash -n` passes and the release dry-run tests pass. - One PR now schedules ~44 fleet runners. If the RunsOn fleet caps concurrency, queueing may absorb part of the gain; watch the first runs. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with tool use (shell, file edits) in Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
663c44cb2b |
fix: continue conversations after confirmed remote runner stop (#13254)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can stop a run and send another message on the same task.
> - Remote runners need evidence from their sandbox provider that
execution stopped.
> - Local process checks cannot prove that a remote process exited.
> - This pull request records provider stop receipts and uses them for
conversation admission.
> - New user messages can proceed after confirmed cleanup without
repeating interrupted actions.
## Linked Issues or Issue Description
Refs #13237 and #13239. Related: #13163 covers app-restart recovery;
this change covers an explicit stop followed by a new user message.
**What happened?**
A stopped remote Claude ACP task kept its execution hold after Daytona
cleanup succeeded. Native runners also rejected remote process
identities and retained stale session cleanup gates. A message sent
during cleanup could stay deferred after the sandbox stopped.
**Expected behavior**
After the provider confirms that the old execution stopped, a new user
message starts a fresh turn. Pending user messages must not need another
message to trigger admission. Prior action outcomes remain recorded.
**Steps to reproduce**
1. Start a long-running task in Daytona with a legacy Claude ACP or
native ACP runner.
2. Cancel the run while its tool is active.
3. Send a new message immediately, or after cleanup completes.
4. Observe the execution hold despite the old sandbox having stopped.
**Paperclip version or commit**
Reproduced on master at
|
||
|
|
c1b55537ba |
fix(paperclip-runner): bump claude-agent-acp pin to 0.73.0 (#13162)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Claude local adapter can run agent turns through an ACP (Agent Client Protocol) server, `claude-agent-acp`, instead of the plain CLI > - Two separate packages each pin their own copy of that dependency: `packages/adapters/claude-local` (the server-side adapter) and `packages/paperclip-runner` (which builds the provider pack baked into every managed sandbox image) > - `claude-local` moved to `^0.73.0` in #12730, but `paperclip-runner` was never bumped past `0.70.0` — nothing keeps the two in sync when only one changes > - That split means a sandbox image built from `paperclip-runner`'s provider pack ships a `claude-agent-acp` the server-side adapter was never actually compatible with > - This pull request bumps `paperclip-runner`'s pin to `0.73.0`, the only version that satisfies both packages' declared ranges at once, and fixes the matching hardcoded version assertion in `docker/daytona-runner/Dockerfile` > - The benefit is one consistent, compatible `claude-agent-acp` version across both the server host and every sandbox image built from this source, instead of a silent split that only surfaces as a runtime failure ## Linked Issues or Issue Description No public issue exists for this specific split; opening directly per CONTRIBUTING.md path B, following the bug report template fields. **What happened?** `packages/paperclip-runner/package.json` pins `@agentclientprotocol/claude-agent-acp` at an exact `0.70.0`. `packages/adapters/claude-local/package.json` requires `^0.73.0` (added in #12730, 2026-09-02). Nobody re-synced `paperclip-runner`'s pin after that change — the two packages' dependency graphs are independent, so a bump in one doesn't propagate to the other. `paperclip-runner`'s copy is what the fleet sandbox image's provider pack actually ships, so every managed sandbox built from current source carries a `claude-agent-acp` version the server-side adapter's own declared compatibility range excludes. **Expected behavior** The two packages' `claude-agent-acp` pins should stay within a mutually compatible range, so a sandbox image built from this source always ships a version the server-side adapter actually supports. **Steps to reproduce** 1. Check `packages/adapters/claude-local/package.json`'s `@agentclientprotocol/claude-agent-acp` range (`^0.73.0`). 2. Check `packages/paperclip-runner/package.json`'s pin for the same package (`0.70.0` before this PR). 3. Note that `^0.73.0` on a `0.x` version only admits patch releases (`>=0.73.0 <0.74.0` per semver caret rules), so `0.70.0` falls outside it. **Paperclip version or commit** `master` as of this PR (paperclip-runner still at `0.70.0` prior to this change; claude-local's `^0.73.0` requirement landed in #12730). **Deployment mode** Any deployment that runs `claude_local` agents through the ACP engine against a sandbox image built from `packages/paperclip-runner`'s provider pack (managed cloud sandboxes in particular). Related PRs for context (not duplicates — none of these touch `paperclip-runner`'s pin): - #12730 — introduced the `^0.73.0` requirement in `claude-local` - #11873 — the last time `paperclip-runner`'s pin moved (`0.69.0` → `0.70.0`) - #13105 — separately made an unavailable ACP engine a hard failure instead of a silent CLI fallback, which is what turned this version split into a visible, run-blocking error rather than a quiet downgrade ## What Changed - Bump `@agentclientprotocol/claude-agent-acp` from `0.70.0` to `0.73.0` (exact pin, matching this package's existing pin style for its other agent-CLI dependencies) in `packages/paperclip-runner/package.json`. - Update the corresponding hardcoded version assertion (`test "$(claude-agent-acp --version)" = "0.70.0"`) in `docker/daytona-runner/Dockerfile` to `0.73.0`, so its own build-time check stays accurate instead of failing on the next build for an unrelated reason. - `pnpm-lock.yaml` is intentionally **not** included — `pr-trusted.yml`'s `Validate dependency resolution and regenerate stale lockfile` step already regenerates it for the merge tree and hands it to downstream `--frozen-lockfile` jobs as an artifact, so a manual lockfile commit here would just be stale the moment CI runs. ## Verification - `0.73.0` is a real published version on npm (confirmed via `npm view @agentclientprotocol/claude-agent-acp versions`), and it's the *only* version satisfying claude-local's `^0.73.0` range, so this isn't a guess at compatibility — it's the unique intersection of both packages' declared ranges. - `grep -rn "0\.70\.0" docker/ packages/paperclip-runner/package.json` after this change shows no remaining stale references to the old pin. - I did not run a full local install/test pass against a hand-updated lockfile, since regenerating one locally would conflict with leaving `pnpm-lock.yaml` untouched per the note above; CI's own lockfile-regeneration step is the intended verification path for a manifest-only dependency bump like this one. - Downstream/full verification (does a sandbox image actually built with this pin work end-to-end) is tracked separately in `paperclip-cloud` — an unrelated internal-only repo, so not linked here — where a sibling fix restores the ACP servers to the runtime `PATH` in the fleet sandbox image itself; both fixes are needed together for a working sandbox, but this PR is scoped to the version pin alone. ## Risks - Low risk: single-line dependency version bump plus a matching test-assertion update, no code changes. `0.73.0` is a patch release within claude-local's own already-declared-safe range, so there's no reason to expect it changes behavior tenants depend on. - The main risk is unknown breaking changes between `claude-agent-acp` 0.70.0 and 0.73.0 that aren't caught by the version-string assertion alone (that check only confirms the binary reports the right version, not that its behavior is unchanged). I have not audited that package's own changelog between those versions. - `docker/daytona-runner/Dockerfile` is a parallel/reference image (per its own header comment, meant to stay aligned with the private `paperclip-cloud/fleet-sandbox-image/Dockerfile`, which is out of scope here) — this PR does not touch that other Dockerfile. ## Model Used Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code, with tool use (file edits, shell/git, `gh` CLI, `npm view` for version verification). No extended-thinking mode. Standard Claude Code context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — see Verification: a manifest-only bump with the lockfile intentionally left to CI's own regeneration step; no local test run applicable - [x] I have added or updated tests where applicable — version-pin bump only, no new behavior to test - [x] I have updated relevant documentation to reflect my changes — none applicable - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — pending CI run on this PR - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — pending review - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3b550c80fa |
fix(codex): correct startup trust, history reads, and resume usage (#13110)
## Thinking Path
> - Paperclip runs Codex locally and in remote sandboxes.
> - The runner must preserve startup configuration and session identity.
> - Missing project trust can disable repository configuration.
> - Full-history requests use deprecated provider fields.
> - Resume usage describes old work and must not become new run usage.
> - This change corrects startup trust, state reads, and usage
classification.
## Linked Issues or Issue Description
**What happened?**
Normal Codex runs could show repository-trust and history-deprecation
warnings.
Resume could report the preceding turn's token snapshot as a late-turn
warning.
The historical last-usage value could also be attributed to the new run.
**Expected behavior**
Trust the server-selected startup root in isolated configuration. Read
lightweight
provider state and paginated evidence. Use historical cumulative usage
as a
baseline without a new charge or user-facing warning.
**Steps to reproduce**
1. Start a native Codex task in a selected repository.
2. Finish the turn and resume the provider thread.
3. Inspect provider notices, history requests, and per-run usage.
4. Repeat startup and cold resume inside a Daytona sandbox.
**Paperclip version or commit**
Codex CLI 0.153.4 is the pinned runtime and reproduced baseline.
Replayed onto master at
|
||
|
|
5bddff0920 |
feat(runner): add guarded API search and call fallback (#13003)
## Thinking Path > - Paperclip manages AI agents and their work. > - The new runner gives agents dedicated tools for common tasks. > - Some API operations and parameters have no dedicated tool. > - Agents need a controlled way to find and use those operations. > - This pull request adds API search and calls through the real server routes. > - Existing tools remain the preferred path. The new tools are disabled by default. > - Paired tests measure correctness, tool choice, cost and time. ## Linked Issues or Issue Description **Subsystem affected** Paperclip Runner contracts, production tool authority and the server API catalog. **Problem or motivation** The runner cannot use much of the API described by the old Paperclip skill. A generic HTTP client would also let agents bypass runner control rules. **Proposed solution** Add `search_api` and `call_api`. Resolve calls from the mounted API catalog. Use server-held, run-bound credentials. Preserve route checks and runner lifecycle rules. Keep the tools disabled until an operator enables selected companies. **Alternatives considered** A dedicated tool for every endpoint would add a large initial prompt. An unrestricted HTTP tool would weaken authorization and replay controls. **Roadmap alignment** This extends the native runner tooling. The repository owner requested this design and implementation. The roadmap and related open PRs were checked. No duplicate API escape-hatch PR was found. ## What Changed - Register two compact fallback tools in canonical contracts and provider projections. - Build deterministic API discovery from OpenAPI, mounted experimental routes and the old skill reference. - Execute bounded JSON, text, file and download requests through authenticated HTTP routes. - Recheck active runs, company access and work modes. Block runner lifecycle, scheduling, credential and approval bypasses. Keep routine annotation collaboration available. - Retain mutation receipts. Report uncertain outcomes without blindly repeating writes. - Add a company rollout gate and a durable eval worker with complete cost accounting checks. - Record child-task creation in the activity log with the agent and run. - Add contract, authorization, file, replay and real runnerd/PRP/HTTP tests. - Document rollout gates and paid coverage limits. The companion eval repository retains immutable attempts and reports. ## Verification - Final app commit `da58370524c3626a744eec20164397c5fb6ba9ef`: all 32 checks passed; the unrelated Storybook visual check was skipped. Greptile 5/5; no unresolved review threads. - Full Linux build and recursive typecheck passed. Repository tests were run by project and serialized shard; all 143 serialized server suites passed. - Runner TypeScript: 1,599 passed, two skipped. Rust release: 451 passing test reports. Conformance and replay parity passed. The required API check passed 837 tests, including runnerd → PRP → authority → real HTTP. - Bindings cannot enable API tools without the explicit deployment flag. Unit and real-authority tests prove the default-off boundary. - The standalone API check builds and stages its own binary. It passed after existing staged and debug binaries were removed from the test container. - UI and CLI tests passed. Initial environment failures (missing jq, Docker overlay file identity, and parallel linker memory pressure) and focused passing reruns are retained. The macOS full runner suite has platform-specific failures; Linux is the qualified full-check platform. - Eval harness: 27 tests passed; existing CI discovery ran 86 tests with two unrelated skips. Credential export rejection is tested against the actual report command. - Luna and OpenRouter Sonnet each passed 60 common-workflow runs: ten workflows, three repetitions per arm, zero unnecessary API fallback. - Sonnet passed 11 selected capability/contract cases after fixes. Gemini passed three smoke cases. DeepSeek exceeded the 120-second limit and remains unqualified. - Luna's two cost flags received focused follow-up. The original flags and a later n=1 latency flag remain visible. Sonnet had no cost or latency increase above 20%. - The catalog contains 785 entries; 58 were exercised across all stages. Most operation probes remain unrun and some need additional fixtures. Authored probes do not establish successful coverage. - Total conservative accounted cost: $9.875960. Active paid-campaign time: 88.16/90 minutes. No missing accounting. Later security and harness fixes have provider-free verification; no paid validation is claimed for those revisions. - Inspect the [qualification report](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/READINESS.md) and [verification record](https://github.com/paperclipai/paperclip-evals/blob/codex/seach-call-api-tools/evals/runner-api-tools/reports/2026-09-07-production/verification.json). ## Risks - This is a broad authenticated API surface. Keep the default-off gate until an operator selects initial rollout companies. - Paid coverage is incomplete. Small regression samples do not prove all workflows are unchanged. - A timeout or server failure can follow a committed mutation. The result reports an unknown outcome and requires state inspection. - The new definitions add prompt tokens. The report retains cost flags and cache variation. - No database migration is required. - Repository rules require code-owner approval before merge. Technical CI and automated review are complete. ## Model Used OpenAI Codex based on GPT-6 assisted with code, tests and review. The exact serving model ID and context window are not exposed in this session. It used reasoning, tool calls and code execution. Eval models: `gpt-5.6-luna` with low reasoning, `openrouter/anthropic/claude-sonnet-5`, `openrouter/google/gemini-3.8-flash`, and `openrouter/deepseek/deepseek-v4-flash-0731`. Attempts retain runtime versions, model identity, usage and source provenance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f6a211479f |
fix: share current CLI runtimes across sandbox adapters (#12994)
## Thinking Path - Paperclip Runner needs its runtime preinstalled for fast sandbox startup. - Native and local adapters should launch one current CLI installation per provider. - An older global copy can shadow that installation, and exact native compatibility pins must match it. - Update the qualified releases and binary digests, expose shared CLI entrypoints from the provider pack, and prefer the image-owned bin directory. - Keep dependency installation in the image build; task startup only discovers, links, and verifies artifacts. ## Linked Issues or Issue Description **What happened?** Remote native startup rejected a stale global Codex, while CLI-only images lacked runnerd entirely. **Expected behavior** An image-baked runtime starts without uploading binaries or installing packages. All adapters share the same current provider CLI. **Steps to reproduce** Start a native remote task with the old global Codex and the updated runtime available only under `/opt/paperclip-runner/bin`. **Paperclip version or commit** Discovery behavior at `54a99d884`. **Deployment mode** Docker with a remote sandbox. ## What Changed - Prefer `/opt/paperclip-runner/bin`, then the user's local bin directory, then PATH. Existing metadata and version validation remains in force. - Qualify Codex 0.153.4, OpenCode 1.18.29, and Claude SDK 0.3.263 / CLI 2.1.263. Update binary digests, TypeScript/Rust checks, registry defaults, and the displayed OpenCode version together. - Share Codex and Claude's native executable with the ACP bridges through exact dependency overrides. Preserve the separately qualified ACP bridge implementations and their security patches. - Expose shared provider-pack CLI launchers; fail the pack build if Codex ACP resolves a separate Codex installation. Update the eval image's other agent CLIs to current stable releases and remove duplicate global provider installs. - Document the single-current-CLI policy in source comments and development guidance. Latest stable releases are resolved at review/build preparation and pinned; task startup never auto-updates. ## Verification - Native-session and adapter-registry suites: 158 tests passed. - Provider suites: 88 tests passed, 7 Linux-only checks skipped on macOS. One existing macOS temporary-path alias assertion passed when rerun with canonical `TMPDIR=/private/tmp`. - Package-contract and OpenCode materialization tests: 11 passed. - Full typecheck, build, and token gates passed. Rust native-provider/recovery tests: 19 passed. - Broad local suite: 5,974 passed, 23 failed, 41 skipped. Failures are in unchanged macOS workspace/path/port and connection suites; focused runtime tests pass. All latest-head Linux PR checks passed, including the full test shards, typecheck, build, runner verification, browser suites, and canary dry run. - The standalone fleet image built with one current provider CLI each and passed native Codex/Claude binary-integrity checks. A disposable Daytona sandbox reported ready in 798 ms; its baked runner completed an API-key `gpt-5.6-luna` turn in 2,430 ms and returned the expected marker with a usage receipt. No runtime artifacts were uploaded or installed. - The normal shared `codex exec` entrypoint also completed an API-key `gpt-5.6-luna` turn in 2,321 ms. - Both image builds verify the complete generated lockfile against a reviewed SHA-256 before package installation or lifecycle execution. Root lockfile changes remain CI-owned. Merge and rollout remain on hold for operator review. ## Risks - Updating provider CLIs changes their behavior for all adapters; version probes and live native smoke testing are required before image promotion. - The image-owned directory takes precedence. Its entries must launch the same shared CLI as the global PATH, not a private older/newer copy. - Application qualification pins and the deployed image must move together. No startup fallback installation is added. - No schema or authentication-policy changes. ## Model Used OpenAI GPT-6 (Codex). The session does not expose a more specific model ID or context-window size. Used reasoning, repository inspection, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
54a99d8840 |
fix(evals): make the chat viewer the default published Evalbook (#12952)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Direct Runner evals retain evidence across model configurations. > - Evalbook already has a grid and a read-only Runner Lab chat viewer. > - Public projection stripped the view and selected a second plain result page. > - This change uses the existing viewer for public and private results. > - The data access differs, but the presentation does not. ## Linked Issues or Issue Description Refs #12931, #12945. Related open runtime-contract PR #11634 does not contain this report-only change. **What happened?** The public direct-eval campaign opened plain result pages. The access-controlled artifact used the chat viewer. Users could not follow the same recorded interaction from the published grid. **Expected behavior** Every newly generated Runner Evalbook opens the existing chat viewer. The grid and durable run history remain. Public evidence has explicit redactions. **Steps to reproduce** Open campaign gha-34062394019-1 from the direct-eval history. Click a result, then compare its plain page with the corresponding Actions artifact. **Paperclip version or commit** Reproduced at |
||
|
|
d96452db05 |
fix(runner): restore Vite 6 viewer compatibility (#12929)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner includes an issue-thread viewer for direct Evalbook reports. > - The runner package uses Vite 6.4.3. > - Dependabot changed the React plugin from version 4.7.0 to version 6.1.1. > - React plugin 6.1.1 requires Vite 7 or Vite 8 package internals. > - The live evaluation workflow found this mismatch when it built the viewer on a clean Linux worker. > - This pull request restores the compatible plugin and adds the viewer build to pull request CI. > - The benefit is that CI detects this class of report-viewer build failure before a paid evaluation campaign starts. ## Linked Issues or Issue Description **What happened?** The direct live evaluation workflow failed before model execution. The `build:issue-thread` command could not load `@vitejs/plugin-react@6.1.1` with Vite 6.4.3. The plugin imported the unavailable `vite/internal` package path. **Expected behavior** The Evalbook issue-thread viewer must build from a clean frozen-lockfile installation before the live evaluation matrix starts. **Steps to reproduce** 1. Check out commit `165ca56a22adb60e5fda56045442d9c8498116a8`. 2. Run `pnpm install --frozen-lockfile --ignore-scripts`. 3. Run `pnpm --filter @paperclipai/paperclip-runner build:issue-thread`. 4. Observe the `ERR_PACKAGE_PATH_NOT_EXPORTED` error for `vite/internal`. **Paperclip version or commit** `165ca56a22adb60e5fda56045442d9c8498116a8` **Deployment mode** Built from source in GitHub Actions on Ubuntu. ## What Changed - Restore `@vitejs/plugin-react` 4.7.0 in the Vite 6 runner package. - Follow the repository policy: trusted PR CI regenerates and verifies the lockfile artifact, and the master refresh workflow commits the lock-only update after merge. - Build the Runner Evalbook viewer in pull request CI. ## Verification - `ci / policy`: regenerated the dependency lock artifact successfully - `pnpm --filter @paperclipai/paperclip-runner build:issue-thread` - `actionlint .github/workflows/pr-trusted.yml .github/workflows/runner-protocol-live-evals.yml` - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs` - `git diff --check` ## Risks Low risk. This change restores the previous React plugin major version for one package. The selected version declares support for Vite 6. Pull request CI now builds the affected viewer directly. > This bug fix does not add or change a roadmap feature. ## Model Used - OpenAI Codex with GPT-5. Tool use and code execution were enabled. The Codex app managed the context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
af8439a70b |
feat(runner): restore direct live eval campaigns and reports (#12909)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner executes agents through native and managed provider drivers. > - The direct live eval layer had drifted from the current Runner contracts. > - The old local workflow did not provide a complete parallel campaign or durable report history. > - The Runner also needed current native OpenCode and OpenRouter qualification. > - This pull request restores the direct campaign, corrects the runtime gaps that the campaign found, and adds safe hosted Evalbook history. > - The benefit is repeatable model comparison against an immutable Runner and eval source revision. ## Linked Issues or Issue Description Refs #11297 Refs #11634 **What existing behavior does this improve?** This improves the direct live `paperclip-runner` eval workflow, provider execution contract, and static Evalbook reporting path. **Current behavior** The direct evals do not have one maintained full campaign on current `master`. OpenCode has no qualified multi-model OpenRouter roster. Parallel provider bursts can compact committed events before the transport observes them. Local reports do not have a separate safe S3 history index. **Proposed behavior** Run one immutable roster-plus-case matrix. Use the shared paid AWS runner fleet. Keep raw artifacts access-controlled. Publish a sanitized canonical Evalbook report under the separate `runner-protocol-evals` S3 prefix. Keep immutable campaign directories plus root history, latest, and latest-green pointers. **Reason and benefit** Maintainers can compare native Codex, native OpenCode, ACPX, Claude Managed, and AWS AgentCore behavior over time. They can inspect failures without mixing this direct protocol layer with browser full-stack E2E. **Breaking changes** None. The new workflow and S3 prefix are additive. The existing Runner full-stack E2E workflow and report remain separate. ## What Changed - Added a trusted two-shard direct live workflow for up to 393 roster-plus-case cells. - Reused the numeric actor allowlist, protected paid environment, and RunsOn fleet controls from Runner full-stack E2E. - Added immutable Runner and eval revision resolution, exact credential boundaries, bounded retries, and cost ceilings. - Added a public report projection that removes sessions, transcripts, tool payloads, state, traces, raw failures, remote profile identities, and credential-shaped values. - Added additive S3 history under `runner-protocol-evals`, with immutable campaigns and mutable root index pointers. - Added native OpenCode model injection and current OpenRouter pricing contracts. - Fixed direct eval completion, workflow execution, semantic discovery, warm-attach state reset, executable binding, and event-burst handling. - Kept Runner browser full-stack E2E behavior and publication separate. - Documented local and hosted direct eval operation. ## Verification - `pnpm --filter @paperclipai/paperclip-runner test:runner-protocol-eval-publish` — 15 passed. - `pnpm --filter @paperclipai/paperclip-runner build:typescript` — passed. - `actionlint .github/workflows/runner-protocol-live-evals.yml .github/workflows/runner-full-stack-e2e.yml` — passed. - Local current matrix at the revision in [paperclip-evals#17](https://github.com/paperclipai/paperclip-evals/pull/17) — 323 cells across 10 enabled configurations completed. - Final local current matrix — 269 passed, 11 behavior failures, and 43 expected macOS-only ACPX platform failures. - Targeted Runner checks — 13/13 eval-session tests, 15/15 publisher/security tests, and package typecheck passed; complete PR CI is green, including all browser E2E shards. ## Risks - Paid live campaigns can consume provider budget. Actor authorization, exact per-cell ceilings, protected environments, and explicit schedule enablement bound this risk. - Public reports can leak provider data. The workflow publishes only a separately projected report and validates every file before upload. - The new workflow cannot publish until it is present on the default branch. This pull request does not change the existing `runner-full-stack-e2e` publication path. - The campaign is large. It uses two GitHub matrices and caps combined concurrency at the shared fleet limit. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex on GPT-5. The exact deployment ID and context-window size are not exposed. The model used reasoning, code editing, browser inspection, repository tools, and live provider execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
af3023f1e3 |
fix(runner): repair paid provider startup paths (#12769)
## Thinking Path > - Paperclip manages AI agents that perform work. > - Paperclip Runner connects durable task runs to local provider processes. > - The full-stack paid matrix exposed failures after the runner integrity repair. > - Verified JavaScript entrypoints lost their relative module graph when Linux executed them through descriptor paths. > - Returned provider startup errors also remained pending and became indeterminate after recovery. > - Sparse Codex tool lifecycle events lost the `write_document` identity before task transcript projection. > - This pull request repairs those three boundaries and makes the structured-question fixture deterministic. > - The benefit is repeatable provider startup, exact failure replay, and correct inline Plan placement. ## Linked Issues or Issue Description Refs #12721 and #12700. **What happened?** The paid runner matrix failed ACPX and OpenCode startup before provider session creation. The runner journal then replaced the original startup error with an indeterminate recovery result. Native Codex saved a Plan but rendered it only as a fallback card. A legacy Claude waiting reply could also echo the reserved terminal marker before the answer arrived. **Expected behavior** Verified JavaScript providers must start from immutable descriptor-backed artifacts. Returned startup failures must persist as terminal failed command results. Native tool lifecycle updates must preserve the `write_document` boundary. Pre-answer fixture output must not contain the reserved terminal marker. **Steps to reproduce** 1. Run the local provider cells in the Runner Full-Stack E2E workflow. 2. Observe ACPX and OpenCode fail during `session.open` before provider execution. 3. Observe recovery report `execution_indeterminate` instead of the original startup error. 4. Run the native Codex Plan cell and observe the fallback Plan card after the tool activity row. 5. Run the legacy Claude structured-question resume cell and observe an early marker echo in waiting prose. **Paperclip version or commit** `0f9452101740835ce0b1488a204bf48acd5bafc3` **Deployment mode** Local development with the paid GitHub Actions acceptance workflow. ## What Changed - Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM entrypoints before hashing and verified descriptor launch. - Anchor ACPX dynamic provider package resolution at a controller-derived provider-pack root and keep that root out of the provider child environment. - Persist executor-returned startup errors as redacted durable failed command results while retaining indeterminate recovery for true process death. - Coalesce sparse native tool items by stable ID so a late `write_document` name, input, and result reach the transcript boundary once. - Forbid the structured-question fixture from spelling or announcing its reserved terminal marker before the user answers. ## Verification - Rust and TypeScript regression tests cover durable failed replay, true crash ambiguity, bundle closure, package-root derivation, environment filtering, exact Codex tool lifecycle coalescing, and prompt determinism. - Local execution is intentionally limited to formatters and static diff checks. GitHub Actions will run tests, type checks, builds, and security checks. - After ordinary CI is green, scoped paid cells will validate one ACPX launch, one OpenCode launch, native Codex Plan projection, and legacy Claude structured resume before a complete matrix rerun. - Prior failing matrix: https://github.com/paperclipai/paperclip/actions/runs/33682434315 ## Risks - Bundling changes the bytes covered by provider launch hashes. Provider-pack generation already hashes the final built files. - ACPX still loads qualified provider packages dynamically. The controller supplies a normalized package root, while existing version, digest, path, and descriptor checks remain active. - Durable `failed` is terminal. Replays return the same redacted result and do not execute the provider effect twice. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5 with agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions coordination. The exact deployed snapshot and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked related public work or described the bug in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] No documentation change is required for this runtime repair - [x] I have considered and documented the risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d95c71027d |
chore(deps-dev): bump @types/react-dom from 19.2.4 to 19.2.5 (#12253)
Bumps [@types/react-dom](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/react-dom) from 19.2.4 to 19.2.5. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/react-dom">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
a0028d7e1b |
chore(deps-dev): bump vitest from 4.1.10 to 4.1.11 (#12262)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.10 to 4.1.11. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v4.1.11</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Revive global concurrency limit for test lifecycle [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> and <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10992">vitest-dev/vitest#10992</a> <a href="https://github.com/vitest-dev/vitest/commit/5146df80b"><!-- raw HTML omitted -->(5146d)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Encode iframeId in tester iframe URL [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a>, <strong>Pduhard</strong> and <strong>Claude Opus 4.8</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10955">vitest-dev/vitest#10955</a> <a href="https://github.com/vitest-dev/vitest/commit/10b2cd201"><!-- raw HTML omitted -->(10b2c)<!-- raw HTML omitted --></a></li> <li>Trigger playwright/chromium gc on lower disk availability [backport to v4] - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>OpenCode</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10951">vitest-dev/vitest#10951</a> <a href="https://github.com/vitest-dev/vitest/commit/9851dbc41"><!-- raw HTML omitted -->(9851d)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>mocker</strong>: <ul> <li>Restrict redirect mocks to the fs allowlist [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10974">vitest-dev/vitest#10974</a> <a href="https://github.com/vitest-dev/vitest/commit/fe5a11d3c"><!-- raw HTML omitted -->(fe5a1)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v4.1.10...v4.1.11">View changes on GitHub</a></h5> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/9bd8d464e6328c567c2dbcd8fdd977d57a9425c2"><code>9bd8d46</code></a> chore: release v4.1.11 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10995">#10995</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/9851dbc41c286a30abfb6b29cce65f3e5b7b40a1"><code>9851dbc</code></a> fix(browser): trigger playwright/chromium gc on lower disk availability [back...</li> <li>See full diff in <a href="https://github.com/vitest-dev/vitest/commits/v4.1.11/packages/vitest">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
313d6ca115 |
fix(runner): materialize pinned OpenCode binary (#12782)
## Thinking Path > - Paperclip manages AI agents and their provider runtimes. > - Paid runner validation installs target dependencies with lifecycle scripts disabled. > - OpenCode leaves a sentinel executable until its package lifecycle script runs. > - Running arbitrary lifecycle code would weaken the paid-secret boundary. > - This pull request materializes one exact pinned binary before secrets are exposed. > - The benefit is working OpenCode validation without trusting dependency install scripts. ## Linked Issues or Issue Description **What happened?** Every local OpenCode paid cell stopped before provider startup because `pnpm install --ignore-scripts` correctly retained `opencode-ai/bin/opencode.exe` as a sentinel. **Expected behavior** The trusted workflow must make the exact lockfile-pinned OpenCode executable available without running package lifecycle scripts. **Steps to reproduce** Run a local legacy or native OpenCode paid cell from the trusted workflow after the target dependency install. The provider health check reports that the OpenCode postinstall script was not run. **Paperclip version or commit** Default branch commit `865b4854fb44d3689f1c0ff17e3e715d52aaea73`. ## What Changed - Materialize only `opencode-linux-x64-baseline@1.18.17` into the matching `opencode-ai@1.18.17` package. - Verify package identity, version, regular-file type, SHA-256 equality, executable permissions, and runtime `--version`. - Invoke the helper for local OpenCode and breadth cells and for remote provider-pack assembly. - Retain `pnpm install --ignore-scripts`. - Add helper and trusted-workflow security regressions. ## Verification - Helper syntax checks passed. - Helper unit tests passed: 2/2. - Workflow-security tests passed: 5/5. - Prettier, actionlint, and diff whitespace checks passed. ## Risks Risk is low and contained to paid runner setup. The helper supports only Linux x64, fails closed on package or version drift, and runs before provider credentials enter the job. ## Model Used OpenAI GPT-5 Codex with repository tools and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change. - [x] I have specified the model used. - [x] I have checked ROADMAP.md and confirmed this does not duplicate planned core work. - [x] I have searched GitHub for duplicate or related PRs and found none. - [x] I have described the issue in this PR with the bug template labels. - [x] I have not referenced internal or instance-local issues. - [x] My branch name describes the change. - [x] Focused local tests pass. - [x] I added tests for the change. - [x] I updated the runner E2E security documentation. - [x] I documented the risks above. |
||
|
|
4e56afec11 |
chore(deps-dev): bump @vitejs/plugin-react from 4.7.0 to 6.1.1 (#12566)
Bumps [@vitejs/plugin-react](https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react) from 4.7.0 to 6.1.1. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitejs/vite-plugin-react/releases">@vitejs/plugin-react's releases</a>.</em></p> <blockquote> <h2>plugin-react@6.1.1</h2> <h3>Add <code>compiler.logDiagnostics</code> option</h3> <p>Recoverable React Compiler diagnostics are no longer logged by default. Set <code>compiler.logDiagnostics</code> to <code>true</code> to log them through Vite. Fatal diagnostics are always logged and fail the transform.</p> <h3>Respect environment sourcemap option for React Compiler transform when <code>builder.sharedPlugins</code> is enabled (<a href="https://redirect.github.com/vitejs/vite-plugin-react/pull/1439">#1439</a>)</h3> <p>The React Compiler transform was using the top-level sourcemap option instead of the environment sourcemap option. This caused a problem when the experimental <code>builder.sharedPlugins</code> was enabled.</p> <h2>plugin-react@6.1.0</h2> <h3>Add experimental native React Compiler support (<a href="https://redirect.github.com/vitejs/vite-plugin-react/pull/1419">#1419</a>)</h3> <p>Add experimental native React Compiler support.</p> <p>You can use it by installing <code>oxc-transform-react</code> and enabling it via the <code>compiler</code> option:</p> <pre lang="sh"><code>npm install -D oxc-transform-react </code></pre> <pre lang="js"><code>import { defineConfig } from 'vite' import react from '@vitejs/plugin-react' <p>export default defineConfig({<br /> plugins: [<br /> react({ compiler: true })<br /> ]<br /> })<br /> </code></pre></p> <h2>plugin-react@6.0.5</h2> <h3>Fixed the react compiler preset filter to be linear (<a href="https://redirect.github.com/vitejs/vite-plugin-react/pull/1353">#1353</a>)</h3> <p>The improved filter in v6.0.3 was non-linear and caused a performance regression (<a href="https://redirect.github.com/vitejs/vite-plugin-react/issues/1349">#1349</a>). The filter was changed to be linear to avoid that.</p> <h2>plugin-react@6.0.4</h2> <h3>Fixed <code>$RefreshSig$ is not defined</code> error when running <code>vite dev</code> with <code>NODE_ENV=production</code></h3> <p>When running <code>vite dev</code> with <code>NODE_ENV=production</code>, the app errored with <code>$RefreshSig$ is not defined</code>. This error is now fixed.</p> <h2>plugin-react@6.0.3</h2> <p>No release notes provided.</p> <h2>plugin-react@6.0.2</h2> <h3>Allow all options in reactCompilerPreset (<a href="https://redirect.github.com/vitejs/vite-plugin-react/pull/1189">#1189</a>)</h3> <p>This is a type only change. Only <code>compilationMode</code> and <code>target</code> options were available for <code>reactCompilerPreset</code>.</p> <h2>plugin-react@6.0.1</h2> <h3>Expand <code>@rolldown/plugin-babel</code> peer dep range (<a href="https://redirect.github.com/vitejs/vite-plugin-react/pull/1146">#1146</a>)</h3> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/vitejs/vite-plugin-react/blob/main/packages/plugin-react/CHANGELOG.md">@vitejs/plugin-react's changelog</a>.</em></p> <blockquote> <h2>6.1.1 (2026-08-28)</h2> <h3>Add <code>compiler.logDiagnostics</code> option</h3> <p>Recoverable React Compiler diagnostics are no longer logged by default. Set <code>compiler.logDiagnostics</code> to <code>true</code> to log them through Vite. Fatal diagnostics are always logged and fail the transform.</p> <h3>Respect environment sourcemap option for React Compiler transform when <code>builder.sharedPlugins</code> is enabled (<a href="https://redirect.github.com/vitejs/vite-plugin-react/pull/1439">#1439</a>)</h3> <p>The React Compiler transform was using the top-level sourcemap option instead of the environment sourcemap option. This caused a problem when the experimental <code>builder.sharedPlugins</code> was enabled.</p> <h2>6.1.0 (2026-08-19)</h2> <h3>Add experimental native React Compiler support (<a href="https://redirect.github.com/vitejs/vite-plugin-react/pull/1419">#1419</a>)</h3> <p>Add experimental native React Compiler support.</p> <p>You can use it by installing <code>oxc-transform-react</code> and enabling it via the <code>compiler</code> option:</p> <pre lang="sh"><code>npm install -D oxc-transform-react </code></pre> <pre lang="js"><code>import { defineConfig } from 'vite' import react from '@vitejs/plugin-react' <p>export default defineConfig({<br /> plugins: [<br /> react({ compiler: true })<br /> ]<br /> })<br /> </code></pre></p> <h2>6.0.5 (2026-07-30)</h2> <h3>Fixed the react compiler preset filter to be linear (<a href="https://redirect.github.com/vitejs/vite-plugin-react/pull/1353">#1353</a>)</h3> <p>The improved filter in v6.0.3 was non-linear and caused a performance regression (<a href="https://redirect.github.com/vitejs/vite-plugin-react/issues/1349">#1349</a>). The filter was changed to be linear to avoid that.</p> <h2>6.0.4 (2026-07-22)</h2> <h3>Fixed <code>$RefreshSig$ is not defined</code> error when running <code>vite dev</code> with <code>NODE_ENV=production</code></h3> <p>When running <code>vite dev</code> with <code>NODE_ENV=production</code>, the app errored with <code>$RefreshSig$ is not defined</code>. This error is now fixed.</p> <h2>6.0.3 (2026-06-23)</h2> <h3>Improve the react compiler preset filter to reduce false-positives (<a href="https://redirect.github.com/vitejs/vite-plugin-react/pull/1138">#1138</a>)</h3> <p>Improved the filter in the react compiler babel preset to reduce the false-positives so that less modules are processed by the react compiler.</p> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/04cac5020e349f452d76c5a4f6d788ad4b38930a"><code>04cac50</code></a> release: plugin-react@6.1.1 (<a href="https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react/issues/1440">#1440</a>)</li> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/82d35abe4946eddd4e6456802bf2b53444e264f2"><code>82d35ab</code></a> fix(react): respect environment sourcemap option when <code>builder.sharedPlugins</code>...</li> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/397e8471a559f18a16dd21bd797ac01a369dabdc"><code>397e847</code></a> fix(react): make logging diagnostics an opt-in for React Compiler (<a href="https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react/issues/1431">#1431</a>)</li> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/61006e6f52124821c24121a78712f7162ae36f5b"><code>61006e6</code></a> fix(deps): update all non-major dependencies (<a href="https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react/issues/1433">#1433</a>)</li> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/e2a649cbaa7334d6991f843563683975667e1be1"><code>e2a649c</code></a> chore: use <code>deps.neverBundle</code> instead of <code>external</code> in tsdown config (<a href="https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react/issues/1430">#1430</a>)</li> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/fb2d6f3635acbb0f3acbd0e9a914f6c620460957"><code>fb2d6f3</code></a> fix(deps): update all non-major dependencies (<a href="https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react/issues/1427">#1427</a>)</li> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/39b31735bf79c2dd380eedaba7ed849256f92a29"><code>39b3173</code></a> release: plugin-react@6.1.0 (<a href="https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react/issues/1428">#1428</a>)</li> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/f1340b0c760b1c16e1b780eeba46fd933ddd52eb"><code>f1340b0</code></a> feat(react): add native React Compiler support (<a href="https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react/issues/1419">#1419</a>)</li> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/9ab698eafc38ffa14861db450291ed2f6f557557"><code>9ab698e</code></a> fix(deps): update all non-major dependencies (<a href="https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react/issues/1375">#1375</a>)</li> <li><a href="https://github.com/vitejs/vite-plugin-react/commit/68c0cb8796ce18bd049c3d05c5210eaf0617eac0"><code>68c0cb8</code></a> release: plugin-react@6.0.5 (<a href="https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react/issues/1362">#1362</a>)</li> <li>Additional commits viewable in <a href="https://github.com/vitejs/vite-plugin-react/commits/plugin-react@6.1.1/packages/plugin-react">compare view</a></li> </ul> </details> <details> <summary>Maintainer changes</summary> <p>This version was pushed to npm by <a href="https://www.npmjs.com/~GitHub%20Actions">GitHub Actions</a>, a new releaser for <code>@vitejs/plugin-react</code> since your current version.</p> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
5716fe907e |
test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner subsystem executes agent work across local and managed provider backends. > - The lower pull requests restore the task runtime, provider backends, and managed-provider control plane. > - The restored system needs repeatable full-stack checks before it can ship safely. > - Paid live checks also need clear access, cost, and secret controls. > - This pull request adds acceptance, live evaluation, chaos, and release gates for the restored runner stack. > - The benefit is measurable runner parity with safer release decisions. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change covers runner tests, release workflows, server contracts, and evaluation tools. **Problem or motivation** The runner stack did not have one complete acceptance surface for native Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could miss provider drift, task-view regressions, cost-policy errors, and destructive cleanup errors. **Proposed solution** Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid workflows. Add live evaluation, chaos, cost-limit, redaction, and release contract checks. Add AWS AgentCore infrastructure and guarded provisioning tools. Keep the native runner experimental flag off by default. **Alternatives considered** We considered manual smoke tests only. They do not give repeatable evidence and they do not protect release branches. We also considered one large pull request. The stacked pull requests keep each review below the Greptile file limit. **Roadmap alignment** This work supports the shipped Cloud / Sandbox agents milestone and the shipped Agent evals & feedback milestone in `ROADMAP.md`. Related stack: - #12699 adds managed provider backends and lifecycle support. - #12691 adds qualified OpenCode and ACPX provider backends. - #12685 restores task runtime rendering and steering. ## What Changed - Add the runner full-stack harness with 57 catalog cells and 60 unit tests. - Add a Daytona runner image with digest-pinned base images and base-aware image-content checks. - Add guarded live evaluation and chaos workflows with a fixed 40-execution matrix; live and full-stack paid schedules now run only on Sundays or by manual dispatch. - Add in-flight reported-usage cost stops, post-turn cost caps, exact-threshold failure classification, secret redaction, retry classification, and actor authorization. - Reattach stream and hard-budget listeners before restart-recovery continuations so restored paid sessions cannot bypass in-flight interruption. - Preserve OpenCode usage and cost across tool-loop messages and turns while exposing an explicit current-run delta to durable accounting. - Keep PNG/WebM evidence in access-controlled artifacts only, reject SVG, and publish only pruned inert structured per-attempt evidence. - Add AWS AgentCore infrastructure, provisioning checks, and smoke tools; reject unsafe model identifiers, require exact stack ownership markers, and make failed-stack replacement explicit. - Add evaluation-session contracts and capability reports. - Add release workflow checks for immutable action pins, frozen dependency installs, exact weekly cron shape, paid-run guards, provider-secret isolation, and chaos test paths. - Reauthorize the original and triggering numeric actor IDs as the first step of every provider-secret job, including partial reruns, before checkout or provider access. - Give each full-stack matrix cell only its matching provider credential, expose Daytona only to Daytona cells, and disable shared dependency caches anywhere paid credentials or OIDC write access are present. - Protect the legacy manual E2E workflow with the same default-branch, allowlist, environment, and per-job authorization boundary. - Rotate live-eval candidates by week and retain 120 days of compatible history so the seven-week trend window remains viable. - Restore the root runner-acceptance commands and reconcile reported snapshots, raw receipts, and terminal usage without double counting or losing late usage. - Mark ACPX token deltas exact only when every budget field is present, keep cumulative cost/request authority separate, reject non-USD cost labeling, and include thought tokens in output-token budgets. - Keep `enableNativeRunner` off by default. The acceptance harness enables it only in its isolated test instance. ## Verification Passed locally: - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm test:runner-acceptance:typecheck` - `pnpm test:runner-acceptance` (19 tests) - focused OpenCode proxy, driver, runnerd transport, live-session, and turn-stream tests (106 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/clean-room-server.test.ts` (22 tests) - `pnpm test:e2e:runner:typecheck` - `pnpm test:e2e:runner:unit` (62 tests) - `node --test scripts/__tests__/release-verify-workflow.test.mjs` - `pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals` (22 tests) - `pnpm -r typecheck` - `pnpm build` - `node --test packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs` (6 tests) - `git diff --check` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core --lib --locked` (161 tests) - focused ACPX provider-event tests (10 tests) - The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged. I did not run paid live provider jobs or provision AWS resources. Those checks need credentials and can create cost. ## Risks The paid workflows can create provider cost. They require an allowlisted original and triggering actor, the protected `runner-e2e-paid` environment, explicit opt-in variables, and cost limits. The four provider credentials exist only in that master-only environment, which requires allowlisted reviewer approval and disables administrator bypass; repository and organization Actions scopes contain no copies. Provider usage arrives after a billable request, so the live guard cannot prevent one request from crossing a threshold. It interrupts immediately on the first reported threshold hit and permits no continuation. Visual evidence can contain secrets rendered as pixels. PNG/WebM remain only in access-controlled workflow artifacts; SVG and per-attempt XML are excluded, and S3/Pages receive a pruned structured dashboard. The AWS scripts can create cloud resources. They use explicit commands, least-privilege roles, KMS encryption, saved nonsecret metadata, and explicit teardown. This pull request does not enable the experimental native runner for existing instances. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The model used extended reasoning, tool use, code execution, and parallel subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5458940a6e |
feat(runner): add offline evaluation tooling (#12653)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs repeatable evaluation contracts. > - Evaluation code must stay separate from provider launch and production orchestration. > - Offline fixtures need stable compatibility, scoring, traceability, and report rules. > - Published Runner consumers need only the supported evaluation contract surface. > - This pull request adds offline evaluation tooling and a workspace-private matrix kernel. > - The benefit is deterministic evaluation without credentials or paid provider calls. ## Linked Issues or Issue Description Refs #11297 This pull request extracts the offline evaluation unit from the earlier aggregate Runner work. ## What Changed - Add a workspace-private, provider-neutral evaluation matrix kernel. - Add the public `@paperclipai/paperclip-runner/evals` compatibility and native execution contracts. - Add fail-closed runnerd artifact and protocol compatibility checks. - Add deterministic workflow catalogs, scoring, traceability, and report generation. - Add sanitized Codex, OpenCode, and ACPX fixtures. - Add package-boundary and clean-consumer checks. - Add the eval package manifest to the Docker dependency stage. - Add the generated protocol fixture digest without changing the lockfile. ## Verification GitHub Actions must run: - Runner TypeScript and Rust type checks. - Runner unit and protocol tests. - Evaluation kernel tests. - Workflow traceability checks. - Clean-consumer and package-boundary checks. - Repository test, type-check, build, policy, and security gates. No local test command was run. The repository owner requested GitHub-only verification. ## Risks This is a large greenfield review surface with 51 files. The code does not launch a live provider or load credentials. Package and protocol drift fail closed. The workspace lockfile remains under the existing CI-owned process. ## Model Used OpenAI Codex with the GPT-5 agent model. The work used high reasoning, repository inspection, tool use, and parallel code review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
560e7e48b5 |
feat(runner): add SDK and developer tooling (#12608)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ad0ad43cf4 |
feat(runner): activate qualified Claude ACPX runtime (#12590)
## Thinking Path > - Paperclip Runner already has a hardened ACPX path for Codex. > - Claude can reuse that protocol only with an exact package/model profile and provider-lifetime fencing. > - Pi needs a separately spawned runtime whose executable does not yet have the descriptor-confined verified launch used by the ACP server. > - This pull request therefore activates Claude only and keeps Pi unavailable before installation or process launch. ## Linked Issues or Issue Description **Subsystem affected** Paperclip Runner ACPX driver, runtime host, sidecar, backend factory, package dependency, and provider conformance tests. **Problem or motivation** The production ACPX backend was Codex-only. Claude needs the same fail-closed model, authorization, cancellation, cleanup, and recovery boundaries without exposing an unsafe secondary runtime path. **Proposed solution** Generalize the hardened ACPX runtime for the exact qualified `claude` profile, add the pinned Claude ACP package and reviewed isolation patch, and reject Pi before installation, backend construction, sidecar initialization, Rust session admission, or process creation. **Alternatives considered** Activating Pi in this PR was rejected after security review because its secondary runtime executable was pathname-based and lacked the verified descriptor/snapshot boundary. Pi is deferred to a dedicated follow-up. Replaying the older generic ACPX implementation was rejected because it predates current hardening. **Roadmap alignment** ROADMAP.md does not list a conflicting ACPX-provider project. This extends the existing Runner provider architecture. ## What Changed - Generalized the ACPX backend, driver, runtime adapter, host, and sidecar for the qualified Claude profile. - Added Claude ACPX activation through its exact pinned package/model pair and isolated-settings patch. - Added provider-lifetime fencing for non-Codex qualified ACPX sessions. - Kept Pi dependencies and its patch out of the package and build configuration. - Added fail-closed Pi rejection at driver validation, backend construction, runtime-host admission, sidecar initialization, and Rust session validation. - Added focused tests for Claude selection, model enforcement, lifecycle fencing, cancellation, recovery, and Pi rejection. - Did not change or commit `pnpm-lock.yaml`; CI regenerates the PR lockfile under the existing repository policy. ## Verification - GitHub Actions is the authoritative verification environment for this PR. - CI runs dependency policy, runner package checks, protocol parity, typecheck, build, security, and stack policy. - Local tests were not run because this checkout is resource constrained, per the requested workflow. ## Risks - Claude package behavior can drift from the qualified protocol; the package and patch are pinned and admission verifies the exact profile. - Unsupported providers and models fail closed. - Pi remains unavailable until descriptor-confined verified launch exists for its separate runtime. - Existing Codex ACPX behavior remains covered by shared conformance tests. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs - [x] I have described the issue in the PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name contains no internal task identifier - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented risks - [ ] All applicable Paperclip CI gates are green - [ ] Greptile is 5/5 with no actionable findings ## Stack - Position: lowest unmerged PR - Base: `master` - Previous: [#12588](https://github.com/paperclipai/paperclip/pull/12588), merged qualified OpenCode runtime - Next: [#12591](https://github.com/paperclipai/paperclip/pull/12591), native application integration |
||
|
|
2e5a24e177 |
feat(runner): add qualified OpenCode runtime (#12588)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner provides a durable execution boundary for supported providers. > - The current production runtime supports Codex but cannot execute OpenCode sessions. > - OpenCode needs a qualified transport, strict input mapping, and normalized events. > - This pull request adds the OpenCode runtime as one isolated provider unit. > - The benefit is a reviewable provider expansion that does not weaken the existing Codex path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: packages/paperclip-runner and the Codex-local adapter configuration contract. **Problem or motivation** Paperclip Runner has provider-neutral contracts, but the production backend factory cannot start a qualified OpenCode session. This blocks OpenCode from using the durable runner path. **Proposed solution** Add the qualified OpenCode app-server proxy, driver, MCP bridge, backend, fixtures, and factory wiring. Keep existing Codex behavior unchanged. **Alternatives considered** Keeping OpenCode only on the direct adapter path would avoid this runtime work, but it would not provide durable runner recovery or normalized provider events. **Roadmap alignment** ROADMAP.md does not list a conflicting provider-runtime project. This change extends the existing Paperclip Runner architecture. ## What Changed - Added the qualified OpenCode app-server proxy and input queue. - Added collaboration-mode and provider-event normalization. - Added the OpenCode MCP bridge and native session backend. - Added strict fixtures and focused unit coverage. - Added only the package exports and adapter configuration required by this runtime. - Kept deferred SDK, lab, eval, and public package surfaces out of this change. ## Verification - GitHub Actions is the authoritative verification environment for this PR. - Run the package type checks and focused OpenCode tests in CI. - Run repository typecheck, test, build, security, and policy gates through the stack-aware workflow. - Local tests were not run because this checkout is resource constrained. ## Risks - OpenCode protocol changes could affect event normalization or recovery. - The driver fails closed on malformed input and unsupported runtime behavior. - Existing Codex selection remains unchanged unless the stored provider is OpenCode. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack - Position: 1 of 4 - Base: master - Next: additional qualified provider runtimes |
||
|
|
9ad8dbffa0 |
feat(runner): add Codex ACPX sidecar (#12410)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner needs a bounded process boundary for each qualified provider runtime. > - The ACPX contract, Codex profile, structured questions, and recovery rules now exist in the package. > - The package does not yet provide an executable that applies those rules to a real Codex ACPX host. > - Provider admission must retain process, credential, and cleanup ownership on every failure path. > - A production selector must not depend on an unreviewed provider-generic sidecar. > - This pull request adds one installable Codex-only ACPX sidecar and leaves it unselected. > - The benefit is a testable package boundary for later runnerd integration without changing current execution selection. ## Linked Issues or Issue Description Refs #12409 Refs #12386 This pull request implements the Codex-only executable for the ACPX sidecar contract merged in #12386. It builds on the question conformance gate merged in #12409. Runnerd and the server do not select this executable in this pull request. ## What Changed - Publish the `paperclip-runner-acpx-sidecar` package binary and document its current boundary. - Add a versioned stdin/stdout sidecar that admits only the qualified Codex ACPX profile and exact initialized model. - Support atomic session open, run attachment, turn start and cancellation, tool and input resolution, session read and snapshot, safe suspension, close, and recovery identity checks. - Bind runtime directory, workspace, permission mode, provider identity, run identity, and semantic tool catalog before use. - Validate completion and blocked results against the PRP result contract. Bound pending tools, pending inputs, messages, events, usage, diagnostics, and errors. - Redact provider output and convert file locations to bounded workspace-relative display data. Do not treat displayed paths as file-access authority. - Harden verified executable loading, module resolution, launch environment filtering, process-group guardianship, and provider termination. - Retain managed credentials and every failed-admission resource until the exact provider cleanup proves ownership was released. - Keep failed-admission cleanup alive with bounded backoff until the provider exits. Do not scrub credentials or admit a replacement while cleanup still owns the provider. - Use one runtime-host cleanup-owner registry and preserve sequential cleanup retries across command timeouts and shutdown. - Arm a credential-free same-group watchdog before provider admission so guardian death reaps even a stopped provider; retain an independent kernel EOF proof before releasing credentials. - Add sidecar process, lifecycle, location, package, driver, credential, installation, runtime-adapter, and runtime-host regression tests. - Add `tsx` as a package test-only development dependency for the real TypeScript sidecar process test. ## Verification - Replay base: `b93ad538b63c81a1e3d24bbb54c02f8effdea787` (`master` after #12409 merged). - Exact replay head: `ca7e93f4385c37289c82b360ca8def0d88c1bd00`. - Stable patch ID for the resolved 18-file delta: `d293717a1e9f2485b61c553c58ea37690fc0a4fa`. - The intended pull request delta contains exactly these 18 files: - `packages/paperclip-runner/README.md` - `packages/paperclip-runner/package.json` - `packages/paperclip-runner/src/cli/acpx-runtime-sidecar.ts` - `packages/paperclip-runner/src/cli/acpx-runtime-sidecar.test.ts` - `packages/paperclip-runner/src/cli/acpx-sidecar-lifecycle.ts` - `packages/paperclip-runner/src/cli/acpx-sidecar-locations.ts` - `packages/paperclip-runner/src/cli/acpx-sidecar-locations.test.ts` - `packages/paperclip-runner/src/drivers/acpx/codex-acpx-driver.ts` - `packages/paperclip-runner/src/drivers/acpx/codex-acpx-driver.test.ts` - `packages/paperclip-runner/src/drivers/acpx/codex-credentials.ts` - `packages/paperclip-runner/src/drivers/acpx/codex-credentials.test.ts` - `packages/paperclip-runner/src/drivers/acpx/codex-runtime-adapter.ts` - `packages/paperclip-runner/src/drivers/acpx/codex-runtime-adapter.test.ts` - `packages/paperclip-runner/src/drivers/acpx/installation-integrity.ts` - `packages/paperclip-runner/src/drivers/acpx/installation-integrity.test.ts` - `packages/paperclip-runner/src/drivers/acpx/runtime-host.ts` - `packages/paperclip-runner/src/drivers/acpx/runtime-host.test.ts` - `packages/paperclip-runner/test/acpx-codex-package-contract.test.mjs` - The resolved combined delta is 4,725 additions and 665 deletions; it preserves the lower-PR runtime-close semantics, repairs stale successful-admission fixtures, deterministically observes renewed reconciliation, accepts authoritative same-host cleanup recovery without dropping pending owners, ignores superseded cleanup failures after a newer owner recovers, and requires both guardian exit and independent provider-lifetime EOF before releasing cleanup or credential ownership. A readiness-gated, credential-free watchdog also reaps a stopped provider if its guardian is externally killed. - This change updates the runner package manifest, README, and test-only dependencies. It does not change `pnpm-lock.yaml`, a workflow, migration, server route, UI path, runnerd selection, or current direct-adapter behavior. - Focused GitHub verification: **PASSED** for the sidecar process, lifecycle, location, package-contract, driver, credential, installation-integrity, runtime-adapter, and runtime-host suites on the replayed head. - Package verification: **PASSED** for the clean tarball, Node shebang, and exact binary mapping on the replayed head. - GitHub Actions and security checks: **PASSED** for the replayed exact head; full CI run `33357846557` completed 23/23 jobs successfully, and Superagent, Socket, Snyk, supply-chain, and contributor-trust checks are green. Storybook was intentionally skipped because this PR does not touch its paths. - Greptile: **5/5** on the exact head with zero unresolved review threads. - No local test result is claimed. GitHub Actions is the authoritative verification environment for the replayed revision. ## Risks This change has medium security and lifecycle risk because the new executable crosses a process, credential, filesystem, and provider boundary. The sidecar fails closed on unsupported providers, models, permissions, identities, catalogs, forms, commands, and persistent-state deletion. A cleanup owner can remain alive until a stubborn provider exits. Its retries use bounded backoff, and admission stays closed while ownership remains. Command and shutdown waits remain bounded without abandoning the underlying cleanup. The new `tsx` dependency is development-only. The package exposes a new binary, but no runnerd, server, UI, or direct-adapter path starts it in this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6, extended reasoning, repository tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9ca24bba3c |
feat(runner): pin the Codex ACPX runtime (#12400)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The package-local host boundary is ready for a concrete ACP implementation, but the first production profile is Codex only. > - ACPX must not inherit the server process environment or choose an executable by pathname after admission. > - Codex must not re-enable ambient apps, memory, skills, MCP configuration, or instructions inside its isolated home. > - This pull request pins only the two required production packages and applies narrowly tested host patches. > - The benefit is a minimal dependency boundary that follows the repository's CI-owned lockfile process. ## Linked Issues or Issue Description **Agent or provider** Codex through `acpx@0.13.1` and `@agentclientprotocol/codex-acp@1.6.2`. **Why this adapter is useful** The injected runtime host needs a concrete ACP session manager and the exact reviewed Codex ACP server. Upstream ACPX does not yet expose a host-owned spawn callback, and upstream Codex ACP does not yet apply Paperclip's isolated instruction, MCP, app, memory, and skill boundary. Both behaviors are required before the dependency can execute inside the runner. **How the agent is invoked** The next pull request will adapt these pinned packages to the private runtime host. ACPX receives a host-owned callback that consumes the already verified executable lease. Codex receives only the isolated environment, explicit base instructions, explicit MCP servers, and the skills rooted in its private `CODEX_HOME`. This pull request alone does not spawn either package or register an adapter. **Additional context** This pull request is stacked on #12399. It adds no Pi, Claude, AWS, SDK, lab, browser, or UI dependency. It intentionally does not commit `pnpm-lock.yaml`: the repository policy job regenerates a manifest-only PR lockfile artifact for downstream frozen installs, and the lockfile bot updates master separately. ## What Changed - Pin `acpx` to `0.13.1` and the Codex ACP server to `1.6.2` in the runner package. - Register both patches in the pnpm 9 root configuration and newer-pnpm workspace configuration. - Preserve the existing embedded-Postgres and ACPX 0.12 patch entries used by other packages. - Patch ACPX to evaluate an allowlisted environment at child-spawn time and keep spawn cwd out of provider-visible session identity. - Patch ACPX to accept a host-owned spawn callback with the resolved arguments and options, allowing the verified command lease to own execution. - Patch Codex ACP to retain runner-owned MCP server identity in permission requests. - Patch Codex ACP to pass explicit Paperclip base instructions on both start and resume. - In isolated mode, disable ambient apps, memory, and existing MCP configuration; load skills only from `CODEX_HOME`; and configure only requested servers. - Add a package contract test that enforces exact versions, Codex-only dependency scope, both pnpm patch registries, and every required patch hook. ## Verification - Both patch files dry-apply successfully to fresh published tarballs for `acpx@0.13.1` and `@agentclientprotocol/codex-acp@1.6.2`. - A local no-lockfile install applied both patches; their runtime markers and exact installed versions were inspected. - Runner TypeScript typecheck — passed against the patched packages. - Runner package tests — passed: 16 Node protocol/package tests and 426 Vitest tests. - `pnpm -r typecheck` — passed for all applicable workspaces. - `pnpm build` — passed, including runner binary, server, UI, and workspace packages. - `git diff --check` — passed. - The diff contains 6 files and does not change `pnpm-lock.yaml`, a GitHub workflow, server selection, or UI behavior. ## Risks The primary risk is drift between published package contents and checked-in compiled patches. Exact versions are pinned, both patches are exercised by package-contract gates, and CI performs the authoritative regenerated-lockfile frozen install. The spawn callback does not grant a new executable path: the following adapter must consume the opaque verified command lease. Codex isolation changes activate only when `PAPERCLIP_ACPX_ISOLATED_CONTEXT=1`, so existing direct Codex adapters are unaffected. ## Model Used OpenAI Codex with GPT-5 and repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing public item or described the issue in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal task identifier - [x] I have run the affected tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the dependency, patch, isolation, and lockfile boundaries - [ ] All applicable GitHub Actions are green - [ ] Greptile is 5/5 with every actionable comment resolved - [x] I will address all review findings before requesting merge |
||
|
|
e13856eb37 |
feat(runner): define ACPX sidecar contract (#12386)
## Thinking Path > - Paperclip Runner now has a complete guarded Codex vertical slice. > - The next provider series must not start by importing a provider implementation or dependency bundle. > - ACPX needs one bounded, versioned process boundary shared by TypeScript and Rust. > - A schema is the authority; checked-in generated inventories keep both languages in lockstep. > - Unknown versions, commands, event types, and properties must fail closed. > - This pull request therefore lands only the sidecar wire contract and its drift gate. > - No ACPX runtime, dependency, executable, package export, or production selection is added. ## Linked Issues or Issue Description This is the first package-local unit in the post-Codex provider series. **What happened?** The integration branch contains an ACPX provider, but its TypeScript sidecar and Rust client need a small shared authority before either implementation can be reviewed safely. Importing the final integration implementation directly would mix the protocol, runtime, third-party dependencies, and production wiring. **Expected behavior** The schema defines every ACPX sidecar request, response, event, command, event type, and protocol version. Generated TypeScript and Rust inventories must drift-check against that schema. No runtime can select or execute ACPX yet. **Steps to reproduce** 1. Change the protocol version, command inventory, or event inventory in the schema. 2. Run the runner protocol type check without regenerating the language inventories. 3. Observe the drift gate fail. **Paperclip version or commit** Stacked on `runner-server-semantic-codex` at `ebd7f9df7`. ## What Changed - Add the internal ACPX sidecar v2 JSON Schema outside the public PRP v1 schema catalog. - Generate one TypeScript inventory and one Rust inventory from that schema. - Add generate and check hooks to the existing runner protocol-type workflow. - Add fail-closed AJV tests for all three message families, version drift, unknown commands, and extra properties. - Keep the generated Rust module unregistered until the Rust ACPX transport exists. ## Compatibility Boundary - Codex remains the only production runner provider. - `paperclip_runner` selection and the default-off rollout flag are unchanged. - No ACPX package, patch, lockfile, binary entry point, root export, server file, UI file, workflow, or dependency is added. - The schema is shipped with the existing `protocol` directory but is not added to the public PRP manifest. - Existing direct adapters continue through their current paths. - Diff against the actual stacked base: 6 files. ## Verification - Runner TypeScript typecheck and both generated-contract drift gates — passed. - Runner TypeScript tests — 37 files and 355 Vitest tests passed; 11 Node contract tests passed. - Rust provider-bridge regression suite after restacking — 14/14 passed. - `pnpm -r typecheck` — passed for all applicable workspaces. - `pnpm build` — passed, including runner binary, server, UI, and workspace packages. - `pnpm test:run` — attempted; the local host reproduced unrelated workspace/Postgres and port-exposure failures in unchanged server suites. The changed runner contract suites pass, and the repository's serialized/sharded GitHub checks remain authoritative for those host-sensitive suites. - Prettier, rustfmt, generated-source drift checks, and `git diff --check` — passed. - `pnpm-lock.yaml` is unchanged. ## Risks The main risk is allowing schema and generated language inventories to diverge. Build and typecheck now fail on any drift. The sidecar implementation and third-party ACPX packages are deliberately absent, so this PR cannot alter runtime behavior or expand the production attack surface. ## Model Used OpenAI Codex with GPT-5 and repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs - [x] I have described the issue and expected behavior in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal task identifier - [x] I have run the affected tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the compatibility and security boundary - [ ] All applicable GitHub Actions are green - [ ] Greptile is 5/5 with every actionable comment resolved - [x] I will address all review findings before requesting merge |
||
|
|
0cedb45df3 |
build(deps-dev): bump typescript from 5.9.3 to 7.0.2 (#11880)
Bumps [typescript](https://github.com/microsoft/TypeScript) from 5.9.3 to 7.0.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/microsoft/TypeScript/releases">typescript's releases</a>.</em></p> <blockquote> <h2>TypeScript 7.0.2</h2> <p><a href="https://devblogs.microsoft.com/typescript/announcing-typescript-7-0/">https://devblogs.microsoft.com/typescript/announcing-typescript-7-0/</a></p> <p>This tag was originally released at: <a href="https://github.com/microsoft/typescript-go/releases/tag/typescript%2Fv7.0.2">https://github.com/microsoft/typescript-go/releases/tag/typescript%2Fv7.0.2</a></p> <h2>TypeScript 6.0.3</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.2%22">fixed issues query for TypeScript 6.0.2 (Stable)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.3%22">fixed issues query for TypeScript 6.0.3 (Stable)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.2%22">fixed issues query for TypeScript 6.0.2 (Stable)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0.1 RC</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0-rc/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0 Beta</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0-beta/">release announcement</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22+is%3Aclosed+">fixed issues query for Typescript 6.0.0 (Beta)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/microsoft/TypeScript/commit/1e4744d68260a7cb91b62b12edc3f6a2187faaf1"><code>1e4744d</code></a> Merge branch 'main' into ts7-release</li> <li><a href="https://github.com/microsoft/TypeScript/commit/a5a219c3b5da0db4fa0ecf6c0b1f588c9af9c669"><code>a5a219c</code></a><code>microsoft/typescript-go#4558</code></li> <li><a href="https://github.com/microsoft/TypeScript/commit/ecfe30dce91368d52c9a49b6095bb0b673a238f8"><code>ecfe30d</code></a> Update status localization</li> <li><a href="https://github.com/microsoft/TypeScript/commit/5de25b5f8fec2ca35eadaed041f1f06d2e214895"><code>5de25b5</code></a> Hide executable name in TypeScript status</li> <li><a href="https://github.com/microsoft/TypeScript/commit/d7ce74a75da2b80e8201506a1599c06549432b93"><code>d7ce74a</code></a> Show bundled TypeScript version for packaged servers</li> <li><a href="https://github.com/microsoft/TypeScript/commit/29be66a607707f90d7a53103a4469bb3015a4d54"><code>29be66a</code></a> Correct TS 7 release version to 7.0.2</li> <li><a href="https://github.com/microsoft/TypeScript/commit/ed2bd1bfa4aac5211ce4bc58fcd1313c7eddc8ff"><code>ed2bd1b</code></a> Merge branch 'main' into ts7-release</li> <li><a href="https://github.com/microsoft/TypeScript/commit/887307575c58ea640dbeba3b4e8fdb6347cd3044"><code>8873075</code></a> Bump the github-actions group across 1 directory with 3 updates (microsoft/ty...</li> <li><a href="https://github.com/microsoft/TypeScript/commit/9427131ae2d4e230a90ee8a09daac4e75da3e311"><code>9427131</code></a> Set up stable / nightly extension split, other prep (microsoft/typescript-go#...</li> <li><a href="https://github.com/microsoft/TypeScript/commit/d4eaca5460a1f5f02a829e62706794b0a6fb903e"><code>d4eaca5</code></a><code>microsoft/typescript-go#4549</code></li> <li>Additional commits viewable in <a href="https://github.com/microsoft/TypeScript/compare/v5.9.3...v7.0.2">compare view</a></li> </ul> </details> <details> <summary>Maintainer changes</summary> <p>This version was pushed to npm by <a href="https://www.npmjs.com/~microsoft1es">microsoft1es</a>, a new releaser for typescript since your current version.</p> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
397de98193 |
feat(runner): add flagged Codex execution adapter (#12188)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner now has protocol, provider, tool, package, persistence, and hidden server boundaries. > - The server still cannot select that path for a real agent heartbeat. > - A new runtime must not change any existing direct adapter. > - An experimental runtime must fail closed when its rollout flag is off. > - This pull request adds one guarded Codex vertical slice through runnerd. > - The benefit is a production-built runner path that users cannot start by default. ## Linked Issues or Issue Description Refs #11962 Refs #12111 Refs #12169 Refs #12176 **Subsystem affected** Cross-cutting. The change affects the runner package, server orchestration, shared settings, and adapter configuration UI. **Problem or motivation** The hidden PRP coordinator cannot execute a real heartbeat. The application also needs an explicit rollout boundary before it can expose the experimental runner. Existing direct adapters must keep their current execution and finalization behavior. **Proposed solution** Add `paperclip_runner` as a Codex-only adapter behind the default-off `enableNativeRunner` instance flag. Select the native runtime only for that adapter. Persist the run binding before runnerd starts. Wait for the durable PRP result and terminal event. Resume the real Codex provider thread on later heartbeats. Keep persisted native runs readable and recoverable after the flag changes. **Alternatives considered** The server could route `codex_local` through runnerd. That option would change an existing adapter and weaken rollback safety. The server could expose all providers now. That option would add unreviewed provider behavior. The build could depend on a prebuilt runner binary. That option would make source builds architecture-dependent and difficult to verify. **Roadmap alignment** This work supports the shipped enforced-outcomes, governed-tool, and self-healing-run milestones. It does not add a new roadmap surface. It is the guarded execution step after the merged hidden runner boundaries. **Additional context** This is the next replacement for the closed large runner pull request. Task-thread presentation remains a separate follow-up so this change can preserve the current direct-adapter UI. ## What Changed - Add `paperclip_runner` as an explicit Codex-only adapter. - Add the default-off `enableNativeRunner` instance flag. - Reject fresh create, hire, import, switch, and execution requests while the flag is off. - Allow edits to persisted runner agents while the flag is off. - Recover an already persisted native run even after the flag is disabled. - Keep every built-in direct adapter on its existing runtime path. - Persist an immutable native run binding and revisioned completion contract before runnerd starts. - Execute server to PRP to runnerd to Codex to server through the hidden coordinator. - Validate the durable result against the terminal event and exact completion criteria before finalization. - Preserve the Codex provider thread ID and use `thread/resume` on the next heartbeat. - Strip unsupported Codex configuration fields from the experimental adapter. - Build a target-native release runner binary from source and vendor it into the server distribution. - Install Rust only in the Docker build stage. Do not add a workflow or lockfile change. - Stop the runner process group on completion, cancellation, and forced shutdown. ## Verification - Run `pnpm --filter @paperclipai/paperclip-runner check:all`. All 69 TypeScript tests and 58 Rust tests pass. Protocol, conformance, replay, formatting, and generated-file checks pass. - Run the 12 focused adapter, settings, runtime-selection, coordinator, direct-isolation, and real Codex integration test files. All 186 tests pass. - The real integration test uses PostgreSQL, HTTP, WebSocket, runnerd, and a fake Codex app server. It proves one `thread/start` followed by one `thread/resume`. - Run `pnpm -r typecheck`. - Run `pnpm build`. - Run `pnpm check:token-gates`. - Build the Docker `build` target from a clean context. Confirm that the server distribution contains an executable `paperclip-runnerd` built with Debian Rust 1.85. - Start the server through the source-mode tsx entry point with the package `dist` directory absent. Confirm the vendor shim resolves source exports and the server boots. - Run `pnpm test:run` twice. On this macOS host, 405 files pass and 1 file skips. Eight untouched workspace and loopback tests fail because macOS resolves `/tmp` and `/var` through `/private` and because PID-derived test ports exceed 65535. Linux CI must pass the full suite. - Confirm that the diff contains 52 files. Confirm that it contains no `.github` or `pnpm-lock.yaml` change. ## Risks - The feature flag is off by default. A fresh native start fails with a stable error while the flag is off. - A persisted native run remains recoverable after the flag changes. This prevents rollout changes from corrupting recorded work. - Only local Codex execution is accepted. Other providers and remote work modes fail closed. - Existing direct adapters do not start runnerd, create native rows, use native status arbitration, or enter native finalization. - The runner receives its one-use bootstrap ticket through the child environment. The server does not put the ticket in command arguments or logs. - The server validates the company, task, agent, run, runner, session, completion contract, result, and terminal binding before it accepts completion. - The build compiles a target-native Rust binary. Cross-platform release packaging remains a later concern. Source builds and Docker builds compile for their current target. - Docker needs enough build memory for the existing server TypeScript compile. The Docker build stage sets a 4 GB V8 heap limit. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment ID and context-window size are not exposed. The model used agentic reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and applicable tests pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ffff1fe6e3 |
feat(runner): define package API and verification boundary (#12129)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package now has protocol, transport, provider, catalog, and authorization foundations. > - Its first upstream package boundary should expose only the implemented runtime and test-helper surfaces. > - Rust correctness belongs in the repository existing build verification, without introducing a parallel release process. > - Direct package creation must build the files declared by the package manifest. > - This pull request defines the minimal package API and verifies the optimized runner binaries in the existing PR and release Build jobs. > - The benefit is a production-ready runner package boundary with minimal build-process change. ## Linked Issues or Issue Description Refs #11962 This pull request replaces one bounded part of the archived large runner change. It follows the package-local authorization change in #12126. ## What Changed - Export only `@paperclipai/paperclip-runner` and `@paperclipai/paperclip-runner/testing`. - Keep Node-only fixture loading and semantic conformance helpers out of the runtime root. - Add a provider-neutral semantic conformance kit with stable JSON comparison and fail-closed input checks. - Keep deferred SDK, eval, browser, React, lab, and command surfaces private. - Pin the runner Rust toolchain to 1.97.1 with the minimal profile and `rustfmt`. - Run the Rust workspace tests in release mode. - Launch the optimized `paperclip-runnerd` and fake-harness binaries in process-level integration coverage. - Add one `pnpm --filter @paperclipai/paperclip-runner check:all` step to each existing PR and release Build job. - Make the existing server `prepack` lifecycle run its existing build after it prepares UI assets. - Document that no production adapter starts runnerd yet. This revision adds no standalone GitHub Actions job. It adds no server runner dependency or runner vendoring. It adds no Docker bootstrap or clean-consumer harness. It does not change `pnpm-lock.yaml`. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` - 66 TypeScript tests - 8 protocol contract tests - 56 Rust unit and integration tests - Release-mode integration coverage launches the optimized runnerd and fake-harness binaries. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/server-package-build-script.test.ts` (2 tests) - Clean `pnpm pack` from `server/` rebuilt the server and produced both `package/dist/index.js` and `package/dist/index.d.ts`. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` (8 tests) - `pnpm -r typecheck` - `pnpm build` - `pnpm check:token-gates` - `git diff --check` - No `pnpm-lock.yaml` diff. - The diff changes 12 files. ## Risks The runner adds Rust work to the existing Build jobs. These jobs can take longer on a cold cache. The pinned toolchain makes contributor and CI behavior reproducible. Cargo tests use `--release` to verify optimized executables. The server prepack lifecycle now performs the build that its published entry points require. This can make direct server packing slower. This pull request does not wire runnerd into the server. It does not select runnerd for any adapter. Existing application execution and finalization paths remain unchanged. ## Model Used OpenAI Codex with GPT-5. Agentic coding mode used repository tools, code execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
23048f1219 |
Add canonical semantic action catalog to Paperclip Runner (#12121)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner now has a durable PRP transport and a Codex provider bridge. > - Codex must use stable, provider-neutral action contracts before Paperclip can grant run-scoped tool access. > - A catalog must describe actions without granting permission to discover or invoke them. > - Generated inventory must stay synchronized with its TypeScript source. > - This pull request adds the canonical Codex-spine semantic action catalog inside the runner package. > - The benefit is a small review unit for schemas and inventory before authorization and dispatch land. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This pull request extends private runner infrastructure in `packages/paperclip-runner`. **Problem or motivation** The Codex provider bridge has no canonical description of the Paperclip actions that a later authorization layer can project into a run. Independent operation lists can drift in names, claims, task modes, effects, and input bounds. **Proposed solution** Add one immutable v1 catalog for the first 27 Codex-spine actions. Give each action a stable identifier, placement, effect, required claims, supported task modes, and JSON Schema input and output contracts. Generate a deterministic JSON inventory from that source and fail package checks on drift. **Alternatives considered** The combined runner branch contains larger live and scenario catalogs with authorization, bindings, labs, and other providers. That change is too large for this review unit. A generic API escape hatch would also bypass the operation-level boundary, so this catalog excludes it. **Roadmap alignment** This work supports the governed tool access direction in `ROADMAP.md`. It does not add a tool gateway, application binding, server endpoint, or production authorization decision. **Additional context** Refs #12111 and #11962. Pull request #12111 was squash-merged first. This branch starts at the resulting `master` commit. Its delta is 10 files. ## What Changed - Added 27 versioned, provider-neutral semantic action declarations for the Codex spine. - Added bounded JSON Schema input contracts and normalized operation receipt output contracts. - Added placement, effect, claim, mode, and role metadata. - Added a deeply frozen public catalog and an operation lookup helper. - Added a deterministic checked-in JSON inventory and generation commands. - Added a byte-for-byte drift gate to the package build. - Added AJV schema compilation, mutation-bound, forged-field, immutability, inventory, and non-executable-boundary tests. - Exported only the catalog types and declarations from the existing package root. - Documented that catalog membership does not grant discovery, authorization, dispatch, or application binding. - Kept server code, UI code, other providers, scenario-only actions, labs, generic API access, authorization, dispatch, and receipts processing out of this pull request. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` passes. - TypeScript protocol tests pass: 8 Node tests and 49 Vitest tests. - All package Rust tests and conformance and replay parity checks pass. - `pnpm --filter @paperclipai/paperclip-runner check:semantic-action-catalog` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm check:token-gates` passes. - Prettier and `git diff --check` pass for the changed source and documentation files. - The generated catalog matches its source byte for byte. - The secret scan is clean. - The delta against `master` is 10 files. `pnpm-lock.yaml` is unchanged. - `pnpm test:run` completed locally with 4,692 passing tests, 19 skipped tests, and 24 failures in 8 unchanged server test files. The failures reproduce the established local macOS path-alias, listener, and workspace-runtime baseline. No changed-file test failed. Linux CI remains the repository handoff authority. - The full Linux PR workflow passes, including the aggregate `verify` gate. - Snyk, Socket, Superagent security, and supply-chain checks pass. - Greptile is 5/5 with no actionable comments, recommendations, or follow-ups. - Storybook visual regression skipped by design because this pull request changes no UI file. - Browser and migration tests are not applicable because this pull request changes no server, UI, database, or migration file. ## Risks Production behavior is unchanged because no consumer projects this catalog into a provider run. The main risks are contract drift, unbounded mutation input, forged scope fields, accidental executable authority, and generated inventory drift. Closed input schemas, explicit bounds, a frozen catalog, tests, and the byte drift gate cover these risks. The later authorization layer must still bind every action to the active run and company before discovery or invocation. I checked `ROADMAP.md`. This change is private contract infrastructure for the governed tool access direction. It does not duplicate a shipped or public product surface. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4ffa8de4e2 |
Add Codex provider bridge to Paperclip Runner (#12111)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The package-local runner now has a durable PRP transport, but it cannot execute a real provider. > - The first provider must preserve PRP identities while using Codex native thread and turn identities. > - Recovery must resume the same Codex thread without starting a duplicate turn. > - Provider output must become bounded and provider-neutral before it crosses PRP. > - Semantic tools must remain unavailable until the catalog and authorization layers exist. > - This pull request adds the Codex provider bridge inside the runner package only. > - The benefit is a reviewable provider slice with no server or user-facing behavior change. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This pull request extends private provider infrastructure in `packages/paperclip-runner`. **Problem or motivation** The durable runner from #12100 has no production provider. It cannot start Codex app-server, map its events, cancel or steer a turn, deliver a structured question, or recover a native thread after process restart. **Proposed solution** Add a supervised Codex app-server transport and a normalized runner backend. Persist the Codex thread and active turn identities. Resume and inspect the exact thread after restart. Convert supported notifications into bounded PRP events. Keep the dynamic tool inventory empty. **Alternatives considered** The combined runner branch implements several providers, semantic tools, server coordination, and UI integration together. That change is too large for one review unit. Reusing the direct `codex_local` adapter would also couple this package layer to the existing server execution path. **Roadmap alignment** This work supports the governed tools and self-healing run direction in `ROADMAP.md`. It does not add a server endpoint, runtime adapter, rollout flag, or user-facing behavior. **Additional context** Refs #12100 and #11962. Pull request #12100 was squash-merged first. This branch starts at the resulting `master` commit. Its current delta is 16 files. ## What Changed - Added a Codex-only app-server process transport with bounded JSONL frames and buffered notifications. - Added strict provider descriptor validation for the Codex driver, working directory, launch arguments, model, instructions, and non-interactive approval policy. - Started new Codex threads with an empty dynamic tool inventory and the named workspace-only permission profile. - Added native turn start, steering, interruption, cancellation, thread reads, and structured question responses. - Added thread and active-turn binding checks for provider requests and notifications. - Added provider-neutral normalization for session, turn, item, plan, usage, tool execution, notice, and structured input events. - Bounded and redacted provider text and process output before durable persistence. - Added private atomic provider state for the descriptor, thread ID, account session ID, active turn ID, and unacknowledged normalized events. - Added exact-thread recovery through `thread/resume` and `thread/read`. Recovery does not issue another `turn/start` for an active turn. - Preserved active native turn identity across unexpected provider exit and reconciled it before later start, interrupt, or snapshot commands. - Added stable provider-event identities, per-event durable commit and acknowledgement, and a bounded fingerprint receipt journal that prevents duplicate delivery across outbox and provider-ack crash windows. - Extended the durable command executor with provider event polling and explicit process shutdown on stop, suspend, revocation, lease expiry, and runtime expiry. - Preserved completed shutdown behavior when the command result is replayed after a disconnect. - Added a fake Codex app-server and integration tests for response buffering, structured questions, interruption, provider exit, unacknowledged-event recovery, durable resume, and duplicate-turn prevention. - Added a focused `test:codex` package command for the provider integration suite. - Kept server code, UI code, other providers, semantic catalogs, tool authorization, and production runtime selection out of this pull request. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` passes. - TypeScript contract tests pass: 8 Node tests and 44 Vitest tests. - Rust tests pass: 43 unit tests, 5 Codex integration tests, 3 public durable-recovery tests, 2 local-runner tests, and 3 process-supervisor tests. - Rust conformance and replay parity checks pass against the shared PRP fixtures. - `cargo clippy --workspace --all-targets -- -A clippy::filter-map-bool-then -D warnings` passes. The narrow allow covers an unchanged replay implementation from the preceding contract pull request. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm check:token-gates` passes. - `git diff --check` passes. - The delta against `master` is 16 files. The package lockfile is unchanged. The PR workflow generates its temporary lockfile artifact from the changed package manifest. - `pnpm test:run` completed locally with 4,690 passing tests, 19 skipped tests, and 26 failures in 8 unchanged server test files. The failures reproduce the established local macOS path-alias, listener, port-range, and workspace-runtime baseline. No changed-file test failed. Linux CI remains the repository handoff authority. - Browser and migration tests are not applicable because this pull request changes no server, UI, database, or migration file. - The full Linux PR workflow passes. One unchanged heartbeat recovery test timed out on the first pass and passed on the failed-only rerun; the aggregate `verify` gate is green. - Greptile is 5/5 on the final commit. All four review threads are resolved. ## Risks Production behavior is unchanged because no server code starts this provider. The main risks are a provider process escape, cross-thread event confusion, secret leakage, duplicated turns, duplicated or lost provider events, lost questions, and unsafe recovery. Process-group supervision, identity binding, private bounded state, redaction, durable command replay, retained event acknowledgements, bounded durable receipts, exact-thread reconciliation, and integration tests cover these risks. Semantic tools remain undiscoverable in this layer. I checked `ROADMAP.md`. This change is private provider infrastructure for planned control-plane work. It does not duplicate a shipped or public product surface. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b76e36d6cf |
Add durable PRP transport and recovery (#12100)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The package-local runner can supervise a local process, but it cannot yet survive a broken controller connection. > - A production transport must authenticate both peers without putting the bootstrap secret on the wire. > - Commands and events must remain bounded, ordered, and recoverable across reconnects and crashes. > - Retrying an uncertain side effect is unsafe, so indeterminate outcomes must fail closed instead of running twice. > - This pull request adds those transport and recovery guarantees inside the runner package only. > - The benefit is a durable PRP boundary that can be reviewed before any provider or server integration exists. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This pull request extends private transport infrastructure in `packages/paperclip-runner`. **Problem or motivation** The local runner introduced by #12095 has no authenticated network handshake, durable outbox, reconnect lease, cumulative acknowledgement, or crash-safe command journal. A dropped connection could otherwise lose an event or tempt a controller to repeat a side effect whose outcome is unknown. **Proposed solution** Add an authenticated PRP v1 WebSocket transport, encrypted frames, lease-based reconnects, a bounded durable event outbox, cumulative acknowledgements, and an idempotent command journal. Preserve pending commands before execution and classify the crash window as indeterminate so an uncertain side effect is never repeated automatically. **Alternatives considered** The combined runner branch implements transport together with Codex, semantic tools, and server coordination. That change is too large for one review unit. Keeping transport in memory would make reconnect and crash recovery unverifiable. Re-running a pending command after restart would weaken the at-most-once side-effect boundary. **Roadmap alignment** This work supports the governed tools and self-healing run direction in `ROADMAP.md`. It does not add a production provider, server endpoint, adapter, feature flag, or user-facing behavior. **Additional context** Refs #12095 and #11962. Pull request #12095 was squash-merged first. This branch has been rebased onto the resulting `master` commit, and its current delta is 13 files. ## What Changed - Added a loopback-only WebSocket connection policy with one-time DNS resolution and pinned reconnect addresses. - Added an HMAC mutual-authentication handshake that never sends the bootstrap ticket over the socket. - Added AES-256-GCM secure frames with per-direction keys, monotonic counters, and session-bound authenticated data. - Added one-use bootstrap-ticket handling and lease-based reconnect validation with expiry, revocation, and epoch checks. - Added a private, symlink-resistant state directory with atomic, synchronized state replacement. - Added a bounded durable event outbox, priority-zero reserve, cumulative acknowledgements, and reconnect replay of only the unacknowledged suffix. - Added a bounded command journal with contiguous sequence enforcement, persistent results, and deterministic duplicate responses. Duplicate replay requires a SHA-256 match over the complete canonical command. - Persisted commands before their effects. A crash after persistence but before result storage returns an indeterminate terminal result and does not execute the command again. - Migrated pre-fingerprint command journals by compacting through their persisted controller cursor. Legacy redelivery fails closed instead of reconstructing an incomplete identity or repeating an uncertain effect. - Added strict limits and validation for frames, state, results, outbox entries, command history, and redacted diagnostics. - Added a transport-only `paperclip-runnerd --connect-url` mode. It handles lifecycle commands and rejects provider commands because no provider is present in this pull request. - Added a full disconnect-before-ack fault test that reconnects with the lease, replays identical command and event state, and proves the effect ran once. - Kept provider transports, semantic tools, server integration, and production runtime selection out of this pull request. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` passes. - TypeScript contract tests pass: 8 Node tests and 44 Vitest tests. - Rust tests pass: 33 unit tests, 3 public durable-recovery integration tests, plus the existing 2 local-runner and 3 process-supervisor tests. - The disconnect-before-ack, lease reconnect, duplicate command, malformed state, unknown command, bounds, and crash-window tests pass. - Rust conformance and replay parity checks pass against the shared PRP fixtures. - `cargo clippy --workspace --all-targets -- -A clippy::filter-map-bool-then -D warnings` passes. The narrow allow covers an unchanged replay implementation from the preceding contract pull request. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm check:token-gates` passes. - `git diff --check` passes. - The delta against `master` is 13 files. The package lockfile is unchanged. - `pnpm test:run` completed locally with 4,686 passing tests, 19 skipped tests, and 30 failures in 8 unchanged server test files. The failures reproduce the established local macOS path-alias, listener, port-range, and workspace-runtime baseline. No changed-file test failed; Linux CI remains the repository handoff authority. - Storybook visual regression is not applicable because this pull request changes no UI or story files. ## Risks Production behavior is unchanged because no server code starts or connects to this transport. The main risks are secret disclosure, forged or replayed frames, state corruption, unbounded disk growth, duplicated side effects, and incorrect recovery. Mutual authentication, encrypted counter-bound frames, private atomic state, explicit bounds, cumulative acknowledgements, a durable command journal, fail-closed indeterminate recovery, and fault-injection tests cover these risks. I checked `ROADMAP.md`. This change is private transport infrastructure for planned control-plane work. It does not duplicate a shipped or public product surface. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6b20cc97cc |
Add local fake runner supervision (#12095)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs a small local process model before it can connect to a production provider or server. > - The TypeScript PRP contracts now define the expected replay behavior. > - A second language implementation must produce the same result from the same fixtures. > - Local child processes also need bounded input, bounded output, and complete descendant cleanup. > - This pull request adds a package-local Rust runner, a scripted fake harness, and deterministic parity checks. > - The benefit is a testable process boundary with no production Paperclip behavior change. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This pull request adds private test infrastructure to `packages/paperclip-runner`. **Problem or motivation** The PRP contracts have no second implementation on `master`. There is also no small harness that can prove process cleanup, command idempotency, terminal reconciliation, or bounded JSONL handling without a production provider. **Proposed solution** Add a minimal Rust workspace. Add a local runner process, a scripted fake harness, a bounded process supervisor, and Rust conformance and replay checks. Keep all binaries package-local. Do not connect them to the Paperclip server. **Alternatives considered** The combined runner branch includes provider transports, durable networking, SDKs, labs, and server behavior. That change is too large for this review unit. A TypeScript-only harness would not test cross-language contract parity. **Roadmap alignment** This work supports the governed tools and self-healing run direction in `ROADMAP.md`. It does not add a user-facing runtime, adapter, endpoint, or rollout flag. **Additional context** Refs #12091 and #11962. Pull request #12091 was merged before this branch opened. This branch is based on the current `master`. Its delta is 25 files. ## What Changed - Added a minimal locked Rust workspace with only `serde` and `serde_json` dependencies. - Added a package-local `paperclip-runnerd` local mode and a scripted fake harness. - Added bounded controller input, harness input, subprocess output queues, line sizes, log retention, script sizes, script steps, and command history. - Added contiguous controller and harness sequence checks and equivalent-command replay handling. - Added process-group supervision that cleans up child processes and remaining descendants after forced or natural harness exit. - Added runner-owned terminal reconciliation for success, failure, interruption, cancellation, controller closure, and protocol failure. - Added Rust conformance output and deterministic replay summaries for the shared PRP fixtures. - Added fake scripts for success, failure, interruption, interaction, duplicate terminal output, process cleanup, and oversized output. - Added package scripts and documentation for the Rust and cross-language checks. - Kept provider transport, server integration, semantic tools, and production runtime selection out of this pull request. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` passes. - TypeScript contract tests pass: 8 Node tests and 44 Vitest tests. - Rust tests pass: 20 unit tests, 2 local-runner tests, and 3 process-supervisor tests. - The Rust conformance and replay parity checks pass against the shared fixtures. - The natural-exit and forced-exit tests confirm that the harness and its worker process are stopped. - The oversized-frame test confirms that a harness frame above the configured limit is rejected. - `pnpm -r typecheck` passes after the final rebase to `master`. - `pnpm build` passes after the final rebase to `master`. - `pnpm check:token-gates` passes. - `git diff --check` passes. - The delta against `master` is 25 files. The package lockfile is unchanged. - `pnpm test:run` completed locally with 4,686 passing tests, 19 skipped tests, and 30 failures in 8 unchanged server test files. The failures are local macOS path-alias, listener, port-range, and workspace-runtime baseline failures. No changed-file test failed, and every applicable Linux CI shard passes. - Storybook visual regression skipped intentionally because this pull request changes no UI or story files. ## Risks Low production risk. No server code invokes the new binaries. The package remains private. The main risks are process leaks, unbounded local input, and cross-language drift. Bounded queues and sizes, process-group cleanup tests, fixture manifests, and parity checks cover these risks. I checked `ROADMAP.md`. This change is private test infrastructure for planned control-plane work. It does not duplicate a shipped or public product surface. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b2d1673b9e |
Add TypeScript PRP replay contracts (#12091)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs one typed interpretation of the language-neutral PRP contract. > - The JSON Schemas and fixtures now exist, but TypeScript consumers cannot validate or replay them yet. > - A deterministic reducer must define how duplicate delivery and source gaps affect the projected session. > - Result and question contracts must also validate untrusted provider and user input before later runtime code uses it. > - This pull request adds those TypeScript contracts and replay oracles without adding a process, provider, endpoint, or production behavior. > - The benefit is a reviewable and testable TypeScript foundation for the local runner and transport pull requests. ## Linked Issues or Issue Description **Subsystem affected** This change affects the private `@paperclipai/paperclip-runner` package. It does not change an existing server or adapter execution path. **Problem or motivation** The PRP v1 schemas do not yet provide TypeScript types, runtime validators, normalized result handling, or a deterministic session projection. Later Rust, transport, provider, and server work needs one tested TypeScript oracle instead of separate interpretations. **Proposed solution** Generate a checked-in TypeScript schema bundle from the PRP v1 sources. Add derived types, AJV validation, result and question validation, deterministic replay, a reducer, and generated golden snapshots. Export only these implemented root-package surfaces. **Alternatives considered** The combined runner branch adds the TypeScript contracts together with Rust, providers, semantic authorization, SDKs, labs, and server behavior. That delta is too large for normal review. Handwritten duplicate protocol types would also create a drift risk. **Roadmap alignment** This work supports the governed tool and control-plane direction in `ROADMAP.md`. It does not enable a new production adapter or endpoint. **Additional context** Refs #12087 and #11962. This pull request was prepared on #12087, then rebased onto its squash merge before opening. The current delta against `master` is 37 files. ## What Changed - Added JSON-Schema-derived PRP v1 types and AJV runtime validation. - Added fail-closed required-version checks and cross-envelope binding checks. - Added provider-neutral completion-result and structured-question contracts. - Added normalization for accepted legacy provider result aliases before strict validation. - Added a deterministic session reducer for replay, duplicate delivery, source gaps, requests, items, results, and terminal state. - Added generated replay snapshots and compact parity summaries for six accepted fixtures. - Added schema-bundle, manifest, and replay-golden drift gates. - Added only the root package export. Deferred testing, SDK, evaluation, lab, provider, and browser entry points remain unavailable. ## Verification - `pnpm --filter @paperclipai/paperclip-runner test` passed with 8 protocol tests and 44 TypeScript tests. - `pnpm --filter @paperclipai/paperclip-runner typecheck` passed. - `pnpm --filter @paperclipai/paperclip-runner check:replay-goldens` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm check:token-gates` passed. - `git diff --check` passed. - The delta against its declared base is 37 files. - `pnpm test:run` was executed locally. The package tests pass, while the macOS repository run retains the unchanged local-environment failures documented on #12087. The complete Linux CI matrix must pass on this commit. - A scoped scan found no secret-like values, internal references, or deferred-provider file names. - Greptile found an unbounded sequence-gap allocation. Commit `4a405c17` caps detailed missing IDs at 256, records the full missing count and truncation state, and rejects sequence values above the exact JavaScript integer range. The focused tests, workspace typecheck, build, and token gates pass after this fix. ## Risks Low production risk. The package remains private. This change adds no process, network endpoint, provider bridge, server integration, database change, or execution selection. The main risk is protocol interpretation drift. Generated schema and replay gates detect that drift. Browser and CSP-specific validator packaging remains deferred to its later package boundary. I checked `ROADMAP.md`. This change defines contracts for planned control-plane work and does not add overlapping product behavior. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fdbc69172d |
feat(runner): add PRP v1 schemas and fixtures (#12087)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs a language-neutral contract between the server and the runner process. > - A shared contract must exist before TypeScript, Rust, transport, or provider implementations can depend on it. > - Required protocol versions must fail closed, while safe optional fields must remain compatible. > - The contract also needs deterministic fixtures and a drift gate for later cross-language work. > - This pull request adds that contract without adding runtime behavior. > - The benefit is a small, reviewable source of truth for the next implementation pull requests. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This pull request adds a private package contract for later server, TypeScript, and Rust work. **Problem or motivation** Paperclip Runner does not have a small language-neutral protocol boundary on `master`. A runtime implementation without this boundary can drift between languages, accept unsupported required versions, or silently change canonical fixtures. **Proposed solution** Add PRP v1 JSON Schemas, accepted and rejected fixtures, a Codex structured-question fixture, and a generated SHA-256 manifest. Run compatibility and manifest checks during the package build. Keep the package private and export nothing in this pull request. **Alternatives considered** The combined runner branch contains schemas together with providers, SDKs, labs, and server behavior. That change is too large for normal review. Generating TypeScript validators in this pull request would also cross into the next review unit. **Roadmap alignment** This contract supports the governed tool and control-plane direction in `ROADMAP.md`. It does not enable a new production adapter or endpoint. **Additional context** Refs #12084 and #11962. This pull request was reviewed as a stack on #12084, then rebased and retargeted to `master` after #12084 merged. The current delta is 38 files. ## What Changed - Added 20 PRP v1 JSON Schemas with stable identifiers and resolved references, including explicit cross-language conformance input and output schemas. - Added canonical replay, cross-language, and Codex question fixtures. - Added accepted cases for additive optional fields and a rejected case for an unsupported required protocol version. - Added a deterministic manifest with SHA-256 digests for every schema and fixture. - Added package-local schema-instance, schema-reference, compatibility, question-ID, conformance-pair, and drift checks. - Added a private workspace package with no public exports and no production runtime behavior. - Added the package manifest to the Docker dependency-stage inventory required for every workspace package. This does not copy or build runner runtime code into the production image. - Kept the provider descriptor and question fixture Codex-only. No deferred provider package or dependency is present. ## Verification - `pnpm install --frozen-lockfile` passed with Node 24.19.0 and pnpm 9.15.4. No lockfile change is committed. - `pnpm --filter @paperclipai/paperclip-runner check:protocol` passed with 8 tests. - The committed AJV 2020-12 gate accepted every canonical v1 replay, question, and cross-language conformance fixture. It rejected the required v2 fixture, a replay fixture with a missing required command ID, and conformance output with a missing session ID. - `pnpm -r typecheck` passed. - `pnpm build` passed and ran the protocol manifest drift check. - `pnpm check:token-gates` passed. - `node ./scripts/check-docker-deps-stage.mjs` passed. - `git diff --check` passed. - The delta against its declared base is 38 files. - `pnpm test:run` completed with 4,687 passing tests, 19 skipped tests, and 29 failures across 9 unchanged server files. The failures reproduce macOS path aliases, local listener probes, workspace-runtime assumptions, and one connection-retry timeout. No changed-file test failed. Linux CI must pass before this pull request is ready. - `pnpm check:tokens` reports existing personal-name references outside this pull request. A scoped scan of `packages/paperclip-runner` found no secret-like values, internal references, or deferred-provider names. - PR #12084 was squash-merged, and this branch was rebased onto that merge and retargeted to `master`. The first master-base policy run correctly caught the missing Docker dependency-stage manifest copy; commit `4fa1ea7c` fixes that gate, and the complete Linux matrix is green. - Serialized server shard 1 initially hit an unchanged heartbeat test-harness timeout and a later assertion in the same file. Its isolated rerun passed in 3m57s. All other shards passed on their first attempt. - Greptile reviewed the final commit at 5/5 with no blocking failure. Both earlier actionable validation threads are resolved, and no review thread remains open. ## Risks Low production risk. The package is private and has no exports, server adapter, endpoint, or process. AJV is a package-only development dependency that the server workspace already uses. The main risk is contract churn before the TypeScript and Rust consumers land. The generated manifest and compatibility fixtures make that churn explicit. I checked `ROADMAP.md`. This change defines a contract for planned control-plane work and does not add overlapping product behavior. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |