mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
29c8fb0b66cb01167f71e7107a89628e36b5e032
29
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
18e8c121d9 |
fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path > - Paperclip manages agents through a shared native runner. > - Built-in harness support should ship with Paperclip's public distribution. > - Grok already speaks ACP; it does not require a new public bridge package. > - Sandbox provisioning owns the native executable and its pinned version. > - The runner must verify that prerequisite without downloading it during npm installation. > - This change separates built-in launcher identity from external runtime identity. > - Clean npm installation and live staging checks verify the distribution boundary. ## Linked Issues or Issue Description Refs #13882, #13973, #13977, #13979. This follow-up now targets master after #13882 was squash-merged. It replaces the private `@paperclipai/grok-acp` workspace package with runner-owned assets. Current master is included so the branch also contains the merged scheduler, complete-event capture, and durable cleanup fixes. ## What Changed - Ship Grok launcher and qualification metadata inside the runner's compiled output and the public server's vendored runner tree. - Remove the separate Grok npm package and all package-manager install hooks for this runtime. - Require the checksum-verified Grok Build 1.0.13 binary at `/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution environment. Provision it explicitly in the Daytona image and CI setup. - Keep native binaries outside the provider pack. Bind the built-in launcher into the pack manifest. - Preserve executable leases, descriptor-backed startup, credential fences, permissions, and exact ACP model admission. - Use `builtin:grok-acp` and `native:grok` as profile identities. Historical package-profile sessions fail closed on resume rather than being silently reinterpreted. - Resolve built-in assets from the authenticated sidecar location, including public server npm layouts. Keep the controller path out of provider environments. - Add clean npm tarball installation verification to the existing trusted canary CI job and the admitted manual EC2 verification path. It stages a unified release version and runs npm lifecycle scripts, then verifies missing-prerequisite rejection and admission after separate provisioning without credentials or inference. - Include the controller-owned provider pack in stamped Cloud images. Unstamped local images omit the pack and remain usable; remote ACPX requires full source provenance. - Correct CLI approval-page metadata for an already authenticated Cloud board user; approval authorization remains unchanged. - Honor explicit native-runner enablement in the Cloud agent picker and direct setup page, keeping the flag disabled by default. - Allow selecting the execution environment before connecting credentials. Include Grok in the existing authenticated hello-probe flow, targeting its pinned native prerequisite for runner setup. - Recover an existing subscription sign-in conflict through an explicit cancel-and-retry action, serialized after cancellation succeeds. - Preserve the selected ACPX harness before normalizing config fields, so new Grok agents use the Grok default model. - Keep the credential-free Cloud provider pack root-owned and readable after runtime UID remapping; verify manifest and referenced asset access under an unrelated unprivileged UID during image builds. - Archive prior failover backups alongside explicitly replaced harness state, preserving evidence while preventing stale backups from blocking a fresh replacement. - Update Daytona image content inputs and contract tests for the built-in assets and explicit provisioner. - Document and regression-test the shared `approve-all` default for Grok setup, saved configuration, and native execution. Explicitly saved restrictions remain unchanged. ## Verification Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05` incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the base PR was squash-merged. All 12 conflicts came from incoming files identical to the tested pre-squash base. The final tree exactly matches a three-way merge using that original base, preserving built-in Grok distribution and removal of the obsolete private package. All 252 focused runner/UI tests, six npm-isolation tests, and token gates pass. Fresh exact-head Greptile review is 5/5 with no outstanding findings; security scans and EC2 native compilation pass. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([run 36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)). The repository owner explicitly authorized bypassing code-owner approval after all checks passed; no CI checks or repository protection settings are bypassed or changed. The only remaining PR was removed from the completed stack metadata to permit native auto-merge. Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)). The initial attempt lost two EC2 runners to shutdown signals and stalled a third shard during dependency preparation; all three passed the same-commit failed-job-only retry. Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. [Final public npm verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764) passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public packages, an executed offline lifecycle sentinel, unchanged consumer lock, built-in launcher, missing-prerequisite rejection, and verified separately provisioned binary/command lease. Provisioning and cleanup require no host privilege elevation; only the positive probe mounts the temporary native binary read-only. The verifier is unchanged by the final master merge. All six isolation tests and an offline npm smoke test pass. The prior head had 56 green CI checks and a 5/5 review after two unchanged tests timed out and passed a failed-job-only retry ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)). All 56 recovery-display/lineage tests pass; re-review cleared the already-covered missed-retry concern. Earlier EC2 failures remain retained: [npm lockfile rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203), [missing compiler in the slim image](https://github.com/paperclipai/paperclip/actions/runs/36440210984), and the aggregate 15-minute test timeouts in those broad runs. Both broad attempts passed typecheck, token gates, Product E2E type/unit checks and build. The focused EC2 lane preserves the existing trusted-actor and immutable-source gates. Earlier documentation/test checkpoint `ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior unchanged. 154 focused tests pass across configuration building, native provider resolution, permission policy, credentials, UI configuration, and new-agent setup (including both Grok auth modes); token gates pass. All fresh CI is green for this head: 56 successful checks/statuses and two intentional skips ([run 36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)). Greptile is 5/5 with no new findings. Grok already inherits the shared `approve-all` default, so unattended setup requires no manual permission change. Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final staging continuation failure before provider startup: explicit replacement archived the old harness but left its failover backups active, which caused `runner_harness_state_mismatch`. The regression fails before the fix and passes after it; all eight adjacent recovery-safety cases also pass. Old backups remain inspectable inside the continuity archive. All fresh CI is green at this head ([run 36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)), with a 5/5 review. One unrelated Cursor test timed out in the initial server shard; the same-commit failed-job rerun passed, and both attempts are retained. Staging deployment is confirmed healthy on this revision. The controller image is `ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`. The final browser-created staging task passed on this exact revision with API authentication: context read → structured human question → controller restart → answer submission → same native provider session resumed → document saved → task Done. The two turns took approximately 119s and 77s. The actual write receipt was applied, and the saved document has exactly one revision containing the selected answer and requested marker. Usage and cost were not reported. [Controller image build](https://github.com/paperclipai/paperclip/actions/runs/36360889243). - Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`: all CI green (53 successful checks/statuses, two intentional skips), including repository typecheck/build/tests, native Runner tests, browser shards, and canary installation checks. [CI run 36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529). Greptile is 5/5 with no unresolved findings. - Focused checks cover Grok credentials, executable admission, launcher assets, provider-pack paths/permissions, workflow contracts, setup defaults, CLI authorization, and subscription conflict recovery. All 39 protocol definitions validate. Final integration checks pass 124 catalog/evidence/cache tests and nine project-form tests; token gates pass. Some local dependency checks could not load the stale installed dependency tree; the corresponding fresh EC2 checks pass. - Clean public npm installation passed on EC2 at `8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run 36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)): 17 unified-version packages, lifecycle scripts enabled, built-in launcher present, no separate Grok package or npm-downloaded binary, missing prerequisite rejected, separately provisioned native executable and command lease verified. No credentials or inference were used. Subsequent changes preserve this npm asset layout. - The immutable Daytona prerequisite image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`, built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous Cloud controller image was `ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`, built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded by the latest image above. Its EC2 build verified provider-pack access under an unrelated unprivileged UID. - Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed full Grok onboarding with the correct `grok-4.7` model, saved credential delivery, and pinned Daytona execution. A browser-created task read context and asked the structured human question. After a controller restart, answering the persisted question resumed the same native provider session, saved the requested document, and completed the task. Actual tool outcomes and durable state agree: one question and one document revision. The two successful turns took 42.7s and 63.1s; usage and cost were not reported. - Restricted policy returned the expected `approval_required` outcome. Functional staging tests explicitly selected `approve-all`; controller authorization and governed approvals remain enforced. Temporary board CLI access was revoked and verified rejected (HTTP 401), and the disposable onboarding agent was paused. Failures remain retained: the pre-fix continuation failure (its task remains blocked; the passing final task is fresh), the original Cloud provider-pack permission failure, the expected restricted-policy denial, the superseded npm staging failure, and an earlier monolithic CI infrastructure timeout. Browser CI exposed a project alias/form race; the final stack uses master's stronger draft-preservation fix and all browser shards pass. Historical full subscription/API protocol and Product rosters retain their original source revisions and do not qualify this packaging revision. No local Docker or Rust build was used. ## Risks The branch includes master’s draft-preservation fix for project URL aliases. It keeps the same project’s edit form mounted and clears prior data when the project or company changes. Custom sandboxes and local execution hosts must provision the pinned binary before Grok starts. Missing, changed, unsupported-platform, and symlinked executables fail admission. The new builtin profile cannot resume sessions created with the former private-package profile. Existing Claude/Codex npm bridge profiles retain their package pins. Grok restricted modes preserve the selected policy but cannot automatically admit Paperclip calls: ACP permission metadata does not independently bind tool authority, so those calls stop with `approval_required`. New Grok configurations default to `approve-all`, including API configurations that omit the mode. Existing explicitly restricted configurations remain restricted; controller authorization and governed approvals remain enforced. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b648d8cdda |
fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path > - Paperclip manages AI agents and their provider connections. > - Product E2E checks real tasks through the browser, server, and runner. > - Grok qualification needs separate API-key and subscription evidence. > - Product subscription tests and direct Grok protocol evals need explicit credential delivery. > - This change supplies each credential only to its selected profile and prepares the pinned binary. > - Maintainer authorization and protected-environment gates remain required. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13882. The Grok feature branch has a manual subscription qualification profile. The trusted master workflow must admit its selected credential and prepare the same verified binary and artifact verifier as the API profile. Direct protocol evals also need the selected xAI key and pinned Grok binary. These prerequisites do not register or schedule the new profiles on master. ## What Changed - Deliver `GROK_AUTH_JSON` from the protected paid environment only when the selected profile requests that credential. - Install the checksum-verified Grok binary for the local subscription profile. - Prepare the pinned artifact verifier for the manual subscription suite. - Extend workflow security assertions to cover the new credential and profile. - Add the ACPX Grok credential mapping to the trusted-master catalog, then deliver only the selected `XAI_API_KEY` to direct protocol cells and install the target’s checksum-verified Grok binary before packaging. - Allow a direct-protocol concurrency override from two cases up to the existing configured ceiling; it can only lower concurrency. - Document the Grok protocol workflow and its API-only credential boundary. - Render missing LLM usage and cost as Unavailable, and label partial observations with coverage. Preserve raw records, grades, and actual zero costs. - Preserve measured campaign source metadata during report regeneration instead of inheriting the renderer checkout or CI event; skip empty legacy source records when recovering older provenance. ## Verification - Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54 reported checks successful, two intentional skips, Greptile 5/5, and zero unresolved review threads. [CI run](https://github.com/paperclipai/paperclip/actions/runs/35890978288). - After merging current master, all 17 workflow security/image tests and 23 catalog/workflow policy tests passed. The trusted catalog also generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per shard. The new policy tests execute the concurrency guard against valid, out-of-range, and malformed values. - The Grok branch separately passed 450 Product harness unit tests, including private company credential staging, cleanup, and token-fragment redaction. - The fresh-login native subscription smoke passed three repetitions of tool execution, session resume, restrictive permissions, and cleanup. These are setup evidence; full subscription Product qualification remains pending. - All 72 focused report/billing/history/catalog tests and the Product harness typecheck passed for the report-display change. The initial sandbox run could not open the tsx IPC socket; the permitted rerun passed. A zero-provider-call replay of the actual 16-cell Grok campaign preserved all result records, grades, timing, and source provenance while correcting missing usage labels. - Review the thirteen-file diff. Provider credentials still enter only the selected paid-test step; default-branch, numeric-actor, and environment restrictions are unchanged. ## Risks This admits a refreshable subscription credential to explicitly selected trusted tests. Store it only in `runner-e2e-paid`, use a test login, and remove it after qualification. Unselected profiles receive an empty value. Pull requests cannot trigger the paid workflow. This PR changes no fleet admission or actor allowlist. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
60c7c9cd1a |
fix(runner-e2e): pass verified lock digest to Daytona image build (#13876)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Product E2E campaigns test the native runner in local and Daytona environments. > - Each campaign resolves one target lockfile and verifies its downloaded artifact. > - The Daytona image job did not pass that artifact digest to Docker. > - Docker used an older default digest and stopped before any selected task ran. > - This pull request passes and validates the campaign digest at the image build boundary. > - The image keeps its checksum check and frozen package installation. ## Linked Issues or Issue Description **What happened?** The merged-master [qualification campaign](https://github.com/paperclipai/paperclip/actions/runs/35863582409) stopped in the Daytona image build. The resolved target lock digest was `e0c928a494f90ddad3c00791e83f09315ee8c82df2a0418a809dcf93649a8ab3`. Docker used its default digest, `57b298aceebc48bb94ea0593347348256475da7b2fddb77025d9e57cc8759420`. The checksum check rejected the mismatch. All 12 selected model cells were skipped. This follows the image provenance work in #13814. **Expected behavior** The image build must check the same lockfile artifact that the campaign restored and verified. A changed lockfile must still fail the checksum check. **Steps to reproduce** 1. Start a Product E2E campaign with a Daytona cell on master `7944ed3d976d1a7cc26a2d0cee51f227f3542084`. 2. Resolve a target lockfile whose digest differs from the Dockerfile default. 3. Observe the provider-pack image stage reject the lockfile before model execution. **Paperclip version or commit** `7944ed3d976d1a7cc26a2d0cee51f227f3542084`. **Deployment mode** GitHub Actions Product E2E campaign with a Daytona image build. **Install method** Built from source with the campaign lockfile artifact. **Agent adapter(s) involved** Native Codex and ACPX Claude cells were selected. No model cell ran in this failed campaign. **Database mode** Not involved. The failure occurs during image creation. **Access context** The authorized default-branch paid workflow. The build receives no provider credentials. ## What Changed - Read the image checksum from the existing target-lock job output. - Require a 64-character lowercase hexadecimal digest before image inspection or build. - Pass the digest as the existing Docker build argument. - Add regression checks and document the campaign checksum handoff. ## Verification - The Daytona image regression fails with the original workflow and passes with the fix. - All six Daytona image contract tests pass. - All 450 Product E2E unit tests pass. - Product E2E typecheck passes. - Actionlint passes for the changed workflow. - A context-shaped resolution probe preserves the downloaded lockfile bytes and digest. - All latest-head CI checks passed on `67d41fd9439b2a9a809ddb05765f8617585072c5` ([run](https://github.com/paperclipai/paperclip/actions/runs/35865739359)). - Greptile gave 5/5 on this head; its test-scoping comment is addressed and resolved. - A hosted Daytona image rebuild and the three remote qualification cells remain pending after merge. ## Risks The campaign digest comes from the existing trusted target-lock job. The restored artifact checks, Docker checksum check, frozen install, content identity, image signing, and verification remain in place. The standalone Docker default remains available. This change does not alter task behavior, prompts, credentials, or dependency versions. ## Model Used OpenAI `gpt-6-astra` through Codex performed diagnosis and review with code execution tools. OpenAI `gpt-5.6-luna` assisted with investigation, implementation, and verification. Context window limits are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ac854a9b59 |
ci: build Daytona eval images on the authorized fleet (#13847)
## Thinking Path > - Paperclip manages AI agents and their work. > - Product E2E campaigns test the browser, server, and runner together. > - The trusted workflow selects an authorized execution fleet. > - Daytona image builds still use a fixed GitHub-hosted runner. > - This pull request applies the existing fleet selection to image builds. > - Campaign builds and tests then use the configured EC2 fleet. ## Linked Issues or Issue Description Refs: https://github.com/paperclipai/paperclip/pull/13845 **What existing behavior does this improve?** The location of Daytona Product E2E image builds. **Current behavior** The image job uses ubuntu-latest even when RUNNER_E2E_AWS_ENABLED selects EC2 for the rest of the campaign. **Proposed behavior** Use the existing authorized runner output for the image job. Preserve the GitHub-hosted fallback when the EC2 switch is off. ## What Changed - Route Daytona image builds through the existing authorized runner selection. - Assert that the image job depends on authorization and receives no provider credentials. - Document the build routing and credential boundary. ## Verification - Product E2E workflow security: 11 tests passed. - git diff --check passed. - Reviewed actor checks, workflow triggers, target commit selection, signing, and package permissions. They are unchanged. - Required CI checks pass on the latest head. The first CI attempt had failures in unchanged chat tests; one diagnostic rerun passed. Both attempts remain in Actions. - Greptile reviewed the latest head at 5/5 with no inline findings. - Live EC2 image verification follows after this trusted workflow change is merged. No local Docker execution. ## Risks EC2 image builds depend on the fleet having working Docker and sufficient disk space. Existing authorization, image digest checks, signing, and the GitHub-hosted fallback remain in place. No provider secrets are added to the image job. No application or database changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool execution. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
950ccb8eef |
ci: enable Grok qualification in the trusted paid workflow (#13845)
## Thinking Path > - Paperclip manages AI agents and their work. > - Product E2E tests verify tasks through the browser, server, and runner. > - These paid tests use a trusted workflow from master and an isolated target commit. > - The Grok target branch selects XAI_API_KEY, but the trusted workflow does not deliver that credential. > - Its artifact test also needs the pinned Python verifier before execution. > - This pull request adds both bindings inside the existing paid boundary. > - The tests can then run on the configured EC2 fleet without laptop Docker. ## Linked Issues or Issue Description Related evaluation infrastructure: https://github.com/paperclipai/paperclip/pull/11297. No duplicate Grok paid-workflow change was found. **What existing behavior does this improve?** Branch-targeted Grok Product E2E qualification in the existing paid workflow. **Current behavior** Grok cells cannot receive their selected API credential. The Grok build-revise case also misses the artifact-verifier setup step. **Proposed behavior** Deliver XAI_API_KEY only when the selected matrix credential is XAI_API_KEY. Prepare the existing pinned verifier for the Grok qualification suite. **Reason and benefit** Run the controller, browser, runner and artifact checks on the EC2 fleet. Preserve default-branch workflow authorization and protected environment secret access. ## What Changed - Bind the selected XAI credential only in the paid test step. - Install the checksum-verified Grok binary for local cells before provider access. - Include Grok qualification in the existing pinned artifact-verifier preparation. - Add an optional max_parallel input that can only lower the configured campaign concurrency. Use 1 for the Grok test key. - Extend security assertions and document setup. ## Verification - Ran the Product E2E workflow-security tests: 11 passed. - Checked the diff for whitespace errors. - Reviewed credential selection, setup ordering, numeric actor gates, target commit pinning, and trusted report checkout. - Live Grok execution follows after this workflow is available on master. This PR does not claim completed Grok qualification. ## Risks The paid test step can use the selected XAI credential and incur provider charges. The credential remains in runner-e2e-paid and is absent from setup, build, and reporting jobs. The default-branch gate and existing environment restrictions remain in place. No database migration or product behavior changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool execution. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3790ca2f13 |
fix(runner): repair approval and Stop races and eval infrastructure (#13750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Runner tasks must continue after approval and stop when the user presses Stop. > - Live evals found races at approval delivery and provider startup. > - Browser readiness and CI setup errors also hid the actual task results. > - This pull request fixes those races and the related test infrastructure. > - Regression tests and saved live reports show which cases now pass. ## Linked Issues or Issue Description Companion eval definitions PR: https://github.com/paperclipai/paperclip-evals/pull/25 (AgentCore paused and provider/environment infrastructure). Related: #13741 now supplies the late-startup Stop fence and warm-attachment recovery; this PR retains that fence and extends startup tracking and regression coverage to both native backend paths. #13539 introduced queued approvals during active runs. #13738 fixes child assignment, task replies, and warm process continuity and is already in the base. #13291 concerns automatic continuation of interrupted legacy sandbox runs; this PR fixes native startup cancellation and does not change that recovery policy. **What happened?** An accepted service approval could wait after its source run stopped. Stop could return success before the provider handle existed. Work could then start after Stop, or a cancelled run could be recorded as failed. Some E2E tests also failed on unloaded browser content or irrelevant reply wording. Runner CI could fail before model work because of dependency or sandbox setup. **Expected behavior** Deliver each settled approval once after its source run stops. Do not start work after an acknowledged Stop. Preserve the audited cancellation. Test the intended product behavior with a ready browser and verified runtime dependencies. **Steps to reproduce** 1. Approve a service request while its source run is active. Let the run finish. Check that its result starts one continuation. 2. Delay provider startup. Press Stop before its handle is available. Check cancellation, then submit `/new`. 3. Run the browser, warm-workspace, and Stop-and-redirect cases from the linked report. **Paperclip version or commit** The branch includes master at `9d19f98b5`. The report records the original source for each focused attempt. **Deployment mode** Isolated local development instances and disposable Daytona sandboxes. ## What Changed - Deliver settled tool-action results for the exact company and source run during final cleanup. Keep the existing idempotent receipt and periodic recovery sweep. - Wait for startup to hand off its provider handle before acknowledging Stop. Reject first-turn admission after cancellation. Preserve a matching audited pending or acknowledged cancellation. - Wait for mounted task history and connector controls in browser tests. Record failure evidence. Grade workspace contents and process continuity separately from exact reply wording. Require each warm-turn marker once and in order, allowing surrounding prose. - Stop-and-redirect now checks that the source file exists and work is active before Stop. - Resolve target dependency locks in an uncredentialed CI job. Verify the lock artifact hash. Keep orchestration and publication on the trusted workflow revision. - Materialize the pinned OpenCode executable and configure the exact Codex executable's user-namespace profile before provider credentials are available. - Compress Daytona directory uploads with gzip. Preserve files, executable modes, symlinks, empty directories, and confinement checks. - Classify file-transfer RPC deadlines as infrastructure. Keep unrelated runner RPC failures visible. ## Verification - [Focused live report with screenshots and original attempts](https://pages.paperclip.ing/runner-reliability-20260921/): 14 of 15 selected Product E2E cases pass across the recorded revisions. Claude and Codex Stop → `/new`, Claude service approval, delegation, both hiring/reuse cases, and native Daytona warm continuity pass. - Two credentialed Runner smoke cases pass. These are not full protocol coverage. - E2E harness after the master merge: 429 tests pass. E2E and server TypeScript checks pass. - Daytona plugin: 239 tests pass, 6 skipped. Plugin TypeScript build passes. The compression test fails against the old code and passes with the change. - Runner backend/runtime regression group: 161 tests pass. Cancellation/startup selection: 26 tests pass. Approval delivery: 34 real-database tests pass. - Workflow security: 7 tests pass. Both edited workflows pass actionlint. Runner TypeScript and Rust builds pass. - After merging master, all 389 native executor tests pass, including both native backend paths and late startup after the Stop deadline. - Post-merge `pnpm -r typecheck` and `pnpm build` pass. The monolithic local `pnpm test:run` was interrupted to integrate master and is inconclusive. The [hosted CI test partitions](https://github.com/paperclipai/paperclip/actions/runs/35620461738) pass on `50a3e43822bcba1e0d07b1b45b0be91cbf9312da`. An unchanged sandbox callback schema test initially received HTTP 503. It passed five isolated local runs, its full local test file, and one failed-job CI retry. No assertion was weakened. ## Risks - Stop can wait for the bounded startup handoff. If it cannot settle, the existing pending-recovery state remains instead of a false acknowledgement. - Immediate approval delivery must remain idempotent across cleanup and recovery sweeps. Tests cover duplicate delivery and company/run boundaries. - The workflow changes still need hosted Linux verification. They retain the trusted workflow and credential boundaries. - Gzip reduces the observed provider upload from about 1.8 GB to 663 MB. It does not yet fix the remaining Claude Daytona transfer timeout. That recovery test never reached Claude, so recovery remains unverified. Use a matching image with the verified provider package preinstalled for the next recovery test; retain cold-upload coverage separately. - The report preserves diagnostic runs with missing source metadata and marks them as such. It does not claim a new full-suite pass. - This PR adds no new prompt policy or historical status reconciliation. ## Model Used OpenAI GPT-6 through Codex performed the primary implementation and review. The exact primary backend model ID is not exposed in this session. OpenAI `gpt-5.6-luna` assisted with bounded infrastructure work and verification. The agents used repository tools, code execution, and browser tests. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 <noreply@openai.com> |
||
|
|
4cfc0e3f6c |
fix(runner-e2e): provision complete worker prerequisites (#13674)
## Thinking Path
> - Paperclip manages work performed by AI agents.
> - Runner E2E tests verify that work in local and remote environments.
> - The full catalog exposed setup failures before agents could execute
tasks.
> - ACPX Codex missed sandbox provisioning, and new artifact cases
missed Docker preparation.
> - A host Claude version probe also stopped native Claude stories that
use a packaged provider.
> - This pull request repairs trusted worker setup and keeps its
selectors covered by catalog tests.
> - The resulting reruns can measure behavior instead of missing
prerequisites.
## Linked Issues or Issue Description
Follow-up to #13655. Full-catalog campaign:
https://github.com/paperclipai/paperclip/actions/runs/35417932353.
**What happened?**
ACPX Codex could not start its sandbox. New artifact stories missed
Docker preparation. Native Claude stories tried to spawn an unrelated
host CLI. The report job failed on trusted lockfile drift.
**What did you expect to happen?**
Prepare each worker's required capabilities before paid execution and
publish the retained results from trusted code.
**Steps to reproduce**
Run local ACPX Codex, everyday agent-review-handoff, or native Claude
everyday cells in the full-stack workflow at
|
||
|
|
43acbcc398 |
fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task state to provider sessions. > - Follow-up turns must retain provider memory and carry new user direction. > - Lost session IDs caused repeated context and extra input tokens. > - Native question answers and approval races could leave valid work blocked. > - This pull request repairs those paths and adds regression coverage. > - Agents can continue accepted work without repeating the conversation or losing the user's answer. ## Linked Issues or Issue Description Refs #13574. That merged PR shortened continuation prompts and moved question instructions into tool documentation. This change preserves sessions and fixes failures exposed by broader testing. Related runtime work: #13408 and #13410. **What happened?** Native follow-up turns could lose the provider session ID. Completion guidance could replace the original task with its latest comment. Claude native questions could remain pending after the user answered. Approval during a running tool call could suspend the run before the tool response arrived. Onboarding and chat handoff instructions also caused repeated planning or missing plan documents. **Expected behavior** Reuse a valid provider session. Send only new events when that session already has the history. Preserve the task requirements and apply later user direction. Store the question answer and deliver it to the waiting run. Finish governed tool responses before suspending. Execute the accepted plan without asking for the same approval again. **Steps to reproduce** Run the continuation, local-session-integrity, first-task, and agent-chat suites with native Codex and Claude. Include provider-question-bridge, accept-while-running, and plan-handoff. **Paperclip version or commit** This branch is based on master |
||
|
|
4577d10029 |
fix: prepare everyday artifact and Codex sandbox prerequisites in CI (#13516)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The runner E2E workflow executes paid everyday workflow stories on disposable CI hosts > - The everyday artifact oracle requires a pinned Python image and fails closed when it is absent > - Fresh CI hosts did not prepare this image before the paid cell, so project stories failed during preflight > - This pull request prepares and verifies the pinned image before the affected everyday cells > - The benefit is reliable artifact isolation checks on fresh trusted CI hosts ## Linked Issues or Issue Description **What happened?** Fresh trusted CI runners did not have the pinned Python artifact oracle image. **Expected behavior:** The workflow prepares and verifies the pinned image before an everyday project story starts. **Steps to reproduce:** Run an everyday project story on a fresh CI host without the image cached. The `everyday-artifact.py --preflight` check fails before task creation. **Paperclip version or commit:** `master` at `bd51f157e`. **Deployment mode:** Other: GitHub Actions trusted paid workflow. ## What Changed - Add a matrix-gated CI step for everyday project and recovery cells. - Check Docker, pull the fixed digest with bounded timeouts, and verify the exact repo digest. - Apply the existing Codex sandbox preparation to both native Codex profiles, including the mini profile. - Add workflow security assertions for ordering, condition, digest, timeouts, and secret isolation. - Document that CI prepares the pinned oracle image. - Check provisioning eligibility against every catalog cell, and scope the Daytona registry inspection assertion to the Daytona image job. ## Verification - `pnpm test:e2e:runner:unit` — 24 files and 222 tests passed. - `pnpm test:e2e:runner:typecheck` — passed. - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts workflow-security.test.ts` — 10 tests passed. - Python artifact oracle calibration — 12/12 passed. - `git diff --check` — passed. ## Risks Low risk. The image step runs only for everyday cells that execute the artifact preflight. The Codex setup now covers both native Codex profiles. It uses a fixed public image digest and has no provider credentials. ## Model Used OpenAI gpt-5.6-luna (implementation subagent) and gpt-6-astra (review fixes and orchestration), using code execution and repository tools. Context window size is not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Paperclip <paperclip@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9ecd93a54d |
test(e2e): link runner campaign summaries (#12927)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip uses a paid full-stack campaign to verify runner behavior across providers and environments. > - The campaign already creates an interactive report, workflow logs, and retained evidence artifacts. > - The merge job summary shows result totals but does not link to those resources. > - Reviewers must search several workflow jobs and artifacts to find the executed cells. > - This pull request adds direct and safe links to the exact campaign, each cell, the workflow logs, and the artifacts. > - The benefit is that a reviewer can inspect a result from the Actions summary with one click. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the `Merge and enforce campaign result` summary in the `Runner Full-Stack E2E` workflow. **Subsystem affected** The runner E2E report generator and its GitHub Actions workflow are affected. **Current behavior** The summary lists each selected cell and its result. It does not link to the published campaign report, the workflow logs, or the evidence artifacts. **Proposed behavior** The summary includes a `View results` section. It links to the exact immutable campaign report, the workflow logs, and the artifacts. Each cell name links to its stable section in the campaign report. **Reason and benefit** The current summary does not show reviewers where to inspect the run. Direct links make the result evidence discoverable without manual URL construction or artifact searches. **Breaking changes** None. This change only adds links and stable HTML anchors to existing report output. **Additional context** Related: #12904. The cited successful campaign is [run 34026735033](https://github.com/paperclipai/paperclip/actions/runs/34026735033). ## What Changed - Add a safe URL builder for public campaign, workflow, and artifact links. - Add a `View results` section to the GitHub Actions campaign summary. - Link each summary table cell to its exact section in the immutable campaign report. - Add stable execution anchors to the generated dashboard. - Reject non-HTTPS, credential-bearing, malformed, and ambiguous link destinations. - Document the new links and their retention or publication timing. ## Verification - `pnpm test:e2e:runner:unit` — 116 tests passed. - `pnpm test:e2e:runner:typecheck` — passed. - `pnpm typecheck` — passed, including migration safety. - `pnpm build` — passed. - `pnpm exec prettier --check ...` for all changed files — passed. - `git diff --check origin/master...HEAD` — passed. - The full local server suite also ran. One unrelated macOS workspace-runtime file passed 157 tests and failed 4 existing path and port assumptions. Two failures compare `/var` with `/private/var`. Two failures cannot reserve a port outside a hard-coded range. This PR does not change that file or its dependencies. ## Risks - The immutable campaign link becomes available after the history publisher completes. The workflow and artifact links remain available while publication runs. - The artifact link requires GitHub access and follows the existing 30-day retention period. - Invalid configured URLs are omitted instead of being rendered into the summary. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex desktop agent with GPT-5. The runtime does not expose the context-window size. The agent used repository inspection, agentic reasoning, code execution, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/no-internal-issue-references`, `fix/sandbox-secret-resolution`, `feat/adapter-retry-backoff`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
87832c48fd |
feat(runner-e2e): publish declared screenshots (#12895)
## Thinking Path > - Paperclip uses runner end-to-end reports to compare agent profiles and execution environments > - The report dashboard shows each reviewed final-state screenshot as a thumbnail and gallery item > - The public history publisher removed all per-attempt images before it regenerated the dashboard > - Therefore the public dashboard had the new layout but could not show the screenshots from the run > - The publisher needs a narrow rule that keeps only screenshots from the exact live fixture issue route > - This pull request keeps those trusted PNG files in every future public S3 and GitHub Pages report > - The benefit is that each future report can show its screenshot gallery without exposing logs, traces, videos, archives, arbitrary images, or generated report trees ## Linked Issues or Issue Description **What happened?** The runner E2E job captured final-state screenshots in its private artifact. The public S3 and GitHub Pages publication step removed those screenshots before it regenerated the dashboard. As a result, the public report showed the new dashboard controls but no screenshot thumbnails or gallery items. **Expected behavior** Each future public runner E2E report must include reviewed PNG screenshots from the live fixture issue. Other captures and active or unsafe evidence must stay private. **Steps to reproduce** 1. Run the runner full-stack E2E workflow on `master` before this change. 2. Open the private `runner-e2e-report-*` artifact and confirm that it contains per-attempt PNG screenshots. 3. Open the public campaign URL and confirm that the dashboard has no screenshot gallery items. **Paperclip version or commit** The issue was reproduced on commit `64d8929`, after the report design change in PR #12889. **Deployment mode** GitHub Actions with the public S3 and GitHub Pages report publishers. Related design work: Refs #12889. ## What Changed - Mark screenshots from the exact server-created live fixture issue route with `public-runner-fixture`. - Keep marked PNG files in both the S3 history bundle and the GitHub Pages bundle. - Keep captures from other issue routes, sensitive routes, and external origins private. - Bind public files to the normalized execution ID, attempt, and safe PNG base name. - Validate every retained image with the existing PNG signature and 12 MiB size checks. - Skip missing-artifact sentinel results with attempt `0` when they have no public screenshots. - Continue to remove unmarked images, videos, traces, archives, generated HTML reports, and other private evidence. - Update publisher tests, workflow checks, report copy, and the public-evidence security documentation. ## Verification - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts tests/runner-e2e/history.test.ts tests/runner-e2e/report.test.ts` - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts tests/runner-e2e/workflow-security.test.ts -t "uses environment-scoped OIDC"` - `pnpm test:e2e:runner:typecheck` - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/runner-e2e-dashboard.spec.ts` - `pnpm -r typecheck` - `pnpm build` - Regenerated the dashboard from retained evidence for Actions run `33968240659` without a paid matrix rerun. The public-stage proof contained 121 screenshot gallery items and thumbnail frames, with zero generated HTML report files. The trusted-fixture marker and route gate have separate focused tests. - All pull request CI checks pass on commit `ccd2b49e1`. ## Risks - This change intentionally makes marked fixture screenshots public at the campaign URL. A screenshot can show data that a raw-byte secret scan cannot detect. - The capture helper marks a screenshot only on the exact loopback issue route for the fixture that the harness created. A different issue, sensitive page, or external origin stays private. - The publisher also requires the marker, a safe normalized path, a valid PNG signature, and the size limit. - The change does not publish videos, logs, traces, archives, arbitrary images, or generated browser report trees. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5.6-sol`, with high reasoning, repository tool use, shell execution, browser inspection, and GitHub CLI access. The working context was the Codex desktop task context. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8430bd897f |
ci: reuse trusted cache for Daytona images (#12862)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The full-stack runner campaign checks local and Daytona runner behavior. > - A Daytona image content miss starts a cold multi-stage Docker build. > - Stable dependency and agent CLI layers take most of the image build time. > - Development targets must not write shared cache state. > - This pull request adds a registry cache with a default-branch write gate. > - It also puts volatile source inputs after stable install layers. > - The benefit is a shorter Daytona image build without weaker secret isolation. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the Daytona runner image stage in the full-stack E2E workflow. **Subsystem affected** The GitHub Actions runner E2E workflow and its Daytona Docker image are affected. **Current behavior** Each new Daytona image content ID starts with an empty BuildKit cache. A runner source change also invalidates dependency and agent CLI install layers because volatile inputs occur before those layers. **Proposed behavior** All authorized campaigns can read one GHCR BuildKit cache. Only a campaign whose target ref is the repository default branch can update that cache. The Dockerfile installs dependencies and agent CLIs before it consumes volatile runner source or revision metadata. **Reason and benefit** The paid runner matrix spends several minutes building the image before any selected cell can start. Cache reuse removes repeated stable setup work and makes focused Daytona iterations faster. **Breaking changes** None. The immutable content tag, digest inspection, Cosign signature, image labels, pinned base images, and provider credential boundary stay unchanged. ## What Changed - Read a registry-backed BuildKit cache for Daytona image content misses. - Export the cache only when the resolved target ref is the default branch. - Keep provider credentials outside the image build and cache. - Install provider-pack dependencies before runner source is copied. - Keep expensive agent CLI installs before source revision metadata. - Add workflow and Docker layer-order contract checks. ## Verification - `prettier --write .github/workflows/runner-full-stack-e2e.yml tests/runner-e2e/daytona-image.test.ts tests/runner-e2e/workflow-security.test.ts` - `actionlint .github/workflows/runner-full-stack-e2e.yml` - `git diff --check` - I did not run a test suite or Docker image build locally. The requested iteration policy reserves those checks for GitHub Actions. ## Risks Low risk. BuildKit can use a cache record only when its content key matches the build instruction and input. Development targets have read-only cache access. The cache contains public source and build outputs, but it does not receive provider credentials or the GitHub token as Docker build inputs. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bcc6fe7a44 |
fix(runner): restore multi-turn remote sessions (#12840)
## Thinking Path > - Paperclip manages AI agents and their work. > - The runner executes agent turns on local and remote providers. > - A remote per-turn session must save its state before Paperclip releases its sandbox. > - The session runtime returned after 100 milliseconds while the remote checkpoint still ran. > - The next turn also checked the local state path instead of the verified remote backup. > - This pull request waits for the bounded remote close and accepts only a verified suspended backup. > - The benefit is reliable multi-turn execution without weaker identity checks. ## Linked Issues or Issue Description **What happened?** A successful remote agent turn released its sandbox before the runner saved the verified continuation backup. The next turn failed with `runner_state_identity_mismatch`. **Expected behavior** Paperclip must finish the bounded remote checkpoint before it releases the sandbox. A later turn must validate and restore the digest-matched suspended backup. **Steps to reproduce** 1. Run a native ACPX Claude Plan test in a non-reusable Daytona sandbox. 2. Reject the first plan to start a second turn. 3. Observe that the second turn fails before provider execution. **Paperclip version or commit** The failure reproduced at `13775a90b078ff64872f50961ea1b83d575e7bc6`. **Deployment mode** GitHub Actions with a Daytona sandbox. ## What Changed - Wait for the internally bounded remote runner close and checkpoint before the host returns. - Preserve the existing short cleanup bound for other providers. - Validate remote continuation lifecycle from a complete digest-verified backup when local runner state is absent. - Keep corrupt, non-suspended, mismatched, and unverified state fail-closed. - Make native Plan completion and accepted-Plan wake prompts deterministic. ## Verification - A prior 45-cell local campaign passed 44 cells. The only failure was the OpenCode Plan prompt variance fixed here. - A focused OpenCode local Plan rerun passed. - ACPX Claude Daytona message and question cells passed. - Focused regressions cover delayed checkpoint close and verified remote backup lifecycle. - GitHub Build and the focused ACPX Claude Daytona Plan cell will validate this exact head. ## Risks Remote runnerd sessions now wait for their internally bounded close/checkpoint path before returning; generic provider cleanup retains the existing 100 millisecond bound. Durable run success still cannot be reversed. The environment release guard still blocks sandbox destruction when no verified backup stamp exists. ## Model Used OpenAI Codex, GPT-5.6, extended reasoning, with code execution and GitHub Actions inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal task id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open findings - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4ef6155aae |
ci: harden paid runner browser and lock repair (#12829)
## Thinking Path Paid cells now reuse the AWS image's system Chrome, but Playwright video recording still resolves its revision-pinned FFmpeg helper from the Playwright cache. Run 33875618534 proved Chrome qualification succeeds and then failed before provider startup because that helper was absent. The same run also exposed that generic lock repair can churn unrelated package platform metadata, so the automated repair paths need resolution-only regeneration rather than lockfile-only metadata refresh. ## What Changed - install Playwright FFmpeg only on the AWS/system-Chrome path - retry the small helper installation up to three times before provider secrets are exposed - keep the GitHub-hosted Chromium fallback unchanged - bind static coverage to the exact FFmpeg step block and its pre-secret ordering - add pnpm `--resolution-only` to all four automated lock-repair paths while retaining full transitive resolution - require resolution-only repair in the shared workflow regression The actual generated lockfile correction remains bot-owned by PR #12828 and is intentionally not committed here. ## Verification - `node --test .github/scripts/tests/lockfile-refresh-workflows.test.mjs` - `actionlint -ignore SC2012` on all modified workflows - focused Prettier checks - `git diff --check` - prior run 33875618534: system Chrome 151 qualified; missing Playwright FFmpeg was the sole cell startup failure ## Risks Low. The new network operation is limited to Playwright's pinned FFmpeg payload, happens before paid credentials are exposed, and leaves the hosted-runner path unchanged. Resolution-only is still a full dependency-resolution pass, unlike lockfile-only, while avoiding unrelated current-platform metadata churn. ## Model Used GPT-5 |
||
|
|
af3023f1e3 |
fix(runner): repair paid provider startup paths (#12769)
## Thinking Path > - Paperclip manages AI agents that perform work. > - Paperclip Runner connects durable task runs to local provider processes. > - The full-stack paid matrix exposed failures after the runner integrity repair. > - Verified JavaScript entrypoints lost their relative module graph when Linux executed them through descriptor paths. > - Returned provider startup errors also remained pending and became indeterminate after recovery. > - Sparse Codex tool lifecycle events lost the `write_document` identity before task transcript projection. > - This pull request repairs those three boundaries and makes the structured-question fixture deterministic. > - The benefit is repeatable provider startup, exact failure replay, and correct inline Plan placement. ## Linked Issues or Issue Description Refs #12721 and #12700. **What happened?** The paid runner matrix failed ACPX and OpenCode startup before provider session creation. The runner journal then replaced the original startup error with an indeterminate recovery result. Native Codex saved a Plan but rendered it only as a fallback card. A legacy Claude waiting reply could also echo the reserved terminal marker before the answer arrived. **Expected behavior** Verified JavaScript providers must start from immutable descriptor-backed artifacts. Returned startup failures must persist as terminal failed command results. Native tool lifecycle updates must preserve the `write_document` boundary. Pre-answer fixture output must not contain the reserved terminal marker. **Steps to reproduce** 1. Run the local provider cells in the Runner Full-Stack E2E workflow. 2. Observe ACPX and OpenCode fail during `session.open` before provider execution. 3. Observe recovery report `execution_indeterminate` instead of the original startup error. 4. Run the native Codex Plan cell and observe the fallback Plan card after the tool activity row. 5. Run the legacy Claude structured-question resume cell and observe an early marker echo in waiting prose. **Paperclip version or commit** `0f9452101740835ce0b1488a204bf48acd5bafc3` **Deployment mode** Local development with the paid GitHub Actions acceptance workflow. ## What Changed - Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM entrypoints before hashing and verified descriptor launch. - Anchor ACPX dynamic provider package resolution at a controller-derived provider-pack root and keep that root out of the provider child environment. - Persist executor-returned startup errors as redacted durable failed command results while retaining indeterminate recovery for true process death. - Coalesce sparse native tool items by stable ID so a late `write_document` name, input, and result reach the transcript boundary once. - Forbid the structured-question fixture from spelling or announcing its reserved terminal marker before the user answers. ## Verification - Rust and TypeScript regression tests cover durable failed replay, true crash ambiguity, bundle closure, package-root derivation, environment filtering, exact Codex tool lifecycle coalescing, and prompt determinism. - Local execution is intentionally limited to formatters and static diff checks. GitHub Actions will run tests, type checks, builds, and security checks. - After ordinary CI is green, scoped paid cells will validate one ACPX launch, one OpenCode launch, native Codex Plan projection, and legacy Claude structured resume before a complete matrix rerun. - Prior failing matrix: https://github.com/paperclipai/paperclip/actions/runs/33682434315 ## Risks - Bundling changes the bytes covered by provider launch hashes. Provider-pack generation already hashes the final built files. - ACPX still loads qualified provider packages dynamically. The controller supplies a normalized package root, while existing version, digest, path, and descriptor checks remain active. - Durable `failed` is terminal. Replays return the same redacted result and do not execute the provider effect twice. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5 with agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions coordination. The exact deployed snapshot and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked related public work or described the bug in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] No documentation change is required for this runtime repair - [x] I have considered and documented the risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a0a78ee609 |
ci: bootstrap Node before pnpm setup (#12808)
## Thinking Path The trusted workflows currently invoke `pnpm/action-setup@v6` before installing the repository Node version. On hosts whose ambient Node is older than 22.13, the action downloads standalone `@pnpm/exe`, which has repeatedly taken several minutes. Supplying Node 24 first lets the same pinned pnpm action use its normal Node-backed path. ## What Changed - install Node 24 before every trusted `pnpm/action-setup` invocation - preserve the existing pnpm-store cache setup, pinned actions, telemetry suppression, conditions, and secret boundaries - enforce ordering, modern Node, condition parity, and cache counts in the workflow security contract ## Verification - workflow security: 7/7 - Prettier - actionlint (excluding one pre-existing SC2129 in an untouched Daytona shell block) - `git diff --check` ## Risks Low. Product code, providers, paid-runner selection, credentials, and pnpm version are unchanged. Jobs that restore pnpm cache run `setup-node` a second time after pnpm becomes available; the first setup is deliberately cache-free. ## Model Used Codex (GPT-5) |
||
|
|
18ea965442 |
ci(runner): stamp paid target provenance (#12805)
## Thinking Path Trusted workflow-dispatch runs execute an authorized target SHA, but GitHub context still describes the default-branch workflow revision. Retained paid results and artifact names were therefore labeling target-branch executions as master. The workflow must explicitly pass its authorized target coordinates to target code and trusted reporting. ## What Changed - emit the canonical authorized target ref alongside the immutable target SHA - pass those coordinates to paid cells and the trusted report - name shared build/provider artifacts with the target SHA rather than workflow SHA - add workflow-security coverage for all trusted provenance wiring ## Verification - focused workflow-security tests: 6/6 passed - Prettier and git diff checks passed - run 33823252706 independently proved the pre-fix defect: functionally green target cells were retained as master SHA |
||
|
|
0ad180b85f |
ci(runner): skip bootstrap registry telemetry (#12797)
## Thinking Path Every trusted PR and paid-workflow job invokes the pinned pnpm setup action. Its internal npm install is currently waiting four to seven minutes on npm audit telemetry before any Paperclip or provider code runs. Audit, funding, and update notifications are not integrity controls for this action; its committed lockfile still verifies installed package bytes. ## What Changed - disable npm audit, funding, and update-notifier telemetry narrowly on all seven pinned setup steps in each of the trusted PR and full-stack workflows - add a workflow security contract proving every setup invocation remains covered and the overrides do not leak elsewhere ## Verification - focused workflow security tests: 6/6 passed - Prettier and git diff checks passed - observed unhealthy setup: 4-7+ minutes; historical healthy setup: about four seconds ## Risks This skips npm vulnerability-report telemetry for the setup action bootstrap only. Repository dependency checks, lockfile integrity, provider-secret authorization, and target-lock verification remain unchanged. ## Model Used Codex (GPT-5) |
||
|
|
03faa644fb |
ci(runner): inspect Daytona image metadata remotely (#12795)
## Thinking Path The reused Daytona image path already verifies the signed immutable digest. It then downloads every filesystem layer only to read OCI config fields. Buildx can retrieve the same config from that immutable digest without pulling the layers. The assertions can therefore stay intact while removing the expensive transfer. ## What Changed - inspect the signed immutable Daytona image config through Buildx after GHCR logout - preserve digest, source revision, content ID, platform, user, and provider-pack assertions - extend the workflow contract test for the metadata-only path ## Verification - Daytona image and workflow security tests: 10 passed - Prettier and git diff checks passed - observed full pull/prune cost: about 4m55s; metadata inspection: about one second ## Risks The current image has one runnable linux/amd64 platform plus its attestation. A future genuinely multi-platform image would need explicit linux/amd64 selection. ## Model Used Codex (GPT-5) |
||
|
|
d3c04d8932 |
fix(runner-e2e): prepare frozen Daytona plugin dependencies (#12791)
## Thinking Path
> - Paperclip manages AI agents and their provider runtimes.
> - The paid runner workflow installs target dependencies with lifecycle
scripts disabled.
> - The bundled Daytona plugin depends on an audited repo-local plugin
SDK link.
> - The lifecycle-safe install path did not create that link.
> - This pull request restores only the trusted Daytona preparation step
before provider secrets are exposed.
> - The benefit is a working Daytona canary without enabling dependency
lifecycle scripts.
## Linked Issues or Issue Description
**What happened?**
The Daytona paid canary stopped before lease or provider startup. The
trusted paid job disabled root lifecycle scripts, so the repo-local
plugin SDK link was absent. The plugin install returned a missing
runtime dependency error for @paperclipai/plugin-sdk.
**Expected behavior**
The trusted workflow must prepare the bundled Daytona plugin without
running untrusted dependency lifecycle scripts. The paid cell must start
only after its runtime dependencies and entrypoints pass validation.
**Steps to reproduce**
1. Dispatch the runner full-stack paid workflow for
core-compatibility.runner-acpx-claude.daytona.message-marker.
2. Let the trusted job install root dependencies with lifecycle scripts
disabled.
3. Observe the Daytona plugin installation fail before a lease or
provider process starts.
**Paperclip version or commit**
Feature head
|
||
|
|
313d6ca115 |
fix(runner): materialize pinned OpenCode binary (#12782)
## Thinking Path > - Paperclip manages AI agents and their provider runtimes. > - Paid runner validation installs target dependencies with lifecycle scripts disabled. > - OpenCode leaves a sentinel executable until its package lifecycle script runs. > - Running arbitrary lifecycle code would weaken the paid-secret boundary. > - This pull request materializes one exact pinned binary before secrets are exposed. > - The benefit is working OpenCode validation without trusting dependency install scripts. ## Linked Issues or Issue Description **What happened?** Every local OpenCode paid cell stopped before provider startup because `pnpm install --ignore-scripts` correctly retained `opencode-ai/bin/opencode.exe` as a sentinel. **Expected behavior** The trusted workflow must make the exact lockfile-pinned OpenCode executable available without running package lifecycle scripts. **Steps to reproduce** Run a local legacy or native OpenCode paid cell from the trusted workflow after the target dependency install. The provider health check reports that the OpenCode postinstall script was not run. **Paperclip version or commit** Default branch commit `865b4854fb44d3689f1c0ff17e3e715d52aaea73`. ## What Changed - Materialize only `opencode-linux-x64-baseline@1.18.17` into the matching `opencode-ai@1.18.17` package. - Verify package identity, version, regular-file type, SHA-256 equality, executable permissions, and runtime `--version`. - Invoke the helper for local OpenCode and breadth cells and for remote provider-pack assembly. - Retain `pnpm install --ignore-scripts`. - Add helper and trusted-workflow security regressions. ## Verification - Helper syntax checks passed. - Helper unit tests passed: 2/2. - Workflow-security tests passed: 5/5. - Prettier, actionlint, and diff whitespace checks passed. ## Risks Risk is low and contained to paid runner setup. The helper supports only Linux x64, fails closed on package or version drift, and runs before provider credentials enter the job. ## Model Used OpenAI GPT-5 Codex with repository tools and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change. - [x] I have specified the model used. - [x] I have checked ROADMAP.md and confirmed this does not duplicate planned core work. - [x] I have searched GitHub for duplicate or related PRs and found none. - [x] I have described the issue in this PR with the bug template labels. - [x] I have not referenced internal or instance-local issues. - [x] My branch name describes the change. - [x] Focused local tests pass. - [x] I added tests for the change. - [x] I updated the runner E2E security documentation. - [x] I documented the risks above. |
||
|
|
865b4854fb |
ci(runner): build paid artifacts once per campaign (#12777)
## Thinking Path > - Paid cells repeated the same TypeScript and Rust builds even when one campaign selected dozens of cells. > - The trusted workflow can compile once without provider credentials and distribute run-scoped, digest-verified artifacts. > - The paid cell can then disable install lifecycle scripts, verify each artifact before extraction, and expose provider credentials only to the final test step. > - Local JS-backed providers also need the setup-node interpreter permission-qualified before Rust verifies the launch artifact. ## Linked Issues or Issue Description Run 33786122875 proved target-lock setup and catalog selection, then failed before provider creation because trusted master did not yet qualify the setup-node interpreter. The same workflow also rebuilt TypeScript and Rust inside every matrix cell. ## What Changed - Build runner TypeScript and native binaries once per campaign in a credential-free job. - Build the remote provider pack once only when selected Daytona cells require it. - Upload run-scoped bundles with SHA-256 manifests and verify before extraction in each paid cell. - Remove repeated TypeScript, provider-pack, and Rust builds from paid cells. - Qualify the local provider Node interpreter before verified launch. - Propagate the resolved target lockfile through all five target-code jobs. - Keep local-only selection off Daytona and exclude Xiaomi from the 67-cell catalog. ## Risks A shared build artifact could fan out a bad payload to many cells. The producing jobs receive no provider credentials, use the exact authorized target SHA and resolved lockfile, and publish run-scoped artifacts. Every consuming job verifies SHA-256 before extraction. Paid dependency setup keeps lifecycle scripts disabled and provider credentials remain scoped to the final test step. ## Verification - Focused runner workflow-security, catalog, and Daytona-image tests: 25/25 passed. - Prettier passed. - Actionlint passed with only the two pre-existing SC2129 style notices ignored. - Git diff check passed. ## Model Used OpenAI Codex, GPT-5. ## Checklist - [x] Build jobs are credential-free. - [x] Paid installs disable lifecycle scripts. - [x] Artifacts are run-scoped and digest-verified before extraction. - [x] Trusted report and history jobs remain isolated from target artifacts. - [x] No Daytona or Xiaomi paid run was started for this change. |
||
|
|
6e50ca9d0a |
ci(runner): prepare target lockfile once for paid validation (#12774)
## Thinking Path > - The trusted target-branch runner workflow checks out PR code before paid tests. > - PR policy intentionally forbids manual lockfile commits. > - Some runner changes legitimately alter pnpm patch hashes. > - Frozen installs therefore fail before test selection. > - Resolve one script-disabled lockfile from the authorized immutable target SHA and distribute it by exact artifact ID and digest. > - Keep provider credentials and trusted reporting outside this resolution job. ## Linked Issues or Issue Description Target-branch paid runner campaigns currently fail frozen install when a PR changes pnpm patch content, even though ordinary PR CI regenerates the lockfile. ## What Changed - Added one credential-free target-lock job that resolves the authorized immutable target SHA with lifecycle scripts disabled. - Uploaded the resolved lockfile with its SHA-256 and restored it by exact artifact ID before every target-code frozen install. - Left trusted reporting and history jobs on the workflow SHA. - Changed the disabled-AWS fallback from unavailable ubuntu-latest-m to ubuntu-latest. ## Risks The workflow evaluates pnpm lockfile resolution from authorized target code. That job receives no provider credentials, disables lifecycle scripts, rejects unrelated workspace mutations, and exposes only a digest-verified lockfile artifact. Paid-secret jobs consume only that lockfile after exact artifact-ID and SHA-256 validation. ## Verification - Runner workflow-security focused tests pass. - actionlint passes. - Prettier and git diff checks pass. ## Model Used OpenAI Codex, GPT-5. ## Checklist - [x] Change is narrowly scoped to paid runner orchestration. - [x] Target lock resolution has no provider credentials and disables lifecycle scripts. - [x] Downloaded artifacts are selected by exact artifact ID and verified by SHA-256. - [x] Trusted reporting and history jobs remain on the workflow SHA. |
||
|
|
98c569b2df |
ci(runner): allow trusted branch targets (#12768)
## Thinking Path > - Paperclip uses paid runner tests to qualify agent execution. > - The runner workflow controls provider secrets and AWS runner access. > - The trusted workflow must stay on the protected default branch. > - The code under test often exists on a branch before merge. > - CODEOWNERS need a safe way to select that branch. > - This pull request separates workflow authority from the code under test. > - The benefit is pre-merge AWS testing without target-controlled workflow code. ## Linked Issues or Issue Description **What existing behavior does this improve?** The manual Runner Full-Stack E2E workflow can test only the default branch. **Subsystem affected** GitHub Actions and the paid runner E2E security boundary. **Current behavior** A CODEOWNER must merge runner changes before the trusted AWS workflow can test them. Selecting another branch as the workflow ref is rejected. **Proposed behavior** A CODEOWNER starts the workflow from `master` and supplies a same-repository branch in `target_branch`. The authorization job resolves the branch to one commit SHA. Catalog, image, and paid test jobs check out that SHA after authorization. Report sanitization and AWS publication use the trusted workflow SHA. **Reason and benefit** This permits paid pre-merge qualification on AWS. It keeps the workflow definition, report sanitizer, history publisher, environment deployment, and runner-group permission on `master`. **Breaking changes** None. The new input is optional. An omitted input still tests the default branch. ## What Changed - Add the optional `target_branch` workflow input. - Resolve only a branch in `paperclipai/paperclip` to an immutable SHA. - Pin catalog, image, paid test, and Daytona provenance to the target SHA. - Pin report sanitization and AWS history publication to the trusted workflow SHA. - Disable persisted checkout credentials in every job. - Key cancellation by the selected target branch. - Add policy regression coverage and operator documentation. ## Verification - `pnpm test:e2e:runner:unit` passes with 65 tests. - `actionlint -ignore SC2129 .github/workflows/runner-full-stack-e2e.yml` passes. - Prettier checks pass for all changed files. - `git diff --check` passes. ## Risks A CODEOWNER can authorize selected branch code to receive a cell-scoped provider credential. This is the intended trust decision. The workflow rejects fork refs and target-controlled workflow definitions. The trusted workflow SHA owns report sanitization and AWS history publication. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5. The exact serving snapshot and context-window size are not exposed. The model used tool-enabled reasoning and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1b74561fea |
ci(runner): route paid matrix to AWS fleet (#12765)
## Thinking Path > - Paperclip manages AI agents that perform work. > - The paid runner matrix verifies complete runner behavior with real providers. > - Each matrix job currently repeats work on GitHub-hosted runners. > - Paperclip has an ephemeral AWS runner fleet for trusted workflows. > - The paid workflow needs a reviewed and fail-closed route to that fleet. > - This pull request adds that route and keeps the existing hosted runner as the disabled-state fallback. > - The benefit is faster paid campaigns with the same actor, environment, and secret boundaries. ## Linked Issues or Issue Description **What happened?** The Runner Full-Stack E2E workflow always uses `ubuntu-latest-m`. It limits the matrix to 57 parallel jobs. The repository AWS fleet can run 100 ephemeral jobs, but the paid workflow cannot select it. **Expected behavior** An explicit repository flag must select the reviewed AWS fleet label. A missing or invalid flag must keep the existing hosted runner. The workflow must authorize the stable actor identity before it routes any paid job. **Steps to reproduce** 1. Dispatch the Runner Full-Stack E2E workflow from `master`. 2. Inspect a paid matrix job. 3. Observe that the job requests `ubuntu-latest-m` even when the AWS fleet should be used. **Paperclip version or commit** `da0947d3582ac7779d6bf11851c9938eca6c5c8c` **Deployment mode** GitHub Actions paid runner campaign. ## What Changed - Add a fail-closed `RUNNER_E2E_AWS_ENABLED` switch. - Select only the reviewed AWS fleet label or the existing hosted label. - Permit up to 100 parallel jobs in AWS mode. - Keep the hosted-runner limit at 57. - Reauthorize paid execution before checkout and provider access. - Stop paid checkouts from storing GitHub credentials. - Cancel superseded validation-ref campaigns while preserving `master` audit runs. - Add workflow policy checks and operator documentation. ## Verification - `git diff --check` - `actionlint -ignore SC2129 .github/workflows/runner-full-stack-e2e.yml` - The organization runner group permits this workflow only from `refs/heads/master`. - The repository AWS switch remains disabled until this pull request is merged and a one-cell probe succeeds. ## Risks - A wrong fleet policy can leave jobs queued. The disabled state keeps the existing hosted runner. - The AWS fleet uses paid compute. The workflow validates a configured maximum of 100 jobs. - The runner group, actor allowlist, and paid environment remain separate enforcement layers. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5 with agentic reasoning, repository inspection, code editing, Git, GitHub API coordination, and static workflow analysis. The exact deployed model identifier and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0f94521017 |
fix(runner): restore local session and task integrity (#12721)
## Thinking Path > - Paperclip is the control plane for agents that perform work. > - Paperclip Runner connects durable provider sessions to individual task runs through PRP. > - Provider continuity and per-run authority are different lifetimes. > - The existing implementation mixed those lifetimes and lost event metadata between provider frames, runnerd, persistence, API sanitization, and the task thread. > - That caused failed continuation, missing progress and Plans, duplicate replies, hidden failures, and unsafe recovery. > - This repair gives every heartbeat fresh authority, preserves qualified provider-session continuity, and restores one lossless presentation path without changing direct adapters. ## Linked Issues or Issue Description **What happened?** A second native heartbeat could reuse tickets, leases, command receipts, sequence state, and run identity from the first heartbeat. Provider phase and item identity could be lost before the UI read them. Redaction could corrupt protocol discriminators while still missing malformed credential tails. The task thread could fold progress into the final response, hide failures, or show more than one final answer. Native Codex also exposed approval modes that do not yet have a durable approval bridge. **Expected behavior** Each heartbeat uses a new PRP authority epoch. Codex and OpenCode preserve exact qualified provider sessions; ACPX emits an explicit continuity event when its qualified process-replacement policy is used. Every accepted provider event is presented, classified as internal, or surfaced as unsupported. The task page shows chronological progress, reasoning summaries, activity, Plans, interactions, terminal failures, and exactly one final reply. Direct adapters retain their existing path. **Steps to reproduce** 1. Enable the unified experimental Paperclip Runner setting. 2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex agent. 3. Run response, Plan, structured-question/resume, restart, cancellation, and failure scenarios. 4. Reload the task while active, waiting, failed, and settled. 5. On the old implementation, observe stale run authority, missing classifications, incomplete output, or duplicated/folded replies. **Paperclip version or commit** The repair is based directly on `master` at `87d05e194b643810d16d20612115acd01d735d43`. **Deployment mode** Local development with the embedded database. Related work: Refs #12616, #12646, #12666, #12685, and #12700. ## What Changed - Rotates PRP control-plane, outbox, ticket, lease, command, receipt, and sequence authority for each heartbeat while carrying forward only a validated provider-session identity. - Reads `control-plane-state.json`, validates both durable schemas and lifecycle values, resumes coherent current runs, archives qualified settled authority, and quarantines malformed or mismatched scoped state without moving ambiguous live legacy state. - Preserves Codex provider phase and stable item identities so commentary remains progress and only `final_answer` becomes final. - Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning lifecycle mapping. - Makes ACPX normalization lossless for visible reasoning, tool lifecycle metadata, stable bounded identities, Plan revisions, structured requests, failures, and qualified process replacement. Only the compatible terminal assistant message is promoted as final. - Applies schema-aware redaction before generic JWT-shaped detection and scans every diagnostic string leaf. Malformed raw/escaped quoted credential tails are redacted in both server and durable Rust state. - Restores snapshot-style chronological task presentation, expandable tool activity, inline Plan cards, visible waiting/resume/cancel/failure states, and exactly one final answer. - Makes `never` the only qualified native Codex permission mode and rejects unsupported persisted native modes with remediation. OpenCode and ACPX policies remain intact. - Keeps the unified experimental Runner setting as the only enablement flag. Onboarding and direct Codex, Claude, and OpenCode stay on their legacy execution/finalization paths. - Adds cross-language goldens, authority/recovery/fault coverage, exact response/count assertions, and native plus legacy acceptance scenarios. ## Verification - Pull-request GitHub Actions run Rust formatting/tests, TypeScript checks, server/UI tests, builds, protocol drift checks, browser E2E, and security scans. - A separate workflow-only validation ref is pinned directly on this PR head and runs the 35-cell paid local matrix: three core scenarios plus structured-question resume and restart/resume for native Codex, native OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode. Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315 - Acceptance requires exact single visible replies, monotonic sequences, matching envelope discriminators, one semantic terminal, one run terminal, no unresolved interaction, no duplicate mutation, no secret leakage, provider continuity, and zero native rows for direct adapters. - Per maintainer direction, tests are running in GitHub Actions rather than on the slower local host. Only formatters and static diff checks were run locally. ## Risks - Recovery from old or partial filesystem state is sensitive. The repair fails closed, preserves active or unverifiable authority, and quarantines only state whose scoped ownership is safe to move. - Provider event formats can change. Closed validators and boundary goldens turn new or malformed events into visible diagnostics instead of silent drops. - Shared task presentation could affect direct adapters. Runtime-fact gating plus the direct-adapter matrix protect the existing path. - Managed and remote providers are not qualified here. Shared code continues to compile and fail safely, but live qualification is deferred. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact deployed snapshot and context-window size are not exposed to this task. It used agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] The paid local-provider matrix is green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5716fe907e |
test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner subsystem executes agent work across local and managed provider backends. > - The lower pull requests restore the task runtime, provider backends, and managed-provider control plane. > - The restored system needs repeatable full-stack checks before it can ship safely. > - Paid live checks also need clear access, cost, and secret controls. > - This pull request adds acceptance, live evaluation, chaos, and release gates for the restored runner stack. > - The benefit is measurable runner parity with safer release decisions. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change covers runner tests, release workflows, server contracts, and evaluation tools. **Problem or motivation** The runner stack did not have one complete acceptance surface for native Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could miss provider drift, task-view regressions, cost-policy errors, and destructive cleanup errors. **Proposed solution** Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid workflows. Add live evaluation, chaos, cost-limit, redaction, and release contract checks. Add AWS AgentCore infrastructure and guarded provisioning tools. Keep the native runner experimental flag off by default. **Alternatives considered** We considered manual smoke tests only. They do not give repeatable evidence and they do not protect release branches. We also considered one large pull request. The stacked pull requests keep each review below the Greptile file limit. **Roadmap alignment** This work supports the shipped Cloud / Sandbox agents milestone and the shipped Agent evals & feedback milestone in `ROADMAP.md`. Related stack: - #12699 adds managed provider backends and lifecycle support. - #12691 adds qualified OpenCode and ACPX provider backends. - #12685 restores task runtime rendering and steering. ## What Changed - Add the runner full-stack harness with 57 catalog cells and 60 unit tests. - Add a Daytona runner image with digest-pinned base images and base-aware image-content checks. - Add guarded live evaluation and chaos workflows with a fixed 40-execution matrix; live and full-stack paid schedules now run only on Sundays or by manual dispatch. - Add in-flight reported-usage cost stops, post-turn cost caps, exact-threshold failure classification, secret redaction, retry classification, and actor authorization. - Reattach stream and hard-budget listeners before restart-recovery continuations so restored paid sessions cannot bypass in-flight interruption. - Preserve OpenCode usage and cost across tool-loop messages and turns while exposing an explicit current-run delta to durable accounting. - Keep PNG/WebM evidence in access-controlled artifacts only, reject SVG, and publish only pruned inert structured per-attempt evidence. - Add AWS AgentCore infrastructure, provisioning checks, and smoke tools; reject unsafe model identifiers, require exact stack ownership markers, and make failed-stack replacement explicit. - Add evaluation-session contracts and capability reports. - Add release workflow checks for immutable action pins, frozen dependency installs, exact weekly cron shape, paid-run guards, provider-secret isolation, and chaos test paths. - Reauthorize the original and triggering numeric actor IDs as the first step of every provider-secret job, including partial reruns, before checkout or provider access. - Give each full-stack matrix cell only its matching provider credential, expose Daytona only to Daytona cells, and disable shared dependency caches anywhere paid credentials or OIDC write access are present. - Protect the legacy manual E2E workflow with the same default-branch, allowlist, environment, and per-job authorization boundary. - Rotate live-eval candidates by week and retain 120 days of compatible history so the seven-week trend window remains viable. - Restore the root runner-acceptance commands and reconcile reported snapshots, raw receipts, and terminal usage without double counting or losing late usage. - Mark ACPX token deltas exact only when every budget field is present, keep cumulative cost/request authority separate, reject non-USD cost labeling, and include thought tokens in output-token budgets. - Keep `enableNativeRunner` off by default. The acceptance harness enables it only in its isolated test instance. ## Verification Passed locally: - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm test:runner-acceptance:typecheck` - `pnpm test:runner-acceptance` (19 tests) - focused OpenCode proxy, driver, runnerd transport, live-session, and turn-stream tests (106 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/clean-room-server.test.ts` (22 tests) - `pnpm test:e2e:runner:typecheck` - `pnpm test:e2e:runner:unit` (62 tests) - `node --test scripts/__tests__/release-verify-workflow.test.mjs` - `pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals` (22 tests) - `pnpm -r typecheck` - `pnpm build` - `node --test packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs` (6 tests) - `git diff --check` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core --lib --locked` (161 tests) - focused ACPX provider-event tests (10 tests) - The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged. I did not run paid live provider jobs or provision AWS resources. Those checks need credentials and can create cost. ## Risks The paid workflows can create provider cost. They require an allowlisted original and triggering actor, the protected `runner-e2e-paid` environment, explicit opt-in variables, and cost limits. The four provider credentials exist only in that master-only environment, which requires allowlisted reviewer approval and disables administrator bypass; repository and organization Actions scopes contain no copies. Provider usage arrives after a billable request, so the live guard cannot prevent one request from crossing a threshold. It interrupts immediately on the first reported threshold hit and permits no continuation. Visual evidence can contain secrets rendered as pixels. PNG/WebM remain only in access-controlled workflow artifacts; SVG and per-attempt XML are excluded, and S3/Pages receive a pruned structured dashboard. The AWS scripts can create cloud resources. They use explicit commands, least-privilege roles, KMS encryption, saved nonsecret metadata, and explicit teardown. This pull request does not enable the experimental native runner for existing instances. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The model used extended reasoning, tool use, code execution, and parallel subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |