mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 21:03:51 +02:00
codex/plugin-task-execution
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
18e8c121d9 |
fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path > - Paperclip manages agents through a shared native runner. > - Built-in harness support should ship with Paperclip's public distribution. > - Grok already speaks ACP; it does not require a new public bridge package. > - Sandbox provisioning owns the native executable and its pinned version. > - The runner must verify that prerequisite without downloading it during npm installation. > - This change separates built-in launcher identity from external runtime identity. > - Clean npm installation and live staging checks verify the distribution boundary. ## Linked Issues or Issue Description Refs #13882, #13973, #13977, #13979. This follow-up now targets master after #13882 was squash-merged. It replaces the private `@paperclipai/grok-acp` workspace package with runner-owned assets. Current master is included so the branch also contains the merged scheduler, complete-event capture, and durable cleanup fixes. ## What Changed - Ship Grok launcher and qualification metadata inside the runner's compiled output and the public server's vendored runner tree. - Remove the separate Grok npm package and all package-manager install hooks for this runtime. - Require the checksum-verified Grok Build 1.0.13 binary at `/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution environment. Provision it explicitly in the Daytona image and CI setup. - Keep native binaries outside the provider pack. Bind the built-in launcher into the pack manifest. - Preserve executable leases, descriptor-backed startup, credential fences, permissions, and exact ACP model admission. - Use `builtin:grok-acp` and `native:grok` as profile identities. Historical package-profile sessions fail closed on resume rather than being silently reinterpreted. - Resolve built-in assets from the authenticated sidecar location, including public server npm layouts. Keep the controller path out of provider environments. - Add clean npm tarball installation verification to the existing trusted canary CI job and the admitted manual EC2 verification path. It stages a unified release version and runs npm lifecycle scripts, then verifies missing-prerequisite rejection and admission after separate provisioning without credentials or inference. - Include the controller-owned provider pack in stamped Cloud images. Unstamped local images omit the pack and remain usable; remote ACPX requires full source provenance. - Correct CLI approval-page metadata for an already authenticated Cloud board user; approval authorization remains unchanged. - Honor explicit native-runner enablement in the Cloud agent picker and direct setup page, keeping the flag disabled by default. - Allow selecting the execution environment before connecting credentials. Include Grok in the existing authenticated hello-probe flow, targeting its pinned native prerequisite for runner setup. - Recover an existing subscription sign-in conflict through an explicit cancel-and-retry action, serialized after cancellation succeeds. - Preserve the selected ACPX harness before normalizing config fields, so new Grok agents use the Grok default model. - Keep the credential-free Cloud provider pack root-owned and readable after runtime UID remapping; verify manifest and referenced asset access under an unrelated unprivileged UID during image builds. - Archive prior failover backups alongside explicitly replaced harness state, preserving evidence while preventing stale backups from blocking a fresh replacement. - Update Daytona image content inputs and contract tests for the built-in assets and explicit provisioner. - Document and regression-test the shared `approve-all` default for Grok setup, saved configuration, and native execution. Explicitly saved restrictions remain unchanged. ## Verification Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05` incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the base PR was squash-merged. All 12 conflicts came from incoming files identical to the tested pre-squash base. The final tree exactly matches a three-way merge using that original base, preserving built-in Grok distribution and removal of the obsolete private package. All 252 focused runner/UI tests, six npm-isolation tests, and token gates pass. Fresh exact-head Greptile review is 5/5 with no outstanding findings; security scans and EC2 native compilation pass. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([run 36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)). The repository owner explicitly authorized bypassing code-owner approval after all checks passed; no CI checks or repository protection settings are bypassed or changed. The only remaining PR was removed from the completed stack metadata to permit native auto-merge. Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)). The initial attempt lost two EC2 runners to shutdown signals and stalled a third shard during dependency preparation; all three passed the same-commit failed-job-only retry. Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. [Final public npm verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764) passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public packages, an executed offline lifecycle sentinel, unchanged consumer lock, built-in launcher, missing-prerequisite rejection, and verified separately provisioned binary/command lease. Provisioning and cleanup require no host privilege elevation; only the positive probe mounts the temporary native binary read-only. The verifier is unchanged by the final master merge. All six isolation tests and an offline npm smoke test pass. The prior head had 56 green CI checks and a 5/5 review after two unchanged tests timed out and passed a failed-job-only retry ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)). All 56 recovery-display/lineage tests pass; re-review cleared the already-covered missed-retry concern. Earlier EC2 failures remain retained: [npm lockfile rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203), [missing compiler in the slim image](https://github.com/paperclipai/paperclip/actions/runs/36440210984), and the aggregate 15-minute test timeouts in those broad runs. Both broad attempts passed typecheck, token gates, Product E2E type/unit checks and build. The focused EC2 lane preserves the existing trusted-actor and immutable-source gates. Earlier documentation/test checkpoint `ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior unchanged. 154 focused tests pass across configuration building, native provider resolution, permission policy, credentials, UI configuration, and new-agent setup (including both Grok auth modes); token gates pass. All fresh CI is green for this head: 56 successful checks/statuses and two intentional skips ([run 36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)). Greptile is 5/5 with no new findings. Grok already inherits the shared `approve-all` default, so unattended setup requires no manual permission change. Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final staging continuation failure before provider startup: explicit replacement archived the old harness but left its failover backups active, which caused `runner_harness_state_mismatch`. The regression fails before the fix and passes after it; all eight adjacent recovery-safety cases also pass. Old backups remain inspectable inside the continuity archive. All fresh CI is green at this head ([run 36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)), with a 5/5 review. One unrelated Cursor test timed out in the initial server shard; the same-commit failed-job rerun passed, and both attempts are retained. Staging deployment is confirmed healthy on this revision. The controller image is `ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`. The final browser-created staging task passed on this exact revision with API authentication: context read → structured human question → controller restart → answer submission → same native provider session resumed → document saved → task Done. The two turns took approximately 119s and 77s. The actual write receipt was applied, and the saved document has exactly one revision containing the selected answer and requested marker. Usage and cost were not reported. [Controller image build](https://github.com/paperclipai/paperclip/actions/runs/36360889243). - Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`: all CI green (53 successful checks/statuses, two intentional skips), including repository typecheck/build/tests, native Runner tests, browser shards, and canary installation checks. [CI run 36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529). Greptile is 5/5 with no unresolved findings. - Focused checks cover Grok credentials, executable admission, launcher assets, provider-pack paths/permissions, workflow contracts, setup defaults, CLI authorization, and subscription conflict recovery. All 39 protocol definitions validate. Final integration checks pass 124 catalog/evidence/cache tests and nine project-form tests; token gates pass. Some local dependency checks could not load the stale installed dependency tree; the corresponding fresh EC2 checks pass. - Clean public npm installation passed on EC2 at `8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run 36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)): 17 unified-version packages, lifecycle scripts enabled, built-in launcher present, no separate Grok package or npm-downloaded binary, missing prerequisite rejected, separately provisioned native executable and command lease verified. No credentials or inference were used. Subsequent changes preserve this npm asset layout. - The immutable Daytona prerequisite image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`, built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous Cloud controller image was `ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`, built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded by the latest image above. Its EC2 build verified provider-pack access under an unrelated unprivileged UID. - Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed full Grok onboarding with the correct `grok-4.7` model, saved credential delivery, and pinned Daytona execution. A browser-created task read context and asked the structured human question. After a controller restart, answering the persisted question resumed the same native provider session, saved the requested document, and completed the task. Actual tool outcomes and durable state agree: one question and one document revision. The two successful turns took 42.7s and 63.1s; usage and cost were not reported. - Restricted policy returned the expected `approval_required` outcome. Functional staging tests explicitly selected `approve-all`; controller authorization and governed approvals remain enforced. Temporary board CLI access was revoked and verified rejected (HTTP 401), and the disposable onboarding agent was paused. Failures remain retained: the pre-fix continuation failure (its task remains blocked; the passing final task is fresh), the original Cloud provider-pack permission failure, the expected restricted-policy denial, the superseded npm staging failure, and an earlier monolithic CI infrastructure timeout. Browser CI exposed a project alias/form race; the final stack uses master's stronger draft-preservation fix and all browser shards pass. Historical full subscription/API protocol and Product rosters retain their original source revisions and do not qualify this packaging revision. No local Docker or Rust build was used. ## Risks The branch includes master’s draft-preservation fix for project URL aliases. It keeps the same project’s edit form mounted and clears prior data when the project or company changes. Custom sandboxes and local execution hosts must provision the pinned binary before Grok starts. Missing, changed, unsupported-platform, and symlinked executables fail admission. The new builtin profile cannot resume sessions created with the former private-package profile. Existing Claude/Codex npm bridge profiles retain their package pins. Grok restricted modes preserve the selected policy but cannot automatically admit Paperclip calls: ACP permission metadata does not independently bind tool authority, so those calls stop with `approval_required`. New Grok configurations default to `approve-all`, including API configurations that omit the mode. Existing explicitly restricted configurations remain restricted; controller authorization and governed approvals remain enforced. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1a394bd30 |
feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |