Commit Graph
4 Commits
Author SHA1 Message Date
DottaandPaperclip 57e977be72 feat: integrate Pi 1.0 into the experimental Runner (#14921)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - The experimental Runner owns provider processes and durable
sessions.
> - Pi needs working task execution and human controls.
> - The five-PR stack must preserve changes already on master.
> - Each layer now carries the complete integrated source for a safe
sequential fallback.
> - This PR belongs to native GitHub stack #15602, ending at #14956.

## Linked Issues or Issue Description

Refs #14436, #14631, #14743 and #14956.

Ship Pi 1.0 through the experimental Paperclip Runner. The five PRs are
#14921, #14922, #14923, #14924 and #14956. The user authorized the
complete merge after checks pass. Existing `pi_local` execution is
unchanged. Accounting and wider provider/platform qualification remain
deferred.

## What Changed

- Recover missing final replies after workspace finalization changes
owners, using accepted-turn evidence without rerunning work or granting
external-chat publication.
- Preserve the admitted Pi instruction root across warm runs, while
retaining changed-root rejection.
- Give Pi a bounded 15-second default shutdown grace so stop, drain
acknowledgement and durable suspension can complete. Explicit deadlines
and other providers retain their existing behavior.
- Integrate the Pi 1.0 runtime and master contracts.
- Use Pi profile 22. Preserve explicit caller-selected models and exact
native thinking levels. Keep Pi's wrapper, helper, extension and
question/control behavior unchanged from the qualified profile-19
runtime.
- Preserve master's Dot lifecycle and consent fields, configured task
environment, status guards and current Codex/Claude dependency versions.
Cursor stays qualified. Copilot stays pending; profile 17 binds the
changed shared protocol validation sources.
- Exclude general AWS IAM credentials from Pi static/custom provider
bindings and selected task projections; preserve the provider-scoped
Bedrock bearer key. Profile 21 is retained as historical provenance.
Rust and cloud install probes use the current declaration.
- Patch bundled brace-expansion 5.0.9 to the exact official 5.0.12
payload. Pin the patch and complete runtime closures. Include the patch
in normal installed setup tooling. Keep the upstream Pi shrinkwrap as
provenance and permit only this exact security correction.
- Include current attestation files in the Docker build context. Keep
the repository lockfile unchanged from master. CI and private image
builds resolve manifest changes before their frozen installation.

## Verification

- Full local `pnpm -r typecheck` passes, including Runner Rust, server
and UI. Focused integration checks pass: 194 Runner
admission/environment tests, 63 profile/credential tests with one
expected skip, 152 Dot/UI configuration tests, and Pi transcript/notice
tests.
- Full local `pnpm build` passes on the final source.
- Fresh final-source checks pass: all 698 Rust workspace tests (32
binaries), 156 credential/profile/controller tests with one expected
skip, Runner TypeScript typecheck, and 20 package/setup/sandbox tests.
- The profile-21 Pi materializer passes on the native host with the
official pinned Node 24.21.0 and its npm. It verifies all 150 locked
packages, the patched dependency and the exact closure. Setup/package
bundle tests and UI token gates pass.
- The old hashes were reproduced for all three supported targets before
calculating the patched graph. New closure hashes are darwin-arm64
`282022db10150c6632b3444df421342e7d534bdf5d5fb1097a2e79d0625a2bcf`,
darwin-x64
`64e251e19009f755c0b04f73ce2138246faab71a961b0f13d75ebfcc34bef12e`, and
linux-x64
`713b1fdff42fb56a1518bdc084f181d70bee8ebadc3e4b1d76321ed9108c8410`.
Independent native platform execution is separate from graph identity
reproduction.
- Historical cloud qualification remains unchanged: all seven core cases
pass on shipping source `10dc43c9ec65d88c2f782d62afb296d09494f215`,
harness `1a4408a48cfb5a1f094a311141c257c92cd7a893`, image
`sha256:5b3a775b383591bda1b0c1889e509acc70ce7f37c53f09733c81d59037f02280`,
and accepted Sonnet 4.6/low fixture. All 215 canonical files and all
seven cleanup checks pass independent verification. These are profile-19
results and are not relabeled as fresh profile-22 runs.
- Current Pi digest:
`sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`.
[The readiness
plan](https://github.com/paperclipai/paperclip/blob/codex/pi-production-readiness/doc/plans/2026-10-02-pi-production-readiness.md)
preserves campaign and failed-attempt provenance.
- Merge only after every PR's current-head CI and fresh review pass.
Linux CI covers the full suites, build and browser tests. The local
embedded Postgres API-authority suite cannot start on this macOS/Node 26
host, so Linux CI must confirm that suite.

### Fresh profile-22 core qualification — 2026-10-08

All seven accepted core cases pass canonically on Pi profile 22, with
`openrouter/anthropic/claude-sonnet-4.6` and native-confirmed low
thinking. This model is a fixture; production accepts the caller's
explicit Pi provider/model.

Runtime/install source: `3241a992f2a7703e59e97ed0fd3e5d6405de4401`.
Frozen accepted harness: `1a4408a48cfb5a1f094a311141c257c92cd7a893`.
Immutable cloud image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:506f22db7edd78f37c0c40bec1cc084af1850455026dbf467194bfbb8fcef141`.
Pi digest:
`sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`.

[Hosted Linux image and clean-install
verification](https://github.com/paperclipai/paperclip/actions/runs/37868328023)
passes, including all 20 source-bound archives, normal CLI/Pi setup,
companion import and the production pack reader. This exact installation
source includes the latest master integration and the corrected Pi warm
instruction-root fence. Full local typecheck/build and current-head
hosted CI verify the final stack. All 13 focused real-root regressions
pass. The full local executor suite passed 662 tests; 15 database tests
could not start the Mac embedded PostgreSQL service. Hosted Linux CI
passes the full required verification and E2E checks. These fresh
results keep their own source identity; profile-19 results remain
historical.

| Core path | Canonical campaign | Retained archive SHA-256 |
| --- | --- | --- |
| File edit, validation, download and Done |
`pi-core22-replyfix-0-1791511228` | 23 files;
`a473e8603a3dd4737863291f8d3d1e392391f0b16d433c3e0e0e9d8baf7a97b0` |
| Pending question and controller restart |
`pi-core22-replyfix-1-1791511376` | 33 files;
`6b829c4eb74e1f32a89c692a4ae7130dbfc1c6d3cf13915effe2103d9e242c8e` |
| Three-turn session/process/workspace continuity |
`pi-core22-replyfix-2-1791511587` | 23 files;
`7a87021f8f9a3fdd3c58bb4467f8d82c635e3ea4795d6e75f144d9aa14818df8` |
| Four typed questions and browser reconnects |
`pi-core22-replyfix-3-1791511881` | 42 files;
`9e31755252be1f4f9cb0626c984c142d4d1ae5f5bee3a7af08444db8d12c280a` |
| Plan approval and completion | `pi-core22-replyfix-4-1791512031` | 22
files;
`a0383ce1aab38e7b5a25ce0e9dd3bebea5c037ebd96ae6b29dae19015da2ae2c` |
| Same-turn steering and permission denial |
`pi-core22-replyfix-5-1791512261` | 39 files;
`c929b8c7070f0b66aedc17e65ca46e6beab1e363926ac9f7e2a75fb250f05949` |
| Stop during pending permission | `pi-core22-replyfix-6-1791512390` |
33 files;
`7f0a58ae0f4d5bfc76149435f4e322537089c5bd16e7ffe9b5ad71f10a621a07` |

All 215 canonical files (28714587 bytes) are independently
hash-verified. All seven cleanup grades pass, with no owned runtime
process or temporary root after each case. Automatic retries are zero.
The owned cloud host stopped normally after retention. The prior
profile-22 warm attempt remains failed and separately retained: archive
SHA-256
`1e54eba5ec72b50cee1534b23d1d1d4f21a090006b8a64501ba70db972abfde5`. Its
original canonical classification is preserved. Diagnosis reproduced a
product bug comparing an agent-files root against an unset
checkpoint-only field. The fix stores the admitted physical root
separately from the adopted per-run collection capability. The real-root
regression fails before the fix and passes afterward, including
rejection of a changed physical root. Fixture, grader, model and all
seven accepted case IDs are unchanged; this fresh campaign tests
final-reply publication after file registration first. The intermediate
restart attempt also remains failed and retained: archive SHA-256
`5dcaefdf1d17cf4cd54fd4cf810f45e736667392339b8ce7caf08bb4e225277f`. Its
original canonical classification is preserved. Pi resumed, wrote the
verified answer and completed its task; exact runner suspension was
proven, but idle stop consumed about 5.2s and left under 3s for the
drain acknowledgement. The Pi-only default shutdown grace is now 15s,
preserving a full 5s drain round trip and a finite suspension reserve.
Explicit caller deadlines, other provider defaults, literal drain
receipts and exact suspension identity checks remain unchanged. The
timing regression fails before this correction and passes afterward; all
18 focused settlement tests and Runner typecheck pass. The final-source
file attempt is also preserved as failed (`candidate_failure`), archive
SHA-256
`db6767b6773ea618997927ac77bdb005a5ac81492c7b9c0ffbc900449f829bc9`.
Native edit, validation, exact downloadable artifact and Done/succeeded
all passed, and the exact final reply was durably recorded. A workspace
recovery owner completed before the live heartbeat reached presentation,
leaving that reply absent from task chat. Recovery now materializes only
a completed final reply from the accepted turn of an ordinary internal
Done task, preserving issue/run/contract binding, suppression,
external-chat authorization and same-run deduplication. The database
regression covers the generated file-preparation receipt, suppression,
unapproved external continuation and replay. Server typecheck and all 49
response-selection tests pass; hosted Linux verifies the database
regression because embedded PostgreSQL cannot start on this Mac.

The delayed-final-answer database regression passes on [the final
root-source Linux server
shard](https://github.com/paperclipai/paperclip/actions/runs/37868262553/job/113628594152),
alongside 1,108 passing tests. The first root Runner shard had one
unchanged durable-resume test exceed its 5-second timeout; the identical
top-source shard and the isolated exact test passed. One rerun of that
failed job and its required aggregate passed without source or test
changes. The original failed job log and the single-rerun receipt remain
retained.

### October 9 merge verification

Current merge head: `5a8fe63512a7166aaef5cf50065a25008aa8b44b`. All
current-head checks pass, including `ci / verify` and `ci / e2e`;
exact-head Greptile review is 5/5 with no unresolved threads. Current
master conflicts are resolved. The user authorized the maintainer
override of the code-owner review gate after these checks. The seven
retained live core cases remain bound to source
`3241a992f2a7703e59e97ed0fd3e5d6405de4401` and its recorded cloud image.

## Risks

- The security correction changes the dependency closure and profile
identity. Old sessions must reopen on the new profile. Exact identities
and credential bindings fail closed.
- The runner remains experimental and requires explicit selection.
Legacy Pi Local is unchanged. Caller model IDs pass through; the E2E
model is a fixture.
- Accounting and the broad platform/provider matrix remain deferred.
This merge does not publish a release or deploy a service.

## Model Used

OpenAI GPT-6 through Codex assisted with reasoning, repository
inspection, editing and tool use. The exact serving ID and context
window are not exposed in this session. Final live qualification uses Pi
1.0.0 with `openrouter/anthropic/claude-sonnet-4.6` and native-confirmed
low thinking.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 08:56:41 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00
Dotta 5716fe907e test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner subsystem executes agent work across local and managed
provider backends.
> - The lower pull requests restore the task runtime, provider backends,
and managed-provider control plane.
> - The restored system needs repeatable full-stack checks before it can
ship safely.
> - Paid live checks also need clear access, cost, and secret controls.
> - This pull request adds acceptance, live evaluation, chaos, and
release gates for the restored runner stack.
> - The benefit is measurable runner parity with safer release
decisions.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting. This change covers runner tests, release workflows,
server contracts, and evaluation tools.

**Problem or motivation**

The runner stack did not have one complete acceptance surface for native
Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could
miss provider drift, task-view regressions, cost-policy errors, and
destructive cleanup errors.

**Proposed solution**

Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid
workflows. Add live evaluation, chaos, cost-limit, redaction, and
release contract checks. Add AWS AgentCore infrastructure and guarded
provisioning tools. Keep the native runner experimental flag off by
default.

**Alternatives considered**

We considered manual smoke tests only. They do not give repeatable
evidence and they do not protect release branches. We also considered
one large pull request. The stacked pull requests keep each review below
the Greptile file limit.

**Roadmap alignment**

This work supports the shipped Cloud / Sandbox agents milestone and the
shipped Agent evals & feedback milestone in `ROADMAP.md`.

Related stack:

- #12699 adds managed provider backends and lifecycle support.
- #12691 adds qualified OpenCode and ACPX provider backends.
- #12685 restores task runtime rendering and steering.

## What Changed

- Add the runner full-stack harness with 57 catalog cells and 60 unit
tests.
- Add a Daytona runner image with digest-pinned base images and
base-aware image-content checks.
- Add guarded live evaluation and chaos workflows with a fixed
40-execution matrix; live and full-stack paid schedules now run only on
Sundays or by manual dispatch.
- Add in-flight reported-usage cost stops, post-turn cost caps,
exact-threshold failure classification, secret redaction, retry
classification, and actor authorization.
- Reattach stream and hard-budget listeners before restart-recovery
continuations so restored paid sessions cannot bypass in-flight
interruption.
- Preserve OpenCode usage and cost across tool-loop messages and turns
while exposing an explicit current-run delta to durable accounting.
- Keep PNG/WebM evidence in access-controlled artifacts only, reject
SVG, and publish only pruned inert structured per-attempt evidence.
- Add AWS AgentCore infrastructure, provisioning checks, and smoke
tools; reject unsafe model identifiers, require exact stack ownership
markers, and make failed-stack replacement explicit.
- Add evaluation-session contracts and capability reports.
- Add release workflow checks for immutable action pins, frozen
dependency installs, exact weekly cron shape, paid-run guards,
provider-secret isolation, and chaos test paths.
- Reauthorize the original and triggering numeric actor IDs as the first
step of every provider-secret job, including partial reruns, before
checkout or provider access.
- Give each full-stack matrix cell only its matching provider
credential, expose Daytona only to Daytona cells, and disable shared
dependency caches anywhere paid credentials or OIDC write access are
present.
- Protect the legacy manual E2E workflow with the same default-branch,
allowlist, environment, and per-job authorization boundary.
- Rotate live-eval candidates by week and retain 120 days of compatible
history so the seven-week trend window remains viable.
- Restore the root runner-acceptance commands and reconcile reported
snapshots,
raw receipts, and terminal usage without double counting or losing late
usage.
- Mark ACPX token deltas exact only when every budget field is present,
keep
cumulative cost/request authority separate, reject non-USD cost
labeling,
  and include thought tokens in output-token budgets.
- Keep `enableNativeRunner` off by default. The acceptance harness
enables it only in its isolated test instance.

## Verification

Passed locally:

- `pnpm --filter @paperclipai/paperclip-runner typecheck`
- `pnpm test:runner-acceptance:typecheck`
- `pnpm test:runner-acceptance` (19 tests)
- focused OpenCode proxy, driver, runnerd transport, live-session, and
turn-stream tests (106 tests)
- `pnpm --filter @paperclipai/paperclip-runner exec vitest run
src/live/clean-room-server.test.ts` (22 tests)
- `pnpm test:e2e:runner:typecheck`
- `pnpm test:e2e:runner:unit` (62 tests)
- `node --test scripts/__tests__/release-verify-workflow.test.mjs`
- `pnpm --filter @paperclipai/paperclip-runner
test:runner-workflow-evals` (22 tests)
- `pnpm -r typecheck`
- `pnpm build`
- `node --test
packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs`
(6 tests)
- `git diff --check`
- `cargo test --manifest-path
packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core
--lib --locked` (161 tests)
- focused ACPX provider-event tests (10 tests)
- The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged.

I did not run paid live provider jobs or provision AWS resources. Those
checks need credentials and can create cost.

## Risks

The paid workflows can create provider cost. They require an allowlisted
original and triggering actor, the protected `runner-e2e-paid`
environment, explicit opt-in variables, and cost limits. The four
provider credentials exist only in that master-only environment, which
requires allowlisted reviewer approval and disables administrator
bypass; repository and organization Actions scopes contain no copies.

Provider usage arrives after a billable request, so the live guard
cannot prevent one request from crossing a threshold. It interrupts
immediately on the first reported threshold hit and permits no
continuation.

Visual evidence can contain secrets rendered as pixels. PNG/WebM remain
only in access-controlled workflow artifacts; SVG and per-attempt XML
are excluded, and S3/Pages receive a pruned structured dashboard.

The AWS scripts can create cloud resources. They use explicit commands,
least-privilege roles, KMS encryption, saved nonsecret metadata, and
explicit teardown.

This pull request does not enable the experimental native runner for
existing instances.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex with GPT-5. The model used extended reasoning, tool use,
code execution, and parallel subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-02 08:55:08 -05:00