Files
PaperClipAI/tests/runner-e2e/FIXTURES.md
T
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00

13 KiB

Runner E2E fixture authoring

The fixture catalog is executable production-contract data. Keep it small, typed, deterministic, and free of raw credentials.

Suites and matrices

A RunnerSuiteFixture declares one durable testing purpose: stable ID, label, description, profiles, environments, cases, expected size, and definition or ranking metadata. Its execution IDs are globally prefixed as <suite>.<profile>.<environment>.<case>. Add a new suite when the testing purpose or desired cross-product differs; do not inflate an existing suite with unrelated dimensions.

The suite definition fingerprint is historical comparison metadata. Any profile, model qualification, environment, task, or ranking-snapshot change must change that fingerprint automatically so the dashboard can annotate the boundary instead of silently joining unlike totals.

Agent profiles

Add RunnerProfileFixture entries in catalog.ts. A profile declares:

  • a stable ID and searchable groups;
  • legacy or native generation;
  • adapter/provider and required credential;
  • a model imported from its adapter constant or qualified runner profile;
  • supported environment IDs;
  • expected runtime metadata; and
  • an agent payload factory.

Do not duplicate model IDs, qualification decisions, CLI versions, or runner artifact rules. Codex profiles import DEFAULT_CODEX_LOCAL_MODEL, OpenCode profiles import QUALIFIED_OPENCODE_MODEL, and ACPX profiles import QUALIFIED_ACPX_PROFILES. Add or qualify models at their owning production source first.

OpenRouter breadth profiles are generated from openrouter-models.json, not written by hand. That reviewed snapshot must contain exactly five unique, available, tool-capable models with rank, canonical ID, display name, supported parameters, source URL, capture time, and verified content hash. Refresh it manually with pnpm test:e2e:runner:models:update; nightly campaigns never change fixture definitions.

Agent adapterConfig.env values must be {type:"secret_ref", secretId, version:"latest"} objects supplied to the factory. A fixture source containing a raw secret-looking value is rejected by catalog validation.

The manual Grok subscription profile uses GROK_AUTH_JSON as an explicit login fixture. It does not put this credential in agent configuration or substitute an API key. Setup seeds a new company-scoped Grok home inside the disposable instance with mode 0700 and an exclusive mode-0600 auth file. Setup rejects redirected, occupied, or nonisolated homes. Production runner discovery and refresh operate on that company login; teardown destroys it after the remote environment is removed. This fixture tests subscription execution, not the interactive browser login flow.

Environments

An EnvironmentFixture declares driver/provider, credential requirements, attempt deadline, lifecycle behavior, expected execution target, and a payload factory validated by the shared environment schema.

The local environment is instance-managed: company creation ensures it exists, and the public API intentionally rejects a second local environment. The setup registry therefore discovers that row through the public environments API. This still provides full isolation because every cell starts a new Paperclip instance and database.

Daytona creates sandbox environments through the public API. The core fixture keeps reuseLease:false and runnerLifecycleMode:"per_turn". The dedicated warm-continuity fixture uses reuseLease:true and runnerLifecycleMode:"warm"; its distinct configurationKey is part of the suite fingerprint even though both fixtures report environmentId:"daytona". Keep short provider cleanup backstops, a Daytona secret reference, and an immutable image digest. Teardown must delete the environment with reusable-lease destruction and must fail the cell if cleanup cannot be confirmed. Keep CPU, memory, and disk explicit: lease metadata and the per-test public-list-price runtime estimate depend on that pinned billable resource shape. Changing it requires updating billing tests and reviewing the versioned Daytona rates in billing.ts.

Usage and billing data

Do not add fixture-authored token or dollar expectations. The live harness reads usage from selected public heartbeat-run records and records coverage per run. Provider-reported dollars remain distinct from runtime estimates. A zero or missing native usage payload is unavailable unless a real token-bearing receipt or provider cost proves otherwise. New execution environments must provide lease/resource metadata for a runtime estimate or explicitly remain unavailable; never infer that missing billing data means free execution.

Future providers (SSH, E2B, Modal, Cloudflare, Kubernetes, Novita, exe.dev) should implement the same setup/probe/cleanup contract before being added to a matrix. Unsupported profile/environment combinations belong in supportedEnvironments, not in ad hoc test conditionals.

Task cases and matchers

A RunnerTaskFixture owns a work mode, a typed flow, expected run count, nonce-based title/prompt/marker factories, per-environment attempt deadlines, deterministic matchers, and expected terminal state. Single-turn prompts should make one bounded request with observable output and no nondeterministic judging. The plan_revision_acceptance flow must also provide revision-request and Plan marker factories. question_resume_completion must define the deterministic browser answer and prove exactly two successful runs with no pending interaction. plan_approval_completion must target the exact two-step canonical Plan revision, capture its pending UI, approve in the browser, and prove exactly two successful runs. warm_three_turn provides exactly two browser follow-up messages, preserves one project/execution-workspace scope, verifies host file contents after every turn, and finishes within three ten-minute turn deadlines. Native turns 1 and 2 include an actionable human review in the completion report's attentionRequests. Paperclip creates the review gate from that report. An explicit question-tool wait yields the turn and suppresses its final prose, so it is not interchangeable with this completion-review fixture. Turn 3 reports Done without another review.

Every selected case runs in its own isolated Paperclip process, and independent cases may run concurrently. Follow-up turns inside one case retain their shared task state. Each case creates and tears down its own company, secrets, environment selection, agent, and browser-created task. The current plan case proves three runs on the same issue: publish a two-step Plan, request a three-step revision through the UI, and accept the exact new revision through the UI before verifying implementation and Done.

The matcher union supports message exact/contains/regex/ordered checks, issue and run state, runtime/environment metadata, files, artifacts, JSON paths, and JSON Schema. The initial cases use normalized message_contains plus state, runtime, and environment assertions; the plan flow additionally verifies canonical document revision IDs, bodies, step counts, interaction targets, and visible previews. Add matcher behavior and credential-free tests together.

Adding a task expands its suite's matrix. Update the suite's intentional size, the complete-catalog size, and credential-free unit tests in the same change. Paid tests never silently skip a missing credential or unsupported artifact.

New Paperclip object fixtures

The explicit-only lifecycle-baseline suite reuses this registry and existing continuation, chat and governed-action flows. Its narrative pairs require actual agent/run-attributed comments or exact visible responses. See the live baseline contract for selectors and proof boundaries.

Register new objects in live-fixtures.ts with explicit dependencies in FixtureRegistry. Setup must use a public API. Teardown runs in reverse order and is invoked after partial setup failures. Direct database writes and private test-only runner endpoints are prohibited.

The expected dependency shape is:

company
└── encrypted secrets
    └── environment
        └── agent
            └── browser-created task

Projects, goals, apps, and configuration fixtures can be inserted into that graph without changing the launcher. Keep returned fixture state to IDs and sanitized metadata; never retain raw secret values.

Required checks

Run before a fixture change is reviewed:

pnpm test:e2e:runner:unit
pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner -- --list

Then run the narrowest paid cell that exercises the fixture. A full matrix is a manual or scheduled campaign, not a PR requirement.

Persistent chat fixtures

chat-cases.ts defines the eight-case agent-chat suite; chat-flow.ts drives the production composer, plan revision/approval controls, questions, reset command, and project cards. Keep its 28 local cells intentional. expectedRunCount counts provider turns, including cancelled and handed-off task runs, but excludes synthetic /new runs. Assertions must inspect all company runs because ordinary issue lists exclude the source conversation. assertChatHandoff rejects missing projects/plans, chat children, wrong assignees, and execution before plan commit.

Retained api-state.json, chat-handoff.json, and plan-revision evidence include persisted comments, session generations, run context and logs, project workspaces, task documents, and ordering. They pass through the normal sanitizer. Screenshots are allowlisted to the exact disposable agent chat. Cleanup cancels all active runs in the isolated company, including handed-off work; usage from failed and cancelled runs must not disappear from campaign totals.

Warm three-turn continuity grades the exact workspace file after each turn, task completion, and sandbox/session identity. It also requires a visible persisted final reply with each turn marker once and in order. It does not grade exact final-reply wording; the hello and continuation fixtures retain those exact-response checks. This separates workspace persistence failures from model response-format variance.

chat-hardening.ts adds the explicit-only agent-chat-hardening journeys. Use the ordinary public APIs to seed source documents and blockers. Keep the answer out of the user's status/review request. Grade the exact source values, latest blocker, preserved task identities, worker-authored output, and real executions. The status request asks for JSON so the grader can distinguish the current blocker from a historical mention and compare active-run count separately from task status. The request must not reveal those expected values. Capture the source after seeding and compare every field in the public issue update contract, plus labels, dependencies, and dedicated-endpoint settings. Derived inbound references may change when the chat legitimately cites a task. The lost-acknowledgement probe may interrupt only the fixture browser's own comment request after the real server has committed it. Retain its request ID and replay that same request through the public API after restarting the server. Never fabricate tool results or repair task state after a failed assertion.

chat-stories.ts seeds an ordinary file wait in the isolated agent's actual home workspace; native Codex intentionally cannot see arbitrary host temp files. The observed run workspace must match the fixture location. This is a deterministic interruption boundary. The real provider command writes the readiness file and waits at most two minutes. The harness must persist the next browser message while the same run is active before supplying the brief. Always release the wait in finally. Save boundary observations independently of the final outcome. The final answer must recover a brief reference absent from both prompts; the revision oracle also reads the actual conversation plan. Fixture setup never enables native API tools for this suite. Do not describe its prepared-agent settings case as a production onboarding qualification.

The agent-chat-qualification local fixtures use public APIs to seed two workers and a task with a saved plan, or read-only tasks with contradictory historical comments. Ordinary Node file waits in the isolated agent workspace establish observable active execution; no provider output or database outcome is fabricated. A worker-crash case sends SIGKILL only to a positively identified running native worker PID, then uses the production Retry button. Each gate is released in a finally block. Source facts and boundary state are retained with the attempt. The lifecycle suite also includes two legacy disposition-repair probes. Their first provider turn intentionally omits task disposition, and their second turn must be an automatic, causally bound repair that records completion. They use public task comments/status APIs and run-detail evidence; no private runtime hooks or database mutations are used by the fixture.