mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its adapters supply task context and access to Paperclip skills and tools. > - The default hire manual and shared prompts also repeat general work procedures. > - Those procedures overlap with stock provider instructions and the Paperclip skill. > - Existing E2E fixtures supply a QA manual, so they do not qualify the production default. > - This pull request reduces the generic instructions and adds real default-hire coverage. > - The benefit is less competing guidance, with inspectable evidence for preserved skills and task context. ## Linked Issues or Issue Description Refs: #14920. That merged change preserves native Codex base instructions. This PR covers the default manual, shared legacy prompts, operational skill guidance, and the narrowly approved ACP skill-discovery/session-environment repair for measured delivery and credential-persistence failures. **What existing behavior does this improve?** New non-CEO hires without a custom bundle and legacy task/chat startup and continuation prompts. **Current behavior** The shipped default manual contains 602 words. Generic task/chat prompts and ordinary resume deltas repeat work procedures already available through the harness and Paperclip skill. **Proposed behavior** The default manual contains only the eight-word company identity. Shared startup prompts retain identity and connection guidance. Ordinary resume deltas retain current work context without the generic execution contract. **Reason and benefit** Let the stock harness guide general work. Keep Paperclip-specific capabilities and independently test default hires, skills, ordered comments, and chat restart. **Breaking changes** New default hires receive less guidance. Existing saved manuals, explicit custom bundles, CEO templates, and specialized wake contracts retain their behavior. The obsolete includeExecutionContract option remains accepted for source compatibility. ## What Changed - Reduce the default hire manual to one sentence. - Reduce shared task/chat defaults and remove the generic ordinary-resume contract. - Keep connection guidance, auth, skills, custom prompts, and specialized wake context. - Add credential-free instruction-boundary gates and 26 explicit Product E2E cells across eight legacy/native profiles, including two focused Paperclip-storage cases. - Capture public hire receipts before providers run, then grade delivered prompts and independent task/chat outcomes. - Add an early legacy skill API recipe for saving a task document, checking the saved revision receipt and linking the document. Improve stock task/heartbeat skill-selection metadata and show a clickable Markdown UI-link example. Keep native tool completion separate. - Advertise bounded routing descriptions and exact successfully staged SKILL.md paths in legacy ACP Claude; keep full bodies on demand and preserve remote path rebasing. - Remove only the provider environment from copied persisted ACP session records, while loading current run credentials and preserving all other options/conversation state. - Regenerate both capability metadata inventories and reject stale manifests/inventories before provider admission. - Publish the original reduction and focused skill-repair comparisons, preserving all failures, automatic recovery, cost coverage and limitations. ## Verification **Behavioral qualification remains pending.** Original legacy ACP Claude loses the issue document only in the reduced cohort beneath an unchanged credential failure. A source-backed diagnosis finds that neither ordinary assignment reads the staged operational skill, while the runtime persists provider environment in session state. The new common repairs expose skill metadata/path and omit persisted env; strict document and credential guards stay intact. [Inspectable diagnosis and retained hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md). Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb` incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged #14961/#15007). All 222 affected adapter tests, adapter-utils/E2E typechecks, and final 96 variant/grader/retry calibrations pass. The frozen historical comparator is `c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths identical, exactly two production instruction paths plus nine declared unit expectations differ. The operational skill/discovery/environment repairs, selected model/profile/task/core grader/auth/permissions/retry policy are identical. Both actual launcher prepare→verify admissions pass with zero providers. [Immutable manifest and exact receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json). One original legacy ACP Claude cell per variant is authorized, with enforced single campaign attempts, 12-minute deadlines and company/agent 1,000-cent hard stops; every product recovery run/cost is counted. Actual live outcomes are pending. Current normal CI has one failed server shard and failed aggregate verify under diagnosis; other normal gates including typecheck/build/Rust/all eight browser shards pass. Fresh review completed successfully; the valid historical startup/resume masking finding was fixed with per-invocation task/chat checks and strict complete-snapshot capture, calibrated and resolved. Prior heads, failures and campaigns below remain historical evidence, not checks on this repair head. - Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52 current-head checks pass with two intentional Storybook skips, including repository typecheck/test/build and the browser shard. Fresh Greptile is 5/5 with zero unresolved review threads. Exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained receipt verification, zero providers/source errors. Fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck and 26-cell stock discovery pass. Canonical contract/inventory checks and the later issue-derived reference calibration are retained; that reference-only follow-up is not live-qualified by earlier frozen runs. - Prior full repository typecheck/build passed. The complete local Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three timing failures. All three affected files passed unchanged narrow reruns; original failures remain retained. Current-head CI now passes the full general checks; the original local failures remain retained. - The original 24-pair default-manual/shared-prompt comparison has two new overall classic Claude/OpenCode document-delivery failures plus an additional legacy ACP Claude document loss beneath an unchanged credential-guard failure (not closed by later runs), two newly passing OpenCode ordered cases, seven unchanged failures and 13 unchanged passes. Equal 15/24 totals do not establish behavioral equivalence. [Complete original report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md). - The skill-only repair holds the eight-word manual/shared prompts and merged #14920 fixed. All four matched profile configurations and 203 fixture/behavior files match. Candidate `abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill sources against baseline `bc83fe030234439ac51279502a28803958963e2e`. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547) and [baseline workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047) each pass 571 exact-source prerequisites before providers; all eight cells clean up successfully. Failed campaigns publish successfully and remain failed. - Repair pairs: Claude original Fail → Pass; Claude explicit Pass → Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff worsens beneath the unchanged failing UI-link grade: baseline gives a clickable API URL, candidate gives a code-formatted path without an anchor. The request's usable-link wording is narrower in the UI-only oracle. [Complete repair report and safe projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md). - The subsequent narrow stock metadata/link correction has two matched Pass → Pass cases, zero new machine failures/passes and no pending pairs. Both original-case handoff links remain deficient: candidate uses a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the preserved original oracle only requires a durable document. Both explicit clickable UI-link cases pass revision/content/link grading. All four exact-source 587-check gates, single assignment runs and cleanup pass. This does not establish fix causality because baseline also succeeds. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401) freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374) freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only comparison with reduced manuals/shared prompts held constant, not a repeat of the historical-manual comparison. Only original and clarified explicit classic OpenCode cases are selected, two per variant/four expected turns. 8,242 other tracked files and both profile hashes match; protected workflows admit each exact source before credentials. [Complete qualification report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md). Candidate original loads Paperclip/reference before saving publicly; baseline original loads it after writing locally, then saves publicly within the same assignment. Reported cost totals are $0.0107824490 candidate / $0.0107909015 baseline, with unmetered runtime. The later reference-only issue-derived link correction is provider-free calibrated and **not live-qualified** by these frozen runs; no further paid runs. - Retained tool calls show the repaired original OpenCode assignment loads only its assigned output skill before writing locally. Operational Paperclip is first loaded during automatic disposition recovery; its early recipe is visible then, but it never saves the missing document. Explicit candidate loads Paperclip and reads the new reference before saving successfully. All nine actual runs are counted. Reported LLM totals are $0.3802537209 baseline and $0.4918990161 candidate; local runtime is unmetered. - Initial setup, packaging, cancelled/missing-cell recovery, callback test and relative-output attempts remain retained. No completed provider failure was rerun. Frozen measurement branches are unchanged by later canonical metadata maintenance. - Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`, and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly for paid execution; it is excluded from `--all`. Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` replays this PR on merged hiring #14985 (`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four explicit custom-CEO-bundle checks, minimal generic manual boundary, and both suites. The combined fixture catalog and hiring calibrations pass 67 assertions; exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained-receipt verification, zero providers/source errors, fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E typecheck and 26-cell stock discovery pass. Fresh current-head CI passes all 52 checks with two intentional skips, and fresh Greptile is 5/5 with zero unresolved review threads. The prior source-plan browser failure is retained: a deterministic process fixture replayed its last `fixture:plan` command on `chat_task_completed`, writing revision 2 with identical body after the approval handoff. This was not paid provider execution. Rebased current-head CI passes the same assertion without an old-head retry or a change to that browser fixture. The merged hiring change was measured separately on immutable matched unions, with this reduced/shared/operational context and native completion guidance held constant. [Complete original two-profile report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md): [candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208) / [historical baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463), 705 provider-free prerequisites each. Both pairs are unchanged Fail → Fail on the exact-five count, with six core delivery checks passing all four cells; 28 actual successful runs include eight automatic completion wakes, zero retries, four successful cleanups. Source-read coverage is uncomparable, actual model charges unknown. Separately versioned provider-free accounting remains analytical work; original verdicts are preserved. This does not rerun or qualify the completed default-manual or native campaigns. ## Risks - Legacy ACP Claude's additional delivery loss is not closed by any later matched run and blocks the no-extra-failing-behavior merge criterion. Legacy document delivery may have relied on the prior manual/shared prompts. The early skill repair improves Claude in one trial; the later OpenCode pairs pass in both variants and cannot establish causality or robust recovery. Both original-case links remain deficient beneath the storage-only grade. The later issue-derived reference correction has only provider-free validation. Native finish/block descriptions must not be supplied to legacy agents. - The comparison holds merged native Codex fix #14920 constant; it cannot measure that fix's before/after task performance. - These bounded skill/context/chat workflows do not measure general coding quality. Unrepresented providers remain unqualified. - Saved manuals and old Codex sessions are not automatically migrated. Codex through ACP still has a separate base-instruction follow-up. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution, and delegated PR/eval work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (relevant suites and all three unchanged narrow reruns pass; complete-run timing failures retained in Verification) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green on the new repair head (prior-head checks retained above) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the new repair head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
339 lines
19 KiB
Markdown
339 lines
19 KiB
Markdown
# Runner E2E fixture authoring
|
|
|
|
The fixture catalog is executable production-contract data. Keep it small,
|
|
typed, deterministic, and free of raw credentials.
|
|
|
|
## Suites and matrices
|
|
|
|
A `RunnerSuiteFixture` declares one durable testing purpose: stable ID, label,
|
|
description, profiles, environments, cases, expected size, and definition or
|
|
ranking metadata. Its execution IDs are globally prefixed as
|
|
`<suite>.<profile>.<environment>.<case>`. Add a new suite when the testing
|
|
purpose or desired cross-product differs; do not inflate an existing suite with
|
|
unrelated dimensions.
|
|
|
|
The suite definition fingerprint is historical comparison metadata. Any
|
|
profile, model qualification, environment, task, or ranking-snapshot change
|
|
must change that fingerprint automatically so the dashboard can annotate the
|
|
boundary instead of silently joining unlike totals.
|
|
|
|
The explicit [stock-harness suite](STOCK-HARNESS.md) wraps existing profiles with
|
|
`productionDefaultHireProfile`: omit only `instructionsBundle` so the public
|
|
hire route loads the shipped default, while preserving runtime, permissions,
|
|
auth, skills, and managed secret references. Do not replace this with a fixture
|
|
copy of the default manual. Public receipts check the exact independently
|
|
specified bundle before provider execution and again during cleanup, along with
|
|
both budget hard stops and actual legacy invocation prompts. Missing evidence
|
|
fails closed. The definition fingerprint includes the helper, graders, journey
|
|
sources, live fixture, and execution integration; editing those sources changes
|
|
the suite revision automatically.
|
|
|
|
## Agent profiles
|
|
|
|
Add `RunnerProfileFixture` entries in `catalog.ts`. A profile declares:
|
|
|
|
- a stable ID and searchable groups;
|
|
- legacy or native generation;
|
|
- adapter/provider and required credential;
|
|
- a model imported from its adapter constant or qualified runner profile;
|
|
- supported environment IDs;
|
|
- expected runtime metadata; and
|
|
- an agent payload factory.
|
|
|
|
Do not duplicate model IDs, qualification decisions, CLI versions, or runner
|
|
artifact rules. Codex profiles import `DEFAULT_CODEX_LOCAL_MODEL`, OpenCode
|
|
profiles import `QUALIFIED_OPENCODE_MODEL`, and ACPX profiles import
|
|
`QUALIFIED_ACPX_PROFILES`. Add or qualify models at their owning production
|
|
source first.
|
|
|
|
OpenRouter breadth profiles are generated from `openrouter-models.json`, not
|
|
written by hand. That reviewed snapshot must contain exactly five unique,
|
|
available, tool-capable models with rank, canonical ID, display name, supported
|
|
parameters, source URL, capture time, and verified content hash. Refresh it
|
|
manually with `pnpm test:e2e:runner:models:update`; nightly campaigns never
|
|
change fixture definitions.
|
|
|
|
Agent `adapterConfig.env` values must be `{type:"secret_ref", secretId,
|
|
version:"latest"}` objects supplied to the factory. A fixture source containing
|
|
a raw secret-looking value is rejected by catalog validation.
|
|
|
|
The manual Grok subscription profile uses `GROK_AUTH_JSON` as an explicit login
|
|
fixture. It does not put this credential in agent configuration or substitute an
|
|
API key. Setup seeds a new company-scoped Grok home inside the disposable instance
|
|
with mode 0700 and an exclusive mode-0600 auth file. Setup rejects redirected,
|
|
occupied, or nonisolated homes. Production runner discovery and refresh operate on
|
|
that company login; teardown destroys it after the remote environment is removed.
|
|
This fixture tests subscription execution, not the interactive browser login flow.
|
|
|
|
## Environments
|
|
|
|
An `EnvironmentFixture` declares driver/provider, credential requirements,
|
|
attempt deadline, lifecycle behavior, expected execution target, and a payload
|
|
factory validated by the shared environment schema.
|
|
|
|
The local environment is instance-managed: company creation ensures it exists,
|
|
and the public API intentionally rejects a second local environment. The setup
|
|
registry therefore discovers that row through the public environments API.
|
|
This still provides full isolation because every cell starts a new Paperclip
|
|
instance and database.
|
|
|
|
Daytona creates sandbox environments through the public API. The core fixture
|
|
keeps `reuseLease:false` and `runnerLifecycleMode:"per_turn"`. The dedicated
|
|
warm-continuity fixture uses `reuseLease:true` and
|
|
`runnerLifecycleMode:"warm"`; its distinct `configurationKey` is part of the
|
|
suite fingerprint even though both fixtures report `environmentId:"daytona"`.
|
|
Keep short provider cleanup backstops, a Daytona secret reference, and an
|
|
immutable image digest. Teardown
|
|
must delete the environment with reusable-lease destruction and must fail the
|
|
cell if cleanup cannot be confirmed. Keep CPU, memory, and disk explicit: lease
|
|
metadata and the per-test public-list-price runtime estimate depend on that
|
|
pinned billable resource shape. Changing it requires updating billing tests and
|
|
reviewing the versioned Daytona rates in `billing.ts`.
|
|
|
|
## Usage and billing data
|
|
|
|
Do not add fixture-authored token or dollar expectations. The live harness
|
|
reads usage from selected public heartbeat-run records and records coverage per
|
|
run. Provider-reported dollars remain distinct from runtime estimates. A zero
|
|
or missing native usage payload is `unavailable` unless a real token-bearing
|
|
receipt or provider cost proves otherwise. New execution environments must
|
|
provide lease/resource metadata for a runtime estimate or explicitly remain
|
|
`unavailable`; never infer that missing billing data means free execution.
|
|
|
|
Future providers (SSH, E2B, Modal, Cloudflare, Kubernetes, Novita, exe.dev)
|
|
should implement the same setup/probe/cleanup contract before being added to a
|
|
matrix. Unsupported profile/environment combinations belong in
|
|
`supportedEnvironments`, not in ad hoc test conditionals.
|
|
|
|
## Task cases and matchers
|
|
|
|
A `RunnerTaskFixture` owns a work mode, a typed flow, expected run count,
|
|
nonce-based title/prompt/marker factories, per-environment attempt deadlines,
|
|
deterministic matchers, and expected terminal state. Single-turn prompts should
|
|
make one bounded request with observable output and no nondeterministic judging.
|
|
The `plan_revision_acceptance` flow must also provide revision-request and Plan
|
|
marker factories. `question_resume_completion` must define the deterministic
|
|
browser answer and prove exactly two successful runs with no pending
|
|
interaction. `plan_approval_completion` must target the exact two-step
|
|
canonical Plan revision, capture its pending UI, approve in the browser, and
|
|
prove exactly two successful runs. `warm_three_turn` provides exactly two
|
|
browser follow-up messages, preserves one project/execution-workspace scope,
|
|
verifies host file contents after every turn, and finishes within three
|
|
ten-minute turn deadlines. The ordinary warm fixture uses managed instructions,
|
|
updates AGENT_HOME each turn, and verifies memory, an unchanged 8 MiB binary and
|
|
a deletion through public file APIs. Native turns 2 and 3 must copy/hash only the
|
|
changed memory file, with a saved receipt and the same provider PID. Journal and
|
|
Git stress fixtures retain fixed external bundles as controls. Keep the stable-PID
|
|
oracle strict; `instruction-persistence` also covers cold restarts and quota handling.
|
|
Native turns 1 and 2 include an actionable human review in the completion report's `attentionRequests`. Paperclip creates the review gate from that report. An explicit question-tool wait yields the turn and suppresses its final prose, so it is not interchangeable with this completion-review fixture. Turn 3 reports Done without another review.
|
|
|
|
Every selected case runs in its own isolated Paperclip process, and independent
|
|
cases may run concurrently. Follow-up turns inside one case retain their shared
|
|
task state. Each case creates and tears down its own company, secrets,
|
|
environment selection, agent, and browser-created task. The current plan case
|
|
proves three runs on the same issue: publish a two-step Plan,
|
|
request a three-step revision through the UI, and accept the exact new revision
|
|
through the UI before verifying implementation and Done.
|
|
|
|
The matcher union supports message exact/contains/regex/ordered checks, issue
|
|
and run state, runtime/environment metadata, files, artifacts, JSON paths, and
|
|
JSON Schema. The initial cases use normalized `message_contains` plus state,
|
|
runtime, and environment assertions; the plan flow additionally verifies
|
|
canonical document revision IDs, bodies, step counts, interaction targets, and
|
|
visible previews. Add matcher behavior and credential-free tests together.
|
|
|
|
Adding a task expands its suite's matrix. Update the suite's intentional size,
|
|
the complete-catalog size, and credential-free unit tests in the same change.
|
|
Paid tests never silently skip a missing credential or unsupported artifact.
|
|
|
|
## Prompt-only task title fixtures
|
|
|
|
`task-titles.ts` defines a bounded ordinary writing request and an independent
|
|
title oracle. Its `single_turn` cases leave the title field empty or supply an
|
|
explicit control title. The harness captures the exact browser creation response
|
|
instead of searching by a title that the agent may already have changed. It
|
|
never patches the title itself. Normal production instructions own the early
|
|
naming behavior; fixture prompts and agent instruction bundles contain no naming
|
|
hints. Existing company/secret/environment/agent registry dependencies are reused,
|
|
with 500-cent company and agent budgets and normal instance teardown.
|
|
|
|
Keep the call input, successful result, execution receipt, saved task, and
|
|
agent/run-attributed audit correlated. Missing evidence must fail. The first-five
|
|
tool-call bound counts calls in the initial provider run, including discovery.
|
|
The title must describe API-key rotation without requiring one exact wording.
|
|
The control must retain its title throughout, not merely restore it at the end.
|
|
The source digest versions the grader and request in catalog metadata. See
|
|
[Automatic task titles](README.md#automatic-task-titles) for live selectors,
|
|
coverage limits, evidence, and calibration.
|
|
|
|
## New Paperclip object fixtures
|
|
|
|
The explicit-only `lifecycle-baseline` suite reuses this registry and existing
|
|
continuation, chat and governed-action flows. Its narrative pairs require actual
|
|
agent/run-attributed comments or exact visible responses. See
|
|
[the live baseline contract](LIFECYCLE-BASELINE.md) for selectors and proof boundaries.
|
|
|
|
Register new objects in `live-fixtures.ts` with explicit dependencies in
|
|
`FixtureRegistry`. Setup must use a public API. Teardown runs in reverse order
|
|
and is invoked after partial setup failures. Direct database writes and private
|
|
test-only runner endpoints are prohibited.
|
|
|
|
The expected dependency shape is:
|
|
|
|
```text
|
|
company
|
|
└── encrypted secrets
|
|
└── environment
|
|
└── agent
|
|
└── browser-created task
|
|
```
|
|
|
|
Projects, goals, apps, and configuration fixtures can be inserted into that
|
|
graph without changing the launcher. Keep returned fixture state to IDs and
|
|
sanitized metadata; never retain raw secret values.
|
|
|
|
## Required checks
|
|
|
|
Run before a fixture change is reviewed:
|
|
|
|
```bash
|
|
pnpm test:e2e:runner:unit
|
|
pnpm test:e2e:runner:typecheck
|
|
pnpm test:e2e:runner -- --list
|
|
```
|
|
|
|
Then run the narrowest paid cell that exercises the fixture. A full matrix is a
|
|
manual or scheduled campaign, not a PR requirement.
|
|
|
|
|
|
## Persistent chat fixtures
|
|
|
|
`chat-cases.ts` defines the eight-case `agent-chat` suite; `chat-flow.ts` drives the
|
|
production composer, plan revision/approval controls, questions, reset command,
|
|
and project cards. Keep its 28 local cells intentional. `expectedRunCount`
|
|
counts provider turns, including cancelled and handed-off task runs, but excludes
|
|
synthetic `/new` runs. Assertions must inspect all company runs because ordinary
|
|
issue lists exclude the source conversation. `assertChatHandoff` rejects missing
|
|
projects/plans, chat children, wrong assignees, and execution before plan commit.
|
|
|
|
Retained `api-state.json`, `chat-handoff.json`, and plan-revision evidence
|
|
include persisted comments, session generations, run context and logs, project
|
|
workspaces, task documents, and ordering. They pass through the normal sanitizer.
|
|
Screenshots are allowlisted to the exact disposable agent chat. Cleanup cancels
|
|
all active runs in the isolated company, including handed-off work; usage from
|
|
failed and cancelled runs must not disappear from campaign totals.
|
|
|
|
|
|
Warm three-turn continuity grades the exact workspace file after each turn,
|
|
task completion, and sandbox/session identity. It also requires a visible
|
|
persisted final reply with each turn marker once and in order. It does not
|
|
grade exact final-reply wording; the hello
|
|
and continuation fixtures retain those exact-response checks. This separates
|
|
workspace persistence failures from model response-format variance.
|
|
|
|
`chat-hardening.ts` adds the explicit-only `agent-chat-hardening` journeys. Use
|
|
the ordinary public APIs to seed source documents and blockers. Keep the answer
|
|
out of the user's status/review request. Grade the exact source values, latest
|
|
blocker, preserved task identities, worker-authored output, and real executions.
|
|
The status request asks for JSON so the grader can distinguish the current
|
|
blocker from a historical mention and compare active-run count separately from
|
|
task status. The request must not reveal those expected values.
|
|
Capture the source after seeding and compare every field in the public issue
|
|
update contract, plus labels, dependencies, and dedicated-endpoint settings.
|
|
Derived inbound references may change when the chat legitimately cites a task.
|
|
The lost-acknowledgement probe may interrupt only the fixture browser's own
|
|
comment request after the real server has committed it. Retain its request ID
|
|
and replay that same request through the public API after restarting the server.
|
|
Never fabricate tool results or repair task state after a failed assertion.
|
|
|
|
`chat-stories.ts` seeds an ordinary file wait in the isolated agent's actual
|
|
home workspace; native Codex intentionally cannot see arbitrary host temp files.
|
|
The observed run workspace must match the fixture location. This is a deterministic interruption
|
|
boundary. The real provider command writes the readiness file and waits at most
|
|
two minutes. The harness must persist the next browser message while the same
|
|
run is active before supplying the brief. Always release the wait in `finally`.
|
|
Save boundary observations independently of the final outcome. The final answer
|
|
must recover a brief reference absent from both prompts; the revision oracle
|
|
also reads the actual conversation plan. Fixture setup never enables native API
|
|
tools for this suite. Do not describe its prepared-agent settings case as a
|
|
production onboarding qualification.
|
|
|
|
The `agent-chat-qualification` local fixtures use public APIs to seed two workers
|
|
and a task with a saved plan, or read-only tasks with contradictory historical
|
|
comments. Ordinary Node file waits in the isolated agent workspace establish
|
|
observable active execution; no provider output or database outcome is fabricated.
|
|
A worker-crash case sends SIGKILL only to a positively identified running native
|
|
worker PID, then uses the production Retry button. Each gate is released in a
|
|
finally block. Source facts and boundary state are retained with the attempt.
|
|
The lifecycle suite also includes two legacy disposition-repair probes. Their
|
|
first provider turn intentionally omits task disposition, and their second turn
|
|
must be an automatic, causally bound repair that records completion. They use
|
|
public task comments/status APIs and run-detail evidence; no private runtime
|
|
hooks or database mutations are used by the fixture.
|
|
|
|
The explicit-only `extended-harnesses` suite uses five bounded journeys for each
|
|
pending ACP candidate on local and Daytona. Candidate profile metadata includes
|
|
the exact authenticated discovery choice without promoting it to a product
|
|
default. Its file case anchors the task to a public project workspace, validates
|
|
the model's claimed result by reading the actual final bytes, and also exercises
|
|
remote copy-back. Keep candidate admission scoped to the selected model and the
|
|
isolated operator environment; ordinary agent configuration must not enable it.
|
|
|
|
## Persistent agent files
|
|
|
|
The `instruction_persistence` flow uses production managed storage and public file
|
|
APIs. The browser creates a supporting file, then a real agent edits its registered
|
|
AGENT_HOME with ordinary filesystem tools. Independent oracles verify instructions,
|
|
nested text, binary download bytes, and a stopped-run save receipt without new
|
|
revision history. The harness restarts the server and creates a fresh browser task
|
|
without disclosing the saved nonces. Its readback oracle downloads and verifies an
|
|
attachment's bytes and SHA-256, rather than accepting a filename or model claim.
|
|
A third task uploads a ready attachment and waits in an ordinary bounded shell
|
|
command while the board changes the current file through the public API. Stopped
|
|
cleanup must preserve the original candidate as a conflict. The browser reviews
|
|
current and incoming files and applies the run edits against the reviewed current
|
|
directory hash. All three tasks' runs count toward billing and teardown. The suite
|
|
is explicit-only. No private control-plane hooks or direct database writes are used.
|
|
|
|
## Direct blocker fixtures
|
|
|
|
`blocker-cases.ts`, `blocker-fixtures.ts`, `blocker-flow.ts`, and
|
|
`blocker-scoring.ts` define the explicit local legacy `blocker-guidance` suite.
|
|
Its fixture registry creates a manager through the public API and assigns the
|
|
production operational skill to worker and manager. Company-wide evidence and
|
|
cleanup include unexpected manager runs. The grader checks saved human input,
|
|
requester identity for scope questions, ownership history, no additional work or
|
|
hires, and the browser-answer continuation. See [Direct blocker guidance](README.md#direct-blocker-guidance)
|
|
for coverage boundaries and run commands.
|
|
|
|
|
|
## Source-derived hiring template fixture
|
|
|
|
The explicit `hiring-templates` suite reuses the public company/agent, personal
|
|
managed account and browser chat fixtures. `hiringTemplateProfile` removes the
|
|
custom instruction bundle from the ordinary profile; the real agent creation
|
|
route selects the evaluated revision's CEO bundle. Keep its two local native
|
|
profiles, five expected runs and 15-minute deadline stable for paired runs.
|
|
`isManagedHiringCase` requests the account fixture and `chatNeedsApiTools`
|
|
enables only the existing API-tool path. It adds no private fixture endpoint,
|
|
provider fake or database write.
|
|
|
|
`hiring-template-flow.ts` reads the production instructions and company skill
|
|
files through public APIs before dispatch and verifies their hashes against the
|
|
checkout. The public run-events API supplies paginated read evidence after
|
|
execution. `hiring-template-scoring.ts` grades deterministic child documents,
|
|
actual worker identity/account, reuse and source coverage independently of the
|
|
agents' claims. `hiring-template.test.ts` calibrates production Codex/ACPX event
|
|
shapes, wrong/missing/late reads, incorrect/default bundles, source mismatch,
|
|
wrong hire/output, missing durable state, and an admissible historical four-file
|
|
CEO with a long coder role.
|
|
|
|
Preserve both dimensions in a comparison: `outcomePassed` describes the work;
|
|
`comparisonStatus` describes whether the expected sources and reads were proven.
|
|
Unprovable provider event shapes are coverage gaps. They must not become a
|
|
passing template comparison or a claimed behavior regression. The existing
|
|
report matcher paths carry the dimension and private final evidence carries the
|
|
explicit status. Provider runs are separately authorized; unit results establish
|
|
oracle calibration only. See the [suite contract](README.md#production-hiring-templates)
|
|
for evidence, budgets, cleanup and exact IDs.
|