Files
PaperClipAI/tests/runner-e2e/FIXTURES.md
T
DottaandPaperclip e34abee670 feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path

> - Paperclip gives teams durable tasks, agent execution, budgets, and
approvals.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - Those assistants need a scoped connection that preserves the
person’s permissions and attribution.
> - Delegating a task must not turn the assistant into the assigned
agent.
> - This PR adds opt-in user OAuth, ten first-party tools, browser
consent, and workflow packages.
> - Paid product evals verify the resulting tasks, documents,
attribution, retries, and access boundaries.
> - The team keeps working after the assistant conversation ends.

## Linked Issues or Issue Description

**Problem or motivation**

A person cannot connect an external assistant to an existing team
through browser consent and safely delegate durable work as themselves.

**Proposed solution**

Expose an opt-in `/mcp/paperclip` endpoint with individually described
first-party operations. Bind every connection to a person, client,
company, resource, and scopes. Reuse domain authorization and
scheduling. Package shared team-review, delegation, and follow-up
workflows for OpenAI/Codex and Claude.

**Alternatives considered**

Related PRs #9393 and #12549 cover earlier remote MCP and board-operator
approaches. This change uses user OAuth and a bounded public catalog. It
does not expose a generic executor, operator administration, static
shared board credentials, or external agent execution. Registry listing
work in #9851 is a separate distribution step.

**Roadmap alignment**

This maintainer-requested implementation extends the governed MCP
gateway, activity attribution, durable work products, and hosted
deployment direction in `ROADMAP.md`. It implements the first release of
the saved design plan; external agent participation and granted
third-party tools remain later releases.

## What Changed

- Add MCP 2.0 discovery and task status/comment/document Events on the
same authenticated endpoint. Persist subscriptions and delivery
receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt
callback material, recheck permissions/Cloud membership, and bound
retries/expiry. Older MCP clients keep their existing tools.
- Add discovery, dynamic client registration, S256 PKCE, resource
validation, rotating refresh tokens, revocation, and company consent.
Store credentials as hashes and recheck membership at execution.
- Add tools for connection identity, agents/projects, task
search/read/create, human comments, documents/deliverables, and
pending-approval links. Preserve current domain permissions and
scheduling.
- Add durable mutation receipts across reconnects. Matching retries
replay results; uncertain outcomes keep the same request ID and require
inspection.
- Add consent and connection-management pages, OAuth log redaction,
shared plugin workflows, and separate OpenAI/Codex and Claude package
outputs.
- Add eight paid Product E2E cases across three models, independent
durable-state grading, usage evidence, cleanup, and report integration.
Add task-document guidance and regenerate the runner capability
inventories.
- Add migrations 0301 and 0302, the dated implementation plan, result
notes, and direct-client setup instructions in `doc/public-mcp.md`.

## Verification

- Merge integration `e180b1948`: resolved conflicts with current master,
preserved both eval registries, regenerated capability catalogs, and
regenerated migrations as 0301/0302 while keeping the original
replay-safe SQL byte-identical. Local migration safety/snapshot tests
(26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval
catalog/grading tests (198) pass. Token and capability gates pass. Full
recursive typecheck passed. Fresh Greptile review is 5/5 with no
unresolved findings. CI is green on this exact head (55 successes, two
intentional skips, one neutral result): one unchanged Cursor sandbox
test timed out at 10 seconds, then passed locally in 856 ms. A single
retry of that failed shard and the aggregate workflow passed. Merge
remains blocked on the repository code-owner approval rule.

Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55
successes, two intentional skips and one neutral result. [The earlier CI
run](https://github.com/paperclipai/paperclip/actions/runs/36901592350)
includes all test shards, browser tests, typecheck, build and canary dry
run. Greptile was 5/5 on that commit with no unresolved review threads.
GitHub still requires code-owner review under the repository merge
rules; passing checks do not bypass that approval. Paid source
fingerprints remain separate below and in the dated result note.

- Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku
4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback,
signature verification and report retrieval in a fresh conversation. A
final Mini regression passes after the quota/status fixes. All evidence
validates. Bounded tunnel startup retries occur before provider calls
and remain visible; failed earlier attempts retain their original
grades.
- The earlier complete seven-case matrix passes **21/21**, with a
separate **3/3** delegation regression. Two preceding matrices also
passed 21/21 each. A complete 24-cell matrix including Events has not
been run. [The dated
results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain
exact source fingerprints, failures, model IDs and partial costs.
- Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass
after merging master. Server typecheck passes after the final
quota/status changes. Eval typecheck and all 892 eval-support tests
pass.
- All 33 real MCP/OAuth tests pass. The preceding combined MCP,
redaction, private-address and DNS-rebinding run passed 129 tests; two
later MCP regressions cover quota reuse and unchanged-status
suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI
then found a null checkout result in the existing concurrent-workspace
path; logging now uses optional status access. All 12 closed-workspace
tests and all 33 MCP tests pass after that correction. The exact-start
event calibration exposed a timestamp gap; scanning now includes the
subscription start, with all 33 MCP tests and server typecheck passing.
These two narrow corrections follow the paid regression.
- A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0
subscription/delivery, current membership loss, unsubscribe, legacy SDK
tools, refresh and revocation. Its callback transport is a fixture with
independent HMAC verification. The paid Events campaigns separately
prove public HTTPS delivery.
- Earlier component qualification passed UI 7,117 tests, CLI 502, shared
832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module
boundaries, migration order and plugin regeneration passed. CI covers
general/serialized suites, eight browser shards, runner checks,
typecheck, build and canary dry run.
- **Local full-suite limitation:** the earlier monolithic run was not
clean. It encountered overlapping schema rebuilding, Mac database
shared-memory limits and isolated CLI/fixture failures. Targeted reruns
passed. The existing >32 MiB Git filename stress test still hit its
300-second Mac timeout. The additional serialized sweep stopped after 62
passing suites once CI passed. Original failures and partial logs
remain; this PR does not claim a wholly green local monolithic run.
- Local Codex CLI and Claude Code OAuth login and MCP SDK
interoperability were verified. Public-store installation, actual
ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted
newcomer provisioning remain release gates.

Enablement is moving to **Settings → Experimental → Assistant
connections (MCP)** in the stacked follow-up
[#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge
both for the intended setup experience. This foundation branch alone
still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set
`PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and
connect to `/mcp/paperclip`. Select a team and allow writes in browser
consent. Configure an available agent and budget, then delegate and
retrieve results later. For Events, rescan the deployed plugin catalog
in ChatGPT Work Cloud; the host supplies its webhook credentials when
the user asks to watch a task. See [the setup
runbook](doc/public-mcp.md).

## Risks

- Events are at-least-once and may arrive out of order. No replay cursor
is advertised. Clients must refresh finite subscriptions, read current
state and avoid comment feedback loops. Callback material uses the
instance secrets master key; hosted subscriptions require the updated
Cloud broker and are bounded to five minutes/the access proof expiry.
- ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging
subscription remain deployment gates. Local signed-webhook and paid
model evidence does not claim those surfaces have been exercised.
- Disabled by default. Merging adds schema and opt-in code; it does not
deploy a public endpoint, publish a store listing, create a team, or
start paid agents.
- Migrations 0301 and 0302 are additive and idempotent. Their SQL is
unchanged from the earlier preview numbers, so hash-aware upgrade
reconciliation preserves prior staging applications. Normal instance
upgrades must apply it before enabling MCP.
- Task creation and comments can schedule paid agent work. Consent and
tool descriptions disclose that effect. Revocation blocks future calls
but does not undo delegated work.
- Public deployments need edge rate limits and credential-safe logging.
Internal dispatch is restricted to the closed catalog and carries a
request-local verified actor.
- Hosted onboarding requires the companion Cloud broker, encryption-key
configuration, and tenant rollout. Self-hosted direct connections can
use this PR alone.
- Store acceptance and agent-mode participation are not claimed.
Checked-in plugin endpoints are development defaults; rebuild packages
for a real deployment before installation.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A
more specific serving version and context-window size were not exposed
by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`,
`claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted/component checks;
full local-run limitations are recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 11:48:53 -05:00

20 KiB

Runner E2E fixture authoring

The public MCP journeys reuse the fixture registry with a real authenticated browser session. RunnerApi.setBrowserSession binds that session to API calls, including encrypted secret provisioning via Node fetch. OAuth setup stays outside model context. The external assistant receives the catalog from tools/list and the shipped workflow skills; its calls execute against the real SDK transport. Client-side loss of a successful response is the sole fault injection in the uncertain-retry case. Task/run/document REST reads own grading.

The fixture catalog is executable production-contract data. Keep it small, typed, deterministic, and free of raw credentials.

Suites and matrices

A RunnerSuiteFixture declares one durable testing purpose: stable ID, label, description, profiles, environments, cases, expected size, and definition or ranking metadata. Its execution IDs are globally prefixed as <suite>.<profile>.<environment>.<case>. Add a new suite when the testing purpose or desired cross-product differs; do not inflate an existing suite with unrelated dimensions.

The suite definition fingerprint is historical comparison metadata. Any profile, model qualification, environment, task, or ranking-snapshot change must change that fingerprint automatically so the dashboard can annotate the boundary instead of silently joining unlike totals.

The explicit stock-harness suite wraps existing profiles with productionDefaultHireProfile: omit only instructionsBundle so the public hire route loads the shipped default, while preserving runtime, permissions, auth, skills, and managed secret references. Do not replace this with a fixture copy of the default manual. Public receipts check the exact independently specified bundle before provider execution and again during cleanup, along with both budget hard stops and actual legacy invocation prompts. Missing evidence fails closed. The definition fingerprint includes the helper, graders, journey sources, live fixture, and execution integration; editing those sources changes the suite revision automatically.

Agent profiles

Add RunnerProfileFixture entries in catalog.ts. A profile declares:

  • a stable ID and searchable groups;
  • legacy or native generation;
  • adapter/provider and required credential;
  • a model imported from its adapter constant or qualified runner profile;
  • supported environment IDs;
  • expected runtime metadata; and
  • an agent payload factory.

Do not duplicate model IDs, qualification decisions, CLI versions, or runner artifact rules. Codex profiles import DEFAULT_CODEX_LOCAL_MODEL, OpenCode profiles import QUALIFIED_OPENCODE_MODEL, and ACPX profiles import QUALIFIED_ACPX_PROFILES. Add or qualify models at their owning production source first.

OpenRouter breadth profiles are generated from openrouter-models.json, not written by hand. That reviewed snapshot must contain exactly five unique, available, tool-capable models with rank, canonical ID, display name, supported parameters, source URL, capture time, and verified content hash. Refresh it manually with pnpm test:e2e:runner:models:update; nightly campaigns never change fixture definitions.

Agent adapterConfig.env values must be {type:"secret_ref", secretId, version:"latest"} objects supplied to the factory. A fixture source containing a raw secret-looking value is rejected by catalog validation.

The manual Grok subscription profile uses GROK_AUTH_JSON as an explicit login fixture. It does not put this credential in agent configuration or substitute an API key. Setup seeds a new company-scoped Grok home inside the disposable instance with mode 0700 and an exclusive mode-0600 auth file. Setup rejects redirected, occupied, or nonisolated homes. Production runner discovery and refresh operate on that company login; teardown destroys it after the remote environment is removed. This fixture tests subscription execution, not the interactive browser login flow.

Environments

An EnvironmentFixture declares driver/provider, credential requirements, attempt deadline, lifecycle behavior, expected execution target, and a payload factory validated by the shared environment schema.

The local environment is instance-managed: company creation ensures it exists, and the public API intentionally rejects a second local environment. The setup registry therefore discovers that row through the public environments API. This still provides full isolation because every cell starts a new Paperclip instance and database.

Daytona creates sandbox environments through the public API. The core fixture keeps reuseLease:false and runnerLifecycleMode:"per_turn". The dedicated warm-continuity fixture uses reuseLease:true and runnerLifecycleMode:"warm"; its distinct configurationKey is part of the suite fingerprint even though both fixtures report environmentId:"daytona". Keep short provider cleanup backstops, a Daytona secret reference, and an immutable image digest. Teardown must delete the environment with reusable-lease destruction and must fail the cell if cleanup cannot be confirmed. Keep CPU, memory, and disk explicit: lease metadata and the per-test public-list-price runtime estimate depend on that pinned billable resource shape. Changing it requires updating billing tests and reviewing the versioned Daytona rates in billing.ts.

Usage and billing data

Do not add fixture-authored token or dollar expectations. The live harness reads usage from selected public heartbeat-run records and records coverage per run. Provider-reported dollars remain distinct from runtime estimates. A zero or missing native usage payload is unavailable unless a real token-bearing receipt or provider cost proves otherwise. New execution environments must provide lease/resource metadata for a runtime estimate or explicitly remain unavailable; never infer that missing billing data means free execution.

Future providers (SSH, E2B, Modal, Cloudflare, Kubernetes, Novita, exe.dev) should implement the same setup/probe/cleanup contract before being added to a matrix. Unsupported profile/environment combinations belong in supportedEnvironments, not in ad hoc test conditionals.

Task cases and matchers

A RunnerTaskFixture owns a work mode, a typed flow, expected run count, nonce-based title/prompt/marker factories, per-environment attempt deadlines, deterministic matchers, and expected terminal state. Single-turn prompts should make one bounded request with observable output and no nondeterministic judging. The plan_revision_acceptance flow must also provide revision-request and Plan marker factories. question_resume_completion must define the deterministic browser answer and prove exactly two successful runs with no pending interaction. plan_approval_completion must target the exact two-step canonical Plan revision, capture its pending UI, approve in the browser, and prove exactly two successful runs. warm_three_turn provides exactly two browser follow-up messages, preserves one project/execution-workspace scope, verifies host file contents after every turn, and finishes within three ten-minute turn deadlines. The ordinary warm fixture uses managed instructions, updates AGENT_HOME each turn, and verifies memory, an unchanged 8 MiB binary and a deletion through public file APIs. Native turns 2 and 3 must copy/hash only the changed memory file, with a saved receipt and the same provider PID. Journal and Git stress fixtures retain fixed external bundles as controls. Keep the stable-PID oracle strict; instruction-persistence also covers cold restarts and quota handling. Native turns 1 and 2 include an actionable human review in the completion report's attentionRequests. Paperclip creates the review gate from that report. An explicit question-tool wait yields the turn and suppresses its final prose, so it is not interchangeable with this completion-review fixture. Turn 3 reports Done without another review.

Every selected case runs in its own isolated Paperclip process, and independent cases may run concurrently. Follow-up turns inside one case retain their shared task state. Each case creates and tears down its own company, secrets, environment selection, agent, and browser-created task. The current plan case proves three runs on the same issue: publish a two-step Plan, request a three-step revision through the UI, and accept the exact new revision through the UI before verifying implementation and Done.

The matcher union supports message exact/contains/regex/ordered checks, issue and run state, runtime/environment metadata, files, artifacts, JSON paths, and JSON Schema. The initial cases use normalized message_contains plus state, runtime, and environment assertions; the plan flow additionally verifies canonical document revision IDs, bodies, step counts, interaction targets, and visible previews. Add matcher behavior and credential-free tests together.

Adding a task expands its suite's matrix. Update the suite's intentional size, the complete-catalog size, and credential-free unit tests in the same change. Paid tests never silently skip a missing credential or unsupported artifact.

Prompt-only task title fixtures

task-titles.ts defines a bounded ordinary writing request and an independent title oracle. Its single_turn cases leave the title field empty or supply an explicit control title. The harness captures the exact browser creation response instead of searching by a title that the agent may already have changed. It never patches the title itself. Normal production instructions own the early naming behavior; fixture prompts and agent instruction bundles contain no naming hints. Existing company/secret/environment/agent registry dependencies are reused, with 500-cent company and agent budgets and normal instance teardown.

Keep the call input, successful result, execution receipt, saved task, and agent/run-attributed audit correlated. Missing evidence must fail. The first-five tool-call bound counts calls in the initial provider run, including discovery. The title must describe API-key rotation without requiring one exact wording. The control must retain its title throughout, not merely restore it at the end. The source digest versions the grader and request in catalog metadata. See Automatic task titles for live selectors, coverage limits, evidence, and calibration.

New Paperclip object fixtures

The explicit-only lifecycle-baseline suite reuses this registry and existing continuation, chat and governed-action flows. Its narrative pairs require actual agent/run-attributed comments or exact visible responses. See the live baseline contract for selectors and proof boundaries.

Register new objects in live-fixtures.ts with explicit dependencies in FixtureRegistry. Setup must use a public API. Teardown runs in reverse order and is invoked after partial setup failures. Direct database writes and private test-only runner endpoints are prohibited.

The expected dependency shape is:

company
└── encrypted secrets
    └── environment
        └── agent
            └── browser-created task

Projects, goals, apps, and configuration fixtures can be inserted into that graph without changing the launcher. Keep returned fixture state to IDs and sanitized metadata; never retain raw secret values.

Required checks

Run before a fixture change is reviewed:

pnpm test:e2e:runner:unit
pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner -- --list

Then run the narrowest paid cell that exercises the fixture. A full matrix is a manual or scheduled campaign, not a PR requirement.

Persistent chat fixtures

chat-cases.ts defines the eight-case agent-chat suite; chat-flow.ts drives the production composer, plan revision/approval controls, questions, reset command, and project cards. Keep its 28 local cells intentional. expectedRunCount counts provider turns, including cancelled and handed-off task runs, but excludes synthetic /new runs. Assertions must inspect all company runs because ordinary issue lists exclude the source conversation. assertChatHandoff rejects missing projects/plans, chat children, wrong assignees, and execution before plan commit.

Retained api-state.json, chat-handoff.json, and plan-revision evidence include persisted comments, session generations, run context and logs, project workspaces, task documents, and ordering. They pass through the normal sanitizer. Screenshots are allowlisted to the exact disposable agent chat. Cleanup cancels all active runs in the isolated company, including handed-off work; usage from failed and cancelled runs must not disappear from campaign totals.

Warm three-turn continuity grades the exact workspace file after each turn, task completion, and sandbox/session identity. It also requires a visible persisted final reply with each turn marker once and in order. It does not grade exact final-reply wording; the hello and continuation fixtures retain those exact-response checks. This separates workspace persistence failures from model response-format variance.

chat-hardening.ts adds the explicit-only agent-chat-hardening journeys. Use the ordinary public APIs to seed source documents and blockers. Keep the answer out of the user's status/review request. Grade the exact source values, latest blocker, preserved task identities, worker-authored output, and real executions. The status request asks for JSON so the grader can distinguish the current blocker from a historical mention and compare active-run count separately from task status. The request must not reveal those expected values. Capture the source after seeding and compare every field in the public issue update contract, plus labels, dependencies, and dedicated-endpoint settings. Derived inbound references may change when the chat legitimately cites a task. The lost-acknowledgement probe may interrupt only the fixture browser's own comment request after the real server has committed it. Retain its request ID and replay that same request through the public API after restarting the server. Never fabricate tool results or repair task state after a failed assertion.

chat-stories.ts seeds an ordinary file wait in the isolated agent's actual home workspace; native Codex intentionally cannot see arbitrary host temp files. The observed run workspace must match the fixture location. This is a deterministic interruption boundary. The real provider command writes the readiness file and waits at most two minutes. The harness must persist the next browser message while the same run is active before supplying the brief. Always release the wait in finally. Save boundary observations independently of the final outcome. The final answer must recover a brief reference absent from both prompts; the revision oracle also reads the actual conversation plan. Fixture setup never enables native API tools for this suite. Do not describe its prepared-agent settings case as a production onboarding qualification.

The agent-chat-qualification local fixtures use public APIs to seed two workers and a task with a saved plan, or read-only tasks with contradictory historical comments. Ordinary Node file waits in the isolated agent workspace establish observable active execution; no provider output or database outcome is fabricated. A worker-crash case sends SIGKILL only to a positively identified running native worker PID, then uses the production Retry button. Each gate is released in a finally block. Source facts and boundary state are retained with the attempt. The lifecycle suite also includes two legacy disposition-repair probes. Their first provider turn intentionally omits task disposition, and their second turn must be an automatic, causally bound repair that records completion. They use public task comments/status APIs and run-detail evidence; no private runtime hooks or database mutations are used by the fixture.

The explicit-only extended-harnesses suite uses five bounded journeys for each pending ACP candidate on local and Daytona. Candidate profile metadata includes the exact authenticated discovery choice without promoting it to a product default. Its file case anchors the task to a public project workspace, validates the model's claimed result by reading the actual final bytes, and also exercises remote copy-back. Keep candidate admission scoped to the selected model and the isolated operator environment; ordinary agent configuration must not enable it.

Persistent agent files

The instruction_persistence flow uses production managed storage and public file APIs. The browser creates a supporting file, then a real agent edits its registered AGENT_HOME with ordinary filesystem tools. Independent oracles verify instructions, nested text, binary download bytes, and a stopped-run save receipt without new revision history. The harness restarts the server and creates a fresh browser task without disclosing the saved nonces. Its readback oracle downloads and verifies an attachment's bytes and SHA-256, rather than accepting a filename or model claim. A third task uploads a ready attachment and waits in an ordinary bounded shell command while the board changes the current file through the public API. Stopped cleanup must preserve the original candidate as a conflict. The browser reviews current and incoming files and applies the run edits against the reviewed current directory hash. All three tasks' runs count toward billing and teardown. The suite is explicit-only. No private control-plane hooks or direct database writes are used.

Direct blocker fixtures

blocker-cases.ts, blocker-fixtures.ts, blocker-flow.ts, and blocker-scoring.ts define the explicit local legacy blocker-guidance suite. Its fixture registry creates a manager through the public API and assigns the production operational skill to worker and manager. Company-wide evidence and cleanup include unexpected manager runs. The grader checks saved human input, requester identity for scope questions, ownership history, no additional work or hires, and the browser-answer continuation. See Direct blocker guidance for coverage boundaries and run commands.

Source-derived hiring template fixture

The explicit hiring-templates suite reuses the public company/agent, personal managed account and browser chat fixtures. hiringTemplateProfile removes the custom instruction bundle from the ordinary profile; the real agent creation route selects the evaluated revision's CEO bundle. Keep its two local native profiles, five expected runs and 15-minute deadline stable for paired runs. isManagedHiringCase requests the account fixture and chatNeedsApiTools enables only the existing API-tool path. It adds no private fixture endpoint, provider fake or database write.

hiring-template-flow.ts reads the production instructions and company skill files through public APIs before dispatch and verifies their hashes against the checkout. The public run-events API supplies paginated read evidence after execution. hiring-template-scoring.ts grades deterministic child documents, actual worker identity/account, reuse and source coverage independently of the agents' claims. hiring-template.test.ts calibrates production Codex/ACPX event shapes, wrong/missing/late reads, incorrect/default bundles, source mismatch, wrong hire/output, missing durable state, and an admissible historical four-file CEO with a long coder role.

Preserve both dimensions in a comparison: outcomePassed describes the work; comparisonStatus describes whether the expected sources and reads were proven. Unprovable provider event shapes are coverage gaps. They must not become a passing template comparison or a claimed behavior regression. The existing report matcher paths carry the dimension and private final evidence carries the explicit status. Provider runs are separately authorized; unit results establish oracle calibration only. See the suite contract for evidence, budgets, cleanup and exact IDs.