Files
DottaandPaperclip b17019e14d fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its adapters supply task context and access to Paperclip skills and
tools.
> - The default hire manual and shared prompts also repeat general work
procedures.
> - Those procedures overlap with stock provider instructions and the
Paperclip skill.
> - Existing E2E fixtures supply a QA manual, so they do not qualify the
production default.
> - This pull request reduces the generic instructions and adds real
default-hire coverage.
> - The benefit is less competing guidance, with inspectable evidence
for preserved skills and task context.

## Linked Issues or Issue Description

Refs: #14920. That merged change preserves native Codex base
instructions. This PR covers the default manual, shared legacy prompts,
operational skill guidance, and the narrowly approved ACP
skill-discovery/session-environment repair for measured delivery and
credential-persistence failures.

**What existing behavior does this improve?**

New non-CEO hires without a custom bundle and legacy task/chat startup
and continuation prompts.

**Current behavior**

The shipped default manual contains 602 words. Generic task/chat prompts
and ordinary resume deltas repeat work procedures already available
through the harness and Paperclip skill.

**Proposed behavior**

The default manual contains only the eight-word company identity. Shared
startup prompts retain identity and connection guidance. Ordinary resume
deltas retain current work context without the generic execution
contract.

**Reason and benefit**

Let the stock harness guide general work. Keep Paperclip-specific
capabilities and independently test default hires, skills, ordered
comments, and chat restart.

**Breaking changes**

New default hires receive less guidance. Existing saved manuals,
explicit custom bundles, CEO templates, and specialized wake contracts
retain their behavior. The obsolete includeExecutionContract option
remains accepted for source compatibility.

## What Changed

- Reduce the default hire manual to one sentence.
- Reduce shared task/chat defaults and remove the generic
ordinary-resume contract.
- Keep connection guidance, auth, skills, custom prompts, and
specialized wake context.
- Add credential-free instruction-boundary gates and 26 explicit Product
E2E cells across eight legacy/native profiles, including two focused
Paperclip-storage cases.
- Capture public hire receipts before providers run, then grade
delivered prompts and independent task/chat outcomes.
- Add an early legacy skill API recipe for saving a task document,
checking the saved revision receipt and linking the document. Improve
stock task/heartbeat skill-selection metadata and show a clickable
Markdown UI-link example. Keep native tool completion separate.
- Advertise bounded routing descriptions and exact successfully staged
SKILL.md paths in legacy ACP Claude; keep full bodies on demand and
preserve remote path rebasing.
- Remove only the provider environment from copied persisted ACP session
records, while loading current run credentials and preserving all other
options/conversation state.
- Regenerate both capability metadata inventories and reject stale
manifests/inventories before provider admission.
- Publish the original reduction and focused skill-repair comparisons,
preserving all failures, automatic recovery, cost coverage and
limitations.

## Verification

**Behavioral qualification remains pending.** Original legacy ACP Claude
loses the issue document only in the reduced cohort beneath an unchanged
credential failure. A source-backed diagnosis finds that neither
ordinary assignment reads the staged operational skill, while the
runtime persists provider environment in session state. The new common
repairs expose skill metadata/path and omit persisted env; strict
document and credential guards stay intact. [Inspectable diagnosis and
retained
hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md).

Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb`
incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged
#14961/#15007). All 222 affected adapter tests, adapter-utils/E2E
typechecks, and final 96 variant/grader/retry calibrations pass. The
frozen historical comparator is
`c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths
identical, exactly two production instruction paths plus nine declared
unit expectations differ. The operational skill/discovery/environment
repairs, selected model/profile/task/core grader/auth/permissions/retry
policy are identical. Both actual launcher prepare→verify admissions
pass with zero providers. [Immutable manifest and exact
receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json).

One original legacy ACP Claude cell per variant is authorized, with
enforced single campaign attempts, 12-minute deadlines and company/agent
1,000-cent hard stops; every product recovery run/cost is counted.
Actual live outcomes are pending. Current normal CI has one failed
server shard and failed aggregate verify under diagnosis; other normal
gates including typecheck/build/Rust/all eight browser shards pass.
Fresh review completed successfully; the valid historical startup/resume
masking finding was fixed with per-invocation task/chat checks and
strict complete-snapshot capture, calibrated and resolved. Prior heads,
failures and campaigns below remain historical evidence, not checks on
this repair head.

- Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on
merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52
current-head checks pass with two intentional Storybook skips, including
repository typecheck/test/build and the browser shard. Fresh Greptile is
5/5 with zero unresolved review threads. Exact-head stock prerequisites
pass 599 assertions (598 TypeScript + 1 Rust), all six gates and
retained receipt verification, zero providers/source errors. Fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`.
Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck
and 26-cell stock discovery pass. Canonical contract/inventory checks
and the later issue-derived reference calibration are retained; that
reference-only follow-up is not live-qualified by earlier frozen runs.
- Prior full repository typecheck/build passed. The complete local
Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three
timing failures. All three affected files passed unchanged narrow
reruns; original failures remain retained. Current-head CI now passes
the full general checks; the original local failures remain retained.
- The original 24-pair default-manual/shared-prompt comparison has two
new overall classic Claude/OpenCode document-delivery failures plus an
additional legacy ACP Claude document loss beneath an unchanged
credential-guard failure (not closed by later runs), two newly passing
OpenCode ordered cases, seven unchanged failures and 13 unchanged
passes. Equal 15/24 totals do not establish behavioral equivalence.
[Complete original
report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md).
- The skill-only repair holds the eight-word manual/shared prompts and
merged #14920 fixed. All four matched profile configurations and 203
fixture/behavior files match. Candidate
`abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill
sources against baseline `bc83fe030234439ac51279502a28803958963e2e`.
[Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547)
and [baseline
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047)
each pass 571 exact-source prerequisites before providers; all eight
cells clean up successfully. Failed campaigns publish successfully and
remain failed.
- Repair pairs: Claude original Fail → Pass; Claude explicit Pass →
Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff
worsens beneath the unchanged failing UI-link grade: baseline gives a
clickable API URL, candidate gives a code-formatted path without an
anchor. The request's usable-link wording is narrower in the UI-only
oracle. [Complete repair report and safe
projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md).
- The subsequent narrow stock metadata/link correction has two matched
Pass → Pass cases, zero new machine failures/passes and no pending
pairs. Both original-case handoff links remain deficient: candidate uses
a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the
preserved original oracle only requires a durable document. Both
explicit clickable UI-link cases pass revision/content/link grading. All
four exact-source 587-check gates, single assignment runs and cleanup
pass. This does not establish fix causality because baseline also
succeeds. [Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401)
freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched
baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374)
freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only
comparison with reduced manuals/shared prompts held constant, not a
repeat of the historical-manual comparison. Only original and clarified
explicit classic OpenCode cases are selected, two per variant/four
expected turns. 8,242 other tracked files and both profile hashes match;
protected workflows admit each exact source before credentials.
[Complete qualification
report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md).
Candidate original loads Paperclip/reference before saving publicly;
baseline original loads it after writing locally, then saves publicly
within the same assignment. Reported cost totals are $0.0107824490
candidate / $0.0107909015 baseline, with unmetered runtime. The later
reference-only issue-derived link correction is provider-free calibrated
and **not live-qualified** by these frozen runs; no further paid runs.
- Retained tool calls show the repaired original OpenCode assignment
loads only its assigned output skill before writing locally. Operational
Paperclip is first loaded during automatic disposition recovery; its
early recipe is visible then, but it never saves the missing document.
Explicit candidate loads Paperclip and reads the new reference before
saving successfully. All nine actual runs are counted. Reported LLM
totals are $0.3802537209 baseline and $0.4918990161 candidate; local
runtime is unmetered.
- Initial setup, packaging, cancelled/missing-cell recovery, callback
test and relative-output attempts remain retained. No completed provider
failure was rerun. Frozen measurement branches are unchanged by later
canonical metadata maintenance.
- Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`,
and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly
for paid execution; it is excluded from `--all`.

Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f`
replays this PR on merged hiring #14985
(`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four
explicit custom-CEO-bundle checks, minimal generic manual boundary, and
both suites. The combined fixture catalog and hiring calibrations pass
67 assertions; exact-head stock prerequisites pass 599 assertions (598
TypeScript + 1 Rust), all six gates and retained-receipt verification,
zero providers/source errors, fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E
typecheck and 26-cell stock discovery pass. Fresh current-head CI passes
all 52 checks with two intentional skips, and fresh Greptile is 5/5 with
zero unresolved review threads.

The prior source-plan browser failure is retained: a deterministic
process fixture replayed its last `fixture:plan` command on
`chat_task_completed`, writing revision 2 with identical body after the
approval handoff. This was not paid provider execution. Rebased
current-head CI passes the same assertion without an old-head retry or a
change to that browser fixture.

The merged hiring change was measured separately on immutable matched
unions, with this reduced/shared/operational context and native
completion guidance held constant. [Complete original two-profile
report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md):
[candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208)
/ [historical
baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463),
705 provider-free prerequisites each. Both pairs are unchanged Fail →
Fail on the exact-five count, with six core delivery checks passing all
four cells; 28 actual successful runs include eight automatic completion
wakes, zero retries, four successful cleanups. Source-read coverage is
uncomparable, actual model charges unknown. Separately versioned
provider-free accounting remains analytical work; original verdicts are
preserved. This does not rerun or qualify the completed default-manual
or native campaigns.

## Risks

- Legacy ACP Claude's additional delivery loss is not closed by any
later matched run and blocks the no-extra-failing-behavior merge
criterion. Legacy document delivery may have relied on the prior
manual/shared prompts. The early skill repair improves Claude in one
trial; the later OpenCode pairs pass in both variants and cannot
establish causality or robust recovery. Both original-case links remain
deficient beneath the storage-only grade. The later issue-derived
reference correction has only provider-free validation. Native
finish/block descriptions must not be supplied to legacy agents.
- The comparison holds merged native Codex fix #14920 constant; it
cannot measure that fix's before/after task performance.
- These bounded skill/context/chat workflows do not measure general
coding quality. Unrepresented providers remain unqualified.
- Saved manuals and old Codex sessions are not automatically migrated.
Codex through ACP still has a separate base-instruction follow-up.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution, and delegated PR/eval
work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites and all
three unchanged narrow reruns pass; complete-run timing failures
retained in Verification)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green on the new repair head
(prior-head checks retained above)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
on the new repair head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:32:42 -05:00

17 KiB

Stock harness with Paperclip

This suite covers the instruction reductions tracked in the working checklist. Existing context-integrity and chat evals supplied a custom QA instruction bundle. Their passing results therefore did not qualify a hire with the tiny production default. The new suite omits that fixture bundle and lets the public agent-creation route materialize the shipped default.

The 2026-10-03 legacy ACP Claude repair compares only the original assigned-skill-explicit-invocation cell on matched current-master sources. Both variants carry identical bounded skill-description/staged-path delivery and env-free session persistence. Only the historical versus reduced default manual/shared prompts and their declared structural unit expectations differ. The public bundle/procedure observations derive from exact admitted source bytes (eight-word identity or the independently pinned 4,249-byte historical manual), never a branch/environment label. Candidate absence assertions stay intact. Independent document/task and exact-credential guards are unchanged. All stock tasks enforce single_attempt in the hosted launcher; automatic product disposition recovery still counts as actual usage. This is not a new 24-cell campaign or qualification of the frozen native-completion context.

Coverage contract

Change Required deterministic evidence Product E2E evidence
SH-1: native Codex preserves vendor base instructions Serialized start/resume requests from the TypeScript driver, recovery paths, runnerd transport, Runner Lab/live sessions, and Rust provider use additive developer instructions. Real new native Codex hires execute the three journeys below. Task success alone cannot prove vendor base preservation.
SH-2: identity-only default hire manual Existing public agent-creation and onboarding-asset tests cover default, custom, and CEO exceptions. The public bundle is exactly one AGENTS.md containing the eight-word shipped identity, checked before provider execution and again during cleanup.
SH-3: reduced shared legacy task/chat defaults and resume delta Shared prompt tests and ACPX, Codex, OpenCode, Pi, Hermes, and Cursor Cloud adapter regressions retain runtime context and exclude removed generic procedures. Actual legacy adapter.invoke prompts retain fresh identity/connection guidance, omit the removed procedures, and have complete receipts for every observed run.

pnpm test:e2e:runner:stock-harness runs these credential-free prerequisites, including the Rust test. It writes JSON reports, the source SHA, a working-source fingerprint, selected checks, test counts, and providerCalls: 0 under results/stock-harness-preflight-<UTC>/. Missing reports or required skipped assertions fail. Unrelated native tests filtered by the name selector remain explicitly skipped; they are not counted as executed coverage. The live launcher runs the prerequisites automatically before loading local credentials or starting an isolated provider instance. Its subprocess receives only allowlisted toolchain and operating-system variables. Disposable GitHub runners may resolve Cargo dependencies; local prerequisites retain offline Cargo execution. The ordinary server SDK dependency builder prepares the shared/SDK outputs needed by route tests on a cold install with lifecycle scripts disabled.

The direct Playwright path also requires the retained prerequisite receipt. It verifies the exact checkout SHA, evaluated-source fingerprint, all requested Vitest reports and required assertions, and the Rust result before creating a company. Missing, stale, partial, or failed prerequisites cannot qualify a cell. Oracle/admission calibration is itself included in the prerequisite gate.

Live matrix

There are 24 explicit local cells: these eight existing profiles each run three existing journeys with independently calibrated graders.

  • Legacy: legacy-codex, legacy-claude, legacy-opencode, legacy-acp-codex, legacy-acp-claude.
  • Native: runner-codex, runner-acpx-claude, runner-opencode.
Journey Expected provider turns Observable outcome Cell deadline
assigned-skill-explicit-invocation 1 Public skill creation/pinning, an explicit skill request, and saved output containing the marker available only in the skill body. 12 minutes
ordered-comment-continuation 2 Initial report followed by three ordered public comments, including repeated wording and a changed scope; final saved report preserves the ledger and requested scope. 12 minutes
continuity-restart 3 Task-backed chat retains the requested context across a server restart and subsequent replies. 15 minutes

A full matrix expects 48 provider turns. Models, credentials, effort, permissions, assigned skills, environment, and managed secret references are inherited from the existing profile; only its QA manual is omitted. Both company and agent receive a 1,000-cent monthly hard stop before provider execution, verified through public records. These limits do not predict final spend: attempts, retries, partial runs, unknown billing, and cleanup remain in the existing campaign accounting and qualification rules.

stock-harness is excluded from --all; select it explicitly. It has no Daytona cells. Pending or unrepresented harnesses are not live-qualified by this matrix. Pi/Hermes/Cursor Cloud rendering coverage is deterministic here. The separate Codex-through-ACP base-instruction patch is still an open checklist item: its legacy cells qualify the common prompt reduction, not vendor base preservation.

Run and inspect

# No credentials or paid providers:
pnpm test:e2e:runner:stock-harness --list
pnpm test:e2e:runner:stock-harness
pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
pnpm test:e2e:runner -- --list --suite stock-harness

# With authorization and the selected profile's required provider credential:
pnpm test:e2e:runner -- --id stock-harness.runner-codex.local.assigned-skill-explicit-invocation

Use an exact cell first, then expand profile/journey selections when its evidence is understood. The ordinary launcher, isolated instance, fixture cleanup, screenshots, attempt history, usage/cost reporting, and dashboard publisher are unchanged. snapshots/stock-harness-hire.json captures the public bundle and budgets before paid execution. snapshots/stock-harness.json captures the final bundle, budgets, run IDs, actual legacy prompts, and matcher verdicts. Evidence API failures retain an error snapshot and cannot pass. All snapshots use the existing secret sanitizer; screenshots remain the original captured pixels.

The oracle's identity is independent of the implementation constant, so changing the shipped manual cannot silently change the expected result. The suite definition digest incorporates the evaluated default manual, shared prompt implementation and connection guidance, fixture, journeys, grader, prerequisite, and execution integration sources. Positive and plausible-negative support tests exercise missing/malformed receipts, manual regrowth, removed startup/resume procedures, absent connection guidance, and budget drift. Existing calibrated lifecycle graders still own task/chat success.

The skill oracle proves the requested pinned skill's output marker reached the saved result without appearing in the task request or follow-up comments. Skill tool/read event detection is retained as supporting evidence; it is not a required cross-provider tool-trace assertion.

Qualification status

Setup was validated locally on 2026-10-02 with the deterministic prerequisites, Product E2E support tests, typecheck, and discovery: 478 executed prerequisite tests passed (477 TypeScript plus one Rust), along with 860 support tests and discovery of all 24 cells. The 313 unrelated native tests filtered by the gate are not counted as passing coverage. Local evidence is retained under results/stock-harness-preflight-2026-10-02T16-15-39.064Z/preflight.json. Earlier interrupted, discovery-failure, and setup/test-timeout attempts remain retained; the final unchanged-assertion retry passed. No live cells or paid providers were run during that initial setup. This establishes executable coverage, not a live reliability result or improved coding quality. A quality claim needs comparable tasks, models, effort, tools, and independently graded before/after results.

The suite exercises fresh isolated hires and their continuations. Existing saved manuals and old Codex sessions are not automatically migrated. The latter still need a provider-session reset to restore a previously replaced vendor base. Private Runner protocol definitions remain separate: they use mock control-plane operations and cannot substitute for this public hiring and assembled-prompt coverage.

GitHub qualification and matched comparison

The instruction reductions and coverage are under review in PR #14948. The first diagnostic native Codex skill cell passed on GitHub at a63437069de58d22ee5adbcb6a6202c007dcf037. All seven independent skill/task checks passed; the provider run lasted 34.745 seconds and the cell 56.660 seconds. Its tiny public hire bundle and budget receipts passed, and cleanup passed. This diagnostic predates enforced prerequisite admission and the expanded source digest, so it is retained separately from the final qualification matrix.

The first full candidate attempt at 36e987246b649927e96ce1184cd616c4e490106e was cancelled during prerequisites. Cold protected installs disable lifecycle scripts, leaving the plugin SDK unbuilt; SH-2/SH-3 could not import it. Provider admission was not reached. This setup failure is retained separately from behavioral results. The prerequisite now runs the ordinary server dependency builder first, retains its output and exit status, and refuses admission when setup fails. A cold legacy pilot must pass before the full matrix retry.

The fixed legacy pilot at 4163dbfd0fd4d145bfa52b4d7f80eb59a362ee36 passed SDK setup, hire/shared prompt/oracle checks and the Rust additive test. It stopped before providers: the real daemon-frame test lacked the cold paperclip-runnerd binary. Setup now builds that daemon from the locked Rust source before TypeScript gates, retains runnerd-build.txt, and requires both setup exits in the admission receipt. The required daemon-frame test selects that built debug binary explicitly, without replacing staged product binaries. The receipt records its SHA-256; verification rejects a changed binary before provider admission. The next cold pilot at ac6ddefb589c18fa9c30946db40e58499e9fb0e0 then found the required fake Codex protocol fixture absent. It also stopped before paid providers. Setup uses the package's ordinary locked workspace --bins build, covering both the daemon and its fixture; verification binds both binaries to the retained receipt.

That final cold pilot passed all 537 prerequisites on GitHub and reached Claude Sonnet 4.6. It saved a workspace task-output.md and completed the issue, while the independent durable-document oracle found no Paperclip issue document. The pinned skill's phrase "task document" does not explicitly name storage in Paperclip; matched historical results must precede any regression attribution. The task, instruction delivery, budget, and cleanup receipts remain retained.

Its trusted merged report rejected the prerequisite folder alongside the campaign root and synthesized a missing-result infrastructure error. Raw packaged cell results were uploaded and remain inspectable. Prerequisites now live under the exact campaign root; the unchanged trusted selector contract is calibrated in the mandatory gate. The initial full matched campaigns at f02d8d0df and 12c5433c6 retain their original raw results and publication outcomes. Any reconstructed comparison must declare directory-layout recovery and preserve every result, hash, original failure, and assertion.

Dispatch the trusted workflow from master, with target_branch naming the same-repository candidate and an exact cell selector first. The workflow resolves that target once to an immutable SHA. Never dispatch target-controlled workflow definitions with protected credentials. Protected environments, scoped provider keys, frozen target dependencies, report sanitization, publication, and existing bounded retry/cleanup policies retain their existing owners.

A temporary codex/stock-harness-previous-instructions branch compares the same 24 cells with the previous default manual and shared startup/resume prompts. It holds merged native Codex fix #14920 constant. Only those two production instruction sources differ. Its explicit historical structural oracle expects the old manual and records old generic procedures; the candidate's reduction assertions remain mandatory. Historical tests verify that prior contract. The independent skill/context/chat journeys, behavioral graders, fixtures, models, effort, tools, permissions, and credentials are identical. The retained comparison manifest records their hashes and the restored instruction revision.

One full campaign per variant initially expects 48 provider turns, plus any existing bounded automatic retries; the earlier one-cell diagnostic remains separate. Compare behavioral results by profile and journey, with missing evidence unqualified. Keep structural instruction differences separate from task success. Report every attempt, failure attribution, provider timing, token usage, reported costs and unknown spend. A reported zero subtotal is not proof of zero provider spending. This small single-trial matrix cannot establish general coding quality or broad performance equivalence, and does not compare #14920 before/after.

Measured comparison and unresolved delivery

The dated live report and its safe JSON projection contain the per-profile/case comparison, exact source hashes, all campaign and recovery links, timing/usage, cost coverage, security failures, clipping limits and publication-layout recovery. Candidate f02d8d0df has 24 retained results: 15 pass and nine fail. Historical 12c5433c6 also has 24 retained results: 15 pass and nine fail. Its single AWS-runner recovery timed out; original missing evidence remains recorded. Classic Claude/OpenCode skill runs save no Paperclip task document where historical runs save one; the original oracle remains failed.

The merged native Codex change is held constant. Same aggregate success counts would not establish equivalence: case outcomes differ, fixture storage wording is ambiguous, ACP credential guards fail, and some public prompt receipts are clipped. Native finish/block tool guidance belongs to native runners; legacy document-delivery guidance must use the actual Paperclip skill/API path. No production or skill instructions have been changed to turn the measured failures into passes. PR #14948 remains draft.

The corrected packaging pilot at 1eb5ba420 passes on GitHub with valid evidence and cleanup. It validates prerequisite nesting under the exact campaign root using the unchanged trusted selector. Its current prerequisite runs 557 checks (556 TypeScript plus one Rust); all 895 E2E support tests pass. The 313 filtered native tests are not counted. Original failed/partial attempts and publication failures remain retained; local reconstructed copies move folders without editing results, graders or usage. Never dispatch a second development campaign for an active target branch: workflow concurrency supersedes the older run. Use a separate frozen-source branch when independent campaigns must overlap.

Focused legacy delivery repair

The original 24 cells and original assigned-skill request/procedure remain unchanged. Two added explicit Paperclip-document cells apply only to classic Claude and OpenCode, making 26 catalog cells (50 expected turns if every cell is selected). The repair campaign selects only the original skill case and new document case for those two profiles: four cells per variant, eight expected turns total. No full matrix rerun is planned.

Both skill sources are recorded in definition/admission digests. The pre-fix baseline restores the old SKILL.md and records the new reference as absent, without copying the new recipe into that baseline. It holds the tiny manual/shared prompts and all fixture/model/auth/effort inputs fixed. The explicit document oracle checks actual public content/revision and an exact same-app document link, with plausible-negative calibrations.