fix(agents): reduce default instructions and qualify stock harnesses (#14948)

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its adapters supply task context and access to Paperclip skills and
tools.
> - The default hire manual and shared prompts also repeat general work
procedures.
> - Those procedures overlap with stock provider instructions and the
Paperclip skill.
> - Existing E2E fixtures supply a QA manual, so they do not qualify the
production default.
> - This pull request reduces the generic instructions and adds real
default-hire coverage.
> - The benefit is less competing guidance, with inspectable evidence
for preserved skills and task context.

## Linked Issues or Issue Description

Refs: #14920. That merged change preserves native Codex base
instructions. This PR covers the default manual, shared legacy prompts,
operational skill guidance, and the narrowly approved ACP
skill-discovery/session-environment repair for measured delivery and
credential-persistence failures.

**What existing behavior does this improve?**

New non-CEO hires without a custom bundle and legacy task/chat startup
and continuation prompts.

**Current behavior**

The shipped default manual contains 602 words. Generic task/chat prompts
and ordinary resume deltas repeat work procedures already available
through the harness and Paperclip skill.

**Proposed behavior**

The default manual contains only the eight-word company identity. Shared
startup prompts retain identity and connection guidance. Ordinary resume
deltas retain current work context without the generic execution
contract.

**Reason and benefit**

Let the stock harness guide general work. Keep Paperclip-specific
capabilities and independently test default hires, skills, ordered
comments, and chat restart.

**Breaking changes**

New default hires receive less guidance. Existing saved manuals,
explicit custom bundles, CEO templates, and specialized wake contracts
retain their behavior. The obsolete includeExecutionContract option
remains accepted for source compatibility.

## What Changed

- Reduce the default hire manual to one sentence.
- Reduce shared task/chat defaults and remove the generic
ordinary-resume contract.
- Keep connection guidance, auth, skills, custom prompts, and
specialized wake context.
- Add credential-free instruction-boundary gates and 26 explicit Product
E2E cells across eight legacy/native profiles, including two focused
Paperclip-storage cases.
- Capture public hire receipts before providers run, then grade
delivered prompts and independent task/chat outcomes.
- Add an early legacy skill API recipe for saving a task document,
checking the saved revision receipt and linking the document. Improve
stock task/heartbeat skill-selection metadata and show a clickable
Markdown UI-link example. Keep native tool completion separate.
- Advertise bounded routing descriptions and exact successfully staged
SKILL.md paths in legacy ACP Claude; keep full bodies on demand and
preserve remote path rebasing.
- Remove only the provider environment from copied persisted ACP session
records, while loading current run credentials and preserving all other
options/conversation state.
- Regenerate both capability metadata inventories and reject stale
manifests/inventories before provider admission.
- Publish the original reduction and focused skill-repair comparisons,
preserving all failures, automatic recovery, cost coverage and
limitations.

## Verification

**Behavioral qualification remains pending.** Original legacy ACP Claude
loses the issue document only in the reduced cohort beneath an unchanged
credential failure. A source-backed diagnosis finds that neither
ordinary assignment reads the staged operational skill, while the
runtime persists provider environment in session state. The new common
repairs expose skill metadata/path and omit persisted env; strict
document and credential guards stay intact. [Inspectable diagnosis and
retained
hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md).

Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb`
incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged
#14961/#15007). All 222 affected adapter tests, adapter-utils/E2E
typechecks, and final 96 variant/grader/retry calibrations pass. The
frozen historical comparator is
`c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths
identical, exactly two production instruction paths plus nine declared
unit expectations differ. The operational skill/discovery/environment
repairs, selected model/profile/task/core grader/auth/permissions/retry
policy are identical. Both actual launcher prepare→verify admissions
pass with zero providers. [Immutable manifest and exact
receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json).

One original legacy ACP Claude cell per variant is authorized, with
enforced single campaign attempts, 12-minute deadlines and company/agent
1,000-cent hard stops; every product recovery run/cost is counted.
Actual live outcomes are pending. Current normal CI has one failed
server shard and failed aggregate verify under diagnosis; other normal
gates including typecheck/build/Rust/all eight browser shards pass.
Fresh review completed successfully; the valid historical startup/resume
masking finding was fixed with per-invocation task/chat checks and
strict complete-snapshot capture, calibrated and resolved. Prior heads,
failures and campaigns below remain historical evidence, not checks on
this repair head.

- Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on
merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52
current-head checks pass with two intentional Storybook skips, including
repository typecheck/test/build and the browser shard. Fresh Greptile is
5/5 with zero unresolved review threads. Exact-head stock prerequisites
pass 599 assertions (598 TypeScript + 1 Rust), all six gates and
retained receipt verification, zero providers/source errors. Fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`.
Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck
and 26-cell stock discovery pass. Canonical contract/inventory checks
and the later issue-derived reference calibration are retained; that
reference-only follow-up is not live-qualified by earlier frozen runs.
- Prior full repository typecheck/build passed. The complete local
Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three
timing failures. All three affected files passed unchanged narrow
reruns; original failures remain retained. Current-head CI now passes
the full general checks; the original local failures remain retained.
- The original 24-pair default-manual/shared-prompt comparison has two
new overall classic Claude/OpenCode document-delivery failures plus an
additional legacy ACP Claude document loss beneath an unchanged
credential-guard failure (not closed by later runs), two newly passing
OpenCode ordered cases, seven unchanged failures and 13 unchanged
passes. Equal 15/24 totals do not establish behavioral equivalence.
[Complete original
report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md).
- The skill-only repair holds the eight-word manual/shared prompts and
merged #14920 fixed. All four matched profile configurations and 203
fixture/behavior files match. Candidate
`abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill
sources against baseline `bc83fe030234439ac51279502a28803958963e2e`.
[Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547)
and [baseline
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047)
each pass 571 exact-source prerequisites before providers; all eight
cells clean up successfully. Failed campaigns publish successfully and
remain failed.
- Repair pairs: Claude original Fail → Pass; Claude explicit Pass →
Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff
worsens beneath the unchanged failing UI-link grade: baseline gives a
clickable API URL, candidate gives a code-formatted path without an
anchor. The request's usable-link wording is narrower in the UI-only
oracle. [Complete repair report and safe
projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md).
- The subsequent narrow stock metadata/link correction has two matched
Pass → Pass cases, zero new machine failures/passes and no pending
pairs. Both original-case handoff links remain deficient: candidate uses
a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the
preserved original oracle only requires a durable document. Both
explicit clickable UI-link cases pass revision/content/link grading. All
four exact-source 587-check gates, single assignment runs and cleanup
pass. This does not establish fix causality because baseline also
succeeds. [Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401)
freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched
baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374)
freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only
comparison with reduced manuals/shared prompts held constant, not a
repeat of the historical-manual comparison. Only original and clarified
explicit classic OpenCode cases are selected, two per variant/four
expected turns. 8,242 other tracked files and both profile hashes match;
protected workflows admit each exact source before credentials.
[Complete qualification
report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md).
Candidate original loads Paperclip/reference before saving publicly;
baseline original loads it after writing locally, then saves publicly
within the same assignment. Reported cost totals are $0.0107824490
candidate / $0.0107909015 baseline, with unmetered runtime. The later
reference-only issue-derived link correction is provider-free calibrated
and **not live-qualified** by these frozen runs; no further paid runs.
- Retained tool calls show the repaired original OpenCode assignment
loads only its assigned output skill before writing locally. Operational
Paperclip is first loaded during automatic disposition recovery; its
early recipe is visible then, but it never saves the missing document.
Explicit candidate loads Paperclip and reads the new reference before
saving successfully. All nine actual runs are counted. Reported LLM
totals are $0.3802537209 baseline and $0.4918990161 candidate; local
runtime is unmetered.
- Initial setup, packaging, cancelled/missing-cell recovery, callback
test and relative-output attempts remain retained. No completed provider
failure was rerun. Frozen measurement branches are unchanged by later
canonical metadata maintenance.
- Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`,
and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly
for paid execution; it is excluded from `--all`.

Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f`
replays this PR on merged hiring #14985
(`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four
explicit custom-CEO-bundle checks, minimal generic manual boundary, and
both suites. The combined fixture catalog and hiring calibrations pass
67 assertions; exact-head stock prerequisites pass 599 assertions (598
TypeScript + 1 Rust), all six gates and retained-receipt verification,
zero providers/source errors, fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E
typecheck and 26-cell stock discovery pass. Fresh current-head CI passes
all 52 checks with two intentional skips, and fresh Greptile is 5/5 with
zero unresolved review threads.

The prior source-plan browser failure is retained: a deterministic
process fixture replayed its last `fixture:plan` command on
`chat_task_completed`, writing revision 2 with identical body after the
approval handoff. This was not paid provider execution. Rebased
current-head CI passes the same assertion without an old-head retry or a
change to that browser fixture.

The merged hiring change was measured separately on immutable matched
unions, with this reduced/shared/operational context and native
completion guidance held constant. [Complete original two-profile
report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md):
[candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208)
/ [historical
baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463),
705 provider-free prerequisites each. Both pairs are unchanged Fail →
Fail on the exact-five count, with six core delivery checks passing all
four cells; 28 actual successful runs include eight automatic completion
wakes, zero retries, four successful cleanups. Source-read coverage is
uncomparable, actual model charges unknown. Separately versioned
provider-free accounting remains analytical work; original verdicts are
preserved. This does not rerun or qualify the completed default-manual
or native campaigns.

## Risks

- Legacy ACP Claude's additional delivery loss is not closed by any
later matched run and blocks the no-extra-failing-behavior merge
criterion. Legacy document delivery may have relied on the prior
manual/shared prompts. The early skill repair improves Claude in one
trial; the later OpenCode pairs pass in both variants and cannot
establish causality or robust recovery. Both original-case links remain
deficient beneath the storage-only grade. The later issue-derived
reference correction has only provider-free validation. Native
finish/block descriptions must not be supplied to legacy agents.
- The comparison holds merged native Codex fix #14920 constant; it
cannot measure that fix's before/after task performance.
- These bounded skill/context/chat workflows do not measure general
coding quality. Unrepresented providers remain unqualified.
- Saved manuals and old Codex sessions are not automatically migrated.
Codex through ACP still has a separate base-instruction follow-up.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution, and delegated PR/eval
work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites and all
three unchanged narrow reruns pass; complete-run timing failures
retained in Verification)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green on the new repair head
(prior-head checks retained above)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
on the new repair head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
DottaandPaperclip authored and GitHub committed 2026-10-03 12:32:42 -05:00
1 parent 569c7203aa
commit b17019e14d
68 files changed
+6635 -1414

No files matched your search

+8
View File
@@ -95,6 +95,14 @@ execution ID because `--all` excludes explicit-only suites. Each cell applies a
1,000-cent company and agent budget hard stop before task creation and records
both limits in its evidence.
The explicit-only [stock-harness suite](../tests/runner-e2e/STOCK-HARNESS.md)
reuses skill, ordered-continuation, and chat-restart journeys across eight local
legacy/native profiles with production-default hires. It closes the custom QA
manual coverage gap. Its required credential-free prerequisite maps vendor
instruction layering, the tiny hire bundle, and shared startup/resume reductions
to executable checks. The 24 live cells are configured; no live qualification is
claimed from their setup or unit calibration.
The explicit-only `agent-chat-stories` suite covers the experimental settings
lifecycle for a configured native agent and follow-ups during active work. Its
fixture-driven file wait and persisted-plan oracle are documented in the
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,51 @@
# Legacy document skill repair qualification — 2026-10-02
**TL;DR: one newly passing case (Claude original document delivery), one unchanged pass (Claude explicit storage), two unchanged OpenCode failures, and no pending pairs or new overall machine failures. OpenCode's explicit-case clickable handoff nevertheless got worse:** baseline supplied a clickable API document URL, while candidate supplied only a code-formatted path. Both fail the original UI-only link oracle, whose intended route was narrower than the request's "usable link" wording. This result does not establish general non-regression.
This focused repair keeps the eight-word manual, reduced shared prompts and merged Codex base fix #14920 fixed. It varies only the early operational Paperclip skill recipe and the presence of the new `issue-documents.md` reference. The earlier combined prompt-removal comparison retains its two new classic Claude/OpenCode storage failures; this trial repairs Claude's original case but leaves OpenCode's original regression unresolved.
| Classic profile | Case | Pre-fix skill → repaired skill | Retained evidence |
| --- | --- | --- | --- |
| Claude | Original assigned skill | Fail → Pass | Baseline lacks a durable task document; candidate saves one containing the skill marker. |
| Claude | Explicit Paperclip document | Pass → Pass | Both save the document/revision and provide a canonical UI link. |
| OpenCode | Original assigned skill | Fail → Fail | Both write a workspace file rather than a public task document. Candidate requires automatic disposition recovery. |
| OpenCode | Explicit Paperclip document | Fail → Fail | Both save the document/revision. Baseline provides a clickable API URL; candidate gives a code-formatted UI path without an anchor. Original UI-link checks fail in both. |
Original machine verdicts and evidence hashes are retained in the [safe evidence projection](2026-10-02-legacy-document-skill-repair.json). No model was rerun to improve these results.
## Frozen sources and public reports
- [Repaired skill campaign](https://github.com/paperclipai/paperclip/actions/runs/37060885547), source `abd0b628ca642c09a54a4edc56a5227402f6686e`, branch `codex/legacy-document-repaired-skill`. [Published candidate report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37060885547-1/index.html).
- [Pre-fix skill campaign](https://github.com/paperclipai/paperclip/actions/runs/37060888047), source `bc83fe030234439ac51279502a28803958963e2e`, branch `codex/legacy-document-pre-fix-skill`. [Published baseline report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37060888047-1/index.html).
- Both trusted default-branch workflows ran concurrently on separate concurrency targets. All eight cells pass their exact-source prerequisite gate before providers; all eight clean up successfully. Failed campaigns are published successfully and remain failed.
- The four model/effort/auth/tool/permission/environment configurations match. Models are `claude-sonnet-4-6` and `openrouter/deepseek/deepseek-v4-flash-0731`. The comparison records 203 byte-identical fixture/behavior sources and explicit present/absent skill-source fingerprints. The new reference is absent from the baseline, rather than silently introduced into it.
- Local admission preparation passed 909 E2E support tests, E2E typecheck and four-cell discovery; the candidate exact-source gate passed 571 checks, including one Rust check. Initial sandbox and obsolete catalog-count attempts remain retained. The original request and behavioral oracle were not weakened.
## What OpenCode actually received
The repaired original assignment invokes only its assigned Context integrity output skill, then writes `task-document.md`. There is no operational Paperclip skill load or issue-document reference read before that local write. A separate automatic `issue_disposition_repair` run then loads Paperclip. Its returned content contains the new early saved-document paragraph and reference pointer even though the overall skill result is marked truncated; it never reads the reference or saves a document, and only finalizes the issue's status.
This is evidence of operational-skill selection after work, not evidence that the assignment saw and ignored the new recipe. The pre-fix original run also writes locally before loading the old Paperclip skill. Whether the adapter should activate operational Paperclip guidance before work is a production-delivery question for review; expanding a skill the first assignment never opened would not address the observed sequence.
The explicit candidate loads the assigned skill, then Paperclip, receives the new early paragraph, reads `issue-documents.md`, and calls the document API successfully. Its saved document has a persisted revision. The remaining defect is its comment's non-clickable path. Baseline supplied an actual Markdown link to the issue-document API GET endpoint. Read-only source review confirms that this endpoint returns the document in board context, and the unprefixed UI issue route redirects through the selected company's prefix while preserving the document hash. Actual authenticated browser navigation against these cleaned-up measured instances was not replayed, so neither URL is classified as a proven broken route. Candidate's missing clickable anchor is directly observable.
A future explicit fixture should name a clickable Paperclip UI document link if that is the intended contract. This clarification would not retroactively pass or fail either original result. A short canonical Markdown UI-link example can be considered separately from operational-skill delivery. No further production instruction change or paid rerun was made for this diagnosis.
## Runs, timing and cost
Eight provider turns were expected; nine actual runs are retained. Candidate OpenCode's original assignment succeeds but leaves the task `in_progress`, triggering automatic disposition recovery. The recovery succeeds and marks it done, without creating the missing document. Both runs, their time and cost are counted rather than selecting a better attempt.
| Profile | Case | Baseline provider seconds | Candidate provider seconds | Baseline reported USD | Candidate reported USD |
| --- | --- | ---: | ---: | ---: | ---: |
| legacy-claude | Original | 23.032 | 27.702 | 0.1668394500 | 0.2267947500 |
| legacy-claude | Explicit storage | 34.999 | 40.513 | 0.1997571000 | 0.2561425500 |
| legacy-opencode | Original | 36.770 | 68.032 | 0.0033927116 | 0.0047148437 |
| legacy-opencode | Explicit storage | 206.191 | 46.300 | 0.0102644593 | 0.0042468724 |
Reported LLM totals are $0.3802537209 baseline and $0.4918990161 candidate. All four baseline and five candidate ledger rows contain token usage and reported cost. Local runtime is unmetered and actual external billing is not established. Single trials, the additional recovery and baseline OpenCode's long explicit-case execution prevent a general timing or spending conclusion.
## PR maintenance and remaining qualifications
The skill insertion shifted generated capability heading anchors. General PR CI detected stale `capabilities.yaml`; the paid workflow's narrower build and every exact-source admission still passed before provider access. Both canonical capability metadata inventories are regenerated after measurement, include the new reference where declared, and adds a credential-free admission test that rejects stale/missing manifests. These maintenance changes do not modify the frozen measured branches or the skill bytes used in this comparison.
Draft [PR #14948](https://github.com/paperclipai/paperclip/pull/14948) keeps this repair and the tiny manual; draft [PR #14961](https://github.com/paperclipai/paperclip/pull/14961) measures native completion documentation separately. OpenCode original delivery remains unresolved, and existing Claude chat-memory and legacy ACP credential/receipt findings remain separate. No broad matrix rerun was launched, and merging remains user-controlled.
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,75 @@
# OpenCode skill routing and clickable delivery — 2026-10-02
**TL;DR: 0 newly failing cells, 0 newly passing cells, 2 unchanged passes, and 0 pending pairs in this matched trial. Both document cases pass in both variants, so this does not establish that the metadata/link change caused the earlier failures to recover. Handoff remains imperfect in the original case:** candidate uses a clickable link with the wrong `PAP` company prefix; baseline supplies only a bare prefix-less slug path. Neither defect is scored by the preserved original document-storage oracle. The explicit clickable-link cases deliver the correct `RUN` link in both variants.
The narrow correction expands the operational skill's stock discovery description to include Paperclip task/heartbeat work and document/file delivery. The early API recipe asks for a clickable Markdown link, and the reference shows a canonical UI-link example. The full skill stays on disk; the eight-word manual and shared prompts remain reduced. No full-body injection, adapter policy override or native tool guidance is added to legacy runs.
| Classic OpenCode case | Pre-correction skill | Corrected skill | Independent evidence and limitation |
| --- | --- | --- | --- |
| Original assigned skill, verbatim request | Pass | Pass | Exactly one saved document with skill-only marker and persisted revision in each. Both handoff paths are deficient and outside this case's link-free oracle. |
| Explicit Paperclip storage with clarified clickable UI link | Pass | Pass | Saved revision/content and exact clickable `/RUN/issues/RUN-1#document-context-integrity-output` in both. |
Original machine results, source/configuration proof, tool-read sequence, costs and evidence hashes are preserved in the [safe evidence projection](2026-10-02-opencode-skill-routing-link-qualification.json). All four cells clean up successfully; none uses automatic disposition recovery or an extra attempt. No provider was rerun to improve these results.
## Sources, admission and inspectable reports
- Candidate `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`, frozen branch `codex/opencode-stock-routing-qualified`: [workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401), [published dashboard and screenshots](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37069547401-1/index.html).
- Baseline `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`, frozen branch `codex/opencode-pre-routing-skill`: [workflow](https://github.com/paperclipai/paperclip/actions/runs/37069552374), [published dashboard and screenshots](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37069552374-1/index.html).
- Trusted default-branch workflows ran concurrently on separate targets. Each of four measured cells passed its own exact-source, credential-free 587-assertion prerequisite before provider credentials. Candidate source fingerprint is `aeb792549146cc29eda779e83ae22d336276e31162f738a1dde85a82e7833902`; baseline is `0315602ac6788cf0d65561a73cc85d790768ef90c6815257077a2381f961b96e`.
- Git tree comparison verifies 8,242 identical tracked files. Only the two operational skill sources and an explicit baseline provenance receipt differ. Models, effort, tools, auth, permissions, budgets, manual/shared prompts and behavioral fixtures/graders are fixed; both fixture configuration hashes match. Suite/source digest differences correctly include the changed skill bytes.
- Local source preparation passed 918 E2E support tests, eight canonical inventory calibrations, skill validation, E2E typecheck, two-cell discovery and both exact-source prerequisites. Both capability inventories use the same declared source list. Initial setup mistakes remain retained and did not reach providers.
The explicit fixture now asks for a clickable company-prefixed Paperclip UI document link, matching its intended oracle. Bare paths and code-formatted paths fail calibration; correct relative/same-app absolute links pass. The original assigned-skill request and procedure remain verbatim. Earlier API-link grades are not retroactively changed, and these explicit results are not a direct repeat of the older “usable link” cohort.
## Which guidance was read before delivery
In the candidate original assignment, OpenCode loads the assigned Context integrity output skill, then operational Paperclip before any document write. The returned Paperclip skill is marked truncated, but the early saved-document recipe and reference pointer are visible. It reads `issue-documents.md`, then writes the public task document through the API. This differs from the previously failing candidate, which first loaded Paperclip only in a separate disposition-recovery run after its local-only output.
The current baseline original assignment first loads the assigned output skill and writes a workspace Markdown file. It later loads Paperclip and reads the issue-document reference, then saves a public document before completing the **same** assignment. Thus its successful delivery did not require improved metadata; stock selection timing varies between trials. Both explicit variants load Paperclip and read the reference before the document API call. Four rendered final-state screenshots were inspected, and public document revisions/content and completion comments agree with the retained results.
The candidate original comment has an actual Markdown href `/PAP/issues/RUN-1#document-context-integrity-output`, despite the measured company prefix being `RUN`. Baseline's original comment has a bare `/issues/runner-e2e-…#document-task-document` path without a Markdown anchor. The candidate's literal prefix matches the reference's example, making example copying a plausible cause, but this trial does not isolate that cause. Authenticated click navigation was not replayed after cleanup; neither original link is presented as a verified usable handoff. A minimal follow-up can show an issue-derived prefix instead of a literal company example. These original machine passes must not be presented as fully correct links.
## Timing, usage and cost
Four provider turns were expected and four assignment runs occurred. All four ledger receipts contain token usage and reported LLM cost. Local runtime remains unmetered; actual external billing is not established.
| Case | Baseline provider seconds | Candidate provider seconds | Baseline cell seconds | Candidate cell seconds | Baseline reported USD | Candidate reported USD |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| Original | 25.870 | 67.897 | 46.285 | 89.455 | 0.0058472545 | 0.0041383700 |
| Explicit clickable storage | 73.984 | 82.751 | 94.804 | 102.901 | 0.0049436470 | 0.0066440790 |
Reported LLM totals are **$0.0107909015 baseline** and **$0.0107824490 candidate**. The candidate original case takes longer in this trial; these single observations, different cache/token receipts and unmetered runtime do not establish general speed, cost or quality equivalence.
## Preserved failures and scope
The [original combined prompt-removal report](2026-10-02-stock-harness-live-comparison.md) still records two new classic Claude/OpenCode document-delivery failures, two OpenCode ordered-case improvements, and separate unchanged credential/chat failures. The [first recipe repair report](2026-10-02-legacy-document-skill-repair.md) still records Claude's improvement, OpenCode's unresolved local-only output, and the candidate's worse non-clickable explicit handoff. New successes do not erase those earlier observations or establish robustness.
This comparison qualifies only classic OpenCode's two selected journeys. Existing classic Codex/Claude/OpenCode and ACP Codex skill mounts were audited; ACP Claude's names/root-only metadata and unsupported custom ACP delivery remain separate. Hermes/Pi/OpenClaw are not live-qualified here. Native completion descriptions and hiring templates have separate changes and measurements. Both PRs remain draft; no merge or further paid breadth is part of this report.
## Provider-free follow-up after measurement
The approved reference-only follow-up replaces the literal `PAP` link with a
JavaScript construction using the current `issue.identifier` and the successful
write receipt's `saved.key`. It is a later source change, **not live-qualified**
by the frozen candidate above. The independent storage grader now calibrates a
wrong-company link as failure, and executes the shipped recipe against two
other company prefixes plus a locked-write redirected document key. All 25
focused document/source-digest tests, eight inventory calibrations, canonical
generated metadata checks and E2E typecheck pass. An initial sandboxed support attempt denied localhost/tsx
pipe creation; the failed attempt is retained, with its affected files rerun
unchanged under permitted local execution: 906 assertions passed in the original
921-assertion attempt, and all 82 selected assertions passed in the permitted
retry, including every originally failed assertion. No new model calls or old verdict
changes are part of this follow-up.
Fresh native-PR review caught a supported edge case in that later recipe:
`issue.identifier` may be null. The recipe now uses the current issue ID in the
unprefixed UI route when no identifier exists. Provider-free tests cover both
null and absent identifiers, preserving the returned document key. Source review
confirms that the board resolves the loaded issue's actual company and preserves
the document hash; no agent-side company fetch is needed. This route logic also
corrects wrong prefixes, so the frozen candidate's `PAP` href is noncanonical,
not a proven broken link. No authenticated click replay or further model run was
made. The previous 4/5 review and local setup/timing failures remain retained;
current focused document/source/manifest/admission checks pass 56 assertions and
E2E typecheck passes. This correction still has no live qualification.
@@ -0,0 +1,357 @@
{
"schema": "paperclip.stock-harness.safe-comparison/v1",
"candidateSha": "f02d8d0df327abb43b20c7e7beb86798239abbf5",
"baselineSha": "12c5433c67dc2e62916b879349c7ba2b6e0431f0",
"nativeCodexFixHeldConstant": "408f70e69f9c5e49cb4377f4886ac2001bfa67a2",
"comparisonManifest": {
"schema": "paperclip.stock-harness-comparison.v1",
"variant": "previous-instructions",
"candidateSha": "f02d8d0df327abb43b20c7e7beb86798239abbf5",
"restoredInstructionsFrom": "e00d10d5d5594f6e2d1e8cf4e0ec81c074b0bd88",
"nativeCodex14920": "held-constant; no before/after performance claim",
"productionDifferences": [
"server/src/onboarding-assets/default/AGENTS.md",
"packages/adapter-utils/src/server-utils.ts"
],
"structuralOracle": "Historical manual exact content and observed generic procedures; independent behavioral graders unchanged",
"expectedMatrixCells": 24,
"expectedProviderTurns": 48,
"maxInitialCampaignsPerVariant": 1,
"behavioralSources": {
"context-integrity-cases.ts": "201d05f8257c1121a94dd9d8a22f08b9db80555e388013ffeabb65e301499fd2",
"context-integrity-scoring.ts": "50d370044e45386ba87fb23ca63c228b96bb8373d67bb7e82e8b2a89eb23f252",
"context-integrity-flow.ts": "63bcb7bdba4c80b15ff5872f16383844456eab0c5a10627da5d3487a26821c6a",
"chat-cases.ts": "754b251bdd42605e147da9081b663329234260d2b13df0df3d44562e362b14ca",
"chat-flow.ts": "11fc506559da61315ce04a877f6b38a055e71e532fecfcf62c749a1d3bf985e7",
"live-fixtures.ts": "1ef80209cc53cc514a8e58af4c2225e38ec15b20acf17ca7564506ca6617ac3e",
"runner.spec.ts": "14d57e0f9131b3f1981ac593464738d01362008eefe85ea836075c794681fa85",
"launch.ts": "5ada6ca5484542f028b29fc603dc7992213d4c06fa9046834bd7d23c4511b847",
"stock-harness-admission.ts": "c9775d77d48f8d59ba1197a7e5769f14db0c7d33f5e2a2968a04466e083d63ff",
"types.ts": "c45dc3b2366457449ec9e7d413f930605dec88ea93a14ce5c5a5f445a27d148c"
},
"instructionSha256": {
"manual": "e4d2375d722602cd744403292d99f6e6c9b7e8014d9cdfa6f9811ace03428e5f",
"shared": "819026782b92f37064fca354ddbca685b3c05f1869ec560e40d8a757b0a8aee9"
}
},
"modelsEffortToolsPermissionsCredentials": "same fixture/source configuration; effort inherits provider default when unspecified",
"expectedCellsPerVariant": 24,
"expectedProviderTurnsPerVariant": 48,
"baseline": [
{"executionId":"stock-harness.legacy-acp-claude.local.assigned-skill-explicit-invocation","profile":"legacy-acp-claude","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":59075,"providerDurationMs":42037,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":15184,"outputTokens":2261,"cachedInputTokens":218341,"totalTokens":235786,"reportedCostUsd":0.16114005000000003,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":42037,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.16114005000000003,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.16114005000000003,"complete":true},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":14137,"reportedPromptChars":13607,"freshPromptChars":4688,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"78f6add6f11a45f58bfc6fee81600871f888750986c27497563a79eba39a101c","evidence-manifest.json":"14f67d036a3663f606cf269890f5a36a18221e5c7dcb466f513f761d882fb6eb","snapshots/context-integrity.json":"93bba435c845af31707d4e4b84b2330cd8ff14045a76c2d1267b958c8133950b","snapshots/stock-harness-hire.json":"e8b85aa909fc6af6a72b180c79572d70a26dd6ee6307c9b985f80b928eb96170","snapshots/stock-harness.json":"fa0eb10eb3939b91d83fed8a8c9f9915bb407ab03a5a9bd7e5df99aa593ffb99","snapshots/stock-harness-preflight.json":"2ac069369e2d26b0e4c213966770962c0282f9a7665fa1297fe3c70df2c73960"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.legacy-acp-claude.local.continuity-restart","profile":"legacy-acp-claude","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["stockHarness.fresh-default-delivered"],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":36961,"providerDurationMs":19492,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":19069,"outputTokens":329,"cachedInputTokens":55518,"totalTokens":74916,"reportedCostUsd":0.1039979,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":19492,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.1039979,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.1039979,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":false},"invocationMetrics":[{"capturedPromptChars":16395,"reportedPromptChars":17606,"freshPromptChars":1462,"retrievalClipped":true},{"capturedPromptChars":16395,"reportedPromptChars":17638,"freshPromptChars":1462,"retrievalClipped":true}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"b00296e015e0d35a213e031aede08732f900cf8aaa0f67dd6adc30310871a9f4","evidence-manifest.json":"101cb2a051bb64d3fc12e3a5389c2f39be4acb85e120b9bca1557163643f32e8","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"df3b7ce86f99b1f31ce3061e637e936d40a68b63d49aa307cf4a9dcfa75b3983","snapshots/stock-harness.json":"cffaf9caec42246841497abf2059ee2d11c23a67971ad14112fa96beb23d302a","snapshots/stock-harness-preflight.json":"7da7c7be4d0c33c6d1c3fadeb500eebe3fceba797edd0f24a9db187091896b6b"}},
{"executionId":"stock-harness.legacy-acp-claude.local.ordered-comment-continuation","profile":"legacy-acp-claude","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["stockHarness.fresh-default-delivered"],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":169818,"providerDurationMs":150645,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":49679,"outputTokens":7557,"cachedInputTokens":465247,"totalTokens":522483,"reportedCostUsd":0.45116585000000003,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":150645,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.45116585000000003,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.45116585000000003,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":false},"invocationMetrics":[{"capturedPromptChars":16433,"reportedPromptChars":21108,"freshPromptChars":4688,"retrievalClipped":true},{"capturedPromptChars":14499,"reportedPromptChars":14021,"freshPromptChars":4688,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"9994318ef5b3cac9143215d580f7551600777864f3d54a0cb2c48941a0799523","evidence-manifest.json":"79ba61a1c2ac3adde8d1905f09a051d0b711be7c618285b728d13f3650fe9d64","snapshots/context-integrity.json":"0566e89e716c742ee4452f5d9bd719354dba968459809e3e362794f55135b705","snapshots/stock-harness-hire.json":"05919c3be635604a40c0ea9bb76ab1ce1a1ce6545417e66e02c5a26881f082ee","snapshots/stock-harness.json":"6fdbee98a58b20e131a46a5bf9220c1d9e6c4cfc02d64baeee700627869b137c","snapshots/stock-harness-preflight.json":"cb4abccc23ebbe0cd4884ca1d5238eef14b82e15987a2d9e7083ccb608504189"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.legacy-acp-codex.local.assigned-skill-explicit-invocation","profile":"legacy-acp-codex","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":59328,"providerDurationMs":35548,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":511,"outputTokens":30,"cachedInputTokens":27646,"totalTokens":28187,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":35548,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":13618,"reportedPromptChars":13642,"freshPromptChars":4687,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"743b4b283ee73f34db4434facefb26012145fa528a1912856725c4f13e5b9247","evidence-manifest.json":"fe7f46cac1d8722c6137ec813638ecd0c5944f1ea25cc24092ff8d22236c807b","snapshots/context-integrity.json":"d247901367c17626a5738743dc300566c758db50fa7c0b60e5f74be5b7083245","snapshots/stock-harness-hire.json":"15866eaa21a98137bd2c4088dd53256360cdf65b27c8475d49d98cd49788fcf6","snapshots/stock-harness.json":"639cd4b7fec29383a18016e5b8be2227f68d8a5c37035318089ace252408a231","snapshots/stock-harness-preflight.json":"f52edc51ee627e9f5f02d3049a83a9b99fc073fc504ca9f238cf95bc691a1ce8"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.legacy-acp-codex.local.continuity-restart","profile":"legacy-acp-codex","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["stockHarness.invocation-evidence-complete","stockHarness.fresh-default-delivered"],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":44754,"providerDurationMs":22118,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":662,"outputTokens":29,"cachedInputTokens":29955,"totalTokens":30646,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":22118,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":false,"historical-generic-procedures-observed":true,"fresh-default-delivered":false},"invocationMetrics":[{"capturedPromptChars":16383,"reportedPromptChars":17671,"freshPromptChars":1461,"retrievalClipped":true}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"3ba472d29268b3e2d2170cd9947d7c89fbe685d902fedacf5c74a1952a97ce3a","evidence-manifest.json":"a5c263cc9afe3836a3c9f46cda537bb07a374c30fd7ad6f136088680b3ade01c","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"8dd9d388cc265387d960fd435c19874601a5f7ca209b328943f8bacb145c5d89","snapshots/stock-harness.json":"f2b69dda7e26a76c952f0b36f10cfd7e9fb364ae3930585961d84093d17c2979","snapshots/stock-harness-preflight.json":"5fc3c8235143766fc6be8c7e12b31b92222594698c966462f944f13fb5929113"}},
{"executionId":"stock-harness.legacy-acp-codex.local.ordered-comment-continuation","profile":"legacy-acp-codex","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["contextIntegrity.ordered-comments","contextIntegrity.changed-scope-preserved","contextIntegrity.packing-report-order","contextIntegrity.final-scope-applied","contextIntegrity.continuation-run-count","contextIntegrity.continuation-wake-comment-ids","contextIntegrity.single-durable-output","contextIntegrity.completed-task","contextIntegrity.successful-runs","stockHarness.invocation-evidence-complete"],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":50896,"providerDurationMs":39443,"cleanup":"passed","billing":{"llm":{"runCount":4,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":992,"outputTokens":46,"cachedInputTokens":24349,"totalTokens":25387,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":39443,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":false,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":14032,"reportedPromptChars":14056,"freshPromptChars":4687,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"33ac59f0a690ed76618f695e658575751e718f53f9d828d8b52c690df9c15afb","evidence-manifest.json":"92c048e01012a77b94003e38f3f73888d725f6942327f7fbf3e6d2a42d12184d","snapshots/context-integrity.json":"2f9a803657068a67bcbf3b190370395187ef13b7a1487705d0305e121339eac7","snapshots/stock-harness-hire.json":"f47bb35d606ae00e6057a2a77769e2091e8ccb12eca08db3a17e1ed297a8ccf4","snapshots/stock-harness.json":"05582c37956b55f6a2bd7cad0c55bc0fa1dbff875d7dbed12b83352f6bd4e04c","snapshots/stock-harness-preflight.json":"badf01a591f5327f0dc2e48865711f8f4d7327888a7688668728173a29b273a8"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":3,"failed":9}},
{"executionId":"stock-harness.legacy-claude.local.assigned-skill-explicit-invocation","profile":"legacy-claude","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":49054,"providerDurationMs":29420,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":29883,"outputTokens":1348,"cachedInputTokens":225730,"totalTokens":256961,"reportedCostUsd":0.19999125,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":29420,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.19999125,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.19999125,"complete":true},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":7512,"reportedPromptChars":7512,"freshPromptChars":4684,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"4fc73f7b1226b9a41b209a2f6dfa2a58805ddc1a7e7b699ac2b42b2e74a189a7","evidence-manifest.json":"440ee14cca3f8558fb47ecf27b17693c0f9dbecdcf5c8532d7e2ccbefe31700b","snapshots/context-integrity.json":"4545847b9c4ab2a4fa3053a384b5e990133705fc9da0d1e9cd98f1275ec05f44","snapshots/stock-harness-hire.json":"a6abfc40d1c788e4b897f2375506d59b2caba9fab0d4df019e7137fe9f66177f","snapshots/stock-harness.json":"755cc9d5f656dc068a08c8c141203eaa5d3d60b53edd9d516887aae108119f45","snapshots/stock-harness-preflight.json":"f52730b4555ab31de952c8c1e38db4067fb40756b6befd36ed4b686e1a273d73"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.legacy-claude.local.continuity-restart","profile":"legacy-claude","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["chat_memory_assertion"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":40202,"providerDurationMs":23072,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":39036,"outputTokens":829,"cachedInputTokens":116171,"totalTokens":156036,"reportedCostUsd":0.1936638,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":23072,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.1936638,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.1936638,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":11451,"reportedPromptChars":11451,"freshPromptChars":1458,"retrievalClipped":false},{"capturedPromptChars":11508,"reportedPromptChars":11508,"freshPromptChars":1458,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"f8e377f617204eabdb6b4e0c24ab3280c51664b429051f7ea21a4746790820a6","evidence-manifest.json":"45818efcb8ed6d663f7ffab6d53605d5191fb983a68e422dadebdae875ed5b4c","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"e3a8735ad4d5f89dc3899668c36791579562157fa25aa32a6e59aa098d0c079a","snapshots/stock-harness.json":"f3d24452ade2f232b0f671768d46e4d65a6a0c2a0740bb1daa17a4871a09aa7b","snapshots/stock-harness-preflight.json":"b85c96ada9801ce6c6569338b68bc810494a0711f01af39a63932e288f7a81a3"}},
{"executionId":"stock-harness.legacy-claude.local.ordered-comment-continuation","profile":"legacy-claude","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":142017,"providerDurationMs":119179,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":66601,"outputTokens":6610,"cachedInputTokens":783643,"totalTokens":856854,"reportedCostUsd":0.58397715,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":119179,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.58397715,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.58397715,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":14473,"reportedPromptChars":14447,"freshPromptChars":4684,"retrievalClipped":false},{"capturedPromptChars":7926,"reportedPromptChars":7926,"freshPromptChars":4684,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"9161c05c288e8620e5a748a28c901e1582efc240ca640203e1d8f7527e22a651","evidence-manifest.json":"296246a83eb0870a53e70f359f1859719f30b655e6ed20908abbfe4213605280","snapshots/context-integrity.json":"17779f4ebc38ed08071e40ac4b939149a18dfceb76c0f1f4f6ed21bf5d0153fa","snapshots/stock-harness-hire.json":"1f5a97e38555f57f47ee1900440d88ee521de75e7589c163f7688c852b661dc8","snapshots/stock-harness.json":"499a0b49c042b7c44f992c56983d207677cf619306ed6f6db9975ef1b84d455b","snapshots/stock-harness-preflight.json":"8f92c4b703533a645196bf03d1353be7196271d25638be7cb3fae6ec92c69c52"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.legacy-codex.local.assigned-skill-explicit-invocation","profile":"legacy-codex","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":50078,"providerDurationMs":27658,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":128946,"outputTokens":1727,"cachedInputTokens":97854,"totalTokens":228527,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":27658,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":12332,"reportedPromptChars":12332,"freshPromptChars":4683,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"ea06fc8b1ec166c9932de8634d08e528e90f7b734388ccad6d47b4b96df9ff6d","evidence-manifest.json":"45e71226e99eaee9bea0b0fae2cca35de573c2ab9d5add13159f321c8480ac0b","snapshots/context-integrity.json":"4cf10f812f258a77b472abd74f6691224dfcba4b9a2dec9c0f2c35c5ae741d7a","snapshots/stock-harness-hire.json":"19db78dbf51a67ff4bc307ef378da9cc024a0cae1c977951b49d089c1ec47fe4","snapshots/stock-harness.json":"a3cb7e2bf4e775cf9058b6be33d8df5321e312df2aed39736719454647610a8f","snapshots/stock-harness-preflight.json":"e9a116f8605a2630ff46075598ac6bae30817892d76efcd9a25891649cac0856"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.legacy-codex.local.continuity-restart","profile":"legacy-codex","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":85550,"providerDurationMs":29332,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":0,"inputTokens":468597,"outputTokens":2919,"cachedInputTokens":373648,"totalTokens":845164,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":29332,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":2787,"reportedPromptChars":2787,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":2730,"reportedPromptChars":2730,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":16326,"reportedPromptChars":16326,"freshPromptChars":1457,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"db7be8127de96e697a9d274b7a18db40695daa234684ec4d8e28afd8665c1f3e","evidence-manifest.json":"2e9151536b8ade2bfbb3c465438e8ace21417258c4037b6923920ee0784b0e47","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"af8b2fe11402a8b87a056818df29a62829ec9425f8a2d955e95692980e15e043","snapshots/stock-harness.json":"739e642582faf89b74c522774c64ea26c6af623539b7a8f99ce332066c965600","snapshots/stock-harness-preflight.json":"8fe9f32d7629a1f6bd632d58e22e39e47baff2b8bc3b74e3593e6411ab3c9a2b"}},
{"executionId":"stock-harness.legacy-codex.local.ordered-comment-continuation","profile":"legacy-codex","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":74114,"providerDurationMs":57150,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":0,"inputTokens":628626,"outputTokens":7098,"cachedInputTokens":563211,"totalTokens":1198935,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":57150,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":9811,"reportedPromptChars":9785,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":12746,"reportedPromptChars":12746,"freshPromptChars":4683,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"56a457fcc7a36809eb4bfcedeff2fa079f5272754ff52fc453e8bf8d813e52f3","evidence-manifest.json":"bf17fb7d2a05f7e99e0723907bdc57f0a88bf6838d68f31f4b84dc6ebc62d3c7","snapshots/context-integrity.json":"a9156b810e32fa1d9d07826276a00c62ebe4f25ac3dc284ada157f2b2d72d800","snapshots/stock-harness-hire.json":"5fc8abd51d401c96d701b0862ac2dc674eb36244fdc25c49b67ba3947ba5a2fb","snapshots/stock-harness.json":"c72b8c53e1e50a82ab6ff815b363942936440cfff75f8f8ca5f54185c723889d","snapshots/stock-harness-preflight.json":"b14d304196cdc3146436c66199243c6da04784c13e45c9bd31e0ef83e97a29bc"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation","profile":"legacy-opencode","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":86384,"providerDurationMs":68956,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":57002,"outputTokens":2365,"cachedInputTokens":223232,"totalTokens":282599,"reportedCostUsd":0.0044563934000000005,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":68956,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0044563934000000005,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0044563934000000005,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":12335,"reportedPromptChars":12335,"freshPromptChars":4686,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"20dbb34bb975f73c3ac2695c047b6cbc8f6382757939ae6210b08a0f13ba9e9d","evidence-manifest.json":"67c3e042be8e0a5a0cd7cca82d1b0887e1392a2806c3b4db67baa3e99365b6be","snapshots/context-integrity.json":"decf2d93769acce2e2990de27c81ee32fd805ca53d3c2be68c56f22a013a7ac0","snapshots/stock-harness-hire.json":"187598ed72837719be8d6f0b77c5f8d35cfc871f8cf41d4fdc5b259f88171025","snapshots/stock-harness.json":"98424edd880cd05e0cbf8c32973fe9fbfc055037a74e76c20f893c3ee263f6a0","snapshots/stock-harness-preflight.json":"86d9f4080a2045ae3fa68468d0ffaf68baa26f27f0528274f13762f093b4c113"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.legacy-opencode.local.continuity-restart","profile":"legacy-opencode","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":126589,"providerDurationMs":78910,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":64286,"outputTokens":1046,"cachedInputTokens":28672,"totalTokens":94004,"reportedCostUsd":0.0022020907,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":78910,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0022020907,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0022020907,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":7611,"reportedPromptChars":7611,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":7554,"reportedPromptChars":7554,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":16335,"reportedPromptChars":16335,"freshPromptChars":1460,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"29dce6418aa35bb8070b07433046bbf39c9a8367a198cdb0aa28a0bea480032c","evidence-manifest.json":"439dc46130b607b7488e9df7736a711b267efa53689fe9fdad2f822525677215","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"7f7ad64827bbe430e06a7e01b1377443653793584264c0fa00ac88a742bd0f49","snapshots/stock-harness.json":"4dd536951a950e7507710a9eb44cdbcd354907397f1e95b8297b28a98fd75dbc","snapshots/stock-harness-preflight.json":"8fa2ec342ad583fb20fed4db88890fc31eab3ef5898a952a85371e53dd433aaf"}},
{"executionId":"stock-harness.legacy-opencode.local.ordered-comment-continuation","profile":"legacy-opencode","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37048838402","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["deadline_or_run_start_timeout"],"failedMatcherPaths":["contextIntegrity.comments-queued-as-batch","contextIntegrity.initial-run-started","contextIntegrity.ordered-comments","contextIntegrity.changed-scope-preserved","contextIntegrity.packing-report-order","contextIntegrity.final-scope-applied","contextIntegrity.continuation-run-count","contextIntegrity.continuation-wake-comment-ids","contextIntegrity.single-durable-output","contextIntegrity.completed-task","contextIntegrity.successful-runs"],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":720616,"providerDurationMs":442902,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":51446,"outputTokens":8911,"cachedInputTokens":967936,"totalTokens":1028293,"reportedCostUsd":0.030657645200000007,"costStatus":"partial"},"runtime":{"provider":"local","agentRunDurationMs":442902,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.030657645200000007,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.030657645200000007,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":16392,"reportedPromptChars":16461,"freshPromptChars":0,"retrievalClipped":true},{"capturedPromptChars":16406,"reportedPromptChars":16879,"freshPromptChars":4686,"retrievalClipped":true},{"capturedPromptChars":12749,"reportedPromptChars":12749,"freshPromptChars":4686,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"6d0e5d54b1742b02eb37346747510c66cb5cd0b80262a6c9590fd2a28861dcb2","evidence-manifest.json":"758931a4d30213d1b320f58cd2aa338e8a6a8aed579b878e14c5e52848bd5f13","snapshots/context-integrity.json":"bcaa4ee278361a90347625044aa49086b27093356abc0f3700b0363d28c8653c","snapshots/stock-harness-hire.json":"21431fa5e7cb6486037c670fc1e4aa64fb24716ad8ab325601ff294ada40263f","snapshots/stock-harness.json":"29c55e52f4adbbc9259cb9e221c08bfbbd7ad4c5f8247c24b3f029362e9000de","snapshots/stock-harness-preflight.json":"f341723074c996c79d8a6c30b9f12270634b20b40da452b13e4ab9ce91b47cac"},"taskDocumentCount":null,"behaviorCheckCounts":{"passed":1,"failed":11}},
{"executionId":"stock-harness.runner-acpx-claude.local.assigned-skill-explicit-invocation","profile":"runner-acpx-claude","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":46264,"providerDurationMs":25372,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":10,"outputTokens":1042,"cachedInputTokens":141254,"totalTokens":142306,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":25372,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"073085f804a58d191bd44972edfb08d914fe622a86f236d0c4dc1b2bf8a03322","evidence-manifest.json":"4fad428c1df72abbfee24a4f7338e25284fc874cff1220e6b631fa86eb2d0bed","snapshots/context-integrity.json":"77e27b8ea01a83e0b65f3d5410e0787fbf84a7f6ea02f4401bd77488999dd4a2","snapshots/stock-harness-hire.json":"8c6c5d56fd068c15f1e19ce9a0b62204443c894d204311bd949d5e22c624fd3d","snapshots/stock-harness.json":"2a6860d7c29ad0d03240bda0a779b83d0455b8ce9fcfb1cf9e6ad559c4dec112","snapshots/stock-harness-preflight.json":"79b11e544d2821d1025476f5fd0d737e822ae5ab731ab32047d1fbf6b76015e4"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.runner-acpx-claude.local.continuity-restart","profile":"runner-acpx-claude","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":77424,"providerDurationMs":36820,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":6,"outputTokens":94,"cachedInputTokens":94997,"totalTokens":95097,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":36820,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"be582fe8af5f19543c16c8590697c5aa9dc688fdb653b0550aba0b55174becd1","evidence-manifest.json":"cfbd1cb13917a96b08dd46aae1fcc63e2b27af5d1d22f4952974d3415c74ffb8","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"254f18a8e8e1d6a5a2e7731728756671c1374cd4d1210469172bed9c2cd6e4c2","snapshots/stock-harness.json":"8efa84d6907de220dd7b9fc2c384c61aae0606e7f5311f0e926e26ae1617d869","snapshots/stock-harness-preflight.json":"c7e3f2342ef9fc824b28e3f22ffbbb326d2ba1ff84e36c4deedc56a7abefbe07"}},
{"executionId":"stock-harness.runner-acpx-claude.local.ordered-comment-continuation","profile":"runner-acpx-claude","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":132587,"providerDurationMs":108895,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":26,"outputTokens":7996,"cachedInputTokens":565030,"totalTokens":573052,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":108895,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"62c6c55047813b0b5aa5334e29dbe3f90225c83ac436eab14afd97d78ed9ec65","evidence-manifest.json":"7d53ad7f80b65a1c7af3fad4fa3d0c652ae25cdaeae3d41cbee1bb888b853fb5","snapshots/context-integrity.json":"2081459b761a84469399bec377d7380c172b64fddae48735c5329771656e28a7","snapshots/stock-harness-hire.json":"f11c90fdae215949458d61a6e1ca6565fe277ea30d52fd922c05cffdf4867e5c","snapshots/stock-harness.json":"f303a3fb5c5e55cefb696d46204aafe6f1cefc5b0a6b07a728f2bf0dbf1dfdcf","snapshots/stock-harness-preflight.json":"79768aa58996d28ee80c8a9e8d5c4ba1a42eba58a408163cd8a6037ac6005b26"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.runner-codex.local.assigned-skill-explicit-invocation","profile":"runner-codex","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":53751,"providerDurationMs":29850,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":162041,"outputTokens":1296,"cachedInputTokens":137822,"totalTokens":301159,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":29850,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"ffcc5203624c19b0acda50dfd1e283c595ae3ed6072e4aaf9781fef4bf08fa45","evidence-manifest.json":"1460fae9f00a241a3ddbe7f623a4c1c6cf0f6434e73e47d0de7f98dc00dec459","snapshots/context-integrity.json":"63d16c6558efcf03097fd2ef8caf1f01c39a38127aee2904eba0cf0edfb711e6","snapshots/stock-harness-hire.json":"61630a0e0521f8ccca30753bde385b731f5c8c15cf6280b000947e80acf3decf","snapshots/stock-harness.json":"bc95a45aab085ce3eee988a9b235281b847cd8cb73c4d294ce576b97e73d879f","snapshots/stock-harness-preflight.json":"4daf4cd83a8b1539ab8a8db2fd5cf204746884d8cde04f6507acfff5be75ef11"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.runner-codex.local.continuity-restart","profile":"runner-codex","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":84362,"providerDurationMs":35132,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":315100,"outputTokens":2339,"cachedInputTokens":231408,"totalTokens":548847,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":35132,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"dbf2e14fcda52d2a33808e5a4c226d45f864752274c823c26beb7c2ea5409f03","evidence-manifest.json":"86a34f1a25854354bee0d9f0ea88cab9ba9b51476cf274d12f5b491a9910188a","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"a38cab643c78fde9d61d15074b82297b94bf46ae5c3296aa06e6958f67188eda","snapshots/stock-harness.json":"2d148d3e55c57619bde56558f7e87e85cbb1305e46eba83bedb21c3ee4b43b54","snapshots/stock-harness-preflight.json":"9517ef436cc83da9078dd75587f01be8a28685131c37773dc092cf89c137ca0c"}},
{"executionId":"stock-harness.runner-codex.local.ordered-comment-continuation","profile":"runner-codex","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":92663,"providerDurationMs":70832,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":428674,"outputTokens":3591,"cachedInputTokens":372482,"totalTokens":804747,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":70832,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"8c7a81172e0d519ffc8aca9a3025c88166d93d12a1744e560813c66063d6d130","evidence-manifest.json":"b495ad141c12f2458e3b66123cb673dec68511bcb5c2d59af78400dc67c417eb","snapshots/context-integrity.json":"166eb5a26fa389529287aa940c4e3caa85665ea67fcce019bc69a8bb8d9f7b41","snapshots/stock-harness-hire.json":"2fe4b0d207648628658312e97b491336929197f826c00dce012b6afc43537564","snapshots/stock-harness.json":"79b83fed61c69589d38b48552b04dce41bf41dca27f6e7f5dd15229024b6338f","snapshots/stock-harness-preflight.json":"689ce3141b2c6970eac5468885e5111daa730bcda25a1459e7bac4177650045e"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.runner-opencode.local.assigned-skill-explicit-invocation","profile":"runner-opencode","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":90297,"providerDurationMs":68028,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":44259,"outputTokens":1165,"cachedInputTokens":62976,"totalTokens":108400,"reportedCostUsd":0.0020380985,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":68028,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0020380985,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0020380985,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"3ffa3bff50fbade8801da7c26ffed426eee9bda9a7e3640793f03a220f74ea82","evidence-manifest.json":"6df0c3b49ef54154721d4d0ba434f893fcba372ba8588fef2bd1a5947284b1d3","snapshots/context-integrity.json":"41b60c88adb0c529148981a55ab3b18964035b47786da50942183b58e8884fed","snapshots/stock-harness-hire.json":"c9a2fe249e217fcd5c205f1dc768446333932de74d87dc6f4a1e7855c5e4d815","snapshots/stock-harness.json":"474b0d6aecd8e8bce23f78c260ad4dc83c286d1d39917ea0920ad4b402d9404b","snapshots/stock-harness-preflight.json":"baeb38686ea4d3233906d83a889fdf3d805846b27f25c6c04da68842822a496d"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.runner-opencode.local.continuity-restart","profile":"runner-opencode","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":155833,"providerDurationMs":101201,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":48028,"outputTokens":96,"cachedInputTokens":23808,"totalTokens":71932,"reportedCostUsd":0.0004892436,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":101201,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0004892436,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0004892436,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"57fe32615f16ff3706044419bb81362a67ec35aed7a5f74a5b8207b5831a5cba","evidence-manifest.json":"07e3be863cc20a1352a094f38e6133004dbfd673f4fd59599431a022ce061ff1","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"d127504476aee2130312fdf092a523060188066abe9f969ae9b178b57374cced","snapshots/stock-harness.json":"0df941763ed917e2c553a6c6a4dfb0c87bdb7181d42a28f614c765906b74f1e9","snapshots/stock-harness-preflight.json":"f57655c970472d2a9b9bebe100d8c31788ed76e06da4363046a5a6d3f09269bb"}},
{"executionId":"stock-harness.runner-opencode.local.ordered-comment-continuation","profile":"runner-opencode","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["deadline_or_run_start_timeout"],"failedMatcherPaths":["contextIntegrity.comments-queued-as-batch","contextIntegrity.initial-run-started","contextIntegrity.ordered-comments","contextIntegrity.changed-scope-preserved","contextIntegrity.packing-report-order","contextIntegrity.final-scope-applied","contextIntegrity.continuation-run-count","contextIntegrity.continuation-wake-comment-ids","contextIntegrity.single-durable-output","contextIntegrity.completed-task","contextIntegrity.successful-runs"],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":721578,"providerDurationMs":446042,"cleanup":"passed","billing":{"llm":{"runCount":11,"runsWithTokenUsage":11,"runsWithReportedCost":11,"inputTokens":159917,"outputTokens":7291,"cachedInputTokens":1077987,"totalTokens":1245195,"reportedCostUsd":0.015645790399999998,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":446042,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.015645790399999998,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.015645790399999998,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"f727dd5f76162c4b995e649e1a7bf55a52381efa550e3bda46ec0a51a899d178","evidence-manifest.json":"6a550c598b67f60fcea70d56907c4a2c84239bc0d01e58737e99a0014144bb70","snapshots/context-integrity.json":"14212826847fbd9d6d63c7b70afb102fdb5d0100f7d8f1368d46ff6f5e32db07","snapshots/stock-harness-hire.json":"ab6dd2377fc251e44ed8f3680d40d10a315b635f3003e45501c7496037fb21ca","snapshots/stock-harness.json":"38759bca194a8038f0d2be4598e50a04d61215ecabdfec95cd7f6cfa4eeec61c","snapshots/stock-harness-preflight.json":"312611057436fd8cbdcc5179ffa5e62144800a6a9244e646bb6e0b3a9a0cb551"},"taskDocumentCount":null,"behaviorCheckCounts":{"passed":1,"failed":11}}
],
"candidate": [
{"executionId":"stock-harness.legacy-acp-claude.local.assigned-skill-explicit-invocation","profile":"legacy-acp-claude","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","no_single_paperclip_document"],"failedMatcherPaths":["contextIntegrity.single-durable-output"],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":123175,"providerDurationMs":111928,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":25333,"outputTokens":5771,"cachedInputTokens":316215,"totalTokens":347319,"reportedCostUsd":0.2838505,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":111928,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.2838505,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.2838505,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":10966,"reportedPromptChars":10449,"freshPromptChars":787,"retrievalClipped":false},{"capturedPromptChars":6028,"reportedPromptChars":5498,"freshPromptChars":787,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"2a5fed84ee0f95a67e2c2b0d4247ffae33e16a7c2aa11dc57b11228e0db90965","evidence-manifest.json":"7e31c707e866e15bf44a276ac252c423a47a63364e76ad85de9a8d5b6a4e1dfd","snapshots/context-integrity.json":"3c36c18d5e257e788514c02b50d63c7c69885201df0a07750129d55c0f583676","snapshots/stock-harness-hire.json":"2d128826b4e5068ad6ed819512bcb17160346becd9344b8ab8e70c984c3241b7","snapshots/stock-harness.json":"49b8da640ab0b9dd2b209ce0cf9f3bec6deceb801bfe98c0166d7702a9732a80","snapshots/stock-harness-preflight.json":"a12fdf9b8b7d761ff223e30dcaeb835ad9bf32ef7f1c8b8983a638d175c2ab5f"},"taskDocumentCount":0,"behaviorCheckCounts":{"passed":6,"failed":1}},
{"executionId":"stock-harness.legacy-acp-claude.local.continuity-restart","profile":"legacy-acp-claude","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":34813,"providerDurationMs":18798,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":16797,"outputTokens":329,"cachedInputTokens":54375,"totalTokens":71501,"reportedCostUsd":0.092943,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":18798,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.092943,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.092943,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":13201,"reportedPromptChars":12723,"freshPromptChars":787,"retrievalClipped":false},{"capturedPromptChars":13233,"reportedPromptChars":12755,"freshPromptChars":787,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"ecab7c9f83668fd4712702bb906adf38e188e34f7f07bd7d746da685339a2ecd","evidence-manifest.json":"544a726ddf3ac212671f7c3d73dafcf94cd5b3abb9a3ad5197d6c793e9d7aaa0","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"78f5c467789921db4d2194ba80459524b04644f2d8d8a3b8678fdf7686004e8f","snapshots/stock-harness.json":"61641e5e25351e748afb2e4533d70b1cc34a8b78e1fbc56a27d96b2017bd0f45","snapshots/stock-harness-preflight.json":"4e522117d8489a16263c3d66ca513207217a7861733bfd11ef2f7149aef680e9"}},
{"executionId":"stock-harness.legacy-acp-claude.local.ordered-comment-continuation","profile":"legacy-acp-claude","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":154765,"providerDurationMs":139409,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":61445,"outputTokens":6979,"cachedInputTokens":562903,"totalTokens":631327,"reportedCostUsd":0.5120724,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":139409,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.5120724,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.5120724,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":13419,"reportedPromptChars":12915,"freshPromptChars":787,"retrievalClipped":false},{"capturedPromptChars":6390,"reportedPromptChars":5912,"freshPromptChars":787,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"9c23fec76e718c1fa5f775ae8cdcee1e5433039ff74d7ffbf306b5be6693700d","evidence-manifest.json":"5284aba88b923322b1f162045d20e36236fb731689527a1a1d06d5305efd4a64","snapshots/context-integrity.json":"8fef43dac791a84db62bfea8977f586afe3e5dc5c36225eaab79faaa15b4dfca","snapshots/stock-harness-hire.json":"4961a6103fcb6b58d5d65981c68afa178ea9e1eb6038ad3b3049d47e4bc737ab","snapshots/stock-harness.json":"4e11dd1b37581fed8628b7ebd953ebf3b6fe389d357ffd2efcd9e6b976e4d64a","snapshots/stock-harness-preflight.json":"7d962ba46f4f02f111ed4e423705009f38ee99096d180b693f4c41ed4c805126"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.legacy-acp-codex.local.assigned-skill-explicit-invocation","profile":"legacy-acp-codex","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":65745,"providerDurationMs":49358,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":1379,"outputTokens":34,"cachedInputTokens":41987,"totalTokens":43400,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":49358,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":5509,"reportedPromptChars":5533,"freshPromptChars":786,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"81c5c35a942c78a2c9d7d77064264504311a191c6dd421e656d4f0751670b6ff","evidence-manifest.json":"bff59a88d627186d18282fe39822dd5a57d3a2396ae9aca5fa178a3b4ac3cf60","snapshots/context-integrity.json":"3809130fd34f0b6035ed5b17bd9e354d19aac02fc2f64aea01811803151c7f82","snapshots/stock-harness-hire.json":"4e0468ee05d56f08faa9ffbe3ed6cf28c07490a8b27a08e80875264dc66daac5","snapshots/stock-harness.json":"c5a82192ff53c6eec7a5497602a9969af3c21ea7fde8fdff4309025f5ccf80d2","snapshots/stock-harness-preflight.json":"12944d93aac9f8688d2832369f7eadc10a9546a6fb6bc2f2d8b3caef38b0cc2f"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.legacy-acp-codex.local.continuity-restart","profile":"legacy-acp-codex","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["stockHarness.invocation-evidence-complete"],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":35779,"providerDurationMs":20059,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":711,"outputTokens":30,"cachedInputTokens":31118,"totalTokens":31859,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":20059,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":false,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":12764,"reportedPromptChars":12788,"freshPromptChars":786,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"1c59fb6b3845ef6bc4f9cc549826e45334b3220ce5832f2fe4b9d23fd45136b6","evidence-manifest.json":"4cdd7cc5bb0a025e7626e947dd72558c13203aa9dc9765f5afbfed5acc907de1","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"dcbe621edcd9437c47361c74676db38a240fdcaff2401d41213a185e7cf9b785","snapshots/stock-harness.json":"8908438e0c7565ebca6d4753e619f87bd0c015210214fbb8c39876d7e13cebaa","snapshots/stock-harness-preflight.json":"3b2de5b161384bebe03f7cad3b3268f98cca291cd8fbf4b8a0c830536daecf87"}},
{"executionId":"stock-harness.legacy-acp-codex.local.ordered-comment-continuation","profile":"legacy-acp-codex","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":["contextIntegrity.comments-queued-as-batch","contextIntegrity.initial-run-started","contextIntegrity.ordered-comments","contextIntegrity.changed-scope-preserved","contextIntegrity.packing-report-order","contextIntegrity.final-scope-applied","contextIntegrity.continuation-run-count","contextIntegrity.continuation-wake-comment-ids","contextIntegrity.single-durable-output","contextIntegrity.completed-task","contextIntegrity.successful-runs"],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":720307,"providerDurationMs":46665,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":540,"outputTokens":28,"cachedInputTokens":32395,"totalTokens":32963,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":46665,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":5923,"reportedPromptChars":5947,"freshPromptChars":786,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"a1285f2fecccdcdac85360784e4be82188b564d3135f7adf6606fa1451307376","evidence-manifest.json":"39864272e9331c88c1266f2b5f05b755aacdc46c2c7015c4d4df8539c0a22718","snapshots/context-integrity.json":"5e663656093c672fb294719edb8636dd15a5ed1441912f32a32c78254663b797","snapshots/stock-harness-hire.json":"35d859103bf2acb7c7f6d212a8ce277e42726da388f7c52ee0b05d600dc09c78","snapshots/stock-harness.json":"483b4efedcc34570ad0ad364d92ba0212186cbc36a704639106e181402470468","snapshots/stock-harness-preflight.json":"5bb04e0c93a3f4ac5d9f41a5775fba85e0687a8161e255257545d54151421c62"},"taskDocumentCount":null,"behaviorCheckCounts":{"passed":1,"failed":11}},
{"executionId":"stock-harness.legacy-claude.local.assigned-skill-explicit-invocation","profile":"legacy-claude","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["no_single_paperclip_document"],"failedMatcherPaths":["contextIntegrity.single-durable-output"],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":37720,"providerDurationMs":25763,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":27376,"outputTokens":1281,"cachedInputTokens":184039,"totalTokens":212696,"reportedCostUsd":0.17707545,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":25763,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.17707545,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.17707545,"complete":true},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":3611,"reportedPromptChars":3611,"freshPromptChars":783,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"8636a50a54d39fd66b1439f13ce5ac125b79f16b2826f76fc297427a1a2ab701","evidence-manifest.json":"d1d7f47272981406d104f5ae6455d5b0a4f6113f3c3d02fc8bc787d7c8fc70fc","snapshots/context-integrity.json":"eac7d7b889e5cc3861b937ffa87c7b92d25ade5b0f2bb73bdd29b5764c2a085a","snapshots/stock-harness-hire.json":"545030ff52d9196a604e099b81b52a3071f582f2467284c695eb62d96589c67c","snapshots/stock-harness.json":"c9ca463a792c7dccd00c0795c1e405495d9e1dff30b2cbae2b5e7bac7114fdd2","snapshots/stock-harness-preflight.json":"c1b5a843459d80e074759600c830c92bb00a0d2430255cda7bed5c22cabc7680"},"taskDocumentCount":0,"behaviorCheckCounts":{"passed":6,"failed":1}},
{"executionId":"stock-harness.legacy-claude.local.continuity-restart","profile":"legacy-claude","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["chat_memory_assertion"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":28251,"providerDurationMs":11843,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":19490,"outputTokens":298,"cachedInputTokens":50061,"totalTokens":69849,"reportedCostUsd":0.09257055,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":11843,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.09257055,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.09257055,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":10776,"reportedPromptChars":10776,"freshPromptChars":783,"retrievalClipped":false},{"capturedPromptChars":10833,"reportedPromptChars":10833,"freshPromptChars":783,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"86af4c61a239c718181ab58fd815cf7a200e7473b2fe9390b58bc15ff9d6ed4c","evidence-manifest.json":"e1c69d7c6044b0e8ba65fa1b961de84989facb3fd65f6bba42632b22d2a1bab9","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"9a320857e3d1fe14a41b0154b8c1d2cfc81f869ed69790bd729fe7c86e9c54fb","snapshots/stock-harness.json":"9a2d546c992bf9d59711e37e81d841b7b1c73e6e1d62a6f27b5302f1d8ea3980","snapshots/stock-harness-preflight.json":"13318522befc6c0ffcc0d6ca7a0dc12112902427ddc1a7ac36f3ac43e7f83d8d"}},
{"executionId":"stock-harness.legacy-claude.local.ordered-comment-continuation","profile":"legacy-claude","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":158440,"providerDurationMs":139386,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":64318,"outputTokens":7515,"cachedInputTokens":669889,"totalTokens":741722,"reportedCostUsd":0.5548662,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":139386,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.5548662,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.5548662,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":11079,"reportedPromptChars":11053,"freshPromptChars":783,"retrievalClipped":false},{"capturedPromptChars":4025,"reportedPromptChars":4025,"freshPromptChars":783,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"59885705b15fc33517ab7f7a82d9ada2e411c918b10fa5b4ce928498a801e739","evidence-manifest.json":"c2a5be637da0d7ff9436b732d8db72ab1fdce788caf6a40c5973c14a27bcc08a","snapshots/context-integrity.json":"fbee7c1a37b964bcc327ebbf37602bc227bccb7fafdc4296614e178e66016599","snapshots/stock-harness-hire.json":"cf2f8ff4b1eb13c5e5f854ecac27c25826e7a4dbafbc277e8702b8beee6d913f","snapshots/stock-harness.json":"3e8e7c06f0918a896341a38e4613cc1f4caaad4b5ab12c03a8c3e41c4707a82a","snapshots/stock-harness-preflight.json":"6d6716a08fc2ea7d27ee08840fd3290a04a8e85dfed5950b690f4c2267d20a6e"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.legacy-codex.local.assigned-skill-explicit-invocation","profile":"legacy-codex","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":44075,"providerDurationMs":27042,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":241944,"outputTokens":2132,"cachedInputTokens":215273,"totalTokens":459349,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":27042,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":4223,"reportedPromptChars":4223,"freshPromptChars":782,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"b0b26f37d2f0c0342a9a275b22716be190891d1b55e79fa8d33e3f495649a26b","evidence-manifest.json":"446abc91b7366507c4100741d71ace903de0b78149b33e45fa90d39183f71b75","snapshots/context-integrity.json":"85197eb92d34334cc2de658d09d7a7e54d44ab3d2fd81ad87732c88c6610b6b2","snapshots/stock-harness-hire.json":"568453725988fc61d459f96efda07d19ec5c00935d4c4a9ff9557cd0062aef29","snapshots/stock-harness.json":"0c8626740b520687d08b0b1f0aab8a9ead2ebd58290252f05d890f524636f461","snapshots/stock-harness-preflight.json":"9a15351a65076deeb77ccd58216def768a658fc743fd57dc89c3234ff6b36542"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.legacy-codex.local.continuity-restart","profile":"legacy-codex","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":62742,"providerDurationMs":10028,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":0,"inputTokens":91929,"outputTokens":475,"cachedInputTokens":45089,"totalTokens":137493,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":10028,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":2787,"reportedPromptChars":2787,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":2730,"reportedPromptChars":2730,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":11443,"reportedPromptChars":11443,"freshPromptChars":782,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"1ec35f3857350d86d3888c4becf8ce898a59641f3eb9825a0d800970ec7090ad","evidence-manifest.json":"908761dcb779dcd0b95ef78b5e9addb77b211acbc502a1c8fdf793202d309472","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"b7b6f4db7cdc84a795ccfa24df58dd8131fe8a44b3ef1cdaf36dfc6685556da7","snapshots/stock-harness.json":"9a0045bbf4240ddbf08ec0444e08ef70e3ef76399db09075b5316b1fc51ac5bf","snapshots/stock-harness-preflight.json":"cfbae9842e2a045b815a9a7f1fa471f8f9494ec999c91e9808933cbc013c84a4"}},
{"executionId":"stock-harness.legacy-codex.local.ordered-comment-continuation","profile":"legacy-codex","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":66403,"providerDurationMs":44175,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":0,"inputTokens":416426,"outputTokens":4483,"cachedInputTokens":353673,"totalTokens":774582,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":44175,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":8456,"reportedPromptChars":8430,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":4637,"reportedPromptChars":4637,"freshPromptChars":782,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"71f99ec1b45a57ae40acbd42c55e835f63d9d2329752f6bffcabd4488ecb2a96","evidence-manifest.json":"0c3347aaec6d7c3498e62c19938762c35a6069d31385e5d617174695ab60462e","snapshots/context-integrity.json":"5c08b2af1b27e9255505e917958e7b97856cfbf48cce3bef3b8e792275823b6e","snapshots/stock-harness-hire.json":"ddc6b182c4d53c7cd7f63c906d2a64e6286c54c8b6e9991e0c78da578586d865","snapshots/stock-harness.json":"11a737f54107eb9a9e55234890bce90003f84b258f6d5c8512d65f080edfc52f","snapshots/stock-harness-preflight.json":"d72a98df5a4e00bfa766e85ef28b866d32b89fe2aa9a67e3000ad50ddea47599"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation","profile":"legacy-opencode","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["no_single_paperclip_document"],"failedMatcherPaths":["contextIntegrity.single-durable-output"],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":212558,"providerDurationMs":199337,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":96260,"outputTokens":6708,"cachedInputTokens":172288,"totalTokens":275256,"reportedCostUsd":0.010608548800000001,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":199337,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.010608548800000001,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.010608548800000001,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":8426,"reportedPromptChars":8439,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":4226,"reportedPromptChars":4226,"freshPromptChars":785,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"435f708b59438d1d344130bc9e23b8ddbc93adf74d4e5ea8f049e0f1d06efa72","evidence-manifest.json":"8591f98df496159f99b98659ffdcbdf2e86718dccb4545d461f592bab0626b05","snapshots/context-integrity.json":"e4271e26bb891c3c15f2d54dd5cfba6605ebb4431456d70b4f146bb8eb4e1788","snapshots/stock-harness-hire.json":"6bb95cb92fbb14976b5f6796f530823f6aacdcb1448f89cb8b2ec093125266e2","snapshots/stock-harness.json":"b8a5a04730324a038869e28f79d55dd06ebf12abc2e6366bbfe9218995a6748b","snapshots/stock-harness-preflight.json":"c17087412694993475276a69162a171859dc54afae5b08ab89573cf0637d6e19"},"taskDocumentCount":0,"behaviorCheckCounts":{"passed":6,"failed":1}},
{"executionId":"stock-harness.legacy-opencode.local.continuity-restart","profile":"legacy-opencode","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":69761,"providerDurationMs":19747,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":12219,"outputTokens":116,"cachedInputTokens":0,"totalTokens":12335,"reportedCostUsd":0.0003185752,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":19747,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0003185752,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0003185752,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":3403,"reportedPromptChars":3403,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":3346,"reportedPromptChars":3346,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":11452,"reportedPromptChars":11452,"freshPromptChars":785,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"0e56bb0c46a05efeeffa6feb09f54fe13a33c2c96ab4e7a7ec548a0d80a1a235","evidence-manifest.json":"0e601f9bde1bdb7f44f000704078343a1478046c9eb488f2c8a785cd1d4cebec","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"ab7d6af088306e1d43d4901a370dffa46b16dd4c17c2838a8e9e48a2b8e7a62f","snapshots/stock-harness.json":"2a6eabd1657c89568f890de12840747b544c7385c4363b3016879b07028961fa","snapshots/stock-harness-preflight.json":"539f106844e9421c284944c0b1526d74b0cc8e0905c5960be14cfa77ee17a023"}},
{"executionId":"stock-harness.legacy-opencode.local.ordered-comment-continuation","profile":"legacy-opencode","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":212184,"providerDurationMs":188508,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":177025,"outputTokens":7495,"cachedInputTokens":256512,"totalTokens":441032,"reportedCostUsd":0.011804638700000002,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":188508,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.011804638700000002,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.011804638700000002,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":10775,"reportedPromptChars":10736,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":4640,"reportedPromptChars":4640,"freshPromptChars":785,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"8ca5a9a446dda03100d8ee66c5b9e02fd19c335dfc8cc34f83662aa1181240b9","evidence-manifest.json":"496a392c39d3a75f0287c19a6fc25e7577dee6bfef0e3c7f16801f61574cfa99","snapshots/context-integrity.json":"da023ad1b2e22b1659573e1829cfb273a4d3e329d1b324506dbc73104901ecf7","snapshots/stock-harness-hire.json":"ec6d9274b80d291e8949c71a9bfb1f6fdc88cd78a2094dcad1c0a599ce9b78cd","snapshots/stock-harness.json":"680d3af8a66f473a95905920123f0760aff53eb4ce09cd01393a78b6292178cd","snapshots/stock-harness-preflight.json":"5b79fae54bd6036dab2ad8bf8e8bacb9087e702b363dbbe4af2c58115eea8ead"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.runner-acpx-claude.local.assigned-skill-explicit-invocation","profile":"runner-acpx-claude","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":50059,"providerDurationMs":28899,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":10,"outputTokens":1332,"cachedInputTokens":136292,"totalTokens":137634,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":28899,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"526b8881ac8dd2b51a788ca6ae535f7cd856e38066fadbd70737071c4f898abf","evidence-manifest.json":"40dc320ecab36c004d5e33bba8ebe1fb4d8444fd23d7ef0f055b188b474fffca","snapshots/context-integrity.json":"0b0b2b42376e1c5cf58b6ae1e9e0960b6b2bc82218e20589e04ca44639207ecb","snapshots/stock-harness-hire.json":"7f47075f11209d29c92091925bb61b2e2e22b06e1ce49b7dba69060242d62fc5","snapshots/stock-harness.json":"06398532782d5e4c9eb76ef55cdc87ed8d49c6f31dc06f09ce28715aa09f76fa","snapshots/stock-harness-preflight.json":"cb1980515ce11e2b55fcea5799277f511d0c6d01bf26b78e93be3ae17fa9b08d"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.runner-acpx-claude.local.continuity-restart","profile":"runner-acpx-claude","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":93902,"providerDurationMs":43211,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":6,"outputTokens":101,"cachedInputTokens":92222,"totalTokens":92329,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":43211,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"332ba05ee7cda6971777af3454e8589652f742f45cd6350da4396a420e8c549f","evidence-manifest.json":"eaff7dd508a80dffef9d638a39e13b21d2ef427086ae478c10520c2b336fe881","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"e8a854b3a0d886b01cb910e5b05baa5f1c791ef57b9a04168a0ce65cc60df7ce","snapshots/stock-harness.json":"1e53449e403bc079887cd183e2b556fbecd238a20130ef07c0c5decf499ff4ae","snapshots/stock-harness-preflight.json":"02283c369f44ab85a6400e0659da00e48bee84727fcebc35a8a691c4529c7cc0"}},
{"executionId":"stock-harness.runner-acpx-claude.local.ordered-comment-continuation","profile":"runner-acpx-claude","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":93050,"providerDurationMs":70393,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":18,"outputTokens":3794,"cachedInputTokens":332880,"totalTokens":336692,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":70393,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"3aae7a7ca62c39c74ac6d12e71b6b75223c76112ea018b8d08db1da52ea72e6a","evidence-manifest.json":"f08a3eee09a8fd4b279865dbe8f9ab2a66cc1a79addb216fe484561fcb9e2467","snapshots/context-integrity.json":"400df6ee5a9294998d4306f6eab6816e86155d8ecf0f39a6fc672381d47d423f","snapshots/stock-harness-hire.json":"604c41ac65944ef25cb7bbbf19941401139ed422342d7d6ebf9a07b79ecf7a75","snapshots/stock-harness.json":"f17d4f47657a0151dbfc585de64e338b7f79deae052f454794c2b934a7d398a2","snapshots/stock-harness-preflight.json":"fa502f0e3a3467dab6b04f810e6001ca635ce2d4bb5bae42a50e7a85f520a356"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.runner-codex.local.assigned-skill-explicit-invocation","profile":"runner-codex","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":48064,"providerDurationMs":26079,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":133429,"outputTokens":1102,"cachedInputTokens":110284,"totalTokens":244815,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":26079,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"690736dbf71179913631255763fd7b8aecd825236532154e54a3e1e53240d6b2","evidence-manifest.json":"75135aa0e7020617f54b5f2f593a5d48669ac7ecc864064d1eefc91af86d431c","snapshots/context-integrity.json":"46e65857a4d82f6ad900bc7e89166c35f2543faa368e48130f83dcea526ca65e","snapshots/stock-harness-hire.json":"2bf564606e974bb38baf50e363edd6d11aa5009fd429602a44fecb337bf9036c","snapshots/stock-harness.json":"2bf943a7cba0fc3886a04fa96f5646677f427e516232b35f6538384f4e6ed39f","snapshots/stock-harness-preflight.json":"635be5d6d578d7ad95b0e87697b939ed10f3cf2d3338caedc94350a1ed67ccc9"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.runner-codex.local.continuity-restart","profile":"runner-codex","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":89776,"providerDurationMs":43862,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":460440,"outputTokens":2336,"cachedInputTokens":378421,"totalTokens":841197,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":43862,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"b32c6c2c10d1eaabd0e2408db313bf9f30a7a60413733ef8ddfd8114ff6f2fdb","evidence-manifest.json":"34a60f60f021a731a98fab708ac31e344d8ecfa5a1d7dc72daa8c957028cbcea","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"855403ffdcaa2f1542a5ebbc0dbd7590c609c170afc657884eeb9a1e79a51310","snapshots/stock-harness.json":"d4997c371e7db1a9e1c13536279480035eee7873e92a3ab3300969d508339481","snapshots/stock-harness-preflight.json":"642913765f2dfb9bbae2f46401448e4c3b6d3f9134cba2a073178d530c395efc"}},
{"executionId":"stock-harness.runner-codex.local.ordered-comment-continuation","profile":"runner-codex","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":85296,"providerDurationMs":64569,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":359345,"outputTokens":2627,"cachedInputTokens":307813,"totalTokens":669785,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":64569,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"1a1a5f9b53dfc8cd2d96bafb8b34f6b1ac006514da70963ea3d8328f1c594309","evidence-manifest.json":"5220dd721a29958a2dbc6139a33bed2c94765801181c795abca6118722204f7b","snapshots/context-integrity.json":"100d34aa2c461b540164a00442100d5f6f438f706d14fb7316d80bf07d680cd8","snapshots/stock-harness-hire.json":"103296d40bc0b4500fa9b45ffb1e8192b770a1d4f87f8d0dd4f9e309384d407c","snapshots/stock-harness.json":"8d1eea8270fd2c29e284db9af321bebe85e635f6ec77d56674466d1480012ba8","snapshots/stock-harness-preflight.json":"be022954d2b5f5ef0bdae4cd892f06163f9d0863d0201aa92a5bab12f56a4dac"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}},
{"executionId":"stock-harness.runner-opencode.local.assigned-skill-explicit-invocation","profile":"runner-opencode","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":69432,"providerDurationMs":45144,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":22870,"outputTokens":1171,"cachedInputTokens":105472,"totalTokens":129513,"reportedCostUsd":0.0021534242,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":45144,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0021534242,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0021534242,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"69ac8808e939a796c40ac3665d56b152b43747fa91b7bbc84fd0a0f2fc68f5d9","evidence-manifest.json":"a0af2ef7ba4653d7ec825ac12de2a9e5f8b7fd7042bcfbe215e8af1b75ad38f8","snapshots/context-integrity.json":"e82ea3b31d4c6c73a95715c1d0e31338eb31a50f760bb6ea99b1d21032c6da32","snapshots/stock-harness-hire.json":"87ec7c1456344be2ed5cf68d4813920e8204fc483c6a34b1b62116bd721d9073","snapshots/stock-harness.json":"11cb4a18967ee59a7d5c202ee590a8dcfa7a9c145ab5ac87cde8da6a69f25497","snapshots/stock-harness-preflight.json":"c044805c4bb0db4c3cac0a4bbf54b937da7d1d376a4ac319f4d2baf3c636f51c"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}},
{"executionId":"stock-harness.runner-opencode.local.continuity-restart","profile":"runner-opencode","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":152945,"providerDurationMs":93715,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":26520,"outputTokens":106,"cachedInputTokens":45312,"totalTokens":71938,"reportedCostUsd":0.0005020232,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":93715,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0005020232,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0005020232,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"ae12b2eee65f45b8e2f530c037002e91db351134b6848da89e80bf7c15eb25a3","evidence-manifest.json":"dce1906b35fbbe6ac471ef922a2f0b1bae9da9438d56e78894e3651cba402aaa","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"a57da515e6454ee8b82b9bf633a8e07097217f30555606538e135c7979de71c0","snapshots/stock-harness.json":"fa0501dbd07065bbbbd9a1f4c6bd82ba27f0fa45238790d531db4587c000309d","snapshots/stock-harness-preflight.json":"a20db2a5478fb7bd568bf6fcfbe7702dbe2a7a0390d9b36ec0fe2a6d07954d1b"}},
{"executionId":"stock-harness.runner-opencode.local.ordered-comment-continuation","profile":"runner-opencode","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":95671,"providerDurationMs":72942,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":24355,"outputTokens":2690,"cachedInputTokens":160512,"totalTokens":187557,"reportedCostUsd":0.0043860217,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":72942,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0043860217,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0043860217,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"54f3029b581b82be87bd30b7670029bd5722915917e4bcab2e8ef0682bc3c0be","evidence-manifest.json":"b793fa707ccc2125b192816d59e922a0d43348c11eb516b8dff7e5d445b5545b","snapshots/context-integrity.json":"105ff5eb813ac3779e3ba04a8115ce1495c185011de4735782d3c4adba99787e","snapshots/stock-harness-hire.json":"4dde5055f48fff1b4cf466b4307739c967e2595e3c2e93d4e07230545741fa16","snapshots/stock-harness.json":"b0ce3eb0edce09ca4ef61af6f622cddd7c9bb1a3286c9ea7756dc228d00c10fa","snapshots/stock-harness-preflight.json":"78f2298b3bb4357b2a3ce6d75a922fc34d6449ef43378bb0dcc40476469daa01"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}
],
"totals": {
"baseline": {
"results": 24,
"passed": 15,
"failed": 9,
"missing": 0,
"cleanupPassed": 24,
"durationMs": 3250195,
"providerDurationMs": 2108034,
"llm": {
"runCount": 58,
"runsWithTokenUsage": 52,
"runsWithReportedCost": 43,
"inputTokens": 2778581,
"outputTokens": 68015,
"cachedInputTokens": 6908917,
"reportedCostUsd": 1.7494252618000001
}
},
"candidate": {
"results": 24,
"passed": 15,
"failed": 9,
"missing": 0,
"cleanupPassed": 24,
"durationMs": 2804913,
"providerDurationMs": 1540860,
"llm": {
"runCount": 47,
"runsWithTokenUsage": 45,
"runsWithReportedCost": 36,
"inputTokens": 2280185,
"outputTokens": 58933,
"cachedInputTokens": 4655025,
"reportedCostUsd": 1.7431513318
}
}
},
"matchedPassingTiming": {
"baseline": {
"cells": 13,
"medianCellDurationMs": 85550,
"medianProviderDurationMs": 57150
},
"candidate": {
"cells": 13,
"medianCellDurationMs": 69761,
"medianProviderDurationMs": 43862
}
},
"missingBaseline": [],
"cancellationRecovery": {
"sourceSha": "f02d8d0df327abb43b20c7e7beb86798239abbf5",
"originalRun": 37042856368,
"reason": "superseding pilot cancelled incomplete cells; completed behavior failures not rerun",
"executionIds": [
"stock-harness.legacy-acp-codex.local.ordered-comment-continuation",
"stock-harness.legacy-opencode.local.ordered-comment-continuation",
"stock-harness.legacy-opencode.local.continuity-restart",
"stock-harness.runner-acpx-claude.local.continuity-restart",
"stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation",
"stock-harness.runner-acpx-claude.local.assigned-skill-explicit-invocation",
"stock-harness.runner-acpx-claude.local.ordered-comment-continuation",
"stock-harness.runner-codex.local.continuity-restart",
"stock-harness.runner-codex.local.ordered-comment-continuation",
"stock-harness.runner-codex.local.assigned-skill-explicit-invocation",
"stock-harness.runner-opencode.local.assigned-skill-explicit-invocation",
"stock-harness.runner-opencode.local.continuity-restart",
"stock-harness.runner-opencode.local.ordered-comment-continuation"
]
},
"baselineInfrastructureRecovery": {
"schema": "paperclip.stock-harness.infra-recovery/v1",
"sourceSha": "12c5433c67dc2e62916b879349c7ba2b6e0431f0",
"originalRun": 37042864888,
"originalJob": 110958906481,
"reason": "AWS runner shutdown, no result artifact; no completed model failure rerun",
"executionIds": [
"stock-harness.legacy-opencode.local.ordered-comment-continuation"
]
},
"layoutRecovery": {
"resultEdits": 0,
"oracleEdits": 0,
"originalEvidenceRetained": true,
"manifests": [
{
"campaign": "candidate-original",
"sha256": "9135457cdc339d2fcaeea52cd06622ecaea6a74561df73d4d7a0baa0d8c0e7cc"
},
{
"campaign": "baseline-final",
"sha256": "85223428a0d473ba83e324621e80d711725ebe069c4f055665b0bf987d1f0c4c"
},
{
"campaign": "candidate-recovery",
"sha256": "7788f879d0cf1d140d452cfa78fe085b34bc23d298d7c31cac9d630b12b2572f"
},
{
"campaign": "baseline-recovery-final",
"sha256": "e39d62dd60f092e76f8776e844041f2458fca1989854d57ab13932a36c937ceb"
}
]
},
"limitations": [
"Single trial per cell; not general coding quality or statistical proof.",
"Skill task document storage wording is ambiguous; original graders unchanged.",
"Persisted credential guard failures are not equivalent to poor task behavior.",
"Report costs are incomplete; zero reported native cost has unknown billing type.",
"Local runtime and interrupted partial attempts are not metered.",
"Historical public prompt retrieval is clipped in three ACP cells; missing suffix does not prove provider omission.",
"Reconstructed reports use separate copies with directory-layout-only repair; original public reports retain infrastructure failures."
],
"fixtureMatchProof": {
"schema": "paperclip.stock-harness.fixture-match-proof/v1",
"cells": [
{
"executionId": "stock-harness.legacy-codex.local.ordered-comment-continuation",
"fixtureConfigSha256": "f727e68de06dec501c49efc41ab941bd6d63826db82e615d5e60f2d20133b488",
"model": "gpt-5.6-sol"
},
{
"executionId": "stock-harness.legacy-codex.local.assigned-skill-explicit-invocation",
"fixtureConfigSha256": "ca51e8b29aab8e6efa32f2433c0f3960e43f9fc2f5edcd0660e753810d71f30d",
"model": "gpt-5.6-sol"
},
{
"executionId": "stock-harness.legacy-codex.local.continuity-restart",
"fixtureConfigSha256": "8b1669039f3285f89a7f93c57e1c0dc12f071fbb2bbc8dcef6bf70bc14a4949a",
"model": "gpt-5.6-sol"
},
{
"executionId": "stock-harness.legacy-claude.local.ordered-comment-continuation",
"fixtureConfigSha256": "94db17927ceddb54fe7766a362adcf78783726c68e1e136e407572aef6fed319",
"model": "claude-sonnet-4-6"
},
{
"executionId": "stock-harness.legacy-claude.local.assigned-skill-explicit-invocation",
"fixtureConfigSha256": "6c21be9bb3fcec6274372fc586d88e42d82bc0d438e26a69ef614f8e6fcb8519",
"model": "claude-sonnet-4-6"
},
{
"executionId": "stock-harness.legacy-claude.local.continuity-restart",
"fixtureConfigSha256": "695071e5453c90288659f54f40fb307486f64d5fb8a401b767d2e9974faab15f",
"model": "claude-sonnet-4-6"
},
{
"executionId": "stock-harness.runner-codex.local.ordered-comment-continuation",
"fixtureConfigSha256": "9e19859684aad0246fd009061101eec59c41179e51e92a9467f506af22558126",
"model": "gpt-5.6-sol"
},
{
"executionId": "stock-harness.runner-codex.local.assigned-skill-explicit-invocation",
"fixtureConfigSha256": "8fc61211aefb339a310469695d716c3faea15c2a547c543d70c599857a7e071d",
"model": "gpt-5.6-sol"
},
{
"executionId": "stock-harness.runner-codex.local.continuity-restart",
"fixtureConfigSha256": "024e90354d760f90a40e2b3923468b72d36a9fabdb4fc58c10a7939a09e5cbc3",
"model": "gpt-5.6-sol"
},
{
"executionId": "stock-harness.runner-opencode.local.ordered-comment-continuation",
"fixtureConfigSha256": "9dd6fef8fea77b07a67681acc1d7866fac9b4f16e23aca69eda32a4656cc260b",
"model": "openrouter/deepseek/deepseek-v4-flash-0731"
},
{
"executionId": "stock-harness.runner-opencode.local.assigned-skill-explicit-invocation",
"fixtureConfigSha256": "891eb6dd154aa5926a3de2eae79780661b436618b6eac7733998750c6c341ba3",
"model": "openrouter/deepseek/deepseek-v4-flash-0731"
},
{
"executionId": "stock-harness.runner-opencode.local.continuity-restart",
"fixtureConfigSha256": "441c2fa8bdd26cf14ac2440311b99a71b7dd6202698cafc3d01ed7a91dc4b004",
"model": "openrouter/deepseek/deepseek-v4-flash-0731"
},
{
"executionId": "stock-harness.runner-acpx-claude.local.ordered-comment-continuation",
"fixtureConfigSha256": "77cc32714b23852d3ef67c687202d89f7513275500b3c522a0a1de3d745c9c4e",
"model": "claude-sonnet-5"
},
{
"executionId": "stock-harness.runner-acpx-claude.local.assigned-skill-explicit-invocation",
"fixtureConfigSha256": "c190a1e731ef450ea6ae0c24b5a8ae0f21db725b47350645cdbb9273415c5acd",
"model": "claude-sonnet-5"
},
{
"executionId": "stock-harness.runner-acpx-claude.local.continuity-restart",
"fixtureConfigSha256": "40876df6f88f34ba8232ccba8816bf7e1173be0042a90a174f79cddcca346c4c",
"model": "claude-sonnet-5"
},
{
"executionId": "stock-harness.legacy-acp-codex.local.ordered-comment-continuation",
"fixtureConfigSha256": "22500c39874cda62a2512ac3d8a43309ce5879f55f8515e31dd8ac632edb6c93",
"model": "gpt-5.6-sol"
},
{
"executionId": "stock-harness.legacy-acp-codex.local.assigned-skill-explicit-invocation",
"fixtureConfigSha256": "896d8602d2990bf7517a4392d66909c6649e0f00f79a5ea0bb8f0e370b7b597a",
"model": "gpt-5.6-sol"
},
{
"executionId": "stock-harness.legacy-acp-codex.local.continuity-restart",
"fixtureConfigSha256": "7b66845e231674cc6e78f1d86e863dc03005998c2674316af50db60ec53b7e4a",
"model": "gpt-5.6-sol"
},
{
"executionId": "stock-harness.legacy-acp-claude.local.ordered-comment-continuation",
"fixtureConfigSha256": "6c1b9d7debb306af405640eb381e1709ae7185e7925bc2bd3ef31acf98fbff8f",
"model": "claude-sonnet-4-6"
},
{
"executionId": "stock-harness.legacy-acp-claude.local.assigned-skill-explicit-invocation",
"fixtureConfigSha256": "53f308d6fbdc30adc60370279ab38679b5d87075899be4db6c0730593e1632c4",
"model": "claude-sonnet-4-6"
},
{
"executionId": "stock-harness.legacy-acp-claude.local.continuity-restart",
"fixtureConfigSha256": "80ec22c5080ef1a71fc97547c3e78f4bc5f54f121f46f982506f2f9f773b27d0",
"model": "claude-sonnet-4-6"
},
{
"executionId": "stock-harness.legacy-opencode.local.ordered-comment-continuation",
"fixtureConfigSha256": "e015ded0177d8d4c893acc658e25699b28937d97cb9d58da023ddff7ba7eac0b",
"model": "openrouter/deepseek/deepseek-v4-flash-0731"
},
{
"executionId": "stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation",
"fixtureConfigSha256": "dd673f469cab583d26bc2e4b484f18b27884ec14a59542177026389181e2da2c",
"model": "openrouter/deepseek/deepseek-v4-flash-0731"
},
{
"executionId": "stock-harness.legacy-opencode.local.continuity-restart",
"fixtureConfigSha256": "3dadc83dff00ef2919cd24c6328ea9507f971e7d294e9ece699ff2ee6fa98230",
"model": "openrouter/deepseek/deepseek-v4-flash-0731"
}
]
},
"workflowProvenance": [
{
"run": 37042856368,
"workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a",
"workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37042856368"
},
{
"run": 37042864888,
"workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a",
"workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37042864888"
},
{
"run": 37045368302,
"workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a",
"workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37045368302"
},
{
"run": 37044967981,
"workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a",
"workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37044967981"
},
{
"run": 37048838402,
"workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a",
"workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37048838402"
}
]
}
@@ -0,0 +1,110 @@
# Stock-harness live comparison — 2026-10-02
**TL;DR:** Two newly failing paired cases: classic Claude/OpenCode document delivery. Two newly passing cases: classic and native OpenCode ordered continuation. Seven unchanged failures and 13 unchanged passes. The extra ACP Claude storage symptom occurs beneath an existing credential-guard failure. All 48 results are available; no general performance equivalence is established.
The reduced instructions are **not yet qualified for merge**. Classic Claude and classic OpenCode pass the historical skill-output case but fail with the reduced instructions: they finish without saving a Paperclip task document. Native Codex and native Claude pass all three journeys in both variants. Single trials and ambiguous fixture storage wording limit causal attribution; the failed oracle remains unchanged.
The reductions and evaluation setup are in draft [PR #14948](https://github.com/paperclipai/paperclip/pull/14948). This report and its [safe evidence projection](2026-10-02-stock-harness-live-comparison.json) record the measured revisions rather than claiming the final documentation head was run through the full matrix.
## Revisions and matched inputs
- Candidate: `f02d8d0df327abb43b20c7e7beb86798239abbf5`, eight-word hire manual plus reduced shared startup/resume instructions.
- Historical baseline: `12c5433c67dc2e62916b879349c7ba2b6e0431f0`, the same evaluation harness with only the prior default manual and shared prompt source restored from `e00d10d5d5594f6e2d1e8cf4e0ec81c074b0bd88`. Its manifest records ten identical behavioral/fixture source hashes.
- Native Codex fix [#14920](https://github.com/paperclipai/paperclip/pull/14920) is held constant. This comparison does not measure its before/after performance.
- Same eight profiles, three journeys, provider models, effort configuration, skills, tools, permissions, credentials, local runtime, and independent graders. Effort inherits provider defaults where the existing profile does not set it.
- Each variant selects 24 cells and initially budgets 48 turns, with existing bounded retries and 1,000-cent company/agent hard stops. Recorded run-ledger counts can differ after early failures, retries and cleanup.
## Per-profile and case outcomes
These are recorded overall qualifications, including security and receipt failures. A credential-guard failure does not imply every behavioral matcher failed.
| Profile | Model | Journey | Historical | Reduced |
| --- | --- | --- | --- | --- |
| `legacy-codex` | `gpt-5.6-sol` | `assigned-skill-explicit-invocation` | Pass | Pass |
| `legacy-codex` | `gpt-5.6-sol` | `ordered-comment-continuation` | Pass | Pass |
| `legacy-codex` | `gpt-5.6-sol` | `continuity-restart` | Pass | Pass |
| `legacy-claude` | `claude-sonnet-4-6` | `assigned-skill-explicit-invocation` | Pass | Fail: no Paperclip document |
| `legacy-claude` | `claude-sonnet-4-6` | `ordered-comment-continuation` | Pass | Pass |
| `legacy-claude` | `claude-sonnet-4-6` | `continuity-restart` | Fail: chat memory | Fail: chat memory |
| `legacy-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `assigned-skill-explicit-invocation` | Pass | Fail: no Paperclip document |
| `legacy-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `ordered-comment-continuation` | Fail: deadline/run-start timeout | Pass |
| `legacy-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `continuity-restart` | Pass | Pass |
| `legacy-acp-codex` | `gpt-5.6-sol` | `assigned-skill-explicit-invocation` | Fail: credential guard | Fail: credential guard |
| `legacy-acp-codex` | `gpt-5.6-sol` | `ordered-comment-continuation` | Fail: credential guard, receipt incomplete | Fail: credential guard |
| `legacy-acp-codex` | `gpt-5.6-sol` | `continuity-restart` | Fail: credential guard, receipt incomplete | Fail: credential guard, receipt incomplete |
| `legacy-acp-claude` | `claude-sonnet-4-6` | `assigned-skill-explicit-invocation` | Fail: credential guard | Fail: credential guard, no Paperclip document |
| `legacy-acp-claude` | `claude-sonnet-4-6` | `ordered-comment-continuation` | Fail: credential guard, receipt incomplete | Fail: credential guard |
| `legacy-acp-claude` | `claude-sonnet-4-6` | `continuity-restart` | Fail: credential guard, receipt incomplete | Fail: credential guard |
| `runner-codex` | `gpt-5.6-sol` | `assigned-skill-explicit-invocation` | Pass | Pass |
| `runner-codex` | `gpt-5.6-sol` | `ordered-comment-continuation` | Pass | Pass |
| `runner-codex` | `gpt-5.6-sol` | `continuity-restart` | Pass | Pass |
| `runner-acpx-claude` | `claude-sonnet-5` | `assigned-skill-explicit-invocation` | Pass | Pass |
| `runner-acpx-claude` | `claude-sonnet-5` | `ordered-comment-continuation` | Pass | Pass |
| `runner-acpx-claude` | `claude-sonnet-5` | `continuity-restart` | Pass | Pass |
| `runner-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `assigned-skill-explicit-invocation` | Pass | Pass |
| `runner-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `ordered-comment-continuation` | Fail: deadline/run-start timeout | Pass |
| `runner-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `continuity-restart` | Pass | Pass |
## Findings and next step
The classic Claude and classic OpenCode skill runs read the pinned skill marker and complete the task, but save zero Paperclip documents; both historical runs save exactly one. The legacy ACP Claude skill run has the same document difference under its credential-guard failure. The earlier reduced-instruction Claude pilot also failed this document oracle. The pinned skill asks for a “task document” without explicitly naming Paperclip storage, so this may reveal both reliance on the former manual and a fixture ambiguity. It remains a failed delivery result, not a regraded pass.
Keep the agreed tiny manual. Next, verify that legacy agents discover and follow the existing **Paperclip skill/API path for durable task documents and artifacts**. The skill already prohibits file-only handoff, lists issue-document routes and gives a plan-document example; audit that delivery before adding a concise generic-document recipe if needed. Run the same preserved case plus a clearly storage-specific case on both variants. Legacy agents do not receive native `paperclip_finish`/`paperclip_block` tools. Native tool descriptions are a separate item 2.3 change; guidance must follow the actual runtime capability. No production or skill instructions were changed to make these runs pass.
Classic Claude chat fails the memory assertion in both variants. Six legacy ACP cells per variant fail the persisted-credential guard: the scanner finds a scoped provider credential in persisted ACP session state. Raw session files and credential values are not published. These are existing qualification failures requiring their own investigation; they cannot be waived or interpreted as instruction-reduction quality evidence.
Historical native OpenCode ordered comments hit the case deadline without the required continuation outcome. The initial historical classic OpenCode ordered cell lost its AWS runner before result upload; its one unchanged-source recovery also hits the 12-minute deadline. Both OpenCode ordered cases pass in the candidate. Completed failures are never rerun to select a better result. Those candidate passes against historical timeouts are observed differences, not proof of a general improvement.
Historical public prompt retrieval is clipped in three legacy ACP cells (four invocations), including identity/connection suffix text. Those structural receipt failures cannot establish that the provider omitted instructions. The candidate ACP Codex chat run also lacks a complete per-run invocation receipt. Instruction presence/absence claims are bounded by the retained public receipts.
## Timing, usage and cost
| Variant | Retained cells | Pass / fail / missing | Cell duration sum | Provider duration sum | Recorded runs with tokens / with reported cost / total | Reported LLM subtotal |
| --- | ---: | --- | ---: | ---: | --- | ---: |
| baseline | 24 | 15 / 9 / 0 | 3250.195s | 2108.034s | 52 / 43 / 58 | $1.749425 |
| candidate | 24 | 15 / 9 / 0 | 2804.913s | 1540.860s | 45 / 36 / 47 | $1.743151 |
For the 13 cells that pass both variants, median cell time is 85.550s historical versus 69.761s reduced; median provider time is 57.150s versus 43.862s. This excludes failures and is descriptive, not a reliable speedup estimate.
Provider billing is incomplete. Legacy Codex/ACP usage is unpriced in several runs; native Codex/Claude often report zero with billing type `unknown`. Reported subtotals do not prove zero native spending or invoice totals. Local runner costs and interrupted partial runs are not metered. Diagnostic/pilot costs and repeated interrupted work are outside the matched-cell subtotal and remain retained. No aggregate cost-saving claim is justified.
## Campaign and evidence history
| Campaign | Source | Outcome / evidence |
| --- | --- | --- |
| [Initial diagnostic Codex](https://github.com/paperclipai/paperclip/actions/runs/37034213743) | `a63437069` | Skill passes; predates mandatory admission. |
| [Cold full attempt](https://github.com/paperclipai/paperclip/actions/runs/37037105491) | `36e987246` | Missing SDK build; cancelled before provider admission. |
| [SDK pilot](https://github.com/paperclipai/paperclip/actions/runs/37039240025) | `4163dbfd0` | Missing daemon; stops before providers. |
| [Daemon pilot](https://github.com/paperclipai/paperclip/actions/runs/37040493183) | `ac6ddefb5` | Missing fake Codex fixture; stops before providers. |
| [Complete-build Claude pilot](https://github.com/paperclipai/paperclip/actions/runs/37041741124) | `f02d8d0df` | 537 prerequisites pass; skill document oracle fails; reported $0.175308. |
| [Initial candidate matrix](https://github.com/paperclipai/paperclip/actions/runs/37042856368) | `f02d8d0df` | 4 pass, 7 fail, 13 cancelled when a same-target-branch pilot superseded it. Original attempts retained. |
| [Candidate cancelled-cell recovery](https://github.com/paperclipai/paperclip/actions/runs/37045368302) | `f02d8d0df` | Only the 13 cancelled cells: 11 pass, 2 fail. Combined candidate has 24 results, 15 pass / 9 fail. |
| [Historical matrix](https://github.com/paperclipai/paperclip/actions/runs/37042864888) | `12c5433c6` | 15 pass, 8 recorded failures, one AWS-runner shutdown without result upload. |
| [Historical interrupted-cell recovery](https://github.com/paperclipai/paperclip/actions/runs/37048838402) | `12c5433c6` | Only legacy OpenCode ordered comments: 720.616s deadline failure; evidence and cleanup valid. Historical cohort now has 24 results, 15 pass / 9 fail. |
| [Current packaging pilot](https://github.com/paperclipai/paperclip/actions/runs/37044967981) | `1eb5ba420` | Attempt 1 stopped before cell/provider execution; attempt 2 passes, evidence valid and cleanup pass. |
The same-target-branch workflow concurrency rule caused the candidate interruption; that orchestration mistake was acknowledged and corrected with a separate unchanged-source recovery branch. No completed model failure was retried. The historical missing cell is recovered once for a runner shutdown. Partial/failed attempts remain part of the history and unknown-spend accounting.
The measured matrix packages contain a prerequisite folder beside the exact campaign root. The trusted selector correctly rejects this layout, so their original published reports synthesize missing-result infrastructure failures. A separate local reconstruction nests that prerequisite folder under the campaign root and verifies byte identity for every copied file; result, scorer, usage, screenshot and prerequisite content are unchanged. The unchanged trusted selector then selects 11/24 original candidate, 13/13 recovery 23/24 original historical packages and 1/1 interrupted-cell recovery; recovered normalization validates every retained result’s evidence. Original missing packages stay missing in their original campaign; the separate recovery result has its own provenance. This is declared directory-layout recovery, not a canonical successful publication or a change to the oracle. The safe JSON records result/receipt/scoring hashes and recovery-manifest hashes.
The fix at `1eb5ba420` writes prerequisites beneath the exact campaign root. Its single selected Codex skill cell passes through the protected report and publication pipeline: [interactive pilot report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37044967981-2/index.html), [GitHub evidence](https://github.com/paperclipai/paperclip/actions/runs/37044967981#artifacts). It has `evidenceValid=true`, cleanup pass, 27.149s provider / 49.523s cell, and a reported zero subtotal with unknown billing type. It does not replace the full matched cohort or qualify untested providers.
Original historical/recovery dashboards retain the publication failure: [historical](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37042864888-1/index.html), [candidate recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37045368302-1/index.html). Inspect their access-controlled raw artifacts for measured cell evidence; do not read synthetic missing-result rows as task performance.
## Verification and remaining qualification
At `1eb5ba420`, all 557 credential-free prerequisites pass (556 TypeScript plus one Rust), along with 895 E2E support tests, E2E typecheck and 24-cell discovery. The cold builder prepares SDK/shared outputs and both daemon/fake protocol binaries, records hashes, and fails closed before credentials. The 313 filtered native tests are not counted as passing coverage. This head has 52 passing CI gates and two skipped gates; Greptile reviewed that exact head at 5/5 with zero open threads. Later documentation heads require fresh checks/review; measured sources remain explicit above.
Repository typecheck and build pass. The complete local Vitest run has 14,870 passed, 83 skipped and three unrelated timing failures; the three affected files pass unchanged narrow reruns (25 auth, 16 reviewed-chat binding, four webhook tests). Original failures are retained rather than calling the first full run green.
PR #14948 stays draft while document delivery is unresolved. Unrepresented harnesses, legacy ACP credential safety, public-receipt clipping, the independent ACP base-replacement issue, existing Claude chat memory behavior and saved-session migration are not qualified away by these results. This small matrix measures skill/context/chat behavior, not broad coding quality or statistical equivalence.
## Retained delivery diagnosis and approved repair
Both failed classic agents loaded the operational Paperclip skill. This is not an observed skill-discovery failure. In the reduced Claude run, the agent wrote a workspace Markdown file before loading Paperclip and then completed the issue without a document API write. The reduced OpenCode run loaded Paperclip, accessed the assigned skill through the public skills API, and wrote a workspace task-document file; it made no document, attachment, or work-product delivery write. Both corresponding historical runs saved a Paperclip issue document.
Claude received the full operational skill, including its existing no-local-only delivery guidance. OpenCode's skill result was explicitly truncated in both variants: the early artifact rule survived, but the later generic document endpoint and plan-only write example were omitted. That truncation is existing behavior, not a newly measured regression. Removing always-on delivery reminders may interact with it, but the combined manual/shared reduction and ambiguous assigned-skill wording prevent a causal attribution to one layer.
Dotta approved a small early API-runtime skill recipe and a generic issue-document reference. The tiny manual remains eight words. The recipe checks the successful write's returned saved revision/content and links the returned key; it does not require another GET after a clear valid receipt. It preserves explicit destinations, downloadable-file delivery, and native document-tool boundaries.
The focused repair comparison will hold the tiny manual/shared prompts fixed and vary only `skills/paperclip/SKILL.md` plus the new `references/issue-documents.md`. Classic Claude and OpenCode each run the preserved original case and an added explicitly Paperclip-storage case: four cells per variant, eight expected provider turns in total. This measures the skill repair separately from the original reduction. The explicit case independently checks public saved body/revision and an agent comment linking the exact same-app document; local-only, missing-revision, wrong-document, local-path, and other-origin outcomes cannot pass. Current state: fixture preparation and source freeze, no repair provider run yet.
@@ -1,7 +1,17 @@
# Stock harness, with Paperclip: working checklist
Created: 2026-10-02. Status: item 1 implemented for native Codex app-server;
PR validation in progress. Other implementation items remain open.
Created: 2026-10-02. Status: item 1 merged for native Codex app-server in
[PR #14920](https://github.com/paperclipai/paperclip/pull/14920). Item 2's default
hire manual is reduced to identity only, and common legacy startup/resume
instructions are in [PR #14948](https://github.com/paperclipai/paperclip/pull/14948).
GitHub live qualification retains earlier document-delivery failures; the latest
focused OpenCode comparison passes both cases in both variants but still exposes
deficient original-case handoff links. PR #14948 remains draft. The original
[live report](2026-10-02-stock-harness-live-comparison.md), completed
[skill repair](2026-10-02-legacy-document-skill-repair.md) and latest
[stock-selection/link comparison](2026-10-02-opencode-skill-routing-link-qualification.md)
retain their separate sources and limitations. Additional carriers and fixed
native instructions remain open; hiring templates are being reduced separately.
Goal: keep the agent's stock harness behavior and add only what it needs to work
with Paperclip. Apply this across legacy adapters, the new Runner, and their
@@ -12,7 +22,11 @@ record the proposed behavior and Dotta's direction, make a bounded change, and
verify the affected paths before moving on. New findings get stable IDs in the
ledger below so they do not disappear into conversation history.
**Current item: 1 — validate and review the native Codex instruction fix.**
For every change, record its executable test/eval coverage before continuing.
Distinguish coverage setup, deterministic results, and measured live results;
configured cells alone do not qualify behavior.
**Current item: 2 — reduce the default operating manual and shared prompt layers.**
## Agreed direction and boundaries
@@ -43,7 +57,7 @@ ledger below so they do not disappear into conversation history.
- [x] Remove default replacement of stock base instructions across those paths.
- [x] Verify the actual app-server request and retained session instructions,
including resume; checking only a prompt builder is insufficient.
- [ ] Verify Paperclip task context, tools, auth, assigned skills, and completion
- [x] Verify Paperclip task context, tools, auth, assigned skills, and completion
still work. Record applicable regressions and eval results.
Starting points: [Codex backend](../../packages/paperclip-runner/src/backends/codex-native-backend.ts),
@@ -61,14 +75,20 @@ and need a provider session reset. Do not reset active sessions automatically.
The separate Codex-through-ACP dependency patch remains a coverage follow-up
under item 6; this change covers the native app-server path.
Verification: 412 focused TypeScript/Rust tests, repository typecheck/build,
55 passing PR checks, and Greptile 5/5. The original full local test command
was stopped after setup failures; affected suites passed on rerun. No paid
live campaign was run, and no task-quality improvement is claimed.
## 2. Reduce the default operating manual and shared prompt layers
- [ ] Inventory what an agent actually receives: hire instructions, shared
prompt template, wake context, runtime prompt, bootstrap, and loaded skills.
Separate always-present text from content loaded on demand.
- [ ] Agree on the tiny common contract and the coordination details that each
runtime still requires. Legacy API coordination and native semantic tools
need appropriate instructions for their respective interfaces.
- [x] Agree on the tiny default: identity only, with no skill or native-tool
pointers. The harness supplies coordination instructions.
- [x] Reduce the generic default hire `AGENTS.md` to the agreed sentence and
update the existing creation test to expect the minimal bundle.
- [ ] Remove repeated workflow rules and stock coding/style/autonomy guidance.
- [ ] Move detailed planning, hiring, artifacts, and exceptional procedures to
discoverable references or tools where feasible.
@@ -80,7 +100,130 @@ Starting points: [default hire instructions](../../server/src/onboarding-assets/
[native runtime contract](../../packages/paperclip-runner/src/contracts/runtime-context.ts),
[Paperclip operational skill](../../skills/paperclip/SKILL.md).
Decision: exact retained paragraph and on-demand boundaries pending.
Decision: Dotta confirmed the harness already handles skill and tool delivery;
the default hire manual is now only "You are an agent in a Paperclip company."
This replaces 602 words with eight. The shared loader supplies this default
across instruction-bundle-capable adapters for non-CEO hires without explicit
instructions. Existing saved bundles retain their content; CEO, first-agent,
and role/team templates remain separate work under item 3. The common
prompt/wake reduction is complete locally below; additional carriers, native
instructions, and skill/reference corrections remain open.
Verification: the existing agent-skills route suite passed all 54 tests,
including default creation, custom bundles, and CEO/first-agent paths. The
onboarding asset suite passed all 10 tests. This initially had only deterministic
coverage. Dotta subsequently requested behavioral eval coverage for every change;
the dedicated production-default-hire suite below closes the custom QA manual
gap. Full repository typecheck/build/test were not rerun for this narrow asset
and existing-test update.
### Shared prompt follow-ups
- [x] **2.1 Reduce common legacy startup/resume instructions.** Keep
identity and connection guidance in the shared task/chat defaults; remove the
generic resumed-wake execution contract. Preserve current task/event data,
specialized wake contracts, custom templates, skills, and auth.
- [ ] **2.2 Review additional legacy carriers.** Reduce Hermes local/gateway
wrappers, review Pi system delivery, and check OpenClaw fresh-wake framing.
Preserve transport facts and user configuration. Dotta deferred this item on
2026-10-02 for a later revisit; it remains open.
- [ ] **2.3 Reduce native Runner instructions and constraints.** Improve
discoverable tool documentation first, then shorten fixed guidance and
consolidate completion rules. Verify prompt revisions, digests, and session
compatibility across native Codex, ACPX, and OpenCode. Dotta approved the
native tool-documentation slice on 2026-10-02 in a separate worktree. Native
`paperclip_finish`/`paperclip_block` guidance must not leak into legacy
completion paths, which use the operational skill and API.
### Follow-up 2.1 implementation and verification
The common task and conversation defaults now share the same identity sentence
and unchanged connection guidance: 113 words each, down from 661 and 205.
Connection guidance accounts for 108 of those words. The ordinary resumed-wake
execution contract is removed (172 words), including the old opt-in used by
OpenClaw on fresh turns. `includeExecutionContract` remains accepted as a
deprecated no-op for adapter/plugin source compatibility.
Current task facts, ordered comments, work modes, approval/review state,
checkout, holds/blockers, recovery/watchdog roles, and external-chat contracts
remain in their existing owners. Skill delivery, runtime authentication,
custom `promptTemplate`, and `bootstrapPromptTemplate` mechanics are unchanged.
Hermes's own wrappers and Pi's system carrier remain follow-up 2.2; native
fixed instructions and full-turn constraints remain follow-up 2.3.
The later read-only 2.2 audit found another OpenClaw-owned HTTP identity,
checkout/status, delegation, plan-approval and task-discovery wrapper in
`buildWakeText`. Its short conversation branch is selected when optional task
Markdown is present, including on ordinary task dispatch. Existing dispatch
coverage does not assert the absence of conversation waiting language. This is
a framing concern to verify, not a measured failure. Pi already uses additive
`--append-system-prompt` delivery and suppresses the duplicate user-prompt copy;
its next step is a delivery audit with minimal changes. These findings are
deferred with 2.2 and are not implemented in PR #14948.
The new defaults apply when prompts are assembled after deployment. Existing
provider sessions can retain earlier startup instructions in their history
until reset; no active session or saved custom template is rewritten here.
Verification: 563 distinct focused tests passed across 18 suites:
- Shared prompt selection/rendering, operational-skill selection, and retained
specialized wake contracts (130 tests).
- ACPX, Pi, OpenCode, Cursor Cloud, and Codex local execution, including
task/chat fresh/resumed/reset/fallback delivery, custom prompts, skill
mounting, and authentication (259 tests).
- Hermes local prompt rendering and gateway execution (41 tests).
- Claude, Gemini, Cursor local, Grok, Kimi, OpenClaw, and Claude/Codex ACP
fallback execution (113 tests).
- Native execution input, preserving shared wake compatibility (20 tests).
`@paperclipai/adapter-utils` typecheck and build passed; `git diff --check`
passed. Full repository typecheck/build/test and live behavioral evals were not
run for this local bounded change. These checks prove prompt and runtime
mechanics, not an improvement in coding-task quality.
### Remaining shared layers inspected (2026-10-02)
The table records the initial inspection before follow-up 2.1. Word counts
cover fixed source text, not complete assembled prompts or token counts. They
exclude task data, assigned agent instructions, and skills. This inspection
establishes instruction mechanics, not a performance result.
| Layer | Delivery at initial inspection | Disposition |
| --- | --- | --- |
| Shared legacy task/chat templates | `DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE` was 661 words and the conversation template was 205, including connection guidance. Used by Claude, Codex, Cursor local/cloud, Gemini, Grok, Kimi, OpenCode, Pi, and Hermes local. | Completed in 2.1: both defaults now 113 words. Repeated procedures removed, connection guidance retained, no replacement skill pointer. |
| Generic resumed wake contract | `renderPaperclipWakePrompt` added a 172-word execution contract on ordinary resumed turns. OpenClaw requested it on fresh turns too because it has no shared template. | Completed in 2.1: generic contract removed, including legacy opt-ins. Task/event data and conditional runtime contracts retained. |
| Additional legacy carriers | Hermes local adds a 171-word identity/API/curl wrapper before the shared task template; Hermes gateway supplies its own four-rule execution contract. Pi carries shared defaults in its system extension and suppresses the wake copy when that extension owns policy. | Separate follow-up 2.2; changing the shared constants alone does not remove every wrapper. Preserve transport facts and custom configuration. |
| Native fixed prompt | The 262-word `paperclip-execution.v5` prompt carries hiring, delegation/dependencies, connection setup, and finalization guidance across native Codex, ACPX, and OpenCode backends. | Keep one short native completion/dependency contract. Relocate uncommon procedures to the relevant tool documentation, improving that documentation where needed before removing instructions. Any change needs prompt revision/digest/session-compatibility verification. |
| Native full-turn constraints | `nativeTaskConstraints` adds assigned-skill limits, current agent-file paths, document/file delivery, answered-question handling, and accepted completion/final-response sequencing. Codex/OpenCode add another completion reminder. Prepared constraints reach the model envelope through `native-session-runtime`; compact continuations avoid replaying prior task constraints. | Consolidate duplicate completion instructions. Keep current paths, answer scope, and the required result protocol. Move detailed file/document procedure to discoverable tool descriptions where adequately covered. |
| Task/event-specific framing | Server task Markdown and wake rendering supply work mode, ordered comments, approved revisions, checkout/holds/blockers, recovery/review/watchdog roles, external-chat bookkeeping/delivery, and conversation handoff. | Retain the facts and mode boundaries in the first reduction. Use one authoritative owner for shared mode directives; review specialized wording separately. |
Custom `promptTemplate` and fresh-only `bootstrapPromptTemplate` configuration
are distinct from shipped defaults and should not be silently rewritten.
The built-in process adapter invokes a configured command; HTTP passes context
as JSON. They do not use the shared text template. External adapter plugins and
hosted configuration variants still require the broader item 6 audit.
An accepted-plan discrepancy needs resolution: task Markdown permits cohesive
implementation on the current issue, while the wake renderer's planning-mode
accepted-confirmation branch requires children and prohibits implementation on
the source issue. Existing tests assert both versions. Production reachability
after the server changes work mode has not yet been established; do not claim
this proves a live plan-continuation failure.
Dotta selected follow-up 2.1 first: reduce the common legacy startup/resume
copies, preserving conditional context. Additional carriers and native fixed
instructions remain distinct follow-ups 2.2 and 2.3. No shared runtime text was
changed by the initial inspection.
Coverage identified during inspection includes shared prompt selection/rendering,
adapter execution tests, native backend/runtime context tests, and server
task-context/native-input tests. Follow-up 2.1 extended the shared and adapter
checks; its results are above. Follow-ups 2.2 and 2.3 need their corresponding
carrier/protocol assertions. Product E2E context-integrity, completion-updates,
and direct blocker guidance are candidate lifecycle checks; their current scope
does not establish coding-task quality. The initial inspection itself ran no
tests or live evals.
## 3. Review every hiring and role template
@@ -107,7 +250,61 @@ Starting points: [onboarding assets](../../server/src/onboarding-assets/),
[hiring skill and references](../../skills/paperclip-create-agent/),
[teams catalog](../../packages/teams-catalog/catalog/).
Decision: per-template disposition and migration policy pending.
Dotta approved and merged [PR #14985](https://github.com/paperclipai/paperclip/pull/14985)
at `862a5758ba0e88a33232c1f1fa645e85c38a3113`, after corrected source
`57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25` passed full CI and fresh Greptile 5/5.
The hiring skill/references, CEO/CoS assets and team catalog are now on master.
Their live qualification remains separate from unit/CI verification.
[Prompt diffs and preserved scope](https://github.com/paperclipai/paperclip/blob/57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25/doc/plans/2026-10-02-hiring-template-prompt-diff.md).
The read-only audit found four onboarding CEO files (1,897 words), a 164-word
CoS manual, four role examples (coder 652, QA 619, UX 1,325, security 1,724 words),
and eight team-catalog manuals. Hiring step 6, the 60–150-line role guide and
review checklist would regenerate the removed operating policies.
- [x] **3.1** Reduce role drafting rules, coder/adapted hire examples and matching
catalog coder, preserving metadata, auth, permissions and skill selections.
- [x] **3.2** Reduce the CEO bundle/copy, including forced delegation, hiring,
memory and repeated API recipes.
- [x] **3.3** Review remaining roles/catalog and CoS copy.
- [ ] **3.4** Review specialized built-in/plugin contracts separately; retain
product-required behavior rather than assuming every rule is redundant.
Custom and existing bundles and governance remain deliberate rollout boundaries.
Specialized Summarizer, Reflection Coach and Wiki Maintainer prompts remain
unchanged for separate product-contract review; optional Content Lead was already
tiny. The generic non-CEO fallback belongs to PR #14948. PR #14985 does not reduce
every shipped specialized agent.
Its explicit `hiring-templates` suite selects native local Codex and ACPX Claude
`hire-coder-template-reuse` cells: a real default CEO, explicitly requested
production hiring skill/coder reference, independently scored saved JSON
artifacts and session reuse. Outcome and source/read coverage are separate;
missing or redacted receipts do not prove no regression. The existing first-task
suite covers actual wizard CoS snapshot selection. The matched union holds the
reduced manual/shared/operational skill and native completion guidance constant,
restoring only 21 historical hiring production/derived sources in baseline.
Frozen candidate `9f5404ad3aacbe76777952759414d34fd381e674` and baseline
`296a4df85e8bcc97a160fc78c291b17adb828196` passed 705 exact-source credential-free
prerequisites each (704 TypeScript + 1 Rust; zero providers), with 8,242 identical
other tracked files and matching fixture/model/auth/permissions configuration.
Only four candidate-specific single-file CEO selection assertions are filtered
symmetrically; each variant's bundle/source hashes and canonical catalog are
independently admitted. Two cells per variant, five expected turns each:
four cells / 20 turns, 15 minutes each, no automatic retries or broader selection.
The [candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208)
and [baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463)
protected campaigns are dispatched. No graded live pairs yet; pending evidence
cannot establish non-regression.
[Inspect the partial report](https://github.com/paperclipai/paperclip/blob/e3720f369df070036526dd295a7a68e49488c51a/doc/plans/2026-10-02-hiring-template-live-comparison.md).
Report-only updates preserve the measured refs and merged hiring PR.
The default-manual PR is replayed on the merged hiring base, preserving custom
CEO bundle checks, the minimal generic boundary and both eval suites. Its prior
browser failure remains retained: a deterministic process fixture replayed its
last plan command on an asynchronous `chat_task_completed` wake and wrote a
second, text-identical source plan revision. It was not paid/provider execution;
fresh rebased CI remains required before readiness.
## 4. Fix repository context while retaining Paperclip configuration
@@ -154,7 +351,7 @@ does not apply.
| Legacy Codex / Claude | Local adapters, managed auth and isolated configuration | Initial inspection only |
| Other legacy local adapters | ACPX, OpenCode, Pi, Cursor, Gemini, Grok, Kimi, Hermes | Pending |
| Other adapter transports | Cursor Cloud, Hermes/OpenClaw gateways, process, HTTP, external adapter plugins | Pending |
| Runner Codex | App-server driver, runnerd bridge, Rust provider, direct/fallback paths | Additive instruction fix implemented; PR validation pending |
| Runner Codex | App-server driver, runnerd bridge, Rust provider, direct/fallback paths | Additive instruction fix merged in PR #14920 |
| Runner ACPX | Enabled profiles, especially Claude/Grok; declared or pending profiles tracked separately | Claude initial inspection; remaining audit pending |
| Runner OpenCode | Native provider and configuration paths | Pending |
| Hosted/remote providers | Claude Managed and AWS AgentCore; identify their own baseline rather than assuming CLI semantics | Pending |
@@ -172,15 +369,18 @@ does not apply.
## Verification and evals — apply to each item
- [ ] Map existing coverage before adding cases; use [doc/evals.md](../evals.md).
- [x] Map existing coverage for all implemented changes before adding cases;
see the SH-1–SH-3 map below and [doc/evals.md](../evals.md).
Keep Runner protocol evals and Product E2E evals distinct.
- [ ] Start with narrow deterministic checks for instruction layering, effective
- [x] Set up narrow deterministic checks for the implemented instruction layering,
hire, and shared-prompt changes. Future changes still need their own map for effective
configuration, skill/auth delivery, and session behavior where appropriate.
- [ ] Use Runner evals for provider/session/tool protocol changes; use Product
E2E for real hiring, repository context, task lifecycle, and artifact delivery.
- [ ] Review existing context-integrity, hiring, completion-updates, and blocker
suites for reusable coverage; record gaps rather than claiming coverage from
a similarly named case.
- [x] Reuse existing protocol tests for native Codex request layering and add Product
E2E for production-default hiring, skill delivery, task lifecycle, and chat
continuity. Repository context and additional artifact cases remain future work.
- [x] Review existing suites for reusable coverage. Their custom QA manuals did
not exercise the tiny default hire; the new suite deliberately omits those
bundles and reuses independently graded skill, continuation, and chat journeys.
- [ ] Compare task quality as well as Paperclip protocol compliance when claiming
that fewer instructions improve agent performance. Keep model, effort,
permissions, tools, and fixture comparable.
@@ -189,19 +389,71 @@ does not apply.
- [ ] Run the relevant checks for each change and the repository's required full
verification before a PR-ready handoff. Record unrun checks and their reasons.
### Coverage for changes implemented so far
The [stock-harness runbook](../../tests/runner-e2e/STOCK-HARNESS.md) defines a
credential-free prerequisite plus 24 explicit local Product E2E cells across
eight legacy/native profiles. Each profile runs assigned-skill invocation,
ordered comment continuation, and chat continuity across restart. Public receipts
verify the identity-only managed bundle before provider execution; final receipts
check actual legacy prompts and both budget hard stops. The existing lifecycle
oracles still own task/chat success. Models, auth, skills, permissions, and
secret-reference plumbing are inherited from existing profiles.
| Coverage ID | Implemented change | Executable coverage | Live status |
| --- | --- | --- | --- |
| SH-1 | Native Codex additive developer instructions, including start/resume/recovery | TypeScript driver, runnerd transport, Runner Lab/live-session, and Rust provider tests in `pnpm test:e2e:runner:stock-harness`; native Codex cells exercise real hires. | Native Codex passes all three matched journeys in both variants; #14920 held constant, so task success is not before/after vendor-base proof. |
| SH-2 | Eight-word default hire `AGENTS.md` | Public creation/onboarding tests; exact independent public bundle oracle before and after provider execution for every stock-harness cell. | Candidate public tiny-bundle and budget receipts pass in all 24 retained cells. Behavioral delivery is not fully qualified; see F14–F17. |
| SH-3 | Reduced shared task/chat defaults and removed generic resume contract | Shared renderer and ACPX/Codex/OpenCode/Pi/Hermes/Cursor Cloud regressions; actual legacy invocation prompts checked in the new suite. | Actual legacy prompts measured. Document-delivery regressions and clipped/missing receipts retained; additional carriers remain deferred item 2.2. |
The suite is explicit-only and excluded from `--all`; it does not add paid work
to ordinary campaigns. Negative calibration covers manual regrowth, missing or
malformed prompt receipts, old startup/resume procedures, missing connection
guidance, wrong budgets, and skipped prerequisite assertions. Source revisions,
attempts, cost, partial failures, cleanup, and sanitized evidence use the existing
report pipeline. No paid providers were launched during the initial setup.
Subsequent GitHub measurements are recorded below; no broad coding-quality
improvement or full live qualification is claimed. Unrepresented harnesses remain
unqualified; Codex-through-ACP vendor-base preservation is still F7.
Local setup verification: 478 prerequisite tests passed (477 TypeScript plus one
Rust), 860 Product E2E support tests passed, E2E typecheck passed, and all 24 cells
were discovered. The 313 unrelated native tests filtered by the gate are not
passing coverage. Final prerequisite evidence is retained at
`tests/runner-e2e/results/stock-harness-preflight-2026-10-02T16-15-39.064Z/preflight.json`;
earlier interrupted, missing-Hermes-discovery, and setup/test-timeout attempts are
retained as failures. Hermes is checked through its package config because the
root Vitest project list omits it. Later exact-head prerequisites expanded to
557 passed checks (556 TypeScript and one Rust), and E2E support to 895 tests.
Repository typecheck/build passed. The complete local Vitest run retained three
unrelated timing failures; all three files passed unchanged reruns. At
`1eb5ba420`, 52 CI gates pass and two skip; Greptile is 5/5 with zero open
threads. Later documentation heads require fresh checks. See the live report
for exact source hashes, campaigns, failures and cost limitations.
## Findings ledger
Confirmed mechanics below do not by themselves establish an effect on task quality.
| ID | Finding | Work item / disposition |
| --- | --- | --- |
| F1 | Native Codex sent Paperclip text as `baseInstructions`. Probes on codex-cli 0.153.4 showed replacement; additive `developerInstructions` retained the stock base on start and cold resume. | 1; app-server fix implemented, PR validation pending |
| F1 | Native Codex sent Paperclip text as `baseInstructions`. Probes on codex-cli 0.153.4 showed replacement; additive `developerInstructions` retained the stock base on start and cold resume. | 1; app-server fix merged in PR #14920 |
| F2 | Default hires and role templates prescribe substantial operating procedures; common prompt and wake layers add further coordination text. | 2–3; mechanics confirmed, performance effect unmeasured |
| F3 | Hiring references require legacy Paperclip skill/comment procedures, while native Runner intentionally omits that operational skill and uses semantic tools. | 2–3, 5; reconcile runtime contracts |
| F4 | Local Claude appends instructions; Runner Claude preserves the Claude Code preset. Runner isolation excludes project/local settings, which can also exclude repository instruction discovery. | 4; selective context fix to design |
| F5 | Some Codex capability settings differ between the direct driver and daemon path; an intermediate configuration does not prove the final provider behavior. | 6; effective-path audit pending |
| F6 | Omitting or nulling `baseInstructions` on an old Codex thread's resume preserves its saved replacement; an empty string produces an empty base. | 1; document the required provider session reset; no automatic migration in this PR |
| F7 | The isolated Codex-through-ACP dependency patch also sets `baseInstructions` on start/resume. It is a separate path from the native app-server backend. | 6; follow-up patch/profile audit pending |
| F8 | The operational skill says target-bound confirmations default `supersedeOnUserComment` to true; the default manual and server normalizer say false. The server uses false. | 2; correct stale skill/reference guidance in a follow-up |
| F9 | The default manual required a comment on every task, while the operational skill's verified external-chat shortcut delegates comments and lifecycle bookkeeping to the harness. | 2; unconditional manual rule removed; shared layers still need review |
| F10 | Removing the default manual left the 661-word legacy task template and its 172-word resumed-wake execution contract. Hermes and Pi have additional policy carriers. | 2.1 complete locally: task/chat defaults 113 words, no generic resume contract; 2.2 wrappers pending |
| F11 | Task Markdown and the wake renderer prescribe different accepted-plan behavior for planning-mode accepted-confirmation payloads; tests currently expect both. | 2; unify the directive owner; production reachability still to trace |
| F12 | Native fixed instructions and full-turn constraints repeat completion and uncommon procedures, while reserved finish/block tool descriptions are only one sentence each. | 2; improve tool documentation before removing needed native protocol guidance |
| F13 | Existing context-integrity/chat fixtures injected a QA manual, so their green results did not qualify the production tiny hire default. | Dedicated stock-harness suite measured real default hires; retained failures prevent blanket qualification. |
| F14 | Historical classic Claude/OpenCode skill runs save one Paperclip document; reduced runs save none while completing the task. Legacy ACP Claude has the same behavior under a separate guard failure. Pinned skill storage wording is ambiguous. | Measured delivery failures; keep original oracle. Improve legacy skill/API delivery guidance separately, then compare both variants with preserved and storage-specific cases. |
| F15 | Classic Claude restart chat fails its memory assertion in both variants. | Existing behavior, not attributable to this reduction from these trials. |
| F16 | All six legacy ACP cells per variant fail the persisted-provider-credential guard. Three historical ACP cells also have clipped public prompt retrieval, and candidate ACP Codex chat lacks a complete invocation receipt. | Security/receipt qualification follow-ups; do not waive guard, expose raw sessions, or infer missing provider instructions from clipped receipts. |
| F17 | Prerequisites beside the campaign root violate trusted artifact selection. A same-target pilot superseded 13 candidate cells; one historical AWS runner shut down without upload. | Packaging fixed at `1eb5ba420` and pilot passes. Directory-layout-only copies preserve every byte; missing cells alone recovered at unchanged source, completed failures never rerun. |
Append new findings with evidence, affected paths, and the numbered item that
will address them. Record intentional behavior explicitly rather than as a bug.
@@ -214,6 +466,26 @@ will address them. Record intentional behavior explicitly rather than as a bug.
| 2026-10-02 | Paperclip-owned MCP isolation and configuration changes/session resets are acceptable. | Preserve Paperclip auth and assigned skills while fixing repository context. |
| 2026-10-02 | Dotta requested implementation and a PR for item 1. Native app-server paths now use additive developer instructions. | 139 targeted TypeScript tests and 91 Rust provider tests passed; repository typecheck/build passed. Remaining test/review results to record. |
| 2026-10-02 | Verified actual Codex instruction layering using a localhost Responses stub, without paid inference. | codex-cli 0.153.4 sent identical 14,732-character stock base instructions on start and cold resume, with the Paperclip marker retained in developer input. This is protocol evidence, not a task-quality eval. |
| 2026-10-02 | Dotta requested and confirmed the merge of item 1. | PR #14920 merged at `408f70e69f9c5e49cb4377f4886ac2001bfa67a2`, with all 55 checks passing and Greptile 5/5. |
| 2026-10-02 | Dotta directed an identity-only default manual with no skill or runtime pointers; the harness already handles coordination. | Reduced the default to eight words. Existing hire and onboarding suites passed all 64 tests. Shared prompts and role templates remain pending. |
| 2026-10-02 | Dotta requested three explicit shared-prompt follow-ups and selected common legacy startup/resume reduction first. | Follow-ups 2.1–2.3 recorded. 2.1 complete locally: both defaults 113 words; generic resume contract removed. 563 focused tests and shared utility typecheck/build passed. Extra carriers and native instruction reduction remain pending. |
| 2026-10-02 | Dotta requested executable eval coverage for everything implemented so far and all subsequent changes before continuing. | Added SH-1–SH-3 coverage map, credential-free prerequisite, and 24 production-default-hire cells. 478 prerequisite and 860 support tests passed; E2E typecheck/discovery passed. Live provider results remain `not_run`; item 2.2/2.3 unchanged. |
| 2026-10-02 | Dotta requested PRs and GitHub-runner before/after qualification while discussing subsequent work separately. | Draft [PR #14948](https://github.com/paperclipai/paperclip/pull/14948); native Codex diagnostic [passed](https://github.com/paperclipai/paperclip/actions/runs/37034213743), 34.745 s provider / 56.660 s cell. Comparison restores only the prior manual/shared prompts and holds #14920 constant; no general coding-quality claim. |
| 2026-10-02 | Full candidate cold setup failed before provider admission; stopped and retained the attempt. | [Run 37037105491](https://github.com/paperclipai/paperclip/actions/runs/37037105491), target `36e987246`: prerequisite imports lacked the plugin SDK build. Added ordinary dependency setup before credential-free prerequisites and exact-source coverage of connection guidance. Cold pilot and full matched campaigns pending. |
| 2026-10-02 | Dotta deferred 2.2 additional legacy carriers and approved the first 2.3 native tool-description slice separately. | Keep wrappers open; native finish/block documentation must not be supplied to legacy skill/API completion paths. |
| 2026-10-02 | Matched default-manual/shared-prompt campaigns measured failures, with #14920 constant. | Candidate `f02d8d0df`: 15/24 pass. Historical `12c5433c6`: 15/24 pass after the one runner-shutdown recovery also timed out. Classic Claude/OpenCode Paperclip document delivery regressed in observed trials; keep PR #14948 draft. [Live report](2026-10-02-stock-harness-live-comparison.md). |
| 2026-10-02 | Cold prerequisite and packaging faults were repaired without weakening admission or behavioral graders. | Three setup attempts stopped before providers. Current `1eb5ba420` pilot passes 557 prerequisite checks and protected report publication; source/hash/cost evidence retained. Candidate cancellation and historical runner shutdown recovered only for missing cells. |
| 2026-10-02 | Final bounded recovery completed; no further model reruns. | Both matched cohorts have all 24 results, 15 pass and nine fail. Two classic skill deliveries regress; two OpenCode ordered cases pass only with reduced instructions. Overall parity is not behavioral equivalence. All 48 retained result/receipt projections are hashed and sanitized; original interruptions and partial unknown spend remain recorded. |
| 2026-10-02 | Dotta approved the measured legacy delivery repair and narrow follow-up qualification. | Early operational skill PUT/receipt/link guidance plus generic issue-document reference; eight-word manual retained. Focused original + explicit Paperclip-storage cases on classic Claude/OpenCode, with only the two skill sources varied. All four pairs completed: Claude original Fail → Pass, Claude explicit Pass → Pass, both OpenCode cases Fail → Fail. Explicit OpenCode handoff worsened beneath the unchanged machine grade (clickable API URL → code-formatted path). [Repair report](2026-10-02-legacy-document-skill-repair.md); no full matrix rerun. Claude chat-memory (F15) and ACP credential/receipt failures (F16) remain separate and unresolved. |
| 2026-10-02 | Native tool-description comparison completed all six paired cases. | Zero newly failing cases, three unchanged completion passes, three unchanged blocker failures. Claude/OpenCode blocker API matchers pass in both variants; UI matcher wrongly demanded marker-only replies. Codex's exact-action punctuation failure is unchanged. Original failures retained; corrected 10-case browser calibration and separate retained-DOM replay pass all six visible replies. Exact Codex action failures remain. No native fixed-prompt removal measured. |
| 2026-10-02 | Dotta approved minimal operational skill selection/link correction after retained OpenCode diagnosis. | Candidate `fe9dc1e3c` and baseline `0d7ecfa96d` have two matched Pass → Pass cases, zero new failures/passes and no pending pairs. All four exact-source gates pass 587 checks, all four provider runs and cleanup pass. Candidate original loads Paperclip/reference before delivery, but uses the wrong PAP prefix; baseline original saves publicly later within the same assignment and gives a bare path. Explicit clickable UI links are correct in both. [Complete report](2026-10-02-opencode-skill-routing-link-qualification.md); prior failures retained, no causal or broad quality claim. |
- F18: Native blocker browser assertion required a marker-only reply despite asking for owner/action/reason. Correct marker-plus-explanation checks symmetrically, calibrate contradictory and future-condition replies, retain original verdicts.
- F19: Original prompt-removal cohorts loaded Paperclip; OpenCode skill truncation omitted the late API recipe. The early repair fixed Claude in one paired trial. In the repaired original OpenCode assignment, Paperclip was first loaded only during disposition recovery after a local-file write. Shared legacy operational-skill delivery/selection needs review before more recipe expansion.
- F20: Explicit repaired OpenCode reads the early recipe/reference and saves a public document/revision, but hands off a code-formatted path rather than a clickable anchor. Baseline provides a clickable API URL, rejected by the UI-only oracle. Preserve both grades; clarify future clickable UI-link fixture and review the smallest link example separately.
- F22: Legacy operational skill mounting is already mandatory, but its discovery description omitted ordinary task/heartbeat work and document delivery. The approved narrow fix expands stock metadata selection and adds a real Markdown link example in the existing reference. Original + clarified-explicit OpenCode pairs both pass in both variants, so improvement causality is not established. Candidate loads the operational skill/reference before public delivery; baseline original writes locally first, then saves publicly in the same assignment. Original candidate copies the example PAP prefix into its link while baseline supplies a bare slug path; neither is scored by the original link-free oracle. No full-body injection or native-tool leakage. ACP Claude names-only metadata, custom ACP unsupported delivery, OpenClaw wrappers and Pi HOME differences remain distinct follow-ups.
- F23: The new original OpenCode candidate copies `/PAP/issues/RUN-1#document-context-integrity-output` from the literal example despite actual prefix RUN. It passes durable-document grading but has an incorrect-prefix handoff. Baseline original also lacks a clickable canonical link. Approved reference-only correction now derives the link from `issue.identifier` and the saved receipt key. Independent provider-free calibration checks two other prefixes, redirected keys and rejects a wrong-company link; 25 focused document/source tests, canonical metadata checks and E2E typecheck pass. The recipe now also handles null/absent identifiers through the supported issue-ID route; fresh review caught the initial nullable edge. 56 focused checks pass. Source review shows the UI corrects wrong prefixes, so the frozen PAP href is noncanonical rather than proven broken. These later corrections are not live-qualified by the preserved runs; no further paid rerun.
- F21: Skill heading insertion shifted both generated capability inventories. General CI caught stale metadata after the frozen repair campaign’s narrower build admitted providers. Regenerate canonical derived files and require stale-manifest/inventory checks before future provider admission.
For each completed item, add the chosen behavior, changed paths, verification
results, remaining exceptions, and follow-ups here before checking it off.