From b17019e14d04273ebff538a85385f2659f6d9ef1 Mon Sep 17 00:00:00 2001 From: Dotta <34892728+cryppadotta@users.noreply.github.com> Date: Sat, 3 Oct 2026 12:32:42 -0500 Subject: [PATCH] fix(agents): reduce default instructions and qualify stock harnesses (#14948) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its adapters supply task context and access to Paperclip skills and tools. > - The default hire manual and shared prompts also repeat general work procedures. > - Those procedures overlap with stock provider instructions and the Paperclip skill. > - Existing E2E fixtures supply a QA manual, so they do not qualify the production default. > - This pull request reduces the generic instructions and adds real default-hire coverage. > - The benefit is less competing guidance, with inspectable evidence for preserved skills and task context. ## Linked Issues or Issue Description Refs: #14920. That merged change preserves native Codex base instructions. This PR covers the default manual, shared legacy prompts, operational skill guidance, and the narrowly approved ACP skill-discovery/session-environment repair for measured delivery and credential-persistence failures. **What existing behavior does this improve?** New non-CEO hires without a custom bundle and legacy task/chat startup and continuation prompts. **Current behavior** The shipped default manual contains 602 words. Generic task/chat prompts and ordinary resume deltas repeat work procedures already available through the harness and Paperclip skill. **Proposed behavior** The default manual contains only the eight-word company identity. Shared startup prompts retain identity and connection guidance. Ordinary resume deltas retain current work context without the generic execution contract. **Reason and benefit** Let the stock harness guide general work. Keep Paperclip-specific capabilities and independently test default hires, skills, ordered comments, and chat restart. **Breaking changes** New default hires receive less guidance. Existing saved manuals, explicit custom bundles, CEO templates, and specialized wake contracts retain their behavior. The obsolete includeExecutionContract option remains accepted for source compatibility. ## What Changed - Reduce the default hire manual to one sentence. - Reduce shared task/chat defaults and remove the generic ordinary-resume contract. - Keep connection guidance, auth, skills, custom prompts, and specialized wake context. - Add credential-free instruction-boundary gates and 26 explicit Product E2E cells across eight legacy/native profiles, including two focused Paperclip-storage cases. - Capture public hire receipts before providers run, then grade delivered prompts and independent task/chat outcomes. - Add an early legacy skill API recipe for saving a task document, checking the saved revision receipt and linking the document. Improve stock task/heartbeat skill-selection metadata and show a clickable Markdown UI-link example. Keep native tool completion separate. - Advertise bounded routing descriptions and exact successfully staged SKILL.md paths in legacy ACP Claude; keep full bodies on demand and preserve remote path rebasing. - Remove only the provider environment from copied persisted ACP session records, while loading current run credentials and preserving all other options/conversation state. - Regenerate both capability metadata inventories and reject stale manifests/inventories before provider admission. - Publish the original reduction and focused skill-repair comparisons, preserving all failures, automatic recovery, cost coverage and limitations. ## Verification **Behavioral qualification remains pending.** Original legacy ACP Claude loses the issue document only in the reduced cohort beneath an unchanged credential failure. A source-backed diagnosis finds that neither ordinary assignment reads the staged operational skill, while the runtime persists provider environment in session state. The new common repairs expose skill metadata/path and omit persisted env; strict document and credential guards stay intact. [Inspectable diagnosis and retained hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md). Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb` incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged #14961/#15007). All 222 affected adapter tests, adapter-utils/E2E typechecks, and final 96 variant/grader/retry calibrations pass. The frozen historical comparator is `c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths identical, exactly two production instruction paths plus nine declared unit expectations differ. The operational skill/discovery/environment repairs, selected model/profile/task/core grader/auth/permissions/retry policy are identical. Both actual launcher prepare→verify admissions pass with zero providers. [Immutable manifest and exact receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json). One original legacy ACP Claude cell per variant is authorized, with enforced single campaign attempts, 12-minute deadlines and company/agent 1,000-cent hard stops; every product recovery run/cost is counted. Actual live outcomes are pending. Current normal CI has one failed server shard and failed aggregate verify under diagnosis; other normal gates including typecheck/build/Rust/all eight browser shards pass. Fresh review completed successfully; the valid historical startup/resume masking finding was fixed with per-invocation task/chat checks and strict complete-snapshot capture, calibrated and resolved. Prior heads, failures and campaigns below remain historical evidence, not checks on this repair head. - Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52 current-head checks pass with two intentional Storybook skips, including repository typecheck/test/build and the browser shard. Fresh Greptile is 5/5 with zero unresolved review threads. Exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained receipt verification, zero providers/source errors. Fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck and 26-cell stock discovery pass. Canonical contract/inventory checks and the later issue-derived reference calibration are retained; that reference-only follow-up is not live-qualified by earlier frozen runs. - Prior full repository typecheck/build passed. The complete local Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three timing failures. All three affected files passed unchanged narrow reruns; original failures remain retained. Current-head CI now passes the full general checks; the original local failures remain retained. - The original 24-pair default-manual/shared-prompt comparison has two new overall classic Claude/OpenCode document-delivery failures plus an additional legacy ACP Claude document loss beneath an unchanged credential-guard failure (not closed by later runs), two newly passing OpenCode ordered cases, seven unchanged failures and 13 unchanged passes. Equal 15/24 totals do not establish behavioral equivalence. [Complete original report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md). - The skill-only repair holds the eight-word manual/shared prompts and merged #14920 fixed. All four matched profile configurations and 203 fixture/behavior files match. Candidate `abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill sources against baseline `bc83fe030234439ac51279502a28803958963e2e`. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547) and [baseline workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047) each pass 571 exact-source prerequisites before providers; all eight cells clean up successfully. Failed campaigns publish successfully and remain failed. - Repair pairs: Claude original Fail → Pass; Claude explicit Pass → Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff worsens beneath the unchanged failing UI-link grade: baseline gives a clickable API URL, candidate gives a code-formatted path without an anchor. The request's usable-link wording is narrower in the UI-only oracle. [Complete repair report and safe projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md). - The subsequent narrow stock metadata/link correction has two matched Pass → Pass cases, zero new machine failures/passes and no pending pairs. Both original-case handoff links remain deficient: candidate uses a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the preserved original oracle only requires a durable document. Both explicit clickable UI-link cases pass revision/content/link grading. All four exact-source 587-check gates, single assignment runs and cleanup pass. This does not establish fix causality because baseline also succeeds. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401) freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374) freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only comparison with reduced manuals/shared prompts held constant, not a repeat of the historical-manual comparison. Only original and clarified explicit classic OpenCode cases are selected, two per variant/four expected turns. 8,242 other tracked files and both profile hashes match; protected workflows admit each exact source before credentials. [Complete qualification report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md). Candidate original loads Paperclip/reference before saving publicly; baseline original loads it after writing locally, then saves publicly within the same assignment. Reported cost totals are $0.0107824490 candidate / $0.0107909015 baseline, with unmetered runtime. The later reference-only issue-derived link correction is provider-free calibrated and **not live-qualified** by these frozen runs; no further paid runs. - Retained tool calls show the repaired original OpenCode assignment loads only its assigned output skill before writing locally. Operational Paperclip is first loaded during automatic disposition recovery; its early recipe is visible then, but it never saves the missing document. Explicit candidate loads Paperclip and reads the new reference before saving successfully. All nine actual runs are counted. Reported LLM totals are $0.3802537209 baseline and $0.4918990161 candidate; local runtime is unmetered. - Initial setup, packaging, cancelled/missing-cell recovery, callback test and relative-output attempts remain retained. No completed provider failure was rerun. Frozen measurement branches are unchanged by later canonical metadata maintenance. - Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`, and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly for paid execution; it is excluded from `--all`. Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` replays this PR on merged hiring #14985 (`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four explicit custom-CEO-bundle checks, minimal generic manual boundary, and both suites. The combined fixture catalog and hiring calibrations pass 67 assertions; exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained-receipt verification, zero providers/source errors, fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E typecheck and 26-cell stock discovery pass. Fresh current-head CI passes all 52 checks with two intentional skips, and fresh Greptile is 5/5 with zero unresolved review threads. The prior source-plan browser failure is retained: a deterministic process fixture replayed its last `fixture:plan` command on `chat_task_completed`, writing revision 2 with identical body after the approval handoff. This was not paid provider execution. Rebased current-head CI passes the same assertion without an old-head retry or a change to that browser fixture. The merged hiring change was measured separately on immutable matched unions, with this reduced/shared/operational context and native completion guidance held constant. [Complete original two-profile report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md): [candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208) / [historical baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463), 705 provider-free prerequisites each. Both pairs are unchanged Fail → Fail on the exact-five count, with six core delivery checks passing all four cells; 28 actual successful runs include eight automatic completion wakes, zero retries, four successful cleanups. Source-read coverage is uncomparable, actual model charges unknown. Separately versioned provider-free accounting remains analytical work; original verdicts are preserved. This does not rerun or qualify the completed default-manual or native campaigns. ## Risks - Legacy ACP Claude's additional delivery loss is not closed by any later matched run and blocks the no-extra-failing-behavior merge criterion. Legacy document delivery may have relied on the prior manual/shared prompts. The early skill repair improves Claude in one trial; the later OpenCode pairs pass in both variants and cannot establish causality or robust recovery. Both original-case links remain deficient beneath the storage-only grade. The later issue-derived reference correction has only provider-free validation. Native finish/block descriptions must not be supplied to legacy agents. - The comparison holds merged native Codex fix #14920 constant; it cannot measure that fix's before/after task performance. - These bounded skill/context/chat workflows do not measure general coding quality. Unrepresented providers remain unqualified. - Saved manuals and old Codex sessions are not automatically migrated. Codex through ACP still has a separate base-instruction follow-up. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution, and delegated PR/eval work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (relevant suites and all three unchanged narrow reruns pass; complete-run timing failures retained in Verification) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green on the new repair head (prior-head checks retained above) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the new repair head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip --- doc/evals.md | 8 + ...26-10-02-legacy-document-skill-repair.json | 1187 ++++++++++ ...2026-10-02-legacy-document-skill-repair.md | 51 + ...code-skill-routing-link-qualification.json | 1418 ++++++++++++ ...encode-skill-routing-link-qualification.md | 75 + ...6-10-02-stock-harness-live-comparison.json | 357 ++++ ...026-10-02-stock-harness-live-comparison.md | 110 + ...10-02-stock-harness-paperclip-checklist.md | 308 ++- package.json | 1 + .../ephemeral-session-environment.test.ts | 165 ++ .../ephemeral-session-environment.ts | 32 + .../src/acpx-engine/execute.test.ts | 104 +- .../adapter-utils/src/acpx-engine/execute.ts | 39 +- .../adapter-utils/src/prompt-sections.test.ts | 13 +- .../adapter-utils/src/server-utils.test.ts | 134 +- packages/adapter-utils/src/server-utils.ts | 57 +- .../cursor-cloud/src/server/execute.test.ts | 5 +- .../src/server/prompt-rendering.test.ts | 6 +- .../opencode-local/src/server/execute.test.ts | 5 +- .../src/server/execute.remote.test.ts | 12 +- .../docs/capability-contract.md | 177 +- .../generated/capability/capabilities.yaml | 143 +- .../capability/capability-contract.md | 4 +- .../check-capability-inventory.test.mjs | 2 +- .../generate-capability-contract.d.mts | 1 + .../scripts/generate-capability-contract.mjs | 2 +- .../scripts/lib/capability-inventory.mjs | 12 +- .../spec/capability/capabilities.yaml | 1904 +++++++++-------- .../spec/capability/source-contract.json | 1 + .../src/generated/capability-contract.ts | 4 +- .../test/capability-contract.test.mjs | 15 +- .../src/__tests__/agent-skills-routes.test.ts | 32 +- .../src/__tests__/codex-local-execute.test.ts | 3 +- .../src/__tests__/workspace-runtime.test.ts | 35 +- .../src/onboarding-assets/default/AGENTS.md | 25 +- server/src/routes/agents.ts | 4 +- .../services/onboarding-first-task-assets.ts | 4 +- skills/paperclip/SKILL.md | 10 +- .../paperclip/references/issue-documents.md | 58 + tests/runner-e2e/FIXTURES.md | 11 + tests/runner-e2e/README.md | 8 + tests/runner-e2e/STOCK-HARNESS.md | 248 +++ tests/runner-e2e/catalog.test.ts | 6 +- tests/runner-e2e/catalog.ts | 28 +- tests/runner-e2e/context-integrity-cases.ts | 24 +- tests/runner-e2e/context-integrity-flow.ts | 16 +- tests/runner-e2e/context-integrity-scoring.ts | 31 +- .../historical-default-agents.md | 24 + tests/runner-e2e/launch.ts | 25 +- tests/runner-e2e/live-fixtures.test.ts | 24 + tests/runner-e2e/live-fixtures.ts | 5 +- .../native-completion-defaults.test.ts | 10 +- tests/runner-e2e/paperclip-document.test.ts | 89 + tests/runner-e2e/runner.spec.ts | 36 + .../runner-e2e/select-rerun-artifacts.test.ts | 13 + .../stock-harness-admission.test.ts | 44 + tests/runner-e2e/stock-harness-admission.ts | 38 + tests/runner-e2e/stock-harness-checks.mjs | 246 +++ .../runner-e2e/stock-harness-checks.test.mjs | 79 + tests/runner-e2e/stock-harness-digest.test.ts | 27 + .../stock-harness-instruction-variant.d.mts | 7 + .../stock-harness-instruction-variant.mjs | 24 + ...stock-harness-instruction-variant.test.mjs | 48 + .../runner-e2e/stock-harness-manifest.test.ts | 21 + tests/runner-e2e/stock-harness-manifest.ts | 19 + tests/runner-e2e/stock-harness.test.ts | 189 ++ tests/runner-e2e/stock-harness.ts | 184 ++ tests/runner-e2e/vitest.config.ts | 2 +- 68 files changed, 6635 insertions(+), 1414 deletions(-) create mode 100644 doc/plans/2026-10-02-legacy-document-skill-repair.json create mode 100644 doc/plans/2026-10-02-legacy-document-skill-repair.md create mode 100644 doc/plans/2026-10-02-opencode-skill-routing-link-qualification.json create mode 100644 doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md create mode 100644 doc/plans/2026-10-02-stock-harness-live-comparison.json create mode 100644 doc/plans/2026-10-02-stock-harness-live-comparison.md create mode 100644 packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.test.ts create mode 100644 packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.ts create mode 100644 packages/paperclip-runner/scripts/generate-capability-contract.d.mts create mode 100644 skills/paperclip/references/issue-documents.md create mode 100644 tests/runner-e2e/STOCK-HARNESS.md create mode 100644 tests/runner-e2e/fixtures/stock-harness/historical-default-agents.md create mode 100644 tests/runner-e2e/paperclip-document.test.ts create mode 100644 tests/runner-e2e/stock-harness-admission.test.ts create mode 100644 tests/runner-e2e/stock-harness-admission.ts create mode 100644 tests/runner-e2e/stock-harness-checks.mjs create mode 100644 tests/runner-e2e/stock-harness-checks.test.mjs create mode 100644 tests/runner-e2e/stock-harness-digest.test.ts create mode 100644 tests/runner-e2e/stock-harness-instruction-variant.d.mts create mode 100644 tests/runner-e2e/stock-harness-instruction-variant.mjs create mode 100644 tests/runner-e2e/stock-harness-instruction-variant.test.mjs create mode 100644 tests/runner-e2e/stock-harness-manifest.test.ts create mode 100644 tests/runner-e2e/stock-harness-manifest.ts create mode 100644 tests/runner-e2e/stock-harness.test.ts create mode 100644 tests/runner-e2e/stock-harness.ts diff --git a/doc/evals.md b/doc/evals.md index f936899203..c529972e76 100644 --- a/doc/evals.md +++ b/doc/evals.md @@ -95,6 +95,14 @@ execution ID because `--all` excludes explicit-only suites. Each cell applies a 1,000-cent company and agent budget hard stop before task creation and records both limits in its evidence. +The explicit-only [stock-harness suite](../tests/runner-e2e/STOCK-HARNESS.md) +reuses skill, ordered-continuation, and chat-restart journeys across eight local +legacy/native profiles with production-default hires. It closes the custom QA +manual coverage gap. Its required credential-free prerequisite maps vendor +instruction layering, the tiny hire bundle, and shared startup/resume reductions +to executable checks. The 24 live cells are configured; no live qualification is +claimed from their setup or unit calibration. + The explicit-only `agent-chat-stories` suite covers the experimental settings lifecycle for a configured native agent and follow-ups during active work. Its fixture-driven file wait and persisted-plan oracle are documented in the diff --git a/doc/plans/2026-10-02-legacy-document-skill-repair.json b/doc/plans/2026-10-02-legacy-document-skill-repair.json new file mode 100644 index 0000000000..b215bc08b8 --- /dev/null +++ b/doc/plans/2026-10-02-legacy-document-skill-repair.json @@ -0,0 +1,1187 @@ +{ + "schema": "paperclip.legacy-document-skill-repair.safe-comparison/v1", + "asOf": "2026-10-02T21:22:00Z", + "partial": false, + "expectedCellsPerVariant": 4, + "expectedProviderTurnsPerVariant": 4, + "candidateSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "baselineSha": "bc83fe030234439ac51279502a28803958963e2e", + "originalMachinePairedDelta": { + "newFailures": 0, + "newPasses": 1, + "unchangedPasses": 1, + "unchangedFailures": 2, + "pending": 0 + }, + "observedBehavioralDeteriorations": [ + "OpenCode explicit case: candidate lacks a clickable link, while baseline provided a clickable API document URL. Both fail original UI-only oracle; equal machine verdicts hide this presentation deterioration." + ], + "baseline": [ + { + "executionId": "stock-harness.legacy-claude.local.assigned-skill-explicit-invocation", + "profileId": "legacy-claude", + "caseId": "assigned-skill-explicit-invocation", + "status": "failed", + "model": "claude-sonnet-4-6", + "durationMs": 34592, + "cleanup": "passed", + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 28439, + "outputTokens": 1137, + "cachedInputTokens": 143829, + "totalTokens": 173405, + "reportedCostUsd": 0.16683944999999997, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 23032, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.16683944999999997, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.16683944999999997, + "complete": true + }, + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "providerDurationMs": 23032, + "resultSha256": "67bcbe152852280b58fec97754601c57c416c2cb7965de55557240b8448dca63", + "contextEvidenceSha256": "9bc59db9678a3f5cfc207d1aeeec1f868f94429a7a10d8d214621f286532a1b7", + "failedBehaviorChecks": [ + "single-durable-output" + ], + "publicDocumentCount": 0, + "savedRevisionEvidence": false, + "runPurposes": [ + "assignment" + ], + "prerequisites": { + "passed": true, + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "sourceFingerprint": "27fbbc1792dca136c9c72dd492ff16c18bce0638ce4979985d798cde85175201", + "providerCallsBeforeAdmission": 0, + "executedChecks": 571 + } + }, + { + "executionId": "stock-harness.legacy-claude.local.assigned-skill-paperclip-document", + "profileId": "legacy-claude", + "caseId": "assigned-skill-paperclip-document", + "status": "passed", + "model": "claude-sonnet-4-6", + "durationMs": 50778, + "cleanup": "passed", + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 27675, + "outputTokens": 2062, + "cachedInputTokens": 216852, + "totalTokens": 246589, + "reportedCostUsd": 0.19975709999999997, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 34999, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.19975709999999997, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.19975709999999997, + "complete": true + }, + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "providerDurationMs": 34999, + "resultSha256": "bca45a760fc5ec30ae8f154d9ec52ab4073b3ce2375256b8718f91200cb79d4e", + "contextEvidenceSha256": "473f3b417e71ee823ae9aad2596c7f1379b3fd5cf8d8fd540c30a4f3a2e10e21", + "failedBehaviorChecks": [], + "publicDocumentCount": 1, + "savedRevisionEvidence": true, + "runPurposes": [ + "assignment" + ], + "prerequisites": { + "passed": true, + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "sourceFingerprint": "27fbbc1792dca136c9c72dd492ff16c18bce0638ce4979985d798cde85175201", + "providerCallsBeforeAdmission": 0, + "executedChecks": 571 + } + }, + { + "executionId": "stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation", + "profileId": "legacy-opencode", + "caseId": "assigned-skill-explicit-invocation", + "status": "failed", + "model": "openrouter/deepseek/deepseek-v4-flash-0731", + "durationMs": 47625, + "cleanup": "passed", + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 24068, + "outputTokens": 1960, + "cachedInputTokens": 149248, + "totalTokens": 175276, + "reportedCostUsd": 0.0033927116000000003, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 36770, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.0033927116000000003, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.0033927116000000003, + "complete": true + }, + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "providerDurationMs": 36770, + "resultSha256": "a322e292fd7688c714cc8e57e7f64ca70c19f0784dc242b3e809a6e4b0be224e", + "contextEvidenceSha256": "3e295ec42b9c7208a7e4b17bf8329d9e1f01ba984c4a586539d7a703ffd59bf6", + "failedBehaviorChecks": [ + "single-durable-output" + ], + "publicDocumentCount": 0, + "savedRevisionEvidence": false, + "runPurposes": [ + "assignment" + ], + "prerequisites": { + "passed": true, + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "sourceFingerprint": "27fbbc1792dca136c9c72dd492ff16c18bce0638ce4979985d798cde85175201", + "providerCallsBeforeAdmission": 0, + "executedChecks": 571 + } + }, + { + "executionId": "stock-harness.legacy-opencode.local.assigned-skill-paperclip-document", + "profileId": "legacy-opencode", + "caseId": "assigned-skill-paperclip-document", + "status": "failed", + "model": "openrouter/deepseek/deepseek-v4-flash-0731", + "durationMs": 217620, + "cleanup": "passed", + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 78795, + "outputTokens": 4331, + "cachedInputTokens": 846848, + "totalTokens": 929974, + "reportedCostUsd": 0.0102644593, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 206191, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.0102644593, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.0102644593, + "complete": true + }, + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "providerDurationMs": 206191, + "resultSha256": "60dfb068ea80d0155d93fdace77a29fd19d64bdb0740e6e6b63f7e83ccfb5ded", + "contextEvidenceSha256": "23eed2766b5e8db9cea0ba5c76da235cf2d4222fa5718420cc60d6e13c86fa06", + "failedBehaviorChecks": [ + "saved-document-link" + ], + "publicDocumentCount": 1, + "savedRevisionEvidence": true, + "runPurposes": [ + "assignment" + ], + "prerequisites": { + "passed": true, + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "sourceFingerprint": "27fbbc1792dca136c9c72dd492ff16c18bce0638ce4979985d798cde85175201", + "providerCallsBeforeAdmission": 0, + "executedChecks": 571 + }, + "observedHandoff": "clickable API document URL; rejected by UI-only link oracle" + } + ], + "candidate": [ + { + "executionId": "stock-harness.legacy-claude.local.assigned-skill-explicit-invocation", + "profileId": "legacy-claude", + "caseId": "assigned-skill-explicit-invocation", + "status": "passed", + "model": "claude-sonnet-4-6", + "durationMs": 46702, + "cleanup": "passed", + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 40330, + "outputTokens": 1336, + "cachedInputTokens": 185090, + "totalTokens": 226756, + "reportedCostUsd": 0.22679475, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 27702, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.22679475, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.22679475, + "complete": true + }, + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "providerDurationMs": 27702, + "resultSha256": "2f66f62dc2f16756400a0f5b0b0d55e5c9f5f7cb2590d84ba05f3d4ee6b4ad66", + "contextEvidenceSha256": "c5619df489213873db64d35424b08f9852b0099bab07e92d3a71f993008de389", + "failedBehaviorChecks": [], + "publicDocumentCount": 1, + "savedRevisionEvidence": true, + "runPurposes": [ + "assignment" + ], + "prerequisites": { + "passed": true, + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "sourceFingerprint": "441999ae2618f1f1e9f430c099bbef63d2948495599f0d8b342d8e5f027a7f4c", + "providerCallsBeforeAdmission": 0, + "executedChecks": 571 + } + }, + { + "executionId": "stock-harness.legacy-claude.local.assigned-skill-paperclip-document", + "profileId": "legacy-claude", + "caseId": "assigned-skill-paperclip-document", + "status": "passed", + "model": "claude-sonnet-4-6", + "durationMs": 56928, + "cleanup": "passed", + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 42017, + "outputTokens": 2468, + "cachedInputTokens": 205226, + "totalTokens": 249711, + "reportedCostUsd": 0.25614254999999997, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 40513, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.25614254999999997, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.25614254999999997, + "complete": true + }, + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "providerDurationMs": 40513, + "resultSha256": "8a0ef693477a03925d8166d1872b5db4f5547ae48eb9df934beabcce9eccf147", + "contextEvidenceSha256": "f7aba5c9653567a0c3ee252d01fbcee20c81c432b4402d76c7e00221b6917f2e", + "failedBehaviorChecks": [], + "publicDocumentCount": 1, + "savedRevisionEvidence": true, + "runPurposes": [ + "assignment" + ], + "prerequisites": { + "passed": true, + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "sourceFingerprint": "441999ae2618f1f1e9f430c099bbef63d2948495599f0d8b342d8e5f027a7f4c", + "providerCallsBeforeAdmission": 0, + "executedChecks": 571 + } + }, + { + "executionId": "stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation", + "profileId": "legacy-opencode", + "caseId": "assigned-skill-explicit-invocation", + "status": "failed", + "model": "openrouter/deepseek/deepseek-v4-flash-0731", + "durationMs": 79578, + "cleanup": "passed", + "billing": { + "llm": { + "runCount": 2, + "runsWithTokenUsage": 2, + "runsWithReportedCost": 2, + "inputTokens": 54946, + "outputTokens": 2264, + "cachedInputTokens": 169472, + "totalTokens": 226682, + "reportedCostUsd": 0.0047148437000000005, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 68032, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.0047148437000000005, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.0047148437000000005, + "complete": true + }, + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "providerDurationMs": 68032, + "resultSha256": "f807c235250d8675d29aa34fc97d0e89be2c7ea885edfc2f3845627037552c0b", + "contextEvidenceSha256": "0dc98ad8ef1f22d69c3a3d7344082a6d5a0effc8d9c73c5d4c1267c92800b57a", + "failedBehaviorChecks": [ + "single-durable-output" + ], + "publicDocumentCount": 0, + "savedRevisionEvidence": false, + "runPurposes": [ + "assignment", + "issue_disposition_repair" + ], + "prerequisites": { + "passed": true, + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "sourceFingerprint": "441999ae2618f1f1e9f430c099bbef63d2948495599f0d8b342d8e5f027a7f4c", + "providerCallsBeforeAdmission": 0, + "executedChecks": 571 + } + }, + { + "executionId": "stock-harness.legacy-opencode.local.assigned-skill-paperclip-document", + "profileId": "legacy-opencode", + "caseId": "assigned-skill-paperclip-document", + "status": "failed", + "model": "openrouter/deepseek/deepseek-v4-flash-0731", + "durationMs": 59941, + "cleanup": "passed", + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 33724, + "outputTokens": 2597, + "cachedInputTokens": 147200, + "totalTokens": 183521, + "reportedCostUsd": 0.0042468724, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 46300, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.0042468724, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.0042468724, + "complete": true + }, + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "providerDurationMs": 46300, + "resultSha256": "cff6e5bf4b6b646ae6f9b5a6a2934c444c37a14f50b591f5226ccac0e4de1dda", + "contextEvidenceSha256": "6c943d78291c1e493cb29ca3a52dcb96d186fd18ed51e4cfb17055649162f6f2", + "failedBehaviorChecks": [ + "saved-document-link" + ], + "publicDocumentCount": 1, + "savedRevisionEvidence": true, + "runPurposes": [ + "assignment" + ], + "prerequisites": { + "passed": true, + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "sourceFingerprint": "441999ae2618f1f1e9f430c099bbef63d2948495599f0d8b342d8e5f027a7f4c", + "providerCallsBeforeAdmission": 0, + "executedChecks": 571 + }, + "observedHandoff": "code-formatted prefix-less UI document path; no clickable anchor" + } + ], + "matchedProfileProof": { + "schema": "paperclip.legacy-document-skill-repair.fixture-match-proof/v1", + "cells": [ + { + "executionId": "stock-harness.legacy-claude.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "6c21be9bb3fcec6274372fc586d88e42d82bc0d438e26a69ef614f8e6fcb8519", + "model": "claude-sonnet-4-6" + }, + { + "executionId": "stock-harness.legacy-claude.local.assigned-skill-paperclip-document", + "fixtureConfigSha256": "ea1cd67b0f68ef3830c011768d1a1133b5828c602d5ae3e021889dcd72708344", + "model": "claude-sonnet-4-6" + }, + { + "executionId": "stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "dd673f469cab583d26bc2e4b484f18b27884ec14a59542177026389181e2da2c", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + }, + { + "executionId": "stock-harness.legacy-opencode.local.assigned-skill-paperclip-document", + "fixtureConfigSha256": "b5173358d664ab6b3daa67d8758f18d59ec99cb9c485c02ccf953a9f8375e01e", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + } + ] + }, + "identicalSourceCount": 203, + "skillSources": [ + { + "path": "skills/paperclip/SKILL.md", + "candidate": { + "present": true, + "sha256": "65e379cab319bfef5911c0f6c9423cb262c9b8f1fef3fd08d45c1e4f32f8971a" + }, + "baseline": { + "present": true, + "sha256": "77aafc0c7b16a82254af642f0ebff723dcef7d7fbb6d3354efc0c11032e0344d" + } + }, + { + "path": "skills/paperclip/references/issue-documents.md", + "candidate": { + "present": true, + "sha256": "353845a406b1b4948487c348c86eb7b8c89a55ae67afe203100eee10aee46336" + }, + "baseline": { + "present": false, + "sha256": null + } + } + ], + "toolDeliveryAudit": { + "schema": "paperclip.legacy-document-repair.tool-audit/v1", + "rows": [ + { + "variant": "candidate", + "case": "assigned-skill-paperclip-document", + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "runOrdinal": 1, + "runPurpose": "assignment", + "startedAt": "2026-10-02T20:37:40.215Z", + "contextEvidenceSha256": "6c943d78291c1e493cb29ca3a52dcb96d186fd18ed51e4cfb17055649162f6f2", + "toolCalls": [ + { + "tool": "skill", + "skill": "context-integrity-output-7b55f897b3b41", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "skill", + "skill": "paperclip", + "operationalSkillLoad": true, + "earlyRecipeVisible": true, + "newReferencePointerVisible": true, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "read", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": true, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": true, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": true + } + ] + }, + { + "variant": "candidate", + "case": "assigned-skill-explicit-invocation", + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "runOrdinal": 1, + "runPurpose": "assignment", + "startedAt": "2026-10-02T20:37:37.587Z", + "contextEvidenceSha256": "0dc98ad8ef1f22d69c3a3d7344082a6d5a0effc8d9c73c5d4c1267c92800b57a", + "toolCalls": [ + { + "tool": "skill", + "skill": "context-integrity-output-60e03959432b1", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "write", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": true, + "documentApiCall": false, + "statusApiCall": false + } + ] + }, + { + "variant": "candidate", + "case": "assigned-skill-explicit-invocation", + "sourceSha": "abd0b628ca642c09a54a4edc56a5227402f6686e", + "runOrdinal": 2, + "runPurpose": "issue_disposition_repair", + "startedAt": "2026-10-02T20:37:52.592Z", + "contextEvidenceSha256": "0dc98ad8ef1f22d69c3a3d7344082a6d5a0effc8d9c73c5d4c1267c92800b57a", + "toolCalls": [ + { + "tool": "skill", + "skill": "paperclip", + "operationalSkillLoad": true, + "earlyRecipeVisible": true, + "newReferencePointerVisible": true, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": true + } + ] + }, + { + "variant": "baseline", + "case": "assigned-skill-explicit-invocation", + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "runOrdinal": 1, + "runPurpose": "assignment", + "startedAt": "2026-10-02T20:37:57.110Z", + "contextEvidenceSha256": "3e295ec42b9c7208a7e4b17bf8329d9e1f01ba984c4a586539d7a703ffd59bf6", + "toolCalls": [ + { + "tool": "skill", + "skill": "context-integrity-output-816e6b4cdd6a1", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "write", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": true, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "skill", + "skill": "paperclip", + "operationalSkillLoad": true, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": true + }, + { + "tool": "write", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": true + } + ] + }, + { + "variant": "baseline", + "case": "assigned-skill-paperclip-document", + "sourceSha": "bc83fe030234439ac51279502a28803958963e2e", + "runOrdinal": 1, + "runPurpose": "assignment", + "startedAt": "2026-10-02T20:38:02.031Z", + "contextEvidenceSha256": "23eed2766b5e8db9cea0ba5c76da235cf2d4222fa5718420cc60d6e13c86fa06", + "toolCalls": [ + { + "tool": "skill", + "skill": "context-integrity-output-06ce9804653e1", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "skill", + "skill": "paperclip", + "operationalSkillLoad": true, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "read", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": true, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "read", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "read", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": true, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": true, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false + }, + { + "tool": "bash", + "skill": null, + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": true, + "statusApiCall": true + } + ] + } + ] + }, + "limitations": [ + "Tiny manual/shared prompts and merged native Codex base fix remain fixed; only the two operational skill sources differ.", + "Original task-document wording is storage-ambiguous; explicit case requests a usable link but its oracle requires the canonical UI route, so the baseline API-link failure has an oracle/request ambiguity.", + "Provider-free source review confirms a board-readable GET document endpoint and an unprefixed UI redirect preserving the hash; actual authenticated browser navigation in the cleaned-up measured instance was not replayed.", + "One paired trial cannot establish broad performance, causality, timing or spending improvement.", + "Candidate OpenCode original has an additional automatic disposition-repair run; all nine actual runs and costs are retained.", + "Local runtime is unmetered; reported LLM cost is not a statement of actual external billing.", + "Subsequent generated capability metadata and pre-provider drift checks are derived maintenance after these measured frozen revisions, not a new provider qualification." + ] +} diff --git a/doc/plans/2026-10-02-legacy-document-skill-repair.md b/doc/plans/2026-10-02-legacy-document-skill-repair.md new file mode 100644 index 0000000000..728e0b7c48 --- /dev/null +++ b/doc/plans/2026-10-02-legacy-document-skill-repair.md @@ -0,0 +1,51 @@ +# Legacy document skill repair qualification — 2026-10-02 + +**TL;DR: one newly passing case (Claude original document delivery), one unchanged pass (Claude explicit storage), two unchanged OpenCode failures, and no pending pairs or new overall machine failures. OpenCode's explicit-case clickable handoff nevertheless got worse:** baseline supplied a clickable API document URL, while candidate supplied only a code-formatted path. Both fail the original UI-only link oracle, whose intended route was narrower than the request's "usable link" wording. This result does not establish general non-regression. + +This focused repair keeps the eight-word manual, reduced shared prompts and merged Codex base fix #14920 fixed. It varies only the early operational Paperclip skill recipe and the presence of the new `issue-documents.md` reference. The earlier combined prompt-removal comparison retains its two new classic Claude/OpenCode storage failures; this trial repairs Claude's original case but leaves OpenCode's original regression unresolved. + +| Classic profile | Case | Pre-fix skill → repaired skill | Retained evidence | +| --- | --- | --- | --- | +| Claude | Original assigned skill | Fail → Pass | Baseline lacks a durable task document; candidate saves one containing the skill marker. | +| Claude | Explicit Paperclip document | Pass → Pass | Both save the document/revision and provide a canonical UI link. | +| OpenCode | Original assigned skill | Fail → Fail | Both write a workspace file rather than a public task document. Candidate requires automatic disposition recovery. | +| OpenCode | Explicit Paperclip document | Fail → Fail | Both save the document/revision. Baseline provides a clickable API URL; candidate gives a code-formatted UI path without an anchor. Original UI-link checks fail in both. | + +Original machine verdicts and evidence hashes are retained in the [safe evidence projection](2026-10-02-legacy-document-skill-repair.json). No model was rerun to improve these results. + +## Frozen sources and public reports + +- [Repaired skill campaign](https://github.com/paperclipai/paperclip/actions/runs/37060885547), source `abd0b628ca642c09a54a4edc56a5227402f6686e`, branch `codex/legacy-document-repaired-skill`. [Published candidate report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37060885547-1/index.html). +- [Pre-fix skill campaign](https://github.com/paperclipai/paperclip/actions/runs/37060888047), source `bc83fe030234439ac51279502a28803958963e2e`, branch `codex/legacy-document-pre-fix-skill`. [Published baseline report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37060888047-1/index.html). +- Both trusted default-branch workflows ran concurrently on separate concurrency targets. All eight cells pass their exact-source prerequisite gate before providers; all eight clean up successfully. Failed campaigns are published successfully and remain failed. +- The four model/effort/auth/tool/permission/environment configurations match. Models are `claude-sonnet-4-6` and `openrouter/deepseek/deepseek-v4-flash-0731`. The comparison records 203 byte-identical fixture/behavior sources and explicit present/absent skill-source fingerprints. The new reference is absent from the baseline, rather than silently introduced into it. +- Local admission preparation passed 909 E2E support tests, E2E typecheck and four-cell discovery; the candidate exact-source gate passed 571 checks, including one Rust check. Initial sandbox and obsolete catalog-count attempts remain retained. The original request and behavioral oracle were not weakened. + +## What OpenCode actually received + +The repaired original assignment invokes only its assigned Context integrity output skill, then writes `task-document.md`. There is no operational Paperclip skill load or issue-document reference read before that local write. A separate automatic `issue_disposition_repair` run then loads Paperclip. Its returned content contains the new early saved-document paragraph and reference pointer even though the overall skill result is marked truncated; it never reads the reference or saves a document, and only finalizes the issue's status. + +This is evidence of operational-skill selection after work, not evidence that the assignment saw and ignored the new recipe. The pre-fix original run also writes locally before loading the old Paperclip skill. Whether the adapter should activate operational Paperclip guidance before work is a production-delivery question for review; expanding a skill the first assignment never opened would not address the observed sequence. + +The explicit candidate loads the assigned skill, then Paperclip, receives the new early paragraph, reads `issue-documents.md`, and calls the document API successfully. Its saved document has a persisted revision. The remaining defect is its comment's non-clickable path. Baseline supplied an actual Markdown link to the issue-document API GET endpoint. Read-only source review confirms that this endpoint returns the document in board context, and the unprefixed UI issue route redirects through the selected company's prefix while preserving the document hash. Actual authenticated browser navigation against these cleaned-up measured instances was not replayed, so neither URL is classified as a proven broken route. Candidate's missing clickable anchor is directly observable. + +A future explicit fixture should name a clickable Paperclip UI document link if that is the intended contract. This clarification would not retroactively pass or fail either original result. A short canonical Markdown UI-link example can be considered separately from operational-skill delivery. No further production instruction change or paid rerun was made for this diagnosis. + +## Runs, timing and cost + +Eight provider turns were expected; nine actual runs are retained. Candidate OpenCode's original assignment succeeds but leaves the task `in_progress`, triggering automatic disposition recovery. The recovery succeeds and marks it done, without creating the missing document. Both runs, their time and cost are counted rather than selecting a better attempt. + +| Profile | Case | Baseline provider seconds | Candidate provider seconds | Baseline reported USD | Candidate reported USD | +| --- | --- | ---: | ---: | ---: | ---: | +| legacy-claude | Original | 23.032 | 27.702 | 0.1668394500 | 0.2267947500 | +| legacy-claude | Explicit storage | 34.999 | 40.513 | 0.1997571000 | 0.2561425500 | +| legacy-opencode | Original | 36.770 | 68.032 | 0.0033927116 | 0.0047148437 | +| legacy-opencode | Explicit storage | 206.191 | 46.300 | 0.0102644593 | 0.0042468724 | + +Reported LLM totals are $0.3802537209 baseline and $0.4918990161 candidate. All four baseline and five candidate ledger rows contain token usage and reported cost. Local runtime is unmetered and actual external billing is not established. Single trials, the additional recovery and baseline OpenCode's long explicit-case execution prevent a general timing or spending conclusion. + +## PR maintenance and remaining qualifications + +The skill insertion shifted generated capability heading anchors. General PR CI detected stale `capabilities.yaml`; the paid workflow's narrower build and every exact-source admission still passed before provider access. Both canonical capability metadata inventories are regenerated after measurement, include the new reference where declared, and adds a credential-free admission test that rejects stale/missing manifests. These maintenance changes do not modify the frozen measured branches or the skill bytes used in this comparison. + +Draft [PR #14948](https://github.com/paperclipai/paperclip/pull/14948) keeps this repair and the tiny manual; draft [PR #14961](https://github.com/paperclipai/paperclip/pull/14961) measures native completion documentation separately. OpenCode original delivery remains unresolved, and existing Claude chat-memory and legacy ACP credential/receipt findings remain separate. No broad matrix rerun was launched, and merging remains user-controlled. diff --git a/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.json b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.json new file mode 100644 index 0000000000..f909e31edb --- /dev/null +++ b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.json @@ -0,0 +1,1418 @@ +{ + "schema": "paperclip.opencode-skill-routing-link-comparison/v1", + "status": "complete", + "candidateSha": "fe9dc1e3c518825242ed889ab9c8352986f8c2ed", + "baselineSha": "0d7ecfa96d72fba79b7f0a25052b42c0686c0488", + "pairedSummary": { + "newFailures": 0, + "newPasses": 0, + "unchangedPasses": 2, + "unchangedFailures": 0, + "pendingPairs": 0, + "limitation": "Both original cases have deficient handoff links outside their machine oracle; candidate uses incorrect PAP prefix, baseline has bare unprefixed slug path." + }, + "sourceProof": { + "identicalTrackedFiles": 8242, + "changedTrackedPaths": [ + "skills/paperclip/SKILL.md", + "skills/paperclip/references/issue-documents.md", + "tests/runner-e2e/opencode-skill-routing-comparison.json" + ], + "onlyProductionDifferences": [ + { + "path": "skills/paperclip/SKILL.md", + "candidateSha256": "6ca593effe18704f8014ac4e8155f7513c0701aa3c61d3cf0b96d3366c92f0e1", + "baselineSha256": "768522e081c1ec43dee4a5794539f7b597ee3708c634d05e388c874df283cc53" + }, + { + "path": "skills/paperclip/references/issue-documents.md", + "candidateSha256": "a6141e8ba4510b683418197d58598d5a14db2641c49c869ae8aca4e2b05aca97", + "baselineSha256": "353845a406b1b4948487c348c86eb7b8c89a55ae67afe203100eee10aee46336" + } + ], + "baselineReceipt": { + "schema": "paperclip.opencode-skill-routing.baseline/v1", + "candidateSha": "fe9dc1e3c518825242ed889ab9c8352986f8c2ed", + "priorSkillSha": "b85668c4065d66e08026218a6bd48a3d22ddd4c9", + "expectedCellsPerVariant": 2, + "expectedProviderTurnsPerVariant": 2, + "heldConstant": [ + "tiny manual/shared prompts", + "native Codex additive base fix", + "future explicit clickable UI-link fixture and calibrated grader", + "models/effort/auth/permissions/tools/budgets", + "all other tracked repository source bytes" + ], + "onlyChangedProductionSources": [ + { + "path": "skills/paperclip/SKILL.md", + "candidateSha256": "6ca593effe18704f8014ac4e8155f7513c0701aa3c61d3cf0b96d3366c92f0e1", + "baselineSha256": "768522e081c1ec43dee4a5794539f7b597ee3708c634d05e388c874df283cc53" + }, + { + "path": "skills/paperclip/references/issue-documents.md", + "candidateSha256": "a6141e8ba4510b683418197d58598d5a14db2641c49c869ae8aca4e2b05aca97", + "baselineSha256": "353845a406b1b4948487c348c86eb7b8c89a55ae67afe203100eee10aee46336" + } + ] + }, + "fixtureConfigurations": { + "schema": "paperclip.opencode-skill-routing.fixture-match-proof/v1", + "cells": [ + { + "executionId": "stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "dd673f469cab583d26bc2e4b484f18b27884ec14a59542177026389181e2da2c", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + }, + { + "executionId": "stock-harness.legacy-opencode.local.assigned-skill-paperclip-document", + "fixtureConfigSha256": "4924cbbcb8768fdbff70d9f53c20835dba28626278da77fb5edcc09c1700fd5f", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + } + ] + } + }, + "cells": [ + { + "variant": "baseline", + "caseId": "assigned-skill-explicit-invocation", + "source": { + "sha": "0d7ecfa96d72fba79b7f0a25052b42c0686c0488", + "ref": "refs/heads/codex/opencode-pre-routing-skill", + "workflowRunUrl": "https://github.com/paperclipai/paperclip/actions/runs/37069552374" + }, + "status": "passed", + "attempt": 1, + "suiteDefinitionHash": "18a8eb87f11ebaa4d5be89f878de669bc32eee6f58a47c3f19d15bb879519dd8", + "model": "openrouter/deepseek/deepseek-v4-flash-0731", + "runtimeMode": "legacy", + "durationMs": 46285, + "providerDurationMs": 25870, + "runCount": 1, + "recoveryRuns": 0, + "cleanup": "passed", + "matcherResults": [ + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.assigned-skill-present", + "expected": true + }, + "passed": true, + "detail": "The task run must receive one pinned skill version through the public assignment state." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.skill-request-explicit", + "expected": true + }, + "passed": true, + "detail": "The task request must explicitly name the assigned skill with a supported slash or dollar reference." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.skill-source-marker", + "expected": true + }, + "passed": true, + "detail": "The assigned pinned skill source must contain the output marker." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.marker-not-in-request", + "expected": true + }, + "passed": true, + "detail": "The output marker must originate from the assigned skill, not the task request or comments." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.single-durable-output", + "expected": true + }, + "passed": true, + "detail": "Exactly one durable task document must contain the skill's marker." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.completed-task", + "expected": true + }, + "passed": true, + "detail": "Final task status: done." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.successful-runs", + "expected": true + }, + "passed": true, + "detail": "All recorded context-integrity runs must succeed." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.budget-hard-stops", + "expected": true + }, + "passed": true, + "detail": "Public company and agent records must both retain the 1,000-cent monthly hard stops." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.default-hire-bundle", + "expected": true + }, + "passed": true, + "detail": "The public managed bundle must contain only the eight-word shipped identity, without an injected QA manual." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.provider-runs-present", + "expected": true + }, + "passed": true, + "detail": "The lifecycle oracle must reach actual provider runs; missing runs cannot pass instruction delivery." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.invocation-evidence-complete", + "expected": true + }, + "passed": true, + "detail": "Every legacy run must have a public adapter.invoke event with its actual nonempty prompt." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.generic-procedures-absent", + "expected": true + }, + "passed": true, + "detail": "Neither startup nor continuation may reintroduce the removed generic manual." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.fresh-default-delivered", + "expected": true + }, + "passed": true, + "detail": "At least one fresh invocation must carry default identity and connection guidance." + } + ], + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 34011, + "outputTokens": 2533, + "cachedInputTokens": 192512, + "totalTokens": 229056, + "reportedCostUsd": 0.0058472545, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 25870, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.0058472545, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.0058472545, + "complete": true + }, + "preflight": { + "sourceSha": "0d7ecfa96d72fba79b7f0a25052b42c0686c0488", + "sourceFingerprint": "0315602ac6788cf0d65561a73cc85d790768ef90c6815257077a2381f961b96e", + "passed": true, + "providerCalls": 0, + "assertions": 587 + }, + "delivery": { + "documentCount": 1, + "documentKey": "task-document", + "savedRevisionNumber": 1, + "savedRevisionPresent": true, + "markerPresent": true, + "markdownLinkTargets": [], + "expectedDocumentUiTarget": "/RUN/issues/RUN-1#document-task-document", + "hasExactDocumentUiLink": false, + "originalCaseRequiresLink": false, + "originalBaselineBarePath": true, + "originalCandidateWrongPrefix": false + }, + "operationalSkillReadOrder": [ + { + "tool": "skill", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 1, + "skill": "assigned-output" + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 2, + "skill": null + }, + { + "tool": "write", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": true, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 3, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 4, + "skill": null + }, + { + "tool": "skill", + "operationalSkillLoad": true, + "earlyRecipeVisible": true, + "newReferencePointerVisible": true, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 5, + "skill": "paperclip" + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 6, + "skill": null + }, + { + "tool": "read", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": true, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 7, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": true, + "statusApiCall": false, + "ordinal": 8, + "skill": null + }, + { + "tool": "read", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 9, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 10, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": true, + "ordinal": 11, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 12, + "skill": null + } + ], + "evidenceHashes": { + "result": "a4953654cba1d6bd3911ad3cfb000d8bfb21041ef3e3ee409e2c19d14af30f1e", + "context": "c3fb84a16e6fe1dead74ffece90035bc99aa247ffa65b4665d5288acbe8a5515", + "preflight": "a0756b19a5c20c1bf600a9abeb1b62307c99f37ffc2ff02063709a870bbe2986", + "finalStateScreenshot": "49068a59e35b2fb0207bd81ca36d4db639d17c8639c0ce885ac5869fc83320e9" + }, + "screenshotInspected": true + }, + { + "variant": "baseline", + "caseId": "assigned-skill-paperclip-document", + "source": { + "sha": "0d7ecfa96d72fba79b7f0a25052b42c0686c0488", + "ref": "refs/heads/codex/opencode-pre-routing-skill", + "workflowRunUrl": "https://github.com/paperclipai/paperclip/actions/runs/37069552374" + }, + "status": "passed", + "attempt": 1, + "suiteDefinitionHash": "18a8eb87f11ebaa4d5be89f878de669bc32eee6f58a47c3f19d15bb879519dd8", + "model": "openrouter/deepseek/deepseek-v4-flash-0731", + "runtimeMode": "legacy", + "durationMs": 94804, + "providerDurationMs": 73984, + "runCount": 1, + "recoveryRuns": 0, + "cleanup": "passed", + "matcherResults": [ + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.assigned-skill-present", + "expected": true + }, + "passed": true, + "detail": "The task run must receive one pinned skill version through the public assignment state." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.skill-request-explicit", + "expected": true + }, + "passed": true, + "detail": "The task request must explicitly name the assigned skill with a supported slash or dollar reference." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.skill-source-marker", + "expected": true + }, + "passed": true, + "detail": "The assigned pinned skill source must contain the output marker." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.marker-not-in-request", + "expected": true + }, + "passed": true, + "detail": "The output marker must originate from the assigned skill, not the task request or comments." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.single-durable-output", + "expected": true + }, + "passed": true, + "detail": "Exactly one durable task document must contain the skill's marker." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.saved-document-revision", + "expected": true + }, + "passed": true, + "detail": "The public saved document must have a persisted revision and the requested content." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.saved-document-link", + "expected": true + }, + "passed": true, + "detail": "An agent completion comment must link the exact saved document on this task; a local path or claimed URL does not suffice." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.completed-task", + "expected": true + }, + "passed": true, + "detail": "Final task status: done." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.successful-runs", + "expected": true + }, + "passed": true, + "detail": "All recorded context-integrity runs must succeed." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.budget-hard-stops", + "expected": true + }, + "passed": true, + "detail": "Public company and agent records must both retain the 1,000-cent monthly hard stops." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.default-hire-bundle", + "expected": true + }, + "passed": true, + "detail": "The public managed bundle must contain only the eight-word shipped identity, without an injected QA manual." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.provider-runs-present", + "expected": true + }, + "passed": true, + "detail": "The lifecycle oracle must reach actual provider runs; missing runs cannot pass instruction delivery." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.invocation-evidence-complete", + "expected": true + }, + "passed": true, + "detail": "Every legacy run must have a public adapter.invoke event with its actual nonempty prompt." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.generic-procedures-absent", + "expected": true + }, + "passed": true, + "detail": "Neither startup nor continuation may reintroduce the removed generic manual." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.fresh-default-delivered", + "expected": true + }, + "passed": true, + "detail": "At least one fresh invocation must carry default identity and connection guidance." + } + ], + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 62834, + "outputTokens": 2127, + "cachedInputTokens": 130304, + "totalTokens": 195265, + "reportedCostUsd": 0.004943647, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 73984, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.004943647, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.004943647, + "complete": true + }, + "preflight": { + "sourceSha": "0d7ecfa96d72fba79b7f0a25052b42c0686c0488", + "sourceFingerprint": "0315602ac6788cf0d65561a73cc85d790768ef90c6815257077a2381f961b96e", + "passed": true, + "providerCalls": 0, + "assertions": 587 + }, + "delivery": { + "documentCount": 1, + "documentKey": "context-integrity-output", + "savedRevisionNumber": 1, + "savedRevisionPresent": true, + "markerPresent": true, + "markdownLinkTargets": [ + "/RUN/issues/RUN-1#document-context-integrity-output" + ], + "expectedDocumentUiTarget": "/RUN/issues/RUN-1#document-context-integrity-output", + "hasExactDocumentUiLink": true, + "originalCaseRequiresLink": true, + "originalBaselineBarePath": false, + "originalCandidateWrongPrefix": false + }, + "operationalSkillReadOrder": [ + { + "tool": "skill", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 1, + "skill": "assigned-output" + }, + { + "tool": "skill", + "operationalSkillLoad": true, + "earlyRecipeVisible": true, + "newReferencePointerVisible": true, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 2, + "skill": "paperclip" + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 3, + "skill": null + }, + { + "tool": "read", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": true, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 4, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 5, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 6, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": true, + "statusApiCall": false, + "ordinal": 7, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": true, + "ordinal": 8, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 9, + "skill": null + } + ], + "evidenceHashes": { + "result": "b09f82a96976fdcd5ddfb1655d783db5076a19fde0a5d751ff3df995cd5fac59", + "context": "2c630300b0f5737c74a9921d25ed67d2ab2612957f843d9709ebfd57fb84306f", + "preflight": "9e6e55c1ca71312f7bb53708bd2f9727e5395c9da8586b285d29e1fee4fe50fe", + "finalStateScreenshot": "128e6e0c7be8b8ca0e3e90601bce8c3d33df8c503636d09ffe514df03881173b" + }, + "screenshotInspected": true + }, + { + "variant": "candidate", + "caseId": "assigned-skill-explicit-invocation", + "source": { + "sha": "fe9dc1e3c518825242ed889ab9c8352986f8c2ed", + "ref": "refs/heads/codex/opencode-stock-routing-qualified", + "workflowRunUrl": "https://github.com/paperclipai/paperclip/actions/runs/37069547401" + }, + "status": "passed", + "attempt": 1, + "suiteDefinitionHash": "62cfca8303c66bf5e29b0665c72b7bb717acf9f96d7a3ef768407afa5bc9870f", + "model": "openrouter/deepseek/deepseek-v4-flash-0731", + "runtimeMode": "legacy", + "durationMs": 89455, + "providerDurationMs": 67897, + "runCount": 1, + "recoveryRuns": 0, + "cleanup": "passed", + "matcherResults": [ + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.assigned-skill-present", + "expected": true + }, + "passed": true, + "detail": "The task run must receive one pinned skill version through the public assignment state." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.skill-request-explicit", + "expected": true + }, + "passed": true, + "detail": "The task request must explicitly name the assigned skill with a supported slash or dollar reference." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.skill-source-marker", + "expected": true + }, + "passed": true, + "detail": "The assigned pinned skill source must contain the output marker." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.marker-not-in-request", + "expected": true + }, + "passed": true, + "detail": "The output marker must originate from the assigned skill, not the task request or comments." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.single-durable-output", + "expected": true + }, + "passed": true, + "detail": "Exactly one durable task document must contain the skill's marker." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.completed-task", + "expected": true + }, + "passed": true, + "detail": "Final task status: done." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.successful-runs", + "expected": true + }, + "passed": true, + "detail": "All recorded context-integrity runs must succeed." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.budget-hard-stops", + "expected": true + }, + "passed": true, + "detail": "Public company and agent records must both retain the 1,000-cent monthly hard stops." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.default-hire-bundle", + "expected": true + }, + "passed": true, + "detail": "The public managed bundle must contain only the eight-word shipped identity, without an injected QA manual." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.provider-runs-present", + "expected": true + }, + "passed": true, + "detail": "The lifecycle oracle must reach actual provider runs; missing runs cannot pass instruction delivery." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.invocation-evidence-complete", + "expected": true + }, + "passed": true, + "detail": "Every legacy run must have a public adapter.invoke event with its actual nonempty prompt." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.generic-procedures-absent", + "expected": true + }, + "passed": true, + "detail": "Neither startup nor continuation may reintroduce the removed generic manual." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.fresh-default-delivered", + "expected": true + }, + "passed": true, + "detail": "At least one fresh invocation must carry default identity and connection guidance." + } + ], + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 71068, + "outputTokens": 1716, + "cachedInputTokens": 97792, + "totalTokens": 170576, + "reportedCostUsd": 0.00413837, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 67897, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.00413837, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.00413837, + "complete": true + }, + "preflight": { + "sourceSha": "fe9dc1e3c518825242ed889ab9c8352986f8c2ed", + "sourceFingerprint": "aeb792549146cc29eda779e83ae22d336276e31162f738a1dde85a82e7833902", + "passed": true, + "providerCalls": 0, + "assertions": 587 + }, + "delivery": { + "documentCount": 1, + "documentKey": "context-integrity-output", + "savedRevisionNumber": 1, + "savedRevisionPresent": true, + "markerPresent": true, + "markdownLinkTargets": [ + "/PAP/issues/RUN-1#document-context-integrity-output" + ], + "expectedDocumentUiTarget": "/RUN/issues/RUN-1#document-context-integrity-output", + "hasExactDocumentUiLink": false, + "originalCaseRequiresLink": false, + "originalBaselineBarePath": false, + "originalCandidateWrongPrefix": true + }, + "operationalSkillReadOrder": [ + { + "tool": "skill", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 1, + "skill": "assigned-output" + }, + { + "tool": "skill", + "operationalSkillLoad": true, + "earlyRecipeVisible": true, + "newReferencePointerVisible": true, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 2, + "skill": "paperclip" + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 3, + "skill": null + }, + { + "tool": "read", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": true, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 4, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": true, + "statusApiCall": false, + "ordinal": 5, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 6, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 7, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": true, + "ordinal": 8, + "skill": null + } + ], + "evidenceHashes": { + "result": "7f13c7f84dbbd662d5689b82ef02d5e9806dc43237604599c183ba807097f403", + "context": "16aa44354be2763149fe4c98d5907230025f909ec2d7eafa93767959c183f6a1", + "preflight": "cd5c451ec7c01ac145b8f20a11facfe4410b6c6ea621358b0d5bfcaa339260cf", + "finalStateScreenshot": "123a366f1d9372bf5249e9a46d16dc322538f7a0c2669504f02a29d97b77cf16" + }, + "screenshotInspected": true + }, + { + "variant": "candidate", + "caseId": "assigned-skill-paperclip-document", + "source": { + "sha": "fe9dc1e3c518825242ed889ab9c8352986f8c2ed", + "ref": "refs/heads/codex/opencode-stock-routing-qualified", + "workflowRunUrl": "https://github.com/paperclipai/paperclip/actions/runs/37069547401" + }, + "status": "passed", + "attempt": 1, + "suiteDefinitionHash": "62cfca8303c66bf5e29b0665c72b7bb717acf9f96d7a3ef768407afa5bc9870f", + "model": "openrouter/deepseek/deepseek-v4-flash-0731", + "runtimeMode": "legacy", + "durationMs": 102901, + "providerDurationMs": 82751, + "runCount": 1, + "recoveryRuns": 0, + "cleanup": "passed", + "matcherResults": [ + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.assigned-skill-present", + "expected": true + }, + "passed": true, + "detail": "The task run must receive one pinned skill version through the public assignment state." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.skill-request-explicit", + "expected": true + }, + "passed": true, + "detail": "The task request must explicitly name the assigned skill with a supported slash or dollar reference." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.skill-source-marker", + "expected": true + }, + "passed": true, + "detail": "The assigned pinned skill source must contain the output marker." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.marker-not-in-request", + "expected": true + }, + "passed": true, + "detail": "The output marker must originate from the assigned skill, not the task request or comments." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.single-durable-output", + "expected": true + }, + "passed": true, + "detail": "Exactly one durable task document must contain the skill's marker." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.saved-document-revision", + "expected": true + }, + "passed": true, + "detail": "The public saved document must have a persisted revision and the requested content." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.saved-document-link", + "expected": true + }, + "passed": true, + "detail": "An agent completion comment must link the exact saved document on this task; a local path or claimed URL does not suffice." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.completed-task", + "expected": true + }, + "passed": true, + "detail": "Final task status: done." + }, + { + "matcher": { + "kind": "json_path", + "path": "contextIntegrity.successful-runs", + "expected": true + }, + "passed": true, + "detail": "All recorded context-integrity runs must succeed." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.budget-hard-stops", + "expected": true + }, + "passed": true, + "detail": "Public company and agent records must both retain the 1,000-cent monthly hard stops." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.default-hire-bundle", + "expected": true + }, + "passed": true, + "detail": "The public managed bundle must contain only the eight-word shipped identity, without an injected QA manual." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.provider-runs-present", + "expected": true + }, + "passed": true, + "detail": "The lifecycle oracle must reach actual provider runs; missing runs cannot pass instruction delivery." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.invocation-evidence-complete", + "expected": true + }, + "passed": true, + "detail": "Every legacy run must have a public adapter.invoke event with its actual nonempty prompt." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.generic-procedures-absent", + "expected": true + }, + "passed": true, + "detail": "Neither startup nor continuation may reintroduce the removed generic manual." + }, + { + "matcher": { + "kind": "json_path", + "path": "stockHarness.fresh-default-delivered", + "expected": true + }, + "passed": true, + "detail": "At least one fresh invocation must carry default identity and connection guidance." + } + ], + "billing": { + "llm": { + "runCount": 1, + "runsWithTokenUsage": 1, + "runsWithReportedCost": 1, + "inputTokens": 92370, + "outputTokens": 2691, + "cachedInputTokens": 185856, + "totalTokens": 280917, + "reportedCostUsd": 0.006644079, + "costStatus": "reported" + }, + "runtime": { + "provider": "local", + "agentRunDurationMs": 82751, + "leaseDurationMs": null, + "leaseCount": 0, + "costStatus": "not_metered", + "costSource": "local_not_metered" + }, + "reportedCostUsd": 0.006644079, + "estimatedRuntimeCostUsd": 0, + "observedAndEstimatedCostUsd": 0.006644079, + "complete": true + }, + "preflight": { + "sourceSha": "fe9dc1e3c518825242ed889ab9c8352986f8c2ed", + "sourceFingerprint": "aeb792549146cc29eda779e83ae22d336276e31162f738a1dde85a82e7833902", + "passed": true, + "providerCalls": 0, + "assertions": 587 + }, + "delivery": { + "documentCount": 1, + "documentKey": "context-integrity-output", + "savedRevisionNumber": 1, + "savedRevisionPresent": true, + "markerPresent": true, + "markdownLinkTargets": [ + "/RUN/issues/RUN-1#document-context-integrity-output" + ], + "expectedDocumentUiTarget": "/RUN/issues/RUN-1#document-context-integrity-output", + "hasExactDocumentUiLink": true, + "originalCaseRequiresLink": true, + "originalBaselineBarePath": false, + "originalCandidateWrongPrefix": false + }, + "operationalSkillReadOrder": [ + { + "tool": "skill", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 1, + "skill": "assigned-output" + }, + { + "tool": "skill", + "operationalSkillLoad": true, + "earlyRecipeVisible": true, + "newReferencePointerVisible": true, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 2, + "skill": "paperclip" + }, + { + "tool": "read", + "operationalSkillLoad": true, + "earlyRecipeVisible": true, + "newReferencePointerVisible": true, + "newReferenceRead": false, + "outputTruncated": true, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 3, + "skill": null + }, + { + "tool": "read", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": true, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 4, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 5, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 6, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 7, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 8, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": true, + "statusApiCall": false, + "ordinal": 9, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": false, + "ordinal": 10, + "skill": null + }, + { + "tool": "bash", + "operationalSkillLoad": false, + "earlyRecipeVisible": false, + "newReferencePointerVisible": false, + "newReferenceRead": false, + "outputTruncated": false, + "localDocumentWrite": false, + "documentApiCall": false, + "statusApiCall": true, + "ordinal": 11, + "skill": null + } + ], + "evidenceHashes": { + "result": "d363c4c394629706e30959682ad9b5ec301330636ec5f8c0dab319082ac3d8e8", + "context": "208a34086b904864fb9d115a9bae910ed49256b812b9e2ba9688459594921530", + "preflight": "08763f3fda4e6c60a23986ebdbba4973e6c10565575593e5797fd24ec8330d89", + "finalStateScreenshot": "982e5645377a39a6aa853795229d6b3e03afe00d567e03bd90a695c0f9115169" + }, + "screenshotInspected": true + } + ], + "providerCalls": 4, + "additionalPaidRuns": 0, + "limitations": [ + "One trial per cell cannot establish fix causality, robust non-regression or general coding quality.", + "Both original case link defects are outside the preserved durable-document oracle.", + "Explicit wording now requires clickable company-prefixed UI link and is not a direct repeat of prior usable-link cohorts.", + "Skill metadata improves stock selection but does not guarantee invocation on every harness.", + "API writes/revisions/comments and screenshot rendering were inspected; authenticated click navigation against cleaned-up instances was not replayed.", + "Runtime is unmetered; reported LLM receipts do not establish actual external bill.", + "Historical combined prompt-removal and early-recipe failures are preserved, alongside unrelated chat/credential findings." + ] +} diff --git a/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md new file mode 100644 index 0000000000..79f0c4b519 --- /dev/null +++ b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md @@ -0,0 +1,75 @@ +# OpenCode skill routing and clickable delivery — 2026-10-02 + +**TL;DR: 0 newly failing cells, 0 newly passing cells, 2 unchanged passes, and 0 pending pairs in this matched trial. Both document cases pass in both variants, so this does not establish that the metadata/link change caused the earlier failures to recover. Handoff remains imperfect in the original case:** candidate uses a clickable link with the wrong `PAP` company prefix; baseline supplies only a bare prefix-less slug path. Neither defect is scored by the preserved original document-storage oracle. The explicit clickable-link cases deliver the correct `RUN` link in both variants. + +The narrow correction expands the operational skill's stock discovery description to include Paperclip task/heartbeat work and document/file delivery. The early API recipe asks for a clickable Markdown link, and the reference shows a canonical UI-link example. The full skill stays on disk; the eight-word manual and shared prompts remain reduced. No full-body injection, adapter policy override or native tool guidance is added to legacy runs. + +| Classic OpenCode case | Pre-correction skill | Corrected skill | Independent evidence and limitation | +| --- | --- | --- | --- | +| Original assigned skill, verbatim request | Pass | Pass | Exactly one saved document with skill-only marker and persisted revision in each. Both handoff paths are deficient and outside this case's link-free oracle. | +| Explicit Paperclip storage with clarified clickable UI link | Pass | Pass | Saved revision/content and exact clickable `/RUN/issues/RUN-1#document-context-integrity-output` in both. | + +Original machine results, source/configuration proof, tool-read sequence, costs and evidence hashes are preserved in the [safe evidence projection](2026-10-02-opencode-skill-routing-link-qualification.json). All four cells clean up successfully; none uses automatic disposition recovery or an extra attempt. No provider was rerun to improve these results. + +## Sources, admission and inspectable reports + +- Candidate `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`, frozen branch `codex/opencode-stock-routing-qualified`: [workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401), [published dashboard and screenshots](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37069547401-1/index.html). +- Baseline `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`, frozen branch `codex/opencode-pre-routing-skill`: [workflow](https://github.com/paperclipai/paperclip/actions/runs/37069552374), [published dashboard and screenshots](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37069552374-1/index.html). +- Trusted default-branch workflows ran concurrently on separate targets. Each of four measured cells passed its own exact-source, credential-free 587-assertion prerequisite before provider credentials. Candidate source fingerprint is `aeb792549146cc29eda779e83ae22d336276e31162f738a1dde85a82e7833902`; baseline is `0315602ac6788cf0d65561a73cc85d790768ef90c6815257077a2381f961b96e`. +- Git tree comparison verifies 8,242 identical tracked files. Only the two operational skill sources and an explicit baseline provenance receipt differ. Models, effort, tools, auth, permissions, budgets, manual/shared prompts and behavioral fixtures/graders are fixed; both fixture configuration hashes match. Suite/source digest differences correctly include the changed skill bytes. +- Local source preparation passed 918 E2E support tests, eight canonical inventory calibrations, skill validation, E2E typecheck, two-cell discovery and both exact-source prerequisites. Both capability inventories use the same declared source list. Initial setup mistakes remain retained and did not reach providers. + +The explicit fixture now asks for a clickable company-prefixed Paperclip UI document link, matching its intended oracle. Bare paths and code-formatted paths fail calibration; correct relative/same-app absolute links pass. The original assigned-skill request and procedure remain verbatim. Earlier API-link grades are not retroactively changed, and these explicit results are not a direct repeat of the older “usable link” cohort. + +## Which guidance was read before delivery + +In the candidate original assignment, OpenCode loads the assigned Context integrity output skill, then operational Paperclip before any document write. The returned Paperclip skill is marked truncated, but the early saved-document recipe and reference pointer are visible. It reads `issue-documents.md`, then writes the public task document through the API. This differs from the previously failing candidate, which first loaded Paperclip only in a separate disposition-recovery run after its local-only output. + +The current baseline original assignment first loads the assigned output skill and writes a workspace Markdown file. It later loads Paperclip and reads the issue-document reference, then saves a public document before completing the **same** assignment. Thus its successful delivery did not require improved metadata; stock selection timing varies between trials. Both explicit variants load Paperclip and read the reference before the document API call. Four rendered final-state screenshots were inspected, and public document revisions/content and completion comments agree with the retained results. + +The candidate original comment has an actual Markdown href `/PAP/issues/RUN-1#document-context-integrity-output`, despite the measured company prefix being `RUN`. Baseline's original comment has a bare `/issues/runner-e2e-…#document-task-document` path without a Markdown anchor. The candidate's literal prefix matches the reference's example, making example copying a plausible cause, but this trial does not isolate that cause. Authenticated click navigation was not replayed after cleanup; neither original link is presented as a verified usable handoff. A minimal follow-up can show an issue-derived prefix instead of a literal company example. These original machine passes must not be presented as fully correct links. + +## Timing, usage and cost + +Four provider turns were expected and four assignment runs occurred. All four ledger receipts contain token usage and reported LLM cost. Local runtime remains unmetered; actual external billing is not established. + +| Case | Baseline provider seconds | Candidate provider seconds | Baseline cell seconds | Candidate cell seconds | Baseline reported USD | Candidate reported USD | +| --- | ---: | ---: | ---: | ---: | ---: | ---: | +| Original | 25.870 | 67.897 | 46.285 | 89.455 | 0.0058472545 | 0.0041383700 | +| Explicit clickable storage | 73.984 | 82.751 | 94.804 | 102.901 | 0.0049436470 | 0.0066440790 | + +Reported LLM totals are **$0.0107909015 baseline** and **$0.0107824490 candidate**. The candidate original case takes longer in this trial; these single observations, different cache/token receipts and unmetered runtime do not establish general speed, cost or quality equivalence. + +## Preserved failures and scope + +The [original combined prompt-removal report](2026-10-02-stock-harness-live-comparison.md) still records two new classic Claude/OpenCode document-delivery failures, two OpenCode ordered-case improvements, and separate unchanged credential/chat failures. The [first recipe repair report](2026-10-02-legacy-document-skill-repair.md) still records Claude's improvement, OpenCode's unresolved local-only output, and the candidate's worse non-clickable explicit handoff. New successes do not erase those earlier observations or establish robustness. + +This comparison qualifies only classic OpenCode's two selected journeys. Existing classic Codex/Claude/OpenCode and ACP Codex skill mounts were audited; ACP Claude's names/root-only metadata and unsupported custom ACP delivery remain separate. Hermes/Pi/OpenClaw are not live-qualified here. Native completion descriptions and hiring templates have separate changes and measurements. Both PRs remain draft; no merge or further paid breadth is part of this report. + +## Provider-free follow-up after measurement + +The approved reference-only follow-up replaces the literal `PAP` link with a +JavaScript construction using the current `issue.identifier` and the successful +write receipt's `saved.key`. It is a later source change, **not live-qualified** +by the frozen candidate above. The independent storage grader now calibrates a +wrong-company link as failure, and executes the shipped recipe against two +other company prefixes plus a locked-write redirected document key. All 25 +focused document/source-digest tests, eight inventory calibrations, canonical +generated metadata checks and E2E typecheck pass. An initial sandboxed support attempt denied localhost/tsx +pipe creation; the failed attempt is retained, with its affected files rerun +unchanged under permitted local execution: 906 assertions passed in the original +921-assertion attempt, and all 82 selected assertions passed in the permitted +retry, including every originally failed assertion. No new model calls or old verdict +changes are part of this follow-up. + +Fresh native-PR review caught a supported edge case in that later recipe: +`issue.identifier` may be null. The recipe now uses the current issue ID in the +unprefixed UI route when no identifier exists. Provider-free tests cover both +null and absent identifiers, preserving the returned document key. Source review +confirms that the board resolves the loaded issue's actual company and preserves +the document hash; no agent-side company fetch is needed. This route logic also +corrects wrong prefixes, so the frozen candidate's `PAP` href is noncanonical, +not a proven broken link. No authenticated click replay or further model run was +made. The previous 4/5 review and local setup/timing failures remain retained; +current focused document/source/manifest/admission checks pass 56 assertions and +E2E typecheck passes. This correction still has no live qualification. diff --git a/doc/plans/2026-10-02-stock-harness-live-comparison.json b/doc/plans/2026-10-02-stock-harness-live-comparison.json new file mode 100644 index 0000000000..7512c94760 --- /dev/null +++ b/doc/plans/2026-10-02-stock-harness-live-comparison.json @@ -0,0 +1,357 @@ +{ + "schema": "paperclip.stock-harness.safe-comparison/v1", + "candidateSha": "f02d8d0df327abb43b20c7e7beb86798239abbf5", + "baselineSha": "12c5433c67dc2e62916b879349c7ba2b6e0431f0", + "nativeCodexFixHeldConstant": "408f70e69f9c5e49cb4377f4886ac2001bfa67a2", + "comparisonManifest": { + "schema": "paperclip.stock-harness-comparison.v1", + "variant": "previous-instructions", + "candidateSha": "f02d8d0df327abb43b20c7e7beb86798239abbf5", + "restoredInstructionsFrom": "e00d10d5d5594f6e2d1e8cf4e0ec81c074b0bd88", + "nativeCodex14920": "held-constant; no before/after performance claim", + "productionDifferences": [ + "server/src/onboarding-assets/default/AGENTS.md", + "packages/adapter-utils/src/server-utils.ts" + ], + "structuralOracle": "Historical manual exact content and observed generic procedures; independent behavioral graders unchanged", + "expectedMatrixCells": 24, + "expectedProviderTurns": 48, + "maxInitialCampaignsPerVariant": 1, + "behavioralSources": { + "context-integrity-cases.ts": "201d05f8257c1121a94dd9d8a22f08b9db80555e388013ffeabb65e301499fd2", + "context-integrity-scoring.ts": "50d370044e45386ba87fb23ca63c228b96bb8373d67bb7e82e8b2a89eb23f252", + "context-integrity-flow.ts": "63bcb7bdba4c80b15ff5872f16383844456eab0c5a10627da5d3487a26821c6a", + "chat-cases.ts": "754b251bdd42605e147da9081b663329234260d2b13df0df3d44562e362b14ca", + "chat-flow.ts": "11fc506559da61315ce04a877f6b38a055e71e532fecfcf62c749a1d3bf985e7", + "live-fixtures.ts": "1ef80209cc53cc514a8e58af4c2225e38ec15b20acf17ca7564506ca6617ac3e", + "runner.spec.ts": "14d57e0f9131b3f1981ac593464738d01362008eefe85ea836075c794681fa85", + "launch.ts": "5ada6ca5484542f028b29fc603dc7992213d4c06fa9046834bd7d23c4511b847", + "stock-harness-admission.ts": "c9775d77d48f8d59ba1197a7e5769f14db0c7d33f5e2a2968a04466e083d63ff", + "types.ts": "c45dc3b2366457449ec9e7d413f930605dec88ea93a14ce5c5a5f445a27d148c" + }, + "instructionSha256": { + "manual": "e4d2375d722602cd744403292d99f6e6c9b7e8014d9cdfa6f9811ace03428e5f", + "shared": "819026782b92f37064fca354ddbca685b3c05f1869ec560e40d8a757b0a8aee9" + } + }, + "modelsEffortToolsPermissionsCredentials": "same fixture/source configuration; effort inherits provider default when unspecified", + "expectedCellsPerVariant": 24, + "expectedProviderTurnsPerVariant": 48, + "baseline": [ + {"executionId":"stock-harness.legacy-acp-claude.local.assigned-skill-explicit-invocation","profile":"legacy-acp-claude","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":59075,"providerDurationMs":42037,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":15184,"outputTokens":2261,"cachedInputTokens":218341,"totalTokens":235786,"reportedCostUsd":0.16114005000000003,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":42037,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.16114005000000003,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.16114005000000003,"complete":true},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":14137,"reportedPromptChars":13607,"freshPromptChars":4688,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"78f6add6f11a45f58bfc6fee81600871f888750986c27497563a79eba39a101c","evidence-manifest.json":"14f67d036a3663f606cf269890f5a36a18221e5c7dcb466f513f761d882fb6eb","snapshots/context-integrity.json":"93bba435c845af31707d4e4b84b2330cd8ff14045a76c2d1267b958c8133950b","snapshots/stock-harness-hire.json":"e8b85aa909fc6af6a72b180c79572d70a26dd6ee6307c9b985f80b928eb96170","snapshots/stock-harness.json":"fa0eb10eb3939b91d83fed8a8c9f9915bb407ab03a5a9bd7e5df99aa593ffb99","snapshots/stock-harness-preflight.json":"2ac069369e2d26b0e4c213966770962c0282f9a7665fa1297fe3c70df2c73960"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.legacy-acp-claude.local.continuity-restart","profile":"legacy-acp-claude","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["stockHarness.fresh-default-delivered"],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":36961,"providerDurationMs":19492,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":19069,"outputTokens":329,"cachedInputTokens":55518,"totalTokens":74916,"reportedCostUsd":0.1039979,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":19492,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.1039979,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.1039979,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":false},"invocationMetrics":[{"capturedPromptChars":16395,"reportedPromptChars":17606,"freshPromptChars":1462,"retrievalClipped":true},{"capturedPromptChars":16395,"reportedPromptChars":17638,"freshPromptChars":1462,"retrievalClipped":true}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"b00296e015e0d35a213e031aede08732f900cf8aaa0f67dd6adc30310871a9f4","evidence-manifest.json":"101cb2a051bb64d3fc12e3a5389c2f39be4acb85e120b9bca1557163643f32e8","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"df3b7ce86f99b1f31ce3061e637e936d40a68b63d49aa307cf4a9dcfa75b3983","snapshots/stock-harness.json":"cffaf9caec42246841497abf2059ee2d11c23a67971ad14112fa96beb23d302a","snapshots/stock-harness-preflight.json":"7da7c7be4d0c33c6d1c3fadeb500eebe3fceba797edd0f24a9db187091896b6b"}}, + {"executionId":"stock-harness.legacy-acp-claude.local.ordered-comment-continuation","profile":"legacy-acp-claude","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["stockHarness.fresh-default-delivered"],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":169818,"providerDurationMs":150645,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":49679,"outputTokens":7557,"cachedInputTokens":465247,"totalTokens":522483,"reportedCostUsd":0.45116585000000003,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":150645,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.45116585000000003,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.45116585000000003,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":false},"invocationMetrics":[{"capturedPromptChars":16433,"reportedPromptChars":21108,"freshPromptChars":4688,"retrievalClipped":true},{"capturedPromptChars":14499,"reportedPromptChars":14021,"freshPromptChars":4688,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"9994318ef5b3cac9143215d580f7551600777864f3d54a0cb2c48941a0799523","evidence-manifest.json":"79ba61a1c2ac3adde8d1905f09a051d0b711be7c618285b728d13f3650fe9d64","snapshots/context-integrity.json":"0566e89e716c742ee4452f5d9bd719354dba968459809e3e362794f55135b705","snapshots/stock-harness-hire.json":"05919c3be635604a40c0ea9bb76ab1ce1a1ce6545417e66e02c5a26881f082ee","snapshots/stock-harness.json":"6fdbee98a58b20e131a46a5bf9220c1d9e6c4cfc02d64baeee700627869b137c","snapshots/stock-harness-preflight.json":"cb4abccc23ebbe0cd4884ca1d5238eef14b82e15987a2d9e7083ccb608504189"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.legacy-acp-codex.local.assigned-skill-explicit-invocation","profile":"legacy-acp-codex","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":59328,"providerDurationMs":35548,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":511,"outputTokens":30,"cachedInputTokens":27646,"totalTokens":28187,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":35548,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":13618,"reportedPromptChars":13642,"freshPromptChars":4687,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"743b4b283ee73f34db4434facefb26012145fa528a1912856725c4f13e5b9247","evidence-manifest.json":"fe7f46cac1d8722c6137ec813638ecd0c5944f1ea25cc24092ff8d22236c807b","snapshots/context-integrity.json":"d247901367c17626a5738743dc300566c758db50fa7c0b60e5f74be5b7083245","snapshots/stock-harness-hire.json":"15866eaa21a98137bd2c4088dd53256360cdf65b27c8475d49d98cd49788fcf6","snapshots/stock-harness.json":"639cd4b7fec29383a18016e5b8be2227f68d8a5c37035318089ace252408a231","snapshots/stock-harness-preflight.json":"f52edc51ee627e9f5f02d3049a83a9b99fc073fc504ca9f238cf95bc691a1ce8"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.legacy-acp-codex.local.continuity-restart","profile":"legacy-acp-codex","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["stockHarness.invocation-evidence-complete","stockHarness.fresh-default-delivered"],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":44754,"providerDurationMs":22118,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":662,"outputTokens":29,"cachedInputTokens":29955,"totalTokens":30646,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":22118,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":false,"historical-generic-procedures-observed":true,"fresh-default-delivered":false},"invocationMetrics":[{"capturedPromptChars":16383,"reportedPromptChars":17671,"freshPromptChars":1461,"retrievalClipped":true}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"3ba472d29268b3e2d2170cd9947d7c89fbe685d902fedacf5c74a1952a97ce3a","evidence-manifest.json":"a5c263cc9afe3836a3c9f46cda537bb07a374c30fd7ad6f136088680b3ade01c","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"8dd9d388cc265387d960fd435c19874601a5f7ca209b328943f8bacb145c5d89","snapshots/stock-harness.json":"f2b69dda7e26a76c952f0b36f10cfd7e9fb364ae3930585961d84093d17c2979","snapshots/stock-harness-preflight.json":"5fc3c8235143766fc6be8c7e12b31b92222594698c966462f944f13fb5929113"}}, + {"executionId":"stock-harness.legacy-acp-codex.local.ordered-comment-continuation","profile":"legacy-acp-codex","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["contextIntegrity.ordered-comments","contextIntegrity.changed-scope-preserved","contextIntegrity.packing-report-order","contextIntegrity.final-scope-applied","contextIntegrity.continuation-run-count","contextIntegrity.continuation-wake-comment-ids","contextIntegrity.single-durable-output","contextIntegrity.completed-task","contextIntegrity.successful-runs","stockHarness.invocation-evidence-complete"],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":50896,"providerDurationMs":39443,"cleanup":"passed","billing":{"llm":{"runCount":4,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":992,"outputTokens":46,"cachedInputTokens":24349,"totalTokens":25387,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":39443,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":false,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":14032,"reportedPromptChars":14056,"freshPromptChars":4687,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"33ac59f0a690ed76618f695e658575751e718f53f9d828d8b52c690df9c15afb","evidence-manifest.json":"92c048e01012a77b94003e38f3f73888d725f6942327f7fbf3e6d2a42d12184d","snapshots/context-integrity.json":"2f9a803657068a67bcbf3b190370395187ef13b7a1487705d0305e121339eac7","snapshots/stock-harness-hire.json":"f47bb35d606ae00e6057a2a77769e2091e8ccb12eca08db3a17e1ed297a8ccf4","snapshots/stock-harness.json":"05582c37956b55f6a2bd7cad0c55bc0fa1dbff875d7dbed12b83352f6bd4e04c","snapshots/stock-harness-preflight.json":"badf01a591f5327f0dc2e48865711f8f4d7327888a7688668728173a29b273a8"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":3,"failed":9}}, + {"executionId":"stock-harness.legacy-claude.local.assigned-skill-explicit-invocation","profile":"legacy-claude","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":49054,"providerDurationMs":29420,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":29883,"outputTokens":1348,"cachedInputTokens":225730,"totalTokens":256961,"reportedCostUsd":0.19999125,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":29420,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.19999125,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.19999125,"complete":true},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":7512,"reportedPromptChars":7512,"freshPromptChars":4684,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"4fc73f7b1226b9a41b209a2f6dfa2a58805ddc1a7e7b699ac2b42b2e74a189a7","evidence-manifest.json":"440ee14cca3f8558fb47ecf27b17693c0f9dbecdcf5c8532d7e2ccbefe31700b","snapshots/context-integrity.json":"4545847b9c4ab2a4fa3053a384b5e990133705fc9da0d1e9cd98f1275ec05f44","snapshots/stock-harness-hire.json":"a6abfc40d1c788e4b897f2375506d59b2caba9fab0d4df019e7137fe9f66177f","snapshots/stock-harness.json":"755cc9d5f656dc068a08c8c141203eaa5d3d60b53edd9d516887aae108119f45","snapshots/stock-harness-preflight.json":"f52730b4555ab31de952c8c1e38db4067fb40756b6befd36ed4b686e1a273d73"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.legacy-claude.local.continuity-restart","profile":"legacy-claude","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["chat_memory_assertion"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":40202,"providerDurationMs":23072,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":39036,"outputTokens":829,"cachedInputTokens":116171,"totalTokens":156036,"reportedCostUsd":0.1936638,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":23072,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.1936638,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.1936638,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":11451,"reportedPromptChars":11451,"freshPromptChars":1458,"retrievalClipped":false},{"capturedPromptChars":11508,"reportedPromptChars":11508,"freshPromptChars":1458,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"f8e377f617204eabdb6b4e0c24ab3280c51664b429051f7ea21a4746790820a6","evidence-manifest.json":"45818efcb8ed6d663f7ffab6d53605d5191fb983a68e422dadebdae875ed5b4c","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"e3a8735ad4d5f89dc3899668c36791579562157fa25aa32a6e59aa098d0c079a","snapshots/stock-harness.json":"f3d24452ade2f232b0f671768d46e4d65a6a0c2a0740bb1daa17a4871a09aa7b","snapshots/stock-harness-preflight.json":"b85c96ada9801ce6c6569338b68bc810494a0711f01af39a63932e288f7a81a3"}}, + {"executionId":"stock-harness.legacy-claude.local.ordered-comment-continuation","profile":"legacy-claude","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":142017,"providerDurationMs":119179,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":66601,"outputTokens":6610,"cachedInputTokens":783643,"totalTokens":856854,"reportedCostUsd":0.58397715,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":119179,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.58397715,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.58397715,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":14473,"reportedPromptChars":14447,"freshPromptChars":4684,"retrievalClipped":false},{"capturedPromptChars":7926,"reportedPromptChars":7926,"freshPromptChars":4684,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"9161c05c288e8620e5a748a28c901e1582efc240ca640203e1d8f7527e22a651","evidence-manifest.json":"296246a83eb0870a53e70f359f1859719f30b655e6ed20908abbfe4213605280","snapshots/context-integrity.json":"17779f4ebc38ed08071e40ac4b939149a18dfceb76c0f1f4f6ed21bf5d0153fa","snapshots/stock-harness-hire.json":"1f5a97e38555f57f47ee1900440d88ee521de75e7589c163f7688c852b661dc8","snapshots/stock-harness.json":"499a0b49c042b7c44f992c56983d207677cf619306ed6f6db9975ef1b84d455b","snapshots/stock-harness-preflight.json":"8f92c4b703533a645196bf03d1353be7196271d25638be7cb3fae6ec92c69c52"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.legacy-codex.local.assigned-skill-explicit-invocation","profile":"legacy-codex","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":50078,"providerDurationMs":27658,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":128946,"outputTokens":1727,"cachedInputTokens":97854,"totalTokens":228527,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":27658,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":12332,"reportedPromptChars":12332,"freshPromptChars":4683,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"ea06fc8b1ec166c9932de8634d08e528e90f7b734388ccad6d47b4b96df9ff6d","evidence-manifest.json":"45e71226e99eaee9bea0b0fae2cca35de573c2ab9d5add13159f321c8480ac0b","snapshots/context-integrity.json":"4cf10f812f258a77b472abd74f6691224dfcba4b9a2dec9c0f2c35c5ae741d7a","snapshots/stock-harness-hire.json":"19db78dbf51a67ff4bc307ef378da9cc024a0cae1c977951b49d089c1ec47fe4","snapshots/stock-harness.json":"a3cb7e2bf4e775cf9058b6be33d8df5321e312df2aed39736719454647610a8f","snapshots/stock-harness-preflight.json":"e9a116f8605a2630ff46075598ac6bae30817892d76efcd9a25891649cac0856"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.legacy-codex.local.continuity-restart","profile":"legacy-codex","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":85550,"providerDurationMs":29332,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":0,"inputTokens":468597,"outputTokens":2919,"cachedInputTokens":373648,"totalTokens":845164,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":29332,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":2787,"reportedPromptChars":2787,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":2730,"reportedPromptChars":2730,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":16326,"reportedPromptChars":16326,"freshPromptChars":1457,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"db7be8127de96e697a9d274b7a18db40695daa234684ec4d8e28afd8665c1f3e","evidence-manifest.json":"2e9151536b8ade2bfbb3c465438e8ace21417258c4037b6923920ee0784b0e47","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"af8b2fe11402a8b87a056818df29a62829ec9425f8a2d955e95692980e15e043","snapshots/stock-harness.json":"739e642582faf89b74c522774c64ea26c6af623539b7a8f99ce332066c965600","snapshots/stock-harness-preflight.json":"8fe9f32d7629a1f6bd632d58e22e39e47baff2b8bc3b74e3593e6411ab3c9a2b"}}, + {"executionId":"stock-harness.legacy-codex.local.ordered-comment-continuation","profile":"legacy-codex","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":74114,"providerDurationMs":57150,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":0,"inputTokens":628626,"outputTokens":7098,"cachedInputTokens":563211,"totalTokens":1198935,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":57150,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":9811,"reportedPromptChars":9785,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":12746,"reportedPromptChars":12746,"freshPromptChars":4683,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"56a457fcc7a36809eb4bfcedeff2fa079f5272754ff52fc453e8bf8d813e52f3","evidence-manifest.json":"bf17fb7d2a05f7e99e0723907bdc57f0a88bf6838d68f31f4b84dc6ebc62d3c7","snapshots/context-integrity.json":"a9156b810e32fa1d9d07826276a00c62ebe4f25ac3dc284ada157f2b2d72d800","snapshots/stock-harness-hire.json":"5fc8abd51d401c96d701b0862ac2dc674eb36244fdc25c49b67ba3947ba5a2fb","snapshots/stock-harness.json":"c72b8c53e1e50a82ab6ff815b363942936440cfff75f8f8ca5f54185c723889d","snapshots/stock-harness-preflight.json":"b14d304196cdc3146436c66199243c6da04784c13e45c9bd31e0ef83e97a29bc"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation","profile":"legacy-opencode","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":86384,"providerDurationMs":68956,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":57002,"outputTokens":2365,"cachedInputTokens":223232,"totalTokens":282599,"reportedCostUsd":0.0044563934000000005,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":68956,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0044563934000000005,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0044563934000000005,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":12335,"reportedPromptChars":12335,"freshPromptChars":4686,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"20dbb34bb975f73c3ac2695c047b6cbc8f6382757939ae6210b08a0f13ba9e9d","evidence-manifest.json":"67c3e042be8e0a5a0cd7cca82d1b0887e1392a2806c3b4db67baa3e99365b6be","snapshots/context-integrity.json":"decf2d93769acce2e2990de27c81ee32fd805ca53d3c2be68c56f22a013a7ac0","snapshots/stock-harness-hire.json":"187598ed72837719be8d6f0b77c5f8d35cfc871f8cf41d4fdc5b259f88171025","snapshots/stock-harness.json":"98424edd880cd05e0cbf8c32973fe9fbfc055037a74e76c20f893c3ee263f6a0","snapshots/stock-harness-preflight.json":"86d9f4080a2045ae3fa68468d0ffaf68baa26f27f0528274f13762f093b4c113"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.legacy-opencode.local.continuity-restart","profile":"legacy-opencode","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":126589,"providerDurationMs":78910,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":64286,"outputTokens":1046,"cachedInputTokens":28672,"totalTokens":94004,"reportedCostUsd":0.0022020907,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":78910,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0022020907,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0022020907,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":7611,"reportedPromptChars":7611,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":7554,"reportedPromptChars":7554,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":16335,"reportedPromptChars":16335,"freshPromptChars":1460,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"29dce6418aa35bb8070b07433046bbf39c9a8367a198cdb0aa28a0bea480032c","evidence-manifest.json":"439dc46130b607b7488e9df7736a711b267efa53689fe9fdad2f822525677215","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"7f7ad64827bbe430e06a7e01b1377443653793584264c0fa00ac88a742bd0f49","snapshots/stock-harness.json":"4dd536951a950e7507710a9eb44cdbcd354907397f1e95b8297b28a98fd75dbc","snapshots/stock-harness-preflight.json":"8fa2ec342ad583fb20fed4db88890fc31eab3ef5898a952a85371e53dd433aaf"}}, + {"executionId":"stock-harness.legacy-opencode.local.ordered-comment-continuation","profile":"legacy-opencode","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37048838402","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["deadline_or_run_start_timeout"],"failedMatcherPaths":["contextIntegrity.comments-queued-as-batch","contextIntegrity.initial-run-started","contextIntegrity.ordered-comments","contextIntegrity.changed-scope-preserved","contextIntegrity.packing-report-order","contextIntegrity.final-scope-applied","contextIntegrity.continuation-run-count","contextIntegrity.continuation-wake-comment-ids","contextIntegrity.single-durable-output","contextIntegrity.completed-task","contextIntegrity.successful-runs"],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":720616,"providerDurationMs":442902,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":51446,"outputTokens":8911,"cachedInputTokens":967936,"totalTokens":1028293,"reportedCostUsd":0.030657645200000007,"costStatus":"partial"},"runtime":{"provider":"local","agentRunDurationMs":442902,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.030657645200000007,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.030657645200000007,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"historical-generic-procedures-observed":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":16392,"reportedPromptChars":16461,"freshPromptChars":0,"retrievalClipped":true},{"capturedPromptChars":16406,"reportedPromptChars":16879,"freshPromptChars":4686,"retrievalClipped":true},{"capturedPromptChars":12749,"reportedPromptChars":12749,"freshPromptChars":4686,"retrievalClipped":false}],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"6d0e5d54b1742b02eb37346747510c66cb5cd0b80262a6c9590fd2a28861dcb2","evidence-manifest.json":"758931a4d30213d1b320f58cd2aa338e8a6a8aed579b878e14c5e52848bd5f13","snapshots/context-integrity.json":"bcaa4ee278361a90347625044aa49086b27093356abc0f3700b0363d28c8653c","snapshots/stock-harness-hire.json":"21431fa5e7cb6486037c670fc1e4aa64fb24716ad8ab325601ff294ada40263f","snapshots/stock-harness.json":"29c55e52f4adbbc9259cb9e221c08bfbbd7ad4c5f8247c24b3f029362e9000de","snapshots/stock-harness-preflight.json":"f341723074c996c79d8a6c30b9f12270634b20b40da452b13e4ab9ce91b47cac"},"taskDocumentCount":null,"behaviorCheckCounts":{"passed":1,"failed":11}}, + {"executionId":"stock-harness.runner-acpx-claude.local.assigned-skill-explicit-invocation","profile":"runner-acpx-claude","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":46264,"providerDurationMs":25372,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":10,"outputTokens":1042,"cachedInputTokens":141254,"totalTokens":142306,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":25372,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"073085f804a58d191bd44972edfb08d914fe622a86f236d0c4dc1b2bf8a03322","evidence-manifest.json":"4fad428c1df72abbfee24a4f7338e25284fc874cff1220e6b631fa86eb2d0bed","snapshots/context-integrity.json":"77e27b8ea01a83e0b65f3d5410e0787fbf84a7f6ea02f4401bd77488999dd4a2","snapshots/stock-harness-hire.json":"8c6c5d56fd068c15f1e19ce9a0b62204443c894d204311bd949d5e22c624fd3d","snapshots/stock-harness.json":"2a6860d7c29ad0d03240bda0a779b83d0455b8ce9fcfb1cf9e6ad559c4dec112","snapshots/stock-harness-preflight.json":"79b11e544d2821d1025476f5fd0d737e822ae5ab731ab32047d1fbf6b76015e4"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.runner-acpx-claude.local.continuity-restart","profile":"runner-acpx-claude","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":77424,"providerDurationMs":36820,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":6,"outputTokens":94,"cachedInputTokens":94997,"totalTokens":95097,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":36820,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"be582fe8af5f19543c16c8590697c5aa9dc688fdb653b0550aba0b55174becd1","evidence-manifest.json":"cfbd1cb13917a96b08dd46aae1fcc63e2b27af5d1d22f4952974d3415c74ffb8","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"254f18a8e8e1d6a5a2e7731728756671c1374cd4d1210469172bed9c2cd6e4c2","snapshots/stock-harness.json":"8efa84d6907de220dd7b9fc2c384c61aae0606e7f5311f0e926e26ae1617d869","snapshots/stock-harness-preflight.json":"c7e3f2342ef9fc824b28e3f22ffbbb326d2ba1ff84e36c4deedc56a7abefbe07"}}, + {"executionId":"stock-harness.runner-acpx-claude.local.ordered-comment-continuation","profile":"runner-acpx-claude","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":132587,"providerDurationMs":108895,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":26,"outputTokens":7996,"cachedInputTokens":565030,"totalTokens":573052,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":108895,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"62c6c55047813b0b5aa5334e29dbe3f90225c83ac436eab14afd97d78ed9ec65","evidence-manifest.json":"7d53ad7f80b65a1c7af3fad4fa3d0c652ae25cdaeae3d41cbee1bb888b853fb5","snapshots/context-integrity.json":"2081459b761a84469399bec377d7380c172b64fddae48735c5329771656e28a7","snapshots/stock-harness-hire.json":"f11c90fdae215949458d61a6e1ca6565fe277ea30d52fd922c05cffdf4867e5c","snapshots/stock-harness.json":"f303a3fb5c5e55cefb696d46204aafe6f1cefc5b0a6b07a728f2bf0dbf1dfdcf","snapshots/stock-harness-preflight.json":"79768aa58996d28ee80c8a9e8d5c4ba1a42eba58a408163cd8a6037ac6005b26"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.runner-codex.local.assigned-skill-explicit-invocation","profile":"runner-codex","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":53751,"providerDurationMs":29850,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":162041,"outputTokens":1296,"cachedInputTokens":137822,"totalTokens":301159,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":29850,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"ffcc5203624c19b0acda50dfd1e283c595ae3ed6072e4aaf9781fef4bf08fa45","evidence-manifest.json":"1460fae9f00a241a3ddbe7f623a4c1c6cf0f6434e73e47d0de7f98dc00dec459","snapshots/context-integrity.json":"63d16c6558efcf03097fd2ef8caf1f01c39a38127aee2904eba0cf0edfb711e6","snapshots/stock-harness-hire.json":"61630a0e0521f8ccca30753bde385b731f5c8c15cf6280b000947e80acf3decf","snapshots/stock-harness.json":"bc95a45aab085ce3eee988a9b235281b847cd8cb73c4d294ce576b97e73d879f","snapshots/stock-harness-preflight.json":"4daf4cd83a8b1539ab8a8db2fd5cf204746884d8cde04f6507acfff5be75ef11"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.runner-codex.local.continuity-restart","profile":"runner-codex","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":84362,"providerDurationMs":35132,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":315100,"outputTokens":2339,"cachedInputTokens":231408,"totalTokens":548847,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":35132,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"dbf2e14fcda52d2a33808e5a4c226d45f864752274c823c26beb7c2ea5409f03","evidence-manifest.json":"86a34f1a25854354bee0d9f0ea88cab9ba9b51476cf274d12f5b491a9910188a","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"a38cab643c78fde9d61d15074b82297b94bf46ae5c3296aa06e6958f67188eda","snapshots/stock-harness.json":"2d148d3e55c57619bde56558f7e87e85cbb1305e46eba83bedb21c3ee4b43b54","snapshots/stock-harness-preflight.json":"9517ef436cc83da9078dd75587f01be8a28685131c37773dc092cf89c137ca0c"}}, + {"executionId":"stock-harness.runner-codex.local.ordered-comment-continuation","profile":"runner-codex","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":92663,"providerDurationMs":70832,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":428674,"outputTokens":3591,"cachedInputTokens":372482,"totalTokens":804747,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":70832,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"8c7a81172e0d519ffc8aca9a3025c88166d93d12a1744e560813c66063d6d130","evidence-manifest.json":"b495ad141c12f2458e3b66123cb673dec68511bcb5c2d59af78400dc67c417eb","snapshots/context-integrity.json":"166eb5a26fa389529287aa940c4e3caa85665ea67fcce019bc69a8bb8d9f7b41","snapshots/stock-harness-hire.json":"2fe4b0d207648628658312e97b491336929197f826c00dce012b6afc43537564","snapshots/stock-harness.json":"79b83fed61c69589d38b48552b04dce41bf41dca27f6e7f5dd15229024b6338f","snapshots/stock-harness-preflight.json":"689ce3141b2c6970eac5468885e5111daa730bcda25a1459e7bac4177650045e"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.runner-opencode.local.assigned-skill-explicit-invocation","profile":"runner-opencode","case":"assigned-skill-explicit-invocation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":90297,"providerDurationMs":68028,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":44259,"outputTokens":1165,"cachedInputTokens":62976,"totalTokens":108400,"reportedCostUsd":0.0020380985,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":68028,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0020380985,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0020380985,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"3ffa3bff50fbade8801da7c26ffed426eee9bda9a7e3640793f03a220f74ea82","evidence-manifest.json":"6df0c3b49ef54154721d4d0ba434f893fcba372ba8588fef2bd1a5947284b1d3","snapshots/context-integrity.json":"41b60c88adb0c529148981a55ab3b18964035b47786da50942183b58e8884fed","snapshots/stock-harness-hire.json":"c9a2fe249e217fcd5c205f1dc768446333932de74d87dc6f4a1e7855c5e4d815","snapshots/stock-harness.json":"474b0d6aecd8e8bce23f78c260ad4dc83c286d1d39917ea0920ad4b402d9404b","snapshots/stock-harness-preflight.json":"baeb38686ea4d3233906d83a889fdf3d805846b27f25c6c04da68842822a496d"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.runner-opencode.local.continuity-restart","profile":"runner-opencode","case":"continuity-restart","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":155833,"providerDurationMs":101201,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":48028,"outputTokens":96,"cachedInputTokens":23808,"totalTokens":71932,"reportedCostUsd":0.0004892436,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":101201,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0004892436,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0004892436,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"57fe32615f16ff3706044419bb81362a67ec35aed7a5f74a5b8207b5831a5cba","evidence-manifest.json":"07e3be863cc20a1352a094f38e6133004dbfd673f4fd59599431a022ce061ff1","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"d127504476aee2130312fdf092a523060188066abe9f969ae9b178b57374cced","snapshots/stock-harness.json":"0df941763ed917e2c553a6c6a4dfb0c87bdb7181d42a28f614c765906b74f1e9","snapshots/stock-harness-preflight.json":"f57655c970472d2a9b9bebe100d8c31788ed76e06da4363046a5a6d3f09269bb"}}, + {"executionId":"stock-harness.runner-opencode.local.ordered-comment-continuation","profile":"runner-opencode","case":"ordered-comment-continuation","sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceRef":"refs/heads/codex/stock-harness-previous-instructions","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042864888","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["deadline_or_run_start_timeout"],"failedMatcherPaths":["contextIntegrity.comments-queued-as-batch","contextIntegrity.initial-run-started","contextIntegrity.ordered-comments","contextIntegrity.changed-scope-preserved","contextIntegrity.packing-report-order","contextIntegrity.final-scope-applied","contextIntegrity.continuation-run-count","contextIntegrity.continuation-wake-comment-ids","contextIntegrity.single-durable-output","contextIntegrity.completed-task","contextIntegrity.successful-runs"],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":721578,"providerDurationMs":446042,"cleanup":"passed","billing":{"llm":{"runCount":11,"runsWithTokenUsage":11,"runsWithReportedCost":11,"inputTokens":159917,"outputTokens":7291,"cachedInputTokens":1077987,"totalTokens":1245195,"reportedCostUsd":0.015645790399999998,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":446042,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.015645790399999998,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.015645790399999998,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"12c5433c67dc2e62916b879349c7ba2b6e0431f0","sourceFingerprint":"2e5f22848450179ac2d573662c718300f86a935d09d15204c9d97e98f65b53f4","passed":true,"providerCallsBeforeAdmission":0,"passedTests":536},"evidenceHashes":{"result.json":"f727dd5f76162c4b995e649e1a7bf55a52381efa550e3bda46ec0a51a899d178","evidence-manifest.json":"6a550c598b67f60fcea70d56907c4a2c84239bc0d01e58737e99a0014144bb70","snapshots/context-integrity.json":"14212826847fbd9d6d63c7b70afb102fdb5d0100f7d8f1368d46ff6f5e32db07","snapshots/stock-harness-hire.json":"ab6dd2377fc251e44ed8f3680d40d10a315b635f3003e45501c7496037fb21ca","snapshots/stock-harness.json":"38759bca194a8038f0d2be4598e50a04d61215ecabdfec95cd7f6cfa4eeec61c","snapshots/stock-harness-preflight.json":"312611057436fd8cbdcc5179ffa5e62144800a6a9244e646bb6e0b3a9a0cb551"},"taskDocumentCount":null,"behaviorCheckCounts":{"passed":1,"failed":11}} + ], + "candidate": [ + {"executionId":"stock-harness.legacy-acp-claude.local.assigned-skill-explicit-invocation","profile":"legacy-acp-claude","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","no_single_paperclip_document"],"failedMatcherPaths":["contextIntegrity.single-durable-output"],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":123175,"providerDurationMs":111928,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":25333,"outputTokens":5771,"cachedInputTokens":316215,"totalTokens":347319,"reportedCostUsd":0.2838505,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":111928,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.2838505,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.2838505,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":10966,"reportedPromptChars":10449,"freshPromptChars":787,"retrievalClipped":false},{"capturedPromptChars":6028,"reportedPromptChars":5498,"freshPromptChars":787,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"2a5fed84ee0f95a67e2c2b0d4247ffae33e16a7c2aa11dc57b11228e0db90965","evidence-manifest.json":"7e31c707e866e15bf44a276ac252c423a47a63364e76ad85de9a8d5b6a4e1dfd","snapshots/context-integrity.json":"3c36c18d5e257e788514c02b50d63c7c69885201df0a07750129d55c0f583676","snapshots/stock-harness-hire.json":"2d128826b4e5068ad6ed819512bcb17160346becd9344b8ab8e70c984c3241b7","snapshots/stock-harness.json":"49b8da640ab0b9dd2b209ce0cf9f3bec6deceb801bfe98c0166d7702a9732a80","snapshots/stock-harness-preflight.json":"a12fdf9b8b7d761ff223e30dcaeb835ad9bf32ef7f1c8b8983a638d175c2ab5f"},"taskDocumentCount":0,"behaviorCheckCounts":{"passed":6,"failed":1}}, + {"executionId":"stock-harness.legacy-acp-claude.local.continuity-restart","profile":"legacy-acp-claude","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":34813,"providerDurationMs":18798,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":16797,"outputTokens":329,"cachedInputTokens":54375,"totalTokens":71501,"reportedCostUsd":0.092943,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":18798,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.092943,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.092943,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":13201,"reportedPromptChars":12723,"freshPromptChars":787,"retrievalClipped":false},{"capturedPromptChars":13233,"reportedPromptChars":12755,"freshPromptChars":787,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"ecab7c9f83668fd4712702bb906adf38e188e34f7f07bd7d746da685339a2ecd","evidence-manifest.json":"544a726ddf3ac212671f7c3d73dafcf94cd5b3abb9a3ad5197d6c793e9d7aaa0","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"78f5c467789921db4d2194ba80459524b04644f2d8d8a3b8678fdf7686004e8f","snapshots/stock-harness.json":"61641e5e25351e748afb2e4533d70b1cc34a8b78e1fbc56a27d96b2017bd0f45","snapshots/stock-harness-preflight.json":"4e522117d8489a16263c3d66ca513207217a7861733bfd11ef2f7149aef680e9"}}, + {"executionId":"stock-harness.legacy-acp-claude.local.ordered-comment-continuation","profile":"legacy-acp-claude","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":154765,"providerDurationMs":139409,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":61445,"outputTokens":6979,"cachedInputTokens":562903,"totalTokens":631327,"reportedCostUsd":0.5120724,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":139409,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.5120724,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.5120724,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":13419,"reportedPromptChars":12915,"freshPromptChars":787,"retrievalClipped":false},{"capturedPromptChars":6390,"reportedPromptChars":5912,"freshPromptChars":787,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"9c23fec76e718c1fa5f775ae8cdcee1e5433039ff74d7ffbf306b5be6693700d","evidence-manifest.json":"5284aba88b923322b1f162045d20e36236fb731689527a1a1d06d5305efd4a64","snapshots/context-integrity.json":"8fef43dac791a84db62bfea8977f586afe3e5dc5c36225eaab79faaa15b4dfca","snapshots/stock-harness-hire.json":"4961a6103fcb6b58d5d65981c68afa178ea9e1eb6038ad3b3049d47e4bc737ab","snapshots/stock-harness.json":"4e11dd1b37581fed8628b7ebd953ebf3b6fe389d357ffd2efcd9e6b976e4d64a","snapshots/stock-harness-preflight.json":"7d962ba46f4f02f111ed4e423705009f38ee99096d180b693f4c41ed4c805126"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.legacy-acp-codex.local.assigned-skill-explicit-invocation","profile":"legacy-acp-codex","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":65745,"providerDurationMs":49358,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":1379,"outputTokens":34,"cachedInputTokens":41987,"totalTokens":43400,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":49358,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":5509,"reportedPromptChars":5533,"freshPromptChars":786,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"81c5c35a942c78a2c9d7d77064264504311a191c6dd421e656d4f0751670b6ff","evidence-manifest.json":"bff59a88d627186d18282fe39822dd5a57d3a2396ae9aca5fa178a3b4ac3cf60","snapshots/context-integrity.json":"3809130fd34f0b6035ed5b17bd9e354d19aac02fc2f64aea01811803151c7f82","snapshots/stock-harness-hire.json":"4e0468ee05d56f08faa9ffbe3ed6cf28c07490a8b27a08e80875264dc66daac5","snapshots/stock-harness.json":"c5a82192ff53c6eec7a5497602a9969af3c21ea7fde8fdff4309025f5ccf80d2","snapshots/stock-harness-preflight.json":"12944d93aac9f8688d2832369f7eadc10a9546a6fb6bc2f2d8b3caef38b0cc2f"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.legacy-acp-codex.local.continuity-restart","profile":"legacy-acp-codex","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard","structural_receipt_incomplete"],"failedMatcherPaths":["stockHarness.invocation-evidence-complete"],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":35779,"providerDurationMs":20059,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":711,"outputTokens":30,"cachedInputTokens":31118,"totalTokens":31859,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":20059,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":false,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":12764,"reportedPromptChars":12788,"freshPromptChars":786,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"1c59fb6b3845ef6bc4f9cc549826e45334b3220ce5832f2fe4b9d23fd45136b6","evidence-manifest.json":"4cdd7cc5bb0a025e7626e947dd72558c13203aa9dc9765f5afbfed5acc907de1","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"dcbe621edcd9437c47361c74676db38a240fdcaff2401d41213a185e7cf9b785","snapshots/stock-harness.json":"8908438e0c7565ebca6d4753e619f87bd0c015210214fbb8c39876d7e13cebaa","snapshots/stock-harness-preflight.json":"3b2de5b161384bebe03f7cad3b3268f98cca291cd8fbf4b8a0c830536daecf87"}}, + {"executionId":"stock-harness.legacy-acp-codex.local.ordered-comment-continuation","profile":"legacy-acp-codex","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"failed","failureClass":"secret_leak","causeCodes":["persisted_credential_guard"],"failedMatcherPaths":["contextIntegrity.comments-queued-as-batch","contextIntegrity.initial-run-started","contextIntegrity.ordered-comments","contextIntegrity.changed-scope-preserved","contextIntegrity.packing-report-order","contextIntegrity.final-scope-applied","contextIntegrity.continuation-run-count","contextIntegrity.continuation-wake-comment-ids","contextIntegrity.single-durable-output","contextIntegrity.completed-task","contextIntegrity.successful-runs"],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":720307,"providerDurationMs":46665,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":540,"outputTokens":28,"cachedInputTokens":32395,"totalTokens":32963,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":46665,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":5923,"reportedPromptChars":5947,"freshPromptChars":786,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"a1285f2fecccdcdac85360784e4be82188b564d3135f7adf6606fa1451307376","evidence-manifest.json":"39864272e9331c88c1266f2b5f05b755aacdc46c2c7015c4d4df8539c0a22718","snapshots/context-integrity.json":"5e663656093c672fb294719edb8636dd15a5ed1441912f32a32c78254663b797","snapshots/stock-harness-hire.json":"35d859103bf2acb7c7f6d212a8ce277e42726da388f7c52ee0b05d600dc09c78","snapshots/stock-harness.json":"483b4efedcc34570ad0ad364d92ba0212186cbc36a704639106e181402470468","snapshots/stock-harness-preflight.json":"5bb04e0c93a3f4ac5d9f41a5775fba85e0687a8161e255257545d54151421c62"},"taskDocumentCount":null,"behaviorCheckCounts":{"passed":1,"failed":11}}, + {"executionId":"stock-harness.legacy-claude.local.assigned-skill-explicit-invocation","profile":"legacy-claude","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["no_single_paperclip_document"],"failedMatcherPaths":["contextIntegrity.single-durable-output"],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":37720,"providerDurationMs":25763,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":27376,"outputTokens":1281,"cachedInputTokens":184039,"totalTokens":212696,"reportedCostUsd":0.17707545,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":25763,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.17707545,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.17707545,"complete":true},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":3611,"reportedPromptChars":3611,"freshPromptChars":783,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"8636a50a54d39fd66b1439f13ce5ac125b79f16b2826f76fc297427a1a2ab701","evidence-manifest.json":"d1d7f47272981406d104f5ae6455d5b0a4f6113f3c3d02fc8bc787d7c8fc70fc","snapshots/context-integrity.json":"eac7d7b889e5cc3861b937ffa87c7b92d25ade5b0f2bb73bdd29b5764c2a085a","snapshots/stock-harness-hire.json":"545030ff52d9196a604e099b81b52a3071f582f2467284c695eb62d96589c67c","snapshots/stock-harness.json":"c9ca463a792c7dccd00c0795c1e405495d9e1dff30b2cbae2b5e7bac7114fdd2","snapshots/stock-harness-preflight.json":"c1b5a843459d80e074759600c830c92bb00a0d2430255cda7bed5c22cabc7680"},"taskDocumentCount":0,"behaviorCheckCounts":{"passed":6,"failed":1}}, + {"executionId":"stock-harness.legacy-claude.local.continuity-restart","profile":"legacy-claude","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["chat_memory_assertion"],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":28251,"providerDurationMs":11843,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":19490,"outputTokens":298,"cachedInputTokens":50061,"totalTokens":69849,"reportedCostUsd":0.09257055,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":11843,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.09257055,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.09257055,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":10776,"reportedPromptChars":10776,"freshPromptChars":783,"retrievalClipped":false},{"capturedPromptChars":10833,"reportedPromptChars":10833,"freshPromptChars":783,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"86af4c61a239c718181ab58fd815cf7a200e7473b2fe9390b58bc15ff9d6ed4c","evidence-manifest.json":"e1c69d7c6044b0e8ba65fa1b961de84989facb3fd65f6bba42632b22d2a1bab9","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"9a320857e3d1fe14a41b0154b8c1d2cfc81f869ed69790bd729fe7c86e9c54fb","snapshots/stock-harness.json":"9a2d546c992bf9d59711e37e81d841b7b1c73e6e1d62a6f27b5302f1d8ea3980","snapshots/stock-harness-preflight.json":"13318522befc6c0ffcc0d6ca7a0dc12112902427ddc1a7ac36f3ac43e7f83d8d"}}, + {"executionId":"stock-harness.legacy-claude.local.ordered-comment-continuation","profile":"legacy-claude","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-4-6","runtimeMode":"legacy","durationMs":158440,"providerDurationMs":139386,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":64318,"outputTokens":7515,"cachedInputTokens":669889,"totalTokens":741722,"reportedCostUsd":0.5548662,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":139386,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.5548662,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.5548662,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":11079,"reportedPromptChars":11053,"freshPromptChars":783,"retrievalClipped":false},{"capturedPromptChars":4025,"reportedPromptChars":4025,"freshPromptChars":783,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"59885705b15fc33517ab7f7a82d9ada2e411c918b10fa5b4ce928498a801e739","evidence-manifest.json":"c2a5be637da0d7ff9436b732d8db72ab1fdce788caf6a40c5973c14a27bcc08a","snapshots/context-integrity.json":"fbee7c1a37b964bcc327ebbf37602bc227bccb7fafdc4296614e178e66016599","snapshots/stock-harness-hire.json":"cf2f8ff4b1eb13c5e5f854ecac27c25826e7a4dbafbc277e8702b8beee6d913f","snapshots/stock-harness.json":"3e8e7c06f0918a896341a38e4613cc1f4caaad4b5ab12c03a8c3e41c4707a82a","snapshots/stock-harness-preflight.json":"6d6716a08fc2ea7d27ee08840fd3290a04a8e85dfed5950b690f4c2267d20a6e"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.legacy-codex.local.assigned-skill-explicit-invocation","profile":"legacy-codex","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":44075,"providerDurationMs":27042,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":0,"inputTokens":241944,"outputTokens":2132,"cachedInputTokens":215273,"totalTokens":459349,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":27042,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["metered_api"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":4223,"reportedPromptChars":4223,"freshPromptChars":782,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"b0b26f37d2f0c0342a9a275b22716be190891d1b55e79fa8d33e3f495649a26b","evidence-manifest.json":"446abc91b7366507c4100741d71ace903de0b78149b33e45fa90d39183f71b75","snapshots/context-integrity.json":"85197eb92d34334cc2de658d09d7a7e54d44ab3d2fd81ad87732c88c6610b6b2","snapshots/stock-harness-hire.json":"568453725988fc61d459f96efda07d19ec5c00935d4c4a9ff9557cd0062aef29","snapshots/stock-harness.json":"0c8626740b520687d08b0b1f0aab8a9ead2ebd58290252f05d890f524636f461","snapshots/stock-harness-preflight.json":"9a15351a65076deeb77ccd58216def768a658fc743fd57dc89c3234ff6b36542"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.legacy-codex.local.continuity-restart","profile":"legacy-codex","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":62742,"providerDurationMs":10028,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":0,"inputTokens":91929,"outputTokens":475,"cachedInputTokens":45089,"totalTokens":137493,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":10028,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":2787,"reportedPromptChars":2787,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":2730,"reportedPromptChars":2730,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":11443,"reportedPromptChars":11443,"freshPromptChars":782,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"1ec35f3857350d86d3888c4becf8ce898a59641f3eb9825a0d800970ec7090ad","evidence-manifest.json":"908761dcb779dcd0b95ef78b5e9addb77b211acbc502a1c8fdf793202d309472","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"b7b6f4db7cdc84a795ccfa24df58dd8131fe8a44b3ef1cdaf36dfc6685556da7","snapshots/stock-harness.json":"9a0045bbf4240ddbf08ec0444e08ef70e3ef76399db09075b5316b1fc51ac5bf","snapshots/stock-harness-preflight.json":"cfbae9842e2a045b815a9a7f1fa471f8f9494ec999c91e9808933cbc013c84a4"}}, + {"executionId":"stock-harness.legacy-codex.local.ordered-comment-continuation","profile":"legacy-codex","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/reduce-default-agent-manual","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37042856368","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"legacy","durationMs":66403,"providerDurationMs":44175,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":0,"inputTokens":416426,"outputTokens":4483,"cachedInputTokens":353673,"totalTokens":774582,"reportedCostUsd":0,"costStatus":"unpriced"},"runtime":{"provider":"local","agentRunDurationMs":44175,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":false},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":8456,"reportedPromptChars":8430,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":4637,"reportedPromptChars":4637,"freshPromptChars":782,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"71f99ec1b45a57ae40acbd42c55e835f63d9d2329752f6bffcabd4488ecb2a96","evidence-manifest.json":"0c3347aaec6d7c3498e62c19938762c35a6069d31385e5d617174695ab60462e","snapshots/context-integrity.json":"5c08b2af1b27e9255505e917958e7b97856cfbf48cce3bef3b8e792275823b6e","snapshots/stock-harness-hire.json":"ddc6b182c4d53c7cd7f63c906d2a64e6286c54c8b6e9991e0c78da578586d865","snapshots/stock-harness.json":"11a737f54107eb9a9e55234890bce90003f84b258f6d5c8512d65f080edfc52f","snapshots/stock-harness-preflight.json":"d72a98df5a4e00bfa766e85ef28b866d32b89fe2aa9a67e3000ad50ddea47599"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation","profile":"legacy-opencode","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"failed","failureClass":"candidate_failure","causeCodes":["no_single_paperclip_document"],"failedMatcherPaths":["contextIntegrity.single-durable-output"],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":212558,"providerDurationMs":199337,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":96260,"outputTokens":6708,"cachedInputTokens":172288,"totalTokens":275256,"reportedCostUsd":0.010608548800000001,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":199337,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.010608548800000001,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.010608548800000001,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":8426,"reportedPromptChars":8439,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":4226,"reportedPromptChars":4226,"freshPromptChars":785,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"435f708b59438d1d344130bc9e23b8ddbc93adf74d4e5ea8f049e0f1d06efa72","evidence-manifest.json":"8591f98df496159f99b98659ffdcbdf2e86718dccb4545d461f592bab0626b05","snapshots/context-integrity.json":"e4271e26bb891c3c15f2d54dd5cfba6605ebb4431456d70b4f146bb8eb4e1788","snapshots/stock-harness-hire.json":"6bb95cb92fbb14976b5f6796f530823f6aacdcb1448f89cb8b2ec093125266e2","snapshots/stock-harness.json":"b8a5a04730324a038869e28f79d55dd06ebf12abc2e6366bbfe9218995a6748b","snapshots/stock-harness-preflight.json":"c17087412694993475276a69162a171859dc54afae5b08ab89573cf0637d6e19"},"taskDocumentCount":0,"behaviorCheckCounts":{"passed":6,"failed":1}}, + {"executionId":"stock-harness.legacy-opencode.local.continuity-restart","profile":"legacy-opencode","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":69761,"providerDurationMs":19747,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":12219,"outputTokens":116,"cachedInputTokens":0,"totalTokens":12335,"reportedCostUsd":0.0003185752,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":19747,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0003185752,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0003185752,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":3403,"reportedPromptChars":3403,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":3346,"reportedPromptChars":3346,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":11452,"reportedPromptChars":11452,"freshPromptChars":785,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"0e56bb0c46a05efeeffa6feb09f54fe13a33c2c96ab4e7a7ec548a0d80a1a235","evidence-manifest.json":"0e601f9bde1bdb7f44f000704078343a1478046c9eb488f2c8a785cd1d4cebec","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"ab7d6af088306e1d43d4901a370dffa46b16dd4c17c2838a8e9e48a2b8e7a62f","snapshots/stock-harness.json":"2a6eabd1657c89568f890de12840747b544c7385c4363b3016879b07028961fa","snapshots/stock-harness-preflight.json":"539f106844e9421c284944c0b1526d74b0cc8e0905c5960be14cfa77ee17a023"}}, + {"executionId":"stock-harness.legacy-opencode.local.ordered-comment-continuation","profile":"legacy-opencode","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"legacy","durationMs":212184,"providerDurationMs":188508,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":177025,"outputTokens":7495,"cachedInputTokens":256512,"totalTokens":441032,"reportedCostUsd":0.011804638700000002,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":188508,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.011804638700000002,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.011804638700000002,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true,"invocation-evidence-complete":true,"generic-procedures-absent":true,"fresh-default-delivered":true},"invocationMetrics":[{"capturedPromptChars":10775,"reportedPromptChars":10736,"freshPromptChars":0,"retrievalClipped":false},{"capturedPromptChars":4640,"reportedPromptChars":4640,"freshPromptChars":785,"retrievalClipped":false}],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"8ca5a9a446dda03100d8ee66c5b9e02fd19c335dfc8cc34f83662aa1181240b9","evidence-manifest.json":"496a392c39d3a75f0287c19a6fc25e7577dee6bfef0e3c7f16801f61574cfa99","snapshots/context-integrity.json":"da023ad1b2e22b1659573e1829cfb273a4d3e329d1b324506dbc73104901ecf7","snapshots/stock-harness-hire.json":"ec6d9274b80d291e8949c71a9bfb1f6fdc88cd78a2094dcad1c0a599ce9b78cd","snapshots/stock-harness.json":"680d3af8a66f473a95905920123f0760aff53eb4ce09cd01393a78b6292178cd","snapshots/stock-harness-preflight.json":"5b79fae54bd6036dab2ad8bf8e8bacb9087e702b363dbbe4af2c58115eea8ead"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.runner-acpx-claude.local.assigned-skill-explicit-invocation","profile":"runner-acpx-claude","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":50059,"providerDurationMs":28899,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":10,"outputTokens":1332,"cachedInputTokens":136292,"totalTokens":137634,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":28899,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"526b8881ac8dd2b51a788ca6ae535f7cd856e38066fadbd70737071c4f898abf","evidence-manifest.json":"40dc320ecab36c004d5e33bba8ebe1fb4d8444fd23d7ef0f055b188b474fffca","snapshots/context-integrity.json":"0b0b2b42376e1c5cf58b6ae1e9e0960b6b2bc82218e20589e04ca44639207ecb","snapshots/stock-harness-hire.json":"7f47075f11209d29c92091925bb61b2e2e22b06e1ce49b7dba69060242d62fc5","snapshots/stock-harness.json":"06398532782d5e4c9eb76ef55cdc87ed8d49c6f31dc06f09ce28715aa09f76fa","snapshots/stock-harness-preflight.json":"cb1980515ce11e2b55fcea5799277f511d0c6d01bf26b78e93be3ae17fa9b08d"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.runner-acpx-claude.local.continuity-restart","profile":"runner-acpx-claude","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":93902,"providerDurationMs":43211,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":6,"outputTokens":101,"cachedInputTokens":92222,"totalTokens":92329,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":43211,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"332ba05ee7cda6971777af3454e8589652f742f45cd6350da4396a420e8c549f","evidence-manifest.json":"eaff7dd508a80dffef9d638a39e13b21d2ef427086ae478c10520c2b336fe881","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"e8a854b3a0d886b01cb910e5b05baa5f1c791ef57b9a04168a0ce65cc60df7ce","snapshots/stock-harness.json":"1e53449e403bc079887cd183e2b556fbecd238a20130ef07c0c5decf499ff4ae","snapshots/stock-harness-preflight.json":"02283c369f44ab85a6400e0659da00e48bee84727fcebc35a8a691c4529c7cc0"}}, + {"executionId":"stock-harness.runner-acpx-claude.local.ordered-comment-continuation","profile":"runner-acpx-claude","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"claude-sonnet-5","runtimeMode":"native","durationMs":93050,"providerDurationMs":70393,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":18,"outputTokens":3794,"cachedInputTokens":332880,"totalTokens":336692,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":70393,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"3aae7a7ca62c39c74ac6d12e71b6b75223c76112ea018b8d08db1da52ea72e6a","evidence-manifest.json":"f08a3eee09a8fd4b279865dbe8f9ab2a66cc1a79addb216fe484561fcb9e2467","snapshots/context-integrity.json":"400df6ee5a9294998d4306f6eab6816e86155d8ecf0f39a6fc672381d47d423f","snapshots/stock-harness-hire.json":"604c41ac65944ef25cb7bbbf19941401139ed422342d7d6ebf9a07b79ecf7a75","snapshots/stock-harness.json":"f17d4f47657a0151dbfc585de64e338b7f79deae052f454794c2b934a7d398a2","snapshots/stock-harness-preflight.json":"fa502f0e3a3467dab6b04f810e6001ca635ce2d4bb5bae42a50e7a85f520a356"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.runner-codex.local.assigned-skill-explicit-invocation","profile":"runner-codex","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":48064,"providerDurationMs":26079,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":133429,"outputTokens":1102,"cachedInputTokens":110284,"totalTokens":244815,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":26079,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"690736dbf71179913631255763fd7b8aecd825236532154e54a3e1e53240d6b2","evidence-manifest.json":"75135aa0e7020617f54b5f2f593a5d48669ac7ecc864064d1eefc91af86d431c","snapshots/context-integrity.json":"46e65857a4d82f6ad900bc7e89166c35f2543faa368e48130f83dcea526ca65e","snapshots/stock-harness-hire.json":"2bf564606e974bb38baf50e363edd6d11aa5009fd429602a44fecb337bf9036c","snapshots/stock-harness.json":"2bf943a7cba0fc3886a04fa96f5646677f427e516232b35f6538384f4e6ed39f","snapshots/stock-harness-preflight.json":"635be5d6d578d7ad95b0e87697b939ed10f3cf2d3338caedc94350a1ed67ccc9"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.runner-codex.local.continuity-restart","profile":"runner-codex","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":89776,"providerDurationMs":43862,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":460440,"outputTokens":2336,"cachedInputTokens":378421,"totalTokens":841197,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":43862,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"b32c6c2c10d1eaabd0e2408db313bf9f30a7a60413733ef8ddfd8114ff6f2fdb","evidence-manifest.json":"34a60f60f021a731a98fab708ac31e344d8ecfa5a1d7dc72daa8c957028cbcea","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"855403ffdcaa2f1542a5ebbc0dbd7590c609c170afc657884eeb9a1e79a51310","snapshots/stock-harness.json":"d4997c371e7db1a9e1c13536279480035eee7873e92a3ab3300969d508339481","snapshots/stock-harness-preflight.json":"642913765f2dfb9bbae2f46401448e4c3b6d3f9134cba2a073178d530c395efc"}}, + {"executionId":"stock-harness.runner-codex.local.ordered-comment-continuation","profile":"runner-codex","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"gpt-5.6-sol","runtimeMode":"native","durationMs":85296,"providerDurationMs":64569,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":359345,"outputTokens":2627,"cachedInputTokens":307813,"totalTokens":669785,"reportedCostUsd":0,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":64569,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"1a1a5f9b53dfc8cd2d96bafb8b34f6b1ac006514da70963ea3d8328f1c594309","evidence-manifest.json":"5220dd721a29958a2dbc6139a33bed2c94765801181c795abca6118722204f7b","snapshots/context-integrity.json":"100d34aa2c461b540164a00442100d5f6f438f706d14fb7316d80bf07d680cd8","snapshots/stock-harness-hire.json":"103296d40bc0b4500fa9b45ffb1e8192b770a1d4f87f8d0dd4f9e309384d407c","snapshots/stock-harness.json":"8d1eea8270fd2c29e284db9af321bebe85e635f6ec77d56674466d1480012ba8","snapshots/stock-harness-preflight.json":"be022954d2b5f5ef0bdae4cd892f06163f9d0863d0201aa92a5bab12f56a4dac"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}}, + {"executionId":"stock-harness.runner-opencode.local.assigned-skill-explicit-invocation","profile":"runner-opencode","case":"assigned-skill-explicit-invocation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":69432,"providerDurationMs":45144,"cleanup":"passed","billing":{"llm":{"runCount":1,"runsWithTokenUsage":1,"runsWithReportedCost":1,"inputTokens":22870,"outputTokens":1171,"cachedInputTokens":105472,"totalTokens":129513,"reportedCostUsd":0.0021534242,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":45144,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0021534242,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0021534242,"complete":true},"billingTypes":["unknown"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"69ac8808e939a796c40ac3665d56b152b43747fa91b7bbc84fd0a0f2fc68f5d9","evidence-manifest.json":"a0af2ef7ba4653d7ec825ac12de2a9e5f8b7fd7042bcfbe215e8af1b75ad38f8","snapshots/context-integrity.json":"e82ea3b31d4c6c73a95715c1d0e31338eb31a50f760bb6ea99b1d21032c6da32","snapshots/stock-harness-hire.json":"87ec7c1456344be2ed5cf68d4813920e8204fc483c6a34b1b62116bd721d9073","snapshots/stock-harness.json":"11cb4a18967ee59a7d5c202ee590a8dcfa7a9c145ab5ac87cde8da6a69f25497","snapshots/stock-harness-preflight.json":"c044805c4bb0db4c3cac0a4bbf54b937da7d1d376a4ac319f4d2baf3c636f51c"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":7,"failed":0}}, + {"executionId":"stock-harness.runner-opencode.local.continuity-restart","profile":"runner-opencode","case":"continuity-restart","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":152945,"providerDurationMs":93715,"cleanup":"passed","billing":{"llm":{"runCount":3,"runsWithTokenUsage":3,"runsWithReportedCost":3,"inputTokens":26520,"outputTokens":106,"cachedInputTokens":45312,"totalTokens":71938,"reportedCostUsd":0.0005020232,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":93715,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0005020232,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0005020232,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"ae12b2eee65f45b8e2f530c037002e91db351134b6848da89e80bf7c15eb25a3","evidence-manifest.json":"dce1906b35fbbe6ac471ef922a2f0b1bae9da9438d56e78894e3651cba402aaa","snapshots/context-integrity.json":null,"snapshots/stock-harness-hire.json":"a57da515e6454ee8b82b9bf633a8e07097217f30555606538e135c7979de71c0","snapshots/stock-harness.json":"fa0501dbd07065bbbbd9a1f4c6bd82ba27f0fa45238790d531db4587c000309d","snapshots/stock-harness-preflight.json":"a20db2a5478fb7bd568bf6fcfbe7702dbe2a7a0390d9b36ec0fe2a6d07954d1b"}}, + {"executionId":"stock-harness.runner-opencode.local.ordered-comment-continuation","profile":"runner-opencode","case":"ordered-comment-continuation","sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceRef":"refs/heads/codex/stock-harness-candidate-cancelled-cells","workflowRunUrl":"https://github.com/paperclipai/paperclip/actions/runs/37045368302","cellAttempt":1,"status":"passed","failureClass":null,"causeCodes":[],"failedMatcherPaths":[],"model":"openrouter/deepseek/deepseek-v4-flash-0731","runtimeMode":"native","durationMs":95671,"providerDurationMs":72942,"cleanup":"passed","billing":{"llm":{"runCount":2,"runsWithTokenUsage":2,"runsWithReportedCost":2,"inputTokens":24355,"outputTokens":2690,"cachedInputTokens":160512,"totalTokens":187557,"reportedCostUsd":0.0043860217,"costStatus":"reported"},"runtime":{"provider":"local","agentRunDurationMs":72942,"leaseDurationMs":null,"leaseCount":0,"costStatus":"not_metered","costSource":"local_not_metered"},"reportedCostUsd":0.0043860217,"estimatedRuntimeCostUsd":0,"observedAndEstimatedCostUsd":0.0043860217,"complete":true},"billingTypes":["unavailable"],"stockChecks":{"budget-hard-stops":true,"default-hire-bundle":true,"provider-runs-present":true},"invocationMetrics":[],"prerequisite":{"sourceSha":"f02d8d0df327abb43b20c7e7beb86798239abbf5","sourceFingerprint":"ea488f1e5f67865c0e08e3628eb3081b86c6b3a8e40631702697ffac08359211","passed":true,"providerCallsBeforeAdmission":0,"passedTests":537},"evidenceHashes":{"result.json":"54f3029b581b82be87bd30b7670029bd5722915917e4bcab2e8ef0682bc3c0be","evidence-manifest.json":"b793fa707ccc2125b192816d59e922a0d43348c11eb516b8dff7e5d445b5545b","snapshots/context-integrity.json":"105ff5eb813ac3779e3ba04a8115ce1495c185011de4735782d3c4adba99787e","snapshots/stock-harness-hire.json":"4dde5055f48fff1b4cf466b4307739c967e2595e3c2e93d4e07230545741fa16","snapshots/stock-harness.json":"b0ce3eb0edce09ca4ef61af6f622cddd7c9bb1a3286c9ea7756dc228d00c10fa","snapshots/stock-harness-preflight.json":"78f2298b3bb4357b2a3ce6d75a922fc34d6449ef43378bb0dcc40476469daa01"},"taskDocumentCount":1,"behaviorCheckCounts":{"passed":12,"failed":0}} + ], + "totals": { + "baseline": { + "results": 24, + "passed": 15, + "failed": 9, + "missing": 0, + "cleanupPassed": 24, + "durationMs": 3250195, + "providerDurationMs": 2108034, + "llm": { + "runCount": 58, + "runsWithTokenUsage": 52, + "runsWithReportedCost": 43, + "inputTokens": 2778581, + "outputTokens": 68015, + "cachedInputTokens": 6908917, + "reportedCostUsd": 1.7494252618000001 + } + }, + "candidate": { + "results": 24, + "passed": 15, + "failed": 9, + "missing": 0, + "cleanupPassed": 24, + "durationMs": 2804913, + "providerDurationMs": 1540860, + "llm": { + "runCount": 47, + "runsWithTokenUsage": 45, + "runsWithReportedCost": 36, + "inputTokens": 2280185, + "outputTokens": 58933, + "cachedInputTokens": 4655025, + "reportedCostUsd": 1.7431513318 + } + } + }, + "matchedPassingTiming": { + "baseline": { + "cells": 13, + "medianCellDurationMs": 85550, + "medianProviderDurationMs": 57150 + }, + "candidate": { + "cells": 13, + "medianCellDurationMs": 69761, + "medianProviderDurationMs": 43862 + } + }, + "missingBaseline": [], + "cancellationRecovery": { + "sourceSha": "f02d8d0df327abb43b20c7e7beb86798239abbf5", + "originalRun": 37042856368, + "reason": "superseding pilot cancelled incomplete cells; completed behavior failures not rerun", + "executionIds": [ + "stock-harness.legacy-acp-codex.local.ordered-comment-continuation", + "stock-harness.legacy-opencode.local.ordered-comment-continuation", + "stock-harness.legacy-opencode.local.continuity-restart", + "stock-harness.runner-acpx-claude.local.continuity-restart", + "stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation", + "stock-harness.runner-acpx-claude.local.assigned-skill-explicit-invocation", + "stock-harness.runner-acpx-claude.local.ordered-comment-continuation", + "stock-harness.runner-codex.local.continuity-restart", + "stock-harness.runner-codex.local.ordered-comment-continuation", + "stock-harness.runner-codex.local.assigned-skill-explicit-invocation", + "stock-harness.runner-opencode.local.assigned-skill-explicit-invocation", + "stock-harness.runner-opencode.local.continuity-restart", + "stock-harness.runner-opencode.local.ordered-comment-continuation" + ] + }, + "baselineInfrastructureRecovery": { + "schema": "paperclip.stock-harness.infra-recovery/v1", + "sourceSha": "12c5433c67dc2e62916b879349c7ba2b6e0431f0", + "originalRun": 37042864888, + "originalJob": 110958906481, + "reason": "AWS runner shutdown, no result artifact; no completed model failure rerun", + "executionIds": [ + "stock-harness.legacy-opencode.local.ordered-comment-continuation" + ] + }, + "layoutRecovery": { + "resultEdits": 0, + "oracleEdits": 0, + "originalEvidenceRetained": true, + "manifests": [ + { + "campaign": "candidate-original", + "sha256": "9135457cdc339d2fcaeea52cd06622ecaea6a74561df73d4d7a0baa0d8c0e7cc" + }, + { + "campaign": "baseline-final", + "sha256": "85223428a0d473ba83e324621e80d711725ebe069c4f055665b0bf987d1f0c4c" + }, + { + "campaign": "candidate-recovery", + "sha256": "7788f879d0cf1d140d452cfa78fe085b34bc23d298d7c31cac9d630b12b2572f" + }, + { + "campaign": "baseline-recovery-final", + "sha256": "e39d62dd60f092e76f8776e844041f2458fca1989854d57ab13932a36c937ceb" + } + ] + }, + "limitations": [ + "Single trial per cell; not general coding quality or statistical proof.", + "Skill task document storage wording is ambiguous; original graders unchanged.", + "Persisted credential guard failures are not equivalent to poor task behavior.", + "Report costs are incomplete; zero reported native cost has unknown billing type.", + "Local runtime and interrupted partial attempts are not metered.", + "Historical public prompt retrieval is clipped in three ACP cells; missing suffix does not prove provider omission.", + "Reconstructed reports use separate copies with directory-layout-only repair; original public reports retain infrastructure failures." + ], + "fixtureMatchProof": { + "schema": "paperclip.stock-harness.fixture-match-proof/v1", + "cells": [ + { + "executionId": "stock-harness.legacy-codex.local.ordered-comment-continuation", + "fixtureConfigSha256": "f727e68de06dec501c49efc41ab941bd6d63826db82e615d5e60f2d20133b488", + "model": "gpt-5.6-sol" + }, + { + "executionId": "stock-harness.legacy-codex.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "ca51e8b29aab8e6efa32f2433c0f3960e43f9fc2f5edcd0660e753810d71f30d", + "model": "gpt-5.6-sol" + }, + { + "executionId": "stock-harness.legacy-codex.local.continuity-restart", + "fixtureConfigSha256": "8b1669039f3285f89a7f93c57e1c0dc12f071fbb2bbc8dcef6bf70bc14a4949a", + "model": "gpt-5.6-sol" + }, + { + "executionId": "stock-harness.legacy-claude.local.ordered-comment-continuation", + "fixtureConfigSha256": "94db17927ceddb54fe7766a362adcf78783726c68e1e136e407572aef6fed319", + "model": "claude-sonnet-4-6" + }, + { + "executionId": "stock-harness.legacy-claude.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "6c21be9bb3fcec6274372fc586d88e42d82bc0d438e26a69ef614f8e6fcb8519", + "model": "claude-sonnet-4-6" + }, + { + "executionId": "stock-harness.legacy-claude.local.continuity-restart", + "fixtureConfigSha256": "695071e5453c90288659f54f40fb307486f64d5fb8a401b767d2e9974faab15f", + "model": "claude-sonnet-4-6" + }, + { + "executionId": "stock-harness.runner-codex.local.ordered-comment-continuation", + "fixtureConfigSha256": "9e19859684aad0246fd009061101eec59c41179e51e92a9467f506af22558126", + "model": "gpt-5.6-sol" + }, + { + "executionId": "stock-harness.runner-codex.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "8fc61211aefb339a310469695d716c3faea15c2a547c543d70c599857a7e071d", + "model": "gpt-5.6-sol" + }, + { + "executionId": "stock-harness.runner-codex.local.continuity-restart", + "fixtureConfigSha256": "024e90354d760f90a40e2b3923468b72d36a9fabdb4fc58c10a7939a09e5cbc3", + "model": "gpt-5.6-sol" + }, + { + "executionId": "stock-harness.runner-opencode.local.ordered-comment-continuation", + "fixtureConfigSha256": "9dd6fef8fea77b07a67681acc1d7866fac9b4f16e23aca69eda32a4656cc260b", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + }, + { + "executionId": "stock-harness.runner-opencode.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "891eb6dd154aa5926a3de2eae79780661b436618b6eac7733998750c6c341ba3", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + }, + { + "executionId": "stock-harness.runner-opencode.local.continuity-restart", + "fixtureConfigSha256": "441c2fa8bdd26cf14ac2440311b99a71b7dd6202698cafc3d01ed7a91dc4b004", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + }, + { + "executionId": "stock-harness.runner-acpx-claude.local.ordered-comment-continuation", + "fixtureConfigSha256": "77cc32714b23852d3ef67c687202d89f7513275500b3c522a0a1de3d745c9c4e", + "model": "claude-sonnet-5" + }, + { + "executionId": "stock-harness.runner-acpx-claude.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "c190a1e731ef450ea6ae0c24b5a8ae0f21db725b47350645cdbb9273415c5acd", + "model": "claude-sonnet-5" + }, + { + "executionId": "stock-harness.runner-acpx-claude.local.continuity-restart", + "fixtureConfigSha256": "40876df6f88f34ba8232ccba8816bf7e1173be0042a90a174f79cddcca346c4c", + "model": "claude-sonnet-5" + }, + { + "executionId": "stock-harness.legacy-acp-codex.local.ordered-comment-continuation", + "fixtureConfigSha256": "22500c39874cda62a2512ac3d8a43309ce5879f55f8515e31dd8ac632edb6c93", + "model": "gpt-5.6-sol" + }, + { + "executionId": "stock-harness.legacy-acp-codex.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "896d8602d2990bf7517a4392d66909c6649e0f00f79a5ea0bb8f0e370b7b597a", + "model": "gpt-5.6-sol" + }, + { + "executionId": "stock-harness.legacy-acp-codex.local.continuity-restart", + "fixtureConfigSha256": "7b66845e231674cc6e78f1d86e863dc03005998c2674316af50db60ec53b7e4a", + "model": "gpt-5.6-sol" + }, + { + "executionId": "stock-harness.legacy-acp-claude.local.ordered-comment-continuation", + "fixtureConfigSha256": "6c1b9d7debb306af405640eb381e1709ae7185e7925bc2bd3ef31acf98fbff8f", + "model": "claude-sonnet-4-6" + }, + { + "executionId": "stock-harness.legacy-acp-claude.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "53f308d6fbdc30adc60370279ab38679b5d87075899be4db6c0730593e1632c4", + "model": "claude-sonnet-4-6" + }, + { + "executionId": "stock-harness.legacy-acp-claude.local.continuity-restart", + "fixtureConfigSha256": "80ec22c5080ef1a71fc97547c3e78f4bc5f54f121f46f982506f2f9f773b27d0", + "model": "claude-sonnet-4-6" + }, + { + "executionId": "stock-harness.legacy-opencode.local.ordered-comment-continuation", + "fixtureConfigSha256": "e015ded0177d8d4c893acc658e25699b28937d97cb9d58da023ddff7ba7eac0b", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + }, + { + "executionId": "stock-harness.legacy-opencode.local.assigned-skill-explicit-invocation", + "fixtureConfigSha256": "dd673f469cab583d26bc2e4b484f18b27884ec14a59542177026389181e2da2c", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + }, + { + "executionId": "stock-harness.legacy-opencode.local.continuity-restart", + "fixtureConfigSha256": "3dadc83dff00ef2919cd24c6328ea9507f971e7d294e9ece699ff2ee6fa98230", + "model": "openrouter/deepseek/deepseek-v4-flash-0731" + } + ] + }, + "workflowProvenance": [ + { + "run": 37042856368, + "workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a", + "workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37042856368" + }, + { + "run": 37042864888, + "workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a", + "workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37042864888" + }, + { + "run": 37045368302, + "workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a", + "workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37045368302" + }, + { + "run": 37044967981, + "workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a", + "workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37044967981" + }, + { + "run": 37048838402, + "workflowSha": "7d59de6113caf5219d34d65e1fb5f0c1a7fbdc2a", + "workflowUrl": "https://github.com/paperclipai/paperclip/actions/runs/37048838402" + } + ] +} diff --git a/doc/plans/2026-10-02-stock-harness-live-comparison.md b/doc/plans/2026-10-02-stock-harness-live-comparison.md new file mode 100644 index 0000000000..42c7ae3fa3 --- /dev/null +++ b/doc/plans/2026-10-02-stock-harness-live-comparison.md @@ -0,0 +1,110 @@ +# Stock-harness live comparison — 2026-10-02 + +**TL;DR:** Two newly failing paired cases: classic Claude/OpenCode document delivery. Two newly passing cases: classic and native OpenCode ordered continuation. Seven unchanged failures and 13 unchanged passes. The extra ACP Claude storage symptom occurs beneath an existing credential-guard failure. All 48 results are available; no general performance equivalence is established. + +The reduced instructions are **not yet qualified for merge**. Classic Claude and classic OpenCode pass the historical skill-output case but fail with the reduced instructions: they finish without saving a Paperclip task document. Native Codex and native Claude pass all three journeys in both variants. Single trials and ambiguous fixture storage wording limit causal attribution; the failed oracle remains unchanged. + +The reductions and evaluation setup are in draft [PR #14948](https://github.com/paperclipai/paperclip/pull/14948). This report and its [safe evidence projection](2026-10-02-stock-harness-live-comparison.json) record the measured revisions rather than claiming the final documentation head was run through the full matrix. + +## Revisions and matched inputs + +- Candidate: `f02d8d0df327abb43b20c7e7beb86798239abbf5`, eight-word hire manual plus reduced shared startup/resume instructions. +- Historical baseline: `12c5433c67dc2e62916b879349c7ba2b6e0431f0`, the same evaluation harness with only the prior default manual and shared prompt source restored from `e00d10d5d5594f6e2d1e8cf4e0ec81c074b0bd88`. Its manifest records ten identical behavioral/fixture source hashes. +- Native Codex fix [#14920](https://github.com/paperclipai/paperclip/pull/14920) is held constant. This comparison does not measure its before/after performance. +- Same eight profiles, three journeys, provider models, effort configuration, skills, tools, permissions, credentials, local runtime, and independent graders. Effort inherits provider defaults where the existing profile does not set it. +- Each variant selects 24 cells and initially budgets 48 turns, with existing bounded retries and 1,000-cent company/agent hard stops. Recorded run-ledger counts can differ after early failures, retries and cleanup. + +## Per-profile and case outcomes + +These are recorded overall qualifications, including security and receipt failures. A credential-guard failure does not imply every behavioral matcher failed. + +| Profile | Model | Journey | Historical | Reduced | +| --- | --- | --- | --- | --- | +| `legacy-codex` | `gpt-5.6-sol` | `assigned-skill-explicit-invocation` | Pass | Pass | +| `legacy-codex` | `gpt-5.6-sol` | `ordered-comment-continuation` | Pass | Pass | +| `legacy-codex` | `gpt-5.6-sol` | `continuity-restart` | Pass | Pass | +| `legacy-claude` | `claude-sonnet-4-6` | `assigned-skill-explicit-invocation` | Pass | Fail: no Paperclip document | +| `legacy-claude` | `claude-sonnet-4-6` | `ordered-comment-continuation` | Pass | Pass | +| `legacy-claude` | `claude-sonnet-4-6` | `continuity-restart` | Fail: chat memory | Fail: chat memory | +| `legacy-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `assigned-skill-explicit-invocation` | Pass | Fail: no Paperclip document | +| `legacy-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `ordered-comment-continuation` | Fail: deadline/run-start timeout | Pass | +| `legacy-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `continuity-restart` | Pass | Pass | +| `legacy-acp-codex` | `gpt-5.6-sol` | `assigned-skill-explicit-invocation` | Fail: credential guard | Fail: credential guard | +| `legacy-acp-codex` | `gpt-5.6-sol` | `ordered-comment-continuation` | Fail: credential guard, receipt incomplete | Fail: credential guard | +| `legacy-acp-codex` | `gpt-5.6-sol` | `continuity-restart` | Fail: credential guard, receipt incomplete | Fail: credential guard, receipt incomplete | +| `legacy-acp-claude` | `claude-sonnet-4-6` | `assigned-skill-explicit-invocation` | Fail: credential guard | Fail: credential guard, no Paperclip document | +| `legacy-acp-claude` | `claude-sonnet-4-6` | `ordered-comment-continuation` | Fail: credential guard, receipt incomplete | Fail: credential guard | +| `legacy-acp-claude` | `claude-sonnet-4-6` | `continuity-restart` | Fail: credential guard, receipt incomplete | Fail: credential guard | +| `runner-codex` | `gpt-5.6-sol` | `assigned-skill-explicit-invocation` | Pass | Pass | +| `runner-codex` | `gpt-5.6-sol` | `ordered-comment-continuation` | Pass | Pass | +| `runner-codex` | `gpt-5.6-sol` | `continuity-restart` | Pass | Pass | +| `runner-acpx-claude` | `claude-sonnet-5` | `assigned-skill-explicit-invocation` | Pass | Pass | +| `runner-acpx-claude` | `claude-sonnet-5` | `ordered-comment-continuation` | Pass | Pass | +| `runner-acpx-claude` | `claude-sonnet-5` | `continuity-restart` | Pass | Pass | +| `runner-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `assigned-skill-explicit-invocation` | Pass | Pass | +| `runner-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `ordered-comment-continuation` | Fail: deadline/run-start timeout | Pass | +| `runner-opencode` | `openrouter/deepseek/deepseek-v4-flash-0731` | `continuity-restart` | Pass | Pass | + +## Findings and next step + +The classic Claude and classic OpenCode skill runs read the pinned skill marker and complete the task, but save zero Paperclip documents; both historical runs save exactly one. The legacy ACP Claude skill run has the same document difference under its credential-guard failure. The earlier reduced-instruction Claude pilot also failed this document oracle. The pinned skill asks for a “task document” without explicitly naming Paperclip storage, so this may reveal both reliance on the former manual and a fixture ambiguity. It remains a failed delivery result, not a regraded pass. + +Keep the agreed tiny manual. Next, verify that legacy agents discover and follow the existing **Paperclip skill/API path for durable task documents and artifacts**. The skill already prohibits file-only handoff, lists issue-document routes and gives a plan-document example; audit that delivery before adding a concise generic-document recipe if needed. Run the same preserved case plus a clearly storage-specific case on both variants. Legacy agents do not receive native `paperclip_finish`/`paperclip_block` tools. Native tool descriptions are a separate item 2.3 change; guidance must follow the actual runtime capability. No production or skill instructions were changed to make these runs pass. + +Classic Claude chat fails the memory assertion in both variants. Six legacy ACP cells per variant fail the persisted-credential guard: the scanner finds a scoped provider credential in persisted ACP session state. Raw session files and credential values are not published. These are existing qualification failures requiring their own investigation; they cannot be waived or interpreted as instruction-reduction quality evidence. + +Historical native OpenCode ordered comments hit the case deadline without the required continuation outcome. The initial historical classic OpenCode ordered cell lost its AWS runner before result upload; its one unchanged-source recovery also hits the 12-minute deadline. Both OpenCode ordered cases pass in the candidate. Completed failures are never rerun to select a better result. Those candidate passes against historical timeouts are observed differences, not proof of a general improvement. + +Historical public prompt retrieval is clipped in three legacy ACP cells (four invocations), including identity/connection suffix text. Those structural receipt failures cannot establish that the provider omitted instructions. The candidate ACP Codex chat run also lacks a complete per-run invocation receipt. Instruction presence/absence claims are bounded by the retained public receipts. + +## Timing, usage and cost + +| Variant | Retained cells | Pass / fail / missing | Cell duration sum | Provider duration sum | Recorded runs with tokens / with reported cost / total | Reported LLM subtotal | +| --- | ---: | --- | ---: | ---: | --- | ---: | +| baseline | 24 | 15 / 9 / 0 | 3250.195s | 2108.034s | 52 / 43 / 58 | $1.749425 | +| candidate | 24 | 15 / 9 / 0 | 2804.913s | 1540.860s | 45 / 36 / 47 | $1.743151 | + +For the 13 cells that pass both variants, median cell time is 85.550s historical versus 69.761s reduced; median provider time is 57.150s versus 43.862s. This excludes failures and is descriptive, not a reliable speedup estimate. + +Provider billing is incomplete. Legacy Codex/ACP usage is unpriced in several runs; native Codex/Claude often report zero with billing type `unknown`. Reported subtotals do not prove zero native spending or invoice totals. Local runner costs and interrupted partial runs are not metered. Diagnostic/pilot costs and repeated interrupted work are outside the matched-cell subtotal and remain retained. No aggregate cost-saving claim is justified. + +## Campaign and evidence history + +| Campaign | Source | Outcome / evidence | +| --- | --- | --- | +| [Initial diagnostic Codex](https://github.com/paperclipai/paperclip/actions/runs/37034213743) | `a63437069` | Skill passes; predates mandatory admission. | +| [Cold full attempt](https://github.com/paperclipai/paperclip/actions/runs/37037105491) | `36e987246` | Missing SDK build; cancelled before provider admission. | +| [SDK pilot](https://github.com/paperclipai/paperclip/actions/runs/37039240025) | `4163dbfd0` | Missing daemon; stops before providers. | +| [Daemon pilot](https://github.com/paperclipai/paperclip/actions/runs/37040493183) | `ac6ddefb5` | Missing fake Codex fixture; stops before providers. | +| [Complete-build Claude pilot](https://github.com/paperclipai/paperclip/actions/runs/37041741124) | `f02d8d0df` | 537 prerequisites pass; skill document oracle fails; reported $0.175308. | +| [Initial candidate matrix](https://github.com/paperclipai/paperclip/actions/runs/37042856368) | `f02d8d0df` | 4 pass, 7 fail, 13 cancelled when a same-target-branch pilot superseded it. Original attempts retained. | +| [Candidate cancelled-cell recovery](https://github.com/paperclipai/paperclip/actions/runs/37045368302) | `f02d8d0df` | Only the 13 cancelled cells: 11 pass, 2 fail. Combined candidate has 24 results, 15 pass / 9 fail. | +| [Historical matrix](https://github.com/paperclipai/paperclip/actions/runs/37042864888) | `12c5433c6` | 15 pass, 8 recorded failures, one AWS-runner shutdown without result upload. | +| [Historical interrupted-cell recovery](https://github.com/paperclipai/paperclip/actions/runs/37048838402) | `12c5433c6` | Only legacy OpenCode ordered comments: 720.616s deadline failure; evidence and cleanup valid. Historical cohort now has 24 results, 15 pass / 9 fail. | +| [Current packaging pilot](https://github.com/paperclipai/paperclip/actions/runs/37044967981) | `1eb5ba420` | Attempt 1 stopped before cell/provider execution; attempt 2 passes, evidence valid and cleanup pass. | + +The same-target-branch workflow concurrency rule caused the candidate interruption; that orchestration mistake was acknowledged and corrected with a separate unchanged-source recovery branch. No completed model failure was retried. The historical missing cell is recovered once for a runner shutdown. Partial/failed attempts remain part of the history and unknown-spend accounting. + +The measured matrix packages contain a prerequisite folder beside the exact campaign root. The trusted selector correctly rejects this layout, so their original published reports synthesize missing-result infrastructure failures. A separate local reconstruction nests that prerequisite folder under the campaign root and verifies byte identity for every copied file; result, scorer, usage, screenshot and prerequisite content are unchanged. The unchanged trusted selector then selects 11/24 original candidate, 13/13 recovery 23/24 original historical packages and 1/1 interrupted-cell recovery; recovered normalization validates every retained result’s evidence. Original missing packages stay missing in their original campaign; the separate recovery result has its own provenance. This is declared directory-layout recovery, not a canonical successful publication or a change to the oracle. The safe JSON records result/receipt/scoring hashes and recovery-manifest hashes. + +The fix at `1eb5ba420` writes prerequisites beneath the exact campaign root. Its single selected Codex skill cell passes through the protected report and publication pipeline: [interactive pilot report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37044967981-2/index.html), [GitHub evidence](https://github.com/paperclipai/paperclip/actions/runs/37044967981#artifacts). It has `evidenceValid=true`, cleanup pass, 27.149s provider / 49.523s cell, and a reported zero subtotal with unknown billing type. It does not replace the full matched cohort or qualify untested providers. + +Original historical/recovery dashboards retain the publication failure: [historical](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37042864888-1/index.html), [candidate recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37045368302-1/index.html). Inspect their access-controlled raw artifacts for measured cell evidence; do not read synthetic missing-result rows as task performance. + +## Verification and remaining qualification + +At `1eb5ba420`, all 557 credential-free prerequisites pass (556 TypeScript plus one Rust), along with 895 E2E support tests, E2E typecheck and 24-cell discovery. The cold builder prepares SDK/shared outputs and both daemon/fake protocol binaries, records hashes, and fails closed before credentials. The 313 filtered native tests are not counted as passing coverage. This head has 52 passing CI gates and two skipped gates; Greptile reviewed that exact head at 5/5 with zero open threads. Later documentation heads require fresh checks/review; measured sources remain explicit above. + +Repository typecheck and build pass. The complete local Vitest run has 14,870 passed, 83 skipped and three unrelated timing failures; the three affected files pass unchanged narrow reruns (25 auth, 16 reviewed-chat binding, four webhook tests). Original failures are retained rather than calling the first full run green. + +PR #14948 stays draft while document delivery is unresolved. Unrepresented harnesses, legacy ACP credential safety, public-receipt clipping, the independent ACP base-replacement issue, existing Claude chat memory behavior and saved-session migration are not qualified away by these results. This small matrix measures skill/context/chat behavior, not broad coding quality or statistical equivalence. + +## Retained delivery diagnosis and approved repair + +Both failed classic agents loaded the operational Paperclip skill. This is not an observed skill-discovery failure. In the reduced Claude run, the agent wrote a workspace Markdown file before loading Paperclip and then completed the issue without a document API write. The reduced OpenCode run loaded Paperclip, accessed the assigned skill through the public skills API, and wrote a workspace task-document file; it made no document, attachment, or work-product delivery write. Both corresponding historical runs saved a Paperclip issue document. + +Claude received the full operational skill, including its existing no-local-only delivery guidance. OpenCode's skill result was explicitly truncated in both variants: the early artifact rule survived, but the later generic document endpoint and plan-only write example were omitted. That truncation is existing behavior, not a newly measured regression. Removing always-on delivery reminders may interact with it, but the combined manual/shared reduction and ambiguous assigned-skill wording prevent a causal attribution to one layer. + +Dotta approved a small early API-runtime skill recipe and a generic issue-document reference. The tiny manual remains eight words. The recipe checks the successful write's returned saved revision/content and links the returned key; it does not require another GET after a clear valid receipt. It preserves explicit destinations, downloadable-file delivery, and native document-tool boundaries. + +The focused repair comparison will hold the tiny manual/shared prompts fixed and vary only `skills/paperclip/SKILL.md` plus the new `references/issue-documents.md`. Classic Claude and OpenCode each run the preserved original case and an added explicitly Paperclip-storage case: four cells per variant, eight expected provider turns in total. This measures the skill repair separately from the original reduction. The explicit case independently checks public saved body/revision and an agent comment linking the exact same-app document; local-only, missing-revision, wrong-document, local-path, and other-origin outcomes cannot pass. Current state: fixture preparation and source freeze, no repair provider run yet. diff --git a/doc/plans/2026-10-02-stock-harness-paperclip-checklist.md b/doc/plans/2026-10-02-stock-harness-paperclip-checklist.md index f71067ff07..38141901dc 100644 --- a/doc/plans/2026-10-02-stock-harness-paperclip-checklist.md +++ b/doc/plans/2026-10-02-stock-harness-paperclip-checklist.md @@ -1,7 +1,17 @@ # Stock harness, with Paperclip: working checklist -Created: 2026-10-02. Status: item 1 implemented for native Codex app-server; -PR validation in progress. Other implementation items remain open. +Created: 2026-10-02. Status: item 1 merged for native Codex app-server in +[PR #14920](https://github.com/paperclipai/paperclip/pull/14920). Item 2's default +hire manual is reduced to identity only, and common legacy startup/resume +instructions are in [PR #14948](https://github.com/paperclipai/paperclip/pull/14948). +GitHub live qualification retains earlier document-delivery failures; the latest +focused OpenCode comparison passes both cases in both variants but still exposes +deficient original-case handoff links. PR #14948 remains draft. The original +[live report](2026-10-02-stock-harness-live-comparison.md), completed +[skill repair](2026-10-02-legacy-document-skill-repair.md) and latest +[stock-selection/link comparison](2026-10-02-opencode-skill-routing-link-qualification.md) +retain their separate sources and limitations. Additional carriers and fixed +native instructions remain open; hiring templates are being reduced separately. Goal: keep the agent's stock harness behavior and add only what it needs to work with Paperclip. Apply this across legacy adapters, the new Runner, and their @@ -12,7 +22,11 @@ record the proposed behavior and Dotta's direction, make a bounded change, and verify the affected paths before moving on. New findings get stable IDs in the ledger below so they do not disappear into conversation history. -**Current item: 1 — validate and review the native Codex instruction fix.** +For every change, record its executable test/eval coverage before continuing. +Distinguish coverage setup, deterministic results, and measured live results; +configured cells alone do not qualify behavior. + +**Current item: 2 — reduce the default operating manual and shared prompt layers.** ## Agreed direction and boundaries @@ -43,7 +57,7 @@ ledger below so they do not disappear into conversation history. - [x] Remove default replacement of stock base instructions across those paths. - [x] Verify the actual app-server request and retained session instructions, including resume; checking only a prompt builder is insufficient. -- [ ] Verify Paperclip task context, tools, auth, assigned skills, and completion +- [x] Verify Paperclip task context, tools, auth, assigned skills, and completion still work. Record applicable regressions and eval results. Starting points: [Codex backend](../../packages/paperclip-runner/src/backends/codex-native-backend.ts), @@ -61,14 +75,20 @@ and need a provider session reset. Do not reset active sessions automatically. The separate Codex-through-ACP dependency patch remains a coverage follow-up under item 6; this change covers the native app-server path. +Verification: 412 focused TypeScript/Rust tests, repository typecheck/build, +55 passing PR checks, and Greptile 5/5. The original full local test command +was stopped after setup failures; affected suites passed on rerun. No paid +live campaign was run, and no task-quality improvement is claimed. + ## 2. Reduce the default operating manual and shared prompt layers - [ ] Inventory what an agent actually receives: hire instructions, shared prompt template, wake context, runtime prompt, bootstrap, and loaded skills. Separate always-present text from content loaded on demand. -- [ ] Agree on the tiny common contract and the coordination details that each - runtime still requires. Legacy API coordination and native semantic tools - need appropriate instructions for their respective interfaces. +- [x] Agree on the tiny default: identity only, with no skill or native-tool + pointers. The harness supplies coordination instructions. +- [x] Reduce the generic default hire `AGENTS.md` to the agreed sentence and + update the existing creation test to expect the minimal bundle. - [ ] Remove repeated workflow rules and stock coding/style/autonomy guidance. - [ ] Move detailed planning, hiring, artifacts, and exceptional procedures to discoverable references or tools where feasible. @@ -80,7 +100,130 @@ Starting points: [default hire instructions](../../server/src/onboarding-assets/ [native runtime contract](../../packages/paperclip-runner/src/contracts/runtime-context.ts), [Paperclip operational skill](../../skills/paperclip/SKILL.md). -Decision: exact retained paragraph and on-demand boundaries pending. +Decision: Dotta confirmed the harness already handles skill and tool delivery; +the default hire manual is now only "You are an agent in a Paperclip company." +This replaces 602 words with eight. The shared loader supplies this default +across instruction-bundle-capable adapters for non-CEO hires without explicit +instructions. Existing saved bundles retain their content; CEO, first-agent, +and role/team templates remain separate work under item 3. The common +prompt/wake reduction is complete locally below; additional carriers, native +instructions, and skill/reference corrections remain open. + +Verification: the existing agent-skills route suite passed all 54 tests, +including default creation, custom bundles, and CEO/first-agent paths. The +onboarding asset suite passed all 10 tests. This initially had only deterministic +coverage. Dotta subsequently requested behavioral eval coverage for every change; +the dedicated production-default-hire suite below closes the custom QA manual +gap. Full repository typecheck/build/test were not rerun for this narrow asset +and existing-test update. + +### Shared prompt follow-ups + +- [x] **2.1 Reduce common legacy startup/resume instructions.** Keep + identity and connection guidance in the shared task/chat defaults; remove the + generic resumed-wake execution contract. Preserve current task/event data, + specialized wake contracts, custom templates, skills, and auth. +- [ ] **2.2 Review additional legacy carriers.** Reduce Hermes local/gateway + wrappers, review Pi system delivery, and check OpenClaw fresh-wake framing. + Preserve transport facts and user configuration. Dotta deferred this item on + 2026-10-02 for a later revisit; it remains open. +- [ ] **2.3 Reduce native Runner instructions and constraints.** Improve + discoverable tool documentation first, then shorten fixed guidance and + consolidate completion rules. Verify prompt revisions, digests, and session + compatibility across native Codex, ACPX, and OpenCode. Dotta approved the + native tool-documentation slice on 2026-10-02 in a separate worktree. Native + `paperclip_finish`/`paperclip_block` guidance must not leak into legacy + completion paths, which use the operational skill and API. + +### Follow-up 2.1 implementation and verification + +The common task and conversation defaults now share the same identity sentence +and unchanged connection guidance: 113 words each, down from 661 and 205. +Connection guidance accounts for 108 of those words. The ordinary resumed-wake +execution contract is removed (172 words), including the old opt-in used by +OpenClaw on fresh turns. `includeExecutionContract` remains accepted as a +deprecated no-op for adapter/plugin source compatibility. + +Current task facts, ordered comments, work modes, approval/review state, +checkout, holds/blockers, recovery/watchdog roles, and external-chat contracts +remain in their existing owners. Skill delivery, runtime authentication, +custom `promptTemplate`, and `bootstrapPromptTemplate` mechanics are unchanged. +Hermes's own wrappers and Pi's system carrier remain follow-up 2.2; native +fixed instructions and full-turn constraints remain follow-up 2.3. + +The later read-only 2.2 audit found another OpenClaw-owned HTTP identity, +checkout/status, delegation, plan-approval and task-discovery wrapper in +`buildWakeText`. Its short conversation branch is selected when optional task +Markdown is present, including on ordinary task dispatch. Existing dispatch +coverage does not assert the absence of conversation waiting language. This is +a framing concern to verify, not a measured failure. Pi already uses additive +`--append-system-prompt` delivery and suppresses the duplicate user-prompt copy; +its next step is a delivery audit with minimal changes. These findings are +deferred with 2.2 and are not implemented in PR #14948. + +The new defaults apply when prompts are assembled after deployment. Existing +provider sessions can retain earlier startup instructions in their history +until reset; no active session or saved custom template is rewritten here. + +Verification: 563 distinct focused tests passed across 18 suites: + +- Shared prompt selection/rendering, operational-skill selection, and retained + specialized wake contracts (130 tests). +- ACPX, Pi, OpenCode, Cursor Cloud, and Codex local execution, including + task/chat fresh/resumed/reset/fallback delivery, custom prompts, skill + mounting, and authentication (259 tests). +- Hermes local prompt rendering and gateway execution (41 tests). +- Claude, Gemini, Cursor local, Grok, Kimi, OpenClaw, and Claude/Codex ACP + fallback execution (113 tests). +- Native execution input, preserving shared wake compatibility (20 tests). + +`@paperclipai/adapter-utils` typecheck and build passed; `git diff --check` +passed. Full repository typecheck/build/test and live behavioral evals were not +run for this local bounded change. These checks prove prompt and runtime +mechanics, not an improvement in coding-task quality. + +### Remaining shared layers inspected (2026-10-02) + +The table records the initial inspection before follow-up 2.1. Word counts +cover fixed source text, not complete assembled prompts or token counts. They +exclude task data, assigned agent instructions, and skills. This inspection +establishes instruction mechanics, not a performance result. + +| Layer | Delivery at initial inspection | Disposition | +| --- | --- | --- | +| Shared legacy task/chat templates | `DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE` was 661 words and the conversation template was 205, including connection guidance. Used by Claude, Codex, Cursor local/cloud, Gemini, Grok, Kimi, OpenCode, Pi, and Hermes local. | Completed in 2.1: both defaults now 113 words. Repeated procedures removed, connection guidance retained, no replacement skill pointer. | +| Generic resumed wake contract | `renderPaperclipWakePrompt` added a 172-word execution contract on ordinary resumed turns. OpenClaw requested it on fresh turns too because it has no shared template. | Completed in 2.1: generic contract removed, including legacy opt-ins. Task/event data and conditional runtime contracts retained. | +| Additional legacy carriers | Hermes local adds a 171-word identity/API/curl wrapper before the shared task template; Hermes gateway supplies its own four-rule execution contract. Pi carries shared defaults in its system extension and suppresses the wake copy when that extension owns policy. | Separate follow-up 2.2; changing the shared constants alone does not remove every wrapper. Preserve transport facts and custom configuration. | +| Native fixed prompt | The 262-word `paperclip-execution.v5` prompt carries hiring, delegation/dependencies, connection setup, and finalization guidance across native Codex, ACPX, and OpenCode backends. | Keep one short native completion/dependency contract. Relocate uncommon procedures to the relevant tool documentation, improving that documentation where needed before removing instructions. Any change needs prompt revision/digest/session-compatibility verification. | +| Native full-turn constraints | `nativeTaskConstraints` adds assigned-skill limits, current agent-file paths, document/file delivery, answered-question handling, and accepted completion/final-response sequencing. Codex/OpenCode add another completion reminder. Prepared constraints reach the model envelope through `native-session-runtime`; compact continuations avoid replaying prior task constraints. | Consolidate duplicate completion instructions. Keep current paths, answer scope, and the required result protocol. Move detailed file/document procedure to discoverable tool descriptions where adequately covered. | +| Task/event-specific framing | Server task Markdown and wake rendering supply work mode, ordered comments, approved revisions, checkout/holds/blockers, recovery/review/watchdog roles, external-chat bookkeeping/delivery, and conversation handoff. | Retain the facts and mode boundaries in the first reduction. Use one authoritative owner for shared mode directives; review specialized wording separately. | + +Custom `promptTemplate` and fresh-only `bootstrapPromptTemplate` configuration +are distinct from shipped defaults and should not be silently rewritten. +The built-in process adapter invokes a configured command; HTTP passes context +as JSON. They do not use the shared text template. External adapter plugins and +hosted configuration variants still require the broader item 6 audit. + +An accepted-plan discrepancy needs resolution: task Markdown permits cohesive +implementation on the current issue, while the wake renderer's planning-mode +accepted-confirmation branch requires children and prohibits implementation on +the source issue. Existing tests assert both versions. Production reachability +after the server changes work mode has not yet been established; do not claim +this proves a live plan-continuation failure. + +Dotta selected follow-up 2.1 first: reduce the common legacy startup/resume +copies, preserving conditional context. Additional carriers and native fixed +instructions remain distinct follow-ups 2.2 and 2.3. No shared runtime text was +changed by the initial inspection. + +Coverage identified during inspection includes shared prompt selection/rendering, +adapter execution tests, native backend/runtime context tests, and server +task-context/native-input tests. Follow-up 2.1 extended the shared and adapter +checks; its results are above. Follow-ups 2.2 and 2.3 need their corresponding +carrier/protocol assertions. Product E2E context-integrity, completion-updates, +and direct blocker guidance are candidate lifecycle checks; their current scope +does not establish coding-task quality. The initial inspection itself ran no +tests or live evals. ## 3. Review every hiring and role template @@ -107,7 +250,61 @@ Starting points: [onboarding assets](../../server/src/onboarding-assets/), [hiring skill and references](../../skills/paperclip-create-agent/), [teams catalog](../../packages/teams-catalog/catalog/). -Decision: per-template disposition and migration policy pending. +Dotta approved and merged [PR #14985](https://github.com/paperclipai/paperclip/pull/14985) +at `862a5758ba0e88a33232c1f1fa645e85c38a3113`, after corrected source +`57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25` passed full CI and fresh Greptile 5/5. +The hiring skill/references, CEO/CoS assets and team catalog are now on master. +Their live qualification remains separate from unit/CI verification. +[Prompt diffs and preserved scope](https://github.com/paperclipai/paperclip/blob/57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25/doc/plans/2026-10-02-hiring-template-prompt-diff.md). + +The read-only audit found four onboarding CEO files (1,897 words), a 164-word +CoS manual, four role examples (coder 652, QA 619, UX 1,325, security 1,724 words), +and eight team-catalog manuals. Hiring step 6, the 60–150-line role guide and +review checklist would regenerate the removed operating policies. + +- [x] **3.1** Reduce role drafting rules, coder/adapted hire examples and matching + catalog coder, preserving metadata, auth, permissions and skill selections. +- [x] **3.2** Reduce the CEO bundle/copy, including forced delegation, hiring, + memory and repeated API recipes. +- [x] **3.3** Review remaining roles/catalog and CoS copy. +- [ ] **3.4** Review specialized built-in/plugin contracts separately; retain + product-required behavior rather than assuming every rule is redundant. + +Custom and existing bundles and governance remain deliberate rollout boundaries. +Specialized Summarizer, Reflection Coach and Wiki Maintainer prompts remain +unchanged for separate product-contract review; optional Content Lead was already +tiny. The generic non-CEO fallback belongs to PR #14948. PR #14985 does not reduce +every shipped specialized agent. + +Its explicit `hiring-templates` suite selects native local Codex and ACPX Claude +`hire-coder-template-reuse` cells: a real default CEO, explicitly requested +production hiring skill/coder reference, independently scored saved JSON +artifacts and session reuse. Outcome and source/read coverage are separate; +missing or redacted receipts do not prove no regression. The existing first-task +suite covers actual wizard CoS snapshot selection. The matched union holds the +reduced manual/shared/operational skill and native completion guidance constant, +restoring only 21 historical hiring production/derived sources in baseline. +Frozen candidate `9f5404ad3aacbe76777952759414d34fd381e674` and baseline +`296a4df85e8bcc97a160fc78c291b17adb828196` passed 705 exact-source credential-free +prerequisites each (704 TypeScript + 1 Rust; zero providers), with 8,242 identical +other tracked files and matching fixture/model/auth/permissions configuration. +Only four candidate-specific single-file CEO selection assertions are filtered +symmetrically; each variant's bundle/source hashes and canonical catalog are +independently admitted. Two cells per variant, five expected turns each: +four cells / 20 turns, 15 minutes each, no automatic retries or broader selection. +The [candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208) +and [baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463) +protected campaigns are dispatched. No graded live pairs yet; pending evidence +cannot establish non-regression. +[Inspect the partial report](https://github.com/paperclipai/paperclip/blob/e3720f369df070036526dd295a7a68e49488c51a/doc/plans/2026-10-02-hiring-template-live-comparison.md). +Report-only updates preserve the measured refs and merged hiring PR. + +The default-manual PR is replayed on the merged hiring base, preserving custom +CEO bundle checks, the minimal generic boundary and both eval suites. Its prior +browser failure remains retained: a deterministic process fixture replayed its +last plan command on an asynchronous `chat_task_completed` wake and wrote a +second, text-identical source plan revision. It was not paid/provider execution; +fresh rebased CI remains required before readiness. ## 4. Fix repository context while retaining Paperclip configuration @@ -154,7 +351,7 @@ does not apply. | Legacy Codex / Claude | Local adapters, managed auth and isolated configuration | Initial inspection only | | Other legacy local adapters | ACPX, OpenCode, Pi, Cursor, Gemini, Grok, Kimi, Hermes | Pending | | Other adapter transports | Cursor Cloud, Hermes/OpenClaw gateways, process, HTTP, external adapter plugins | Pending | -| Runner Codex | App-server driver, runnerd bridge, Rust provider, direct/fallback paths | Additive instruction fix implemented; PR validation pending | +| Runner Codex | App-server driver, runnerd bridge, Rust provider, direct/fallback paths | Additive instruction fix merged in PR #14920 | | Runner ACPX | Enabled profiles, especially Claude/Grok; declared or pending profiles tracked separately | Claude initial inspection; remaining audit pending | | Runner OpenCode | Native provider and configuration paths | Pending | | Hosted/remote providers | Claude Managed and AWS AgentCore; identify their own baseline rather than assuming CLI semantics | Pending | @@ -172,15 +369,18 @@ does not apply. ## Verification and evals — apply to each item -- [ ] Map existing coverage before adding cases; use [doc/evals.md](../evals.md). +- [x] Map existing coverage for all implemented changes before adding cases; + see the SH-1–SH-3 map below and [doc/evals.md](../evals.md). Keep Runner protocol evals and Product E2E evals distinct. -- [ ] Start with narrow deterministic checks for instruction layering, effective +- [x] Set up narrow deterministic checks for the implemented instruction layering, + hire, and shared-prompt changes. Future changes still need their own map for effective configuration, skill/auth delivery, and session behavior where appropriate. -- [ ] Use Runner evals for provider/session/tool protocol changes; use Product - E2E for real hiring, repository context, task lifecycle, and artifact delivery. -- [ ] Review existing context-integrity, hiring, completion-updates, and blocker - suites for reusable coverage; record gaps rather than claiming coverage from - a similarly named case. +- [x] Reuse existing protocol tests for native Codex request layering and add Product + E2E for production-default hiring, skill delivery, task lifecycle, and chat + continuity. Repository context and additional artifact cases remain future work. +- [x] Review existing suites for reusable coverage. Their custom QA manuals did + not exercise the tiny default hire; the new suite deliberately omits those + bundles and reuses independently graded skill, continuation, and chat journeys. - [ ] Compare task quality as well as Paperclip protocol compliance when claiming that fewer instructions improve agent performance. Keep model, effort, permissions, tools, and fixture comparable. @@ -189,19 +389,71 @@ does not apply. - [ ] Run the relevant checks for each change and the repository's required full verification before a PR-ready handoff. Record unrun checks and their reasons. +### Coverage for changes implemented so far + +The [stock-harness runbook](../../tests/runner-e2e/STOCK-HARNESS.md) defines a +credential-free prerequisite plus 24 explicit local Product E2E cells across +eight legacy/native profiles. Each profile runs assigned-skill invocation, +ordered comment continuation, and chat continuity across restart. Public receipts +verify the identity-only managed bundle before provider execution; final receipts +check actual legacy prompts and both budget hard stops. The existing lifecycle +oracles still own task/chat success. Models, auth, skills, permissions, and +secret-reference plumbing are inherited from existing profiles. + +| Coverage ID | Implemented change | Executable coverage | Live status | +| --- | --- | --- | --- | +| SH-1 | Native Codex additive developer instructions, including start/resume/recovery | TypeScript driver, runnerd transport, Runner Lab/live-session, and Rust provider tests in `pnpm test:e2e:runner:stock-harness`; native Codex cells exercise real hires. | Native Codex passes all three matched journeys in both variants; #14920 held constant, so task success is not before/after vendor-base proof. | +| SH-2 | Eight-word default hire `AGENTS.md` | Public creation/onboarding tests; exact independent public bundle oracle before and after provider execution for every stock-harness cell. | Candidate public tiny-bundle and budget receipts pass in all 24 retained cells. Behavioral delivery is not fully qualified; see F14–F17. | +| SH-3 | Reduced shared task/chat defaults and removed generic resume contract | Shared renderer and ACPX/Codex/OpenCode/Pi/Hermes/Cursor Cloud regressions; actual legacy invocation prompts checked in the new suite. | Actual legacy prompts measured. Document-delivery regressions and clipped/missing receipts retained; additional carriers remain deferred item 2.2. | + +The suite is explicit-only and excluded from `--all`; it does not add paid work +to ordinary campaigns. Negative calibration covers manual regrowth, missing or +malformed prompt receipts, old startup/resume procedures, missing connection +guidance, wrong budgets, and skipped prerequisite assertions. Source revisions, +attempts, cost, partial failures, cleanup, and sanitized evidence use the existing +report pipeline. No paid providers were launched during the initial setup. +Subsequent GitHub measurements are recorded below; no broad coding-quality +improvement or full live qualification is claimed. Unrepresented harnesses remain +unqualified; Codex-through-ACP vendor-base preservation is still F7. + +Local setup verification: 478 prerequisite tests passed (477 TypeScript plus one +Rust), 860 Product E2E support tests passed, E2E typecheck passed, and all 24 cells +were discovered. The 313 unrelated native tests filtered by the gate are not +passing coverage. Final prerequisite evidence is retained at +`tests/runner-e2e/results/stock-harness-preflight-2026-10-02T16-15-39.064Z/preflight.json`; +earlier interrupted, missing-Hermes-discovery, and setup/test-timeout attempts are +retained as failures. Hermes is checked through its package config because the +root Vitest project list omits it. Later exact-head prerequisites expanded to +557 passed checks (556 TypeScript and one Rust), and E2E support to 895 tests. +Repository typecheck/build passed. The complete local Vitest run retained three +unrelated timing failures; all three files passed unchanged reruns. At +`1eb5ba420`, 52 CI gates pass and two skip; Greptile is 5/5 with zero open +threads. Later documentation heads require fresh checks. See the live report +for exact source hashes, campaigns, failures and cost limitations. + ## Findings ledger Confirmed mechanics below do not by themselves establish an effect on task quality. | ID | Finding | Work item / disposition | | --- | --- | --- | -| F1 | Native Codex sent Paperclip text as `baseInstructions`. Probes on codex-cli 0.153.4 showed replacement; additive `developerInstructions` retained the stock base on start and cold resume. | 1; app-server fix implemented, PR validation pending | +| F1 | Native Codex sent Paperclip text as `baseInstructions`. Probes on codex-cli 0.153.4 showed replacement; additive `developerInstructions` retained the stock base on start and cold resume. | 1; app-server fix merged in PR #14920 | | F2 | Default hires and role templates prescribe substantial operating procedures; common prompt and wake layers add further coordination text. | 2–3; mechanics confirmed, performance effect unmeasured | | F3 | Hiring references require legacy Paperclip skill/comment procedures, while native Runner intentionally omits that operational skill and uses semantic tools. | 2–3, 5; reconcile runtime contracts | | F4 | Local Claude appends instructions; Runner Claude preserves the Claude Code preset. Runner isolation excludes project/local settings, which can also exclude repository instruction discovery. | 4; selective context fix to design | | F5 | Some Codex capability settings differ between the direct driver and daemon path; an intermediate configuration does not prove the final provider behavior. | 6; effective-path audit pending | | F6 | Omitting or nulling `baseInstructions` on an old Codex thread's resume preserves its saved replacement; an empty string produces an empty base. | 1; document the required provider session reset; no automatic migration in this PR | | F7 | The isolated Codex-through-ACP dependency patch also sets `baseInstructions` on start/resume. It is a separate path from the native app-server backend. | 6; follow-up patch/profile audit pending | +| F8 | The operational skill says target-bound confirmations default `supersedeOnUserComment` to true; the default manual and server normalizer say false. The server uses false. | 2; correct stale skill/reference guidance in a follow-up | +| F9 | The default manual required a comment on every task, while the operational skill's verified external-chat shortcut delegates comments and lifecycle bookkeeping to the harness. | 2; unconditional manual rule removed; shared layers still need review | +| F10 | Removing the default manual left the 661-word legacy task template and its 172-word resumed-wake execution contract. Hermes and Pi have additional policy carriers. | 2.1 complete locally: task/chat defaults 113 words, no generic resume contract; 2.2 wrappers pending | +| F11 | Task Markdown and the wake renderer prescribe different accepted-plan behavior for planning-mode accepted-confirmation payloads; tests currently expect both. | 2; unify the directive owner; production reachability still to trace | +| F12 | Native fixed instructions and full-turn constraints repeat completion and uncommon procedures, while reserved finish/block tool descriptions are only one sentence each. | 2; improve tool documentation before removing needed native protocol guidance | +| F13 | Existing context-integrity/chat fixtures injected a QA manual, so their green results did not qualify the production tiny hire default. | Dedicated stock-harness suite measured real default hires; retained failures prevent blanket qualification. | +| F14 | Historical classic Claude/OpenCode skill runs save one Paperclip document; reduced runs save none while completing the task. Legacy ACP Claude has the same behavior under a separate guard failure. Pinned skill storage wording is ambiguous. | Measured delivery failures; keep original oracle. Improve legacy skill/API delivery guidance separately, then compare both variants with preserved and storage-specific cases. | +| F15 | Classic Claude restart chat fails its memory assertion in both variants. | Existing behavior, not attributable to this reduction from these trials. | +| F16 | All six legacy ACP cells per variant fail the persisted-provider-credential guard. Three historical ACP cells also have clipped public prompt retrieval, and candidate ACP Codex chat lacks a complete invocation receipt. | Security/receipt qualification follow-ups; do not waive guard, expose raw sessions, or infer missing provider instructions from clipped receipts. | +| F17 | Prerequisites beside the campaign root violate trusted artifact selection. A same-target pilot superseded 13 candidate cells; one historical AWS runner shut down without upload. | Packaging fixed at `1eb5ba420` and pilot passes. Directory-layout-only copies preserve every byte; missing cells alone recovered at unchanged source, completed failures never rerun. | Append new findings with evidence, affected paths, and the numbered item that will address them. Record intentional behavior explicitly rather than as a bug. @@ -214,6 +466,26 @@ will address them. Record intentional behavior explicitly rather than as a bug. | 2026-10-02 | Paperclip-owned MCP isolation and configuration changes/session resets are acceptable. | Preserve Paperclip auth and assigned skills while fixing repository context. | | 2026-10-02 | Dotta requested implementation and a PR for item 1. Native app-server paths now use additive developer instructions. | 139 targeted TypeScript tests and 91 Rust provider tests passed; repository typecheck/build passed. Remaining test/review results to record. | | 2026-10-02 | Verified actual Codex instruction layering using a localhost Responses stub, without paid inference. | codex-cli 0.153.4 sent identical 14,732-character stock base instructions on start and cold resume, with the Paperclip marker retained in developer input. This is protocol evidence, not a task-quality eval. | +| 2026-10-02 | Dotta requested and confirmed the merge of item 1. | PR #14920 merged at `408f70e69f9c5e49cb4377f4886ac2001bfa67a2`, with all 55 checks passing and Greptile 5/5. | +| 2026-10-02 | Dotta directed an identity-only default manual with no skill or runtime pointers; the harness already handles coordination. | Reduced the default to eight words. Existing hire and onboarding suites passed all 64 tests. Shared prompts and role templates remain pending. | +| 2026-10-02 | Dotta requested three explicit shared-prompt follow-ups and selected common legacy startup/resume reduction first. | Follow-ups 2.1–2.3 recorded. 2.1 complete locally: both defaults 113 words; generic resume contract removed. 563 focused tests and shared utility typecheck/build passed. Extra carriers and native instruction reduction remain pending. | +| 2026-10-02 | Dotta requested executable eval coverage for everything implemented so far and all subsequent changes before continuing. | Added SH-1–SH-3 coverage map, credential-free prerequisite, and 24 production-default-hire cells. 478 prerequisite and 860 support tests passed; E2E typecheck/discovery passed. Live provider results remain `not_run`; item 2.2/2.3 unchanged. | +| 2026-10-02 | Dotta requested PRs and GitHub-runner before/after qualification while discussing subsequent work separately. | Draft [PR #14948](https://github.com/paperclipai/paperclip/pull/14948); native Codex diagnostic [passed](https://github.com/paperclipai/paperclip/actions/runs/37034213743), 34.745 s provider / 56.660 s cell. Comparison restores only the prior manual/shared prompts and holds #14920 constant; no general coding-quality claim. | +| 2026-10-02 | Full candidate cold setup failed before provider admission; stopped and retained the attempt. | [Run 37037105491](https://github.com/paperclipai/paperclip/actions/runs/37037105491), target `36e987246`: prerequisite imports lacked the plugin SDK build. Added ordinary dependency setup before credential-free prerequisites and exact-source coverage of connection guidance. Cold pilot and full matched campaigns pending. | +| 2026-10-02 | Dotta deferred 2.2 additional legacy carriers and approved the first 2.3 native tool-description slice separately. | Keep wrappers open; native finish/block documentation must not be supplied to legacy skill/API completion paths. | +| 2026-10-02 | Matched default-manual/shared-prompt campaigns measured failures, with #14920 constant. | Candidate `f02d8d0df`: 15/24 pass. Historical `12c5433c6`: 15/24 pass after the one runner-shutdown recovery also timed out. Classic Claude/OpenCode Paperclip document delivery regressed in observed trials; keep PR #14948 draft. [Live report](2026-10-02-stock-harness-live-comparison.md). | +| 2026-10-02 | Cold prerequisite and packaging faults were repaired without weakening admission or behavioral graders. | Three setup attempts stopped before providers. Current `1eb5ba420` pilot passes 557 prerequisite checks and protected report publication; source/hash/cost evidence retained. Candidate cancellation and historical runner shutdown recovered only for missing cells. | +| 2026-10-02 | Final bounded recovery completed; no further model reruns. | Both matched cohorts have all 24 results, 15 pass and nine fail. Two classic skill deliveries regress; two OpenCode ordered cases pass only with reduced instructions. Overall parity is not behavioral equivalence. All 48 retained result/receipt projections are hashed and sanitized; original interruptions and partial unknown spend remain recorded. | +| 2026-10-02 | Dotta approved the measured legacy delivery repair and narrow follow-up qualification. | Early operational skill PUT/receipt/link guidance plus generic issue-document reference; eight-word manual retained. Focused original + explicit Paperclip-storage cases on classic Claude/OpenCode, with only the two skill sources varied. All four pairs completed: Claude original Fail → Pass, Claude explicit Pass → Pass, both OpenCode cases Fail → Fail. Explicit OpenCode handoff worsened beneath the unchanged machine grade (clickable API URL → code-formatted path). [Repair report](2026-10-02-legacy-document-skill-repair.md); no full matrix rerun. Claude chat-memory (F15) and ACP credential/receipt failures (F16) remain separate and unresolved. | +| 2026-10-02 | Native tool-description comparison completed all six paired cases. | Zero newly failing cases, three unchanged completion passes, three unchanged blocker failures. Claude/OpenCode blocker API matchers pass in both variants; UI matcher wrongly demanded marker-only replies. Codex's exact-action punctuation failure is unchanged. Original failures retained; corrected 10-case browser calibration and separate retained-DOM replay pass all six visible replies. Exact Codex action failures remain. No native fixed-prompt removal measured. | +| 2026-10-02 | Dotta approved minimal operational skill selection/link correction after retained OpenCode diagnosis. | Candidate `fe9dc1e3c` and baseline `0d7ecfa96d` have two matched Pass → Pass cases, zero new failures/passes and no pending pairs. All four exact-source gates pass 587 checks, all four provider runs and cleanup pass. Candidate original loads Paperclip/reference before delivery, but uses the wrong PAP prefix; baseline original saves publicly later within the same assignment and gives a bare path. Explicit clickable UI links are correct in both. [Complete report](2026-10-02-opencode-skill-routing-link-qualification.md); prior failures retained, no causal or broad quality claim. | + +- F18: Native blocker browser assertion required a marker-only reply despite asking for owner/action/reason. Correct marker-plus-explanation checks symmetrically, calibrate contradictory and future-condition replies, retain original verdicts. +- F19: Original prompt-removal cohorts loaded Paperclip; OpenCode skill truncation omitted the late API recipe. The early repair fixed Claude in one paired trial. In the repaired original OpenCode assignment, Paperclip was first loaded only during disposition recovery after a local-file write. Shared legacy operational-skill delivery/selection needs review before more recipe expansion. +- F20: Explicit repaired OpenCode reads the early recipe/reference and saves a public document/revision, but hands off a code-formatted path rather than a clickable anchor. Baseline provides a clickable API URL, rejected by the UI-only oracle. Preserve both grades; clarify future clickable UI-link fixture and review the smallest link example separately. +- F22: Legacy operational skill mounting is already mandatory, but its discovery description omitted ordinary task/heartbeat work and document delivery. The approved narrow fix expands stock metadata selection and adds a real Markdown link example in the existing reference. Original + clarified-explicit OpenCode pairs both pass in both variants, so improvement causality is not established. Candidate loads the operational skill/reference before public delivery; baseline original writes locally first, then saves publicly in the same assignment. Original candidate copies the example PAP prefix into its link while baseline supplies a bare slug path; neither is scored by the original link-free oracle. No full-body injection or native-tool leakage. ACP Claude names-only metadata, custom ACP unsupported delivery, OpenClaw wrappers and Pi HOME differences remain distinct follow-ups. +- F23: The new original OpenCode candidate copies `/PAP/issues/RUN-1#document-context-integrity-output` from the literal example despite actual prefix RUN. It passes durable-document grading but has an incorrect-prefix handoff. Baseline original also lacks a clickable canonical link. Approved reference-only correction now derives the link from `issue.identifier` and the saved receipt key. Independent provider-free calibration checks two other prefixes, redirected keys and rejects a wrong-company link; 25 focused document/source tests, canonical metadata checks and E2E typecheck pass. The recipe now also handles null/absent identifiers through the supported issue-ID route; fresh review caught the initial nullable edge. 56 focused checks pass. Source review shows the UI corrects wrong prefixes, so the frozen PAP href is noncanonical rather than proven broken. These later corrections are not live-qualified by the preserved runs; no further paid rerun. +- F21: Skill heading insertion shifted both generated capability inventories. General CI caught stale metadata after the frozen repair campaign’s narrower build admitted providers. Regenerate canonical derived files and require stale-manifest/inventory checks before future provider admission. For each completed item, add the chosen behavior, changed paths, verification results, remaining exceptions, and follow-ups here before checking it off. diff --git a/package.json b/package.json index 2fbd9d95a9..29223ed74e 100644 --- a/package.json +++ b/package.json @@ -74,6 +74,7 @@ "test:e2e:browser-context": "playwright test --config tests/e2e/playwright-company-context.config.ts", "test:e2e:runner:browser-support": "playwright test --config tests/runner-e2e/playwright-support.config.ts", "test:e2e:runner:unit": "vitest run --config tests/runner-e2e/vitest.config.ts", + "test:e2e:runner:stock-harness": "node tests/runner-e2e/stock-harness-checks.mjs", "test:e2e:runner:typecheck": "tsc -p tests/runner-e2e/tsconfig.json", "test:e2e:runner:report": "node cli/node_modules/tsx/dist/cli.mjs tests/runner-e2e/report.ts", "test:runner-workflow-evals": "pnpm --filter @paperclipai/paperclip-eval-kernel build && pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals", diff --git a/packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.test.ts b/packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.test.ts new file mode 100644 index 0000000000..40d13e7b40 --- /dev/null +++ b/packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.test.ts @@ -0,0 +1,165 @@ +import { mkdtemp, readFile, rm } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { createRuntimeStore, type AcpSessionRecord, type AcpSessionStore } from "acpx/runtime"; +import { afterEach, describe, expect, it, vi } from "vitest"; +import { createEphemeralSessionEnvironmentStore } from "./ephemeral-session-environment.js"; + +const currentEnvironment = Object.freeze({ + PAPERCLIP_API_KEY: "current-credential-fixture", + PAPERCLIP_RUN_ID: "current-run-fixture", + PATH: "/fixture/current-bin", +}); +const oldEnvironment = { PAPERCLIP_API_KEY: "old-credential-fixture", OLD_RUN_PATH: "/fixture/old-run" }; + +function record(): AcpSessionRecord { + return { + schema: "acpx.session.v1", + acpxRecordId: "fixture-record", + acpSessionId: "fixture-session", + agentSessionId: "fixture-agent-session", + agentCommand: "fixture-provider", + cwd: "/fixture/project", + name: "fixture-name", + createdAt: "2026-10-03T00:00:00.000Z", + lastUsedAt: "2026-10-03T00:00:01.000Z", + lastSeq: 3, + eventLog: { active_path: "/fixture/events.jsonl", segment_count: 1, max_segment_bytes: 1024, max_segments: 2 }, + messages: [{ User: { id: "message-1", content: [{ Text: "Retained conversation" }] } }], + updated_at: "2026-10-03T00:00:01.000Z", + cumulative_token_usage: { input_tokens: 11, output_tokens: 7 }, + request_token_usage: { request_1: { input_tokens: 11 } }, + importedFrom: { recordId: "origin-record", cwdOriginal: "/fixture/origin", exportedBy: "fixture-caller", exportedAt: "2026-10-03T00:00:00.000Z" }, + acpx: { + current_mode_id: "code", + desired_config_options: { effort: "high" }, + session_options: { model: "fixture-model", allowed_tools: ["Read", "Write"], max_turns: 4, system_prompt: { append: "Retained caller option" }, env: { ...currentEnvironment } }, + }, + }; +} +function freeze(value: T): T { + for (const child of Object.values(value)) if (typeof child === "object" && child !== null) freeze(child); + return Object.freeze(value); +} +function stubStore(loaded: AcpSessionRecord | undefined) { + const load = vi.fn().mockResolvedValue(loaded); + const save = vi.fn().mockResolvedValue(undefined); + return { load, save }; +} +const directories: string[] = []; +afterEach(async () => { await Promise.all(directories.splice(0).map(directory => rm(directory, { recursive: true, force: true }))); }); +async function diskStore() { + const stateDir = await mkdtemp(join(tmpdir(), "paperclip-session-env-")); directories.push(stateDir); + return { store: createRuntimeStore({ stateDir }), file: join(stateDir, "sessions", "fixture-record.json") }; +} + +describe("ephemeral ACPX session environment", () => { + it("saves a fresh session to the actual file store without persisting its live environment", async () => { + const disk = await diskStore(), input = freeze(record()); + const wrapper = createEphemeralSessionEnvironmentStore(disk.store, currentEnvironment); + await wrapper.save(input); + const text = await readFile(disk.file, "utf8"), saved = JSON.parse(text); + expect(saved.acpx.session_options).not.toHaveProperty("env"); + for (const value of Object.values(currentEnvironment)) expect(text).not.toContain(value); + expect(input.acpx?.session_options?.env).toEqual(currentEnvironment); + expect(saved.acpx.current_mode_id).toBe("code"); + expect(saved.acpx.desired_config_options).toEqual(input.acpx?.desired_config_options); + expect(saved.acpx.session_options).toEqual({ model: "fixture-model", allowed_tools: ["Read", "Write"], max_turns: 4, system_prompt: { append: "Retained caller option" } }); + expect(saved.messages).toEqual(input.messages); + expect(saved.imported_from.record_id).toBe("origin-record"); + const loaded = await wrapper.load(input.acpxRecordId); + expect(loaded?.acpx?.session_options?.env).toEqual(currentEnvironment); + expect(loaded?.messages).toEqual(input.messages); + }); + + it("loads an old env-bearing file with rotated credentials without writing, then saves neither environment", async () => { + const disk = await diskStore(), legacy = record(); legacy.acpx!.session_options!.env = { ...oldEnvironment }; + await disk.store.save(legacy); + const before = await readFile(disk.file, "utf8"); + const wrapper = createEphemeralSessionEnvironmentStore(disk.store, currentEnvironment); + const loaded = await wrapper.load(legacy.acpxRecordId); + expect(loaded?.acpx?.session_options?.env).toEqual(currentEnvironment); + expect(loaded?.acpx?.session_options?.env).not.toHaveProperty("OLD_RUN_PATH"); + expect(await readFile(disk.file, "utf8")).toBe(before); + expect(legacy.acpx?.session_options?.env).toEqual(oldEnvironment); + await wrapper.save(loaded!); + const after = await readFile(disk.file, "utf8"); + for (const value of [...Object.values(oldEnvironment), ...Object.values(currentEnvironment)]) expect(after).not.toContain(value); + expect(JSON.parse(after).acpx.session_options).not.toHaveProperty("env"); + expect(loaded?.acpx?.session_options?.env).toEqual(currentEnvironment); + }); + + it("loads an env-free persisted session into the current environment without writing", async () => { + const disk = await diskStore(), input = record(); delete input.acpx!.session_options!.env; + await disk.store.save(input); + const before = await readFile(disk.file, "utf8"); + const loaded = await createEphemeralSessionEnvironmentStore(disk.store, currentEnvironment).load(input.acpxRecordId); + expect(loaded?.acpx?.session_options?.env).toEqual(currentEnvironment); + expect(loaded?.acpx?.session_options?.model).toBe("fixture-model"); + expect(await readFile(disk.file, "utf8")).toBe(before); + }); + + it("copies loaded records, options and current env without mutating stored or caller-owned objects", async () => { + const stored = record(); stored.acpx!.session_options!.env = { ...oldEnvironment }; freeze(stored); + const backing = stubStore(stored), wrapper = createEphemeralSessionEnvironmentStore(backing, currentEnvironment); + const loaded = await wrapper.load("fixture-record"); + expect(backing.load).toHaveBeenCalledWith("fixture-record"); expect(backing.save).not.toHaveBeenCalled(); + expect(loaded).not.toBe(stored); expect(loaded?.acpx).not.toBe(stored.acpx); + expect(loaded?.acpx?.session_options).not.toBe(stored.acpx?.session_options); + expect(loaded?.acpx?.session_options?.env).not.toBe(currentEnvironment); + expect(stored.acpx?.session_options?.env).toEqual(oldEnvironment); + loaded!.acpx!.session_options!.env!.PAPERCLIP_API_KEY = "isolated-loaded-fixture"; + expect(currentEnvironment.PAPERCLIP_API_KEY).toBe("current-credential-fixture"); + }); + + it("saves only a copied env-free projection, preserving unrelated and caller metadata", async () => { + const input = freeze(Object.assign(record(), { caller_metadata: { keep: "caller-state", env: { preserve: "unrelated caller metadata" } } })); + const backing = stubStore(undefined), wrapper = createEphemeralSessionEnvironmentStore(backing, currentEnvironment); + await wrapper.save(input); + const saved = backing.save.mock.calls[0]![0]; + const { env: _environment, ...expectedOptions } = input.acpx!.session_options!; + expect(saved).toEqual({ ...input, acpx: { ...input.acpx, session_options: expectedOptions } }); + expect(saved).not.toBe(input); expect(saved.acpx).not.toBe(input.acpx); + expect(saved.acpx?.session_options).not.toBe(input.acpx?.session_options); + expect(saved.acpx?.session_options).not.toHaveProperty("env"); + expect(input.acpx?.session_options?.env).toEqual(currentEnvironment); + }); + + for (const [name, acpx] of [ + ["absent acpx", undefined], + ["explicit undefined acpx", undefined], + ["absent session options", { current_mode_id: "code" }], + ["undefined session options", { session_options: undefined }], + ["empty session options", { session_options: {} }], + ["undefined env", { session_options: { env: undefined, model: "fixture-model" } }], + ] as const) it(`preserves ${name} when saving and retains existing load injection semantics`, async () => { + const input = record(); + if (name === "absent acpx") delete input.acpx; else input.acpx = acpx; + freeze(input); + const backing = stubStore(input), wrapper = createEphemeralSessionEnvironmentStore(backing, currentEnvironment); + await wrapper.save(input); + const saved = backing.save.mock.calls[0]![0]; + if (name === "undefined env") expect(saved.acpx?.session_options).toEqual({ model: "fixture-model" }); + else expect(saved).toEqual(input); + expect(Object.hasOwn(saved, "acpx")).toBe(Object.hasOwn(input, "acpx")); + if (saved.acpx && input.acpx) expect(Object.hasOwn(saved.acpx, "session_options")).toBe(Object.hasOwn(input.acpx, "session_options")); + const loaded = await wrapper.load(input.acpxRecordId); + expect(loaded?.acpx?.session_options?.env).toEqual(currentEnvironment); + expect(backing.save).toHaveBeenCalledTimes(1); + }); + + it("returns a missing record without saving or creating conversation state", async () => { + const backing = stubStore(undefined); + expect(await createEphemeralSessionEnvironmentStore(backing, currentEnvironment).load("missing")).toBeUndefined(); + expect(backing.save).not.toHaveBeenCalled(); + }); + + it("propagates load and save errors unchanged", async () => { + const failure = new Error("fixture store failure"), backing = stubStore(undefined); + backing.load.mockRejectedValue(failure); backing.save.mockRejectedValue(failure); + const wrapper = createEphemeralSessionEnvironmentStore(backing, currentEnvironment); + await expect(wrapper.load("fixture-record")).rejects.toBe(failure); + const input = freeze(record()); await expect(wrapper.save(input)).rejects.toBe(failure); + expect(input.acpx?.session_options?.env).toEqual(currentEnvironment); + }); +}); diff --git a/packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.ts b/packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.ts new file mode 100644 index 0000000000..3bd59942ec --- /dev/null +++ b/packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.ts @@ -0,0 +1,32 @@ +import type { AcpSessionStore } from "acpx/runtime"; + +/** Keep the current run's environment in memory, outside persisted conversation state. */ +export function createEphemeralSessionEnvironmentStore( + persistedStore: AcpSessionStore, + currentEnvironment: Readonly>, +): AcpSessionStore { + return { + async load(id) { + const record = await persistedStore.load(id); + if (!record) return undefined; + return { + ...record, + acpx: { + ...record.acpx, + session_options: { + ...record.acpx?.session_options, + env: { ...currentEnvironment }, + }, + }, + }; + }, + save(record) { + const persisted = { ...record }; + if (record.acpx?.session_options !== undefined) { + const { env: _environment, ...sessionOptions } = record.acpx.session_options; + persisted.acpx = { ...record.acpx, session_options: sessionOptions }; + } + return persistedStore.save(persisted); + }, + }; +} diff --git a/packages/adapter-utils/src/acpx-engine/execute.test.ts b/packages/adapter-utils/src/acpx-engine/execute.test.ts index ffc724b4d9..03a42f9645 100644 --- a/packages/adapter-utils/src/acpx-engine/execute.test.ts +++ b/packages/adapter-utils/src/acpx-engine/execute.test.ts @@ -2,7 +2,7 @@ import fs from "node:fs/promises"; import os from "node:os"; import path from "node:path"; import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; -import type { AcpRuntimeOptions } from "acpx/runtime"; +import { createRuntimeStore, type AcpRuntimeOptions, type AcpSessionRecord, type AcpSessionStore } from "acpx/runtime"; import type { AdapterExecutionContext, AdapterRuntimeMcpAccess } from "@paperclipai/adapter-utils"; import { DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC, @@ -377,6 +377,42 @@ const ALLOWED_TURN_SPAN_ATTRIBUTE_KEYS = new Set([ ]); describe("shared ACPX engine runtime behavior", () => { + it("integrates the env-free store while delivering current credentials and options to the provider", async () => { + const root = await makeTempRoot(); + const stateDir = path.join(root, "state"); + const first = await runExecutor({ agent: "claude", agentCommand: "node ./fake-acp.js", cwd: root, stateDir, + env: { ANTHROPIC_API_KEY: "first-fixture-credential" } }); + const liveEnvironment = (first.sessionInputs[0]!.sessionOptions as { env: Record }).env; + expect(liveEnvironment.ANTHROPIC_API_KEY).toBe("first-fixture-credential"); + const store = first.runtimeOptions[0]!.sessionStore as AcpSessionStore; + const record: AcpSessionRecord = { + schema: "acpx.session.v1", acpxRecordId: "integration-record", acpSessionId: "integration-session", + agentCommand: "fixture-provider", cwd: root, name: "fixture", createdAt: "2026-10-03T00:00:00.000Z", + lastUsedAt: "2026-10-03T00:00:01.000Z", lastSeq: 0, + eventLog: { active_path: path.join(stateDir, "fixture.jsonl"), segment_count: 1, max_segment_bytes: 1024, max_segments: 2 }, + messages: [], updated_at: "2026-10-03T00:00:01.000Z", cumulative_token_usage: {}, request_token_usage: {}, + acpx: { session_options: { env: { ...liveEnvironment }, model: "retained-model", max_turns: 4 } }, + }; + await store.save(record); + const file = path.join(stateDir, "sessions", "integration-record.json"); + const saved = await fs.readFile(file, "utf8"); + expect(saved).not.toContain("first-fixture-credential"); + expect(JSON.parse(saved).acpx.session_options).toEqual({ model: "retained-model", max_turns: 4 }); + expect(record.acpx?.session_options?.env).toEqual(liveEnvironment); + // A file written by the prior engine is loaded with the current run's env, + // without altering the provider launch or deleting unrelated options. + record.acpx!.session_options!.env = { ANTHROPIC_API_KEY: "legacy-fixture-credential" }; + await createRuntimeStore({ stateDir }).save(record); + const second = await runExecutor({ agent: "claude", agentCommand: "node ./fake-acp.js", cwd: root, stateDir, + env: { ANTHROPIC_API_KEY: "rotated-fixture-credential" } }, { runtime: { sessionParams: first.result.sessionParams } }); + expect((second.sessionInputs[0]!.sessionOptions as { env: Record }).env.ANTHROPIC_API_KEY).toBe("rotated-fixture-credential"); + const rotatedStore = second.runtimeOptions[0]!.sessionStore as AcpSessionStore; + const loaded = await rotatedStore.load(record.acpxRecordId); + expect(loaded?.acpx?.session_options).toMatchObject({ env: { ANTHROPIC_API_KEY: "rotated-fixture-credential" }, model: "retained-model", max_turns: 4 }); + await rotatedStore.save(loaded!); + const resaved = await fs.readFile(file, "utf8"); + for (const value of ["first-fixture-credential", "legacy-fixture-credential", "rotated-fixture-credential"]) expect(resaved).not.toContain(value); + }); it.each(["claude", "codex", "gemini", "kimi", "custom"])("defaults the legacy %s engine to full auto on fresh and resumed runs", async (agent) => { const root = await makeTempRoot(); const config = { @@ -597,6 +633,9 @@ describe("shared ACPX engine runtime behavior", () => { const fresh = await runExecutor(config, { context }); expect(fresh.turnInputs).toHaveLength(1); const prompt = String(fresh.turnInputs[0]?.text ?? ""); + expect(prompt).toContain("You are agent agent-1"); + expect(prompt).toContain("Connection tools:"); + expect(prompt).not.toContain("Execution contract:"); expect(prompt).toContain(context.paperclipTaskMarkdownAssignment); expect(prompt).toContain(context.paperclipTaskCommunicationGuidance); expect(prompt).not.toContain('"objective":'); @@ -610,10 +649,18 @@ describe("shared ACPX engine runtime behavior", () => { const resumed = await runExecutor(config, { context, runtime: { sessionParams: fresh.result.sessionParams } }); expect(resumed.sessionInputs[0]?.resumeSessionId).toBe(fresh.result.sessionId); const resumedPrompt = String(resumed.turnInputs[0]?.text ?? ""); + expect(resumedPrompt).not.toContain("Execution contract:"); expect(resumedPrompt).toContain(context.paperclipTaskMarkdownAssignmentCompact); expect(resumedPrompt).not.toContain(context.paperclipTaskCommunicationGuidance); expect(resumedPrompt).not.toContain('"id":"comment-first"'); expect(resumedPrompt).toContain('"id":"comment-second"'); + const reset = await runExecutor(config, { context }); + expect(reset.sessionInputs[0]?.resumeSessionId).toBeUndefined(); + const resetPrompt = String(reset.turnInputs[0]?.text ?? ""); + expect(resetPrompt).toContain("You are agent agent-1"); + expect(resetPrompt).toContain("Connection tools:"); + expect(resetPrompt).toContain(context.paperclipTaskMarkdownAssignment); + expect(resetPrompt).not.toContain("Execution contract:"); }); it("includes Paperclip env and API access notes in the ACPX prompt without leaking the token", async () => { @@ -747,10 +794,10 @@ describe("shared ACPX engine runtime behavior", () => { expect(prompt).not.toContain("Create child issues"); expect(prompt).not.toContain("Use child issues"); } - expect(String(fresh.meta[0]?.prompt)).toContain(custom ? "Custom agent instructions." : "Continue your Paperclip conversation"); - expect(String(reset.meta[0]?.prompt)).toContain(custom ? "Custom agent instructions." : "Continue your Paperclip conversation"); + expect(String(fresh.meta[0]?.prompt)).toContain(custom ? "Custom agent instructions." : "You are agent agent-1"); + expect(String(reset.meta[0]?.prompt)).toContain(custom ? "Custom agent instructions." : "You are agent agent-1"); const ordinary = await runExecutor({ ...config, promptTemplate: "" }, { context: { ...context, conversationMode: false } }); - expect(String(ordinary.meta[0]?.prompt)).toContain("Execution contract:"); + expect(String(ordinary.meta[0]?.prompt)).not.toContain("Execution contract:"); expect(String(ordinary.meta[0]?.prompt)).toContain("Create child issues from the approved plan"); }); @@ -3773,9 +3820,54 @@ describe("ACPX engine Claude skill bundle staging (remote ACP lane)", () => { }; } + it("advertises folded routing metadata and exact staged paths without injecting skill bodies", async () => { + const root = await makeTempRoot(); + const cwd = path.join(root, "worktree"); + const stateDir = path.join(root, "state"); + await fs.mkdir(cwd, { recursive: true }); + const skill = await createSkill(path.join(root, "sources"), "paperclip--versioned", [ + "---", "description: >-", " Use for Paperclip-managed tasks", " and durable task documents.", + "---", "PRIVATE_SKILL_BODY_RECIPE", "", + ].join("\n")); + const { meta, result } = await runExecutor({ + agent: "claude", agentCommand: "node ./fake-acp.js", stateDir, cwd, + paperclipRuntimeSkills: [skill], paperclipSkillSync: { desiredSkills: [skill.key] }, + }); + const bundle = await onlyChildDir(path.join(stateDir, "runtime-skills", "claude")); + const stagedFile = path.join(bundle, ".claude", "skills", skill.runtimeName, "SKILL.md"); + const prompt = String(meta[0]?.prompt ?? ""); + expect(prompt).toContain(`- ${skill.runtimeName}: Use for Paperclip-managed tasks and durable task documents. (file: ${stagedFile})`); + expect(prompt).not.toContain("PRIVATE_SKILL_BODY_RECIPE"); + expect(prompt).not.toContain(skill.source); + expect(result.sessionParams?.skills).toMatchObject({ selectedSkills: [skill.runtimeName] }); + await expect(fs.readFile(stagedFile, "utf8")).resolves.toContain("PRIVATE_SKILL_BODY_RECIPE"); + }); + + it("bounds routing descriptions and keeps files discoverable when metadata is absent or malformed", async () => { + const root = await makeTempRoot(); + const cwd = path.join(root, "worktree"); + const stateDir = path.join(root, "state"); + await fs.mkdir(cwd, { recursive: true }); + const sources = path.join(root, "sources"); + const bounded = await createSkill(sources, "bounded", `---\ndescription: ${"x".repeat(600)}\n---\nBOUNDED_BODY\n`); + const absent = await createSkill(sources, "absent", "# ABSENT_BODY\n"); + const malformed = await createSkill(sources, "malformed", "---\ndescription: incomplete header\nMALFORMED_BODY\n"); + const skills = [bounded, absent, malformed]; + const { meta } = await runExecutor({ + agent: "claude", agentCommand: "node ./fake-acp.js", stateDir, cwd, + paperclipRuntimeSkills: skills, paperclipSkillSync: { desiredSkills: skills.map((skill) => skill.key) }, + }); + const prompt = String(meta[0]?.prompt ?? ""); + expect(prompt).toContain(`- bounded: ${"x".repeat(512)} (file: `); + expect(prompt).not.toContain("x".repeat(513)); + expect(prompt).toMatch(/- absent \(file: .*\/absent\/SKILL\.md\)/); + expect(prompt).toMatch(/- malformed \(file: .*\/malformed\/SKILL\.md\)/); + for (const body of ["BOUNDED_BODY", "ABSENT_BODY", "MALFORMED_BODY"]) expect(prompt).not.toContain(body); + }); + it("rewrites the prompt and skill identity onto the in-sandbox skill root once the bundle is staged", async () => { const { stateDir, localCwd, executionTarget } = await setupRemoteSandbox(); - const skill = await createSkill(path.join(localCwd, "skills"), "review"); + const skill = await createSkill(path.join(localCwd, "skills"), "review", "---\ndescription: Review requested task documents.\n---\n# review\nREMOTE_BODY_ONLY\n"); const { meta, result } = await runExecutor( { @@ -3798,6 +3890,8 @@ describe("ACPX engine Claude skill bundle staging (remote ACP lane)", () => { const inSandboxSkillsRoot = prompt.match(/Skill root: (\S+)/)![1]!; expect(inSandboxSkillsRoot).not.toBe(hostSkillsHome); + expect(prompt).toContain(`- review: Review requested task documents. (file: ${path.join(inSandboxSkillsRoot, "review", "SKILL.md")})`); + expect(prompt).not.toContain("REMOTE_BODY_ONLY"); await expect( fs.readFile(path.join(inSandboxSkillsRoot, skill.runtimeName, "SKILL.md"), "utf8"), ).resolves.toContain("# review"); diff --git a/packages/adapter-utils/src/acpx-engine/execute.ts b/packages/adapter-utils/src/acpx-engine/execute.ts index 5c575ca8c5..9f25c6ba5f 100644 --- a/packages/adapter-utils/src/acpx-engine/execute.ts +++ b/packages/adapter-utils/src/acpx-engine/execute.ts @@ -7,6 +7,7 @@ import path from "node:path"; import { execFile } from "node:child_process"; import { promisify } from "node:util"; import { createHash, randomUUID } from "node:crypto"; +import { parseFrontmatterMarkdown } from "@paperclipai/shared"; import { fileURLToPath } from "node:url"; import type { AdapterBillingType, @@ -85,6 +86,7 @@ import { type PaperclipSkillEntry, } from "@paperclipai/adapter-utils/server-utils"; import { shellQuote } from "@paperclipai/adapter-utils/ssh"; +import { createEphemeralSessionEnvironmentStore } from "./ephemeral-session-environment.js"; import { createAcpRuntime, createAgentRegistry, @@ -100,7 +102,6 @@ import { type AcpRuntimeTurnResult, type AcpRuntimeUsageBreakdown, type AcpRuntimeUsageCost, - type AcpSessionStore, } from "acpx/runtime"; import { ACPX_DUPLEX_LOSS_CANCEL_DEADLINE_MS, @@ -1183,11 +1184,24 @@ async function prepareClaudeSkillRuntime(input: { } const selectedNames = materializedNames.sort(); + const skillDescriptions = await Promise.all(selectedNames.map(async (name) => { + const file = path.join(skillsHome, name, "SKILL.md"); + let description = ""; + try { + const parsed = parseFrontmatterMarkdown(await fs.readFile(file, "utf8")); + description = asString(parsed.frontmatter.description, "").replace(/\s+/g, " ").trim().slice(0, 512); + } catch { + // A readable skill without valid routing metadata remains available. + // Do not substitute its body for a missing description. + } + return `- ${name}${description ? `: ${description}` : ""} (file: ${file})`; + })); const promptInstructions = selectedNames.length > 0 ? [ "Paperclip has materialized selected runtime skills for this ACPX Claude session.", `Skill root: ${skillsHome}`, `Selected skills: ${selectedNames.join(", ")}`, + ...skillDescriptions, "When a task calls for one of these skills, read its SKILL.md from that root and follow it.", ].join("\n") : ""; @@ -4249,26 +4263,9 @@ export function createAcpxEngineExecutor(deps: AcpxEngineExecutorOptions = {}) { flushChildStderr(childStderrState); childStderrState.logPath = prepared.childStderrLogPath; const persistedRuntimeStore = createRuntimeStore({ stateDir: prepared.stateDir }); - const runtimeStore: AcpSessionStore = { - async load(id) { - const record = await persistedRuntimeStore.load(id); - if (!record) return undefined; - // ACPX resumes from the stored session options rather than the - // options passed to ensureSession. Keep conversation state, but - // launch the provider with this run's credentials and scratch paths. - return { - ...record, - acpx: { - ...record.acpx, - session_options: { - ...record.acpx?.session_options, - env: { ...prepared.env }, - }, - }, - }; - }, - save: (record) => persistedRuntimeStore.save(record), - }; + // Resume with this run's launch environment; keep it out of the saved + // conversation record without mutating the live runtime's options. + const runtimeStore = createEphemeralSessionEnvironmentStore(persistedRuntimeStore, prepared.env); const runtimeOptions: PaperclipAcpRuntimeOptions = { cwd: prepared.cwd, // Host-only spawn cwd for the relay proxy on the remote process-session diff --git a/packages/adapter-utils/src/prompt-sections.test.ts b/packages/adapter-utils/src/prompt-sections.test.ts index 6c8e67192a..bc660ccfde 100644 --- a/packages/adapter-utils/src/prompt-sections.test.ts +++ b/packages/adapter-utils/src/prompt-sections.test.ts @@ -3,13 +3,14 @@ import { selectPaperclipPromptSections as selectSections } from "./server-utils. import { createPromptContextFixture } from "./test-fixtures/prompt-context.js"; describe("task and event section ownership", () => { - it("lets a separate instruction carrier own the execution contract on a resumed turn", () => { + it("preserves resumed wake data without reintroducing generic procedures", () => { const context = createPromptContextFixture(); - const sections = selectSections(context, { resumedSession: true, includeExecutionContract: false }); - expect(sections.taskContextNote).toBe(context.paperclipTaskMarkdownAssignmentCompact); - expect(sections.wakePrompt).toContain('"id":"comment-second"'); - expect(sections.wakePrompt).not.toContain("Execution contract:"); - expect(selectSections(context, { resumedSession: true }).wakePrompt).toContain("Execution contract:"); + for (const includeExecutionContract of [undefined, false, true]) { + const sections = selectSections(context, { resumedSession: true, includeExecutionContract }); + expect(sections.taskContextNote).toBe(context.paperclipTaskMarkdownAssignmentCompact); + expect(sections.wakePrompt).toContain('"id":"comment-second"'); + expect(sections.wakePrompt).not.toContain("Execution contract:"); + } }); it("preserves user repetition and distinct same-body comments under their source owners", () => { diff --git a/packages/adapter-utils/src/server-utils.test.ts b/packages/adapter-utils/src/server-utils.test.ts index 95a2085bdf..c35e4e2a8f 100644 --- a/packages/adapter-utils/src/server-utils.test.ts +++ b/packages/adapter-utils/src/server-utils.test.ts @@ -925,7 +925,7 @@ describe("renderPaperclipWakePrompt", () => { fallbackFetchNeeded: false, }; const ordinary = renderPaperclipWakePrompt(payload, { resumedSession: true }); - expect(ordinary).toContain("Execution contract:"); + expect(ordinary).not.toContain("Execution contract:"); expect(ordinary).toContain("Create child issues from the approved plan"); for (const resumedSession of [false, true]) { const chat = renderPaperclipWakePrompt(payload, { @@ -1171,7 +1171,7 @@ describe("renderPaperclipWakePrompt", () => { expect(incompleteResumePrompt).toContain( "[continuation summary truncated]", ); - expect(incompleteResumePrompt).toContain( + expect(incompleteResumePrompt).not.toContain( "a successful process exit or final response is not sufficient", ); }); @@ -1410,76 +1410,18 @@ describe("renderPaperclipWakePrompt", () => { ); }); - it("keeps the default local-agent prompt action-oriented", () => { - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "Start actionable work in this heartbeat", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "do not stop at a plan", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "clear final disposition", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "evidence, not valid liveness paths by themselves", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "keep `in_progress` only when a live continuation path exists", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "Prefer the smallest verification that proves the change", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "After 2 consecutive failures of the same control-plane write", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "adapter/runtime status channel as the sanctioned fallback", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "Use child issues", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "instead of polling agents, sessions, or processes", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "Create child issues directly when you know what needs to be done", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "POST /api/issues/$PAPERCLIP_TASK_ID/interactions", - ); - // URL paths in prompt text carry real ids or env vars, never brace - // placeholders: agents paste these lines verbatim, and a literal {issueId} - // reaches the server as /api/issues/%7BissueId%7D and 404s. - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).not.toContain( - "/api/issues/{issueId}", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).not.toContain( - "/api/issues/{id}", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "kind suggest_tasks, ask_user_questions, or request_confirmation", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "Use continuationPolicy wake_assignee when you need to resume after a response (it wakes on acceptance and rejection alike; only expiry does not wake); use wake_assignee_on_accept when you want to resume only after acceptance", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).not.toContain( - "for request_confirmation this resumes only after acceptance", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "Never create probe or throwaway issue-thread interactions to discover the interactions API shape or your permissions", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "confirmation:{issueId}:plan:{revisionId}", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "Wait for acceptance before creating implementation subtasks", - ); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( - "Respect budget, pause/cancel, approval gates, and company boundaries", - ); + it("keeps task and chat defaults to identity and connection guidance", () => { + for (const template of [ + DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE, + DEFAULT_PAPERCLIP_CONVERSATION_PROMPT_TEMPLATE, + ]) { + expect(template).toBe( + `You are agent {{agent.id}} ({{agent.name}}).\n\n${CONNECTION_INTENT_AGENT_GUIDANCE}`, + ); + } }); - it("leaves the execution contract to the heartbeat template on fresh scoped wake prompts", () => { + it("keeps current task data without generic procedures on fresh scoped wake prompts", () => { const prompt = renderPaperclipWakePrompt({ reason: "issue_assigned", issue: { @@ -1499,12 +1441,12 @@ describe("renderPaperclipWakePrompt", () => { expect(prompt).toContain("## Paperclip Wake Payload"); expect(prompt).not.toContain("Execution contract:"); - expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).toContain( + expect(DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE).not.toContain( "Execution contract:", ); }); - it("adds the execution contract to resume delta prompts and opted-in fresh prompts", () => { + it("does not restore generic procedures on resume or with the legacy opt-in", () => { const payload = { reason: "issue_assigned", issue: { @@ -1522,36 +1464,19 @@ describe("renderPaperclipWakePrompt", () => { fallbackFetchNeeded: false, }; - for (const prompt of [ - renderPaperclipWakePrompt(payload, { resumedSession: true }), - renderPaperclipWakePrompt(payload, { includeExecutionContract: true }), - ]) { - expect(prompt).toContain( - "Execution contract: take concrete action in this heartbeat", - ); - expect(prompt).toContain("clear final disposition"); - expect(prompt).toContain( - "Immediately before returning, verify that Paperclip records one of those dispositions", - ); - expect(prompt).toContain( - "a successful process exit or final response is not sufficient", - ); - expect(prompt).toContain( - "If no valid disposition is recorded, record it now and do not end the run", - ); - expect(prompt).toContain( - "After 2 consecutive failures of the same control-plane write", - ); - expect(prompt).toContain( - "adapter/runtime status channel as the sanctioned fallback", - ); - expect(prompt).toContain( - "evidence, not valid liveness paths by themselves", - ); - expect(prompt).toContain( - "Use child issues for long or parallel delegated work instead of polling", - ); - expect(prompt).toContain("named unblock owner/action"); + for (const resumedSession of [false, true]) { + for (const includeExecutionContract of [undefined, false, true]) { + const prompt = renderPaperclipWakePrompt(payload, { + resumedSession, includeExecutionContract, + }); + expect(prompt).toContain(resumedSession ? "## Paperclip Resume Delta" : "## Paperclip Wake Payload"); + expect(prompt).toContain("- reason: issue_assigned"); + expect(prompt).toContain("- issue: PAP-1580 Update prompts"); + expect(prompt).toContain("- issue status: in_progress"); + expect(prompt).not.toContain("Execution contract:"); + expect(prompt).not.toContain("clear final disposition"); + expect(prompt).not.toContain("do not end the run"); + } } }); @@ -1947,7 +1872,7 @@ describe("renderPaperclipWakePrompt", () => { ); }); - it("keeps exactly one execution contract in a composed fresh heartbeat prompt", () => { + it("keeps connection guidance without a generic manual in a composed fresh prompt", () => { const wakePrompt = renderPaperclipWakePrompt({ reason: "issue_assigned", issue: { @@ -1967,7 +1892,8 @@ describe("renderPaperclipWakePrompt", () => { const composed = [wakePrompt, DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE].join( "\n\n", ); - expect(composed.match(/Execution contract/g)).toHaveLength(1); + expect(composed).toContain(CONNECTION_INTENT_AGENT_GUIDANCE); + expect(composed).not.toContain("Execution contract:"); }); it("trims comment-batch boilerplate on fresh wakes with zero pending comments", () => { diff --git a/packages/adapter-utils/src/server-utils.ts b/packages/adapter-utils/src/server-utils.ts index 71fda10d4a..6c9b2a2b13 100644 --- a/packages/adapter-utils/src/server-utils.ts +++ b/packages/adapter-utils/src/server-utils.ts @@ -208,41 +208,15 @@ export function resolvePaperclipInstanceRootForAdapter( } export const DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE = [ - "You are agent {{agent.id}} ({{agent.name}}). Continue your Paperclip work.", - "", - "Execution contract:", - "- Start actionable work in this heartbeat; do not stop at a plan unless the issue asks for planning.", - "- Leave durable progress in comments, documents, or work products, then update the issue to a clear final disposition before ending the heartbeat.", - "- Comments, documents, screenshots, work products, and `Remaining` bullets are evidence, not valid liveness paths by themselves.", - "- Final disposition checklist: mark `done` when complete; use `in_review` only with a real reviewer, approval, interaction, or monitor path; use `blocked` only with first-class blockers or a named unblock owner/action; create delegated follow-up issues with blockers when another agent owns the next step; keep `in_progress` only when a live continuation path exists.", - "- Prefer the smallest verification that proves the change; do not default to full workspace typecheck/build/test on every heartbeat unless the task scope warrants it.", - "- After 2 consecutive failures of the same control-plane write, stop retrying that write for the rest of the heartbeat. Continue useful work, report the failure in the final response, and rely on the adapter/runtime status channel as the sanctioned fallback.", - "- Use child issues for parallel or long delegated work instead of polling agents, sessions, or processes.", - "- If woken by a human comment on a dependency-blocked issue, respond or triage the comment without treating the blocked deliverable work as unblocked.", - "- Create child issues directly when you know what needs to be done; use issue-thread interactions when the board/user must choose suggested tasks, answer structured questions, or confirm a proposal.", - "- Use `PAPERCLIP_SCRATCH_DIR` / `PAPERCLIP_RUN_SCRATCH_DIR` for temporary scratch files instead of ad hoc `/tmp` paths; Paperclip removes that run-owned directory after the run ends.", - "- To ask for that input, create an interaction on the current issue with POST /api/issues/$PAPERCLIP_TASK_ID/interactions using kind suggest_tasks, ask_user_questions, or request_confirmation. Use continuationPolicy wake_assignee when you need to resume after a response (it wakes on acceptance and rejection alike; only expiry does not wake); use wake_assignee_on_accept when you want to resume only after acceptance.", - "- Never create probe or throwaway issue-thread interactions to discover the interactions API shape or your permissions; schema discovery goes through the OpenAPI spec and explicit validation errors, not placeholder cards. Every ask_user_questions, suggest_tasks, or request_confirmation you post must carry a real, answerable prompt; withdraw one you no longer need instead of leaving it pending.", - "- When you intentionally restart follow-up work on a completed assigned issue, include structured `resume: true` with the POST /api/issues/$PAPERCLIP_TASK_ID/comments or PATCH /api/issues/$PAPERCLIP_TASK_ID comment payload (substitute that issue's real id when it is not the current task). Generic agent comments on closed issues are inert by default.", - "- For plan approval, update the plan document first, then create request_confirmation targeting the latest plan revision with idempotencyKey confirmation:{issueId}:plan:{revisionId}. Wait for acceptance before creating implementation subtasks, and create a fresh confirmation after superseding board/user comments if approval is still needed.", - "- If blocked, mark the issue blocked and name the unblock owner and action.", - "- Respect budget, pause/cancel, approval gates, and company boundaries.", - "- When the server-authenticated wake payload includes an External chat response contract, that narrower contract replaces the generic Paperclip comment, status, checkout, and final-disposition steps above for that turn. Follow the external-chat contract exactly; it does not relax any permission, approval, execution-policy, containment, budget, pause/cancel, or company boundary.", + "You are agent {{agent.id}} ({{agent.name}}).", "", CONNECTION_INTENT_AGENT_GUIDANCE, ].join("\n"); -// Chat behavior is supplied centrally by the server's task-context markdown. -// Keep the ordinary task's completion/delegation contract out of this template. -export const DEFAULT_PAPERCLIP_CONVERSATION_PROMPT_TEMPLATE = [ - "You are agent {{agent.id}} ({{agent.name}}). Continue your Paperclip conversation using the supplied chat mode directive.", - "Use available tools and assigned skills as needed; respect budget, pause/cancel, approval gates, and company boundaries.", - "Prefer the smallest verification that proves the action. Use PAPERCLIP_SCRATCH_DIR / PAPERCLIP_RUN_SCRATCH_DIR for temporary scratch files.", - "After 2 consecutive failures of the same control-plane write, stop retrying that write for the rest of the turn. Report the failure honestly; never claim an unconfirmed mutation succeeded.", - "Never create probe or throwaway issue-thread interactions. Every interaction must carry a real, answerable prompt; withdraw one you no longer need.", - "", - CONNECTION_INTENT_AGENT_GUIDANCE, -].join("\n"); +// Task/chat modes arrive in server-owned context; operational procedures live +// in the harness-delivered Paperclip skill rather than either default template. +export const DEFAULT_PAPERCLIP_CONVERSATION_PROMPT_TEMPLATE = + DEFAULT_PAPERCLIP_AGENT_PROMPT_TEMPLATE; export const WATCHDOG_DEFAULT_MANDATE = [ "You are running as a task watchdog, not as the original deliverable worker.", @@ -2242,6 +2216,7 @@ export function selectPaperclipPromptSections( options: { resumedSession?: boolean; includeCommunicationGuidance?: boolean; + /** @deprecated Generic execution instructions are no longer injected. */ includeExecutionContract?: boolean; nativeWakeReaderAvailable?: boolean; } = {}, @@ -2277,6 +2252,7 @@ function renderPaperclipWakePromptBody( value: unknown, options: { resumedSession?: boolean; + /** @deprecated Generic execution instructions are no longer injected. */ includeExecutionContract?: boolean; // Conversation policy arrives in the server-owned task markdown. Generic // task disposition and child-delegation instructions conflict with it. @@ -2301,12 +2277,6 @@ function renderPaperclipWakePromptBody( externalChatTurn || externalChatReaderTurn || externalChatQuestionResponseTurn; - // The heartbeat prompt template already carries the execution contract on - // fresh sessions; only resume deltas (which replace the template) and - // template-less adapters need the wake-payload copy. An explicit false means - // another delivery carrier owns the contract, including on resume. - const includeExecutionContract = options.conversationMode !== true && - (options.includeExecutionContract ?? resumedSession); const hasWakeCommentBatch = normalized.comments.length > 0 || normalized.includedCount > 0 || @@ -2397,7 +2367,7 @@ function renderPaperclipWakePromptBody( } }; - const executionContractLines = externalChatContract + const wakeContractLines = externalChatContract ? [ externalChatQuestionResponseTurn ? "## External chat answered-question contract" @@ -2444,12 +2414,7 @@ function renderPaperclipWakePromptBody( `Fallback preference order: (1) send back to ${originalAssigneeLabel} with a retry instruction; (2) fix the runtime/adapter/workspace problem, then send it back; (3) reassign to another agent with the right specialty; (4) convert to an explicit manual-review state for the board.`, "", ] - : includeExecutionContract - ? [ - "Execution contract: take concrete action in this heartbeat when the issue is actionable; do not stop at a plan unless planning was requested. Leave durable progress and then give the issue a clear final disposition before ending the heartbeat: `done`, `in_review` with a real reviewer/approval/interaction path, `blocked` with first-class blockers or a named unblock owner/action, delegated follow-up issues with blockers, or `in_progress` only when a live continuation path exists. Immediately before returning, verify that Paperclip records one of those dispositions; a successful process exit or final response is not sufficient. If no valid disposition is recorded, record it now and do not end the run. After 2 consecutive failures of the same control-plane write, stop retrying it for the rest of the heartbeat, continue useful work, report the failure in the final response, and rely on the adapter/runtime status channel as the sanctioned fallback. Use child issues for long or parallel delegated work instead of polling. Comments, documents, screenshots, work products, and `Remaining` bullets are evidence, not valid liveness paths by themselves.", - "", - ] - : []; + : []; const wakeSummaryLines = [ `- reason: ${normalized.reason ?? "unknown"}`, `- issue: ${normalized.issue?.identifier ?? normalized.issue?.id ?? "unknown"}${normalized.issue?.title ? ` ${normalized.issue.title}` : ""}`, @@ -2507,7 +2472,7 @@ function renderPaperclipWakePromptBody( ]), "", ...externalInteractionContinuationLines, - ...executionContractLines, + ...wakeContractLines, ...wakeSummaryLines, ] : [ @@ -2537,7 +2502,7 @@ function renderPaperclipWakePromptBody( : []), "", ...externalInteractionContinuationLines, - ...executionContractLines, + ...wakeContractLines, ...wakeSummaryLines, ]; diff --git a/packages/adapters/cursor-cloud/src/server/execute.test.ts b/packages/adapters/cursor-cloud/src/server/execute.test.ts index db4617fa81..b57f820781 100644 --- a/packages/adapters/cursor-cloud/src/server/execute.test.ts +++ b/packages/adapters/cursor-cloud/src/server/execute.test.ts @@ -164,7 +164,7 @@ describe("cursor_cloud execute", () => { expect(result.exitCode).toBe(0); const prompt = String(sdkAgent.send.mock.calls[0]?.[0]); expect(prompt).toContain(directive); - expect(prompt).toContain(custom ? "Do the work for" : "Continue your Paperclip conversation"); + expect(prompt).toContain(custom ? "Do the work for" : "You are agent agent-1"); expect(prompt).not.toContain("Execution contract:"); expect(prompt).not.toContain("Create child issues"); }); @@ -178,6 +178,9 @@ describe("cursor_cloud execute", () => { expect(result.exitCode).toBe(0); const prompt = String(sdkAgent.send.mock.calls[0]?.[0]); expect(prompt).toContain("## Owned assignment"); + expect(prompt).toContain("You are agent agent-1 (Cursor Cloud Agent)."); + expect(prompt).toContain("Connection tools:"); + expect(prompt).not.toContain("Execution contract:"); expect(prompt.indexOf("Append the same ledger entry.")).toBeLessThan(prompt.indexOf("Change the final scope to the launch checklist.")); expect(prompt.split("Append the same ledger entry.")).toHaveLength(3); }); diff --git a/packages/adapters/hermes/src/server/prompt-rendering.test.ts b/packages/adapters/hermes/src/server/prompt-rendering.test.ts index d12da18758..b47dca0361 100644 --- a/packages/adapters/hermes/src/server/prompt-rendering.test.ts +++ b/packages/adapters/hermes/src/server/prompt-rendering.test.ts @@ -82,7 +82,7 @@ test("renders standard assignment wake with task authority and no backlog discov expect(prompt).toContain("Paperclip task context:"); expect(prompt).toContain("Add focused unit tests for assignment wake and custom prompt rendering."); expect(prompt).toContain("The harness already checked out this issue for the current run."); - expect(prompt).toContain("clear final disposition"); + expect(prompt).not.toContain("clear final disposition"); expect(prompt).not.toContain("check for unassigned issues"); expect(prompt).not.toContain("status=backlog"); }); @@ -130,8 +130,8 @@ test("renders scoped planning wake authority before the Hermes default workflow" expect(prompt).toContain("- checkout: already claimed by the harness for this run"); expect(prompt).toContain("The harness already checked out this issue for the current run."); expect(prompt).toContain("Issue description:\n```text\nUse the wake payload as runtime authority.\n```"); - expect(prompt).toContain("clear final disposition"); - expect(prompt).toContain("keep `in_progress` only when a live continuation path exists"); + expect(prompt).not.toContain("clear final disposition"); + expect(prompt).not.toContain("keep `in_progress` only when a live continuation path exists"); expect(prompt).not.toContain("check for unassigned issues"); expect(prompt).not.toContain("status=backlog"); }); diff --git a/packages/adapters/opencode-local/src/server/execute.test.ts b/packages/adapters/opencode-local/src/server/execute.test.ts index d1480d95b2..a754b9d4b9 100644 --- a/packages/adapters/opencode-local/src/server/execute.test.ts +++ b/packages/adapters/opencode-local/src/server/execute.test.ts @@ -77,7 +77,7 @@ describe("OpenCode local skill injection", () => { }); expect(result.exitCode).toBe(0); expect(prompt).toContain(directive); - expect(prompt).toContain(custom ? "Custom agent instruction." : "Continue your Paperclip conversation"); + expect(prompt).toContain(custom ? "Custom agent instruction." : "You are agent agent-1"); expect(prompt).not.toContain("Execution contract:"); expect(prompt).not.toContain("Create child issues"); }); @@ -102,6 +102,9 @@ describe("OpenCode local skill injection", () => { expect(prompts).toHaveLength(2); expect(prompts[0]).toContain("## Compact assignment"); expect(prompts[1]).toContain("## Owned assignment"); + for (const prompt of prompts) expect(prompt).not.toContain("Execution contract:"); + expect(prompts[1]).toContain("You are agent agent-1 (OpenCode)."); + expect(prompts[1]).toContain("Connection tools:"); }); it("injects runtime skills into the configured child HOME", async () => { diff --git a/packages/adapters/pi-local/src/server/execute.remote.test.ts b/packages/adapters/pi-local/src/server/execute.remote.test.ts index 99cd81345c..6f443902ff 100644 --- a/packages/adapters/pi-local/src/server/execute.remote.test.ts +++ b/packages/adapters/pi-local/src/server/execute.remote.test.ts @@ -696,7 +696,7 @@ describe("pi remote execution", () => { expect(userPrompt).toBe("CUSTOM POLICY run-custom-policy"); }); - it("keeps the resumed default execution contract in the system carrier only", async () => { + it("keeps resumed default identity and connection guidance in the system carrier", async () => { const rootDir = await mkdtemp(path.join(os.tmpdir(), "paperclip-pi-resumed-policy-")); cleanupDirs.push(rootDir); const sessionPath = path.join(rootDir, "session.jsonl"); @@ -726,11 +726,13 @@ describe("pi remote execution", () => { const args = call?.[2] ?? []; const systemPrompt = args[args.indexOf("--append-system-prompt") + 1] ?? ""; const userPrompt = args.at(-1) ?? ""; - expect(systemPrompt).toContain("Execution contract:"); + expect(systemPrompt).toContain("You are agent agent-1 (Pi Builder)."); + expect(systemPrompt).toContain("Connection tools:"); + expect(systemPrompt).not.toContain("Execution contract:"); expect(userPrompt).not.toContain("Execution contract:"); }); - it("keeps the default contract in system input when custom prompt uses loaded instructions", async () => { + it("keeps default identity in system input when custom prompt uses loaded instructions", async () => { const rootDir = await mkdtemp(path.join(os.tmpdir(), "paperclip-pi-instructions-policy-")); cleanupDirs.push(rootDir); const sessionPath = path.join(rootDir, "session.jsonl"); @@ -769,7 +771,9 @@ describe("pi remote execution", () => { const systemPrompt = args[args.indexOf("--append-system-prompt") + 1] ?? ""; const userPrompt = args.at(-1) ?? ""; expect(systemPrompt).toContain("Loaded instructions for this run."); - expect(systemPrompt).toContain("Execution contract:"); + expect(systemPrompt).toContain("You are agent agent-1 (Pi Builder)."); + expect(systemPrompt).toContain("Connection tools:"); + expect(systemPrompt).not.toContain("Execution contract:"); expect(userPrompt).not.toContain("CUSTOM POLICY run-instructions-policy"); expect(userPrompt).not.toContain("Execution contract:"); }); diff --git a/packages/paperclip-runner/docs/capability-contract.md b/packages/paperclip-runner/docs/capability-contract.md index 15e6d7d4c1..cf14bc5996 100644 --- a/packages/paperclip-runner/docs/capability-contract.md +++ b/packages/paperclip-runner/docs/capability-contract.md @@ -8,9 +8,9 @@ The skill/reference inventory and eval cases are the only normative behavior sou ## Baseline Counts -- Skill/reference headings: 157 +- Skill/reference headings: 160 - Eval cases: 106 across 16 groups -- Total normative rows: 263 +- Total normative rows: 266 - Legacy MCP aliases folded into normative rows: 42 | Eval group | Cases | @@ -43,91 +43,35 @@ The skill/reference inventory and eval cases are the only normative behavior sou | --- | --- | --- | | skill:skills/paperclip/SKILL.md:paperclip-skill:10 | optional_agent_tool | skills/paperclip/SKILL.md:10 | | skill:skills/paperclip/SKILL.md:terminology:14 | optional_agent_tool | skills/paperclip/SKILL.md:14 | -| skill:skills/paperclip/SKILL.md:authentication:18 | control_plane_owned | skills/paperclip/SKILL.md:18 | -| skill:skills/paperclip/SKILL.md:conversation-tasks:30 | optional_agent_tool | skills/paperclip/SKILL.md:30 | -| skill:skills/paperclip/SKILL.md:server-verified-external-chat-turns:47 | control_plane_owned | skills/paperclip/SKILL.md:47 | -| skill:skills/paperclip/SKILL.md:the-heartbeat-procedure:87 | optional_agent_tool | skills/paperclip/SKILL.md:87 | -| skill:skills/paperclip/SKILL.md:generated-artifacts-and-work-products:196 | always_agent_tool | skills/paperclip/SKILL.md:196 | -| skill:skills/paperclip/SKILL.md:status-quick-guide:244 | control_plane_owned | skills/paperclip/SKILL.md:244 | -| skill:skills/paperclip/SKILL.md:monitors-and-watchers-say-only-what-you-actually-scheduled:254 | optional_agent_tool | skills/paperclip/SKILL.md:254 | -| skill:skills/paperclip/SKILL.md:delegating-review-tasks:267 | always_agent_tool | skills/paperclip/SKILL.md:267 | -| skill:skills/paperclip/SKILL.md:managing-a-user-s-inbox:278 | control_plane_owned | skills/paperclip/SKILL.md:278 | -| skill:skills/paperclip/SKILL.md:issue-dependencies-blockers:286 | control_plane_owned | skills/paperclip/SKILL.md:286 | -| skill:skills/paperclip/SKILL.md:requesting-board-approval:311 | optional_agent_tool | skills/paperclip/SKILL.md:311 | -| skill:skills/paperclip/SKILL.md:issue-thread-interactions:332 | optional_agent_tool | skills/paperclip/SKILL.md:332 | -| skill:skills/paperclip/SKILL.md:standalone-decisions:361 | optional_agent_tool | skills/paperclip/SKILL.md:361 | -| skill:skills/paperclip/SKILL.md:mcp-tool-approval-gates:465 | optional_agent_tool | skills/paperclip/SKILL.md:465 | -| skill:skills/paperclip/SKILL.md:niche-workflow-pointers:507 | optional_agent_tool | skills/paperclip/SKILL.md:507 | -| skill:skills/paperclip/SKILL.md:cases:517 | optional_agent_tool | skills/paperclip/SKILL.md:517 | -| skill:skills/paperclip/SKILL.md:company-skills-workflow:522 | optional_agent_tool | skills/paperclip/SKILL.md:522 | -| skill:skills/paperclip/SKILL.md:routines:533 | optional_agent_tool | skills/paperclip/SKILL.md:533 | -| skill:skills/paperclip/SKILL.md:issue-workspace-runtime-controls:544 | optional_agent_tool | skills/paperclip/SKILL.md:544 | -| skill:skills/paperclip/SKILL.md:proposing-credentials-safely:551 | optional_agent_tool | skills/paperclip/SKILL.md:551 | -| skill:skills/paperclip/SKILL.md:reading-granted-secrets:558 | optional_agent_tool | skills/paperclip/SKILL.md:558 | -| skill:skills/paperclip/SKILL.md:critical-rules:584 | optional_agent_tool | skills/paperclip/SKILL.md:584 | -| skill:skills/paperclip/SKILL.md:comment-style-required:605 | always_agent_tool | skills/paperclip/SKILL.md:605 | -| skill:skills/paperclip/SKILL.md:update:637 | optional_agent_tool | skills/paperclip/SKILL.md:637 | -| skill:skills/paperclip/SKILL.md:planning-required-when-planning-requested:647 | optional_agent_tool | skills/paperclip/SKILL.md:647 | -| skill:skills/paperclip/SKILL.md:key-endpoints-hot-routes:680 | optional_agent_tool | skills/paperclip/SKILL.md:680 | -| skill:skills/paperclip/SKILL.md:searching-issues:709 | optional_agent_tool | skills/paperclip/SKILL.md:709 | -| skill:skills/paperclip/SKILL.md:full-reference:719 | optional_agent_tool | skills/paperclip/SKILL.md:719 | -| skill:skills/paperclip/SKILL.md:conversational-confirmation-answers:723 | always_agent_tool | skills/paperclip/SKILL.md:723 | -| skill:skills/paperclip/references/artifacts.md:generated-artifacts-and-work-products:1 | always_agent_tool | skills/paperclip/references/artifacts.md:1 | -| skill:skills/paperclip/references/artifacts.md:workspace-only-file-references:15 | optional_agent_tool | skills/paperclip/references/artifacts.md:15 | -| skill:skills/paperclip/references/cases.md:cases:1 | optional_agent_tool | skills/paperclip/references/cases.md:1 | -| skill:skills/paperclip/references/cases.md:core-model:12 | optional_agent_tool | skills/paperclip/references/cases.md:12 | -| skill:skills/paperclip/references/cases.md:upsert-semantics:29 | optional_agent_tool | skills/paperclip/references/cases.md:29 | -| skill:skills/paperclip/references/cases.md:read-and-search:67 | optional_agent_tool | skills/paperclip/references/cases.md:67 | -| skill:skills/paperclip/references/cases.md:documents:90 | always_agent_tool | skills/paperclip/references/cases.md:90 | -| skill:skills/paperclip/references/cases.md:fields:118 | optional_agent_tool | skills/paperclip/references/cases.md:118 | -| skill:skills/paperclip/references/cases.md:issue-links:151 | optional_agent_tool | skills/paperclip/references/cases.md:151 | -| skill:skills/paperclip/references/cases.md:child-cases:176 | optional_agent_tool | skills/paperclip/references/cases.md:176 | -| skill:skills/paperclip/references/cases.md:attachments:195 | optional_agent_tool | skills/paperclip/references/cases.md:195 | -| skill:skills/paperclip/references/cases.md:lifecycle:208 | optional_agent_tool | skills/paperclip/references/cases.md:208 | -| skill:skills/paperclip/references/cases.md:worked-blog-post-example:222 | optional_agent_tool | skills/paperclip/references/cases.md:222 | -| skill:skills/paperclip/references/company-skills.md:company-skills-workflow:1 | optional_agent_tool | skills/paperclip/references/company-skills.md:1 | -| skill:skills/paperclip/references/company-skills.md:what-exists:5 | optional_agent_tool | skills/paperclip/references/company-skills.md:5 | -| skill:skills/paperclip/references/company-skills.md:permission-model:22 | optional_agent_tool | skills/paperclip/references/company-skills.md:22 | -| skill:skills/paperclip/references/company-skills.md:core-endpoints:29 | optional_agent_tool | skills/paperclip/references/company-skills.md:29 | -| skill:skills/paperclip/references/company-skills.md:install-a-skill-into-the-company:65 | optional_agent_tool | skills/paperclip/references/company-skills.md:65 | -| skill:skills/paperclip/references/company-skills.md:app-shipped-catalog:75 | optional_agent_tool | skills/paperclip/references/company-skills.md:75 | -| skill:skills/paperclip/references/company-skills.md:external-source-import:101 | optional_agent_tool | skills/paperclip/references/company-skills.md:101 | -| skill:skills/paperclip/references/company-skills.md:source-types-in-order-of-preference:105 | optional_agent_tool | skills/paperclip/references/company-skills.md:105 | -| skill:skills/paperclip/references/company-skills.md:example-skills-sh-import-preferred:116 | optional_agent_tool | skills/paperclip/references/company-skills.md:116 | -| skill:skills/paperclip/references/company-skills.md:example-github-import:138 | optional_agent_tool | skills/paperclip/references/company-skills.md:138 | -| skill:skills/paperclip/references/company-skills.md:inspect-what-was-installed:164 | optional_agent_tool | skills/paperclip/references/company-skills.md:164 | -| skill:skills/paperclip/references/company-skills.md:assign-skills-to-an-existing-agent:181 | optional_agent_tool | skills/paperclip/references/company-skills.md:181 | -| skill:skills/paperclip/references/company-skills.md:include-skills-during-hire-or-create:216 | optional_agent_tool | skills/paperclip/references/company-skills.md:216 | -| skill:skills/paperclip/references/company-skills.md:notes:256 | optional_agent_tool | skills/paperclip/references/company-skills.md:256 | -| skill:skills/paperclip/references/issue-workspaces.md:issue-workspace-runtime-controls:1 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:1 | -| skill:skills/paperclip/references/issue-workspaces.md:discover-the-workspace:5 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:5 | -| skill:skills/paperclip/references/issue-workspaces.md:control-services:23 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:23 | -| skill:skills/paperclip/references/issue-workspaces.md:start-all-configured-services-waits-for-configured-readiness-checks:28 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:28 | -| skill:skills/paperclip/references/issue-workspaces.md:restart-all-configured-services:36 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:36 | -| skill:skills/paperclip/references/issue-workspaces.md:stop-all-running-services:44 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:44 | -| skill:skills/paperclip/references/issue-workspaces.md:read-the-url:63 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:63 | -| skill:skills/paperclip/references/issue-workspaces.md:mcp-tools:72 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:72 | -| skill:skills/paperclip/references/routines.md:paperclip-routines:1 | optional_agent_tool | skills/paperclip/references/routines.md:1 | -| skill:skills/paperclip/references/routines.md:lifecycle:16 | optional_agent_tool | skills/paperclip/references/routines.md:16 | -| skill:skills/paperclip/references/routines.md:creating-a-routine:27 | optional_agent_tool | skills/paperclip/references/routines.md:27 | -| skill:skills/paperclip/references/routines.md:concurrency-policies:64 | optional_agent_tool | skills/paperclip/references/routines.md:64 | -| skill:skills/paperclip/references/routines.md:catch-up-policies:76 | optional_agent_tool | skills/paperclip/references/routines.md:76 | -| skill:skills/paperclip/references/routines.md:activity-gated-scheduled-runs:87 | optional_agent_tool | skills/paperclip/references/routines.md:87 | -| skill:skills/paperclip/references/routines.md:example-skip-quiet-nights:107 | optional_agent_tool | skills/paperclip/references/routines.md:107 | -| skill:skills/paperclip/references/routines.md:adding-triggers:126 | optional_agent_tool | skills/paperclip/references/routines.md:126 | -| skill:skills/paperclip/references/routines.md:schedule-cron:136 | optional_agent_tool | skills/paperclip/references/routines.md:136 | -| skill:skills/paperclip/references/routines.md:webhook:150 | optional_agent_tool | skills/paperclip/references/routines.md:150 | -| skill:skills/paperclip/references/routines.md:api-manual-only:167 | optional_agent_tool | skills/paperclip/references/routines.md:167 | -| skill:skills/paperclip/references/routines.md:updating-and-deleting-triggers:179 | optional_agent_tool | skills/paperclip/references/routines.md:179 | -| skill:skills/paperclip/references/routines.md:manual-run:196 | optional_agent_tool | skills/paperclip/references/routines.md:196 | -| skill:skills/paperclip/references/routines.md:updating-a-routine:212 | optional_agent_tool | skills/paperclip/references/routines.md:212 | -| skill:skills/paperclip/references/routines.md:reading-routines-and-runs:223 | optional_agent_tool | skills/paperclip/references/routines.md:223 | -| skill:skills/paperclip/references/workflows.md:paperclip-workflow-playbooks:1 | optional_agent_tool | skills/paperclip/references/workflows.md:1 | -| skill:skills/paperclip/references/workflows.md:project-setup-ceo-manager:7 | optional_agent_tool | skills/paperclip/references/workflows.md:7 | -| skill:skills/paperclip/references/workflows.md:openclaw-invite-ceo:22 | optional_agent_tool | skills/paperclip/references/workflows.md:22 | -| skill:skills/paperclip/references/workflows.md:setting-agent-instructions-path:50 | optional_agent_tool | skills/paperclip/references/workflows.md:50 | -| skill:skills/paperclip/references/workflows.md:company-import-export:79 | optional_agent_tool | skills/paperclip/references/workflows.md:79 | -| skill:skills/paperclip/references/workflows.md:self-test-playbook-app-level:106 | optional_agent_tool | skills/paperclip/references/workflows.md:106 | +| skill:skills/paperclip/SKILL.md:authentication:20 | control_plane_owned | skills/paperclip/SKILL.md:20 | +| skill:skills/paperclip/SKILL.md:conversation-tasks:32 | optional_agent_tool | skills/paperclip/SKILL.md:32 | +| skill:skills/paperclip/SKILL.md:server-verified-external-chat-turns:49 | control_plane_owned | skills/paperclip/SKILL.md:49 | +| skill:skills/paperclip/SKILL.md:the-heartbeat-procedure:89 | optional_agent_tool | skills/paperclip/SKILL.md:89 | +| skill:skills/paperclip/SKILL.md:generated-artifacts-and-work-products:198 | always_agent_tool | skills/paperclip/SKILL.md:198 | +| skill:skills/paperclip/SKILL.md:status-quick-guide:246 | control_plane_owned | skills/paperclip/SKILL.md:246 | +| skill:skills/paperclip/SKILL.md:monitors-and-watchers-say-only-what-you-actually-scheduled:256 | optional_agent_tool | skills/paperclip/SKILL.md:256 | +| skill:skills/paperclip/SKILL.md:delegating-review-tasks:269 | always_agent_tool | skills/paperclip/SKILL.md:269 | +| skill:skills/paperclip/SKILL.md:managing-a-user-s-inbox:280 | control_plane_owned | skills/paperclip/SKILL.md:280 | +| skill:skills/paperclip/SKILL.md:issue-dependencies-blockers:288 | control_plane_owned | skills/paperclip/SKILL.md:288 | +| skill:skills/paperclip/SKILL.md:requesting-board-approval:313 | optional_agent_tool | skills/paperclip/SKILL.md:313 | +| skill:skills/paperclip/SKILL.md:issue-thread-interactions:334 | optional_agent_tool | skills/paperclip/SKILL.md:334 | +| skill:skills/paperclip/SKILL.md:standalone-decisions:363 | optional_agent_tool | skills/paperclip/SKILL.md:363 | +| skill:skills/paperclip/SKILL.md:mcp-tool-approval-gates:467 | optional_agent_tool | skills/paperclip/SKILL.md:467 | +| skill:skills/paperclip/SKILL.md:niche-workflow-pointers:509 | optional_agent_tool | skills/paperclip/SKILL.md:509 | +| skill:skills/paperclip/SKILL.md:cases:519 | optional_agent_tool | skills/paperclip/SKILL.md:519 | +| skill:skills/paperclip/SKILL.md:company-skills-workflow:524 | optional_agent_tool | skills/paperclip/SKILL.md:524 | +| skill:skills/paperclip/SKILL.md:routines:535 | optional_agent_tool | skills/paperclip/SKILL.md:535 | +| skill:skills/paperclip/SKILL.md:issue-workspace-runtime-controls:546 | optional_agent_tool | skills/paperclip/SKILL.md:546 | +| skill:skills/paperclip/SKILL.md:proposing-credentials-safely:553 | optional_agent_tool | skills/paperclip/SKILL.md:553 | +| skill:skills/paperclip/SKILL.md:reading-granted-secrets:560 | optional_agent_tool | skills/paperclip/SKILL.md:560 | +| skill:skills/paperclip/SKILL.md:critical-rules:586 | optional_agent_tool | skills/paperclip/SKILL.md:586 | +| skill:skills/paperclip/SKILL.md:comment-style-required:607 | always_agent_tool | skills/paperclip/SKILL.md:607 | +| skill:skills/paperclip/SKILL.md:update:639 | optional_agent_tool | skills/paperclip/SKILL.md:639 | +| skill:skills/paperclip/SKILL.md:planning-required-when-planning-requested:649 | optional_agent_tool | skills/paperclip/SKILL.md:649 | +| skill:skills/paperclip/SKILL.md:key-endpoints-hot-routes:682 | optional_agent_tool | skills/paperclip/SKILL.md:682 | +| skill:skills/paperclip/SKILL.md:searching-issues:711 | optional_agent_tool | skills/paperclip/SKILL.md:711 | +| skill:skills/paperclip/SKILL.md:full-reference:721 | optional_agent_tool | skills/paperclip/SKILL.md:721 | +| skill:skills/paperclip/SKILL.md:conversational-confirmation-answers:725 | always_agent_tool | skills/paperclip/SKILL.md:725 | | skill:skills/paperclip/references/api-reference.md:paperclip-api-reference:1 | optional_agent_tool | skills/paperclip/references/api-reference.md:1 | | skill:skills/paperclip/references/api-reference.md:response-schemas:9 | optional_agent_tool | skills/paperclip/references/api-reference.md:9 | | skill:skills/paperclip/references/api-reference.md:agent-record-get-api-agents-me-or-get-api-agents-agentid:11 | optional_agent_tool | skills/paperclip/references/api-reference.md:11 | @@ -198,6 +142,65 @@ The skill/reference inventory and eval cases are the only normative behavior sou | skill:skills/paperclip/references/api-reference.md:agent-secret-proposals:1521 | optional_agent_tool | skills/paperclip/references/api-reference.md:1521 | | skill:skills/paperclip/references/api-reference.md:agent-secret-access:1621 | optional_agent_tool | skills/paperclip/references/api-reference.md:1621 | | skill:skills/paperclip/references/api-reference.md:common-mistakes:1661 | optional_agent_tool | skills/paperclip/references/api-reference.md:1661 | +| skill:skills/paperclip/references/artifacts.md:generated-artifacts-and-work-products:1 | always_agent_tool | skills/paperclip/references/artifacts.md:1 | +| skill:skills/paperclip/references/artifacts.md:workspace-only-file-references:15 | optional_agent_tool | skills/paperclip/references/artifacts.md:15 | +| skill:skills/paperclip/references/cases.md:cases:1 | optional_agent_tool | skills/paperclip/references/cases.md:1 | +| skill:skills/paperclip/references/cases.md:core-model:12 | optional_agent_tool | skills/paperclip/references/cases.md:12 | +| skill:skills/paperclip/references/cases.md:upsert-semantics:29 | optional_agent_tool | skills/paperclip/references/cases.md:29 | +| skill:skills/paperclip/references/cases.md:read-and-search:67 | optional_agent_tool | skills/paperclip/references/cases.md:67 | +| skill:skills/paperclip/references/cases.md:documents:90 | always_agent_tool | skills/paperclip/references/cases.md:90 | +| skill:skills/paperclip/references/cases.md:fields:118 | optional_agent_tool | skills/paperclip/references/cases.md:118 | +| skill:skills/paperclip/references/cases.md:issue-links:151 | optional_agent_tool | skills/paperclip/references/cases.md:151 | +| skill:skills/paperclip/references/cases.md:child-cases:176 | optional_agent_tool | skills/paperclip/references/cases.md:176 | +| skill:skills/paperclip/references/cases.md:attachments:195 | optional_agent_tool | skills/paperclip/references/cases.md:195 | +| skill:skills/paperclip/references/cases.md:lifecycle:208 | optional_agent_tool | skills/paperclip/references/cases.md:208 | +| skill:skills/paperclip/references/cases.md:worked-blog-post-example:222 | optional_agent_tool | skills/paperclip/references/cases.md:222 | +| skill:skills/paperclip/references/company-skills.md:company-skills-workflow:1 | optional_agent_tool | skills/paperclip/references/company-skills.md:1 | +| skill:skills/paperclip/references/company-skills.md:what-exists:5 | optional_agent_tool | skills/paperclip/references/company-skills.md:5 | +| skill:skills/paperclip/references/company-skills.md:permission-model:22 | optional_agent_tool | skills/paperclip/references/company-skills.md:22 | +| skill:skills/paperclip/references/company-skills.md:core-endpoints:29 | optional_agent_tool | skills/paperclip/references/company-skills.md:29 | +| skill:skills/paperclip/references/company-skills.md:install-a-skill-into-the-company:65 | optional_agent_tool | skills/paperclip/references/company-skills.md:65 | +| skill:skills/paperclip/references/company-skills.md:app-shipped-catalog:75 | optional_agent_tool | skills/paperclip/references/company-skills.md:75 | +| skill:skills/paperclip/references/company-skills.md:external-source-import:101 | optional_agent_tool | skills/paperclip/references/company-skills.md:101 | +| skill:skills/paperclip/references/company-skills.md:source-types-in-order-of-preference:105 | optional_agent_tool | skills/paperclip/references/company-skills.md:105 | +| skill:skills/paperclip/references/company-skills.md:example-skills-sh-import-preferred:116 | optional_agent_tool | skills/paperclip/references/company-skills.md:116 | +| skill:skills/paperclip/references/company-skills.md:example-github-import:138 | optional_agent_tool | skills/paperclip/references/company-skills.md:138 | +| skill:skills/paperclip/references/company-skills.md:inspect-what-was-installed:164 | optional_agent_tool | skills/paperclip/references/company-skills.md:164 | +| skill:skills/paperclip/references/company-skills.md:assign-skills-to-an-existing-agent:181 | optional_agent_tool | skills/paperclip/references/company-skills.md:181 | +| skill:skills/paperclip/references/company-skills.md:include-skills-during-hire-or-create:216 | optional_agent_tool | skills/paperclip/references/company-skills.md:216 | +| skill:skills/paperclip/references/company-skills.md:notes:256 | optional_agent_tool | skills/paperclip/references/company-skills.md:256 | +| skill:skills/paperclip/references/issue-documents.md:issue-documents-through-the-api:1 | always_agent_tool | skills/paperclip/references/issue-documents.md:1 | +| skill:skills/paperclip/references/issue-documents.md:create-a-document:8 | always_agent_tool | skills/paperclip/references/issue-documents.md:8 | +| skill:skills/paperclip/references/issue-documents.md:update-or-resolve-an-unclear-write:49 | optional_agent_tool | skills/paperclip/references/issue-documents.md:49 | +| skill:skills/paperclip/references/issue-workspaces.md:issue-workspace-runtime-controls:1 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:1 | +| skill:skills/paperclip/references/issue-workspaces.md:discover-the-workspace:5 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:5 | +| skill:skills/paperclip/references/issue-workspaces.md:control-services:23 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:23 | +| skill:skills/paperclip/references/issue-workspaces.md:start-all-configured-services-waits-for-configured-readiness-checks:28 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:28 | +| skill:skills/paperclip/references/issue-workspaces.md:restart-all-configured-services:36 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:36 | +| skill:skills/paperclip/references/issue-workspaces.md:stop-all-running-services:44 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:44 | +| skill:skills/paperclip/references/issue-workspaces.md:read-the-url:63 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:63 | +| skill:skills/paperclip/references/issue-workspaces.md:mcp-tools:72 | optional_agent_tool | skills/paperclip/references/issue-workspaces.md:72 | +| skill:skills/paperclip/references/routines.md:paperclip-routines:1 | optional_agent_tool | skills/paperclip/references/routines.md:1 | +| skill:skills/paperclip/references/routines.md:lifecycle:16 | optional_agent_tool | skills/paperclip/references/routines.md:16 | +| skill:skills/paperclip/references/routines.md:creating-a-routine:27 | optional_agent_tool | skills/paperclip/references/routines.md:27 | +| skill:skills/paperclip/references/routines.md:concurrency-policies:64 | optional_agent_tool | skills/paperclip/references/routines.md:64 | +| skill:skills/paperclip/references/routines.md:catch-up-policies:76 | optional_agent_tool | skills/paperclip/references/routines.md:76 | +| skill:skills/paperclip/references/routines.md:activity-gated-scheduled-runs:87 | optional_agent_tool | skills/paperclip/references/routines.md:87 | +| skill:skills/paperclip/references/routines.md:example-skip-quiet-nights:107 | optional_agent_tool | skills/paperclip/references/routines.md:107 | +| skill:skills/paperclip/references/routines.md:adding-triggers:126 | optional_agent_tool | skills/paperclip/references/routines.md:126 | +| skill:skills/paperclip/references/routines.md:schedule-cron:136 | optional_agent_tool | skills/paperclip/references/routines.md:136 | +| skill:skills/paperclip/references/routines.md:webhook:150 | optional_agent_tool | skills/paperclip/references/routines.md:150 | +| skill:skills/paperclip/references/routines.md:api-manual-only:167 | optional_agent_tool | skills/paperclip/references/routines.md:167 | +| skill:skills/paperclip/references/routines.md:updating-and-deleting-triggers:179 | optional_agent_tool | skills/paperclip/references/routines.md:179 | +| skill:skills/paperclip/references/routines.md:manual-run:196 | optional_agent_tool | skills/paperclip/references/routines.md:196 | +| skill:skills/paperclip/references/routines.md:updating-a-routine:212 | optional_agent_tool | skills/paperclip/references/routines.md:212 | +| skill:skills/paperclip/references/routines.md:reading-routines-and-runs:223 | optional_agent_tool | skills/paperclip/references/routines.md:223 | +| skill:skills/paperclip/references/workflows.md:paperclip-workflow-playbooks:1 | optional_agent_tool | skills/paperclip/references/workflows.md:1 | +| skill:skills/paperclip/references/workflows.md:project-setup-ceo-manager:7 | optional_agent_tool | skills/paperclip/references/workflows.md:7 | +| skill:skills/paperclip/references/workflows.md:openclaw-invite-ceo:22 | optional_agent_tool | skills/paperclip/references/workflows.md:22 | +| skill:skills/paperclip/references/workflows.md:setting-agent-instructions-path:50 | optional_agent_tool | skills/paperclip/references/workflows.md:50 | +| skill:skills/paperclip/references/workflows.md:company-import-export:79 | optional_agent_tool | skills/paperclip/references/workflows.md:79 | +| skill:skills/paperclip/references/workflows.md:self-test-playbook-app-level:106 | optional_agent_tool | skills/paperclip/references/workflows.md:106 | ## Legacy MCP Alias Index diff --git a/packages/paperclip-runner/generated/capability/capabilities.yaml b/packages/paperclip-runner/generated/capability/capabilities.yaml index 61bba278fd..a2557fcd05 100644 --- a/packages/paperclip-runner/generated/capability/capabilities.yaml +++ b/packages/paperclip-runner/generated/capability/capabilities.yaml @@ -20,261 +20,261 @@ "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:18", + "id": "skill:skills/paperclip/SKILL.md:20", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L18:authentication", + "sourceAnchor": "skills/paperclip/SKILL.md#L20:authentication", "heading": "Authentication", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:30", + "id": "skill:skills/paperclip/SKILL.md:32", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L30:conversation-tasks", + "sourceAnchor": "skills/paperclip/SKILL.md#L32:conversation-tasks", "heading": "Conversation tasks", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:47", + "id": "skill:skills/paperclip/SKILL.md:49", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L47:server-verified-external-chat-turns", + "sourceAnchor": "skills/paperclip/SKILL.md#L49:server-verified-external-chat-turns", "heading": "Server-Verified External Chat Turns", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:87", + "id": "skill:skills/paperclip/SKILL.md:89", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L87:the-heartbeat-procedure", + "sourceAnchor": "skills/paperclip/SKILL.md#L89:the-heartbeat-procedure", "heading": "The Heartbeat Procedure", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:196", + "id": "skill:skills/paperclip/SKILL.md:198", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L196:generated-artifacts-and-work-products", + "sourceAnchor": "skills/paperclip/SKILL.md#L198:generated-artifacts-and-work-products", "heading": "Generated Artifacts and Work Products", "primaryDisposition": "always_agent_tool", "semanticOperation": "register_deliverable", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:244", + "id": "skill:skills/paperclip/SKILL.md:246", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L244:status-quick-guide", + "sourceAnchor": "skills/paperclip/SKILL.md#L246:status-quick-guide", "heading": "Status Quick Guide", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:254", + "id": "skill:skills/paperclip/SKILL.md:256", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L254:monitors-and-watchers-say-only-what-you-actually-scheduled", + "sourceAnchor": "skills/paperclip/SKILL.md#L256:monitors-and-watchers-say-only-what-you-actually-scheduled", "heading": "Monitors and Watchers (say only what you actually scheduled)", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:267", + "id": "skill:skills/paperclip/SKILL.md:269", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L267:delegating-review-tasks", + "sourceAnchor": "skills/paperclip/SKILL.md#L269:delegating-review-tasks", "heading": "Delegating review tasks", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:278", + "id": "skill:skills/paperclip/SKILL.md:280", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L278:managing-a-user-s-inbox", + "sourceAnchor": "skills/paperclip/SKILL.md#L280:managing-a-user-s-inbox", "heading": "Managing A User's Inbox", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:286", + "id": "skill:skills/paperclip/SKILL.md:288", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L286:issue-dependencies-blockers", + "sourceAnchor": "skills/paperclip/SKILL.md#L288:issue-dependencies-blockers", "heading": "Issue Dependencies (Blockers)", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:311", + "id": "skill:skills/paperclip/SKILL.md:313", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L311:requesting-board-approval", + "sourceAnchor": "skills/paperclip/SKILL.md#L313:requesting-board-approval", "heading": "Requesting Board Approval", "primaryDisposition": "optional_agent_tool", "semanticOperation": "scoped_discovery", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:332", + "id": "skill:skills/paperclip/SKILL.md:334", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L332:issue-thread-interactions", + "sourceAnchor": "skills/paperclip/SKILL.md#L334:issue-thread-interactions", "heading": "Issue-Thread Interactions", "primaryDisposition": "always_agent_tool", "semanticOperation": "request_human_input", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:361", + "id": "skill:skills/paperclip/SKILL.md:363", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L361:standalone-decisions", + "sourceAnchor": "skills/paperclip/SKILL.md#L363:standalone-decisions", "heading": "Standalone Decisions", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:465", + "id": "skill:skills/paperclip/SKILL.md:467", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L465:mcp-tool-approval-gates", + "sourceAnchor": "skills/paperclip/SKILL.md#L467:mcp-tool-approval-gates", "heading": "MCP Tool Approval Gates", "primaryDisposition": "optional_agent_tool", "semanticOperation": "scoped_discovery", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:507", + "id": "skill:skills/paperclip/SKILL.md:509", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L507:niche-workflow-pointers", + "sourceAnchor": "skills/paperclip/SKILL.md#L509:niche-workflow-pointers", "heading": "Niche Workflow Pointers", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:517", + "id": "skill:skills/paperclip/SKILL.md:519", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L517:cases", + "sourceAnchor": "skills/paperclip/SKILL.md#L519:cases", "heading": "Cases", "primaryDisposition": "optional_agent_tool", "semanticOperation": "scoped_discovery", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:522", + "id": "skill:skills/paperclip/SKILL.md:524", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L522:company-skills-workflow", + "sourceAnchor": "skills/paperclip/SKILL.md#L524:company-skills-workflow", "heading": "Company Skills Workflow", "primaryDisposition": "optional_agent_tool", "semanticOperation": "scoped_discovery", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:533", + "id": "skill:skills/paperclip/SKILL.md:535", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L533:routines", + "sourceAnchor": "skills/paperclip/SKILL.md#L535:routines", "heading": "Routines", "primaryDisposition": "optional_agent_tool", "semanticOperation": "scoped_discovery", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:544", + "id": "skill:skills/paperclip/SKILL.md:546", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L544:issue-workspace-runtime-controls", + "sourceAnchor": "skills/paperclip/SKILL.md#L546:issue-workspace-runtime-controls", "heading": "Issue Workspace Runtime Controls", "primaryDisposition": "optional_agent_tool", "semanticOperation": "scoped_discovery", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:551", + "id": "skill:skills/paperclip/SKILL.md:553", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L551:proposing-credentials-safely", + "sourceAnchor": "skills/paperclip/SKILL.md#L553:proposing-credentials-safely", "heading": "Proposing Credentials Safely", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:558", + "id": "skill:skills/paperclip/SKILL.md:560", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L558:reading-granted-secrets", + "sourceAnchor": "skills/paperclip/SKILL.md#L560:reading-granted-secrets", "heading": "Reading Granted Secrets", "primaryDisposition": "optional_agent_tool", "semanticOperation": "scoped_discovery", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:584", + "id": "skill:skills/paperclip/SKILL.md:586", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L584:critical-rules", + "sourceAnchor": "skills/paperclip/SKILL.md#L586:critical-rules", "heading": "Critical Rules", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:605", + "id": "skill:skills/paperclip/SKILL.md:607", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L605:comment-style-required", + "sourceAnchor": "skills/paperclip/SKILL.md#L607:comment-style-required", "heading": "Comment Style (Required)", "primaryDisposition": "always_agent_tool", "semanticOperation": "report_progress", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:637", + "id": "skill:skills/paperclip/SKILL.md:639", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L637:update", + "sourceAnchor": "skills/paperclip/SKILL.md#L639:update", "heading": "Update", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:647", + "id": "skill:skills/paperclip/SKILL.md:649", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L647:planning-required-when-planning-requested", + "sourceAnchor": "skills/paperclip/SKILL.md#L649:planning-required-when-planning-requested", "heading": "Planning (Required when planning requested)", "primaryDisposition": "always_agent_tool", "semanticOperation": "write_document", "expectedMockState": "operation_result" }, { - "id": "skill:skills/paperclip/SKILL.md:680", + "id": "skill:skills/paperclip/SKILL.md:682", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L680:key-endpoints-hot-routes", + "sourceAnchor": "skills/paperclip/SKILL.md#L682:key-endpoints-hot-routes", "heading": "Key Endpoints (Hot Routes)", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:709", + "id": "skill:skills/paperclip/SKILL.md:711", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L709:searching-issues", + "sourceAnchor": "skills/paperclip/SKILL.md#L711:searching-issues", "heading": "Searching Issues", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:719", + "id": "skill:skills/paperclip/SKILL.md:721", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L719:full-reference", + "sourceAnchor": "skills/paperclip/SKILL.md#L721:full-reference", "heading": "Full Reference", "primaryDisposition": "control_plane_owned", "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, { - "id": "skill:skills/paperclip/SKILL.md:723", + "id": "skill:skills/paperclip/SKILL.md:725", "kind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md#L723:conversational-confirmation-answers", + "sourceAnchor": "skills/paperclip/SKILL.md#L725:conversational-confirmation-answers", "heading": "Conversational confirmation answers", "primaryDisposition": "always_agent_tool", "semanticOperation": "call_api", @@ -1162,6 +1162,33 @@ "semanticOperation": "runtime_reconciliation", "expectedMockState": "runtime_decision_record" }, + { + "id": "skill:skills/paperclip/references/issue-documents.md:1", + "kind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-documents.md#L1:issue-documents-through-the-api", + "heading": "Issue documents through the API", + "primaryDisposition": "always_agent_tool", + "semanticOperation": "write_document", + "expectedMockState": "operation_result" + }, + { + "id": "skill:skills/paperclip/references/issue-documents.md:8", + "kind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-documents.md#L8:create-a-document", + "heading": "Create a document", + "primaryDisposition": "always_agent_tool", + "semanticOperation": "write_document", + "expectedMockState": "operation_result" + }, + { + "id": "skill:skills/paperclip/references/issue-documents.md:49", + "kind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-documents.md#L49:update-or-resolve-an-unclear-write", + "heading": "Update or resolve an unclear write", + "primaryDisposition": "control_plane_owned", + "semanticOperation": "runtime_reconciliation", + "expectedMockState": "runtime_decision_record" + }, { "id": "skill:skills/paperclip/references/issue-workspaces.md:1", "kind": "skill_heading", diff --git a/packages/paperclip-runner/generated/capability/capability-contract.md b/packages/paperclip-runner/generated/capability/capability-contract.md index 907bf749bd..4d12461301 100644 --- a/packages/paperclip-runner/generated/capability/capability-contract.md +++ b/packages/paperclip-runner/generated/capability/capability-contract.md @@ -2,9 +2,9 @@ Generated by `scripts/generate-capability-contract.mjs`; do not edit generated files. -- Skill/reference headings: 158 +- Skill/reference headings: 161 - Legacy MCP tools: 42 - Eval cases: 106 across 16 groups -- Deterministic content SHA-256: `2707de482cfbe6c7a483f6dee3bf4a2fc66e2b944bb50c3818ca6d5ef777b1de` +- Deterministic content SHA-256: `e577ca3ad6e6b045a7413ac11008013c8c12c2dc954f5a7876b86d96fbe017e8` Every row has exactly one primary disposition, a source anchor, a semantic operation, and a mock-state expectation. diff --git a/packages/paperclip-runner/scripts/check-capability-inventory.test.mjs b/packages/paperclip-runner/scripts/check-capability-inventory.test.mjs index f81047ccf4..b15372d329 100644 --- a/packages/paperclip-runner/scripts/check-capability-inventory.test.mjs +++ b/packages/paperclip-runner/scripts/check-capability-inventory.test.mjs @@ -42,7 +42,7 @@ function validInventories() { schemaVersion: 2, inventoryRole: "normative", generatedFrom: ["skills/paperclip/SKILL.md"], - rows: Array.from({ length: 157 }, (_, index) => row(`capability-${index}`)), + rows: Array.from({ length: 160 }, (_, index) => row(`capability-${index}`)), }, evaluations: { schemaVersion: 2, diff --git a/packages/paperclip-runner/scripts/generate-capability-contract.d.mts b/packages/paperclip-runner/scripts/generate-capability-contract.d.mts new file mode 100644 index 0000000000..401886022f --- /dev/null +++ b/packages/paperclip-runner/scripts/generate-capability-contract.d.mts @@ -0,0 +1 @@ +export function buildContract(): Promise>; diff --git a/packages/paperclip-runner/scripts/generate-capability-contract.mjs b/packages/paperclip-runner/scripts/generate-capability-contract.mjs index 840b94521c..7903efd7ca 100644 --- a/packages/paperclip-runner/scripts/generate-capability-contract.mjs +++ b/packages/paperclip-runner/scripts/generate-capability-contract.mjs @@ -159,7 +159,7 @@ function renderHandoff() { ].join("\n") + "\n"; } -async function buildContract() { +export async function buildContract() { const contract = JSON.parse(await readFile(contractPath, "utf8")); const capabilities = await readSkillHeadings(contract.skillSources); const discoveredTools = await readLegacyTools(); diff --git a/packages/paperclip-runner/scripts/lib/capability-inventory.mjs b/packages/paperclip-runner/scripts/lib/capability-inventory.mjs index be9dd2ce7d..5f04ef61ee 100644 --- a/packages/paperclip-runner/scripts/lib/capability-inventory.mjs +++ b/packages/paperclip-runner/scripts/lib/capability-inventory.mjs @@ -221,16 +221,16 @@ export async function buildInventories({ repoRoot, evalRoot }) { } export async function buildSkillInventory(repoRoot) { - const skillRoot = resolve(repoRoot, "skills/paperclip"); - const skillFiles = ["SKILL.md", "references/artifacts.md", "references/cases.md", "references/company-skills.md", "references/issue-workspaces.md", "references/routines.md", "references/workflows.md", "references/api-reference.md"]; + const { skillSources: skillFiles } = JSON.parse(await readFile(resolve(repoRoot, + "packages/paperclip-runner/spec/capability/source-contract.json"), "utf8")); const rows = (await Promise.all(skillFiles.map(async (file) => parseSkillHeadings( - await readFile(resolve(skillRoot, file), "utf8"), - `skills/paperclip/${file}`, + await readFile(resolve(repoRoot, file), "utf8"), + file, )))).flat(); return { schemaVersion: 2, inventoryRole: "normative", - generatedFrom: skillFiles.map((file) => `skills/paperclip/${file}`), + generatedFrom: skillFiles, rows, }; } @@ -247,7 +247,7 @@ export async function buildMcpInventory(repoRoot) { export function validateInventories(inventories) { const errors = []; - const expectedCounts = { capabilities: 157, evaluations: 106, legacyMcpAliases: 42 }; + const expectedCounts = { capabilities: 160, evaluations: 106, legacyMcpAliases: 42 }; const normativeNames = ["capabilities", "evaluations"]; const normativeRows = new Map(); const globalNormativeIds = new Set(); diff --git a/packages/paperclip-runner/spec/capability/capabilities.yaml b/packages/paperclip-runner/spec/capability/capabilities.yaml index 5e07d65358..4170503c4b 100644 --- a/packages/paperclip-runner/spec/capability/capabilities.yaml +++ b/packages/paperclip-runner/spec/capability/capabilities.yaml @@ -4,13 +4,14 @@ "inventoryRole": "normative", "generatedFrom": [ "skills/paperclip/SKILL.md", + "skills/paperclip/references/api-reference.md", "skills/paperclip/references/artifacts.md", "skills/paperclip/references/cases.md", "skills/paperclip/references/company-skills.md", + "skills/paperclip/references/issue-documents.md", "skills/paperclip/references/issue-workspaces.md", "skills/paperclip/references/routines.md", - "skills/paperclip/references/workflows.md", - "skills/paperclip/references/api-reference.md" + "skills/paperclip/references/workflows.md" ], "rows": [ { @@ -44,9 +45,9 @@ ] }, { - "id": "skill:skills/paperclip/SKILL.md:authentication:18", + "id": "skill:skills/paperclip/SKILL.md:authentication:20", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:18", + "sourceAnchor": "skills/paperclip/SKILL.md:20", "title": "Authentication", "expectedSemantics": "Skill guidance headed “Authentication”.", "primaryDisposition": "control_plane_owned", @@ -55,13 +56,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:18" + "skill:skills/paperclip/SKILL.md:20" ] }, { - "id": "skill:skills/paperclip/SKILL.md:conversation-tasks:30", + "id": "skill:skills/paperclip/SKILL.md:conversation-tasks:32", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:30", + "sourceAnchor": "skills/paperclip/SKILL.md:32", "title": "Conversation tasks", "expectedSemantics": "Skill guidance headed “Conversation tasks”.", "primaryDisposition": "optional_agent_tool", @@ -70,13 +71,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:30" + "skill:skills/paperclip/SKILL.md:32" ] }, { - "id": "skill:skills/paperclip/SKILL.md:server-verified-external-chat-turns:47", + "id": "skill:skills/paperclip/SKILL.md:server-verified-external-chat-turns:49", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:47", + "sourceAnchor": "skills/paperclip/SKILL.md:49", "title": "Server-Verified External Chat Turns", "expectedSemantics": "Skill guidance headed “Server-Verified External Chat Turns”.", "primaryDisposition": "control_plane_owned", @@ -85,13 +86,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:47" + "skill:skills/paperclip/SKILL.md:49" ] }, { - "id": "skill:skills/paperclip/SKILL.md:the-heartbeat-procedure:87", + "id": "skill:skills/paperclip/SKILL.md:the-heartbeat-procedure:89", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:87", + "sourceAnchor": "skills/paperclip/SKILL.md:89", "title": "The Heartbeat Procedure", "expectedSemantics": "Skill guidance headed “The Heartbeat Procedure”.", "primaryDisposition": "optional_agent_tool", @@ -100,13 +101,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:87" + "skill:skills/paperclip/SKILL.md:89" ] }, { - "id": "skill:skills/paperclip/SKILL.md:generated-artifacts-and-work-products:196", + "id": "skill:skills/paperclip/SKILL.md:generated-artifacts-and-work-products:198", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:196", + "sourceAnchor": "skills/paperclip/SKILL.md:198", "title": "Generated Artifacts and Work Products", "expectedSemantics": "Skill guidance headed “Generated Artifacts and Work Products”.", "primaryDisposition": "always_agent_tool", @@ -115,13 +116,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:196" + "skill:skills/paperclip/SKILL.md:198" ] }, { - "id": "skill:skills/paperclip/SKILL.md:status-quick-guide:244", + "id": "skill:skills/paperclip/SKILL.md:status-quick-guide:246", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:244", + "sourceAnchor": "skills/paperclip/SKILL.md:246", "title": "Status Quick Guide", "expectedSemantics": "Skill guidance headed “Status Quick Guide”.", "primaryDisposition": "control_plane_owned", @@ -130,13 +131,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:244" + "skill:skills/paperclip/SKILL.md:246" ] }, { - "id": "skill:skills/paperclip/SKILL.md:monitors-and-watchers-say-only-what-you-actually-scheduled:254", + "id": "skill:skills/paperclip/SKILL.md:monitors-and-watchers-say-only-what-you-actually-scheduled:256", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:254", + "sourceAnchor": "skills/paperclip/SKILL.md:256", "title": "Monitors and Watchers (say only what you actually scheduled)", "expectedSemantics": "Skill guidance headed “Monitors and Watchers (say only what you actually scheduled)”.", "primaryDisposition": "optional_agent_tool", @@ -145,13 +146,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:254" + "skill:skills/paperclip/SKILL.md:256" ] }, { - "id": "skill:skills/paperclip/SKILL.md:delegating-review-tasks:267", + "id": "skill:skills/paperclip/SKILL.md:delegating-review-tasks:269", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:267", + "sourceAnchor": "skills/paperclip/SKILL.md:269", "title": "Delegating review tasks", "expectedSemantics": "Skill guidance headed “Delegating review tasks”.", "primaryDisposition": "always_agent_tool", @@ -160,13 +161,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:267" + "skill:skills/paperclip/SKILL.md:269" ] }, { - "id": "skill:skills/paperclip/SKILL.md:managing-a-user-s-inbox:278", + "id": "skill:skills/paperclip/SKILL.md:managing-a-user-s-inbox:280", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:278", + "sourceAnchor": "skills/paperclip/SKILL.md:280", "title": "Managing A User's Inbox", "expectedSemantics": "Skill guidance headed “Managing A User's Inbox”.", "primaryDisposition": "control_plane_owned", @@ -175,13 +176,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:278" + "skill:skills/paperclip/SKILL.md:280" ] }, { - "id": "skill:skills/paperclip/SKILL.md:issue-dependencies-blockers:286", + "id": "skill:skills/paperclip/SKILL.md:issue-dependencies-blockers:288", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:286", + "sourceAnchor": "skills/paperclip/SKILL.md:288", "title": "Issue Dependencies (Blockers)", "expectedSemantics": "Skill guidance headed “Issue Dependencies (Blockers)”.", "primaryDisposition": "control_plane_owned", @@ -190,13 +191,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:286" + "skill:skills/paperclip/SKILL.md:288" ] }, { - "id": "skill:skills/paperclip/SKILL.md:requesting-board-approval:311", + "id": "skill:skills/paperclip/SKILL.md:requesting-board-approval:313", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:311", + "sourceAnchor": "skills/paperclip/SKILL.md:313", "title": "Requesting Board Approval", "expectedSemantics": "Skill guidance headed “Requesting Board Approval”.", "primaryDisposition": "optional_agent_tool", @@ -205,13 +206,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:311" + "skill:skills/paperclip/SKILL.md:313" ] }, { - "id": "skill:skills/paperclip/SKILL.md:issue-thread-interactions:332", + "id": "skill:skills/paperclip/SKILL.md:issue-thread-interactions:334", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:332", + "sourceAnchor": "skills/paperclip/SKILL.md:334", "title": "Issue-Thread Interactions", "expectedSemantics": "Skill guidance headed “Issue-Thread Interactions”.", "primaryDisposition": "optional_agent_tool", @@ -220,13 +221,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:332" + "skill:skills/paperclip/SKILL.md:334" ] }, { - "id": "skill:skills/paperclip/SKILL.md:standalone-decisions:361", + "id": "skill:skills/paperclip/SKILL.md:standalone-decisions:363", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:361", + "sourceAnchor": "skills/paperclip/SKILL.md:363", "title": "Standalone Decisions", "expectedSemantics": "Skill guidance headed “Standalone Decisions”.", "primaryDisposition": "optional_agent_tool", @@ -235,13 +236,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:361" + "skill:skills/paperclip/SKILL.md:363" ] }, { - "id": "skill:skills/paperclip/SKILL.md:mcp-tool-approval-gates:465", + "id": "skill:skills/paperclip/SKILL.md:mcp-tool-approval-gates:467", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:465", + "sourceAnchor": "skills/paperclip/SKILL.md:467", "title": "MCP Tool Approval Gates", "expectedSemantics": "Skill guidance headed “MCP Tool Approval Gates”.", "primaryDisposition": "optional_agent_tool", @@ -250,13 +251,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:465" + "skill:skills/paperclip/SKILL.md:467" ] }, { - "id": "skill:skills/paperclip/SKILL.md:niche-workflow-pointers:507", + "id": "skill:skills/paperclip/SKILL.md:niche-workflow-pointers:509", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:507", + "sourceAnchor": "skills/paperclip/SKILL.md:509", "title": "Niche Workflow Pointers", "expectedSemantics": "Skill guidance headed “Niche Workflow Pointers”.", "primaryDisposition": "optional_agent_tool", @@ -265,13 +266,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:507" + "skill:skills/paperclip/SKILL.md:509" ] }, { - "id": "skill:skills/paperclip/SKILL.md:cases:517", + "id": "skill:skills/paperclip/SKILL.md:cases:519", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:517", + "sourceAnchor": "skills/paperclip/SKILL.md:519", "title": "Cases", "expectedSemantics": "Skill guidance headed “Cases”.", "primaryDisposition": "optional_agent_tool", @@ -280,13 +281,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:517" + "skill:skills/paperclip/SKILL.md:519" ] }, { - "id": "skill:skills/paperclip/SKILL.md:company-skills-workflow:522", + "id": "skill:skills/paperclip/SKILL.md:company-skills-workflow:524", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:522", + "sourceAnchor": "skills/paperclip/SKILL.md:524", "title": "Company Skills Workflow", "expectedSemantics": "Skill guidance headed “Company Skills Workflow”.", "primaryDisposition": "optional_agent_tool", @@ -295,13 +296,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:522" + "skill:skills/paperclip/SKILL.md:524" ] }, { - "id": "skill:skills/paperclip/SKILL.md:routines:533", + "id": "skill:skills/paperclip/SKILL.md:routines:535", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:533", + "sourceAnchor": "skills/paperclip/SKILL.md:535", "title": "Routines", "expectedSemantics": "Skill guidance headed “Routines”.", "primaryDisposition": "optional_agent_tool", @@ -310,13 +311,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:533" + "skill:skills/paperclip/SKILL.md:535" ] }, { - "id": "skill:skills/paperclip/SKILL.md:issue-workspace-runtime-controls:544", + "id": "skill:skills/paperclip/SKILL.md:issue-workspace-runtime-controls:546", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:544", + "sourceAnchor": "skills/paperclip/SKILL.md:546", "title": "Issue Workspace Runtime Controls", "expectedSemantics": "Skill guidance headed “Issue Workspace Runtime Controls”.", "primaryDisposition": "optional_agent_tool", @@ -325,13 +326,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:544" + "skill:skills/paperclip/SKILL.md:546" ] }, { - "id": "skill:skills/paperclip/SKILL.md:proposing-credentials-safely:551", + "id": "skill:skills/paperclip/SKILL.md:proposing-credentials-safely:553", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:551", + "sourceAnchor": "skills/paperclip/SKILL.md:553", "title": "Proposing Credentials Safely", "expectedSemantics": "Skill guidance headed “Proposing Credentials Safely”.", "primaryDisposition": "optional_agent_tool", @@ -340,13 +341,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:551" + "skill:skills/paperclip/SKILL.md:553" ] }, { - "id": "skill:skills/paperclip/SKILL.md:reading-granted-secrets:558", + "id": "skill:skills/paperclip/SKILL.md:reading-granted-secrets:560", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:558", + "sourceAnchor": "skills/paperclip/SKILL.md:560", "title": "Reading Granted Secrets", "expectedSemantics": "Skill guidance headed “Reading Granted Secrets”.", "primaryDisposition": "optional_agent_tool", @@ -355,13 +356,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:558" + "skill:skills/paperclip/SKILL.md:560" ] }, { - "id": "skill:skills/paperclip/SKILL.md:critical-rules:584", + "id": "skill:skills/paperclip/SKILL.md:critical-rules:586", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:584", + "sourceAnchor": "skills/paperclip/SKILL.md:586", "title": "Critical Rules", "expectedSemantics": "Skill guidance headed “Critical Rules”.", "primaryDisposition": "optional_agent_tool", @@ -370,13 +371,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:584" + "skill:skills/paperclip/SKILL.md:586" ] }, { - "id": "skill:skills/paperclip/SKILL.md:comment-style-required:605", + "id": "skill:skills/paperclip/SKILL.md:comment-style-required:607", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:605", + "sourceAnchor": "skills/paperclip/SKILL.md:607", "title": "Comment Style (Required)", "expectedSemantics": "Skill guidance headed “Comment Style (Required)”.", "primaryDisposition": "always_agent_tool", @@ -385,13 +386,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:605" + "skill:skills/paperclip/SKILL.md:607" ] }, { - "id": "skill:skills/paperclip/SKILL.md:update:637", + "id": "skill:skills/paperclip/SKILL.md:update:639", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:637", + "sourceAnchor": "skills/paperclip/SKILL.md:639", "title": "Update", "expectedSemantics": "Skill guidance headed “Update”.", "primaryDisposition": "optional_agent_tool", @@ -400,13 +401,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:637" + "skill:skills/paperclip/SKILL.md:639" ] }, { - "id": "skill:skills/paperclip/SKILL.md:planning-required-when-planning-requested:647", + "id": "skill:skills/paperclip/SKILL.md:planning-required-when-planning-requested:649", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:647", + "sourceAnchor": "skills/paperclip/SKILL.md:649", "title": "Planning (Required when planning requested)", "expectedSemantics": "Skill guidance headed “Planning (Required when planning requested)”.", "primaryDisposition": "optional_agent_tool", @@ -415,13 +416,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:647" + "skill:skills/paperclip/SKILL.md:649" ] }, { - "id": "skill:skills/paperclip/SKILL.md:key-endpoints-hot-routes:680", + "id": "skill:skills/paperclip/SKILL.md:key-endpoints-hot-routes:682", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:680", + "sourceAnchor": "skills/paperclip/SKILL.md:682", "title": "Key Endpoints (Hot Routes)", "expectedSemantics": "Skill guidance headed “Key Endpoints (Hot Routes)”.", "primaryDisposition": "optional_agent_tool", @@ -430,13 +431,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:680" + "skill:skills/paperclip/SKILL.md:682" ] }, { - "id": "skill:skills/paperclip/SKILL.md:searching-issues:709", + "id": "skill:skills/paperclip/SKILL.md:searching-issues:711", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:709", + "sourceAnchor": "skills/paperclip/SKILL.md:711", "title": "Searching Issues", "expectedSemantics": "Skill guidance headed “Searching Issues”.", "primaryDisposition": "optional_agent_tool", @@ -445,13 +446,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:709" + "skill:skills/paperclip/SKILL.md:711" ] }, { - "id": "skill:skills/paperclip/SKILL.md:full-reference:719", + "id": "skill:skills/paperclip/SKILL.md:full-reference:721", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:719", + "sourceAnchor": "skills/paperclip/SKILL.md:721", "title": "Full Reference", "expectedSemantics": "Skill guidance headed “Full Reference”.", "primaryDisposition": "optional_agent_tool", @@ -460,13 +461,13 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:719" + "skill:skills/paperclip/SKILL.md:721" ] }, { - "id": "skill:skills/paperclip/SKILL.md:conversational-confirmation-answers:723", + "id": "skill:skills/paperclip/SKILL.md:conversational-confirmation-answers:725", "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/SKILL.md:723", + "sourceAnchor": "skills/paperclip/SKILL.md:725", "title": "Conversational confirmation answers", "expectedSemantics": "Skill guidance headed “Conversational confirmation answers”.", "primaryDisposition": "always_agent_tool", @@ -475,847 +476,7 @@ "control_plane_invariant" ], "evidenceIds": [ - "skill:skills/paperclip/SKILL.md:723" - ] - }, - { - "id": "skill:skills/paperclip/references/artifacts.md:generated-artifacts-and-work-products:1", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/artifacts.md:1", - "title": "Generated Artifacts and Work Products", - "expectedSemantics": "Skill guidance headed “Generated Artifacts and Work Products”.", - "primaryDisposition": "always_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/artifacts.md:1" - ] - }, - { - "id": "skill:skills/paperclip/references/artifacts.md:workspace-only-file-references:15", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/artifacts.md:15", - "title": "Workspace-Only File References", - "expectedSemantics": "Skill guidance headed “Workspace-Only File References”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/artifacts.md:15" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:cases:1", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:1", - "title": "Cases", - "expectedSemantics": "Skill guidance headed “Cases”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:1" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:core-model:12", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:12", - "title": "Core Model", - "expectedSemantics": "Skill guidance headed “Core Model”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:12" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:upsert-semantics:29", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:29", - "title": "Upsert Semantics", - "expectedSemantics": "Skill guidance headed “Upsert Semantics”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:29" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:read-and-search:67", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:67", - "title": "Read And Search", - "expectedSemantics": "Skill guidance headed “Read And Search”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:67" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:documents:90", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:90", - "title": "Documents", - "expectedSemantics": "Skill guidance headed “Documents”.", - "primaryDisposition": "always_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:90" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:fields:118", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:118", - "title": "Fields", - "expectedSemantics": "Skill guidance headed “Fields”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:118" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:issue-links:151", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:151", - "title": "Issue Links", - "expectedSemantics": "Skill guidance headed “Issue Links”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:151" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:child-cases:176", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:176", - "title": "Child Cases", - "expectedSemantics": "Skill guidance headed “Child Cases”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:176" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:attachments:195", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:195", - "title": "Attachments", - "expectedSemantics": "Skill guidance headed “Attachments”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:195" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:lifecycle:208", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:208", - "title": "Lifecycle", - "expectedSemantics": "Skill guidance headed “Lifecycle”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:208" - ] - }, - { - "id": "skill:skills/paperclip/references/cases.md:worked-blog-post-example:222", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/cases.md:222", - "title": "Worked Blog Post Example", - "expectedSemantics": "Skill guidance headed “Worked Blog Post Example”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/cases.md:222" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:company-skills-workflow:1", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:1", - "title": "Company Skills Workflow", - "expectedSemantics": "Skill guidance headed “Company Skills Workflow”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:1" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:what-exists:5", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:5", - "title": "What Exists", - "expectedSemantics": "Skill guidance headed “What Exists”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:5" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:permission-model:22", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:22", - "title": "Permission Model", - "expectedSemantics": "Skill guidance headed “Permission Model”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:22" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:core-endpoints:29", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:29", - "title": "Core Endpoints", - "expectedSemantics": "Skill guidance headed “Core Endpoints”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:29" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:install-a-skill-into-the-company:65", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:65", - "title": "Install A Skill Into The Company", - "expectedSemantics": "Skill guidance headed “Install A Skill Into The Company”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:65" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:app-shipped-catalog:75", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:75", - "title": "App-shipped catalog", - "expectedSemantics": "Skill guidance headed “App-shipped catalog”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:75" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:external-source-import:101", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:101", - "title": "External source import", - "expectedSemantics": "Skill guidance headed “External source import”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:101" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:source-types-in-order-of-preference:105", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:105", - "title": "Source types (in order of preference)", - "expectedSemantics": "Skill guidance headed “Source types (in order of preference)”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:105" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:example-skills-sh-import-preferred:116", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:116", - "title": "Example: skills.sh import (preferred)", - "expectedSemantics": "Skill guidance headed “Example: skills.sh import (preferred)”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:116" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:example-github-import:138", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:138", - "title": "Example: GitHub import", - "expectedSemantics": "Skill guidance headed “Example: GitHub import”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:138" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:inspect-what-was-installed:164", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:164", - "title": "Inspect What Was Installed", - "expectedSemantics": "Skill guidance headed “Inspect What Was Installed”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:164" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:assign-skills-to-an-existing-agent:181", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:181", - "title": "Assign Skills To An Existing Agent", - "expectedSemantics": "Skill guidance headed “Assign Skills To An Existing Agent”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:181" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:include-skills-during-hire-or-create:216", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:216", - "title": "Include Skills During Hire Or Create", - "expectedSemantics": "Skill guidance headed “Include Skills During Hire Or Create”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:216" - ] - }, - { - "id": "skill:skills/paperclip/references/company-skills.md:notes:256", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/company-skills.md:256", - "title": "Notes", - "expectedSemantics": "Skill guidance headed “Notes”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/company-skills.md:256" - ] - }, - { - "id": "skill:skills/paperclip/references/issue-workspaces.md:issue-workspace-runtime-controls:1", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:1", - "title": "Issue Workspace Runtime Controls", - "expectedSemantics": "Skill guidance headed “Issue Workspace Runtime Controls”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/issue-workspaces.md:1" - ] - }, - { - "id": "skill:skills/paperclip/references/issue-workspaces.md:discover-the-workspace:5", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:5", - "title": "Discover the Workspace", - "expectedSemantics": "Skill guidance headed “Discover the Workspace”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/issue-workspaces.md:5" - ] - }, - { - "id": "skill:skills/paperclip/references/issue-workspaces.md:control-services:23", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:23", - "title": "Control Services", - "expectedSemantics": "Skill guidance headed “Control Services”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/issue-workspaces.md:23" - ] - }, - { - "id": "skill:skills/paperclip/references/issue-workspaces.md:start-all-configured-services-waits-for-configured-readiness-checks:28", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:28", - "title": "Start all configured services; waits for configured readiness checks.", - "expectedSemantics": "Skill guidance headed “Start all configured services; waits for configured readiness checks.”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/issue-workspaces.md:28" - ] - }, - { - "id": "skill:skills/paperclip/references/issue-workspaces.md:restart-all-configured-services:36", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:36", - "title": "Restart all configured services.", - "expectedSemantics": "Skill guidance headed “Restart all configured services.”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/issue-workspaces.md:36" - ] - }, - { - "id": "skill:skills/paperclip/references/issue-workspaces.md:stop-all-running-services:44", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:44", - "title": "Stop all running services.", - "expectedSemantics": "Skill guidance headed “Stop all running services.”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/issue-workspaces.md:44" - ] - }, - { - "id": "skill:skills/paperclip/references/issue-workspaces.md:read-the-url:63", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:63", - "title": "Read the URL", - "expectedSemantics": "Skill guidance headed “Read the URL”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/issue-workspaces.md:63" - ] - }, - { - "id": "skill:skills/paperclip/references/issue-workspaces.md:mcp-tools:72", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:72", - "title": "MCP Tools", - "expectedSemantics": "Skill guidance headed “MCP Tools”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/issue-workspaces.md:72" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:paperclip-routines:1", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:1", - "title": "Paperclip Routines", - "expectedSemantics": "Skill guidance headed “Paperclip Routines”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:1" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:lifecycle:16", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:16", - "title": "Lifecycle", - "expectedSemantics": "Skill guidance headed “Lifecycle”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:16" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:creating-a-routine:27", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:27", - "title": "Creating a Routine", - "expectedSemantics": "Skill guidance headed “Creating a Routine”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:27" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:concurrency-policies:64", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:64", - "title": "Concurrency Policies", - "expectedSemantics": "Skill guidance headed “Concurrency Policies”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:64" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:catch-up-policies:76", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:76", - "title": "Catch-Up Policies", - "expectedSemantics": "Skill guidance headed “Catch-Up Policies”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:76" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:activity-gated-scheduled-runs:87", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:87", - "title": "Activity-Gated Scheduled Runs", - "expectedSemantics": "Skill guidance headed “Activity-Gated Scheduled Runs”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:87" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:example-skip-quiet-nights:107", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:107", - "title": "Example: skip quiet nights", - "expectedSemantics": "Skill guidance headed “Example: skip quiet nights”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:107" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:adding-triggers:126", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:126", - "title": "Adding Triggers", - "expectedSemantics": "Skill guidance headed “Adding Triggers”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:126" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:schedule-cron:136", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:136", - "title": "Schedule (cron)", - "expectedSemantics": "Skill guidance headed “Schedule (cron)”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:136" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:webhook:150", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:150", - "title": "Webhook", - "expectedSemantics": "Skill guidance headed “Webhook”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:150" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:api-manual-only:167", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:167", - "title": "API (manual only)", - "expectedSemantics": "Skill guidance headed “API (manual only)”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:167" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:updating-and-deleting-triggers:179", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:179", - "title": "Updating and Deleting Triggers", - "expectedSemantics": "Skill guidance headed “Updating and Deleting Triggers”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:179" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:manual-run:196", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:196", - "title": "Manual Run", - "expectedSemantics": "Skill guidance headed “Manual Run”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:196" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:updating-a-routine:212", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:212", - "title": "Updating a Routine", - "expectedSemantics": "Skill guidance headed “Updating a Routine”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:212" - ] - }, - { - "id": "skill:skills/paperclip/references/routines.md:reading-routines-and-runs:223", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/routines.md:223", - "title": "Reading Routines and Runs", - "expectedSemantics": "Skill guidance headed “Reading Routines and Runs”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/routines.md:223" - ] - }, - { - "id": "skill:skills/paperclip/references/workflows.md:paperclip-workflow-playbooks:1", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/workflows.md:1", - "title": "Paperclip Workflow Playbooks", - "expectedSemantics": "Skill guidance headed “Paperclip Workflow Playbooks”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/workflows.md:1" - ] - }, - { - "id": "skill:skills/paperclip/references/workflows.md:project-setup-ceo-manager:7", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/workflows.md:7", - "title": "Project Setup (CEO/Manager)", - "expectedSemantics": "Skill guidance headed “Project Setup (CEO/Manager)”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/workflows.md:7" - ] - }, - { - "id": "skill:skills/paperclip/references/workflows.md:openclaw-invite-ceo:22", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/workflows.md:22", - "title": "OpenClaw Invite (CEO)", - "expectedSemantics": "Skill guidance headed “OpenClaw Invite (CEO)”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/workflows.md:22" - ] - }, - { - "id": "skill:skills/paperclip/references/workflows.md:setting-agent-instructions-path:50", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/workflows.md:50", - "title": "Setting Agent Instructions Path", - "expectedSemantics": "Skill guidance headed “Setting Agent Instructions Path”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/workflows.md:50" - ] - }, - { - "id": "skill:skills/paperclip/references/workflows.md:company-import-export:79", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/workflows.md:79", - "title": "Company Import / Export", - "expectedSemantics": "Skill guidance headed “Company Import / Export”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/workflows.md:79" - ] - }, - { - "id": "skill:skills/paperclip/references/workflows.md:self-test-playbook-app-level:106", - "sourceKind": "skill_heading", - "sourceAnchor": "skills/paperclip/references/workflows.md:106", - "title": "Self-Test Playbook (App-Level)", - "expectedSemantics": "Skill guidance headed “Self-Test Playbook (App-Level)”.", - "primaryDisposition": "optional_agent_tool", - "requiredGrants": [], - "assertionClasses": [ - "control_plane_invariant" - ], - "evidenceIds": [ - "skill:skills/paperclip/references/workflows.md:106" + "skill:skills/paperclip/SKILL.md:725" ] }, { @@ -2367,6 +1528,891 @@ "evidenceIds": [ "skill:skills/paperclip/references/api-reference.md:1661" ] + }, + { + "id": "skill:skills/paperclip/references/artifacts.md:generated-artifacts-and-work-products:1", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/artifacts.md:1", + "title": "Generated Artifacts and Work Products", + "expectedSemantics": "Skill guidance headed “Generated Artifacts and Work Products”.", + "primaryDisposition": "always_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/artifacts.md:1" + ] + }, + { + "id": "skill:skills/paperclip/references/artifacts.md:workspace-only-file-references:15", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/artifacts.md:15", + "title": "Workspace-Only File References", + "expectedSemantics": "Skill guidance headed “Workspace-Only File References”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/artifacts.md:15" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:cases:1", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:1", + "title": "Cases", + "expectedSemantics": "Skill guidance headed “Cases”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:1" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:core-model:12", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:12", + "title": "Core Model", + "expectedSemantics": "Skill guidance headed “Core Model”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:12" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:upsert-semantics:29", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:29", + "title": "Upsert Semantics", + "expectedSemantics": "Skill guidance headed “Upsert Semantics”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:29" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:read-and-search:67", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:67", + "title": "Read And Search", + "expectedSemantics": "Skill guidance headed “Read And Search”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:67" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:documents:90", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:90", + "title": "Documents", + "expectedSemantics": "Skill guidance headed “Documents”.", + "primaryDisposition": "always_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:90" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:fields:118", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:118", + "title": "Fields", + "expectedSemantics": "Skill guidance headed “Fields”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:118" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:issue-links:151", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:151", + "title": "Issue Links", + "expectedSemantics": "Skill guidance headed “Issue Links”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:151" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:child-cases:176", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:176", + "title": "Child Cases", + "expectedSemantics": "Skill guidance headed “Child Cases”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:176" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:attachments:195", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:195", + "title": "Attachments", + "expectedSemantics": "Skill guidance headed “Attachments”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:195" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:lifecycle:208", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:208", + "title": "Lifecycle", + "expectedSemantics": "Skill guidance headed “Lifecycle”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:208" + ] + }, + { + "id": "skill:skills/paperclip/references/cases.md:worked-blog-post-example:222", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/cases.md:222", + "title": "Worked Blog Post Example", + "expectedSemantics": "Skill guidance headed “Worked Blog Post Example”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/cases.md:222" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:company-skills-workflow:1", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:1", + "title": "Company Skills Workflow", + "expectedSemantics": "Skill guidance headed “Company Skills Workflow”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:1" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:what-exists:5", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:5", + "title": "What Exists", + "expectedSemantics": "Skill guidance headed “What Exists”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:5" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:permission-model:22", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:22", + "title": "Permission Model", + "expectedSemantics": "Skill guidance headed “Permission Model”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:22" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:core-endpoints:29", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:29", + "title": "Core Endpoints", + "expectedSemantics": "Skill guidance headed “Core Endpoints”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:29" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:install-a-skill-into-the-company:65", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:65", + "title": "Install A Skill Into The Company", + "expectedSemantics": "Skill guidance headed “Install A Skill Into The Company”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:65" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:app-shipped-catalog:75", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:75", + "title": "App-shipped catalog", + "expectedSemantics": "Skill guidance headed “App-shipped catalog”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:75" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:external-source-import:101", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:101", + "title": "External source import", + "expectedSemantics": "Skill guidance headed “External source import”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:101" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:source-types-in-order-of-preference:105", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:105", + "title": "Source types (in order of preference)", + "expectedSemantics": "Skill guidance headed “Source types (in order of preference)”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:105" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:example-skills-sh-import-preferred:116", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:116", + "title": "Example: skills.sh import (preferred)", + "expectedSemantics": "Skill guidance headed “Example: skills.sh import (preferred)”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:116" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:example-github-import:138", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:138", + "title": "Example: GitHub import", + "expectedSemantics": "Skill guidance headed “Example: GitHub import”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:138" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:inspect-what-was-installed:164", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:164", + "title": "Inspect What Was Installed", + "expectedSemantics": "Skill guidance headed “Inspect What Was Installed”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:164" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:assign-skills-to-an-existing-agent:181", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:181", + "title": "Assign Skills To An Existing Agent", + "expectedSemantics": "Skill guidance headed “Assign Skills To An Existing Agent”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:181" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:include-skills-during-hire-or-create:216", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:216", + "title": "Include Skills During Hire Or Create", + "expectedSemantics": "Skill guidance headed “Include Skills During Hire Or Create”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:216" + ] + }, + { + "id": "skill:skills/paperclip/references/company-skills.md:notes:256", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/company-skills.md:256", + "title": "Notes", + "expectedSemantics": "Skill guidance headed “Notes”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/company-skills.md:256" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-documents.md:issue-documents-through-the-api:1", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-documents.md:1", + "title": "Issue documents through the API", + "expectedSemantics": "Skill guidance headed “Issue documents through the API”.", + "primaryDisposition": "always_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-documents.md:1" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-documents.md:create-a-document:8", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-documents.md:8", + "title": "Create a document", + "expectedSemantics": "Skill guidance headed “Create a document”.", + "primaryDisposition": "always_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-documents.md:8" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-documents.md:update-or-resolve-an-unclear-write:49", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-documents.md:49", + "title": "Update or resolve an unclear write", + "expectedSemantics": "Skill guidance headed “Update or resolve an unclear write”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-documents.md:49" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-workspaces.md:issue-workspace-runtime-controls:1", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:1", + "title": "Issue Workspace Runtime Controls", + "expectedSemantics": "Skill guidance headed “Issue Workspace Runtime Controls”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-workspaces.md:1" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-workspaces.md:discover-the-workspace:5", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:5", + "title": "Discover the Workspace", + "expectedSemantics": "Skill guidance headed “Discover the Workspace”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-workspaces.md:5" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-workspaces.md:control-services:23", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:23", + "title": "Control Services", + "expectedSemantics": "Skill guidance headed “Control Services”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-workspaces.md:23" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-workspaces.md:start-all-configured-services-waits-for-configured-readiness-checks:28", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:28", + "title": "Start all configured services; waits for configured readiness checks.", + "expectedSemantics": "Skill guidance headed “Start all configured services; waits for configured readiness checks.”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-workspaces.md:28" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-workspaces.md:restart-all-configured-services:36", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:36", + "title": "Restart all configured services.", + "expectedSemantics": "Skill guidance headed “Restart all configured services.”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-workspaces.md:36" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-workspaces.md:stop-all-running-services:44", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:44", + "title": "Stop all running services.", + "expectedSemantics": "Skill guidance headed “Stop all running services.”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-workspaces.md:44" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-workspaces.md:read-the-url:63", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:63", + "title": "Read the URL", + "expectedSemantics": "Skill guidance headed “Read the URL”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-workspaces.md:63" + ] + }, + { + "id": "skill:skills/paperclip/references/issue-workspaces.md:mcp-tools:72", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/issue-workspaces.md:72", + "title": "MCP Tools", + "expectedSemantics": "Skill guidance headed “MCP Tools”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/issue-workspaces.md:72" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:paperclip-routines:1", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:1", + "title": "Paperclip Routines", + "expectedSemantics": "Skill guidance headed “Paperclip Routines”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:1" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:lifecycle:16", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:16", + "title": "Lifecycle", + "expectedSemantics": "Skill guidance headed “Lifecycle”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:16" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:creating-a-routine:27", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:27", + "title": "Creating a Routine", + "expectedSemantics": "Skill guidance headed “Creating a Routine”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:27" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:concurrency-policies:64", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:64", + "title": "Concurrency Policies", + "expectedSemantics": "Skill guidance headed “Concurrency Policies”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:64" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:catch-up-policies:76", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:76", + "title": "Catch-Up Policies", + "expectedSemantics": "Skill guidance headed “Catch-Up Policies”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:76" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:activity-gated-scheduled-runs:87", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:87", + "title": "Activity-Gated Scheduled Runs", + "expectedSemantics": "Skill guidance headed “Activity-Gated Scheduled Runs”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:87" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:example-skip-quiet-nights:107", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:107", + "title": "Example: skip quiet nights", + "expectedSemantics": "Skill guidance headed “Example: skip quiet nights”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:107" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:adding-triggers:126", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:126", + "title": "Adding Triggers", + "expectedSemantics": "Skill guidance headed “Adding Triggers”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:126" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:schedule-cron:136", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:136", + "title": "Schedule (cron)", + "expectedSemantics": "Skill guidance headed “Schedule (cron)”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:136" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:webhook:150", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:150", + "title": "Webhook", + "expectedSemantics": "Skill guidance headed “Webhook”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:150" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:api-manual-only:167", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:167", + "title": "API (manual only)", + "expectedSemantics": "Skill guidance headed “API (manual only)”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:167" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:updating-and-deleting-triggers:179", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:179", + "title": "Updating and Deleting Triggers", + "expectedSemantics": "Skill guidance headed “Updating and Deleting Triggers”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:179" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:manual-run:196", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:196", + "title": "Manual Run", + "expectedSemantics": "Skill guidance headed “Manual Run”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:196" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:updating-a-routine:212", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:212", + "title": "Updating a Routine", + "expectedSemantics": "Skill guidance headed “Updating a Routine”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:212" + ] + }, + { + "id": "skill:skills/paperclip/references/routines.md:reading-routines-and-runs:223", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/routines.md:223", + "title": "Reading Routines and Runs", + "expectedSemantics": "Skill guidance headed “Reading Routines and Runs”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/routines.md:223" + ] + }, + { + "id": "skill:skills/paperclip/references/workflows.md:paperclip-workflow-playbooks:1", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/workflows.md:1", + "title": "Paperclip Workflow Playbooks", + "expectedSemantics": "Skill guidance headed “Paperclip Workflow Playbooks”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/workflows.md:1" + ] + }, + { + "id": "skill:skills/paperclip/references/workflows.md:project-setup-ceo-manager:7", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/workflows.md:7", + "title": "Project Setup (CEO/Manager)", + "expectedSemantics": "Skill guidance headed “Project Setup (CEO/Manager)”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/workflows.md:7" + ] + }, + { + "id": "skill:skills/paperclip/references/workflows.md:openclaw-invite-ceo:22", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/workflows.md:22", + "title": "OpenClaw Invite (CEO)", + "expectedSemantics": "Skill guidance headed “OpenClaw Invite (CEO)”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/workflows.md:22" + ] + }, + { + "id": "skill:skills/paperclip/references/workflows.md:setting-agent-instructions-path:50", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/workflows.md:50", + "title": "Setting Agent Instructions Path", + "expectedSemantics": "Skill guidance headed “Setting Agent Instructions Path”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/workflows.md:50" + ] + }, + { + "id": "skill:skills/paperclip/references/workflows.md:company-import-export:79", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/workflows.md:79", + "title": "Company Import / Export", + "expectedSemantics": "Skill guidance headed “Company Import / Export”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/workflows.md:79" + ] + }, + { + "id": "skill:skills/paperclip/references/workflows.md:self-test-playbook-app-level:106", + "sourceKind": "skill_heading", + "sourceAnchor": "skills/paperclip/references/workflows.md:106", + "title": "Self-Test Playbook (App-Level)", + "expectedSemantics": "Skill guidance headed “Self-Test Playbook (App-Level)”.", + "primaryDisposition": "optional_agent_tool", + "requiredGrants": [], + "assertionClasses": [ + "control_plane_invariant" + ], + "evidenceIds": [ + "skill:skills/paperclip/references/workflows.md:106" + ] } ] } diff --git a/packages/paperclip-runner/spec/capability/source-contract.json b/packages/paperclip-runner/spec/capability/source-contract.json index ebc924bad3..5ff5ae8ce5 100644 --- a/packages/paperclip-runner/spec/capability/source-contract.json +++ b/packages/paperclip-runner/spec/capability/source-contract.json @@ -7,6 +7,7 @@ "skills/paperclip/references/artifacts.md", "skills/paperclip/references/cases.md", "skills/paperclip/references/company-skills.md", + "skills/paperclip/references/issue-documents.md", "skills/paperclip/references/issue-workspaces.md", "skills/paperclip/references/routines.md", "skills/paperclip/references/workflows.md" diff --git a/packages/paperclip-runner/src/generated/capability-contract.ts b/packages/paperclip-runner/src/generated/capability-contract.ts index 0d145a06c4..70aaf0ac06 100644 --- a/packages/paperclip-runner/src/generated/capability-contract.ts +++ b/packages/paperclip-runner/src/generated/capability-contract.ts @@ -6,9 +6,9 @@ export type CapabilityPrimaryDisposition = | "optional_agent_tool"; export const capabilityInventoryCounts = { - "skillReferenceCapabilities": 157, + "skillReferenceCapabilities": 160, "evalCases": 106, - "normativeRows": 263, + "normativeRows": 266, "legacyMcpAliases": 42 } as const; diff --git a/packages/paperclip-runner/test/capability-contract.test.mjs b/packages/paperclip-runner/test/capability-contract.test.mjs index 234259a91f..ac0d7e0d13 100644 --- a/packages/paperclip-runner/test/capability-contract.test.mjs +++ b/packages/paperclip-runner/test/capability-contract.test.mjs @@ -18,7 +18,8 @@ test("generated Capability inventory has full source coverage", async () => { readRows("eval-traceability.yaml"), ]); - assert.equal(capabilities.length, 158); + assert.equal(capabilities.length, 161); + assert.equal(capabilities.filter(row => row.sourceAnchor.startsWith("skills/paperclip/references/issue-documents.md#")).length, 3); assert.equal(tools.length, 42); assert.equal(evals.length, 106); assert.equal(new Set(evals.map((row) => row.group)).size, 16); @@ -51,3 +52,15 @@ test("conversational answer guidance has the same agent-operation classification for (const row of [...generated, ...spec]) assert.equal(row.primaryDisposition, "always_agent_tool"); for (const row of generated) assert.equal(row.semanticOperation, "call_api"); }); + +test("both capability inventories cover every declared skill source", async () => { + const contract = JSON.parse(await readFile(resolve(import.meta.dirname, "../spec/capability/source-contract.json"), "utf8")); + const generated = await readRows("capabilities.yaml"); + const normative = decodeInventory(await readFile(resolve(import.meta.dirname, "../spec/capability/capabilities.yaml"), "utf8")); + assert.deepEqual(normative.generatedFrom, contract.skillSources); + for (const source of contract.skillSources) { + assert.ok(generated.some(row => row.sourceAnchor.startsWith(`${source}#`)), `Missing generated source ${source}`); + assert.ok(normative.rows.some(row => row.sourceAnchor.startsWith(`${source}:`)), `Missing normative source ${source}`); + } + assert.equal(normative.rows.filter(row => row.sourceAnchor.startsWith("skills/paperclip/references/issue-documents.md:")).length, 3); +}); diff --git a/server/src/__tests__/agent-skills-routes.test.ts b/server/src/__tests__/agent-skills-routes.test.ts index cbf6a260f0..f64229e3d4 100644 --- a/server/src/__tests__/agent-skills-routes.test.ts +++ b/server/src/__tests__/agent-skills-routes.test.ts @@ -1184,7 +1184,7 @@ describe.sequential("agent skill routes", () => { expect(returnedAgent.adapterConfig).not.toHaveProperty("promptTemplate"); }); - it("materializes the bundled default instruction set for non-CEO agents with no prompt template", async () => { + it("materializes minimal default instructions for non-CEO agents with no prompt template", async () => { const res = await requestApp(await createApp(), (baseUrl) => request(baseUrl) .post("/api/companies/company-1/agents") .send({ @@ -1204,38 +1204,10 @@ describe.sequential("agent skill routes", () => { adapterType: "claude_local", }), expect.objectContaining({ - "AGENTS.md": expect.stringMatching(/Start actionable work in the same heartbeat\.[\s\S]*Keep the work moving until it is done\./), + "AGENTS.md": "You are an agent in a Paperclip company.\n", }), { entryFile: "AGENTS.md", replaceExisting: false }, ); - expect(mockAgentInstructionsService.materializeManagedBundle).toHaveBeenCalledWith( - expect.any(Object), - expect.objectContaining({ - "AGENTS.md": expect.stringContaining('kind: "request_confirmation"'), - }), - expect.any(Object), - ); - expect(mockAgentInstructionsService.materializeManagedBundle).toHaveBeenCalledWith( - expect.any(Object), - expect.objectContaining({ - "AGENTS.md": expect.stringContaining("confirmation:{issueId}:plan:{revisionId}"), - }), - expect.any(Object), - ); - expect(mockAgentInstructionsService.materializeManagedBundle).toHaveBeenCalledWith( - expect.any(Object), - expect.objectContaining({ - "AGENTS.md": expect.stringMatching(/PUT \/issues\/\{id\}\/documents\/plan[\s\S]*Re-`GET \/documents\/plan`, assert it returns `200`[\s\S]*latestRevisionId[\s\S]*target=\{ type: 'issue_document', key: 'plan', revisionId: latestRevisionId \}[\s\S]*Never present a plan only in a thread comment or through `ask_user_questions`/), - }), - expect.any(Object), - ); - expect(mockAgentInstructionsService.materializeManagedBundle).toHaveBeenCalledWith( - expect.any(Object), - expect.objectContaining({ - "AGENTS.md": expect.stringContaining("skills/paperclip/scripts/paperclip-upload-artifact.sh"), - }), - expect.any(Object), - ); }); }); diff --git a/server/src/__tests__/codex-local-execute.test.ts b/server/src/__tests__/codex-local-execute.test.ts index 983fa6e87e..29a3a2b225 100644 --- a/server/src/__tests__/codex-local-execute.test.ts +++ b/server/src/__tests__/codex-local-execute.test.ts @@ -1443,9 +1443,8 @@ process.exit(1); expect(invocationPrompt).toContain("baseRevisionId set to that latestRevisionId"); expect(capture.prompt).not.toContain("Execution contract:"); expect(capture.prompt).not.toContain("Use child issues"); - } else { - expect(capture.prompt).toContain("Execution contract:"); } + expect(capture.prompt).not.toContain("Execution contract:"); expect(capture.prompt).not.toContain("Follow the paperclip heartbeat."); if (resumedSession) { expect(capture.prompt).not.toContain("You are managed instructions."); diff --git a/server/src/__tests__/workspace-runtime.test.ts b/server/src/__tests__/workspace-runtime.test.ts index 187135d195..66b9731dfa 100644 --- a/server/src/__tests__/workspace-runtime.test.ts +++ b/server/src/__tests__/workspace-runtime.test.ts @@ -6743,35 +6743,44 @@ describeEmbeddedPostgres("workspace runtime service control persistence", () => executionWorkspaceId, serviceName: "web", status: "provisioning", + healthStatus: "unknown", providerRef: null, }); expect(existsSync(markerPath)).toBe(false); await waitForMarker(markerPath); - const startingRow = await waitForPersistedStatus("starting"); - expect(startingRow).toMatchObject({ + // Readiness begins before the PID-bearing transaction commits. Keep the + // listener independent of DB reads and verify the completed records. + const services = await startPromise; + expect(services).toHaveLength(1); + expect(services[0]).toMatchObject({ + id: provisioningRow.id, companyId, projectId, projectWorkspaceId, executionWorkspaceId, issueId, serviceName: "web", - status: "starting", - healthStatus: "unknown", - }); - expect(startingRow.providerRef).toMatch(/^\d+$/); - expect(startingRow.port).toEqual(expect.any(Number)); - - const services = await startPromise; - expect(services).toHaveLength(1); - expect(services[0]).toMatchObject({ - id: startingRow.id, status: "running", healthStatus: "healthy", }); + expect(services[0]!.providerRef).toMatch(/^\d+$/); + expect(services[0]!.port).toEqual(expect.any(Number)); const runningRow = await waitForPersistedStatus("running"); - expect(runningRow.id).toBe(startingRow.id); + expect(runningRow).toMatchObject({ + id: provisioningRow.id, + companyId, + projectId, + projectWorkspaceId, + executionWorkspaceId, + issueId, + serviceName: "web", + status: "running", + healthStatus: "healthy", + providerRef: services[0]!.providerRef, + port: services[0]!.port, + }); await expect(fetch(services[0]!.url!)).resolves.toMatchObject({ ok: true }); const runtimeProvisionOperations = await db .select() diff --git a/server/src/onboarding-assets/default/AGENTS.md b/server/src/onboarding-assets/default/AGENTS.md index 1472483914..8c6df421c7 100644 --- a/server/src/onboarding-assets/default/AGENTS.md +++ b/server/src/onboarding-assets/default/AGENTS.md @@ -1,24 +1 @@ -You are an agent at Paperclip company. - -## Execution Contract - -- Start actionable work in the same heartbeat. Do not stop at a plan unless the issue explicitly asks for planning. -- Keep the work moving until it is done. If you need QA to review it, ask them. If you need your boss to review it, ask them. -- Leave durable progress in task comments, documents, or work products, then update the issue to a clear final disposition before you exit. -- When your work produces a user-inspectable deliverable file, follow the Paperclip skill's "Generated Artifacts and Work Products" workflow before final disposition. Use `skills/paperclip/scripts/paperclip-upload-artifact.sh` when working in this repo, create/update an artifact work product when the file is the deliverable, and link the uploaded attachment in the final comment. Do not rely on local filesystem paths as the only access path. If an important file intentionally remains workspace-only, create/update a work product with `metadata.resourceRef.kind: "workspace_file"` and a workspace-relative path, then name that work product and path in the final comment. Treat browse/search as a fallback for recovering workspace files, not the preferred deliverable path. -- When your work produces or updates an operator-facing engineering output, create/update the matching work product: `pull_request` for opened PRs, `preview_url` for published previews, `runtime_service` for managed preview/dev services, `commit` for notable pushed commits, and `branch` when the branch itself is the handoff. A comment is not a substitute for the work product access path. -- Comments, documents, screenshots, work products, and `Remaining` bullets are evidence, not valid liveness paths by themselves. -- Final disposition checklist: mark `done` when complete and verified; use `in_review` only with a real reviewer, approval, interaction, or monitor path; use `blocked` only with first-class blockers or a named unblock owner/action; create delegated follow-up issues with blockers when another agent owns the next step; keep `in_progress` only when a live continuation path exists. -- Use child issues for parallel or long delegated work instead of polling agents, sessions, or processes. -- Create child issues directly when you know what needs to be done. If the board/user needs to choose suggested tasks, answer structured questions, or confirm a proposal first, create an issue-thread interaction on the current issue with `POST /api/issues/{issueId}/interactions` using `kind: "suggest_tasks"`, `kind: "ask_user_questions"`, or `kind: "request_confirmation"`. -- Use `request_confirmation` instead of asking for yes/no decisions in markdown. Before presenting a plan for review, you MUST complete this publish contract: - 1. `PUT /issues/{id}/documents/plan` with `{ format: 'markdown', body, changeSummary }`. - 2. Re-`GET /documents/plan`, assert it returns `200`, and capture its `latestRevisionId`. - 3. Only then create `request_confirmation` with `target={ type: 'issue_document', key: 'plan', revisionId: latestRevisionId }` and `idempotencyKey=confirmation:{issueId}:plan:{revisionId}`. - 4. Wait for acceptance before creating implementation subtasks. - Never present a plan only in a thread comment or through `ask_user_questions`; comments are supporting context and questions are for gathering input, not plan review. -- `ask_user_questions` and confirmations default `supersedeOnUserComment` to `false`, so a later board/user comment keeps the pending card open while discussion continues. Set it to `true` when a new comment should replace the pending request. If you wake up from a superseding comment, revise the artifact, question set, or proposal and create a fresh interaction if input is still needed. -- For human input, save a pending question/confirmation interaction and set `in_review`; prose alone does not create a waiting path. Use `blockedByIssueIds` for issue dependencies. An agent may set an `unblockDescriptor` only for itself (`owner: { "agentId": "" }` plus `action`), not for the board/user or another agent. -- Respect budget, pause/cancel, approval gates, and company boundaries. - -Do not let work sit here. You must always update your task with a comment. +You are an agent in a Paperclip company. diff --git a/server/src/routes/agents.ts b/server/src/routes/agents.ts index 094a02a656..11fdc1ab3c 100644 --- a/server/src/routes/agents.ts +++ b/server/src/routes/agents.ts @@ -2759,8 +2759,8 @@ export function agentRoutes( // first-task/chief-of-staff/AGENTS.md, placeholders filled) over the agent's // entry file instead of the generic default. Honored only for board-authored // requests — the onboarding wizard runs as the board — so a client marker - // alone cannot swap another actor's instructions. The generic execution - // contract (default/AGENTS.md) is still appended on every run, unchanged. + // alone cannot swap another actor's instructions. Runtime coordination is + // supplied separately by the harness. async function resolveOnboardingFirstAgentBundle(params: { onboardingFirstAgent: unknown; actorType: string; diff --git a/server/src/services/onboarding-first-task-assets.ts b/server/src/services/onboarding-first-task-assets.ts index 2030faa519..67e6aff2ae 100644 --- a/server/src/services/onboarding-first-task-assets.ts +++ b/server/src/services/onboarding-first-task-assets.ts @@ -139,8 +139,8 @@ export async function renderChiefOfStaffPersona( } // The instruction bundle for the onboarding first agent: the chief-of-staff -// persona as the entry AGENTS.md. The generic execution contract -// (default/AGENTS.md) is still appended on every run by the runner, unchanged. +// persona as the entry AGENTS.md. Runtime coordination is supplied separately +// by the harness. export async function buildOnboardingFirstAgentInstructionsBundle( placeholders: OnboardingFirstTaskPlaceholders, ): Promise<{ files: Record; entryFile: string }> { diff --git a/skills/paperclip/SKILL.md b/skills/paperclip/SKILL.md index 695f91abf4..69a1398837 100644 --- a/skills/paperclip/SKILL.md +++ b/skills/paperclip/SKILL.md @@ -1,10 +1,10 @@ --- name: paperclip description: > - Interact with the Paperclip control plane API for task coordination and - governance. Use when checking assignments, updating issue status, posting - comments, delegating work, managing routines, or calling Paperclip API - endpoints. + Use for Paperclip-managed tasks and heartbeats: reading task context, delivering + task documents or files, updating completion or blockers, coordinating or + delegating work, and following company governance. Includes control plane API + operations for assignments, comments, approvals, and routines. --- # Paperclip Skill @@ -15,6 +15,8 @@ You run in **heartbeats** — short execution windows triggered by Paperclip. Ea In Paperclip, **task** and **issue** refer to the same work item. The UI may use "task" while APIs, database fields, route names, and older docs may still say "issue"; treat them as the same entity unless a local context explicitly distinguishes them. +**Task documents (API runtimes).** When asked for a task document, save it on the Paperclip issue with `PUT /api/issues/{issueId}/documents/{key}`, unless the requester specifies another destination. Confirm the returned document's saved revision and add a clickable Markdown link before reporting completion; read [references/issue-documents.md](references/issue-documents.md) for the payload and revision-safe updates. Downloadable files follow [Generated Artifacts and Work Products](#generated-artifacts-and-work-products). + ## Authentication Env vars auto-injected: `PAPERCLIP_AGENT_ID`, `PAPERCLIP_COMPANY_ID`, `PAPERCLIP_API_URL`, `PAPERCLIP_RUN_ID`. Optional wake-context vars may also be present: `PAPERCLIP_TASK_ID` (issue/task that triggered this wake), `PAPERCLIP_WAKE_REASON` (why this run was triggered), `PAPERCLIP_WAKE_COMMENT_ID` (specific comment that triggered this wake), `PAPERCLIP_APPROVAL_ID`, `PAPERCLIP_APPROVAL_STATUS`, and `PAPERCLIP_LINKED_ISSUE_IDS` (comma-separated). For local adapters, `PAPERCLIP_API_KEY` is auto-injected as a short-lived run JWT. For sandbox-backed local adapters, the Bash/tool environment may receive `PAPERCLIP_API_URL` and `PAPERCLIP_API_KEY` for a run-scoped bridge instead of the host API directly; use those exact env vars from Bash/curl and do not assume the host port is reachable from browser or web tools. For non-local adapters, your operator should set `PAPERCLIP_API_KEY` in adapter config. All requests use `Authorization: Bearer $PAPERCLIP_API_KEY`. All endpoints are under `/api`. Use JSON except for multipart attachment uploads and binary content downloads. Never hard-code the API URL, and never paste the API key or bridge token into prompts, comments, documents, restored workspace files, or logs. diff --git a/skills/paperclip/references/issue-documents.md b/skills/paperclip/references/issue-documents.md new file mode 100644 index 0000000000..4033dcc942 --- /dev/null +++ b/skills/paperclip/references/issue-documents.md @@ -0,0 +1,58 @@ +# Issue documents through the API + +Use this procedure for requested Paperclip task documents in API-based runtimes. +Honor a destination explicitly chosen by the requester. Native Runner runtimes +use their document tool and its contract. Downloadable files follow +[artifacts.md](artifacts.md). + +## Create a document + +Use the injected `PAPERCLIP_API_URL` and bearer `PAPERCLIP_API_KEY`. Include +`X-Paperclip-Run-Id: $PAPERCLIP_RUN_ID` on the write and send JSON with +`Content-Type: application/json`. + +Choose a descriptive document key such as `report`; keys use lowercase letters, +numbers, underscores, or hyphens and are at most 64 characters. Plans use `plan`. + +```text +PUT /api/issues/{issueId}/documents/{key} +``` + +```json +{ + "title": "Task report", + "format": "markdown", + "body": "# Task report\n\nThe requested findings.", + "baseRevisionId": null +} +``` + +A successful write returns HTTP `201` for creation or `200` for an update, +with the saved document JSON. Check its `key`, `body`, and `latestRevisionId` +before claiming delivery. This receipt is sufficient; another GET is unnecessary +when it confirms the intended content and a saved revision. + +Construct a clickable Markdown link from the current issue and successful write +receipt, then add it to the completion comment or response: + +```javascript +const issuePath = issue.identifier + ? `/${issue.identifier.split("-")[0]}/issues/${issue.identifier}` + : `/issues/${issue.id}`; +const documentUrl = `${issuePath}#document-${saved.key}`; +const comment = `Saved [Task report](${documentUrl}).`; +``` + +Use the returned key: a locked document can redirect an agent's write to a new +document. Unnumbered issues use their ID; the board resolves their company. + +## Update or resolve an unclear write + +For an existing document, first `GET /api/issues/{issueId}/documents/{key}` +and read its body and `latestRevisionId`. Send the updated body with +`baseRevisionId` set to that revision. On `409`, fetch the latest document and +reconcile changes before retrying; preserve the skill's bounded-write retry rule. + +If a write's status or receipt is unclear, GET the document to check the saved +content and revision before retrying. If delivery cannot be confirmed, report +that limitation instead of claiming the requested document was saved. diff --git a/tests/runner-e2e/FIXTURES.md b/tests/runner-e2e/FIXTURES.md index 2ff1eace50..d65997921c 100644 --- a/tests/runner-e2e/FIXTURES.md +++ b/tests/runner-e2e/FIXTURES.md @@ -17,6 +17,17 @@ profile, model qualification, environment, task, or ranking-snapshot change must change that fingerprint automatically so the dashboard can annotate the boundary instead of silently joining unlike totals. +The explicit [stock-harness suite](STOCK-HARNESS.md) wraps existing profiles with +`productionDefaultHireProfile`: omit only `instructionsBundle` so the public +hire route loads the shipped default, while preserving runtime, permissions, +auth, skills, and managed secret references. Do not replace this with a fixture +copy of the default manual. Public receipts check the exact independently +specified bundle before provider execution and again during cleanup, along with +both budget hard stops and actual legacy invocation prompts. Missing evidence +fails closed. The definition fingerprint includes the helper, graders, journey +sources, live fixture, and execution integration; editing those sources changes +the suite revision automatically. + ## Agent profiles Add `RunnerProfileFixture` entries in `catalog.ts`. A profile declares: diff --git a/tests/runner-e2e/README.md b/tests/runner-e2e/README.md index 0da021312d..b652b810ab 100644 --- a/tests/runner-e2e/README.md +++ b/tests/runner-e2e/README.md @@ -444,6 +444,14 @@ pnpm test:e2e:runner -- --list --suite context-integrity pnpm test:e2e:runner -- --id context-integrity.runner-codex.local.ordered-comment-continuation ``` +`stock-harness` reuses ordered continuation, assigned-skill invocation, and chat +restart journeys with production-default hires instead of the custom QA manual. +Its 24 explicit local cells cover eight legacy/native profiles and are excluded +from `--all`. Run `pnpm test:e2e:runner:stock-harness` for the credential-free +instruction-layering, hire, and shared-prompt prerequisites. The +[suite contract](STOCK-HARNESS.md) maps each change to its graders, budgets, +evidence, and remaining qualification limits. + Each hardening oracle has positive and plausible-negative calibration tests. The review grader parses the worker's saved JSON and compares both source values and the consistency verdict. Hiring requires one identity, correct reporting diff --git a/tests/runner-e2e/STOCK-HARNESS.md b/tests/runner-e2e/STOCK-HARNESS.md new file mode 100644 index 0000000000..88e06c2cfb --- /dev/null +++ b/tests/runner-e2e/STOCK-HARNESS.md @@ -0,0 +1,248 @@ +# Stock harness with Paperclip + +This suite covers the instruction reductions tracked in the +[working checklist](../../doc/plans/2026-10-02-stock-harness-paperclip-checklist.md). +Existing context-integrity and chat evals supplied a custom QA instruction +bundle. Their passing results therefore did not qualify a hire with the tiny +production default. The new suite omits that fixture bundle and lets the public +agent-creation route materialize the shipped default. + +The 2026-10-03 legacy ACP Claude repair compares only the original +`assigned-skill-explicit-invocation` cell on matched current-master sources. +Both variants carry identical bounded skill-description/staged-path delivery +and env-free session persistence. Only the historical versus reduced default +manual/shared prompts and their declared structural unit expectations differ. +The public bundle/procedure observations derive from exact admitted source +bytes (eight-word identity or the independently pinned 4,249-byte historical +manual), never a branch/environment label. Candidate absence assertions stay +intact. Independent document/task and exact-credential guards are unchanged. +All stock tasks enforce `single_attempt` in the hosted launcher; automatic +product disposition recovery still counts as actual usage. This is not a new +24-cell campaign or qualification of the frozen native-completion context. + +## Coverage contract + +| Change | Required deterministic evidence | Product E2E evidence | +| --- | --- | --- | +| SH-1: native Codex preserves vendor base instructions | Serialized start/resume requests from the TypeScript driver, recovery paths, runnerd transport, Runner Lab/live sessions, and Rust provider use additive developer instructions. | Real new native Codex hires execute the three journeys below. Task success alone cannot prove vendor base preservation. | +| SH-2: identity-only default hire manual | Existing public agent-creation and onboarding-asset tests cover default, custom, and CEO exceptions. | The public bundle is exactly one `AGENTS.md` containing the eight-word shipped identity, checked before provider execution and again during cleanup. | +| SH-3: reduced shared legacy task/chat defaults and resume delta | Shared prompt tests and ACPX, Codex, OpenCode, Pi, Hermes, and Cursor Cloud adapter regressions retain runtime context and exclude removed generic procedures. | Actual legacy `adapter.invoke` prompts retain fresh identity/connection guidance, omit the removed procedures, and have complete receipts for every observed run. | + +`pnpm test:e2e:runner:stock-harness` runs these credential-free prerequisites, +including the Rust test. It writes JSON reports, the source SHA, a working-source +fingerprint, selected checks, test counts, and `providerCalls: 0` under +`results/stock-harness-preflight-/`. Missing reports or required skipped +assertions fail. Unrelated native tests filtered by the name selector remain +explicitly skipped; they are not counted as executed coverage. The live launcher +runs the prerequisites automatically before loading local credentials or starting +an isolated provider instance. Its subprocess receives only allowlisted toolchain +and operating-system variables. Disposable GitHub runners may resolve Cargo +dependencies; local prerequisites retain offline Cargo execution. +The ordinary server SDK dependency builder prepares the shared/SDK outputs +needed by route tests on a cold install with lifecycle scripts disabled. + +The direct Playwright path also requires the retained prerequisite receipt. It +verifies the exact checkout SHA, evaluated-source fingerprint, all requested +Vitest reports and required assertions, and the Rust result before creating a +company. Missing, stale, partial, or failed prerequisites cannot qualify a cell. +Oracle/admission calibration is itself included in the prerequisite gate. + +## Live matrix + +There are 24 explicit local cells: these eight existing profiles each run three +existing journeys with independently calibrated graders. + +- Legacy: `legacy-codex`, `legacy-claude`, `legacy-opencode`, + `legacy-acp-codex`, `legacy-acp-claude`. +- Native: `runner-codex`, `runner-acpx-claude`, `runner-opencode`. + +| Journey | Expected provider turns | Observable outcome | Cell deadline | +| --- | --- | --- | --- | +| `assigned-skill-explicit-invocation` | 1 | Public skill creation/pinning, an explicit skill request, and saved output containing the marker available only in the skill body. | 12 minutes | +| `ordered-comment-continuation` | 2 | Initial report followed by three ordered public comments, including repeated wording and a changed scope; final saved report preserves the ledger and requested scope. | 12 minutes | +| `continuity-restart` | 3 | Task-backed chat retains the requested context across a server restart and subsequent replies. | 15 minutes | + +A full matrix expects 48 provider turns. Models, credentials, effort, +permissions, assigned skills, environment, and managed secret references are +inherited from the existing profile; only its QA manual is omitted. Both company +and agent receive a 1,000-cent monthly hard stop before provider execution, +verified through public records. These limits do not predict final spend: +attempts, retries, partial runs, unknown billing, and cleanup remain in the +existing campaign accounting and qualification rules. + +`stock-harness` is excluded from `--all`; select it explicitly. It has no Daytona +cells. Pending or unrepresented harnesses are not live-qualified by this matrix. +Pi/Hermes/Cursor Cloud rendering coverage is deterministic here. The separate +Codex-through-ACP base-instruction patch is still an open checklist item: its +legacy cells qualify the common prompt reduction, not vendor base preservation. + +## Run and inspect + +```sh +# No credentials or paid providers: +pnpm test:e2e:runner:stock-harness --list +pnpm test:e2e:runner:stock-harness +pnpm test:e2e:runner:typecheck +pnpm test:e2e:runner:unit +pnpm test:e2e:runner -- --list --suite stock-harness + +# With authorization and the selected profile's required provider credential: +pnpm test:e2e:runner -- --id stock-harness.runner-codex.local.assigned-skill-explicit-invocation +``` + +Use an exact cell first, then expand profile/journey selections when its evidence +is understood. The ordinary launcher, isolated instance, fixture cleanup, +screenshots, attempt history, usage/cost reporting, and dashboard publisher are +unchanged. `snapshots/stock-harness-hire.json` captures the public bundle and +budgets before paid execution. `snapshots/stock-harness.json` captures the final +bundle, budgets, run IDs, actual legacy prompts, and matcher verdicts. Evidence +API failures retain an error snapshot and cannot pass. All snapshots use the +existing secret sanitizer; screenshots remain the original captured pixels. + +The oracle's identity is independent of the implementation constant, so changing +the shipped manual cannot silently change the expected result. The suite +definition digest incorporates the evaluated default manual, shared prompt +implementation and connection guidance, fixture, journeys, grader, prerequisite, and execution integration +sources. Positive and plausible-negative support tests exercise +missing/malformed receipts, manual regrowth, removed startup/resume procedures, +absent connection guidance, and budget drift. Existing calibrated lifecycle +graders still own task/chat success. + +The skill oracle proves the requested pinned skill's output marker reached the +saved result without appearing in the task request or follow-up comments. Skill +tool/read event detection is retained as supporting evidence; it is not a +required cross-provider tool-trace assertion. + +## Qualification status + +Setup was validated locally on 2026-10-02 with the deterministic prerequisites, +Product E2E support tests, typecheck, and discovery: 478 executed prerequisite +tests passed (477 TypeScript plus one Rust), along with 860 support tests and +discovery of all 24 cells. The 313 unrelated native tests filtered by the gate +are not counted as passing coverage. Local evidence is retained under +`results/stock-harness-preflight-2026-10-02T16-15-39.064Z/preflight.json`. +Earlier interrupted, discovery-failure, and setup/test-timeout attempts remain +retained; the final unchanged-assertion retry passed. No live cells or paid +providers were run during that initial setup. This establishes executable coverage, not a live reliability +result or improved coding quality. A quality claim needs comparable tasks, +models, effort, tools, and independently graded before/after results. + +The suite exercises fresh isolated hires and their continuations. Existing saved +manuals and old Codex sessions are not automatically migrated. The latter still +need a provider-session reset to restore a previously replaced vendor base. +Private Runner protocol definitions remain separate: they use mock control-plane +operations and cannot substitute for this public hiring and assembled-prompt +coverage. + +## GitHub qualification and matched comparison + +The instruction reductions and coverage are under review in +[PR #14948](https://github.com/paperclipai/paperclip/pull/14948). +The first diagnostic native Codex skill cell +[passed on GitHub](https://github.com/paperclipai/paperclip/actions/runs/37034213743) +at `a63437069de58d22ee5adbcb6a6202c007dcf037`. All seven independent skill/task +checks passed; the provider run lasted 34.745 seconds and the cell 56.660 seconds. +Its tiny public hire bundle and budget receipts passed, and cleanup passed. +This diagnostic predates enforced prerequisite admission and the expanded source +digest, so it is retained separately from the final qualification matrix. + +The first full candidate attempt at `36e987246b649927e96ce1184cd616c4e490106e` +[was cancelled during prerequisites](https://github.com/paperclipai/paperclip/actions/runs/37037105491). +Cold protected installs disable lifecycle scripts, leaving the plugin SDK unbuilt; +SH-2/SH-3 could not import it. Provider admission was not reached. This setup +failure is retained separately from behavioral results. The prerequisite now runs +the ordinary server dependency builder first, retains its output and exit status, +and refuses admission when setup fails. A cold legacy pilot must pass before the +full matrix retry. + +The [fixed legacy pilot](https://github.com/paperclipai/paperclip/actions/runs/37039240025) +at `4163dbfd0fd4d145bfa52b4d7f80eb59a362ee36` passed SDK setup, hire/shared +prompt/oracle checks and the Rust additive test. It stopped before providers: +the real daemon-frame test lacked the cold `paperclip-runnerd` binary. Setup now +builds that daemon from the locked Rust source before TypeScript gates, retains +`runnerd-build.txt`, and requires both setup exits in the admission receipt. +The required daemon-frame test selects that built debug binary explicitly, +without replacing staged product binaries. The receipt records its SHA-256; +verification rejects a changed binary before provider admission. +The [next cold pilot](https://github.com/paperclipai/paperclip/actions/runs/37040493183) +at `ac6ddefb589c18fa9c30946db40e58499e9fb0e0` then found the required fake Codex +protocol fixture absent. It also stopped before paid providers. Setup uses the +package's ordinary locked workspace `--bins` build, covering both the daemon +and its fixture; verification binds both binaries to the retained receipt. + +That final cold pilot passed all 537 prerequisites on GitHub and reached Claude +Sonnet 4.6. It saved a workspace `task-output.md` and completed the issue, while +the independent durable-document oracle found no Paperclip issue document. +The pinned skill's phrase "task document" does not explicitly name storage in +Paperclip; matched historical results must precede any regression attribution. +The task, instruction delivery, budget, and cleanup receipts remain retained. + +Its trusted merged report rejected the prerequisite folder alongside the campaign +root and synthesized a missing-result infrastructure error. Raw packaged cell +results were uploaded and remain inspectable. Prerequisites now live under the +exact campaign root; the unchanged trusted selector contract is calibrated in +the mandatory gate. The initial full matched campaigns at `f02d8d0df` and +`12c5433c6` retain their original raw results and publication outcomes. Any +reconstructed comparison must declare directory-layout recovery and preserve +every result, hash, original failure, and assertion. + +Dispatch the trusted workflow from `master`, with `target_branch` naming the +same-repository candidate and an exact cell selector first. The workflow resolves +that target once to an immutable SHA. Never dispatch target-controlled workflow +definitions with protected credentials. Protected environments, scoped provider +keys, frozen target dependencies, report sanitization, publication, and existing +bounded retry/cleanup policies retain their existing owners. + +A temporary `codex/stock-harness-previous-instructions` branch compares the same +24 cells with the previous default manual and shared startup/resume prompts. +It holds merged native Codex fix #14920 constant. Only those two production +instruction sources differ. Its explicit historical structural oracle expects +the old manual and records old generic procedures; the candidate's reduction +assertions remain mandatory. Historical tests verify that prior contract. +The independent skill/context/chat journeys, behavioral graders, fixtures, +models, effort, tools, permissions, and credentials are identical. The retained +comparison manifest records their hashes and the restored instruction revision. + +One full campaign per variant initially expects 48 provider turns, plus any +existing bounded automatic retries; the earlier one-cell diagnostic remains +separate. Compare behavioral results by profile and journey, with missing +evidence unqualified. Keep structural instruction differences separate from task +success. Report every attempt, failure attribution, provider timing, token usage, +reported costs and unknown spend. A reported zero subtotal is not proof of zero +provider spending. This small single-trial matrix cannot establish general coding +quality or broad performance equivalence, and does not compare #14920 before/after. + +## Measured comparison and unresolved delivery + +The [dated live report](../../doc/plans/2026-10-02-stock-harness-live-comparison.md) +and its safe JSON projection contain the per-profile/case comparison, exact +source hashes, all campaign and recovery links, timing/usage, cost coverage, +security failures, clipping limits and publication-layout recovery. Candidate +`f02d8d0df` has 24 retained results: 15 pass and nine fail. Historical +`12c5433c6` also has 24 retained results: 15 pass and nine fail. Its single +AWS-runner recovery timed out; original missing evidence remains recorded. Classic Claude/OpenCode skill runs save no Paperclip task document +where historical runs save one; the original oracle remains failed. + +The merged native Codex change is held constant. Same aggregate success counts +would not establish equivalence: case outcomes differ, fixture storage wording +is ambiguous, ACP credential guards fail, and some public prompt receipts are +clipped. Native finish/block tool guidance belongs to native runners; legacy +document-delivery guidance must use the actual Paperclip skill/API path. No +production or skill instructions have been changed to turn the measured failures +into passes. PR #14948 remains draft. + +The corrected packaging pilot at `1eb5ba420` passes on GitHub with valid +evidence and cleanup. It validates prerequisite nesting under the exact campaign +root using the unchanged trusted selector. Its current prerequisite runs 557 +checks (556 TypeScript plus one Rust); all 895 E2E support tests pass. The 313 +filtered native tests are not counted. Original failed/partial attempts and +publication failures remain retained; local reconstructed copies move folders +without editing results, graders or usage. Never dispatch a second development +campaign for an active target branch: workflow concurrency supersedes the older +run. Use a separate frozen-source branch when independent campaigns must overlap. + +## Focused legacy delivery repair + +The original 24 cells and original assigned-skill request/procedure remain unchanged. Two added explicit Paperclip-document cells apply only to classic Claude and OpenCode, making 26 catalog cells (50 expected turns if every cell is selected). The repair campaign selects only the original skill case and new document case for those two profiles: four cells per variant, eight expected turns total. No full matrix rerun is planned. + +Both skill sources are recorded in definition/admission digests. The pre-fix baseline restores the old SKILL.md and records the new reference as absent, without copying the new recipe into that baseline. It holds the tiny manual/shared prompts and all fixture/model/auth/effort inputs fixed. The explicit document oracle checks actual public content/revision and an exact same-app document link, with plausible-negative calibrations. diff --git a/tests/runner-e2e/catalog.test.ts b/tests/runner-e2e/catalog.test.ts index 49748c8a6f..735e90e06f 100644 --- a/tests/runner-e2e/catalog.test.ts +++ b/tests/runner-e2e/catalog.test.ts @@ -136,10 +136,10 @@ describe("runner E2E catalog", () => { expect(localIntegrityTasks).toHaveLength(2); expect(openRouterBreadthTasks).toHaveLength(3); expect(runnerSuites.map((suite) => suite.expectedMatrixSize)).toEqual([ - 6, 30, 3, 16, 16, 2, 6, 8, 46, 23, 50, 6, 20, 52, 28, 18, 2, 6, 6, 12, 10, 48, 16, 10, 2, 1, 1, + 6, 30, 3, 16, 16, 2, 6, 8, 46, 23, 50, 6, 20, 26, 52, 28, 18, 2, 6, 6, 12, 10, 48, 16, 10, 2, 1, 1, ]); - expect(validateRunnerCatalog()).toHaveLength(444); - expect(new Set(runnerMatrix.map((entry) => entry.id)).size).toBe(444); + expect(validateRunnerCatalog()).toHaveLength(470); + expect(new Set(runnerMatrix.map((entry) => entry.id)).size).toBe(470); expect( runnerMatrix.filter((entry) => entry.suite.id === "core-compatibility"), ).toHaveLength(48); diff --git a/tests/runner-e2e/catalog.ts b/tests/runner-e2e/catalog.ts index 89301b7f11..8aa562705e 100644 --- a/tests/runner-e2e/catalog.ts +++ b/tests/runner-e2e/catalog.ts @@ -8,7 +8,8 @@ import { taskTitleTasks, taskTitleDefinitionDigest, TASK_TITLE_BUDGET_CENTS } fr import { blockerTasks, blockerProfile } from "./blocker-cases.js"; import { accountingTasks } from "./accounting-cases.js"; import { continuationTasks } from "./continuation-cases.js"; -import { contextIntegrityTasks } from "./context-integrity-cases.js"; +import { contextIntegrityTasks, paperclipDocumentTask } from "./context-integrity-cases.js"; +import { productionDefaultHireProfile, stockHarnessSourceDigest, stockHarnessSkillSources } from "./stock-harness.js"; import { lifecycleLiveTasks, lifecycleLiveDefinitionDigest } from "./lifecycle-live-cases.js"; import { everydayTasks, productionStoryProfile } from "./everyday-cases.js"; @@ -54,6 +55,7 @@ const SELECTABLE_GROUPS = [ "chat", "onboarding", "context-integrity", + "stock-harness", ] as const; const SAMPLE_UUID = "11111111-1111-4111-8111-111111111111"; @@ -1224,6 +1226,30 @@ export const runnerSuites: readonly RunnerSuiteFixture[] = [ prerequisiteGate: PENDING_PROFILE_PREREQUISITES, }, }, + { + id: "stock-harness", + label: "Stock harness with Paperclip", + manualOnly: true, + description: "Production-default hires, reduced shared prompts, assigned skills, continuation, and persistent chat.", + groups: ["stock-harness", "native", "legacy"], + profiles: [...contextIntegrityProfiles.filter(profile => !pendingContextIntegrityProfiles.includes(profile)), + runnerProfiles.find(profile => profile.id === "legacy-opencode")!].map(productionDefaultHireProfile), + environments: [localEnvironment], + tasks: [...contextIntegrityTasks, chatTasks.find(task => task.id === "continuity-restart")!, paperclipDocumentTask] + .map(task => ({ ...task, automaticRetryPolicy: "single_attempt" as const })), + expectedMatrixSize: 26, + excludedExecutionIds: ["legacy-codex", "legacy-acp-codex", "legacy-acp-claude", "runner-codex", "runner-acpx-claude", "runner-opencode"] + .map(profile => `stock-harness.${profile}.local.${paperclipDocumentTask.id}`), + definitionMetadata: { + version: 2, instructions: "production-default-hire", scheduling: "explicit-only", + automaticRetryPolicy: "single_attempt", maximumAttemptsPerCell: 1, + sourceDigest: stockHarnessSourceDigest(), + operationalSkillSources: stockHarnessSkillSources(), + grading: "public-default-bundle-and-delivered-prompts-plus-independent-lifecycle-oracles", + vendorBaseEvidence: "required deterministic Codex driver/runnerd/Rust gate; task success is not vendor-base proof", + paidCalls: "one skill turn, two ordered-comment turns, three chat turns per profile", + }, + }, { id: "first-task", label: "First-task onboarding", description: "Production onboarding, first replies, approval, and durable task execution.", diff --git a/tests/runner-e2e/context-integrity-cases.ts b/tests/runner-e2e/context-integrity-cases.ts index ac28da1cad..388b87f859 100644 --- a/tests/runner-e2e/context-integrity-cases.ts +++ b/tests/runner-e2e/context-integrity-cases.ts @@ -5,10 +5,14 @@ export const CONTEXT_INTEGRITY_CASES = [ "assigned-skill-explicit-invocation", ] as const; -export type ContextIntegrityCase = (typeof CONTEXT_INTEGRITY_CASES)[number]; +export const PAPERCLIP_DOCUMENT_CASE = "assigned-skill-paperclip-document" as const; +export type ContextIntegrityCase = (typeof CONTEXT_INTEGRITY_CASES)[number] | typeof PAPERCLIP_DOCUMENT_CASE; +export function isAssignedSkillContext(id: string): boolean { + return id === "assigned-skill-explicit-invocation" || id === PAPERCLIP_DOCUMENT_CASE; +} export function contextIntegrityScenario(id: string, nonce: string) { - if (!CONTEXT_INTEGRITY_CASES.includes(id as ContextIntegrityCase)) + if (!CONTEXT_INTEGRITY_CASES.includes(id as (typeof CONTEXT_INTEGRITY_CASES)[number]) && id !== PAPERCLIP_DOCUMENT_CASE) throw new Error(`Unknown context-integrity case: ${id}`); const marker = `CONTEXT_OK_${nonce.replace(/[^a-z0-9]/gi, "")}`; const repeated = "Append an entry for every request, even identical wording."; @@ -23,15 +27,17 @@ export function contextIntegrityScenario(id: string, nonce: string) { prompt: id === "ordered-comment-continuation" ? `Maintain a packing list report document. Start with the initial scope: passport and charger. Save the initial report and leave this task waiting for follow-up comments. When follow-up comments arrive, copy each user's exact wording verbatim into one ordered ledger, including identical wording repeated more than once. After the follow-up batch, save one report containing both initial items and every verbatim request in arrival order. Add a separate section headed "## Final scope" outside the quoted ledger. Put only the final requested scope name as plain text in that section. Then finish. Do not invent extra comments, child tasks, or deliverables.` - : "Use the assigned Context integrity output skill for this task, then follow its instructions and finish the task. Do not create child tasks or unrelated deliverables.", + : id === PAPERCLIP_DOCUMENT_CASE + ? "Use the assigned Context integrity output skill. Save its requested document as a Paperclip document on this task, check the saved revision and content, and post a completion comment with a clickable Markdown link to the saved document in the Paperclip task UI, including the company prefix. A workspace file alone does not deliver this requested Paperclip document. Finish after delivery. Do not create child tasks or unrelated deliverables." + : "Use the assigned Context integrity output skill for this task, then follow its instructions and finish the task. Do not create child tasks or unrelated deliverables.", comments: [repeated, repeated, changed] as const, }; } -export const contextIntegrityTasks: readonly RunnerTaskFixture[] = - CONTEXT_INTEGRITY_CASES.map((id) => ({ +function contextIntegrityTask(id: ContextIntegrityCase): RunnerTaskFixture { + return { id, - label: id === "ordered-comment-continuation" ? "Ordered comment continuation" : "Assigned skill invocation", + label: id === "ordered-comment-continuation" ? "Ordered comment continuation" : id === PAPERCLIP_DOCUMENT_CASE ? "Explicit Paperclip document delivery" : "Assigned skill invocation", groups: ["context-integrity"], workMode: "standard", flow: "context_integrity", @@ -42,4 +48,8 @@ export const contextIntegrityTasks: readonly RunnerTaskFixture[] = buildPrompt: (nonce) => contextIntegrityScenario(id, nonce).prompt, buildVisibleMarker: (nonce) => contextIntegrityScenario(id, nonce).marker, buildMatchers: () => [], - })); + }; +} + +export const contextIntegrityTasks: readonly RunnerTaskFixture[] = CONTEXT_INTEGRITY_CASES.map(contextIntegrityTask); +export const paperclipDocumentTask = contextIntegrityTask(PAPERCLIP_DOCUMENT_CASE); diff --git a/tests/runner-e2e/context-integrity-flow.ts b/tests/runner-e2e/context-integrity-flow.ts index b1a2dabb94..ad1e240111 100644 --- a/tests/runner-e2e/context-integrity-flow.ts +++ b/tests/runner-e2e/context-integrity-flow.ts @@ -1,6 +1,6 @@ import type { Page } from "@playwright/test"; import { randomUUID } from "node:crypto"; -import { contextIntegrityScenario } from "./context-integrity-cases.js"; +import { contextIntegrityScenario, isAssignedSkillContext } from "./context-integrity-cases.js"; import { gradeContextIntegrity, type ContextIntegrityCheckpoint } from "./context-integrity-scoring.js"; import { pollUntil, type RunnerApi } from "./api.js"; import { createTaskThroughUi } from "./user-actions.js"; @@ -102,7 +102,7 @@ export async function runContextIntegrityFlow(input: { let assignedVersionId: string | null = null; const checkpoints: ContextIntegrityCheckpoint[] = []; - if (scenario.id === "assigned-skill-explicit-invocation") { + if (isAssignedSkillContext(scenario.id)) { await api.patch("/api/instance/settings/experimental", { enableBetaSkills: true }); const markdown = `---\nname: ${scenario.skillKey}\ndescription: Context integrity output procedure.\n---\n\n# Context integrity output procedure\n\nWrite exactly one task document whose body contains the marker ${scenario.marker}. Finish the task after saving that document.`; if (!markdown.startsWith(`---\nname: ${scenario.skillKey}\ndescription:`) || markdown.includes("\\n")) { @@ -141,18 +141,18 @@ export async function runContextIntegrityFlow(input: { api.get(`/api/issues/${issue!.id}/documents`), api.get(`${companyPath}/skills`), api.get(`/api/issues/${issue!.id}/queued-comments`), - scenario.id === "assigned-skill-explicit-invocation" ? api.get(`/api/agents/${fixtures.agent.id}/skills?companyId=${fixtures.company.id}`) : Promise.resolve({} as Row), + isAssignedSkillContext(scenario.id) ? api.get(`/api/agents/${fixtures.agent.id}/skills?companyId=${fixtures.company.id}`) : Promise.resolve({} as Row), ]); const detailedDocuments = await Promise.all(documents.map((document) => api.get(`/api/issues/${issue!.id}/documents/${encodeURIComponent(String(document.key))}`))); - const assignedSkill = scenario.id === "assigned-skill-explicit-invocation" + const assignedSkill = isAssignedSkillContext(scenario.id) ? skillRows.find((skill) => skill.key === scenario.assignedSkill?.key || skill.slug === scenario.assignedSkill?.key) : undefined; const desiredSkill = (assignedState.desiredSkillEntries as Array | undefined)?.find((entry) => entry.key === scenario.assignedSkill?.key); const runEvents = await Promise.all(runs.map((run) => api.get(`/api/heartbeat-runs/${run.id}/events?limit=1000`))); const runLogs = await Promise.all(runs.map((run) => readContextIntegrityRunLog(api, run.id))); - const skillInvocationEvidence = scenario.id === "assigned-skill-explicit-invocation" && runEvents.some((events) => events.some((event) => containsExplicitSkillInput(event, String(assignedSkill?.slug ?? scenario.skillKey)))); - skillRequestText = scenario.id === "assigned-skill-explicit-invocation" ? String(issue?.description ?? "") : ""; - checkpoints.push({ phase, issue: { id: issue!.id, status: String(issue!.status) }, comments, queuedComments, documents: detailedDocuments as Array<{ key: string; body?: string | null }>, runs, runEvents: runEvents.flat(), runLogs, assignedSkill: assignedSkill ? { key: String(desiredSkill?.key ?? assignedSkill.key ?? assignedSkill.slug), runtimeName: String(assignedSkill.slug ?? ""), versionId: desiredSkill?.versionId === assignedVersionId ? assignedVersionId : null, markdown: String(assignedSkill.markdown ?? scenario.assignedSkill?.markdown ?? "") } : undefined, skillRequestText: scenario.id === "assigned-skill-explicit-invocation" ? skillRequestText : undefined, skillInvocationEvidence }); + const skillInvocationEvidence = isAssignedSkillContext(scenario.id) && runEvents.some((events) => events.some((event) => containsExplicitSkillInput(event, String(assignedSkill?.slug ?? scenario.skillKey)))); + skillRequestText = isAssignedSkillContext(scenario.id) ? String(issue?.description ?? "") : ""; + checkpoints.push({ phase, issue: { id: issue!.id, status: String(issue!.status), identifier: String(issue!.identifier ?? ""), issuePrefix: String(fixtures.company.issuePrefix ?? ""), appOrigin: new URL(page.url()).origin }, comments, queuedComments, documents: detailedDocuments as Array<{ key: string; body?: string | null }>, runs, runEvents: runEvents.flat(), runLogs, assignedSkill: assignedSkill ? { key: String(desiredSkill?.key ?? assignedSkill.key ?? assignedSkill.slug), runtimeName: String(assignedSkill.slug ?? ""), versionId: desiredSkill?.versionId === assignedVersionId ? assignedVersionId : null, markdown: String(assignedSkill.markdown ?? scenario.assignedSkill?.markdown ?? "") } : undefined, skillRequestText: isAssignedSkillContext(scenario.id) ? skillRequestText : undefined, skillInvocationEvidence }); const checks = gradeContextIntegrity({ id: scenario.id, marker: scenario.marker, comments: scenario.comments, checkpoints }); input.observe(issue!, runs, checks); await input.evidence("context-integrity.json", { schema: "paperclip.context-integrity.v1", scenario, budgetGuard, checkpoints, checks }); @@ -175,7 +175,7 @@ export async function runContextIntegrityFlow(input: { await api.patch(`/api/agents/${fixtures.agent.id}/budgets`, { budgetMonthlyCents: budgetGuard.agentMonthlyCents, }); - const taskPrompt = scenario.id === "assigned-skill-explicit-invocation" + const taskPrompt = isAssignedSkillContext(scenario.id) ? `${scenario.prompt}\n\nUse /${scenario.skillKey} for this request.` : scenario.prompt; skillRequestText = taskPrompt; diff --git a/tests/runner-e2e/context-integrity-scoring.ts b/tests/runner-e2e/context-integrity-scoring.ts index e307a66bc0..c1aae9c266 100644 --- a/tests/runner-e2e/context-integrity-scoring.ts +++ b/tests/runner-e2e/context-integrity-scoring.ts @@ -1,11 +1,11 @@ -import type { ContextIntegrityCase } from "./context-integrity-cases.js"; +import { isAssignedSkillContext, PAPERCLIP_DOCUMENT_CASE, type ContextIntegrityCase } from "./context-integrity-cases.js"; export interface ContextIntegrityCheckpoint { phase: "initial" | "comment-1" | "comment-2" | "comment-3" | "final"; - issue: { id: string; status: string }; + issue: { id: string; status: string; identifier?: string; issuePrefix?: string; appOrigin?: string }; comments: Array>; queuedComments?: Record; - documents: Array<{ key: string; body?: string | null }>; + documents: Array<{ key: string; body?: string | null; latestRevisionId?: string | null; latestRevisionNumber?: number | null }>; runs: Array>; assignedSkill?: { key: string; runtimeName?: string; versionId?: string | null; markdown?: string }; skillRequestText?: string; @@ -117,7 +117,7 @@ export function gradeContextIntegrity(input: { const expectedCommentIds = humanRows.slice(0, input.comments.length).map((comment) => String(comment.id ?? "")); check("continuation-wake-comment-ids", Boolean(continuation) && JSON.stringify(wakeCommentIds(continuation)) === JSON.stringify(expectedCommentIds), "The deferred continuation wake must carry all public comment IDs in arrival order."); } - if (input.id === "assigned-skill-explicit-invocation") { + if (isAssignedSkillContext(input.id)) { const runtimeName = String(initial?.assignedSkill?.runtimeName ?? initial?.assignedSkill?.key ?? ""); const requestText = String(initial?.skillRequestText ?? ""); const references = requestText.split(/\s+/).map((word) => word.replace(/[.,]$/, "")); @@ -126,8 +126,27 @@ export function gradeContextIntegrity(input: { check("skill-source-marker", Boolean(initial?.assignedSkill?.markdown?.includes(input.marker)), "The assigned pinned skill source must contain the output marker."); check("marker-not-in-request", !requestText.includes(input.marker) && !finalComments.some((comment) => comment.includes(input.marker)), "The output marker must originate from the assigned skill, not the task request or comments."); } - const output = final ? oneOutput(final, input.marker, input.id === "assigned-skill-explicit-invocation") : undefined; - check("single-durable-output", Boolean(output) && final!.documents.length === 1, input.id === "assigned-skill-explicit-invocation" ? "Exactly one durable task document must contain the skill's marker." : "Exactly one durable packing report document must be saved."); + const output = final ? oneOutput(final, input.marker, isAssignedSkillContext(input.id)) : undefined; + check("single-durable-output", Boolean(output) && final!.documents.length === 1, isAssignedSkillContext(input.id) ? "Exactly one durable task document must contain the skill's marker." : "Exactly one durable packing report document must be saved."); + if (input.id === PAPERCLIP_DOCUMENT_CASE) { + check("saved-document-revision", Boolean(output?.latestRevisionId) && Number.isInteger(output?.latestRevisionNumber) && Number(output?.latestRevisionNumber) > 0, + "The public saved document must have a persisted revision and the requested content."); + const route = output && final?.issue.identifier && final.issue.issuePrefix + ? `/${final.issue.issuePrefix}/issues/${final.issue.identifier}#document-${output.key}` : undefined; + const agentComments = (final?.comments ?? []).filter(comment => comment.authorType === "agent" || Boolean(comment.authorAgentId)); + const origin = final?.issue.appOrigin; + const usableLink = Boolean(route && origin && agentComments.some(comment => { + const links = String(comment.body ?? "").matchAll(/\[[^\]]*\]\((?:<([^>]+)>|([^\s)]+))(?:\s+["'][^"']*["'])?\)/g); + return [...links].some(link => { + try { + const url = new URL(link[1] ?? link[2]!, origin); + return url.origin === origin && `${url.pathname}${url.hash}` === route && !url.search; + } catch { return false; } + }); + })); + check("saved-document-link", usableLink, + "An agent completion comment must link the exact saved document on this task; a local path or claimed URL does not suffice."); + } check("completed-task", final?.issue.status === "done", `Final task status: ${final?.issue.status ?? "missing"}.`); check("successful-runs", Boolean(final?.runs.length) && final!.runs.every((run) => run.status === "succeeded"), "All recorded context-integrity runs must succeed."); return checks; diff --git a/tests/runner-e2e/fixtures/stock-harness/historical-default-agents.md b/tests/runner-e2e/fixtures/stock-harness/historical-default-agents.md new file mode 100644 index 0000000000..1472483914 --- /dev/null +++ b/tests/runner-e2e/fixtures/stock-harness/historical-default-agents.md @@ -0,0 +1,24 @@ +You are an agent at Paperclip company. + +## Execution Contract + +- Start actionable work in the same heartbeat. Do not stop at a plan unless the issue explicitly asks for planning. +- Keep the work moving until it is done. If you need QA to review it, ask them. If you need your boss to review it, ask them. +- Leave durable progress in task comments, documents, or work products, then update the issue to a clear final disposition before you exit. +- When your work produces a user-inspectable deliverable file, follow the Paperclip skill's "Generated Artifacts and Work Products" workflow before final disposition. Use `skills/paperclip/scripts/paperclip-upload-artifact.sh` when working in this repo, create/update an artifact work product when the file is the deliverable, and link the uploaded attachment in the final comment. Do not rely on local filesystem paths as the only access path. If an important file intentionally remains workspace-only, create/update a work product with `metadata.resourceRef.kind: "workspace_file"` and a workspace-relative path, then name that work product and path in the final comment. Treat browse/search as a fallback for recovering workspace files, not the preferred deliverable path. +- When your work produces or updates an operator-facing engineering output, create/update the matching work product: `pull_request` for opened PRs, `preview_url` for published previews, `runtime_service` for managed preview/dev services, `commit` for notable pushed commits, and `branch` when the branch itself is the handoff. A comment is not a substitute for the work product access path. +- Comments, documents, screenshots, work products, and `Remaining` bullets are evidence, not valid liveness paths by themselves. +- Final disposition checklist: mark `done` when complete and verified; use `in_review` only with a real reviewer, approval, interaction, or monitor path; use `blocked` only with first-class blockers or a named unblock owner/action; create delegated follow-up issues with blockers when another agent owns the next step; keep `in_progress` only when a live continuation path exists. +- Use child issues for parallel or long delegated work instead of polling agents, sessions, or processes. +- Create child issues directly when you know what needs to be done. If the board/user needs to choose suggested tasks, answer structured questions, or confirm a proposal first, create an issue-thread interaction on the current issue with `POST /api/issues/{issueId}/interactions` using `kind: "suggest_tasks"`, `kind: "ask_user_questions"`, or `kind: "request_confirmation"`. +- Use `request_confirmation` instead of asking for yes/no decisions in markdown. Before presenting a plan for review, you MUST complete this publish contract: + 1. `PUT /issues/{id}/documents/plan` with `{ format: 'markdown', body, changeSummary }`. + 2. Re-`GET /documents/plan`, assert it returns `200`, and capture its `latestRevisionId`. + 3. Only then create `request_confirmation` with `target={ type: 'issue_document', key: 'plan', revisionId: latestRevisionId }` and `idempotencyKey=confirmation:{issueId}:plan:{revisionId}`. + 4. Wait for acceptance before creating implementation subtasks. + Never present a plan only in a thread comment or through `ask_user_questions`; comments are supporting context and questions are for gathering input, not plan review. +- `ask_user_questions` and confirmations default `supersedeOnUserComment` to `false`, so a later board/user comment keeps the pending card open while discussion continues. Set it to `true` when a new comment should replace the pending request. If you wake up from a superseding comment, revise the artifact, question set, or proposal and create a fresh interaction if input is still needed. +- For human input, save a pending question/confirmation interaction and set `in_review`; prose alone does not create a waiting path. Use `blockedByIssueIds` for issue dependencies. An agent may set an `unblockDescriptor` only for itself (`owner: { "agentId": "" }` plus `action`), not for the board/user or another agent. +- Respect budget, pause/cancel, approval gates, and company boundaries. + +Do not let work sit here. You must always update your task with a comment. diff --git a/tests/runner-e2e/launch.ts b/tests/runner-e2e/launch.ts index 7cafc2c8f0..2e28e2a975 100644 --- a/tests/runner-e2e/launch.ts +++ b/tests/runner-e2e/launch.ts @@ -52,6 +52,8 @@ import { } from "./types.js"; import { assertRunnerE2EPrerequisites } from "./prerequisites.js"; import { assertNativeCompletionSelection, prepareNativeCompletionPreflight, NATIVE_COMPLETION_PREFLIGHT_ENV } from "./native-completion-admission.js"; +import { prepareStockHarnessPreflight, STOCK_PREFLIGHT_ENV } from "./stock-harness-admission.js"; + import { reapNewDetachedDarwinSharedMemory, snapshotDarwinSharedMemory, @@ -1084,13 +1086,17 @@ async function main() { // profiles remain discoverable, but cannot reach a provider. assertRunnerE2EPrerequisites(executions); assertNativeCompletionSelection(executions); - let nativeCampaignId: string | undefined; + const campaignId = cleanId( + process.env.PAPERCLIP_E2E_CAMPAIGN_ID ?? + `local-${new Date().toISOString().replace(/[:.]/g, "-")}`, + ); + const summaryDir = path.join(resultsRoot, campaignId); + await mkdir(summaryDir, { recursive: true }); if (executions.some(execution => execution.suite.id === "native-completion")) { - nativeCampaignId = cleanId(process.env.PAPERCLIP_E2E_CAMPAIGN_ID ?? - `local-${new Date().toISOString().replace(/[:.]/g, "-")}`); - const nativeCampaignDirectory = path.join(resultsRoot, nativeCampaignId); - await mkdir(nativeCampaignDirectory, { recursive: true }); - process.env[NATIVE_COMPLETION_PREFLIGHT_ENV] = prepareNativeCompletionPreflight(nativeCampaignDirectory); + process.env[NATIVE_COMPLETION_PREFLIGHT_ENV] = prepareNativeCompletionPreflight(summaryDir); + } + if (executions.some(execution => execution.suite.id === "stock-harness")) { + process.env[STOCK_PREFLIGHT_ENV] = prepareStockHarnessPreflight(summaryDir); } await loadLocalEnvironment(process.env); @@ -1113,12 +1119,7 @@ async function main() { ); } - const campaignId = nativeCampaignId ?? cleanId( - process.env.PAPERCLIP_E2E_CAMPAIGN_ID ?? - `local-${new Date().toISOString().replace(/[:.]/g, "-")}`, - ); - const summaryDir = path.join(resultsRoot, campaignId); - await mkdir(summaryDir, { recursive: true }); + await writeFile( path.join(summaryDir, "invocation-policy.json"), `${JSON.stringify( diff --git a/tests/runner-e2e/live-fixtures.test.ts b/tests/runner-e2e/live-fixtures.test.ts index cf99914f16..04be71d847 100644 --- a/tests/runner-e2e/live-fixtures.test.ts +++ b/tests/runner-e2e/live-fixtures.test.ts @@ -4,6 +4,30 @@ import { runnerMatrix } from "./catalog.js"; import { setupLiveFixtures } from "./live-fixtures.js"; describe("live runner fixtures", () => { + it.each(["runner-codex", "legacy-codex", "runner-acpx-claude", "legacy-opencode"])( + "creates a production-default %s hire with company and agent budget stops", async (profile) => { + const execution = runnerMatrix.find(row => row.suite.id === "stock-harness" && row.profile.id === profile)!; + let agentBody: any; + const api = { + async get() { return [{ id: "local", driver: "local" }]; }, + async postSensitive() { return { id: "secret" }; }, + async post(url: string, data: any) { + if (url === "/api/companies") { + expect(data.budgetMonthlyCents).toBe(1_000); + return { id: "company", name: "Fixture" }; + } + if (url.endsWith("/agents")) { agentBody = data; return { id: "agent", ...data }; } + throw new Error(`Unexpected POST ${url}`); + }, + } as unknown as RunnerApi; + const fixtures = await setupLiveFixtures({ api, execution, executionNonce: "nonce", workspacePath: "/tmp/fixture", + credentials: { [execution.profile.credential]: "fixture-key" } }); + expect(agentBody).not.toHaveProperty("instructionsBundle"); + expect(agentBody.budgetMonthlyCents).toBe(1_000); + expect(agentBody.adapterConfig.env[execution.profile.credential]).toEqual({ type: "secret_ref", secretId: "secret", version: "latest" }); + await fixtures.teardown(); + }, + ); it("anchors extended file validation to a public project workspace for local and remote copy-back", async () => { const execution = runnerMatrix.find(e => e.id === "extended-harnesses.runner-acpx-pi.local.file-edit-validate")!; let projectBody: any; diff --git a/tests/runner-e2e/live-fixtures.ts b/tests/runner-e2e/live-fixtures.ts index da0aa8e19c..391bdc6448 100644 --- a/tests/runner-e2e/live-fixtures.ts +++ b/tests/runner-e2e/live-fixtures.ts @@ -136,7 +136,9 @@ export async function setupLiveFixtures(input: { return api.post("/api/companies", { name: `Runner E2E ${execution.id} ${input.executionNonce}`, description: "Ephemeral paid full-stack runner acceptance fixture", - budgetMonthlyCents: execution.suite.id === "native-completion" ? NATIVE_COMPLETION_BUDGET_CENTS : execution.suite.id === "task-titles" ? TASK_TITLE_BUDGET_CENTS : 0, + budgetMonthlyCents: execution.suite.id === "native-completion" ? NATIVE_COMPLETION_BUDGET_CENTS + : execution.suite.id === "task-titles" ? TASK_TITLE_BUDGET_CENTS + : execution.suite.id === "stock-harness" ? 1_000 : 0, }); }, async teardown() { @@ -288,6 +290,7 @@ export async function setupLiveFixtures(input: { secretRefs, executionId: input.executionNonce, }); + if (execution.suite.id === "stock-harness") agent.budgetMonthlyCents = 1_000; if (managedHiring) { const account = value(resolved, "ai-connection"); const config = agent.adapterConfig as Record; diff --git a/tests/runner-e2e/native-completion-defaults.test.ts b/tests/runner-e2e/native-completion-defaults.test.ts index 21c5c6dc08..8777a2cb39 100644 --- a/tests/runner-e2e/native-completion-defaults.test.ts +++ b/tests/runner-e2e/native-completion-defaults.test.ts @@ -1,5 +1,6 @@ import { describe, expect, it } from "vitest"; import { readFileSync } from "node:fs"; +import { createHash } from "node:crypto"; import { captureNativeDefault, gradeNativeDefault, nativeCompletionProfile, nativeCompletionWorkspaceDigest, NATIVE_MASTER_DEFAULT_SHA256, type NativeDefaultReceipt } from "./native-completion-defaults.js"; import { mkdtemp, mkdir, writeFile, symlink, rm } from "node:fs/promises"; import { tmpdir } from "node:os"; @@ -58,7 +59,7 @@ describe("native production defaults", () => { expect(original.instructionsBundle).toBeDefined(); expect(() => nativeCompletionProfile({ ...profile, generation: "legacy" })).toThrow(); }); - it("hashes actual public bundle bytes before execution without publishing the body", async () => { + it("captures current default configuration independently of the frozen native qualification contract", async () => { const content = readFileSync(new URL("../../server/src/onboarding-assets/default/AGENTS.md", import.meta.url), "utf8"); const called: string[] = []; const api = { async get(path: string): Promise { @@ -67,7 +68,12 @@ describe("native production defaults", () => { : path.includes("instructions-bundle/file") ? { content } : { budgetMonthlyCents: 1000 }) as T; } }; const actual = await captureNativeDefault({ api, agentId: "agent", companyId: "company" }); - expect(gradeNativeDefault(actual).passed).toBe(true); + const currentSha256 = createHash("sha256").update(content).digest("hex"); + expect(actual.files).toEqual([{ path: "AGENTS.md", sha256: currentSha256, bytes: Buffer.byteLength(content) }]); + // The archived native comparison intentionally pins its original master + // context. Accurate capture of a later configuration does not qualify it + // against those historical live results. + expect(gradeNativeDefault(actual).passed).toBe(currentSha256 === NATIVE_MASTER_DEFAULT_SHA256); expect(called).toContain("/api/agents/agent/instructions-bundle/file?path=AGENTS.md"); expect(JSON.stringify(actual)).not.toContain(content); }); diff --git a/tests/runner-e2e/paperclip-document.test.ts b/tests/runner-e2e/paperclip-document.test.ts new file mode 100644 index 0000000000..4fd62239c0 --- /dev/null +++ b/tests/runner-e2e/paperclip-document.test.ts @@ -0,0 +1,89 @@ +import { readFileSync } from "node:fs"; +import { runInNewContext } from "node:vm"; +import { describe, expect, it } from "vitest"; +import { contextIntegrityScenario, PAPERCLIP_DOCUMENT_CASE } from "./context-integrity-cases.js"; +import { gradeContextIntegrity, type ContextIntegrityCheckpoint } from "./context-integrity-scoring.js"; + +function recording() { + const scenario = contextIntegrityScenario(PAPERCLIP_DOCUMENT_CASE, "probe"); + const initial: ContextIntegrityCheckpoint = { + phase: "initial", issue: { id: "issue", identifier: "DOC-1", issuePrefix: "DOC", appOrigin: "https://paperclip.example", status: "in_progress" }, + documents: [], comments: [], runs: [{ id: "run", status: "running" }], + assignedSkill: { key: scenario.skillKey, runtimeName: scenario.skillKey, versionId: "version", markdown: `Write ${scenario.marker}` }, + skillRequestText: `${scenario.prompt}\nUse /${scenario.skillKey}`, + }; + const final: ContextIntegrityCheckpoint = { + ...initial, phase: "final", issue: { ...initial.issue, status: "done" }, + documents: [{ key: "report", body: `Requested ${scenario.marker}`, latestRevisionId: "revision", latestRevisionNumber: 1 }], + comments: [{ authorAgentId: "agent", body: "Saved [report](/DOC/issues/DOC-1#document-report)." }], + runs: [{ id: "run", status: "succeeded" }], + }; + return { scenario, initial, final }; +} + +describe("explicit Paperclip document delivery", () => { + it("checks the actual saved content, revision and exact document link", () => { + const { scenario, initial, final } = recording(); + expect(gradeContextIntegrity({ ...scenario, checkpoints: [initial, final] }).every(row => row.passed)).toBe(true); + expect(scenario.prompt).toContain("Paperclip document"); + expect(scenario.prompt).toContain("clickable Markdown link"); + expect(scenario.prompt).toContain("Paperclip task UI"); + expect(scenario.prompt).not.toContain(scenario.marker); + }); + it.each(["https://paperclip.example/DOC/issues/DOC-1#document-report", ""])("accepts the same-app document link %s", target => { + const { scenario, initial, final } = recording(); + final.comments[0]!.body = `Saved [report](${target}).`; + expect(gradeContextIntegrity({ ...scenario, checkpoints: [initial, final] }).every(row => row.passed)).toBe(true); + }); + it.each(["local-only", "missing-revision", "wrong-content", "local-link", "wrong-document-link", "wrong-company-link", "other-app-link", "user-link-only", "bare-path", "code-formatted-path"])( + "rejects plausible %s delivery", variant => { + const { scenario, initial, final } = recording(); + if (variant === "local-only") { final.documents = []; final.comments[0]!.body = "Saved report.md in the workspace."; } + if (variant === "missing-revision") final.documents[0]!.latestRevisionId = null; + if (variant === "wrong-content") final.documents[0]!.body = "A different report"; + if (variant === "bare-path") final.comments[0]!.body = "Saved /DOC/issues/DOC-1#document-report."; + if (variant === "code-formatted-path") final.comments[0]!.body = "Saved `/DOC/issues/DOC-1#document-report`."; + if (variant === "local-link") final.comments[0]!.body = "Saved [report](./report.md)."; + if (variant === "wrong-document-link") final.comments[0]!.body = "Saved [report](/DOC/issues/DOC-1#document-other)."; + if (variant === "wrong-company-link") final.comments[0]!.body = "Saved [report](/PAP/issues/DOC-1#document-report)."; + if (variant === "other-app-link") final.comments[0]!.body = "Saved [report](https://other.example/DOC/issues/DOC-1#document-report)."; + if (variant === "user-link-only") { delete final.comments[0]!.authorAgentId; final.comments[0]!.authorUserId = "user"; } + expect(gradeContextIntegrity({ ...scenario, checkpoints: [initial, final] }).some(row => !row.passed)).toBe(true); + }, + ); + it.each([ + { identifier: "ACME-17", issuePrefix: "ACME", key: "report" }, + { identifier: "TEAM-902", issuePrefix: "TEAM", key: "report-agent-redirected" }, + ])("the shipped recipe delivers the current $identifier and returned $key", ({ identifier, issuePrefix, key }) => { + const reference = readFileSync(new URL("../../skills/paperclip/references/issue-documents.md", import.meta.url), "utf8"); + const recipe = reference.match(/```javascript\n([\s\S]*?)```/)?.[1]; + expect(recipe).toBeDefined(); + const { scenario, initial, final } = recording(); + initial.issue = { ...initial.issue, identifier, issuePrefix }; + final.issue = { ...final.issue, identifier, issuePrefix }; + final.documents[0]!.key = key; + final.comments[0]!.body = runInNewContext(`${recipe}\ncomment`, { + issue: { identifier }, saved: { key }, + }) as string; + expect(gradeContextIntegrity({ ...scenario, checkpoints: [initial, final] }).every(row => row.passed)).toBe(true); + }); + it.each([null, undefined])("the shipped recipe supports an unnumbered issue with identifier %s", identifier => { + const reference = readFileSync(new URL("../../skills/paperclip/references/issue-documents.md", import.meta.url), "utf8"); + const recipe = reference.match(/```javascript\n([\s\S]*?)```/)?.[1]; + expect(recipe).toBeDefined(); + const comment = runInNewContext(`${recipe}\ncomment`, { + issue: { id: "849c6474-ab4e-4428-981a-0b9a0207fa9b", identifier }, + saved: { key: "report-agent-redirected" }, + }) as string; + const target = comment.match(/\]\(([^)]+)\)/)?.[1]; + expect(target).toBeDefined(); + const url = new URL(target!, "https://paperclip.example"); + expect(url.pathname).toBe("/issues/849c6474-ab4e-4428-981a-0b9a0207fa9b"); + expect(url.hash).toBe("#document-report-agent-redirected"); + }); + it("preserves the original ambiguous request and assigned output procedure", () => { + expect(contextIntegrityScenario("assigned-skill-explicit-invocation", "probe").prompt).toBe( + "Use the assigned Context integrity output skill for this task, then follow its instructions and finish the task. Do not create child tasks or unrelated deliverables.", + ); + }); +}); diff --git a/tests/runner-e2e/runner.spec.ts b/tests/runner-e2e/runner.spec.ts index 46cc8e96aa..f94562fbb4 100644 --- a/tests/runner-e2e/runner.spec.ts +++ b/tests/runner-e2e/runner.spec.ts @@ -18,6 +18,8 @@ import { runContinuationFlow } from "./continuation-flow.js"; import { runEverydayFlow } from "./everyday-flow.js"; import { gradeTaskTitle } from "./task-titles.js"; import { runContextIntegrityFlow } from "./context-integrity-flow.js"; +import { captureStockHarness, gradeStockHarness, gradeStockHire } from "./stock-harness.js"; +import { verifyStockHarnessPreflight, STOCK_PREFLIGHT_ENV } from "./stock-harness-admission.js"; import { createTaskThroughUi, submitTaskReply } from "./user-actions.js"; import { runFirstTaskFlow, setupFirstTaskFixtures } from "./first-task-flow.js"; @@ -825,6 +827,10 @@ for (const execution of executions) { }); try { + if (execution.suite.id === "stock-harness") { + const receipt = verifyStockHarnessPreflight(process.env[STOCK_PREFLIGHT_ENV]); + await writeSanitizedJson(snapshotsDir, "stock-harness-preflight.json", receipt, secrets); + } const experimental = await api.patch<{ enableNativeRunner: boolean; }>("/api/instance/settings/experimental", { @@ -865,6 +871,18 @@ for (const execution of executions) { nativeInitial = { issueIds: issues.map(value => value.id), agentIds: agents.map(value => value.id), workspaceDigest }; } + if (execution.suite.id === "stock-harness") { + const hire = await captureStockHarness({ api, companyId: fixtures.company.id, + agentId: fixtures.agent.id, generation: execution.profile.generation, runIds: [] }); + const checks = gradeStockHire(hire); + await writeSanitizedJson(snapshotsDir, "stock-harness-hire.json", { + capturePhase: "before-provider", ...hire, checks, + }, secrets); + const failed = checks.filter(check => !check.passed); + if (failed.length) throw new Error(`Stock hire matcher failures: ${failed.map(check => check.id).join(", ")}`); + + } + if (execution.suite.id === "api-response-reading") { const source = await api.post<{ id: string }>(`/api/companies/${fixtures.company.id}/issues`, { title: `Synthetic diagnostic evidence ${nonce}`, @@ -2764,6 +2782,24 @@ for (const execution of executions) { selectedRuns = await Promise.all(companyRuns.map(run => api.get(`/api/heartbeat-runs/${run.id}`))); await writeSanitizedJson(snapshotsDir, execution.task.flow === "first_task" ? "first-task-final-run-ledger.json" : "chat-final-run-ledger.json", selectedRuns, secrets); } + if (execution.suite.id === "stock-harness") { + try { + const stock = await captureStockHarness({ api, companyId: fixtures.company.id, + agentId: fixtures.agent.id, generation: execution.profile.generation, + runIds: selectedRuns.map(run => run.id) }); + const checks = gradeStockHarness(stock); + matcherResults.push(...checks.map(check => ({ matcher: { kind: "json_path" as const, + path: `stockHarness.${check.id}`, expected: true }, passed: check.passed, detail: check.detail }))); + await writeSanitizedJson(snapshotsDir, "stock-harness.json", { ...stock, checks }, secrets); + const failures = checks.filter(check => !check.passed); + if (failures.length && !primaryError) primaryError = new Error(`Stock harness matcher failures: ${failures.map(check => check.id).join(", ")}`); + } catch (error) { + await writeSanitizedJson(snapshotsDir, "stock-harness-evidence-error.json", { + error: error instanceof Error ? error.message : String(error), + }, secrets); + if (!primaryError) primaryError = error; + } + } await fixtures.teardown(); cleanup = "passed"; } catch (error) { diff --git a/tests/runner-e2e/select-rerun-artifacts.test.ts b/tests/runner-e2e/select-rerun-artifacts.test.ts index b6793b2885..657fc2cea1 100644 --- a/tests/runner-e2e/select-rerun-artifacts.test.ts +++ b/tests/runner-e2e/select-rerun-artifacts.test.ts @@ -160,6 +160,19 @@ function singletonSelectionInput(paths: Awaited>) { } describe("runner E2E workflow rerun artifact selection", () => { + it("retains credential-free prerequisites inside the exact campaign root", async () => { + const paths = await fixture(); + const { artifactName, campaignName } = await addArtifact({ root: paths.artifactRoot, + executionId: RERUN, workflowAttempt: 2, status: "passed" }); + const prerequisite = path.join(paths.artifactRoot, artifactName, campaignName, + "stock-harness-prerequisites", "fixture", "preflight.json"); + await mkdir(path.dirname(prerequisite), { recursive: true }); + await writeFile(prerequisite, JSON.stringify({ passed: true, providerCalls: 0 })); + await selectRerunArtifacts(singletonSelectionInput(paths)); + expect(await readFile(path.join(paths.selectedRoot, artifactName, campaignName, + "stock-harness-prerequisites", "fixture", "preflight.json"), "utf8")) + .toContain('"providerCalls":0'); + }); it.each(["before", "after"])("ignores a queued placeholder %s a started replacement without hiding its failure", async order => { const paths = await fixture(); diff --git a/tests/runner-e2e/stock-harness-admission.test.ts b/tests/runner-e2e/stock-harness-admission.test.ts new file mode 100644 index 0000000000..20a07639ae --- /dev/null +++ b/tests/runner-e2e/stock-harness-admission.test.ts @@ -0,0 +1,44 @@ +import { beforeEach, describe, expect, it, vi } from "vitest"; +import { spawnSync } from "node:child_process"; +import { prepareStockHarnessPreflight, stockPreflightEnvironment, verifyStockHarnessPreflight } from "./stock-harness-admission.js"; +import { CREDENTIAL_NAMES } from "./types.js"; + +vi.mock("node:child_process", () => ({ spawnSync: vi.fn() })); +const spawn = vi.mocked(spawnSync); +const campaignDirectory = "/fixture/results/gha-1-1-stock-harness.fixture"; +beforeEach(() => { vi.resetAllMocks(); vi.unstubAllEnvs(); }); + +describe("stock harness credential-free admission", () => { + it("allows toolchain paths and excludes every present or future credential", () => { + const source = { PATH: "/bin", HOME: "/home/fixture", CARGO_HOME: "/cargo", CI: "true", + GH_TOKEN: "secret", FUTURE_PROVIDER_API_KEY: "secret", ...Object.fromEntries(CREDENTIAL_NAMES.map(name => [name, "secret"])) }; + expect(stockPreflightEnvironment(source)).toEqual({ PATH: "/bin", HOME: "/home/fixture", CARGO_HOME: "/cargo", CI: "true" }); + }); + it("rejects missing receipts without invoking a process", () => { + expect(() => verifyStockHarnessPreflight(undefined)).toThrow("no prerequisite receipt"); + expect(spawn).not.toHaveBeenCalled(); + }); + it("fails before provider execution when the prerequisite subprocess fails", () => { + spawn.mockReturnValue({ status: 1 } as ReturnType); + expect(() => prepareStockHarnessPreflight(campaignDirectory)).toThrow("before provider execution"); + expect(spawn).toHaveBeenCalledTimes(1); + }); + it("verifies retained evidence after a passing prerequisite subprocess", () => { + spawn.mockReturnValueOnce({ status: 0 } as ReturnType); + spawn.mockReturnValueOnce({ status: 0, stdout: '{"passed":true}' } as ReturnType); + const receipt = prepareStockHarnessPreflight(campaignDirectory); + expect(receipt).toMatch(/^\/fixture\/results\/gha-1-1-stock-harness\.fixture\/stock-harness-prerequisites\/[^/]+\/preflight.json$/); + expect(spawn.mock.calls[1]?.[1]).toContain(`--verify=${receipt}`); + }); + it("allows Cargo dependency resolution only on the disposable GitHub runner", () => { + vi.stubEnv("GITHUB_ACTIONS", "true"); + spawn.mockReturnValueOnce({ status: 0 } as ReturnType); + spawn.mockReturnValueOnce({ status: 0, stdout: '{}' } as ReturnType); + prepareStockHarnessPreflight(campaignDirectory); + expect(spawn.mock.calls[0]?.[1]).toContain("--allow-rust-network"); + }); + it("refuses failed receipt verification", () => { + spawn.mockReturnValue({ status: 1, stderr: "stale source" } as ReturnType); + expect(() => verifyStockHarnessPreflight("/fixture/preflight.json")).toThrow("stale source"); + }); +}); diff --git a/tests/runner-e2e/stock-harness-admission.ts b/tests/runner-e2e/stock-harness-admission.ts new file mode 100644 index 0000000000..c1a0e40761 --- /dev/null +++ b/tests/runner-e2e/stock-harness-admission.ts @@ -0,0 +1,38 @@ +import { spawnSync } from "node:child_process"; +import { randomUUID } from "node:crypto"; +import path from "node:path"; + +const root = path.resolve(import.meta.dirname, "../.."); +export const STOCK_PREFLIGHT_ENV = "PAPERCLIP_RUNNER_E2E_STOCK_PREFLIGHT"; + +// An allowlist keeps every provider credential, ambient auth override, and GH +// token out of the prerequisite subprocess, including secrets added later. +export function stockPreflightEnvironment(source: NodeJS.ProcessEnv): NodeJS.ProcessEnv { + return Object.fromEntries([ + "PATH", "HOME", "TMPDIR", "TMP", "TEMP", "SYSTEMROOT", "LANG", "LC_ALL", + "CARGO_HOME", "RUSTUP_HOME", "CI", "GITHUB_ACTIONS", + ].flatMap(name => source[name] === undefined ? [] : [[name, source[name]]])); +} + +export function verifyStockHarnessPreflight(receiptPath: string | undefined) { + if (!receiptPath) throw new Error("Stock harness has no prerequisite receipt; use the live launcher."); + const run = spawnSync(process.execPath, [path.join(root, "tests/runner-e2e/stock-harness-checks.mjs"), `--verify=${receiptPath}`], { + cwd: root, env: stockPreflightEnvironment(process.env), encoding: "utf8", timeout: 30_000, + }); + if (run.status !== 0) throw new Error(`Stock harness prerequisite verification failed: ${run.stderr || run.error?.message || run.status}`); + return JSON.parse(run.stdout) as Record; +} + +export function prepareStockHarnessPreflight(campaignDirectory: string) { + // The trusted artifact selector admits exactly one campaign root. Keep all + // prerequisite evidence inside that root, including pre-provider failures. + const output = path.join(campaignDirectory, "stock-harness-prerequisites", randomUUID()); + const receipt = path.join(output, "preflight.json"); + const run = spawnSync(process.execPath, [path.join(root, "tests/runner-e2e/stock-harness-checks.mjs"), + `--output-dir=${output}`, ...(process.env.GITHUB_ACTIONS === "true" ? ["--allow-rust-network"] : [])], { + cwd: root, env: stockPreflightEnvironment(process.env), stdio: "inherit", timeout: 50 * 60_000, + }); + if (run.status !== 0) throw new Error(`Stock harness prerequisites failed before provider execution. Retained receipt: ${receipt}`); + verifyStockHarnessPreflight(receipt); + return receipt; +} diff --git a/tests/runner-e2e/stock-harness-checks.mjs b/tests/runner-e2e/stock-harness-checks.mjs new file mode 100644 index 0000000000..949593873f --- /dev/null +++ b/tests/runner-e2e/stock-harness-checks.mjs @@ -0,0 +1,246 @@ +// Deterministic prerequisite for the stock-harness Product E2E suite. No model +// calls or credentials. Reuse the existing protocol assertions, not model +// self-reports about what its system instructions contain. +import { spawnSync } from "node:child_process"; +import { createHash } from "node:crypto"; +import { mkdirSync, readFileSync, writeFileSync } from "node:fs"; +import { join, resolve } from "node:path"; +import { readStockInstructionVariant } from "./stock-harness-instruction-variant.mjs"; + +const root = resolve(import.meta.dirname, "../.."); +export const stockHarnessGates = [ + { id: "SH-1", name: "Native Codex additive instructions", cwd: "packages/paperclip-runner", files: [ + "src/drivers/codex/codex-app-server-driver.test.ts", + "src/drivers/codex/codex-app-server-driver.lifecycle.test.ts", + "src/live/runnerd-codex-transport.test.ts", + "src/live/live-session.test.ts", + "src/cli/eval-provider-runtime.test.ts", + ], testPattern: "adds Paperclip developer instructions|preserves stock Codex instructions|captures exact provider frames and correlates|direct eval provider runtime|exposes native completion consistently on fresh and resumed sessions|passes caller-supplied native system instructions", + required: ["adds Paperclip developer instructions", "preserves stock Codex instructions on task recovery", + "preserves stock Codex instructions on prepared recovery", "preserves stock Codex instructions on direct recovery", + "captures exact provider frames and correlates", "exposes native completion consistently on fresh and resumed sessions", + "passes caller-supplied native system instructions"] }, + { id: "SH-2", name: "Production-default hire bundle", cwd: ".", files: [ + "server/src/__tests__/agent-skills-routes.test.ts", + "server/src/services/onboarding-first-task-assets.test.ts", + ], required: ["materializes minimal default instructions for non-CEO agents with no prompt template"] }, + { id: "SH-3", name: "Shared legacy startup and continuation", cwd: ".", files: [ + "packages/adapter-utils/src/server-utils.test.ts", + "packages/adapter-utils/src/prompt-sections.test.ts", + "packages/adapter-utils/src/acpx-engine/execute.test.ts", + "packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.test.ts", + "packages/adapters/pi-local/src/server/execute.remote.test.ts", + "packages/adapters/opencode-local/src/server/execute.test.ts", + "packages/adapters/cursor-cloud/src/server/execute.test.ts", + "server/src/__tests__/codex-local-execute.test.ts", + ], required: ["keeps task and chat defaults to identity and connection guidance", + "does not restore generic procedures on resume or with the legacy opt-in", + "integrates the env-free store", "advertises folded routing metadata", "bounds routing descriptions", + "loads an old env-bearing file with rotated credentials", "drops a skill that fails to materialize"] }, + // Hermes is not in the root Vitest project list. Run its package config so + // the requested file cannot silently disappear from discovery. + { id: "SH-3-hermes", name: "Hermes shared-prompt delivery", cwd: "packages/adapters/hermes", + files: ["src/server/prompt-rendering.test.ts"], + required: ["renders standard assignment wake with task authority", "renders scoped planning wake authority"] }, + { id: "SH-eval", name: "Independent oracle and qualification admission", cwd: ".", + config: "tests/runner-e2e/vitest.config.ts", + files: ["tests/runner-e2e/stock-harness-manifest.test.ts", "tests/runner-e2e/paperclip-document.test.ts", "tests/runner-e2e/stock-harness.test.ts", "tests/runner-e2e/stock-harness-checks.test.mjs", + "tests/runner-e2e/stock-harness-admission.test.ts", "tests/runner-e2e/stock-harness-digest.test.ts", + "tests/runner-e2e/select-rerun-artifacts.test.ts", "tests/runner-e2e/stock-harness-instruction-variant.test.mjs", "tests/runner-e2e/automatic-retry.test.ts"], + required: ["requires generated capability manifests for the current skill sources", "the shipped recipe delivers the current", "the shipped recipe supports an unnumbered issue", "rejects old SHA before providers", "allows toolchain paths and excludes every present or future credential", + "changes when the evaluated server/src/onboarding-assets/default/AGENTS.md changes", + "changes when the evaluated packages/adapter-utils/src/server-utils.ts changes", + "changes when the evaluated packages/shared/src/connection-intent-guidance.ts changes", + "retains credential-free prerequisites inside the exact campaign root"] }, +]; + +export function stockHarnessGatesForVariant(variant) { + if (variant !== "reduced" && variant !== "historical") throw new Error("Unknown stock instruction variant."); + return stockHarnessGates.map(gate => variant !== "historical" ? gate : { + ...gate, + required: gate.required.map(name => name === "materializes minimal default instructions for non-CEO agents with no prompt template" + ? "materializes the bundled default instruction set for non-CEO agents with no prompt template" + : name === "keeps task and chat defaults to identity and connection guidance" + ? "keeps the default local-agent prompt action-oriented" + : name === "does not restore generic procedures on resume or with the legacy opt-in" + ? "adds the execution contract to resume delta prompts and opted-in fresh prompts" : name), + }); +} + +export function gradeGate(gate, report, exitCode) { + const assertions = (report?.testResults ?? []).flatMap(file => file.assertionResults ?? []); + const requirements = (gate.required ?? []).map(name => ({ name, + passed: assertions.some(assertion => assertion.fullName?.includes(name) && assertion.status === "passed") })); + const files = gate.files.map(file => ({ file, + passed: (report?.testResults ?? []).some(result => result.name?.endsWith(file) && result.status === "passed" && + result.assertionResults?.some(assertion => assertion.status === "passed") && + result.assertionResults.every(assertion => assertion.status === "passed" || + (gate.testPattern && assertion.status === "skipped" && !(gate.required ?? []).some(name => assertion.fullName?.includes(name))))) })); + return { id: gate.id, name: gate.name, passed: exitCode === 0 && requirements.every(row => row.passed) && files.every(row => row.passed), + exitCode, files, requirements, total: report?.numTotalTests ?? 0, passedTests: report?.numPassedTests ?? 0, + failedTests: report?.numFailedTests ?? 0, pendingTests: report?.numPendingTests ?? 0 }; +} + +export function sourceFingerprint() { + const hash = createHash("sha256"); + const sources = new Set([ + ...stockHarnessGates.flatMap(gate => gate.files.map(file => join(gate.cwd, file))), + "tests/runner-e2e/stock-harness.ts", "tests/runner-e2e/stock-harness-checks.mjs", "tests/runner-e2e/catalog.ts", + "tests/runner-e2e/stock-harness-admission.ts", "tests/runner-e2e/launch.ts", "tests/runner-e2e/runner.spec.ts", + "tests/runner-e2e/stock-harness-instruction-variant.mjs", "tests/runner-e2e/stock-harness-instruction-variant.d.mts", "tests/runner-e2e/fixtures/stock-harness/historical-default-agents.md", + "tests/runner-e2e/automatic-retry.ts", "tests/runner-e2e/types.ts", + "packages/adapter-utils/src/acpx-engine/execute.ts", "packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.ts", + "tests/runner-e2e/context-integrity-cases.ts", "tests/runner-e2e/context-integrity-flow.ts", "tests/runner-e2e/context-integrity-scoring.ts", + "tests/runner-e2e/stock-harness-manifest.ts", "packages/paperclip-runner/scripts/generate-capability-contract.mjs", + "packages/paperclip-runner/spec/capability/source-contract.json", + "packages/paperclip-runner/scripts/check-capability-inventory.mjs", "packages/paperclip-runner/scripts/lib/capability-inventory.mjs", + "packages/paperclip-runner/spec/capability/capabilities.yaml", "packages/paperclip-runner/spec/capability/eval-traceability.yaml", + "packages/paperclip-runner/spec/capability/mcp-tool-map.yaml", "packages/paperclip-runner/spec/capability/inventory.schema.json", + "packages/paperclip-runner/src/generated/capability-contract.ts", "packages/paperclip-runner/docs/capability-contract.md", + ...["capabilities.yaml", "mcp-tool-map.yaml", "eval-traceability.yaml", "capability-contract.md", "downstream-handoff.md"] + .map(file => `packages/paperclip-runner/generated/capability/${file}`), + "packages/adapter-utils/src/server-utils.ts", "packages/shared/src/connection-intent-guidance.ts", + "server/src/onboarding-assets/default/AGENTS.md", "server/src/routes/agents.ts", "scripts/ensure-plugin-build-deps.mjs", + "packages/paperclip-runner/src/drivers/codex/codex-app-server-driver-impl.ts", + "packages/paperclip-runner/src/live/runnerd-codex-transport.ts", + "packages/paperclip-runner/src/live/live-session.ts", + "packages/paperclip-runner/runner/crates/runner-core/src/codex_provider.rs", + "packages/paperclip-runner/runner/crates/runner-core/src/bin/fake-codex-app-server.rs", + "packages/paperclip-runner/runner/Cargo.toml", "packages/paperclip-runner/runner/Cargo.lock", + "packages/paperclip-runner/runner/crates/runner-core/Cargo.toml", + "packages/paperclip-runner/runner/crates/runner-core/tests/codex_provider.rs", + ]); + const sourceErrors = []; + sources.add("skills/paperclip/SKILL.md"); + sources.add("skills/paperclip/references/issue-documents.md"); + for (const source of [...sources].sort()) { + hash.update(source); + try { hash.update("present\0").update(readFileSync(join(root, source))); } + catch (error) { + if (source === "skills/paperclip/references/issue-documents.md" && error.code === "ENOENT") hash.update("absent\0"); + else sourceErrors.push(source); + } + } + return { fingerprint: hash.digest("hex"), sourceErrors }; +} + +export function assertPreflightReceipt(report, current) { + const instructionVariant = readStockInstructionVariant(); + const expected = [...stockHarnessGates.map(gate => gate.id), "SH-1-rust"]; + if (report?.schema !== "paperclip.stock-harness-preflight.v3" || report.passed !== true || + report.setup?.passed !== true || report.setup?.exitCode !== 0 || + report.setup?.sdkExitCode !== 0 || report.setup?.runnerdExitCode !== 0 || + report.setup?.runnerdSha256 !== current.runnerdSha256 || + report.setup?.fakeCodexSha256 !== current.fakeCodexSha256 || + report.providerCalls !== 0 || report.sourceSha !== current.sha || + report.sourceFingerprint !== current.fingerprint || report.sourceErrors?.length !== 0 || + report.instructionVariant?.variant !== instructionVariant.variant || report.instructionVariant?.sha256 !== instructionVariant.sha256 || + !Array.isArray(report.gates) || report.gates.length !== expected.length || + expected.some(id => report.gates.filter(gate => gate.id === id && gate.passed === true && gate.exitCode === 0).length !== 1)) { + throw new Error("Stock harness requires passing prerequisites for this exact source SHA and fingerprint."); + } + return report; +} + +export function main(args = process.argv.slice(2)) { + const { variant, sha256 } = readStockInstructionVariant(); + const instructionVariant = { variant, sha256 }; + const gates = stockHarnessGatesForVariant(variant); + if (args.includes("--list")) { + console.log(JSON.stringify({ gates, instructionVariant, rust: "runtime_instructions_are_additive_for_codex_on_start_and_resume", live: "not invoked" }, null, 2)); + return; + } + if (args.some(arg => !arg.startsWith("--output-dir=") && !arg.startsWith("--verify=") && arg !== "--allow-rust-network")) + throw new Error("Use --list, --output-dir=, --verify=, or --allow-rust-network; never paid providers."); + const git = spawnSync("git", ["rev-parse", "HEAD"], { cwd: root, encoding: "utf8" }); + const verify = args.find(arg => arg.startsWith("--verify="))?.slice("--verify=".length); + if (verify) { + const source = sourceFingerprint(); + if (git.status !== 0 || source.sourceErrors.length) throw new Error("Cannot verify stock harness source provenance."); + const report = assertPreflightReceipt(JSON.parse(readFileSync(verify, "utf8")), { + sha: git.stdout.trim(), fingerprint: source.fingerprint, + runnerdSha256: createHash("sha256").update(readFileSync(join(root, + "packages/paperclip-runner/runner/target/debug", `paperclip-runnerd${process.platform === "win32" ? ".exe" : ""}`))).digest("hex"), + fakeCodexSha256: createHash("sha256").update(readFileSync(join(root, + "packages/paperclip-runner/runner/target/debug/fake-codex-app-server"))).digest("hex"), + }); + const output = resolve(verify, ".."); + for (const gate of gates) { + const grade = gradeGate(gate, JSON.parse(readFileSync(join(output, `${gate.id}.json`), "utf8")), 0); + if (!grade.passed) throw new Error(`Missing or failed retained prerequisite assertions: ${gate.id}`); + } + if (!/test runtime_instructions_are_additive_for_codex_on_start_and_resume \.\.\. ok/.test(readFileSync(join(output, "rust.txt"), "utf8"))) + throw new Error("Missing passing retained Rust prerequisite."); + console.log(JSON.stringify(report)); + return report; + } + const output = args.find(arg => arg.startsWith("--output-dir="))?.slice("--output-dir=".length) ?? + join(root, "tests/runner-e2e/results", `stock-harness-preflight-${new Date().toISOString().replaceAll(":", "-")}`); + mkdirSync(output, { recursive: true }); + const env = Object.fromEntries([ + "PATH", "HOME", "TMPDIR", "TMP", "TEMP", "SYSTEMROOT", "LANG", "LC_ALL", + "CARGO_HOME", "RUSTUP_HOME", "CI", "GITHUB_ACTIONS", + ].flatMap(name => process.env[name] === undefined ? [] : [[name, process.env[name]]])); + // A protected cold install disables lifecycle scripts. Use the same ordinary + // SDK dependency builder as server startup, before importing server tests. + const setupRun = spawnSync(process.execPath, [join(root, "scripts/ensure-plugin-build-deps.mjs")], { + cwd: root, env, encoding: "utf8", timeout: 5 * 60_000, + }); + writeFileSync(join(output, "setup.txt"), `${setupRun.stdout ?? ""}\n${setupRun.stderr ?? ""}`); + // The exact daemon-frame test uses the real local Rust daemon. Legacy cold + // cells do not download native artifacts, so compile it before TS discovery. + const runnerd = setupRun.status === 0 ? spawnSync("cargo", ["build", "--locked", + ...(args.includes("--allow-rust-network") ? [] : ["--offline"]), + "--workspace", "--bins"], { + cwd: join(root, "packages/paperclip-runner/runner"), env, encoding: "utf8", timeout: 10 * 60_000, + }) : null; + writeFileSync(join(output, "runnerd-build.txt"), `${runnerd?.stdout ?? ""}\n${runnerd?.stderr ?? ""}`); + const setup = { passed: setupRun.status === 0 && runnerd?.status === 0, + exitCode: setupRun.status !== 0 ? setupRun.status : runnerd?.status ?? null, + sdkExitCode: setupRun.status, runnerdExitCode: runnerd?.status ?? null }; + if (!setup.passed) { + const source = sourceFingerprint(); + writeFileSync(join(output, "preflight.json"), JSON.stringify({ + schema: "paperclip.stock-harness-preflight.v3", instructionVariant, sourceSha: git.stdout?.trim() || null, + sourceFingerprint: source.fingerprint, measuredAt: new Date().toISOString(), providerCalls: 0, + live: "not_run", passed: false, sourceErrors: source.sourceErrors, setup, gates: [], + }, null, 2) + "\n"); + process.exitCode = 1; + return; + } + const runnerdBinary = join(root, "packages/paperclip-runner/runner/target/debug", + `paperclip-runnerd${process.platform === "win32" ? ".exe" : ""}`); + setup.runnerdSha256 = createHash("sha256").update(readFileSync(runnerdBinary)).digest("hex"); + setup.fakeCodexSha256 = createHash("sha256").update(readFileSync(join(root, + "packages/paperclip-runner/runner/target/debug/fake-codex-app-server"))).digest("hex"); + const results = []; + for (const gate of gates) { + console.log(`Checking ${gate.id}: ${gate.name}`); + const file = join(output, `${gate.id}.json`); + const run = spawnSync(process.execPath, [join(root, "node_modules/vitest/vitest.mjs"), "run", ...gate.files, + ...(gate.config ? ["--config", gate.config] : []), + ...(gate.testPattern ? ["--testNamePattern", gate.testPattern] : []), + "--reporter=default", "--reporter=json", `--outputFile.json=${file}`], + { cwd: resolve(root, gate.cwd), + env: gate.id === "SH-1" ? { ...env, PAPERCLIP_STOCK_PREFLIGHT_RUNNERD: runnerdBinary } : env, + stdio: "inherit", timeout: 10 * 60_000 }); + let report; + try { report = JSON.parse(readFileSync(file, "utf8")); } catch { /* missing evidence fails closed below */ } + results.push(gradeGate(gate, report, run.status)); + } + const rust = spawnSync("cargo", ["test", ...(args.includes("--allow-rust-network") ? [] : ["--offline"]), "-p", "paperclip-runner-core", "--test", "codex_provider", + "runtime_instructions_are_additive_for_codex_on_start_and_resume", "--", "--exact"], + { cwd: join(root, "packages/paperclip-runner/runner"), env, encoding: "utf8", timeout: 10 * 60_000 }); + const rustOutput = `${rust.stdout ?? ""}\n${rust.stderr ?? ""}`; + writeFileSync(join(output, "rust.txt"), rustOutput); + results.push({ id: "SH-1-rust", passed: rust.status === 0 && /test runtime_instructions_are_additive_for_codex_on_start_and_resume \.\.\. ok/.test(rustOutput), exitCode: rust.status }); + const { fingerprint, sourceErrors } = sourceFingerprint(); + const report = { schema: "paperclip.stock-harness-preflight.v3", instructionVariant, sourceSha: git.stdout?.trim() || null, + sourceFingerprint: fingerprint, measuredAt: new Date().toISOString(), providerCalls: 0, + live: "not_run", passed: git.status === 0 && sourceErrors.length === 0 && results.every(row => row.passed), sourceErrors, setup, gates: results }; + writeFileSync(join(output, "preflight.json"), JSON.stringify(report, null, 2) + "\n"); + console.log(`Stock harness prerequisite evidence: ${output}/preflight.json`); + if (!report.passed) process.exitCode = 1; +} + +if (process.argv[1] && resolve(process.argv[1]) === resolve(import.meta.filename)) main(); diff --git a/tests/runner-e2e/stock-harness-checks.test.mjs b/tests/runner-e2e/stock-harness-checks.test.mjs new file mode 100644 index 0000000000..e4a4139455 --- /dev/null +++ b/tests/runner-e2e/stock-harness-checks.test.mjs @@ -0,0 +1,79 @@ +import { describe, expect, it } from "vitest"; +import { assertPreflightReceipt, gradeGate, stockHarnessGates, stockHarnessGatesForVariant } from "./stock-harness-checks.mjs"; +import { readStockInstructionVariant } from "./stock-harness-instruction-variant.mjs"; + +const gate = { id: "SH-test", name: "Boundary", files: ["boundary.test.ts"], required: ["must preserve instructions"] }; +const report = () => ({ testResults: [{ name: "/repo/boundary.test.ts", status: "passed", + assertionResults: [{ fullName: "Boundary must preserve instructions", status: "passed" }] }], +numTotalTests: 1, numPassedTests: 1, numFailedTests: 0, numPendingTests: 0 }); +describe("stock harness prerequisite coverage", () => { + it("maps every implemented change to an executable gate", () => { + expect(stockHarnessGates.map(gate => gate.id)).toEqual(["SH-1", "SH-2", "SH-3", "SH-3-hermes", "SH-eval"]); + expect(stockHarnessGates.every(gate => gate.files.length > 0 && gate.required.length > 0)).toBe(true); + }); + it("accepts an executed passing boundary", () => expect(gradeGate(gate, report(), 0).passed).toBe(true)); + it("preserves candidate assertions and admits only declared historical structural alternatives", () => { + expect(stockHarnessGatesForVariant("reduced")).toEqual(stockHarnessGates); + const historical = stockHarnessGatesForVariant("historical"); + expect(historical.map(value => value.files)).toEqual(stockHarnessGates.map(value => value.files)); + expect(historical.find(value => value.id === "SH-2").required).toContain("materializes the bundled default instruction set for non-CEO agents with no prompt template"); + expect(historical.find(value => value.id === "SH-3").required).toContain("adds the execution contract to resume delta prompts and opted-in fresh prompts"); + expect(historical.find(value => value.id === "SH-3").required).toContain("integrates the env-free store"); + expect(() => stockHarnessGatesForVariant("other")).toThrow("Unknown"); + }); + it.each([null, undefined, { testResults: [] }])("rejects unavailable evidence %s", (report) => { + expect(gradeGate(gate, report, 0).passed).toBe(false); + }); + it.each(["pending", "skipped", "failed", "todo"])("rejects a %s required assertion", (status) => { + const observation = report(); + observation.testResults[0].assertionResults[0].status = status; + expect(gradeGate(gate, observation, 0).passed).toBe(false); + }); + it("rejects a passing different assertion and a nonzero launcher exit", () => { + const different = report(); + different.testResults[0].assertionResults[0].fullName = "Unrelated pass"; + expect(gradeGate(gate, different, 0).passed).toBe(false); + expect(gradeGate(gate, report(), 1).passed).toBe(false); + }); + it("allows explicitly filtered unrelated tests while rejecting a skipped required test", () => { + const filtered = { ...gate, testPattern: "must preserve instructions" }; + const observation = report(); + observation.testResults[0].assertionResults.push({ fullName: "Unrelated", status: "skipped" }); + expect(gradeGate(filtered, observation, 0).passed).toBe(true); + observation.testResults[0].assertionResults[0].status = "skipped"; + expect(gradeGate(filtered, observation, 0).passed).toBe(false); + }); + it("rejects a requested file absent from Vitest discovery even when required names pass", () => { + expect(gradeGate({ ...gate, files: [...gate.files, "undiscovered.test.ts"] }, report(), 0).passed).toBe(false); + expect(stockHarnessGates.find(gate => gate.id === "SH-3-hermes").cwd).toBe("packages/adapters/hermes"); + }); +}); + +describe("stock harness prerequisite admission", () => { + const current = { sha: "a".repeat(40), fingerprint: "b".repeat(64), runnerdSha256: "c".repeat(64), fakeCodexSha256: "e".repeat(64) }; + const receipt = () => ({ schema: "paperclip.stock-harness-preflight.v3", passed: true, + instructionVariant: (({ variant, sha256 }) => ({ variant, sha256 }))(readStockInstructionVariant()), + setup: { passed: true, exitCode: 0, sdkExitCode: 0, runnerdExitCode: 0, runnerdSha256: current.runnerdSha256, + fakeCodexSha256: current.fakeCodexSha256 }, + providerCalls: 0, sourceSha: current.sha, sourceFingerprint: current.fingerprint, + sourceErrors: [], gates: [...stockHarnessGates.map(g => g.id), "SH-1-rust"].map(id => ({ id, passed: true, exitCode: 0 })) }); + it("admits the same passing source revision", () => expect(assertPreflightReceipt(receipt(), current).passed).toBe(true)); + it.each([ + ["old SHA", r => { r.sourceSha = "c".repeat(40); }], + ["changed source", r => { r.sourceFingerprint = "d".repeat(64); }], + ["wrong instruction variant", r => { r.instructionVariant.variant = "other"; }], + ["wrong manual hash", r => { r.instructionVariant.sha256 = "f".repeat(64); }], + ["failed boundary", r => { r.gates[0].passed = false; }], + ["nonzero exit", r => { r.gates[0].exitCode = 1; }], + ["missing Rust", r => { r.gates.pop(); }], + ["duplicate boundary", r => { r.gates[1] = r.gates[0]; }], + ["provider calls", r => { r.providerCalls = 1; }], + ["missing source", r => { r.sourceErrors.push("missing.ts"); }], + ["failed cold setup", r => { r.setup.passed = false; r.setup.exitCode = 1; }], + ["missing daemon build", r => { delete r.setup.runnerdExitCode; }], + ["changed daemon binary", r => { r.setup.runnerdSha256 = "d".repeat(64); }], + ["changed protocol fixture", r => { r.setup.fakeCodexSha256 = "d".repeat(64); }], + ])("rejects %s before providers", (_name, mutate) => { + const r = receipt(); mutate(r); expect(() => assertPreflightReceipt(r, current)).toThrow("exact source SHA and fingerprint"); + }); +}); diff --git a/tests/runner-e2e/stock-harness-digest.test.ts b/tests/runner-e2e/stock-harness-digest.test.ts new file mode 100644 index 0000000000..6fbfd3597a --- /dev/null +++ b/tests/runner-e2e/stock-harness-digest.test.ts @@ -0,0 +1,27 @@ +import { describe, expect, it, vi } from "vitest"; +import { readFileSync } from "node:fs"; +import { stockHarnessSourceDigest, stockHarnessSkillSources } from "./stock-harness.js"; + +vi.mock("node:fs", async importOriginal => ({ ...await importOriginal(), readFileSync: vi.fn() })); + +describe("stock harness instruction revision", () => { + it.each(["server/src/onboarding-assets/default/AGENTS.md", "packages/adapter-utils/src/server-utils.ts", "packages/shared/src/connection-intent-guidance.ts", "skills/paperclip/SKILL.md", "skills/paperclip/references/issue-documents.md", "packages/paperclip-runner/generated/capability/capabilities.yaml", "packages/paperclip-runner/spec/capability/capabilities.yaml", "tests/runner-e2e/stock-harness-manifest.ts", "packages/adapter-utils/src/acpx-engine/execute.ts", "packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.ts", "tests/runner-e2e/stock-harness-instruction-variant.mjs", "tests/runner-e2e/automatic-retry.ts", "tests/runner-e2e/catalog.ts"])( + "changes when the evaluated %s changes", source => { + vi.mocked(readFileSync).mockImplementation(() => Buffer.from("unchanged")); + const original = stockHarnessSourceDigest(); + vi.mocked(readFileSync).mockImplementation(file => Buffer.from(String(file).endsWith(source) ? "changed instructions" : "unchanged")); + expect(stockHarnessSourceDigest()).not.toBe(original); + }); + it("records an absent historical recipe without introducing its content or hiding other read errors", () => { + vi.mocked(readFileSync).mockImplementation(() => Buffer.from("unchanged")); + const present = stockHarnessSourceDigest(); + vi.mocked(readFileSync).mockImplementation(file => { + if (String(file).endsWith("references/issue-documents.md")) throw Object.assign(new Error("absent"), { code: "ENOENT" }); + return Buffer.from("unchanged"); + }); + expect(stockHarnessSkillSources()[1]).toEqual({ path: "skills/paperclip/references/issue-documents.md", present: false, sha256: null }); + expect(stockHarnessSourceDigest()).not.toBe(present); + vi.mocked(readFileSync).mockImplementation(() => { throw Object.assign(new Error("unreadable"), { code: "EACCES" }); }); + expect(stockHarnessSourceDigest).toThrow("unreadable"); + }); +}); diff --git a/tests/runner-e2e/stock-harness-instruction-variant.d.mts b/tests/runner-e2e/stock-harness-instruction-variant.d.mts new file mode 100644 index 0000000000..3c8a570dfb --- /dev/null +++ b/tests/runner-e2e/stock-harness-instruction-variant.d.mts @@ -0,0 +1,7 @@ +export interface StockInstructionVariant { + variant: "reduced" | "historical"; + sha256: string; + content: string; +} +export function classifyStockInstructionManual(content: unknown): StockInstructionVariant; +export function readStockInstructionVariant(): StockInstructionVariant; diff --git a/tests/runner-e2e/stock-harness-instruction-variant.mjs b/tests/runner-e2e/stock-harness-instruction-variant.mjs new file mode 100644 index 0000000000..ec2bed7072 --- /dev/null +++ b/tests/runner-e2e/stock-harness-instruction-variant.mjs @@ -0,0 +1,24 @@ +import { createHash } from "node:crypto"; +import { readFileSync } from "node:fs"; + +const reduced = "You are an agent in a Paperclip company.\n"; +const historicalSha256 = "e4d2375d722602cd744403292d99f6e6c9b7e8014d9cdfa6f9811ace03428e5f"; + +/** + * @param {unknown} content + * @returns {{ variant: 'reduced' | 'historical', sha256: string, content: string }} + */ +export function classifyStockInstructionManual(content) { + if (typeof content !== "string") throw new Error("Stock harness instruction manual must be exact UTF-8 text."); + const sha256 = createHash("sha256").update(content, "utf8").digest("hex"); + if (content === reduced) return { variant: "reduced", sha256, content }; + if (Buffer.byteLength(content, "utf8") === 4249 && sha256 === historicalSha256) + return { variant: "historical", sha256, content }; + throw new Error("Stock harness instruction manual is not an admitted reduced or historical source."); +} + +export function readStockInstructionVariant() { + return classifyStockInstructionManual(readFileSync(new URL( + "../../server/src/onboarding-assets/default/AGENTS.md", import.meta.url, + ), "utf8")); +} diff --git a/tests/runner-e2e/stock-harness-instruction-variant.test.mjs b/tests/runner-e2e/stock-harness-instruction-variant.test.mjs new file mode 100644 index 0000000000..10f027d8ff --- /dev/null +++ b/tests/runner-e2e/stock-harness-instruction-variant.test.mjs @@ -0,0 +1,48 @@ +import { createHash } from "node:crypto"; +import { readFileSync } from "node:fs"; +import { afterEach, describe, expect, it, vi } from "vitest"; +import { classifyStockInstructionManual, readStockInstructionVariant } from "./stock-harness-instruction-variant.mjs"; + +const reduced = "You are an agent in a Paperclip company.\n"; +const historicalBytes = readFileSync(new URL("./fixtures/stock-harness/historical-default-agents.md", import.meta.url)); +const historical = historicalBytes.toString("utf8"); +const historicalSha256 = "e4d2375d722602cd744403292d99f6e6c9b7e8014d9cdfa6f9811ace03428e5f"; + +afterEach(() => vi.unstubAllEnvs()); +describe("closed stock harness instruction variants", () => { + it("admits only the exact reduced manual text", () => { + expect(classifyStockInstructionManual(reduced)).toEqual({ variant: "reduced", content: reduced, + sha256: createHash("sha256").update(reduced).digest("hex") }); + }); + it("admits the independently pinned 4249-byte historical manual", () => { + expect(historicalBytes.length).toBe(4249); + expect(createHash("sha256").update(historicalBytes).digest("hex")).toBe(historicalSha256); + expect(classifyStockInstructionManual(historical)).toEqual({ variant: "historical", sha256: historicalSha256, content: historical }); + }); + it.each([ + ["missing newline", reduced.trimEnd()], + ["appended reduced instructions", `${reduced}Follow additional procedures.\n`], + ["altered reduced wording", reduced.replace("agent", "coder")], + ["CRLF reduced manual", reduced.replace("\n", "\r\n")], + ["leading byte order mark", `\uFEFF${reduced}`], + ["appended historical instructions", `${historical}Extra instructions.\n`], + ["same-size altered historical content", `X${historical.slice(1)}`], + ["unknown manual", "Historical instructions"], + ["empty text", ""], + ])("rejects %s", (_name, content) => { + expect(() => classifyStockInstructionManual(content)).toThrow("not an admitted"); + }); + it.each([undefined, null, 42, {}, [], Buffer.from(reduced)])("rejects malformed non-string content %#", content => { + expect(() => classifyStockInstructionManual(content)).toThrow("must be exact UTF-8 text"); + }); + it("reads and classifies the exact current repository source without selecting by environment or branch labels", () => { + const source = readFileSync(new URL("../../server/src/onboarding-assets/default/AGENTS.md", import.meta.url), "utf8"); + const expected = classifyStockInstructionManual(source); + for (const label of ["historical", "reduced", "unknown"]) { + vi.stubEnv("PAPERCLIP_STOCK_HARNESS_VARIANT", label); + vi.stubEnv("PAPERCLIP_RUNNER_E2E_SOURCE_REF", `refs/heads/${label}`); + vi.stubEnv("GITHUB_HEAD_REF", label); + expect(readStockInstructionVariant()).toEqual(expected); + } + }); +}); diff --git a/tests/runner-e2e/stock-harness-manifest.test.ts b/tests/runner-e2e/stock-harness-manifest.test.ts new file mode 100644 index 0000000000..e8ce32b645 --- /dev/null +++ b/tests/runner-e2e/stock-harness-manifest.test.ts @@ -0,0 +1,21 @@ +import { describe, expect, it } from "vitest"; +import { readFileSync } from "node:fs"; +import { verifyCapabilityManifest } from "./stock-harness-manifest.js"; + +describe("stock harness generated instruction metadata", () => { + it("requires generated capability manifests for the current skill sources", async () => { + expect(await verifyCapabilityManifest()).toBe(true); + }); + it("rejects a stale manifest before provider admission", async () => { + await expect(verifyCapabilityManifest(path => path.endsWith("capabilities.yaml") + ? "stale skill anchors" : readFileSync(path, "utf8"))).rejects.toThrow("manifest drift"); + }); + it("rejects unavailable generated evidence", async () => { + await expect(verifyCapabilityManifest(() => { throw new Error("missing generated evidence"); })) + .rejects.toThrow("missing generated evidence"); + }); + it("rejects a stale source inventory before provider admission", async () => { + await expect(verifyCapabilityManifest(undefined, () => { throw new Error("source inventory drift"); })) + .rejects.toThrow("source inventory drift"); + }); +}); diff --git a/tests/runner-e2e/stock-harness-manifest.ts b/tests/runner-e2e/stock-harness-manifest.ts new file mode 100644 index 0000000000..dbc260923b --- /dev/null +++ b/tests/runner-e2e/stock-harness-manifest.ts @@ -0,0 +1,19 @@ +import { readFileSync } from "node:fs"; +import { execFileSync } from "node:child_process"; +import { fileURLToPath } from "node:url"; +import { buildContract } from "../../packages/paperclip-runner/scripts/generate-capability-contract.mjs"; + +/** Canonical derived metadata must match the instruction sources before providers. */ +export async function verifyCapabilityManifest( + read = (path: string) => readFileSync(path, "utf8"), + checkInventory = () => execFileSync(process.execPath, [fileURLToPath(new URL( + "../../packages/paperclip-runner/scripts/check-capability-inventory.mjs", import.meta.url, + ))], { stdio: "pipe" }), +) { + const expected = await buildContract(); + for (const [path, content] of Object.entries(expected)) { + if (read(path) !== content) throw new Error(`Generated capability manifest drift: ${path}`); + } + checkInventory(); + return true; +} diff --git a/tests/runner-e2e/stock-harness.test.ts b/tests/runner-e2e/stock-harness.test.ts new file mode 100644 index 0000000000..b8ef4087c0 --- /dev/null +++ b/tests/runner-e2e/stock-harness.test.ts @@ -0,0 +1,189 @@ +import { describe, expect, it } from "vitest"; +import { runnerMatrix, runnerProfiles, suiteDefinitionHash } from "./catalog.js"; +import { captureStockHarness, gradeStockHarness as gradeSourceHarness, gradeStockHire as gradeSourceHire, STOCK_HIRE_IDENTITY, type StockHarnessEvidence } from "./stock-harness.js"; +import { readFileSync } from "node:fs"; +import { classifyStockInstructionManual, readStockInstructionVariant } from "./stock-harness-instruction-variant.mjs"; + +const reduced = classifyStockInstructionManual(STOCK_HIRE_IDENTITY); +const gradeStockHire = (evidence: Parameters[0]) => gradeSourceHire(evidence, reduced); +const gradeStockHarness = (evidence: StockHarnessEvidence) => gradeSourceHarness(evidence, reduced); + +function recording(generation: "legacy" | "native" = "legacy"): StockHarnessEvidence { + return { + schema: "paperclip.stock-harness.v1", generation, agentId: "agent", + budgets: { companyMonthlyCents: 1_000, agentMonthlyCents: 1_000 }, + bundle: { entryFile: "AGENTS.md", files: [{ path: "AGENTS.md", content: STOCK_HIRE_IDENTITY }] }, + runIds: ["fresh", "resumed"], + invocations: generation === "native" ? [] : [ + { runId: "fresh", conversationMode: false, prompt: "You are agent agent (QA).\nConnection tools:\nUse connections_search.\nCurrent assignment.", promptMetrics: { heartbeatPromptChars: 113 } }, + { runId: "resumed", conversationMode: false, prompt: "## Paperclip Resume Delta\nCurrent ordered comments.", promptMetrics: { heartbeatPromptChars: 0 } }, + ], + }; +} + +describe("stock harness Product E2E", () => { + it("registers 26 explicit local cells with the original cases and two focused delivery repair cells", () => { + const cells = runnerMatrix.filter(row => row.suite.id === "stock-harness"); + expect(cells).toHaveLength(26); + expect(new Set(cells.map(row => row.profile.id))).toEqual(new Set([ + "legacy-codex", "legacy-claude", "legacy-opencode", "legacy-acp-codex", "legacy-acp-claude", + "runner-codex", "runner-acpx-claude", "runner-opencode", + ])); + expect(new Set(cells.map(row => row.task.id))).toEqual(new Set([ + "ordered-comment-continuation", "assigned-skill-explicit-invocation", "continuity-restart", + "assigned-skill-paperclip-document", + ])); + expect(cells.every(row => row.suite.manualOnly && row.environment.id === "local")).toBe(true); + expect(cells.every(row => row.task.automaticRetryPolicy === "single_attempt")).toBe(true); + expect(cells[0]!.suite.definitionMetadata).toMatchObject({ automaticRetryPolicy: "single_attempt", maximumAttemptsPerCell: 1 }); + expect(cells.reduce((turns, row) => turns + row.task.expectedRunCount, 0)).toBe(50); + expect(cells.filter(row => row.task.id === "assigned-skill-paperclip-document").map(row => row.profile.id).sort()).toEqual(["legacy-claude", "legacy-opencode"]); + expect(cells[0]!.suite.definitionMetadata?.sourceDigest).toMatch(/^[a-f0-9]{64}$/); + }); + + it.each(["legacy-codex", "legacy-claude", "legacy-opencode", "runner-codex", "runner-acpx-claude", "runner-opencode"])( + "omits the custom QA bundle for %s while preserving runtime, auth and skills configuration", (profileId) => { + const profile = runnerMatrix.find(row => row.suite.id === "stock-harness" && row.profile.id === profileId)!.profile; + const original = runnerProfiles.find(row => row.id === profileId)!; + const input = { executionId: "test", workspacePath: "/workspace", environmentId: "environment", environmentFixtureId: "local" as const, + secretRefs: { [profile.credential]: { type: "secret_ref" as const, secretId: "secret", version: "latest" as const } } }; + const expected = original.buildAgent(input); + const { instructionsBundle: _manual, ...withoutManual } = expected; + expect(profile.buildAgent(input)).toEqual(withoutManual); + expect(profile.buildAgent(input)).not.toHaveProperty("instructionsBundle"); + }, + ); + + it.each(["legacy", "native"] as const)("accepts complete %s evidence without interpreting task success as vendor-base proof", (generation) => { + expect(gradeStockHarness(recording(generation)).every(check => check.passed)).toBe(true); + expect(gradeStockHarness(recording(generation)).some(check => /vendor|base/.test(check.id))).toBe(false); + }); + + it("checks the hire before paid execution without treating that receipt as provider evidence", () => { + const hire = recording(); + hire.runIds = []; + hire.invocations = []; + expect(gradeStockHire(hire).every(check => check.passed)).toBe(true); + expect(gradeStockHarness(hire)).toContainEqual(expect.objectContaining({ id: "provider-runs-present", passed: false })); + hire.bundle.files[0]!.content += "Old operating manual."; + expect(gradeStockHire(hire)).toContainEqual(expect.objectContaining({ id: "default-hire-bundle", passed: false })); + }); + + it("keeps historical instruction observations separate from the unchanged behavioral oracle", () => { + const historical = classifyStockInstructionManual(readFileSync(new URL("fixtures/stock-harness/historical-default-agents.md", import.meta.url), "utf8")); + const evidence = recording(); + evidence.bundle.files[0]!.content = historical.content; + evidence.invocations[0]!.prompt += "\nExecution contract:\nFinal disposition checklist:"; + evidence.invocations[1]!.prompt += "\nExecution contract: take concrete action\na successful process exit or final response is not sufficient"; + expect(gradeSourceHarness(evidence, historical).every(check => check.passed)).toBe(true); + expect(gradeSourceHarness(evidence, historical)).toContainEqual(expect.objectContaining({ id: "historical-continuation-contract", passed: true })); + expect(gradeSourceHarness(evidence, reduced)).toContainEqual(expect.objectContaining({ id: "default-hire-bundle", passed: false })); + evidence.invocations[0]!.prompt = "You are agent agent. Connection tools: connections_search"; + expect(gradeSourceHarness(evidence, historical)).toContainEqual(expect.objectContaining({ id: "historical-startup-contract", passed: false })); + const current = readStockInstructionVariant(); + evidence.bundle.files[0]!.content = current.content; + expect(gradeSourceHire(evidence).every(check => check.passed)).toBe(true); + }); + + it("cannot use valid historical startup to conceal a missing task continuation or misclassified mode", () => { + const historical = classifyStockInstructionManual(readFileSync(new URL("fixtures/stock-harness/historical-default-agents.md", import.meta.url), "utf8")); + const evidence = recording(); + evidence.bundle.files[0]!.content = historical.content; + evidence.invocations[0]!.prompt += "\nExecution contract:\nFinal disposition checklist:"; + expect(gradeSourceHarness(evidence, historical)).toContainEqual(expect.objectContaining({ id: "historical-startup-contract", passed: true })); + expect(gradeSourceHarness(evidence, historical)).toContainEqual(expect.objectContaining({ id: "historical-continuation-contract", passed: false })); + evidence.invocations[1]!.prompt += "\nExecution contract: take concrete action\na successful process exit or final response is not sufficient"; + delete evidence.invocations[1]!.conversationMode; + expect(gradeSourceHarness(evidence, historical)).toContainEqual(expect.objectContaining({ id: "historical-invocation-classification", passed: false })); + }); + + it("checks historical conversation startup and resume separately without injecting the task contract", () => { + const historical = classifyStockInstructionManual(readFileSync(new URL("fixtures/stock-harness/historical-default-agents.md", import.meta.url), "utf8")); + const evidence = recording(); + evidence.bundle.files[0]!.content = historical.content; + for (const row of evidence.invocations) row.conversationMode = true; + evidence.invocations[0]!.prompt += "\nContinue your Paperclip conversation using the supplied chat mode directive.\nAfter 2 consecutive failures of the same control-plane write"; + expect(gradeSourceHarness(evidence, historical).every(check => check.passed)).toBe(true); + evidence.invocations[1]!.prompt += "\nExecution contract:"; + expect(gradeSourceHarness(evidence, historical)).toContainEqual(expect.objectContaining({ id: "historical-continuation-contract", passed: false })); + }); + + it.each([ + ["manual-regrew", (e: StockHarnessEvidence) => { e.bundle.files[0]!.content += "Always comment and delegate."; }, "default-hire-bundle"], + ["extra-file", (e: StockHarnessEvidence) => { e.bundle.files.push({ path: "HEARTBEAT.md", content: "Do everything again." }); }, "default-hire-bundle"], + ["wrong-entry", (e: StockHarnessEvidence) => { e.bundle.entryFile = "OTHER.md"; }, "default-hire-bundle"], + ["missing-runs", (e: StockHarnessEvidence) => { e.runIds = []; }, "provider-runs-present"], + ["missing-invocation", (e: StockHarnessEvidence) => { e.invocations.pop(); }, "invocation-evidence-complete"], + ["empty-prompt", (e: StockHarnessEvidence) => { e.invocations[1]!.prompt = ""; }, "invocation-evidence-complete"], + ["malformed-prompt", (e: StockHarnessEvidence) => { e.invocations[1]!.prompt = { text: "hidden" }; }, "invocation-evidence-complete"], + ["old-startup", (e: StockHarnessEvidence) => { e.invocations[0]!.prompt += "\nExecution contract:"; }, "generic-procedures-absent"], + ["old-resume", (e: StockHarnessEvidence) => { e.invocations[1]!.prompt += "\na successful process exit or final response is not sufficient"; }, "generic-procedures-absent"], + ["missing-connections", (e: StockHarnessEvidence) => { e.invocations[0]!.prompt = "You are agent agent."; }, "fresh-default-delivered"], + ["missing-fresh-evidence", (e: StockHarnessEvidence) => { e.invocations[0]!.promptMetrics = {}; }, "fresh-default-delivered"], + ["wrong-budget", (e: StockHarnessEvidence) => { e.budgets.agentMonthlyCents = 0; }, "budget-hard-stops"], + ] as const)("fails closed for %s", (_name, mutate, failedId) => { + const evidence = recording(); + mutate(evidence); + expect(gradeStockHarness(evidence)).toContainEqual(expect.objectContaining({ id: failedId, passed: false })); + }); + + it("captures actual public bundle and invocation receipts, ignoring unrelated events", async () => { + const paths: string[] = []; + const api = { async get(url: string): Promise { + paths.push(url); + if (url.endsWith("/instructions-bundle")) return { entryFile: "AGENTS.md", files: [{ path: "AGENTS.md" }] } as T; + if (url.includes("/file?")) return { content: STOCK_HIRE_IDENTITY } as T; + if (url.includes("/events?")) return [ + { eventType: "adapter.invoke", payload: { prompt: "Actual prompt", promptMetrics: { heartbeatPromptChars: 10 } } }, + { eventType: "assistant", payload: { prompt: "Model claims about instructions" } }, + ] as T; + if (url.endsWith("/heartbeat-runs/run")) return { id: "run", companyId: "company", agentId: "agent", contextSnapshot: { conversationMode: false } } as T; + return { budgetMonthlyCents: 1_000 } as T; + } }; + const evidence = await captureStockHarness({ api, companyId: "company", agentId: "agent", generation: "legacy", runIds: ["run"] }); + expect(evidence.bundle.files[0]?.content).toBe(STOCK_HIRE_IDENTITY); + expect(evidence.invocations).toHaveLength(1); + expect(evidence.invocations[0]).toMatchObject({ runId: "run", prompt: "Actual prompt" }); + expect(paths).toContain("/api/heartbeat-runs/run/events?limit=1000"); + }); + + it("does not hide evidence API failures", async () => { + const api = { async get(): Promise { throw new Error("Public API unavailable"); } }; + await expect(captureStockHarness({ api, companyId: "company", agentId: "agent", generation: "legacy", runIds: ["run"] })) + .rejects.toThrow("Public API unavailable"); + }); + + it("normalizes an omitted conversationMode only from a complete scoped task snapshot", async () => { + const api = { async get(url: string): Promise { + return (url.endsWith("/instructions-bundle") ? { entryFile: "AGENTS.md", files: [] } + : url.endsWith("/heartbeat-runs/run") ? { id: "run", companyId: "company", agentId: "agent", contextSnapshot: { taskId: "task", issueId: "task" } } + : url.includes("/events?") ? [{ eventType: "adapter.invoke", payload: { prompt: "Task context", promptMetrics: { heartbeatPromptChars: 1 } } }] + : { budgetMonthlyCents: 1000 }) as T; + } }; + const result = await captureStockHarness({ api, companyId: "company", agentId: "agent", generation: "legacy", runIds: ["run"] }); + expect(result.invocations[0]?.conversationMode).toBe(false); + }); + + it.each([ + { id: "other", companyId: "company", agentId: "agent", contextSnapshot: {} }, + { id: "run", companyId: "other", agentId: "agent", contextSnapshot: {} }, + { id: "run", companyId: "company", agentId: "other", contextSnapshot: {} }, + { id: "run", companyId: "company", agentId: "agent" }, + { id: "run", companyId: "company", agentId: "agent", contextSnapshot: "invalid" }, + { id: "run", companyId: "company", agentId: "agent", contextSnapshot: [] }, + { id: "run", companyId: "company", agentId: "agent", contextSnapshot: { conversationMode: "unknown" } }, + ])("rejects missing or mismatched public invocation mode attribution %#", async run => { + const api = { async get(url: string): Promise { + return (url.endsWith("/instructions-bundle") ? { entryFile: "AGENTS.md", files: [] } + : url.endsWith("/heartbeat-runs/run") ? run : url.includes("/events?") ? [] : { budgetMonthlyCents: 1000 }) as T; + } }; + await expect(captureStockHarness({ api, companyId: "company", agentId: "agent", generation: "legacy", runIds: ["run"] })) + .rejects.toThrow("mode receipt"); + }); + + it("changes the definition hash when the grader digest changes", () => { + const suite = runnerMatrix.find(row => row.suite.id === "stock-harness")!.suite; + expect(suiteDefinitionHash({ ...suite, definitionMetadata: { ...suite.definitionMetadata, sourceDigest: "changed" } })) + .not.toBe(suiteDefinitionHash(suite)); + }); +}); diff --git a/tests/runner-e2e/stock-harness.ts b/tests/runner-e2e/stock-harness.ts new file mode 100644 index 0000000000..7e7fd1dac1 --- /dev/null +++ b/tests/runner-e2e/stock-harness.ts @@ -0,0 +1,184 @@ +import { createHash } from "node:crypto"; +import { readFileSync } from "node:fs"; +import type { RunnerProfileFixture } from "./types.js"; +import { readStockInstructionVariant } from "./stock-harness-instruction-variant.mjs"; + +// These are an independent product contract, not imports of the implementation +// constants. Otherwise a larger shipped manual could silently update the oracle. +export const STOCK_HIRE_IDENTITY = "You are an agent in a Paperclip company.\n"; +export const STOCK_TEMPLATE_IDENTITY = "You are agent"; +const SKILL_SOURCES = ["../../skills/paperclip/SKILL.md", "../../skills/paperclip/references/issue-documents.md"]; +export function stockHarnessSkillSources() { + return SKILL_SOURCES.map(source => { + try { + return { path: source.replace(/^\.\.\/\.\.\//, ""), present: true, + sha256: createHash("sha256").update(readFileSync(new URL(source, import.meta.url))).digest("hex") }; + } catch (error) { + if (source.endsWith("/references/issue-documents.md") && (error as NodeJS.ErrnoException).code === "ENOENT") + return { path: source.replace(/^\.\.\/\.\.\//, ""), present: false, sha256: null }; + throw error; + } + }); +} +const REMOVED_PROCEDURES = [ + "Execution contract:", + "Start actionable work in this heartbeat", + "clear final disposition", + "After 2 consecutive failures of the same control-plane write", + "Use child issues for parallel or long delegated work", + "a successful process exit or final response is not sufficient", +]; + +export function productionDefaultHireProfile(profile: RunnerProfileFixture): RunnerProfileFixture { + return { + ...profile, + buildAgent(input) { + const { instructionsBundle: _fixtureManual, ...agent } = profile.buildAgent(input); + return agent; + }, + }; +} + +export interface StockHarnessEvidence { + schema: "paperclip.stock-harness.v1"; + generation: "legacy" | "native"; + agentId: string; + budgets: { companyMonthlyCents: unknown; agentMonthlyCents: unknown }; + bundle: { entryFile?: string; files: Array<{ path: string; content: string }> }; + invocations: Array<{ runId: string; prompt: unknown; promptMetrics?: Record; conversationMode?: boolean }>; + runIds: string[]; +} + +export function gradeStockHire(evidence: Pick, + instructionVariant = readStockInstructionVariant()) { + const checks: Array<{ id: string; passed: boolean; detail: string }> = []; + const check = (id: string, passed: boolean, detail: string) => checks.push({ id, passed, detail }); + check("budget-hard-stops", evidence.budgets.companyMonthlyCents === 1_000 && evidence.budgets.agentMonthlyCents === 1_000, + "Public company and agent records must both retain the 1,000-cent monthly hard stops."); + check("default-hire-bundle", evidence.bundle.entryFile === "AGENTS.md" && + evidence.bundle.files.length === 1 && evidence.bundle.files[0]?.path === "AGENTS.md" && + evidence.bundle.files[0]?.content === instructionVariant.content, + instructionVariant.variant === "reduced" + ? "The public managed bundle must contain only the eight-word shipped identity, without an injected QA manual." + : "The comparison bundle must exactly match the independently pinned historical manual; this is structural evidence, not a task outcome."); + return checks; +} + +export function gradeStockHarness(evidence: StockHarnessEvidence, instructionVariant = readStockInstructionVariant()) { + const checks = gradeStockHire(evidence, instructionVariant); + const check = (id: string, passed: boolean, detail: string) => checks.push({ id, passed, detail }); + check("provider-runs-present", evidence.runIds.length > 0 && evidence.runIds.every(Boolean), + "The lifecycle oracle must reach actual provider runs; missing runs cannot pass instruction delivery."); + if (evidence.generation === "legacy") { + const prompts = evidence.invocations.filter(row => typeof row.prompt === "string" && row.prompt.length > 0); + check("invocation-evidence-complete", evidence.runIds.length > 0 && evidence.runIds.every(runId => + prompts.some(row => row.runId === runId)), + "Every legacy run must have a public adapter.invoke event with its actual nonempty prompt."); + if (instructionVariant.variant === "reduced") { + check("generic-procedures-absent", prompts.length > 0 && prompts.every(row => + REMOVED_PROCEDURES.every(procedure => !String(row.prompt).includes(procedure))), + "Neither startup nor continuation may reintroduce the removed generic manual."); + } else { + const classified = prompts.every(row => typeof row.conversationMode === "boolean" && + typeof row.promptMetrics?.heartbeatPromptChars === "number" && + Number.isFinite(row.promptMetrics.heartbeatPromptChars) && row.promptMetrics.heartbeatPromptChars >= 0); + const fresh = prompts.filter(row => Number(row.promptMetrics?.heartbeatPromptChars) > 0); + const resumed = prompts.filter(row => row.promptMetrics?.heartbeatPromptChars === 0); + check("historical-invocation-classification", prompts.length > 0 && classified, + "Every invocation needs public run-scoped conversation mode and explicit template-delivery metrics."); + check("historical-startup-contract", classified && fresh.length > 0 && fresh.every(row => row.conversationMode + ? String(row.prompt).includes("Continue your Paperclip conversation using the supplied chat mode directive.") && + String(row.prompt).includes("After 2 consecutive failures of the same control-plane write") + : String(row.prompt).includes("Execution contract:") && String(row.prompt).includes("Final disposition checklist:")), + `Check each of ${fresh.length} fresh invocations against its historical task or conversation template.`); + check("historical-continuation-contract", classified && resumed.every(row => + String(row.prompt).includes("## Paperclip Resume Delta") && (row.conversationMode + ? !String(row.prompt).includes("Execution contract:") + : String(row.prompt).includes("Execution contract: take concrete action") && + String(row.prompt).includes("a successful process exit or final response is not sufficient"))), + `Check all ${resumed.length} observed continuation invocations separately; zero observations do not claim live continuation coverage.`); + } + const fresh = prompts.filter(row => Number(row.promptMetrics?.heartbeatPromptChars) > 0); + check("fresh-default-delivered", fresh.length > 0 && fresh.every(row => + String(row.prompt).includes(STOCK_TEMPLATE_IDENTITY) && String(row.prompt).includes("Connection tools:") && + String(row.prompt).includes("connections_search")), + "At least one fresh invocation must carry default identity and connection guidance."); + } + return checks; +} + +/** Public API receipts only; no provider-memory claim or private runner hook. */ +export async function captureStockHarness(input: { + api: { get(path: string): Promise }; + agentId: string; + companyId: string; + generation: "legacy" | "native"; + runIds: string[]; +}): Promise { + const bundle = await input.api.get<{ entryFile: string; files: Array<{ path: string }> }>( + `/api/agents/${input.agentId}/instructions-bundle`, + ); + const [company, agent] = await Promise.all([ + input.api.get<{ budgetMonthlyCents: unknown }>(`/api/companies/${input.companyId}`), + input.api.get<{ budgetMonthlyCents: unknown }>(`/api/agents/${input.agentId}`), + ]); + const files = await Promise.all(bundle.files.map(async file => { + const detail = await input.api.get<{ content: string }>( + `/api/agents/${input.agentId}/instructions-bundle/file?path=${encodeURIComponent(file.path)}`, + ); + return { path: file.path, content: detail.content }; + })); + const events = await Promise.all(input.runIds.map(async runId => ({ + runId, + run: await input.api.get<{ id: string; companyId: string; agentId: string; contextSnapshot?: { conversationMode?: boolean } }>( + `/api/heartbeat-runs/${runId}`, + ), + rows: await input.api.get }>>( + `/api/heartbeat-runs/${runId}/events?limit=1000`, + ), + }))); + if (events.some(({ runId, run }) => run.id !== runId || run.companyId !== input.companyId || run.agentId !== input.agentId || + !run.contextSnapshot || typeof run.contextSnapshot !== "object" || Array.isArray(run.contextSnapshot) || + (run.contextSnapshot.conversationMode !== undefined && typeof run.contextSnapshot.conversationMode !== "boolean"))) { + throw new Error("Stock invocation mode receipt is missing or bound to a different run/company/agent."); + } + return { + schema: "paperclip.stock-harness.v1", generation: input.generation, + budgets: { companyMonthlyCents: company.budgetMonthlyCents, agentMonthlyCents: agent.budgetMonthlyCents }, + agentId: input.agentId, bundle: { entryFile: bundle.entryFile, files }, runIds: input.runIds, + invocations: events.flatMap(({ runId, rows, run }) => rows.filter(row => row.eventType === "adapter.invoke") + .map(row => ({ runId, prompt: row.payload?.prompt, promptMetrics: row.payload?.promptMetrics as Record | undefined, + conversationMode: run.contextSnapshot?.conversationMode === true }))), + }; +} + +export function stockHarnessSourceDigest() { + const hash = createHash("sha256"); + for (const source of [ + "stock-harness.ts", "context-integrity-cases.ts", "context-integrity-scoring.ts", + "context-integrity-flow.ts", "chat-cases.ts", "chat-flow.ts", "live-fixtures.ts", "runner.spec.ts", + "stock-harness-checks.mjs", "stock-harness-admission.ts", "stock-harness-manifest.ts", + "stock-harness-instruction-variant.mjs", "stock-harness-instruction-variant.d.mts", "stock-harness-instruction-variant.test.mjs", + "fixtures/stock-harness/historical-default-agents.md", "automatic-retry.ts", "automatic-retry.test.ts", "types.ts", "catalog.ts", "launch.ts", + "../../packages/adapter-utils/src/acpx-engine/execute.ts", + "../../packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.ts", + "../../packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.test.ts", + "../../packages/paperclip-runner/scripts/generate-capability-contract.mjs", + "../../packages/paperclip-runner/scripts/check-capability-inventory.mjs", + "../../packages/paperclip-runner/scripts/lib/capability-inventory.mjs", + "../../packages/paperclip-runner/spec/capability/source-contract.json", + "../../packages/paperclip-runner/spec/capability/capabilities.yaml", + "../../packages/paperclip-runner/spec/capability/eval-traceability.yaml", + "../../packages/paperclip-runner/spec/capability/mcp-tool-map.yaml", + "../../packages/paperclip-runner/spec/capability/inventory.schema.json", + "../../packages/paperclip-runner/src/generated/capability-contract.ts", + "../../packages/paperclip-runner/docs/capability-contract.md", + ...["capabilities.yaml", "mcp-tool-map.yaml", "eval-traceability.yaml", "capability-contract.md", "downstream-handoff.md"] + .map(file => `../../packages/paperclip-runner/generated/capability/${file}`), + "../../server/src/onboarding-assets/default/AGENTS.md", + "../../packages/adapter-utils/src/server-utils.ts", + "../../packages/shared/src/connection-intent-guidance.ts", + ]) hash.update(source).update(readFileSync(new URL(source, import.meta.url))); + hash.update(JSON.stringify(stockHarnessSkillSources())); + return hash.digest("hex"); +} diff --git a/tests/runner-e2e/vitest.config.ts b/tests/runner-e2e/vitest.config.ts index 38dcf1b369..fd13ba8fd3 100644 --- a/tests/runner-e2e/vitest.config.ts +++ b/tests/runner-e2e/vitest.config.ts @@ -4,7 +4,7 @@ export default defineConfig({ root: import.meta.dirname, test: { name: "runner-e2e-support", - include: ["**/*.test.ts"], + include: ["**/*.test.ts", "**/*.test.mjs"], testTimeout: 30_000, }, });