Files
PaperClipAI/tests/runner-e2e/STOCK-HARNESS.md
DottaandPaperclip 71cd0a2621 fix(skills): honor the current run harness checkout (#15548)
## Thinking Path

> - Paperclip manages work for AI agents.
> - The runtime claims eligible assigned tasks before it starts an
agent.
> - The wake tells the agent when the runtime already holds that claim.
> - The legacy skill still requires another checkout in every case.
> - This PR makes the skill honor the current task and run claim.
> - Manual checkout and server ownership checks remain in place for
other cases.

## Linked Issues or Issue Description

**Where is the issue?**

`skills/paperclip/SKILL.md`, in the scoped wake procedure and Step 5.

**What's wrong?**

The wake can say that the harness already checked out the issue. The
skill still tells the agent that it must call checkout. These
instructions conflict.

**Suggested fix**

Skip the second checkout only when the runtime wake explicitly confirms
the claim for this issue and run. Retain manual checkout when that
statement is absent or the agent selects another task. Refs #14948 for
the existing shared prompt reduction.

## What Changed

- Honor the explicit runtime claim in the scoped wake procedure and Step
5.
- Keep context reads, status writes, deliverable handling and conflict
rules.
- Add checks for normal and resumed wake text and excluded automatic
claims.
- Retain successful checkout HTTP activity for legacy stock-task evals.
Bind each receipt to the exact company, task, agent and run. Keep this
observation separate from the original task grades.

## Verification

- Checkout observation calibration: nine tests pass.
- Focused skill, wake and database ownership tests: in progress.
- Full repository build, typecheck and tests: in progress.
- Planned live comparison: the existing assigned-skill document case on
legacy Codex and Claude. One attempt per variant and profile. No
automatic retries. The baseline and candidate share the observation code
and task oracle.
- Live results are pending. This draft does not claim behavioral
qualification.

## Risks

- Agents may misread prompt guidance. The API still enforces ownership;
the text grants no new authority.
- The exception is specific to the current issue and run. It does not
remove ordinary legacy completion writes or authorize another task.
- Activity measures successful checkout HTTP calls. Failed attempts
require separate run-log inspection. Missing or mismatched observations
cannot count as zero calls.
- One trial per profile cannot establish general reliability, speed or
cost trends.

## Model Used

OpenAI Codex, GPT-6 family. The exact model build and context window are
not exposed in this session. Used code editing, shell tools and test
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 06:38:09 -05:00

17 KiB

Stock harness with Paperclip

This suite covers the instruction reductions tracked in the working checklist. Existing context-integrity and chat evals supplied a custom QA instruction bundle. Their passing results therefore did not qualify a hire with the tiny production default. The new suite omits that fixture bundle and lets the public agent-creation route materialize the shipped default.

The 2026-10-03 legacy ACP Claude repair compares only the original assigned-skill-explicit-invocation cell on matched current-master sources. Both variants carry identical bounded skill-description/staged-path delivery and env-free session persistence. Only the historical versus reduced default manual/shared prompts and their declared structural unit expectations differ. The public bundle/procedure observations derive from exact admitted source bytes (eight-word identity or the independently pinned 4,249-byte historical manual), never a branch/environment label. Candidate absence assertions stay intact. Independent document/task and exact-credential guards are unchanged. All stock tasks enforce single_attempt in the hosted launcher; automatic product disposition recovery still counts as actual usage. This is not a new 24-cell campaign or qualification of the frozen native-completion context.

Coverage contract

Change Required deterministic evidence Product E2E evidence
SH-1: native Codex preserves vendor base instructions Serialized start/resume requests from the TypeScript driver, recovery paths, runnerd transport, Runner Lab/live sessions, and Rust provider use additive developer instructions. Real new native Codex hires execute the three journeys below. Task success alone cannot prove vendor base preservation.
SH-2: identity-only default hire manual Existing public agent-creation and onboarding-asset tests cover default, custom, and CEO exceptions. The public bundle is exactly one AGENTS.md containing the eight-word shipped identity, checked before provider execution and again during cleanup.
SH-3: reduced shared legacy task/chat defaults and resume delta Shared prompt tests and ACPX, Codex, OpenCode, Pi, Hermes, and Cursor Cloud adapter regressions retain runtime context and exclude removed generic procedures. Actual legacy adapter.invoke prompts retain fresh identity/connection guidance, omit the removed procedures, and have complete receipts for every observed run.

pnpm test:e2e:runner:stock-harness runs these credential-free prerequisites, including the Rust test. It writes JSON reports, the source SHA, a working-source fingerprint, selected checks, test counts, and providerCalls: 0 under results/stock-harness-preflight-<UTC>/. Missing reports or required skipped assertions fail. Unrelated native tests filtered by the name selector remain explicitly skipped; they are not counted as executed coverage. The live launcher runs the prerequisites automatically before loading local credentials or starting an isolated provider instance. Its subprocess receives only allowlisted toolchain and operating-system variables. Disposable GitHub runners may resolve Cargo dependencies; local prerequisites retain offline Cargo execution. The ordinary server SDK dependency builder prepares the shared/SDK outputs needed by route tests on a cold install with lifecycle scripts disabled.

The direct Playwright path also requires the retained prerequisite receipt. It verifies the exact checkout SHA, evaluated-source fingerprint, all requested Vitest reports and required assertions, and the Rust result before creating a company. Missing, stale, partial, or failed prerequisites cannot qualify a cell. Oracle/admission calibration is itself included in the prerequisite gate.

Checkout observations

Legacy context-integrity journeys retain checkout-activity.json from the public issue activity API. Each receipt must match the exact company, issue, agent and run. An idempotent agent checkout still records a successful route call; the server's pre-dispatch checkout does not. Missing or mismatched evidence is unavailable or uncomparable, never a zero. This observation leaves the original outcome grades unchanged. It counts successful HTTP checkouts only; inspect retained executed commands/results for failed attempts before claiming no calls.

Live matrix

There are 24 explicit local cells: these eight existing profiles each run three existing journeys with independently calibrated graders.

  • Legacy: legacy-codex, legacy-claude, legacy-opencode, legacy-acp-codex, legacy-acp-claude.
  • Native: runner-codex, runner-acpx-claude, runner-opencode.
Journey Expected provider turns Observable outcome Cell deadline
assigned-skill-explicit-invocation 1 Public skill creation/pinning, an explicit skill request, and saved output containing the marker available only in the skill body. 12 minutes
ordered-comment-continuation 2 Initial report followed by three ordered public comments, including repeated wording and a changed scope; final saved report preserves the ledger and requested scope. 12 minutes
continuity-restart 3 Task-backed chat retains the requested context across a server restart and subsequent replies. 15 minutes

A full matrix expects 48 provider turns. Models, credentials, effort, permissions, assigned skills, environment, and managed secret references are inherited from the existing profile; only its QA manual is omitted. Both company and agent receive a 1,000-cent monthly hard stop before provider execution, verified through public records. These limits do not predict final spend: attempts, retries, partial runs, unknown billing, and cleanup remain in the existing campaign accounting and qualification rules.

stock-harness is excluded from --all; select it explicitly. It has no Daytona cells. Pending or unrepresented harnesses are not live-qualified by this matrix. Pi/Hermes/Cursor Cloud rendering coverage is deterministic here. The separate Codex-through-ACP base-instruction patch is still an open checklist item: its legacy cells qualify the common prompt reduction, not vendor base preservation.

Run and inspect

# No credentials or paid providers:
pnpm test:e2e:runner:stock-harness --list
pnpm test:e2e:runner:stock-harness
pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
pnpm test:e2e:runner -- --list --suite stock-harness

# With authorization and the selected profile's required provider credential:
pnpm test:e2e:runner -- --id stock-harness.runner-codex.local.assigned-skill-explicit-invocation

Use an exact cell first, then expand profile/journey selections when its evidence is understood. The ordinary launcher, isolated instance, fixture cleanup, screenshots, attempt history, usage/cost reporting, and dashboard publisher are unchanged. snapshots/stock-harness-hire.json captures the public bundle and budgets before paid execution. snapshots/stock-harness.json captures the final bundle, budgets, run IDs, actual legacy prompts, and matcher verdicts. Evidence API failures retain an error snapshot and cannot pass. All snapshots use the existing secret sanitizer; screenshots remain the original captured pixels.

The oracle's identity is independent of the implementation constant, so changing the shipped manual cannot silently change the expected result. The suite definition digest incorporates the evaluated default manual, shared prompt implementation and connection guidance, fixture, journeys, grader, prerequisite, and execution integration sources. Positive and plausible-negative support tests exercise missing/malformed receipts, manual regrowth, removed startup/resume procedures, absent connection guidance, and budget drift. Existing calibrated lifecycle graders still own task/chat success.

The skill oracle proves the requested pinned skill's output marker reached the saved result without appearing in the task request or follow-up comments. Skill tool/read event detection is retained as supporting evidence; it is not a required cross-provider tool-trace assertion.

Qualification status

Setup was validated locally on 2026-10-02 with the deterministic prerequisites, Product E2E support tests, typecheck, and discovery: 478 executed prerequisite tests passed (477 TypeScript plus one Rust), along with 860 support tests and discovery of all 24 cells. The 313 unrelated native tests filtered by the gate are not counted as passing coverage. Local evidence is retained under results/stock-harness-preflight-2026-10-02T16-15-39.064Z/preflight.json. Earlier interrupted, discovery-failure, and setup/test-timeout attempts remain retained; the final unchanged-assertion retry passed. No live cells or paid providers were run during that initial setup. This establishes executable coverage, not a live reliability result or improved coding quality. A quality claim needs comparable tasks, models, effort, tools, and independently graded before/after results.

The suite exercises fresh isolated hires and their continuations. Existing saved manuals and old Codex sessions are not automatically migrated. The latter still need a provider-session reset to restore a previously replaced vendor base. Private Runner protocol definitions remain separate: they use mock control-plane operations and cannot substitute for this public hiring and assembled-prompt coverage.

GitHub qualification and matched comparison

The instruction reductions and coverage are under review in PR #14948. The first diagnostic native Codex skill cell passed on GitHub at a63437069de58d22ee5adbcb6a6202c007dcf037. All seven independent skill/task checks passed; the provider run lasted 34.745 seconds and the cell 56.660 seconds. Its tiny public hire bundle and budget receipts passed, and cleanup passed. This diagnostic predates enforced prerequisite admission and the expanded source digest, so it is retained separately from the final qualification matrix.

The first full candidate attempt at 36e987246b649927e96ce1184cd616c4e490106e was cancelled during prerequisites. Cold protected installs disable lifecycle scripts, leaving the plugin SDK unbuilt; SH-2/SH-3 could not import it. Provider admission was not reached. This setup failure is retained separately from behavioral results. The prerequisite now runs the ordinary server dependency builder first, retains its output and exit status, and refuses admission when setup fails. A cold legacy pilot must pass before the full matrix retry.

The fixed legacy pilot at 4163dbfd0fd4d145bfa52b4d7f80eb59a362ee36 passed SDK setup, hire/shared prompt/oracle checks and the Rust additive test. It stopped before providers: the real daemon-frame test lacked the cold paperclip-runnerd binary. Setup now builds that daemon from the locked Rust source before TypeScript gates, retains runnerd-build.txt, and requires both setup exits in the admission receipt. The required daemon-frame test selects that built debug binary explicitly, without replacing staged product binaries. The receipt records its SHA-256; verification rejects a changed binary before provider admission. The next cold pilot at ac6ddefb589c18fa9c30946db40e58499e9fb0e0 then found the required fake Codex protocol fixture absent. It also stopped before paid providers. Setup uses the package's ordinary locked workspace --bins build, covering both the daemon and its fixture; verification binds both binaries to the retained receipt.

That final cold pilot passed all 537 prerequisites on GitHub and reached Claude Sonnet 4.6. It saved a workspace task-output.md and completed the issue, while the independent durable-document oracle found no Paperclip issue document. The pinned skill's phrase "task document" does not explicitly name storage in Paperclip; matched historical results must precede any regression attribution. The task, instruction delivery, budget, and cleanup receipts remain retained.

Its trusted merged report rejected the prerequisite folder alongside the campaign root and synthesized a missing-result infrastructure error. Raw packaged cell results were uploaded and remain inspectable. Prerequisites now live under the exact campaign root; the unchanged trusted selector contract is calibrated in the mandatory gate. The initial full matched campaigns at f02d8d0df and 12c5433c6 retain their original raw results and publication outcomes. Any reconstructed comparison must declare directory-layout recovery and preserve every result, hash, original failure, and assertion.

Dispatch the trusted workflow from master, with target_branch naming the same-repository candidate and an exact cell selector first. The workflow resolves that target once to an immutable SHA. Never dispatch target-controlled workflow definitions with protected credentials. Protected environments, scoped provider keys, frozen target dependencies, report sanitization, publication, and existing bounded retry/cleanup policies retain their existing owners.

A temporary codex/stock-harness-previous-instructions branch compares the same 24 cells with the previous default manual and shared startup/resume prompts. It holds merged native Codex fix #14920 constant. Only those two production instruction sources differ. Its explicit historical structural oracle expects the old manual and records old generic procedures; the candidate's reduction assertions remain mandatory. Historical tests verify that prior contract. The independent skill/context/chat journeys, behavioral graders, fixtures, models, effort, tools, permissions, and credentials are identical. The retained comparison manifest records their hashes and the restored instruction revision.

One full campaign per variant initially expects 48 provider turns, plus any existing bounded automatic retries; the earlier one-cell diagnostic remains separate. Compare behavioral results by profile and journey, with missing evidence unqualified. Keep structural instruction differences separate from task success. Report every attempt, failure attribution, provider timing, token usage, reported costs and unknown spend. A reported zero subtotal is not proof of zero provider spending. This small single-trial matrix cannot establish general coding quality or broad performance equivalence, and does not compare #14920 before/after.

Measured comparison and unresolved delivery

The dated live report and its safe JSON projection contain the per-profile/case comparison, exact source hashes, all campaign and recovery links, timing/usage, cost coverage, security failures, clipping limits and publication-layout recovery. Candidate f02d8d0df has 24 retained results: 15 pass and nine fail. Historical 12c5433c6 also has 24 retained results: 15 pass and nine fail. Its single AWS-runner recovery timed out; original missing evidence remains recorded. Classic Claude/OpenCode skill runs save no Paperclip task document where historical runs save one; the original oracle remains failed.

The merged native Codex change is held constant. Same aggregate success counts would not establish equivalence: case outcomes differ, fixture storage wording is ambiguous, ACP credential guards fail, and some public prompt receipts are clipped. Native finish/block tool guidance belongs to native runners; legacy document-delivery guidance must use the actual Paperclip skill/API path. No production or skill instructions have been changed to turn the measured failures into passes. PR #14948 remains draft.

The corrected packaging pilot at 1eb5ba420 passes on GitHub with valid evidence and cleanup. It validates prerequisite nesting under the exact campaign root using the unchanged trusted selector. Its current prerequisite runs 557 checks (556 TypeScript plus one Rust); all 895 E2E support tests pass. The 313 filtered native tests are not counted. Original failed/partial attempts and publication failures remain retained; local reconstructed copies move folders without editing results, graders or usage. Never dispatch a second development campaign for an active target branch: workflow concurrency supersedes the older run. Use a separate frozen-source branch when independent campaigns must overlap.

Focused legacy delivery repair

The original 24 cells and original assigned-skill request/procedure remain unchanged. Two added explicit Paperclip-document cells apply only to classic Claude and OpenCode, making 26 catalog cells (50 expected turns if every cell is selected). The repair campaign selects only the original skill case and new document case for those two profiles: four cells per variant, eight expected turns total. No full matrix rerun is planned.

Both skill sources are recorded in definition/admission digests. The pre-fix baseline restores the old SKILL.md and records the new reference as absent, without copying the new recipe into that baseline. It holds the tiny manual/shared prompts and all fixture/model/auth/effort inputs fixed. The explicit document oracle checks actual public content/revision and an exact same-app document link, with plausible-negative calibrations.