Files
PaperClipAI/tests/runner-e2e/QUALIFICATION-2026-09-21.md
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00

11 KiB

Native onboarding and Agent Chat qualification — 21 September 2026

This record distinguishes live behavior from the existence of an eval. No native onboarding default or production prompt was changed. All paid work ran on isolated GitHub Actions workers with real providers, Chromium, Paperclip, and public APIs.

Native onboarding

Campaign 35656761484 on master 846336e5a0d3003e1b938e8c983aa5a65bda81ff: 26/26 passed, with 26/26 cleanup passes. The 13 cases ran on both native Codex and native Claude: interviews, clear/ambiguous requests, plain messages, explicit plans, ordinary-task isolation, plan acceptance, card/reply acceptance, acceptance during active work, clarification, revision, and rejection. The observed completed models were gpt-5.6-sol and claude-sonnet-5; some Codex first-response cases stop on an answerable native question before model metadata is emitted.

The real wizard creates the agent. The existing fixture switches its runtime through the public API before the first provider run, preserving the wizard's model choice, persona, skills, and task. This qualifies the pre-default native first-task process. It does not qualify native selection in the onboarding UI or production rollout, neither of which is enabled by this change.

Initial new-case campaign

Campaign 35657128077 on 3ca6d25e0d20dd36e74c0868872802856e5503fc retained all six failures:

  • Active reassignment, both providers: the fixture sent an instructions bundle to the general agent PATCH endpoint. It did not update the worker's managed file, so the worker finished before the intended wait. The active boundary was never exercised. Use the instructions-file endpoint and verify its persisted content.
  • Answer quality, both providers: the prompt did not specify whether the blocker field should contain only its reference or also its description, and did not explicitly exclude the current chat from the active-run count. Codex returned the reference plus its correct description; Claude correctly counted this chat and explained that no other work was active. Claude also put the explanation outside the JSON. These attempts do not establish a factual reasoning failure. The fixture now specifies the JSON shape, reference-only field, and task scope.
  • Worker crash, both providers: after the verified worker loss, Retry was visible and accepted. The second run immediately failed with native_session_cleanup_quarantined. The saved plan and original message survived, but no usable recovered answer was produced. This is a real recovery boundary, not a model failure. Reported billing is incomplete for these crashed runs; absent usage must not be described as zero cost.

The fault helper now uses a Linux pidfd, checks the command's exact run ID and start ticks, and signals the owned handle. Tests reject changed identity and dead handles. Linux CI also exercises a real disposable child. macOS does not run this fault case; the supported qualification host is Linux.

Recovery correction

A quarantined failure can precede creation of its coordinator row. The execution projection now preserves recovery_needed even without that row or with a stale retryable coordinator. The generic failed-run retry endpoint refuses this exact native quarantine instead of admitting another doomed run. An explicitly reconciled successor still clears the historical projection correctly.

The worker-crash eval remains non-passing when recovery is unavailable. It records that saved work survives, verifies the unavailable Retry UI and HTTP 409, and reports the missing usable recovery. Passing a guard test is not recovery qualification. Selecting verified cleanup plus a fresh attempt versus exact-session resume remains a product decision; neither is silently enabled here.

Answer-quality review

Review method: inspect retained user comments, agent comments, question/confirmation cards, and durable task outcomes. This is an agent-authored semantic review, not an independent LLM-judge score and not a claim about all possible questions.

The 26 onboarding recordings show focused questions for ambiguous requests, concrete proposals for clear requests, preservation of the supplied day/fee/audience, and respect for revision and rejection. The first status answers in the new campaign correctly identify the resolved budget issue, current venue blocker, deferred printing, and unknown attendance count. The second-turn counterfactual questions were not reached in that initial campaign.

Two user-facing limitations remain worth deciding:

  • One Claude acceptance reply exposed internal wording: “per the confirmation proposal mode.” This conflicts with the existing skill's wording preference.
  • Five of the six Claude acceptance journeys left a future-tense handoff as the latest parent-thread reply, even though the child later completed. The saved child output and task status passed. Codex posted completion follow-ups in the corresponding journeys. Decide whether a final parent-thread completion notice should be a required outcome; the current onboarding oracle does not require it.

Production instructions were preserved. These observations must not be hidden by changing prompts solely to make the benchmark green.

Grounded status answers: corrected live proof

Campaign 35658262695 on 4a26f10be7dd6aeabf2ca7b44b3d6b2ece7817e0: 2/2 passed, both cleanup passes. Each provider answered both turns, preserved both source tasks, and started no execution on either task. The original failed attempts above are retained; this is a new campaign with an explicit output contract.

Semantic review of all four retained replies against the task descriptions and chronological comments found:

  • Both identified venue confirmation as the current blocker, printing as deferred, no task execution, and unknown attendance. Both gave the useful next step of confirming the venue before printing.
  • Both rejected the obsolete budget claim and unsupported printing claim on the follow-up, and distinguished a planned Friday from a guaranteed calendar date. Neither invented a venue, date, or attendance count.
  • Codex's explanations were compact and clear. Claude's second explanation was longer but readable and grounded in named task records. Its phrase “no venue has been confirmed” is slightly stronger than “no confirmation is recorded”; its immediately following quotation and null fact make the evidence limitation clear.

This qualifies these two-turn grounding stories, not general answer quality, statistical reliability, multilingual behavior, or arbitrary long conversations. The five review dimensions remain factual grounding, stale-premise correction, honest uncertainty, useful next step, and clear prose. No production prompt changed.

Active reassignment: corrected live proof

Campaign 35659014397 on cf6d4ae3a8576d822861bc916e7f536d0a0b7bc7: 2/2 passed, both cleanup passes, using gpt-5.6-sol and claude-sonnet-5.

The boundary snapshot proves the original worker was running with a saved draft. Chat then reassigns the existing task to a second agent. The original stops with issue_reassigned before the successor starts; exactly one successor completes the same task. The plan, scope, and original draft revision survive. The successor may revise the canonical document, and its final contribution must be attributed to that successor. The audit records the native reassignment tool.

The preceding campaign 35657945095 on 6039b02ed remains failed: both handoffs actually completed, but the draft oracle incorrectly required the latest document to remain frozen. Both successors legitimately revised that document. The new oracle requires the exact original revision to remain retrievable through the history API while allowing progress. Negative calibration rejects deletion or alteration of that revision. Both crash cases in the preceding campaign independently reproduced cleanup quarantine.

Measurement limits

All selected environments here are local Linux GitHub Actions workers, not Daytona-hosted execution or the user's staging company. The result JSON retains source SHA, suite definition hash, profile/model, attempt, timing, token usage, and billing coverage; initial failures are never regraded as passes.

Onboarding reports complete billing for only 4/26 results; active handoff reports incomplete billing for both interrupted workers; the two answer-quality results report complete coverage. Provider-reported monetary cost is zero while billing type is unknown, and local runtime is not metered. This does not establish zero actual spend. Recorded total tokens (including cached input) are 9,812,045 for onboarding, 825,131 for final handoffs, and 1,273,098 for grounded answers.

Public campaign viewers use the standard Product E2E history with the gha-<workflow-run-id>-<attempt> campaign identifier. These are bounded story qualifications, not a claim that all native-runner reliability is solved.

Crash guard follow-up

Campaign 35658772755 did not reach provider cases: GitHub artifact finalization returned HTTP 403. A replacement campaign 35659580100 on a11bd236e33833dc081cc3702baa3d3f98d8d12f retained two failed recovery attempts, both with successful disposable cleanup. The API correctly refused Retry with 409 and created no second run, but the eval then waited for a run that had never been admitted.

Retained network evidence explains the transition: provider_transport_failed first schedules a same-run retry; that attempt then reaches cleanup quarantine. The fixture now waits for this recovery classification before requesting a user retry. Screenshots also exposed a separate Retry button in the task-recovery banner. That banner previously recognized only native continuation reconciliation; it now also recognizes cleanup quarantine and links to Inspect run. Its regression test failed before the fix and passes afterward. The live guard must verify this link is rendered, no Retry remains, the API refuses retry, and saved work survives.

These corrections do not supply a successful post-crash continuation. Worker-crash recovery remains unqualified until an agreed recovery policy is implemented and its successful outcome passes the eval.