mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 14:10:50 +02:00
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat needs reliable native execution before native runners become the onboarding default. > - Existing stories covered idle reassignment and controller restart, but not an executing worker handoff or worker process loss. > - Status answer tests also need to reject stale claims and invented facts. > - This pull request adds six opt-in full-stack cells with independent state assertions and retained evidence. > - The probes exposed a misleading Retry across server projection and recovery-banner paths; the fix reports the blocked recovery honestly. > - The tests preserve failures without changing recovery policy, production prompts, or onboarding defaults. ## Linked Issues or Issue Description Refs: #13762. Related: #13765 (Retry targets the latest failed attempt), #13753 (task context ownership), #13746 (native recovery work). ## What Changed - Add active reassignment with saved draft and plan preservation, old-worker cancellation, and successor completion checks. - Preserve recovery-needed projection when native cleanup fails before its coordinator exists, refuse a generic retry that would immediately fail again, and replace the recovery banner's misleading Retry with Inspect run. - Add verified local worker process loss with a required successful continuation; retain a failing qualification result when recovery is unavailable, while independently verifying the UI/API refuse doomed retries. - Add two-turn factual answer checks for current blockers, stale claims, inactive backlog work, and unknown facts. Retain prose for separate semantic review. - Add positive and negative oracle calibration and document fault isolation, cleanup, billing, and qualification limits. ## Verification - Eval TypeScript check passes. - All 442 eval support tests pass locally. The 89 focused server tests and server typecheck pass. Six recovery-banner UI tests and token gates pass. - Initial new-cell campaign: https://github.com/paperclipai/paperclip/actions/runs/35657128077. All six results are retained; four failed on fixture-contract issues and two exposed real worker cleanup quarantine. - All 26 existing native onboarding cells: https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26 passed on master846336e5a, all cleanup passed). - Intermediate handoff/fault campaign: https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four retained failures: two overly strict draft oracles, two real crash quarantines). - Final active handoff: https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2 passed on cf6d4ae3a; both cleanup passed). - Clarified answer-quality fixtures: https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2 passed on 4a26f10be; both cleanup passed; all four answers semantically reviewed). - Quarantine guard regression campaign: https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both API requests correctly refused with 409/no second run, but exposed a separate misleading Retry in the recovery banner and a fixture wait on a non-admitted run). - Final quarantine guard verification: https://github.com/paperclipai/paperclip/actions/runs/35661067305 (147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one retained run, unchanged saved plan, and successful disposable cleanup. Both evals intentionally remain red with `worker_crash_recovery_unqualified`; no successful continuation exists). The preceding campaign 35658772755 never ran provider cases because GitHub artifact finalization returned HTTP 403. - Full repository CI passes on147e42f7e: typecheck, tests, build, and browser gates. One unchanged local-service-supervisor readiness test failed initially; its six-test file passed in isolation and the failed shard passed on its single rerun. Latest-head rollup: 54 successful, 2 intentionally skipped, no failed or pending checks. Greptile is 5/5 with zero unresolved findings. - See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts and semantic review. Published reports: [onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/), [handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/), [grounded answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/), [crash guards and unqualified recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/). ## Risks - Paid cells are explicit-only and local-only. The fault fixture signals only the exact native run PID after checking its identity. - Live worker-loss probes currently fail on cleanup quarantine for both providers. The eval must remain red until there is a usable recovery, even when preservation and refusal checks pass. Verified cleanup with a fresh attempt versus exact-session resume remains a product decision. - Structured facts alone do not qualify prose quality; semantic review remains separate. - Onboarding uses the existing runtime switch after the real wizard and before provider execution. Native UI selection and public defaults remain unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact served snapshot and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
181 lines
11 KiB
Markdown
181 lines
11 KiB
Markdown
# Native onboarding and Agent Chat qualification — 21 September 2026
|
|
|
|
This record distinguishes live behavior from the existence of an eval. No native
|
|
onboarding default or production prompt was changed. All paid work ran on isolated
|
|
GitHub Actions workers with real providers, Chromium, Paperclip, and public APIs.
|
|
|
|
## Native onboarding
|
|
|
|
[Campaign 35656761484](https://github.com/paperclipai/paperclip/actions/runs/35656761484)
|
|
on master `846336e5a0d3003e1b938e8c983aa5a65bda81ff`: **26/26 passed**,
|
|
with **26/26 cleanup passes**. The 13 cases ran on both native Codex and native
|
|
Claude: interviews, clear/ambiguous requests, plain messages, explicit plans,
|
|
ordinary-task isolation, plan acceptance, card/reply acceptance, acceptance during
|
|
active work, clarification, revision, and rejection. The observed completed models
|
|
were `gpt-5.6-sol` and `claude-sonnet-5`; some Codex first-response cases stop on an
|
|
answerable native question before model metadata is emitted.
|
|
|
|
The real wizard creates the agent. The existing fixture switches its runtime
|
|
through the public API before the first provider run, preserving the wizard's
|
|
model choice, persona, skills, and task. This qualifies the pre-default native
|
|
first-task process. It does not qualify native selection in the onboarding UI or
|
|
production rollout, neither of which is enabled by this change.
|
|
|
|
## Initial new-case campaign
|
|
|
|
[Campaign 35657128077](https://github.com/paperclipai/paperclip/actions/runs/35657128077)
|
|
on `3ca6d25e0d20dd36e74c0868872802856e5503fc` retained all six failures:
|
|
|
|
- Active reassignment, both providers: the fixture sent an instructions bundle to
|
|
the general agent PATCH endpoint. It did not update the worker's managed file,
|
|
so the worker finished before the intended wait. The active boundary was never
|
|
exercised. Use the instructions-file endpoint and verify its persisted content.
|
|
- Answer quality, both providers: the prompt did not specify whether the blocker
|
|
field should contain only its reference or also its description, and did not
|
|
explicitly exclude the current chat from the active-run count. Codex returned
|
|
the reference plus its correct description; Claude correctly counted this chat
|
|
and explained that no other work was active. Claude also put the explanation
|
|
outside the JSON. These attempts do not establish a factual reasoning failure.
|
|
The fixture now specifies the JSON shape, reference-only field, and task scope.
|
|
- Worker crash, both providers: after the verified worker loss, Retry was visible
|
|
and accepted. The second run immediately failed with
|
|
`native_session_cleanup_quarantined`. The saved plan and original message
|
|
survived, but no usable recovered answer was produced. This is a real recovery
|
|
boundary, not a model failure. Reported billing is incomplete for these crashed
|
|
runs; absent usage must not be described as zero cost.
|
|
|
|
The fault helper now uses a Linux pidfd, checks the command's exact run ID and
|
|
start ticks, and signals the owned handle. Tests reject changed identity and dead
|
|
handles. Linux CI also exercises a real disposable child. macOS does not run this
|
|
fault case; the supported qualification host is Linux.
|
|
|
|
## Recovery correction
|
|
|
|
A quarantined failure can precede creation of its coordinator row. The execution
|
|
projection now preserves `recovery_needed` even without that row or with a stale
|
|
retryable coordinator. The generic failed-run retry endpoint refuses this exact
|
|
native quarantine instead of admitting another doomed run. An explicitly
|
|
reconciled successor still clears the historical projection correctly.
|
|
|
|
The worker-crash eval remains **non-passing** when recovery is unavailable. It
|
|
records that saved work survives, verifies the unavailable Retry UI and HTTP 409,
|
|
and reports the missing usable recovery. Passing a guard test is not recovery
|
|
qualification. Selecting verified cleanup plus a fresh attempt versus exact-session
|
|
resume remains a product decision; neither is silently enabled here.
|
|
|
|
## Answer-quality review
|
|
|
|
Review method: inspect retained user comments, agent comments, question/confirmation
|
|
cards, and durable task outcomes. This is an agent-authored semantic review, not an
|
|
independent LLM-judge score and not a claim about all possible questions.
|
|
|
|
The 26 onboarding recordings show focused questions for ambiguous requests,
|
|
concrete proposals for clear requests, preservation of the supplied day/fee/audience,
|
|
and respect for revision and rejection. The first status answers in the new
|
|
campaign correctly identify the resolved budget issue, current venue blocker,
|
|
deferred printing, and unknown attendance count. The second-turn counterfactual
|
|
questions were not reached in that initial campaign.
|
|
|
|
Two user-facing limitations remain worth deciding:
|
|
|
|
- One Claude acceptance reply exposed internal wording: “per the confirmation
|
|
proposal mode.” This conflicts with the existing skill's wording preference.
|
|
- Five of the six Claude acceptance journeys left a future-tense handoff as the
|
|
latest parent-thread reply, even though the child later completed. The saved
|
|
child output and task status passed. Codex posted completion follow-ups in the
|
|
corresponding journeys. Decide whether a final parent-thread completion notice
|
|
should be a required outcome; the current onboarding oracle does not require it.
|
|
|
|
Production instructions were preserved. These observations must not be hidden by
|
|
changing prompts solely to make the benchmark green.
|
|
|
|
## Grounded status answers: corrected live proof
|
|
|
|
[Campaign 35658262695](https://github.com/paperclipai/paperclip/actions/runs/35658262695)
|
|
on `4a26f10be7dd6aeabf2ca7b44b3d6b2ece7817e0`: **2/2 passed**, both
|
|
cleanup passes. Each provider answered both turns, preserved both source tasks,
|
|
and started no execution on either task. The original failed attempts above are
|
|
retained; this is a new campaign with an explicit output contract.
|
|
|
|
Semantic review of all four retained replies against the task descriptions and
|
|
chronological comments found:
|
|
|
|
- Both identified venue confirmation as the current blocker, printing as deferred,
|
|
no task execution, and unknown attendance. Both gave the useful next step of
|
|
confirming the venue before printing.
|
|
- Both rejected the obsolete budget claim and unsupported printing claim on the
|
|
follow-up, and distinguished a planned Friday from a guaranteed calendar date.
|
|
Neither invented a venue, date, or attendance count.
|
|
- Codex's explanations were compact and clear. Claude's second explanation was
|
|
longer but readable and grounded in named task records. Its phrase “no venue has
|
|
been confirmed” is slightly stronger than “no confirmation is recorded”; its
|
|
immediately following quotation and null fact make the evidence limitation clear.
|
|
|
|
This qualifies these two-turn grounding stories, not general answer quality,
|
|
statistical reliability, multilingual behavior, or arbitrary long conversations.
|
|
The five review dimensions remain factual grounding, stale-premise correction,
|
|
honest uncertainty, useful next step, and clear prose. No production prompt changed.
|
|
|
|
## Active reassignment: corrected live proof
|
|
|
|
[Campaign 35659014397](https://github.com/paperclipai/paperclip/actions/runs/35659014397)
|
|
on `cf6d4ae3a8576d822861bc916e7f536d0a0b7bc7`: **2/2 passed**, both
|
|
cleanup passes, using `gpt-5.6-sol` and `claude-sonnet-5`.
|
|
|
|
The boundary snapshot proves the original worker was running with a saved draft.
|
|
Chat then reassigns the existing task to a second agent. The original stops with
|
|
`issue_reassigned` before the successor starts; exactly one successor completes
|
|
the same task. The plan, scope, and original draft revision survive. The successor
|
|
may revise the canonical document, and its final contribution must be attributed
|
|
to that successor. The audit records the native reassignment tool.
|
|
|
|
The preceding [campaign 35657945095](https://github.com/paperclipai/paperclip/actions/runs/35657945095)
|
|
on `6039b02ed` remains failed: both handoffs actually completed, but the draft
|
|
oracle incorrectly required the latest document to remain frozen. Both successors
|
|
legitimately revised that document. The new oracle requires the exact original
|
|
revision to remain retrievable through the history API while allowing progress.
|
|
Negative calibration rejects deletion or alteration of that revision. Both crash
|
|
cases in the preceding campaign independently reproduced cleanup quarantine.
|
|
|
|
## Measurement limits
|
|
|
|
All selected environments here are local Linux GitHub Actions workers, not
|
|
Daytona-hosted execution or the user's staging company. The result JSON retains
|
|
source SHA, suite definition hash, profile/model, attempt, timing, token usage,
|
|
and billing coverage; initial failures are never regraded as passes.
|
|
|
|
Onboarding reports complete billing for only 4/26 results; active handoff reports
|
|
incomplete billing for both interrupted workers; the two answer-quality results
|
|
report complete coverage. Provider-reported monetary cost is zero while billing
|
|
type is unknown, and local runtime is not metered. This does **not** establish
|
|
zero actual spend. Recorded total tokens (including cached input) are 9,812,045
|
|
for onboarding, 825,131 for final handoffs, and 1,273,098 for grounded answers.
|
|
|
|
Public campaign viewers use the standard
|
|
[Product E2E history](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/)
|
|
with the `gha-<workflow-run-id>-<attempt>` campaign identifier. These are bounded
|
|
story qualifications, not a claim that all native-runner reliability is solved.
|
|
|
|
## Crash guard follow-up
|
|
|
|
[Campaign 35658772755](https://github.com/paperclipai/paperclip/actions/runs/35658772755)
|
|
did not reach provider cases: GitHub artifact finalization returned HTTP 403.
|
|
A replacement [campaign 35659580100](https://github.com/paperclipai/paperclip/actions/runs/35659580100)
|
|
on `a11bd236e33833dc081cc3702baa3d3f98d8d12f` retained two failed recovery
|
|
attempts, both with successful disposable cleanup. The API correctly refused
|
|
Retry with 409 and created no second run, but the eval then waited for a run
|
|
that had never been admitted.
|
|
|
|
Retained network evidence explains the transition: `provider_transport_failed`
|
|
first schedules a same-run retry; that attempt then reaches cleanup quarantine.
|
|
The fixture now waits for this recovery classification before requesting a user
|
|
retry. Screenshots also exposed a separate Retry button in the task-recovery
|
|
banner. That banner previously recognized only native continuation reconciliation;
|
|
it now also recognizes cleanup quarantine and links to Inspect run. Its regression
|
|
test failed before the fix and passes afterward. The live guard must verify this
|
|
link is rendered, no Retry remains, the API refuses retry, and saved work survives.
|
|
|
|
These corrections do not supply a successful post-crash continuation. Worker-crash
|
|
recovery remains unqualified until an agreed recovery policy is implemented and
|
|
its successful outcome passes the eval.
|