Commit Graph
3 Commits
Author SHA1 Message Date
DottaandPaperclip 83076d7e7c feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.

Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 10:25:52 -05:00
DottaandPaperclip 74a9730acb fix: continue native agent chats after worker loss (#13813)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat uses native workers to run Claude and Codex
conversations.
> - A worker crash leaves a cleanup hold because its provider did not
acknowledge suspension.
> - A new user message must not reuse that unverified session or repeat
old tool calls.
> - The existing continuation path can preserve history and start a
fresh session, but local cleanup ownership remained held.
> - This pull request verifies the stopped local owners and releases
only their cleanup hold for a new user turn.

## Linked Issues or Issue Description

Refs #13775.

**What happened?**

After a native worker crashed, both providers retained cleanup
quarantine. A saved plan survived, but the conversation could not
produce another answer.

**Expected behavior**

Once the old worker and provider process groups have stopped, a new user
message can continue in a fresh session with the saved work and prior
action history.

**Steps to reproduce**

Run the opt-in `agent-chat-qualification` suite with case
`worker-crash-retry` on native Codex and native Claude. The fixture
saves a plan, kills the exact worker through a Linux pidfd, releases a
local read-only brief, and sends a new message.

## What Changed

- Verify the exact local worker stop receipt, provider identity
receipts, released leases, and retained state before retiring a native
cleanup hold.
- Recheck process liveness and state before admission. Keep the old run
and durable session files intact.
- Use the existing explicit conversation continuation path. Generic
Retry remains blocked for cleanup quarantine, including on the old
failed-run marker after a successful continuation.
- Extend the live oracle to require a successful fresh session, correct
predecessor context, unchanged plan, one original message, and one
answer containing a reference introduced after the crash.
- Add physical-proof and database-backed admission tests. Document the
precise qualification scope.

## Verification

- Live Product E2E: **2/2 passed**, **2/2 cleanup passed**, with real
native `gpt-5.6-sol` and `claude-sonnet-5`, Chromium, server, database,
and public APIs.
- Core recovery proof source:
`3592b04c2bc76e23795fcdf964720e38a409dc4d`. Suite definition version 8:
`9867367994d81a0c726956d91f2c7fddab6417a12f41b5cef3f7e62f3be417da`.
- [Core recovery
campaign](https://github.com/paperclipai/paperclip/actions/runs/35741746990)
· [Public evidence
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35741746990-1/).
- Both cells verify the real crash boundary, blocked generic Retry,
unchanged saved plan, one original prompt, one fresh successor with
predecessor context, and one run-attributed answer containing the
post-crash reference. Billing coverage is partial because the crashed
runs did not report complete usage; missing cost is not zero cost.
- [Master
baseline](https://github.com/paperclipai/paperclip/actions/runs/35737364443):
both providers stopped at quarantine, with cleanup passing.
- [Follow-up baseline without the
fix](https://github.com/paperclipai/paperclip/actions/runs/35738638866):
Codex produced a complete red result. Claude reached the same error, but
its artifact upload was canceled.
- [Complete Claude baseline with the same version-8
definition](https://github.com/paperclipai/paperclip/actions/runs/35740555177):
red at cleanup quarantine, cleanup passed, source
`6479a90c5754044356b39a9278b9a3e92ce8e55e`.
- Earlier candidate attempts remain retained:
[first](https://github.com/paperclipai/paperclip/actions/runs/35738449214)
passed Codex and found a Claude fixture wait race;
[second](https://github.com/paperclipai/paperclip/actions/runs/35740408065)
exposed the normalized session-open receipt mismatch. Both corrections
are in the final source.
- Targeted server suites: 570 passed before the final two additional
receipt regression cases. The physical-proof suite, including those
cases, passed 46/46. Eval oracle and catalog: 40 passed. Server and
Product E2E typechecks passed.
- **All 54 PR checks passed on final head `db6f775df`**, including
repository typecheck, build, tests, browser shards, and canary dry run.
The canary job required one retry after its runner received a shutdown
signal. On the earlier core proof head, two timing-sensitive tests
passed in isolation and on a single CI retry.
- Final UI regression checks: 5 passed; UI typecheck and token gates
passed. Updated eval oracle/catalog: 40 passed; eval typecheck passed.
- Final version-9 two-provider campaign: **2/2 passed, 2/2 cleanup
passed**, including the browser assertion that the quarantined
historical run never regains Try again. [Final
campaign](https://github.com/paperclipai/paperclip/actions/runs/35747416013),
source `db6f775dfff405e1514ec02fedb0450d42c7dad2`, definition hash
`bf5abf1cc45cb6dad4e082fbf818b8fa0f4d8c282776a7e98762b22919eadab7`.
[Final public evidence
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35747416013-1/).
Final screenshots and retained state inspected for both providers; cost
coverage remains partial.

## Risks

- Missing or conflicting stop evidence keeps the conversation blocked.
This change does not kill an unverified process.
- This qualifies new user input after local worker loss. It does not
enable automatic replay, exact-session recovery, remote crash recovery,
or native onboarding defaults.
- Old action outcomes remain part of the continuation. A process exit is
not proof that an action did not happen.
- No schema migration or production prompt change.

## Model Used

OpenAI Codex, GPT-6. The runtime does not expose a more specific model
identifier or context-window size. Used reasoning, repository
inspection, code editing, shell tools, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 13:01:45 -05:00
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00