mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A user can send a new message after a native run fails. > - The server checks that the old execution has stopped before it starts a fresh turn. > - A failed run can retain a result accepted before checkpoint or cleanup failed. > - The continuation gate treated that saved result as active recovery and held the new message forever. > - This pull request removes that false liveness signal while retaining controller, process, environment, and authorization checks. > - Live staging then exposed a second defect: chat admission created a successor without consuming the original deferred receipt, so completion delivered the message again. > - Consume that exact receipt atomically with admission, while preserving separate turns for later chat messages. ## Linked Issues or Issue Description **What happened?** A new user message stayed in the queue with `controller_settling` after the previous run had reached `terminal_failure`. The old coordinator had no lease owner but still had a `resultId`. Its remote environment had a verified stop receipt. **Expected behavior** Start one fresh turn after execution has stopped and normal admission checks pass. Preserve the failed run and its accepted result as history. **Steps to reproduce** 1. Accept a native result, then fail checkpoint or cleanup and exhaust recovery. 2. Retain the result ID on the terminal failure record and stop the execution environment. 3. Send a new user message. Before this fix, it waits forever for the finished controller. **Paperclip version or commit** Reproduced in a database-backed regression test on `26900655b`. **Deployment mode** Server with a native runner and remote sandbox. Local process stop checks also apply. Related: https://github.com/paperclipai/paperclip/pull/14775. Searched existing PRs for retained-result continuation fixes; no duplicate found. ## What Changed - Remove the retained-result veto for terminal failures. - Keep controller ownership, successor, process, environment cleanup, pending decision, and ordinary admission checks. - Add regressions for retained results, active execution, missing stop evidence, and delayed remote cleanup. - Atomically consume the resumed receipt in agent chat, even though chat does not coalesce other queued messages. - Reproduce completion-time duplicate promotion, race cleanup against periodic recovery, and prove a subsequent chat message keeps its own turn. - Document that a saved result does not make a terminal failure active. - Keep exhausted workspace export on its separate repair path, tested through the production finalizer. ## Verification - Red: retained-result admission failed with `controller_settling` before the original fix. The new chat-specific regression then reproduced duplicate promotion when the first reply finished. - Green: 406 tests across native continuation, workspace-export recovery, and the wake-queue module passed on `cbc531cc0`. - The chat regressions exercise real Postgres transactions, simultaneous recovery callbacks, successful completion, the production queue-drain use case, and repeated drain attempts. A distinct follow-up remains a separate turn. - `pnpm -r typecheck` and `pnpm build` passed on `cbc531cc0`. - The earlier full local test run encountered a timeout and follow-on failure in unchanged AI connection-adoption tests; all 50 tests passed on isolated rerun. That local run was stopped after the full CI test matrix passed on the earlier head. - All 54 CI checks passed on `cbc531cc0` (2 skipped), including the full test matrix and browser shards. One unchanged interaction-route test returned HTTP 500 on its first CI attempt; its full 84-test file passed locally, and the failed shard passed on one targeted rerun. - Greptile reviewed `cbc531cc0`: 5/5, no unresolved findings. - Live staging first verified that the original saved message resumes and receives a successful response; that test exposed the duplicate now covered above. - Deployed exact commit `cbc531cc0410e1ef6e8811c6c5c014c3528351ed` to the affected staging workspace; deployment verification, health, authentication, and startup recovery passed. - Submitted a fresh message through the browser. The agent replied in 39 seconds; server records show exactly one successful run, native phase `committed`, no error, and an empty queue. A later check more than a minute after completion found no duplicate run. ## Risks The change affects admission after native execution failure and consumption of a resumed deferred receipt. A fresh turn must never overlap the prior execution, and consuming one chat receipt must not absorb later messages. Tests retain the controller, process, and remote-stop guards. This change does not migrate data, apply an old result, or reset the old retry budget. ## Model Used OpenAI Codex (GPT-6). The exact runtime model identifier and context window are not exposed in this session. Used reasoning, repository inspection, code execution, database-backed tests, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>