mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner can execute a task inside a Daytona sandbox. > - The sandbox can keep running when the Paperclip controller restarts. > - Recovery treated sandbox process IDs as local process IDs and selected the wrong recovery path. > - Live verification also found races between startup, shutdown, and queued task cleanup. > - This pull request verifies the existing remote owner and orders those transitions. > - Users can continue the same task and provider session after a controller restart. ## Linked Issues or Issue Description **What happened?** The Daytona `recover-controller` cases failed with `runner_state_identity_mismatch`. Remote process IDs can be absent on the controller or collide with unrelated local processes. Recovery then looked for remote state in the local runner directory. Later turns could also start before the previous executor released its sandbox resources. **Expected behavior** Reconnect to the original sandbox and authenticated runner. Preserve the task, provider session, and queued comments. Reject a replacement sandbox or mismatched identity. Do not start another provider during reattachment. **Steps to reproduce** Run the `everyday-workflows` `recover-controller` case for `runner-codex` or `runner-acpx-claude` in Daytona. The browser creates a Python tool, requests a revision, restarts the controller during execution, and queues another revision. It then downloads and tests the final ZIP. Related: #13682 is the preceding operational fix. #13291 addresses legacy sandbox conversation recovery, a different execution path. #13666 includes broader run-capacity work; this change guards cleanup of an existing native task executor. ## What Changed - Add remote runner recovery without interpreting sandbox PIDs on the controller. - Verify the original provider lease, remote workspace, durable state, process marker, and authenticated PRP authority before adoption. - Compare the process marker with live Linux boot identity and start ticks to reject PID reuse. Read virtual proc files through the guaranteed Node runtime; unavailable proof blocks adoption without blocking a fresh launch. - Make the E2E supervisor own the actual server process so forced restart cannot leave a late database closer behind. - Scope the chat delivery lease test to its own fixture instead of draining other tests’ pending deliveries. - Preserve provider-attempt counts and recorded evidence during reattachment. - Serialize an idle-session checkpoint with admission of the next native turn. - Wait for an in-progress startup to acknowledge restart detachment. Fail after a bounded deadline if it cannot. - Keep a queued comment waiting until the previous native task executor releases its resources. Allow unrelated tasks to continue. - Update the Daytona image's resolved lock digest to match current dependency manifests. - Add classifier, ownership, process, startup, checkpoint, and queued-admission regression tests. Document recovery behavior. ## Verification - 415 focused tests passed across native execution, restart recovery, workspace synchronization, queued admission, and real-process restart tests. The final Node-based fingerprint change passed all 375 native-session tests. - Runner harness unit tests: 394 passed. Chat integration shard 2: 335 passed after fixture isolation. - The exact fingerprint command succeeded twice in a disposable Daytona sandbox and returned the same identity; the sandbox was deleted. - 11 real-process restart integration tests passed, including absent and colliding remote PIDs. - Repository typecheck and final build passed. Broad local checks found machine-dependent database startup and timing failures; focused retries passed. The final-revision PR pipeline is green. One unrelated browser shard hit a five-second blank-page timeout on the first run and passed its targeted retry. - Final-revision local headed browser E2E: `everyday-workflows.runner-acpx-claude.daytona.recover-controller` passed on attempt 1 in 4.7 minutes, **40/40 checks**. Manual browser inspection confirmed Done, all three ZIPs, and delivery of the queued follow-up. All three runs succeeded using the same provider session. The harness downloaded and independently tested the final artifact. - Final-revision Daytona campaign: https://github.com/paperclipai/paperclip/actions/runs/35463999611 — **Codex passed first attempt (4.8 minutes); ACPX Claude passed first attempt (6.1 minutes)**. Campaign aggregation/publication is finishing; both test jobs succeeded. - Greptile reviewed `beb08d8493b3286f5bb988dead369ff8c96a395d`: **5/5**, no open findings. - Staging browser verification is pending selection of a disposable staging instance and removal of a Chrome extension UI block. ## Risks - Recovery now depends on the original sandbox remaining available. A replacement or mismatched identity still blocks adoption. - Shutdown waits up to 30 seconds for a native startup to reach a safe detach point. An unfinished startup returns a clear failure instead of a false detach receipt. - Queued native work on the same task waits for cleanup. Unrelated tasks remain eligible. - The image digest update rebuilds the Daytona runtime image. No database migration or public API change is included. ## Model Used OpenAI Codex, GPT-6, with repository inspection, code execution, and browser tools. The runtime does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>