Files
PaperClipAI/packages
Devin FoleyandPaperclip f2e0f19630 Defer agent directory cleanup until stop proof is available (#14866)
## Thinking Path

> - Paperclip manages agents and their persistent files.
> - Each run owns a temporary agent directory and a save receipt.
> - Cleanup needs independent proof that the owning process stopped.
> - A cleanup call without that proof currently waits for the directory
lock anyway.
> - A second lock failure can prevent environment release after the run
already reported a failed save.
> - This change skips cleanup that has no authority and retries
unavailable remote copies after exact destruction proof.
> - The save failure stays visible. Existing lock owners remain
protected.

## Linked Issues or Issue Description

Related work: Refs #14787 (lock diagnostics), #14695 (warm instruction
ownership), #9667 (stale lock proposal), and #9872 (control-plane
ownership proposal). I checked open PRs and issues. This change leaves
the shared filesystem lock protocol in place and does not duplicate the
warm-retention work in #14695.

**What happened?**

Heartbeat cleanup records an explicit unavailable instruction-save
warning, then calls directory release before releasing the environment
lease. Release can wait for a lock even though the copy has no
process-stop proof and cannot be removed. That secondary timeout
prevents the following lease-release step. If destruction proof arrives
later, the unavailable copy is excluded from both recovery queries.

**Expected behavior**

Skip a release that cannot remove anything. Preserve the failed-save
receipt and candidate fields. Once exact remote destruction is recorded,
recover remote cleanup without running a provider command. Unavailable
local copies retain their potentially uncollected edits even if local
stop proof arrives later. A blocked cleanup must not prevent cleanup for
other agents.

**Steps to reproduce**

1. Prepare an agent directory, report its save unavailable, and leave
process-stop proof absent.
2. Hold the shared directory lock and call release. Before this change,
release waits and fails although removal is not authorized.
3. Record destruction of the copy's exact remote lease. Before this
change, neither recovery sweep selects the unavailable copy.

**Paperclip version or commit**

Reproduced against `efc2e6810e9bc0dc8cb412b0e7647c0db9821caa`.

**Deployment mode**

Local and remote execution with persistent agent directories. Tests use
an isolated embedded PostgreSQL database and fixture transports.

## What Changed

- Re-read receipts and skip release before lock acquisition when stop
proof is absent, the copy is superseded, or cleanup is complete. Keep
the same checks inside the lock.
- Recover unavailable remote copies only after exact destruction proof.
Preserve their unavailable state, errors, candidate hash, and candidate
bytes. Keep unavailable local copies and their uncollected edits
unchanged.
- Store destruction-only cleanup authority with the stop proof. Later
cleanup honors it after a lost database response or restart, including
when a transport remains cached.
- Defer failed or unproven cleanup with bounded batches and a retry
delay. Keep failed cleanup visible in logs and its receipt.
- Serialize preparation of an existing run with cleanup. Fresh run
preparation keeps its existing admission path.
- Cover held locks, receipt scope, delayed proof, batch fairness, lost
update responses, cached transports, and concurrent same-run preparation
with database regressions.

## Verification

- Focused directory, legacy instruction-copy, shared lock, and bounded
diagnostic suites: 169 tests passed across four files.
- `pnpm -r typecheck`: passed on the final source.
- `pnpm build`: passed on the final source.
- Completed all selected local `pnpm test:run` groups: 733 general
server suites, 149 serialized suites, and 14 workspace projects. There
are 13 known macOS `EACCES` failures in the unchanged runtime skill
cache tests. Their exact signatures match earlier clean-base results,
and the cache source and test blobs match both that base and this PR
base (existing fix: #14290). One CLI import test timed out under
concurrent load; its full file passed separately (17 tests). Broad
coverage began before the review corrections; the final source has the
focused 169-test run, typecheck, and build. This is a local verification
limit, not a passing full local suite.
- `git diff --check` and local Gitleaks plus private-identifier/PII diff
scans passed.
- Independent review of the final source found no remaining actionable
issue. Its 17 targeted tests cover crash recovery, cached transports,
same-run preparation, real local edit preservation, proof scope, and
batch fairness. The main focused run also covers contained scheduling
failures.
- Final commit `35a24085f7`: Greptile 5/5 with no recommendations and
zero unresolved review threads.
- Final commit `35a24085f7`: all 53 checks passed, including Canary Dry
Run and the security scan; two visual checks were intentionally skipped.
The workspace shard passed on retry after GitHub reported that its first
runner lost communication. An earlier Canary runner shut down after the
release dry run passed. Neither interruption recorded an application
assertion failure; the exact final-head checks are now green.

## Risks

- This repairs cleanup ordering and recovery eligibility. It does not
repair an ambiguous legacy lock owner or restore unsaved files. Actual
collection still fails visibly when its lock cannot be acquired.
- An unavailable remote copy is recovered only after exact destruction
proof. A stopped but retained environment stays protected; recovery does
not execute a command that could restart it.
- Unavailable local copies with later stop proof still retain
potentially uncollected edits. A general local recollection or
reclamation policy remains outside this change.
- Existing-run preparation now waits for the same lock as cleanup. The
fresh-run path is unchanged.
- The cleanup mode is stored in the existing private receipt JSON. No
schema migration or public API change is required.
- No deployment, task replay, or runtime lock deletion was performed.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and test
execution. The runtime does not expose a more specific model suffix or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 14:28:47 -07:00
..