## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - A workspace restore merges a directory, and a cross-process lock
serializes that merge
> - The lock must recover after the process that holds it crashes
> - One test proves that recovery: it kills the holder process and then
acquires the lock
> - That test failed intermittently for two independent reasons, and
this pull request removes both
> - First, it spawned the holder through the tsx command-line entry
point, which re-spawns the evaluated code in a further child process, so
the kill signal reached only the wrapper and the real holder kept the
lock
> - Second, it replaced the global clock to force a timeout, which left
the acquisition with zero real retries, so a single transient busy
result failed the test
> - The benefit is a deterministic crash-recovery test and a reliable
continuous-integration signal
## Linked Issues or Issue Description
**What happened?**
The test `recovers a killed holder even when its recorded PID has been
reused` in `packages/adapter-utils/src/directory-merge-lock.test.ts`
failed intermittently in continuous integration. The failure reported
`ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT` with `waitMs: 3` and
`knownLocalHolder: false`. A rerun of the same job on the same commit
passed.
**Expected behavior**
The test must pass every run. It must acquire the lock after the holder
process dies.
**Steps to reproduce**
1. Check out `master`.
2. Run `npx vitest run
packages/adapter-utils/src/directory-merge-lock.test.ts`.
3. Repeat the run. The named test fails intermittently.
**Paperclip version or commit**
`32e9f3ba0ec000578936731990d23bb0e77493fa`
**Deployment mode**
Built from source. The failure appears in the general test job of
continuous integration.
**Agent adapter(s) involved**
Not adapter-specific (core bug).
**Relevant logs or output**
```
ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT
workspaceRestoreLock: { ownerState: 'alive', knownLocalHolder: false, waitMs: 3,
ownerSameProcess: true, ownerAgeMs: 60030, ownerPredatesProcess: true }
```
## What Changed
The test file had two independent defects. This pull request removes
both.
**1. The kill signal did not reach the real lock holder.**
The test spawned its holder through the tsx command-line entry point.
That entry point re-spawns the evaluated code in a further child
process. `SIGKILL` therefore killed only the wrapper, and the process
that had opened the lock database survived as an orphan that still held
the lock. The test now loads tsx as an `--import` hook, so the spawned
process is the real holder and the kill releases the lock at once. This
also stops the test from leaking an orphan process.
**2. The forced clock left the acquisition with zero retries.**
A test helper replaced `Date.now` to force a timeout. The implementation
reads `Date.now()` one time, to compute its deadline, so that single
read consumed the forced value and every later read returned a time
already past the deadline. The retry loop therefore got one attempt and
no retries. That is correct for a test that asserts a timeout, but the
crash-recovery test asserts a *successful* acquisition, so any transient
busy result on the first attempt failed it.
The fix removes the clock replacement from the whole file and gives each
test a real, short, explicit wait budget:
- `withDirectoryMergeLock` takes a new optional wait-budget parameter.
It threads through to the lock acquisition function. The production
default is the existing 30-second budget, and no production call site
changed.
- The five tests that assert a timeout pass a real 200-millisecond
budget. Each one still times out for the real reason, because the lock
is genuinely held or the legacy lock directory genuinely exists. Each
one now exercises at least four real retries of the 50-millisecond retry
interval.
- The crash-recovery test passes a real 5-second budget. A failure now
reports the structured `ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT` diagnostic
well inside the test timeout, instead of a bare test timeout.
**No test timeout increased.** Every `it(..., N)` timeout in the file
equals its value on `master`.
## Verification
- Measured the first cause rather than assumed it: the spawned wrapper
process reported one process id, and the process that opened the lock
database reported a different process id and named the wrapper as its
parent. The real holder kept the lock for about 50 to 60 milliseconds
after the kill.
- Reproduced the failure deterministically before the change, with no
artificial processor load: 15 of 15 runs failed. Confirmed the fix: 15
of 15 runs passed.
- Ran the lock test file 15 times in series: 12 of 12 tests passed every
time.
- Confirmed the clock replacement is gone: a search for a `Date.now`
override in the file returns nothing.
- Confirmed the production default is unchanged at 30 seconds, and that
the diff touches no production call site.
- Proved the diagnostic still surfaces: with a temporary edit that held
the lock with a genuine live holder, the test failed with
`ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT` and the full `workspaceRestoreLock`
block at about 5 seconds, inside the 15-second test timeout. The
temporary edit was reverted.
- `workspace-restore-merge.test.ts` passed 56 of 56. The adapter test
files that cover every production caller passed 116 of 116 and 52 of 52.
`agent-directory-working-copies.test.ts` passed 70 of 70.
- The `adapter-utils` and `server` type-checks passed with no error.
- Confirmed that no spawned process survives the test run.
## Risks
Low risk. The production change is one optional parameter with the
existing default, so every production caller keeps the real 30-second
budget and no production call site changed. The remaining change is
limited to one test file. The `--import` form of the tsx hook is already
used elsewhere in this repository, in the container image command and in
an end-to-end test configuration. Test coverage does not drop: the owner
record is diagnostic only, the SQLite reserved lock remains the
authority that the tests exercise, and the timeout-asserting tests now
exercise the real retry loop instead of a replaced clock. The file costs
about 0.5 to 0.9 seconds more wall clock than `master`, which is the
cost of the short real waits that replace the instant forced timeout.
## Model Used
Claude Sonnet 5 (`claude-sonnet-5`), used with extended thinking and
tool use for the diagnosis, the measurement, and the change.
## Checklist
Check every box that the state of the pull request satisfies. The local
test runs and the type checks are complete. Reconcile the
continuous-integration and review boxes after the checks reach their
terminal state.
## Test plan
- [x] Continuous integration is green on every check, including the
general test job.
- [x] The general test job passes the file
`packages/adapter-utils/src/directory-merge-lock.test.ts`.
- [x] Greptile returns 5 of 5 with no open item.
- [x] `mergeable: MERGEABLE` is terminal.
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
@paperclipai/adapter-utils
Shared utilities for Paperclip adapters: process spawning, environment injection, sandbox/SSH transport, workspace sync, and the round-trip helpers that move code between the local execution-workspace cwd and wherever the agent actually runs.
For the adapter-author guide see
docs/adapters/creating-an-adapter.md
and the in-repo notes at packages/adapters/AUTHORING.md.
No-remote-git contract
The local execution-workspace cwd is the only persistence boundary across runs. No adapter may depend on a git remote for cross-run state.
Adapters that run the agent on a different host should use the SSH round-trip
helpers in src/ssh.ts:
prepareWorkspaceForSshExecution({ spec, localDir, remoteDir })— bundles the local cwd (tracked files, dirty edits, untracked additions, and the git history needed to reconstruct it) toremoteDirbefore the run starts. Runs with nogit remoteconfigured.restoreWorkspaceFromSshExecution({ spec, localDir, remoteDir, ... })— syncs the remote cwd back intolocalDirafter the run, including any new commits the agent created. Also runs with nogit remoteconfigured.
prepareRemoteManagedRuntime in
src/remote-managed-runtime.ts wraps both
calls for adapters that want a per-run remote workspace and an automatic
restoreWorkspace() finally hook.
The invariant is pinned by the no-remote-git contract case in
src/ssh-fixture.test.ts, which asserts that a
remote-only commit propagates to the local worktree through the
prepare → restore round-trip with no git remote configured at any point. Do
not regress that test.