## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents on remote sandbox targets reach the Paperclip API through the sandbox callback bridge: a loopback HTTP gateway inside the sandbox queues request files for a host-side worker > - The gateway process has no supervisor: nothing inside the sandbox respawns it, so a crash leaves a dead loopback port for the rest of the run > - The gateway also never cleaned up request files whose responses never arrived, so a stalled host wedged the queue at its depth cap and every later request got an immediate 503 > - #12052 made the host-side worker survive transient faults; this pull request hardens the other half of the relay > - The benefit is that a gateway fault degrades one request instead of severing the agent from the control plane until run end ## Linked Issues or Issue Description Refs #12052 (host-side worker half of the same relay). Refs #9904 and #8977 (adjacent bridge behavior). No public issue exists for this defect. The description below follows the bug report template. **What happened?** During a staging run, an agent's API calls to the bridge's loopback port began failing at the connection level (curl reported HTTP 000) partway through the run. A dead gateway process is the only mechanism that produces connection-level failures on that port, and nothing restarts it. Separately, request files for timed-out requests stayed in the queue; after 64 accumulated, the gateway answered every request with `503 Bridge request queue is full.` until the run ended. **Expected behavior** An uncaught fault in the gateway must not kill the loopback listener. A request that times out must not leave its file counting toward the queue-depth cap. A queue full of orphaned files must recover instead of rejecting until run end. **Steps to reproduce** 1. Start a remote-sandbox run and stop the host-side bridge worker. 2. Send requests to the gateway until they time out; the request files stay in `requests/`. 3. After 64 such files, every request gets an immediate 503, even after the host recovers. 4. Independently, raise any uncaught exception in the gateway process; the loopback port dies for the rest of the run. ## What Changed - The generated gateway source installs global `uncaughtException` / `unhandledRejection` handlers that log to stderr (already redirected to `logs/bridge.log`) and keep serving. The relay holds no state a fault can corrupt beyond the one request it interrupted. - Survival is gated on readiness: before the gateway has written its readiness file (file mode) or sent its READY frame (duplex mode), the same handlers exit(1) instead. A startup fault (failed bind, failed readiness write) means the process can never serve, and surviving there would only leave an un-ready zombie while the host waits out its readiness poll. - The file gateway attaches an explicit `error` listener to its server and pins the event loop with a keepalive until the bind settles. Newer Node runtimes do not reliably surface a failed bind through `uncaughtException` in this shape: the process can drain and exit 0 before the error event is delivered (reproduced on Node 24/25; Node 22 delivered it). The duplex gateway already had an explicit listener. - A request that times out waiting for the host now deletes its own request file. The host's response write is guarded on that file, so the removal also signals that no caller waits anymore. - At the queue-depth cap, the gateway sweeps request files older than the response deadline (orphans from killed callers or a previous gateway process) before rejecting with 503. - Host-side, `processRequestFile` treats a request file that vanished before the read as the benign caller-gave-up race and skips it quietly instead of escalating into the recovery pass. ## Verification - `npx vitest run packages/adapter-utils/src/sandbox-callback-bridge.test.ts` — 43 passed, verified on both Node 22 and Node 25. - New end-to-end test: with no worker running, a request times out (502), its file is cleaned, and the same gateway then serves a 200 once a worker starts — no wedge, no dead port. - New end-to-end test: with `maxQueueDepth: 1` and a backdated orphan file at the cap, the gateway sweeps the orphan and admits the request instead of answering 503. - New worker test: a request file that vanishes before the read is skipped without a handler call, a response write, or a run-level error. - New generated-source test: spawned directly against an already-occupied port, the gateway exits 1 promptly with the `EADDRINUSE` fault on stderr instead of lingering un-ready (or exiting 0 silently, the pre-existing behavior on Node 24/25). - A pin keeps the crash handlers, the readiness gate, and the sweep in the generated source. - `pnpm --filter @paperclipai/adapter-utils typecheck`. ## Risks - Keeping a Node process alive after `uncaughtException` is normally suspect; here the alternative is a dead loopback port for the rest of the run, and the gateway is a stateless per-request relay. The fault is logged with its stack to `bridge.log`, and survival applies only after readiness — startup faults still fail fast. - Deleting a timed-out request file could race a host that is mid-processing. The host's response write is already guarded on request-file existence, and the new host-side skip treats the vanished file as a no-op, so no duplicate mutation path is introduced. - The stale sweep runs only at the depth cap and only removes files older than the response deadline plus a 2 s grace, so a live caller's file is never swept. - Orphaned response files (host responded after the caller gave up) still linger; that pre-existing minor leak is unchanged here. ## Model Used - Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended thinking enabled, agentic tool use via Claude Code (CLI harness), 200k context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
@paperclipai/adapter-utils
Shared utilities for Paperclip adapters: process spawning, environment injection, sandbox/SSH transport, workspace sync, and the round-trip helpers that move code between the local execution-workspace cwd and wherever the agent actually runs.
For the adapter-author guide see
docs/adapters/creating-an-adapter.md
and the in-repo notes at packages/adapters/AUTHORING.md.
No-remote-git contract
The local execution-workspace cwd is the only persistence boundary across runs. No adapter may depend on a git remote for cross-run state.
Adapters that run the agent on a different host should use the SSH round-trip
helpers in src/ssh.ts:
prepareWorkspaceForSshExecution({ spec, localDir, remoteDir })— bundles the local cwd (tracked files, dirty edits, untracked additions, and the git history needed to reconstruct it) toremoteDirbefore the run starts. Runs with nogit remoteconfigured.restoreWorkspaceFromSshExecution({ spec, localDir, remoteDir, ... })— syncs the remote cwd back intolocalDirafter the run, including any new commits the agent created. Also runs with nogit remoteconfigured.
prepareRemoteManagedRuntime in
src/remote-managed-runtime.ts wraps both
calls for adapters that want a per-run remote workspace and an automatic
restoreWorkspace() finally hook.
The invariant is pinned by the no-remote-git contract case in
src/ssh-fixture.test.ts, which asserts that a
remote-only commit propagates to the local worktree through the
prepare → restore round-trip with no git remote configured at any point. Do
not regress that test.