Files
PaperClipAI/packages
Devin Foley 87d68f476b fix: harden the sandbox bridge gateway against crashes and queue wedge (#12060)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents on remote sandbox targets reach the Paperclip API through the
sandbox callback bridge: a loopback HTTP gateway inside the sandbox
queues request files for a host-side worker
> - The gateway process has no supervisor: nothing inside the sandbox
respawns it, so a crash leaves a dead loopback port for the rest of the
run
> - The gateway also never cleaned up request files whose responses
never arrived, so a stalled host wedged the queue at its depth cap and
every later request got an immediate 503
> - #12052 made the host-side worker survive transient faults; this pull
request hardens the other half of the relay
> - The benefit is that a gateway fault degrades one request instead of
severing the agent from the control plane until run end

## Linked Issues or Issue Description

Refs #12052 (host-side worker half of the same relay). Refs #9904 and
#8977 (adjacent bridge behavior).

No public issue exists for this defect. The description below follows
the bug report template.

**What happened?**

During a staging run, an agent's API calls to the bridge's loopback port
began failing at the connection level (curl reported HTTP 000) partway
through the run. A dead gateway process is the only mechanism that
produces connection-level failures on that port, and nothing restarts
it. Separately, request files for timed-out requests stayed in the
queue; after 64 accumulated, the gateway answered every request with
`503 Bridge request queue is full.` until the run ended.

**Expected behavior**

An uncaught fault in the gateway must not kill the loopback listener. A
request that times out must not leave its file counting toward the
queue-depth cap. A queue full of orphaned files must recover instead of
rejecting until run end.

**Steps to reproduce**

1. Start a remote-sandbox run and stop the host-side bridge worker.
2. Send requests to the gateway until they time out; the request files
stay in `requests/`.
3. After 64 such files, every request gets an immediate 503, even after
the host recovers.
4. Independently, raise any uncaught exception in the gateway process;
the loopback port dies for the rest of the run.

## What Changed

- The generated gateway source installs global `uncaughtException` /
`unhandledRejection` handlers that log to stderr (already redirected to
`logs/bridge.log`) and keep serving. The relay holds no state a fault
can corrupt beyond the one request it interrupted.
- Survival is gated on readiness: before the gateway has written its
readiness file (file mode) or sent its READY frame (duplex mode), the
same handlers exit(1) instead. A startup fault (failed bind, failed
readiness write) means the process can never serve, and surviving there
would only leave an un-ready zombie while the host waits out its
readiness poll.
- The file gateway attaches an explicit `error` listener to its server
and pins the event loop with a keepalive until the bind settles. Newer
Node runtimes do not reliably surface a failed bind through
`uncaughtException` in this shape: the process can drain and exit 0
before the error event is delivered (reproduced on Node 24/25; Node 22
delivered it). The duplex gateway already had an explicit listener.
- A request that times out waiting for the host now deletes its own
request file. The host's response write is guarded on that file, so the
removal also signals that no caller waits anymore.
- At the queue-depth cap, the gateway sweeps request files older than
the response deadline (orphans from killed callers or a previous gateway
process) before rejecting with 503.
- Host-side, `processRequestFile` treats a request file that vanished
before the read as the benign caller-gave-up race and skips it quietly
instead of escalating into the recovery pass.

## Verification

- `npx vitest run
packages/adapter-utils/src/sandbox-callback-bridge.test.ts` — 43 passed,
verified on both Node 22 and Node 25.
- New end-to-end test: with no worker running, a request times out
(502), its file is cleaned, and the same gateway then serves a 200 once
a worker starts — no wedge, no dead port.
- New end-to-end test: with `maxQueueDepth: 1` and a backdated orphan
file at the cap, the gateway sweeps the orphan and admits the request
instead of answering 503.
- New worker test: a request file that vanishes before the read is
skipped without a handler call, a response write, or a run-level error.
- New generated-source test: spawned directly against an
already-occupied port, the gateway exits 1 promptly with the
`EADDRINUSE` fault on stderr instead of lingering un-ready (or exiting 0
silently, the pre-existing behavior on Node 24/25).
- A pin keeps the crash handlers, the readiness gate, and the sweep in
the generated source.
- `pnpm --filter @paperclipai/adapter-utils typecheck`.

## Risks

- Keeping a Node process alive after `uncaughtException` is normally
suspect; here the alternative is a dead loopback port for the rest of
the run, and the gateway is a stateless per-request relay. The fault is
logged with its stack to `bridge.log`, and survival applies only after
readiness — startup faults still fail fast.
- Deleting a timed-out request file could race a host that is
mid-processing. The host's response write is already guarded on
request-file existence, and the new host-side skip treats the vanished
file as a no-op, so no duplicate mutation path is introduced.
- The stale sweep runs only at the depth cap and only removes files
older than the response deadline plus a 2 s grace, so a live caller's
file is never swept.
- Orphaned response files (host responded after the caller gave up)
still linger; that pre-existing minor leak is unchanged here.

## Model Used

- Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended
thinking enabled, agentic tool use via Claude Code (CLI harness), 200k
context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-24 08:52:05 -07:00
..