Files
PaperClipAI/doc/acp-run-lifecycle.md
Devin FoleyandPaperclip 8326e33ada Fix oversized sandbox process launch payloads (#13793)
## Thinking Path

> - Paperclip manages agent work and preserves context across retries.
> - Sandbox ACP runs encode their command and environment into one
launch value.
> - A long continuation can make that encoded value exceed Linux's exec
limit.
> - The launch shell then exits before the agent can initialize.
> - This PR transfers large command envelopes through a private
temporary file.
> - The agent receives the complete environment and can start normally.

## Linked Issues or Issue Description

Refs #13777.

**Bug description**

Sandbox tasks with long retry context fail during ACP initialization
with exit 127. The streamed bridge also drops the shell error that
explains the failure.

**Steps to reproduce**

Launch the streamed sandbox process bridge on Linux with one valid
110,000-byte environment value. Its base64 command envelope exceeds the
limit on one exec argument or environment string. Daytona reports
`argument list too long: env`, then exit 127.

**Expected behavior**

The command envelope must not make a valid child environment too large
to launch. A shell startup failure must retain its diagnostic in the run
log.

## What Changed

- Keep envelopes up to 64 KiB on the existing launch path. Upload larger
envelopes in bounded chunks inside a mode-0700 session directory. Set
the final payload file to mode 0600.
- Read the file without following symlinks and delete it before spawning
the child. Remove incomplete uploads on failure. Both streamed and
polled bridges use the same envelope.
- Preserve stderr when the launch shell fails before the wrapper emits a
terminal event. Emit a fixed terminal error and shutdown acknowledgement
if a payload cannot be read or parsed, without exposing its contents.
- Cover large environments, file permissions, payload deletion,
interrupted uploads, missing or malformed payloads, and startup
diagnostics. Document the transfer and cleanup behavior.

## Verification

- Both large-envelope regressions fail before the fix and pass after it.
- Targeted bridge, ACP engine, real-spawn, and stdin-race checks pass:
370 tests.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passes.
- A gated live Daytona probe on the current sandbox image reproduced
exit 127 with the old bridge. The fixed bridge launched the same command
successfully. A second probe completed real Claude ACP initialization
with a 110,000-byte context value. It did not run an agent task. All
temporary sandboxes were deleted.
- `pnpm -r typecheck` and `pnpm build` reach the unchanged Rust runner
step and stop because this machine has no `cargo` executable.
- Full CI passes on `ea816cf58d`: [run
35683853754](https://github.com/paperclipai/paperclip/actions/runs/35683853754).
All 53 checks pass; two optional checks are skipped. The local
full-suite run was stopped after equivalent CI suites passed; it has no
final local result.
- Greptile is 5/5 on `ea816cf58d`, with no unresolved review threads.
The branch is mergeable.

## Risks

Large envelopes require extra upload calls during startup. The temporary
data stays inside the private session directory and is removed before
child startup or during failure cleanup. Individual child environment
values still obey the operating system's native limits. No migration or
configuration change is required.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks;
full-workspace limits described above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 20:54:07 -07:00

5.7 KiB

ACP run lifecycle

This document describes how one ACP agent run starts, runs its turn, and releases its resources. It covers the run coordinator, the six run resources, the settlement order, the server-owned staging lease, and the known limitations.

The run coordinator

One run() coordinator owns one attempt. It sequences four steps in a fixed order:

  1. Startup. Bring up the runtime and establish the session.
  2. Turn. Run the agent turn against the ready runtime.
  3. Settlement. Release every resource the run acquired.
  4. Result reproduction. Return the external result to the caller.

The coordinator holds the routing table as data. It reads the startup outcome and routes on it:

  • A build failure or a partial-bridge failure throws from startup. The coordinator replays the startup-rollback entries and rethrows the original error.
  • A runtime-create, handshake, missing-handle, or configuration failure produces a settle result. The coordinator settles the resources, then reproduces the error result.
  • Every other startup outcome produces a ready result. The coordinator runs the turn, settles the resources, then reproduces the turn result.

The coordinator has no catch around the turn: the turn never rejects and returns a typed completion. The coordinator reproduces the external result AFTER settlement, so a caller never observes a result before the run releases its resources.

The six run resources

A run resource ledger holds the resources between acquisition and settlement. The ledger is the single owner. The closed set of run resources is:

  • acp_runtime — the composite of the runtime, the session handle, and the child process. runtime.close() is the one release boundary for all three.
  • staged_runtime — the workspace and the runtime staged into the sandbox, plus the one-time host-side cleanup of the staged temporary files.
  • control_bridge — the control-plane callback bridge.
  • agent_bridge — the agent process-session bridge.
  • managed_home — the per-run managed-home copy-back hook.
  • staging_lease — the per-session staging lease.

The host lane acquires only acp_runtime. The sandbox lane also acquires the staging and transport resources. Each resource enters the ledger the moment it exists, so the settlement always closes it, even on an early failure.

The settlement order

The sandbox agent bridge passes small command envelopes through its launch environment. When the encoded envelope exceeds 64 KiB, it uploads the envelope in bounded chunks to a private session directory instead. This avoids the Linux limit on one argument or environment string when retry context grows. The wrapper reads and deletes the file before spawning the agent; bridge teardown removes it if startup fails. The child's environment values remain unchanged. If the launch shell fails before the wrapper emits a protocol event, the run log retains the shell's stderr alongside the exit code.

The settlement sequence is the one live cleanup owner for every settled path. It claims the ledger once, makes the pure reuse decision, then runs the ordered steps:

  1. end_session — close every runtime the reuse decision did not transfer, and drop the warm entry when it closes.
  2. settle_reuse — perform the decision. A save transfers the reuse candidate to the site store; every other case discards it.
  3. stop_transport — stop both bridges in one settled batch.
  4. sync_back — run the site sync-back (the managed-home copy-back).
  5. release_staging_lease — release the staging lease last (see below).

An error policy governs every step: a step records its error, and the later steps still run. Every step is a no-op on an empty resource slot.

The reuse decision is pure and runs before any close. A save is eligible only when a candidate exists, the settlement cause permits a save, and no run-scoped credential that the candidate would carry into the store remains valid.

The server-owned staging lease outer context

The staging lease is a per-session lease. Only one run of a session may stage into the same remote workspace at a time. A second run of the same session waits on the lease until the first run releases it.

The lease releases as the run's final act, after the coordinator settles every other resource AND reproduces the result. This ordering keeps a same-session second run blocked on the lease until the first run fully returns, so the second run never re-stages into a workspace the first run still uses. The release runs in a finally, so an earlier teardown fault never strands the lease.

Per-phase run-log events

The run writes one run-log event per named lifecycle phase, to the heartbeat_run_events table. This event is not a Paperclip Telemetry event and not an OpenTelemetry export. Each event carries only the phase name, the wall-time duration, and the outcome (ok or failed). The phase name is one member of a closed allowlist. An event never carries a command, an argument, a path, an environment value, or a raw identifier. A run-log write failure never fails the run.

Known limitations and deferred work

  • Host-lane runtime reuse is disabled. A run-minted API key is a stateless token that the control plane never revokes on run end. A warm host runtime would carry a still-valid credential into the next run. The host lane therefore closes and relaunches on every run, rather than keeping the runtime warm, until a run-scoped credential rebind protocol exists.
  • The sandbox staged-files reuse stays enabled. Its reuse payload carries no credential, so a compatible resume reuses the already-staged runtime.
  • The per-phase run-log events record the duration and the outcome only. They never change run control flow.