Files
Devin FoleyandPaperclip 1815474597 fix: report pending execution phase at Stop timeout (#14990)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server owns each adapter execution and waits for it to settle
after Stop.
> - A Stop timeout reports that termination remains unverified.
> - Existing phase timings arrive only after their work completes, so a
stalled await has no timing.
> - This pull request samples the pending phase when the Stop timer
expires.
> - Operators can identify the pending operation without treating
diagnostics as stop proof.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The opt-in Sentry context for an unconfirmed adapter Stop timeout.

**Current behavior**

The timeout includes execution identity but no pending phase. A session
close, instruction collection, workspace restore, or diagnostic write
can remain pending without producing its completion timing.

**Proposed behavior**

Add a closed-list phase and elapsed milliseconds from the exact live
execution control. Sample them when the timeout fires. Report `unknown`
and a null age when the control or attribution is unavailable.

**Reason and benefit**

The next timeout can identify which operation is still pending. It does
not require task text, paths, provider output, or additional database
writes.

**Breaking changes**

No API or execution behavior change. The existing opt-in error context
gains two fields. Related public work: #14639 added Stop identity
diagnostics; #14866 and #14945 cover instruction cleanup and teardown
outcomes. This change adds pending attribution to those paths. The
native Stop work in #14802 remains separate.

## What Changed

- Add a bounded tracker per execution control. Token scopes support
nested and overlapping awaits. A late release cannot clear a newer
scope.
- Track adapter execution, ACP cancellation and settlement, diagnostic
writes, and host cleanup. Keep a coarse host scope until the executor
finishes.
- Sample only the matching current control and settlement promise at
timeout. Freeze the sanitized result. Use a monotonic clock and cap
elapsed time at one day.
- Test stalled operations, repeated Stop calls, stale and wrong-run
controls, callback failures, scope bounds, and the real Sentry SDK
context.

## Verification

- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- Six focused suites passed: 79 tests. They cover pending scopes, Stop
control ownership, real ACP settlement stalls, and Sentry context
isolation.
- The real Sentry SDK contract ran with the audited optional peer
`@sentry/node@10.71.0` installed outside the workspace. Valid phase and
elapsed values were exported; arbitrary labels and nonfinite elapsed
values were rejected.
- An independent agent reviewed the production diff and ran the focused
tests without blockers.
- `git diff --check` and a redacted Gitleaks scan passed. The local full
`pnpm test:run` was stopped during its large serial server batch to
avoid duplicating the sharded CI suite. No complete local broad-suite
pass is claimed.
- All CI gates passed on `c223237b58aa8d479d1c66d7d399de30270bac4f`: 54
successful checks and two expected Storybook skips. This includes the
full sharded test suite, typecheck, build, real Sentry SDK isolation,
browser tests, canary dry run, and the security scan after the PR became
ready for review.
- Greptile scored the same commit 5/5 with no actionable findings or
unresolved review threads.

## Risks

This is diagnostic instrumentation. Cancellation, deadlines, teardown
order, termination proof, and file recovery proof remain unchanged.
Unsupported or uninstrumented work uses a coarse phase. Tracker overflow
fails closed to `unknown`. The tracker emits no new run-log or Telemetry
event. The existing Sentry opt-in gate remains in place.

## Model Used

OpenAI Codex, GPT-6. The agent used code inspection, local command
execution, automated tests, and an independent agent review. The runtime
did not expose a more specific model identifier or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 16:55:38 -07:00
..

@paperclipai/adapter-utils

Shared utilities for Paperclip adapters: process spawning, environment injection, sandbox/SSH transport, workspace sync, and the round-trip helpers that move code between the local execution-workspace cwd and wherever the agent actually runs.

For the adapter-author guide see docs/adapters/creating-an-adapter.md and the in-repo notes at packages/adapters/AUTHORING.md.

Sandbox bridge deadlines

Command-managed bridge control operations (queue reads, input delivery, setup, and cleanup) have a host-enforced deadline of at most 30 seconds per shell command. A shorter configured timeout still applies. These operations cannot inherit the agent run's hours-long lifetime or wait forever on a provider that ignores its timeout. Long-lived agent session commands keep their own limits.

A control timeout reports a transport failure. It does not prove that the remote command stopped and does not replay an uncertain write. The process session closes failed input delivery and uses its existing shutdown path; normal execution settlement must still verify termination.

No-remote-git contract

The local execution-workspace cwd is the only persistence boundary across runs. No adapter may depend on a git remote for cross-run state.

Adapters that run the agent on a different host should use the SSH round-trip helpers in src/ssh.ts:

  • prepareWorkspaceForSshExecution({ spec, localDir, remoteDir }) — bundles the local cwd (tracked files, dirty edits, untracked additions, and the git history needed to reconstruct it) to remoteDir before the run starts. Runs with no git remote configured.
  • restoreWorkspaceFromSshExecution({ spec, localDir, remoteDir, ... }) — syncs the remote cwd back into localDir after the run, including any new commits the agent created. Also runs with no git remote configured.

prepareRemoteManagedRuntime in src/remote-managed-runtime.ts wraps both calls for adapters that want a per-run remote workspace and an automatic restoreWorkspace() finally hook.

The invariant is pinned by the no-remote-git contract case in src/ssh-fixture.test.ts, which asserts that a remote-only commit propagates to the local worktree through the prepare → restore round-trip with no git remote configured at any point. Do not regress that test.