Commit Graph
86 Commits
Author SHA1 Message Date
Devin FoleyandPaperclip 2e67ea8dc7 fix: renew explicit continuation authorization for bounded retries (#14987)
Preserve exact user-authorized continuation receipts across bounded retries and retain the task claim through owned retry admission. Stop, reassignment, superseding input, and active cleanup continue to block unsafe continuation.

Validation: exact-head Greptile 5/5, passing CI, no unresolved review threads, and a clean merge. Detailed verification and limitations are recorded in the pull request.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 10:57:57 -07:00
DottaandPaperclip ffe5e9e2a8 fix: retain execution evidence for Retry and saved input (#15033)
## Thinking Path

> - Paperclip lets people steer and recover AI-agent conversations.
> - Recovery eligibility depends on retained cancellation receipts.
> - Run presentation intentionally omits result JSON on SQL_ASCII
databases and reduces oversized output.
> - Retry and the saved-input sweep mistakenly used that presentation
read for admission.
> - The banner could offer Retry while the endpoint rejected the same
stopped run.
> - Read the narrow execution-evidence fields for internal admission and
keep normal presentation unchanged.

## Linked Issues or Issue Description

Refs #15024 and #15015. Searched existing recovery and redaction PRs; no
duplicate fix found.

**What happened?**

On a SQL_ASCII instance, a verified pre-dispatch review-wait
cancellation offers Retry in the recovery notice. Clicking it returns an
eligibility conflict, and a saved user message remains deferred. The
notice reads the retained database receipt, but the endpoint and sweep
read a presentation projection where `resultJson` is null.

**Expected behavior**

Retry and saved user input use the recorded execution evidence and
ordinary admission gates, independent of presentation redaction. Public
run reads retain their existing encoding and output-size protections.

**Steps to reproduce**

1. Record a cancelled, unclaimed review-wait continuation and its
recovery hold.
2. Use the SQL_ASCII presentation projection, where run result JSON is
omitted.
3. Click Retry or save a new user message and let the recovery sweep
inspect it.
4. Verify a fresh turn starts once, with no replay of consumed input.

**Paperclip version or commit**

Reproduced on `9ae3d8db3`.

**Deployment mode**

Authenticated private self-hosted server with a SQL_ASCII database.

## What Changed

- Add an explicit internal read of cancellation, startup, review-wait,
tool-inventory, and Stop evidence; omit provider diagnostics.
- Use that read in the Retry route, wakeup validation, and saved-input
continuation checks.
- Preserve the distinction between absent result JSON and an
unrecognized stored result.
- Cover the reproduced SQL_ASCII Retry and saved-input failures,
retained public redaction, native Stop behavior, and excluded provider
output.
- Document the presentation and admission distinction.

## Verification

- Red: four selected assertions fail before the fix, including the
SQL_ASCII eligibility conflict and saved input remaining deferred.
- Targeted green regressions and existing native Stop cases pass.
- Workspace `pnpm -r typecheck` and `pnpm build` pass.
- All 626 affected recovery, continuation, and route tests pass,
including 22 focused admission and native Stop cases. Current-head CI
has 54 passing gates and 2 skipped optional Storybook checks. Greptile
reviewed `9b94af91e295a4e8007dfc6bff6a8d532d945e36` at 5/5 with no
findings or open review threads. No complete local monolithic pass is
claimed; the complete suite runs in sharded CI.

## Risks

Admission still checks recorded process and controller ownership,
provider events, cleanup, company scope, user authority, pending
decisions, and task holds. The evidence projection must retain every
field used by these eligibility predicates; existing native Stop cases
guard against dropping its acknowledgement receipt. No schema,
dependency, UI, or public response change.

## Model Used

OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and
context window are not exposed in this session. Used reasoning,
repository tools, code execution, and browser inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 01:38:57 -05:00
DottaandPaperclip 9ae3d8db3d fix: keep pre-dispatch review waits out of execution recovery (#15024)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The queued-run gate can cancel a continuation that must wait for
review.
> - This cancellation happens before execution authority or a provider
starts.
> - Recovery currently treats the gate receipt as unknown provider
execution.
> - That mistake blocks a conversation after a successful reply and
hides Retry.
> - This pull request recognizes only the recorded, unclaimed
review-wait state.
> - Review waits keep their normal disposition path, and new user input
can recover older false holds.

## Linked Issues or Issue Description

Refs #15015, #15020, and #15022. Related #11614 narrows the review
posture that causes cancellation; this change corrects recovery after a
valid cancellation.

**What happened?**

After a successful agent reply, the queued-run gate cancelled an
automatic continuation with `issue_continuation_waiting_on_review`. The
gate retained `timeoutSource: stale_queued_run_gate` and a matching stop
reason. No execution authority or provider started. Recovery still
created an unknown-action hold, moved the task to Blocked, and hid
Retry.

**Expected behavior**

Use normal review-wait disposition repair for this recorded state. Allow
Retry, a new user message, or saved undelivered input to recover an
older false hold after ordinary admission checks pass. Preserve the
cancelled run and do not replay its input.

**Steps to reproduce**

1. Finish an agent turn on an open task that has a real review target.
2. Let the automatic continuation reach the queued-run review gate.
3. Refresh after the cancelled run is checked by recovery.
4. Confirm a review wait is handled as a wait rather than unknown
provider work.
5. Reproduce an older false hold for the same receipt, then request
Retry or send a new message.
6. Confirm only one fresh turn starts, and contradictory execution or
cleanup evidence retains the hold.

**Paperclip version or commit**

Reproduced on `215586d127`.

**Deployment mode**

Authenticated private self-hosted server, built from source.

## What Changed

- Recognize the exact review-wait dispatch receipt only while all
execution claims remain unset.
- Exempt recovery only after checking retained launch and provider
events, coordinators, and environment cleanup. Keep the synchronous
classifier conservative without that database proof.
- Apply the verified classification to automatic recovery, heartbeat
retries, and stranded-queue release, so saved user input starts once.
- Reuse guarded startup admission for Retry, new input, and saved input
on older false holds.
- Show a precise review-wait notice and keep continuation guidance
consistent with Retry availability.
- Verify provider events, launch events, coordinators, and cleanup
before admission.
- Add five initial red regressions, three additional red recovery
evidence regressions, concurrent saved-input coverage, and negative
evidence checks.
- Document the review-wait contract.

## Verification

- Red: five regressions fail on unchanged master. Existing review-wait
and contradictory-evidence cases still pass.
- Workspace `pnpm -r typecheck` passes.
- All 752 affected recovery, continuation, queue, classification, and
retry-scheduling tests pass.
- Three additional review regressions failed before the database proof
was added; all 12 focused review-wait cases then pass.
- Workspace `pnpm build` passes.
- Saved-input promotion and recovery notice regressions fail before
their fixes and pass afterward.
- Current-head CI has 54 passing checks and 2 skipped optional Storybook
checks. Greptile reviewed `0bf1437110e1617ae4544afa41f4319eac763dba` at
5/5 with zero open findings. The complete suite runs in sharded CI. No
complete local monolithic pass is claimed; an earlier long run was
stopped, and its unrelated failing case passed in isolation.

## Risks

The error code alone cannot establish that no provider started. This
exception also requires the server gate receipt, matching stop reason,
and null execution authority fields. User admission separately checks
coordinator, launch and provider events, environment cleanup, pending
decisions, ownership, task holds, budget, and active execution. No
migration or dependency change.

## Model Used

OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and
context window are not exposed in this session. Used reasoning,
repository tools, code execution, and browser inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 01:00:42 -05:00
DottaandPaperclip 215586d127 fix: settle interrupted preparation with a retained cancellation fence (#15022)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A cancelled preparation must preserve history while allowing a new
user turn.
> - Older preparation can retain its cancellation fence but omit the
unwind marker.
> - The controller from that older boot is gone, its lease expired, and
no provider started.
> - Use the same narrow preparation proof for this retained-fence state.
> - The benefit is working Retry and saved-message recovery after
interrupted startup.

## Linked Issues or Issue Description

Refs #15020 and #15015. Related #13293 addresses post-launch recovery.

**What happened?**

A preparation interrupted before native selection retained
`startupCancellation.beforeNativeSelection: true` but no
preparation-settled marker. Its owner expired across restart and it had
no environment leases. The missing-receipt compatibility rule did not
recognize the retained fence, leaving Retry absent and saved messages
deferred.

**Expected behavior**

Admit one new user turn when the preparation evidence agrees, the old
controller expired, and cleanup is complete. Preserve previous results
and do not replay the cancelled input.

**Steps to reproduce**

1. Cancel Paperclip Runner preparation before runtime selection.
2. Retain the cancellation fence without an unwind marker or environment
leases.
3. Restart after its old controller lease expires.
4. Send a new message, select Retry, or allow the saved-message worker
to reconsider new input.
5. Confirm one successor receives only undelivered input. Repeat with
live ownership, invocation evidence, or pending cleanup and confirm
execution stays held.

**Paperclip version or commit**

Reproduced on `94f6f3eb4`.

**Deployment mode**

Authenticated private self-hosted server, built from source.

## What Changed

- Accept a retained before-selection cancellation fence in the expired
historical preparation proof.
- Retain runtime, stage, adapter, ownership expiry, process, event, and
cleanup requirements.
- Extend Retry, fresh-message, partial-queue, concurrent-worker, and
contradictory-evidence tests to both receipt states.
- Document recovery when the preparation-settled marker was not
retained.

## Verification

- Red: five exact-state regressions fail on the parent commit.
- Green: all 203 continuation tests pass, including the five new
regressions and 13 additional negative evidence cases. Process recovery
and queued-comment routes add 413 passing tests.
- A read-only candidate service check against the retained live run
returns Retry eligibility without changing any task state.
- An unrelated containment assertion failed once in CI. All 85 tests in
that route suite pass locally and the CI shard passes on rerun.
- Workspace `pnpm -r typecheck` and `pnpm build` pass.
- Final head `48240de3d`: all 54 CI checks pass, two optional Storybook
checks skip, and Greptile is 5/5 with zero unresolved threads. The
complete local monolithic suite is covered by sharded CI; the earlier
local run was stopped after a workspace case failed, and that case
passed in isolation.
- After merge, deploy the exact merged commit and verify an actual agent
reply through the task composer.

## Risks

A cancellation receipt alone must not certify that provider execution
stopped. This path still requires unresolved preparation, no native
identity or coordinator, an expired controller from another boot, no
launch or provider evidence, and completed environment cleanup.
Current-boot preparation remains held until it settles. No schema or
dependency change.

## Model Used

OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and
context window are not exposed in this session. Used reasoning,
repository tools, code execution, and browser inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 00:04:21 -05:00
DottaandPaperclip 94f6f3eb47 fix: recover historical interrupted native preparation (#15020)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A task conversation must accept new user input after an interrupted
startup.
> - Native startup begins with a legacy preparation row before runtime
selection.
> - Older builds did not retain the startup cancellation receipt on that
row.
> - The immutable adapter claim and expired controller can still prove
that no provider started.
> - This pull request uses that narrow proof for explicit Retry and
saved user input.
> - The benefit is a conversation that recovers after an upgrade without
repeating old work.

## Linked Issues or Issue Description

Refs #15015. Related #13293 covers retained process evidence after
provider startup. This change covers interrupted preparation before
native runtime selection.

**What happened?**

After an upgrade, a task stopped during native preparation can still
show automatic recovery stopped. A new user message saves but does not
start. Retry is absent because the historical run has no cancellation
receipt.

**Expected behavior**

Offer Retry and admit new user input when immutable run evidence proves
that no provider started and cleanup is complete. Preserve incomplete or
contradictory evidence as a recovery hold.

**Steps to reproduce**

1. Retain a cancelled run with the Paperclip Runner adapter claim, an
unresolved runtime, and the preparing stage.
2. Keep its old controller boot ID and expired lease. Retain no native
identity, coordinator, result, process identity, or provider events.
3. Upgrade from a build that did not save the startup cancellation
receipt.
4. Send a new user message or select Retry. Confirm that one fresh turn
starts.
5. Repeat with an active controller, a provider launch, or unfinished
cleanup. Confirm that execution stays held.

**Paperclip version or commit**

Reproduced on `cc67d4e1d` with a historical interrupted preparation row.

**Deployment mode**

Authenticated private self-hosted server, built from source.

## What Changed

- Recognize historical interrupted native preparation from immutable run
evidence and an expired controller from another server boot.
- Reject adapter invocation evidence in the startup proof. Keep process
and environment cleanup checks.
- Apply the same proof to saved user messages in the bounded recovery
worker.
- Recheck saved-message eligibility under the existing task and run
locks. Keep normal ownership, decision, pause, and budget gates.
- Add regression tests for Retry, a new message, concurrent
saved-message recovery, and contradictory evidence. Document the
compatibility rule.

## Verification

- Red: Retry, new-message recovery, and saved-message recovery fail on
the parent commit.
- Green: 185 continuation tests, 335 process-recovery tests, and 78
queued-comment route tests pass. Two additional red regressions cover
partially delivered saved queues and pass after the admission fix.
- Workspace `pnpm -r typecheck` and `pnpm build` pass.
- Final head `48db53f4f`: all 54 CI checks pass; two optional Storybook
checks skip. Greptile is 5/5 with zero unresolved threads.
- The local monolithic `pnpm test:run` was stopped after CI passed. One
unrelated workspace case failed in that long run; all three workspace
reconciliation cases pass in isolation. No complete local monolithic
pass is claimed.
- After merge, deploy the exact merged commit and verify recovery
through the normal task composer.

## Risks

Historical compatibility could grant a new turn without enough startup
evidence. The proof requires an immutable native adapter claim, an
unresolved preparing stage, no result or native identity, an expired
controller from another server boot, and no invocation or provider
evidence. Environment cleanup remains mandatory. The fix does not replay
old input or change historical run results. No schema or dependency
change.

## Model Used

OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and
context window are not exposed in this session. Used reasoning,
repository tools, code execution, and browser inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 23:08:05 -05:00
DottaandPaperclip cc67d4e1d8 fix: preserve steering and recover stopped task conversations (#15015)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A task conversation must let a user guide a running agent and resume
stopped work.
> - The active run owns its input protocol, even when the user changes
the next model or effort.
> - Queue delivery waits for a provider receipt, which must be able to
persist during the request.
> - A stopped startup also needs a clear user action that passes normal
task admission.
> - This pull request fixes steering delivery, makes queue actions
immediate, and restores explicit continuation.
> - The benefit is a responsive conversation that can recover without
losing saved input.

## Linked Issues or Issue Description

**What happened?**

A queued message could change from Steer to Interrupt while a native run
prepared. A steer request could wait on its own database lock and fail
to deliver. A stopped startup could then leave the conversation without
a working Retry or message continuation. Interrupt also waited for the
server and showed a toast.

**Expected behavior**

The active run keeps its input protocol. Steer delivers input to that
run. Steer and Interrupt clear the submitted queue rows and show the
input in the conversation immediately. Failed delivery restores the
latest queue with an inline error. An eligible stopped run offers Retry,
and authenticated user input can start a fresh turn through normal task
admission.

**Steps to reproduce**

1. Start a task with a native Paperclip Runner.
2. Change the selected model or effort while that run prepares.
3. Queue a message and press Steer.
4. Observe the provider receipt and queue state during the request.
5. Stop a startup before its provider process begins, then try Retry or
send a new message.
6. Repeat queued delivery with a legacy runner and press Interrupt.

**Paperclip version or commit**

Reproduced on the parent of this branch, `59c07ede7`.

**Deployment mode**

Authenticated private deployment. The fixes also cover local task
conversations.

Related work: Refs #12834, Refs #13354, Refs #13275. The open refactor
in #13160 moves the same queue route; it does not fix the receipt lock
or stopped-run continuation addressed here.

## What Changed

- Select queue behavior from the active run's immutable dispatch and
runtime resolution.
- Leave the run row unlocked during provider acknowledgement, then lock
and read it before merging the receipt.
- Retain queued input if the target run stops during that wait. Keep
inline delivery errors visible after empty queue updates.
- Permit exact Retry and authenticated continuation after verified
native startup cancellation. Preserve pause, approval, budget,
ownership, and process-stop gates.
- Carry undelivered native queue input into a fresh turn once the old
execution is confirmed stopped.
- Show Steer and Interrupt input in the conversation and clear submitted
composer rows immediately. Restore the latest queue inline on failure.
Remove delivery toasts.
- Keep optimistic delivery stable across stale polls, empty queues, and
paginated history. Preserve classic Interrupt error handling.
- Document recovery and optimistic delivery behavior. Add regression
tests across server, shared queue projection, and UI boundaries.

## Verification

- Red-green regression tests reproduced the queue protocol, receipt
lock, stopped-startup continuation, and optimistic delivery failures.
- The focused server route, continuation, queue, and runner boundary
suites passed during implementation.
- The queue-route suite passes with 78 tests. The three complete
conversation UI suites pass with 347 tests.
- UI typecheck, production build, and `pnpm check:token-gates` pass.
- Workspace `pnpm -r typecheck` and `pnpm build` pass. The local
monolithic `pnpm test:run` is still running; remote CI verifies the
complete suite on the latest commit.
- All CI gates pass on `bd9031ad56abfcde13d13a13488c1b9217c2fd3a`,
including the full test shards, runner verification, browser E2E,
typecheck, release registry, and canary dry run.
- Greptile reports 5/5 for that commit. Both review threads are
resolved.

## Risks

This changes queue display and explicit continuation admission. The UI
must restore rejected delivery without losing other-session edits. The
server must preserve concurrent provider result updates and must not
resume a process whose stop is uncertain. Focused tests cover these
boundaries. This change has no database migration.

## Model Used

OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and
context window are not exposed in this session. Used reasoning,
repository tools, code execution, and browser inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 22:07:14 -05:00
Devin FoleyandPaperclip 22cea6b2e6 fix: bound sandbox bridge waits and flag silent runs sooner (#14979)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Sandbox agents exchange input and output through bridge control
commands.
> - A provider can stop responding to a command even when it receives a
timeout.
> - These small commands can inherit a four-hour agent lifetime and
block input or teardown.
> - The board also calls a silent run healthy for the first hour.
> - This pull request bounds bridge control waits and surfaces silence
sooner.

## Linked Issues or Issue Description

**What happened?**

A sandbox run can remain active when a bridge control command never
returns. The shared helper passes a timeout to the provider but does not
enforce it on the host. It also accepts the agent's hours-long timeout.
Output silence remains `ok` for an hour and becomes `critical` only
after four hours.

**Expected behavior**

Bound short bridge operations even if the provider never settles. Report
failed input delivery through the existing shutdown path. Warn after
five silent minutes and escalate after fifteen. Keep normal agent
command limits and require verified termination before releasing
execution ownership.

**Steps to reproduce**

1. Use a sandbox runner whose bridge read or input-upload promise never
settles.
2. Set its configured timeout to four hours.
3. Observe that the old queue client never returns or rejects.
4. Inspect a running task with 35 minutes of output silence. The old
summary still reports `ok`.

**Paperclip version or commit**

Base commit `d6d88b9de2`.

**Deployment mode**

Self-hosted server with sandbox execution.

**Agent adapter(s) involved**

Shared command-managed sandbox bridge, including Codex ACP sessions. The
informational silence thresholds apply to active runs across adapters.

Related: #14889 recovers stalled Daytona output streams; #14485 retries
explicit gateway failures during input delivery. This change bounds
short control operations whose provider promises never settle. It does
not add tool replay or automatic cancellation for output silence. #6297
proposes configurable per-agent silence thresholds; this patch only
changes the existing defaults.

## What Changed

- Enforce at most 30 seconds per bridge control shell command on the
host and provider, including callback startup and shutdown,
process-session launch, and payload setup. Preserve shorter configured
deadlines and launch environments.
- Keep the long-lived agent command outside this deadline. Use a fixed
timeout diagnostic without command payloads.
- Surface suspicious output silence after five minutes and critical
silence after fifteen minutes.
- Decouple the shared-workspace holder cutoff from warning thresholds
and preserve its existing one-hour value.
- Add regressions for hung reads, a late upload response, failed input
delivery, exact warning boundaries, and fresh output clearing warnings.
- Update the adapter guide and execution contract.

## Verification

- The three new queue-client regressions fail on the unchanged base and
pass with this patch.
- Final callback bridge and sandbox session suites: 214 passed. These
cover hung reads, writes, startup, shutdown, process-session launch,
payload setup, and the separate long-running agent limit.
- Stdin ordering and shutdown suite: 56 passed after the lifecycle
change.
- Daytona and watchdog coverage passed in the earlier focused runs.
Across the focused suites, 602 distinct tests pass.
- `pnpm -r typecheck` and `pnpm build`: passed. Server and adapter
typecheck/build also passed after their respective follow-up changes.
- `pnpm test:run`: attempted and stopped after known local failures.
Four chat/email cases used an external ancestor skill path, three
skill-cache cases failed on macOS, and one wakeup case timed out. The
wakeup case passes alone (1 passed, 27 skipped). This run spanned the
workspace-cutoff follow-up and also failed its new holder case; a fresh
final-head workspace suite passes all 19 tests. The interrupted run is
not a full local-suite pass or final-head verification.
- A filesystem queue-drain test failed once during the lifecycle rerun
and passed on the complete two-suite rerun. It uses the filesystem
client, outside the changed command-runner path.
- Complete CI on `ff2212c235`: 53 successful checks and two expected
skips, including the full test suite and canary packaging dry run. No
failed or pending checks.
- Greptile reviewed `ff2212c235` at 5/5. All review findings are
addressed, no threads remain unresolved, and the branch has no merge
conflicts with `master`.
- `git diff --check` and a scan of added text for secrets and private
identifiers passed.

## Risks

- A bridge control operation that needs more than 30 seconds now fails,
even if the caller selected a longer run lifetime. Agent commands retain
their own limits.
- Timing out a provider promise does not cancel the remote operation or
prove it stopped. Existing execution settlement still owns termination
verification. No uncertain tool action is replayed.
- Quiet healthy runs display warnings sooner. Existing snooze, continue,
and false-positive dismissal controls still apply. Silence alone does
not cancel a run, create review work, or change assignments.
- No schema or API shape change.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run change-specific tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 15:29:41 -07:00
DottaandPaperclip 7a52dcdc74 fix: repair MCP validation and cancelled execution recovery (#14951)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The tool gateway gives agents access to connected services. Recovery
controls what happens when a run stops.
> - Generated tool names can exceed the provider limit after the MCP
client adds its prefix.
> - The same invalid definition can fail each automatic retry. A
cancelled run can also hold saved messages without showing its cause.
> - This pull request bounds tool names, stops configuration retries,
and retains cancellation evidence.
> - It shows the stopped run and admits saved input only after the
existing safety checks pass.
> - The benefit is a clear recovery path that preserves operator Stop
and prevents duplicate message delivery.

## Linked Issues or Issue Description

**What happened?**

A long connected MCP tool name makes the provider reject the entire
request. Automatic recovery repeats the invalid request. Separately,
unexpected legacy cancellations can leave saved input behind a recovery
hold. The notice does not identify the stopped run or its cause.

**Expected behavior**

Complete MCP names fit the provider limit. Tool-definition errors
require configuration repair. Cancelled runs retain their source and
reason. The recovery notice shows the cause and saved-message count.
Verified unexpected cancellations can start a fresh turn through the
existing admission checks.

**Steps to reproduce**

1. Assign an App gallery connection with a long application key and tool
name to a Claude agent.
2. Start a run. The provider rejects a name over 128 characters,
including its MCP prefix.
3. For cancellation recovery, stop a legacy provider turn without an
operator Stop request and send a user message while the recovery hold is
active.
4. Inspect the recovery notice and the deferred message queue.

**Paperclip version or commit**

Rebased onto master at `cf8ad63c806685bfd7c48e3ed4a919d61a7c55f1`.

**Deployment mode**

Hosted or self-hosted server with legacy Claude or Codex execution.

Related public work:

- Refs #14017. That PR caps name segments. This PR preserves existing
short names and uses stable hash aliases for long complete names. It
also covers classification and recovery.
- Refs #4510. That PR adds a cancellation-source column. This PR records
bounded evidence in the existing run result, without a migration.
- Refs #12552 and #4506. Those PRs suppress recovery after operator
cancellation. This PR preserves operator intent and uses the existing
continuation gates.

## What Changed

- Bound gateway names with the full provider prefix in the 128-character
budget. Retain the original upstream tool name for dispatch and
permissions.
- Classify invalid tool definitions as configuration failures before
diagnostic redaction. Stop automatic retries and continuation attempts
for that error code.
- Persist cancellation source, expectedness, initiator, reason, and
time. Preserve recorded Stop intent when adapter results arrive. Report
unexpected started cancellations with closed diagnostic labels.
- Show the run cause, saved-message count, and Inspect run link. Offer
Continue for eligible unexpected cancellations. Require verified
provider stop, empty tool inventory, ownership, and the existing pause,
budget, approval, and dependency gates. Use the existing queue for
single delivery.
- Add regression coverage and update the execution, MCP gateway, and
run-log documentation.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- `pnpm check:token-gates` passed.
- Ran `pnpm test:run` and completed its workspace and serialized groups.
Initial resource and timing failures passed on isolated reruns. All 149
serialized route suites passed.
- Reran the changed server, adapter, and UI suites after the rebase.
Coverage includes long-name upstream dispatch, configuration retry
suppression, cancellation evidence retention, privacy labels, oversized
run projection, and concurrent saved-message delivery.
- `pnpm test:e2e tests/e2e/legacy-failure-continuation.spec.ts` passed
all six browser scenarios. The recovery notice shows the run cause and
inspection link, and each recovery entry point reaches one new response.
- Added database-backed checks for active, removed, paused, unavailable,
and disabled chat connections. The final continuation and
recovery-notice suites passed 167 tests. Externally bound chats hide
board Continue and show a usable next action.
- All 55 GitHub checks passed on
`42afbf1371dcaeb72646e3d8f65c19ff7cddf8de`. Two unrelated Storybook jobs
were skipped by their normal conditions. Greptile reviewed that commit
at 5/5 with no findings and no open review threads.

## Risks

- Long tool names change to aliases. Existing short names stay
compatible. The original connection and upstream name remain the
dispatch authority.
- Invalid tool definitions no longer get automatic retries. An operator
must repair the configuration before a new attempt.
- Continuation changes apply only to positively identified unexpected
legacy cancellations with complete empty tool inventory. Operator Stop,
unknown historical cancellations, outstanding tools, and unverified
provider termination keep their holds.
- No database migration. The added projection fields are optional.
Cancellation reason and initiator IDs remain local run evidence; Sentry
receives only closed source and initiator-type labels and expectedness.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository editing, shell
execution, and GitHub tool use. The runtime does not expose the exact
model variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 13:47:59 -05:00
DottaandPaperclip 8ec4b84e1c fix(chat): resume messages after failed runs without duplicate delivery (#14857)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A user can send a new message after a native run fails.
> - The server checks that the old execution has stopped before it
starts a fresh turn.
> - A failed run can retain a result accepted before checkpoint or
cleanup failed.
> - The continuation gate treated that saved result as active recovery
and held the new message forever.
> - This pull request removes that false liveness signal while retaining
controller, process, environment, and authorization checks.
> - Live staging then exposed a second defect: chat admission created a
successor without consuming the original deferred receipt, so completion
delivered the message again.
> - Consume that exact receipt atomically with admission, while
preserving separate turns for later chat messages.

## Linked Issues or Issue Description

**What happened?**

A new user message stayed in the queue with `controller_settling` after
the previous run had reached `terminal_failure`. The old coordinator had
no lease owner but still had a `resultId`. Its remote environment had a
verified stop receipt.

**Expected behavior**

Start one fresh turn after execution has stopped and normal admission
checks pass. Preserve the failed run and its accepted result as history.

**Steps to reproduce**

1. Accept a native result, then fail checkpoint or cleanup and exhaust
recovery.
2. Retain the result ID on the terminal failure record and stop the
execution environment.
3. Send a new user message. Before this fix, it waits forever for the
finished controller.

**Paperclip version or commit**

Reproduced in a database-backed regression test on `26900655b`.

**Deployment mode**

Server with a native runner and remote sandbox. Local process stop
checks also apply.

Related: https://github.com/paperclipai/paperclip/pull/14775. Searched
existing PRs for retained-result continuation fixes; no duplicate found.

## What Changed

- Remove the retained-result veto for terminal failures.
- Keep controller ownership, successor, process, environment cleanup,
pending decision, and ordinary admission checks.
- Add regressions for retained results, active execution, missing stop
evidence, and delayed remote cleanup.
- Atomically consume the resumed receipt in agent chat, even though chat
does not coalesce other queued messages.
- Reproduce completion-time duplicate promotion, race cleanup against
periodic recovery, and prove a subsequent chat message keeps its own
turn.
- Document that a saved result does not make a terminal failure active.
- Keep exhausted workspace export on its separate repair path, tested
through the production finalizer.

## Verification

- Red: retained-result admission failed with `controller_settling`
before the original fix. The new chat-specific regression then
reproduced duplicate promotion when the first reply finished.
- Green: 406 tests across native continuation, workspace-export
recovery, and the wake-queue module passed on `cbc531cc0`.
- The chat regressions exercise real Postgres transactions, simultaneous
recovery callbacks, successful completion, the production queue-drain
use case, and repeated drain attempts. A distinct follow-up remains a
separate turn.
- `pnpm -r typecheck` and `pnpm build` passed on `cbc531cc0`.
- The earlier full local test run encountered a timeout and follow-on
failure in unchanged AI connection-adoption tests; all 50 tests passed
on isolated rerun. That local run was stopped after the full CI test
matrix passed on the earlier head.
- All 54 CI checks passed on `cbc531cc0` (2 skipped), including the full
test matrix and browser shards. One unchanged interaction-route test
returned HTTP 500 on its first CI attempt; its full 84-test file passed
locally, and the failed shard passed on one targeted rerun.
- Greptile reviewed `cbc531cc0`: 5/5, no unresolved findings.
- Live staging first verified that the original saved message resumes
and receives a successful response; that test exposed the duplicate now
covered above.
- Deployed exact commit `cbc531cc0410e1ef6e8811c6c5c014c3528351ed` to
the affected staging workspace; deployment verification, health,
authentication, and startup recovery passed.
- Submitted a fresh message through the browser. The agent replied in 39
seconds; server records show exactly one successful run, native phase
`committed`, no error, and an empty queue. A later check more than a
minute after completion found no duplicate run.

## Risks

The change affects admission after native execution failure and
consumption of a resumed deferred receipt. A fresh turn must never
overlap the prior execution, and consuming one chat receipt must not
absorb later messages. Tests retain the controller, process, and
remote-stop guards. This change does not migrate data, apply an old
result, or reset the old retry budget.

## Model Used

OpenAI Codex (GPT-6). The exact runtime model identifier and context
window are not exposed in this session. Used reasoning, repository
inspection, code execution, database-backed tests, and browser
inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 14:43:02 -05:00
DottaandPaperclip dd7fc1f90a fix: raise the native journal read limit to 256 MiB (#14711)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native sessions persist control-plane state so they can resume
safely.
> - The state includes committed provider history needed for recovery.
> - The server, runnerd recovery, and durable control plane validate
this state before trusting its identity.
> - Their differing 64 MiB and 192 MiB limits can reject a valid journal
before recovery.
> - This pull request aligns all three local state limits at 256 MiB.
> - Larger files remain bounded, while recovery can read larger valid
histories.

## Linked Issues or Issue Description

Refs #13882
Refs #14312

## What Changed

- Raise the server and runnerd recovery limits from 64 MiB to 256 MiB,
and align the durable control-plane limit from 192 MiB to 256 MiB.
- Add coverage for a valid history above 64 MiB and rejection above 256
MiB.

## Verification

- Matching server recovery passed with more than 64 MiB of actual
committed event payloads (128 events with 512 KiB deltas).
- The real runnerd exact-authority resume regression with the test Codex
provider passed with 193 MiB of valid JSON whitespace appended. It
crosses the former 192 MiB core limit and confirms the same provider
identity. This exercises the runner process and durable control plane
with a simulated provider, not a live OpenAI API call. This test used
approximately 1.15 GiB peak RSS.
- The actual runnerd reader accepted valid 256 MiB JSON and rejected
valid 256 MiB + 1 byte. The reader call took 231 ms; the fresh process
peaked at 1,244 MiB RSS.
- Server tests reject mismatched identity above 64 MiB and files above
256 MiB.
- `pnpm -r typecheck`, `pnpm build`, and `git diff --check` passed.
- Full local `pnpm test:run`: 13,730 passed, 575 skipped, 7 failed
across 6 files. All failures were embedded PostgreSQL startup errors
after five attempts. They affected agent hiring, instruction revisions,
environment images, reviewed chat bindings, issue monitoring, and legacy
continuation authority. The focused journal tests passed; the latest
pushed head passed all ordinary CI checks. Superagent is the only
blocking check.

## Risks

- **Open review concern:** Greptile is 5/5, but Superagent is
`ACTION_REQUIRED` with two P2 findings on the server and runnerd
readers. Both flag the increased synchronous parsing and memory cost.
This PR keeps the requested fixed-limit change small. It does not add a
worker parser or a process-wide memory budget. This resource tradeoff
needs review before merge.

- Large state parsing is synchronous and can consume several times the
file size in memory.
- Remote checkpoint archive and expanded-size limits remain 64 MiB, so
this change alone does not make larger remote checkpoint transfers
portable.

## Model Used

- OpenAI Codex, GPT-6, with delegated assistance from `gpt-6-luna` at
high reasoning effort; tool use and code execution. The GPT-6 context
window is not exposed in this task runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:08:25 -05:00
DottaandPaperclip 1778075155 fix(server): continue unfinished tasks after status replies (#14626)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs report their task outcome through a structured finish
result.
> - A Board comment can permit a passive wait for the next response.
> - That exception accepted reports that also admitted blocking
unfinished work.
> - The task then stayed In Progress without a runner, and recovery
treated the wait as healthy.
> - This pull request rejects that contradiction and uses the existing
bounded continuation path.
> - Real wait conditions and protection against obsolete requests remain
in force.

## Linked Issues or Issue Description

**What happened?**

An agent answered a status inquiry with `yielded` and `response_wake`.
The same result listed blocking remaining work. The server accepted an
indefinite wait without a question, approval, dependency, or pause. No
further run was queued.

**Expected behavior**

Unfinished ordinary tasks must continue or have a recorded reason to
wait. A status reply alone must not suspend the work.

**Steps to reproduce**

1. Add a Board status inquiry to an unfinished assigned task.
2. Submit a successful native result with `yielded`, `response_wake`,
and `remainingWork[].blocksCompletion: true`.
3. Leave the task without any real wait condition.
4. Observe that the old policy preserves In Progress with no
continuation and suppresses recovery.

**Paperclip version or commit**

Reproduced in the native status policy at `da887ea3e`. The branch is
based on current master.

**Deployment mode**

Authenticated server with the native Paperclip Runner.

Related work: Refs #13338 (native response waits and recovery). Refs
#12071 (separate legacy recovery and retry-state work). This change
fixes the native unfinished-response-wait exception.

## What Changed

- Reject contradictory finish reports while the provider can still
correct them.
- Route accepted unfinished response waits through the existing
one-follow-up continuation budget. Repeated incomplete results create a
visible recovery action.
- Preserve questions, approvals, dependencies, pauses, conversation
lifecycles, and superseded Board requests.
- Recheck the current Board source in the decision transaction before
queuing repair.
- Let normal recovery reconsider old committed waits that report
blocking work and still have a current source.
- Add policy and database regression tests. Document the rule.

## Verification

- `pnpm exec vitest run
server/src/services/native-runtime/status-arbiter.test.ts
server/src/__tests__/heartbeat-process-recovery.test.ts
--no-file-parallelism`: 346 tests passed, no skips.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `git diff --check origin/master...HEAD`: passed.
- Final head `8f0824d9bb4d631f2347f8afffe5fff6d574cb69`: all CI checks
green (54 passed, 2 intentionally skipped), including all general tests,
serialized server suites, runner tests, eight E2E shards, and the canary
dry run.
- Greptile: 5/5 on the final head, with no unresolved review threads.
- The broad local `pnpm test:run` encountered an unrelated timing
failure in `workspace-runtime.test.ts` (waiting for managed process-tree
listeners). That test passed on an isolated rerun. The duplicate broad
run was stopped; the complete CI suite passed on the final commit.
- The transaction-race regression reproduced an obsolete continuation
before the fix. All 12 source-change cases now pass, covering passive
waits, corrective continuations, and exhausted-repair decisions.

## Risks

- Agents that previously parked unfinished ordinary tasks must now
continue or record an actual wait condition.
- Old contradictory waits become eligible for normal recovery. Existing
ownership, budget, pause, and supersession checks still apply.
- The rule uses the structured blocking-work flag. It does not infer
omitted work from prose.
- No schema, dependency, or UI change.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
editing, and test execution. The session does not expose a more specific
deployment identifier or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 16:51:50 -05:00
DottaandPaperclip 2de43fc909 fix(issues): keep agent mentions as context and defer personal app authorization (#14577)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Each task has one assignee. Explicit assignment and review requests
select who should act.
> - An agent mention started another agent on a task it did not own.
Native attachment staging then rejected that run.
> - Allowing that run through startup could also let two agents work on
the same task.
> - Mentions should identify relevant context. They should not start
work or forward comments to other tasks.
> - A personal app installed on a shared agent must also wait until tool
use to resolve the current user's grant.
> - This pull request removes mention dispatch and keeps missing
personal app credentials from blocking startup.

## Linked Issues or Issue Description

**What happened?**

A native agent mentioned on another agent's task failed with
`paperclip_runner_attachment_staging_not_authorized`. The source task
could already be complete. A nearby optional-app warning was a separate
problem: personal app tools were excluded when their shared health state
required attention.

**Expected behavior**

An agent mention is context only. It does not wake the agent, take
ownership, or copy a comment onto another task. Normal feedback still
reaches the assignee. Assignment and explicit review requests still
dispatch work. An unavailable personal app does not block startup or
produce a startup warning. Tool use requests the current user's
authorization and never uses another user's grant.

**Steps to reproduce**

1. Assign a task to agent A. Post a comment that mentions agent B,
including a comment that closes A's task or references B's child task.
2. Confirm the comment retains its agent link and B receives no run or
deferred wake. A can still receive normal feedback.
3. Install an active personal MCP connection on B. Give only Alice a
grant and leave shared health at `error`.
4. Explicitly assign work to B for another user. Confirm it can finish
without using the app.
5. Ask B to use the app. Confirm its tool call shows an inline
connection request for the current user.

Related work: Refs #11144. This change uses the existing execution-time
personal grant resolution.

## What Changed

- Remove mention dispatch from standalone comments and issue updates.
Remove implicit forwarding of parent comments to a mentioned worker's
child task.
- Ignore new requests with the legacy mention wake reason before
creating a run or deferred request. Preserve already accepted queue
entries, which can combine assignments and feedback with a later
mention.
- Remove the native mention admission, staging, and finalization
exceptions from this PR. Native task ownership checks remain intact.
- Keep active, installed personal app tools available despite shared
health errors. Remove optional-app startup warnings. Tool execution
retains the current user's grant and policy checks.
- Update agent instructions and product/API docs. Refresh generated
capability source anchors.

## Verification

- Red: comment-route regressions reproduced extra agent wakes and child
comment forwarding. A separate regression proved that cancelling by the
last coalesced reason could drop an accepted assignment.
- Green: the targeted route, wake queue, heartbeat, workspace,
responsible-user, MCP discovery, and HTTP gateway suites passed. The
final queue and heartbeat rerun passed 104 tests, the restored queue
adapter passed 56, and both comment-route suites passed 135. These
include accepted assignment preservation, rejection of new mention
requests, and normal assignee feedback.
- `pnpm -r typecheck` and `pnpm build` passed locally. The full local
`pnpm test:run` attempt was interrupted for review/CI fixes, so it is
not claimed as a completed local pass. It exposed a cleanup timing race
in the concurrent-mention assertion, now fixed and verified across 10
repetitions. CI also exposed an obsolete test waiting for the removed
mention lookup; it was reproduced and fixed, then both comment suites
passed. Final full-suite verification is through CI.
- Final head `bd9ea4cb05a8f081c54e017760a8999f9ea6ef44`: 54 checks
passed, 2 Storybook checks intentionally skipped; no pending or failing
checks. Full CI includes general and serialized suites, all 8 browser
shards, runner verification, typecheck, build, and canary dry run.
Greptile is 5/5 on this exact commit, with no unresolved findings.
- One unchanged Cursor adapter test hit its 10-second CI timeout. All 5
tests in that file passed locally; one retry of its CI shard passed all
674 tests (3 skipped). The aggregate verification gate then passed. No
code or timeout was changed for that retry.
- Live browser check: inserted a structured mention with the picker on a
human-owned task. The saved link remained visible. Database checks found
zero new runs and zero wake requests.
- Live Codex runner check: explicitly assigned that task with the
unavailable personal app attached. The run succeeded and committed
completion without using the app or creating a connection card.
- Live browser follow-up: asked the assignee to call PostHog and
mentioned another enabled agent as context. Only the assignee ran. It
succeeded and displayed the existing inline connection card. Only
Alice's grant existed; the run belonged to a different user.
- The HTTP regression covers tool discovery with no provider calls or
connection cards, first use returning the current user's authorization
request, and successful retry after that user's grant exists.
- App checks use an isolated local fixture and a fake MCP provider. They
do not use production app credentials.

## Risks

- Intentional behavior change: workflows that used mentions to wake
agents must use assignment, a bounded child task, or an explicit review
request.
- Already accepted queue entries retain their prior rules. An old entry
can combine assignment or feedback with a later mention; its last reason
cannot safely identify mention-only work. New mention requests create no
run or deferred wake.
- Personal apps with a shared health error remain discoverable. Actual
tool use still requires the responsible user's grant and existing policy
gates.
- No database migration or public API schema change.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact serving model ID and
context-window size are not exposed in this session.
- Live native-run verification used `gpt-6-astra` through the Codex
provider.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 12:49:01 -05:00
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00
DottaandPaperclip 8751e2de46 fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task board shows which recovery actions are active.
> - A native run can stop while a person must repair its workspace.
> - The board previously called that state “Recovery in progress.”
> - The label implied that work would continue without operator action.
> - This pull request derives the label from the recovery owner and live
continuation.
> - Operators can distinguish scheduled recovery from a repair that
needs attention.

## Linked Issues or Issue Description

**What happened?**
A blocked task showed “Recovery in progress” after native finalization
stopped and no automatic continuation remained.

**Expected behavior**
Show “Recovery needed” for an idle board repair or when no live recovery
path exists. Show “Recovery in progress” while the recorded continuation
can run, including an explicitly admitted export whose exact callback is
executing even if the old recovery action remains board-owned.

**Steps to reproduce**
1. Complete a native run whose workspace export cannot be recovered
automatically.
2. Inspect the task recovery action and badge.
3. Compare the board-owned action with the old “Recovery in progress”
label.

Related work: #14314 serializes native workspace finalization and fences
stale recovery outcomes. This change reports the recovery action that
currently owns the task.

## What Changed

- Show “Recovery needed” for board-owned active-run recovery unless the
exact native export callback is positively verified as executing.
- Require the native continuation run and a live or future continuation
before showing progress.
- Project native activity from the exact company, source issue, and run.
Include active workspace export while the original heartbeat remains
failed.
- Require an executing callback before using a running export row as
evidence. Preserve activity for long exports and clear it when the
callback joins.
- Give native resume its own card explanation. Preserve ordinary
watchdog observation behavior.
- Remove the redundant ownership sentence from all six recovery-card
explanations that used it.
- Document the labels and add regression cases for stopped, scheduled,
and active recovery.

## Verification

- Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests
and token gates pass. No UI occurrence of the removed sentence remains.
`pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54
successful checks, two optional Storybook checks skipped). Greptile is
5/5 with no open review threads. The duplicate local `pnpm test:run` was
stopped after the full CI suite passed; it did not complete locally.
- Original regressions: seven failures before the change, then 52
focused cases pass.
- Review regressions: seven UI failures and nine database failures
before the follow-up. All 75 database/API recovery tests, 167 UI tests,
and 18 workspace lifecycle/finalizer tests pass. A further three RED
cases cover explicit board retry activity; one RED case rejects orphaned
running export rows after controller loss. Wrong company, issue, run,
service, phase, and completed-operation cases remain inactive.
- Final recursive typecheck, production build, token gates, and complete
local suite coverage pass. Embedded PostgreSQL startup/socket failures
passed in isolated retries with the canonical test environment; no
expected behavior was weakened. The recovery regression added during the
earlier full run passed in its final complete 75-case file.
- Before the copy-only follow-up, all 56 CI checks passed on
`c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no
unresolved review threads. The final native-activity staging repeat
passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`.
- Verified on a separate staging instance: a real failed Daytona
workspace export retains its board-owned repair action and displays
“Recovery needed” in the task list. The repair card remains actionable
without starting another provider turn.

- Real staged export-only repair: the actual task list showed “Recovery
in progress” while the original run had a positively identified running
export operation, then Done after exact copyback of all 20,000
nonce-bound files. The accepted result and full provider
session/turn/terminal envelopes remained unchanged. All 17 independent
final checks passed; the browser downloaded the exact 19-byte result.
The separate fixture was cleaned up with independent provider-absence
verification. A control transport process restarted during repair; it
did not submit another provider turn.

- Final deployed-source repeat on
`d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral
Daytona allocation, actual browser export repair, and a saved full
activity projection referencing the exact executing export operation
with no scheduled retry. The task list showed recovery in progress, then
Done; all 21 final checks passed, including 20,000 exact host files,
unchanged provider provenance, and provider deletion only after
committed copyback. The downloaded 19-byte result matched independently.
This integrates #14334; no source change was required here.

## Risks

- The label depends on the persisted recovery action. A separate runtime
defect can still stop work; this change makes that condition visible.
- Future recovery kinds must supply a valid continuation path before
they can display progress.

## Model Used

OpenAI Codex, based on GPT-6, with code execution, browser testing, and
subagent tool use. The runtime does not expose an exact serving model
variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:33:39 -05:00
Devin FoleyandPaperclip be43c23e2d fix(recovery): preserve pending native result finalization (#14219)
Preserve native coordinator ownership while accepted results await workspace
copy-back, assessment, or arbitration. Check that ownership in the terminal
update so a coordinator recorded after the liveness read is also protected.
Keep terminal task authority and exhausted-retry cleanup unchanged.

Validation: 28 focused recovery tests, local typecheck/build, 52 passing CI
checks, and Greptile 5/5 with no open findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-27 00:22:49 -07:00
DottaandPaperclip 640dee1802 fix(runner): acquire reusable leases when switching to warm sessions (#14187)
## Thinking Path
> - Paperclip coordinates agent work across task turns.
> - Warm runners need reusable sandbox leases.
> - Lease acquisition read only the environment reuse setting.
> - Startup then read the agent's new warm setting and rejected the
ephemeral lease.
> - This fix requests reuse before acquisition, with provider checks
intact.

## Linked Issues or Issue Description
**What happened?**
Switching an existing native runner task from per-turn to warm fails
with `runner_warm_environment_requires_reusable_lease`. Environment
`reuseLease` defaults to false. Acquisition therefore creates an
ephemeral lease, and capability narrowing correctly denies reuse.

**Expected behavior**
Changing lifecycle between turns requests a compatible lease without
replacing the task workspace or changing the shared environment.

**Steps to reproduce**
1. Start a native runner task with inherited lifecycle and environment
reuse disabled.
2. Change the agent from per-turn to warm.
3. Continue the task on a reusable-capable sandbox provider.

**Paperclip version or commit**
Reproduced on `a6c4e7a`. Related: #12904 established warm workspace
continuity; this fixes the earlier lease-selection mismatch.

## What Changed
- Pass agent settings into acquisition and derive a run-scoped reuse
request.
- Respect explicit environment lifecycle overrides; leave other adapters
unchanged.
- Keep provider capability, ownership, cleanup and restore gates intact.
- Add orchestration and database-backed transition tests; document the
contract.

## Verification
- Red: three new assertions failed before the fix.
- Green: 156 tests pass in environment-run-orchestrator,
environment-runtime, native-sandbox-lifecycle, and
environment-execution-target-capabilities.
- Database-backed regression proves ephemeral → reusable → resumed
lease, same workspace, and unchanged stored environment. Provider RPCs
are mocked, not live Daytona.
- `git diff --check` passes.
- Latest-head CI passes typecheck, build, tests, runner checks, E2E, and
canary dry run. Greptile: 5/5; no open threads.
- Recovery uses the persisted lifecycle. Direct regression: red before,
six cases green after. CI runs the added Vitest cases.
- Full local checks were unavailable (missing dependencies and earlier
memory limits).

## Risks
Warm mode now requests retained resources despite an environment's
default `reuseLease:false`. Existing idle timeout and cleanup still
apply. Unsupported providers remain denied. No migration, credential
[REDACTED], or shared-environment mutation.

## Model Used
OpenAI Codex agent; reasoning, code editing and test execution. Exact
model ID and context size were not exposed to this run.

## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details available)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket [REDACTED] or instance-derived details
- [x] I have run targeted tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 21:23:33 -05:00
Devin FoleyandPaperclip b2e9e82f05 fix: stop remote Grok runs before continuing queued messages (#14100)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The execution service owns each run and saves messages sent while it
runs.
> - Interrupt must stop the current executor before it delivers those
messages.
> - Remote Grok commands did not register the host cancellation control.
> - A cancelled task run could still write Done and prevent queue
recovery.
> - This pull request connects remote cancellation and revokes cancelled
run writes.
> - Saved input can use the existing queue admission rules after
verified cleanup.

## Linked Issues or Issue Description

**What happened?**

Interrupting a queued message marked a remote Grok run cancelled before
its sandbox stopped. The old run could still post a reply and mark the
task Done. Its saved follow-up remained deferred behind execution
recovery.

**Expected behavior**

Stop revokes run write authority and waits for verified termination.
Saved messages remain durable and enter one successor through normal
admission after cleanup.

**Steps to reproduce**

1. Run a task with `grok_local` in a remote sandbox.
2. Send a follow-up and use Interrupt while the command runs.
3. Let the old command attempt a task status update after cancellation.
4. Observe the task disposition and the saved message queue.

**Paperclip version or commit**

The gap is present in master at `d3e0f0a238`.

**Deployment mode**

Authenticated server with a Daytona sandbox.

Related work: #14028 and #14046 handle bounded continuation. #13291
covers infrastructure interruption and verified remote cleanup. #13332
addresses atomic recovery holds. This change handles direct Grok
operator cancellation and stale task writes.

## What Changed

- Register remote Grok cancellation before preparation. Keep command
ownership until the host confirms sandbox termination.
- Reuse the sandbox cancellation boundary for the direct CLI invocation.
Reject fresh attempts after cancellation and preserve workspace restore
failure evidence.
- Reject writes from cancelled task JWTs and runs with a pending stop.
Preserve diagnostic reads and existing conversation error codes.
- Recheck run authority under a database lock before task updates and
interaction responses commit.
- Preserve authorized handoffs that stop their own run. Only the
server-issued stop receipt for that request permits the final task
update.
- Add tests for hung commands, unverified stops, early cancellation,
copy-back failures, late Done, late interaction responses, authorized
handoffs, exact lease receipts, and one queue successor across
concurrent restart sweeps.
- Document the cancellation and write-authority contract.

## Verification

- Targeted adapter, cancellation-boundary, authentication,
queued-message, interaction-service, and activity-route tests passed.
The expanded run passed 214 tests; one new test had an incomplete
fixture. After correcting the fixture, all 8 selected follow-up cases
passed.
- `pnpm -r typecheck`: passed on
`179c86caf1bf0d89914a503d46e24af7e4b8c557`.
- `pnpm build`: passed on the same commit.
- `pnpm test:run`: the general-server group completed with 13,521
passed, 99 skipped, and 18 failed tests. It then stopped, so the
remaining local groups did not run. Five Slack, email, and wake-batching
failures passed on focused reruns after correcting the local
environment. The remaining 13 failures reproduce as `EACCES` on rename
in unchanged skill-cache code on macOS. Two custom-image suite setup
hooks also failed to start embedded PostgreSQL after the machine
exhausted shared-memory slots; all 31 tests in that file passed on rerun
after the local resource issue was resolved. CI covers all test groups.
- CI: 53 checks passed and 2 were skipped on the latest commit,
including the aggregate verification gate. The last server shard passed
on its single rerun after a preview-server startup timeout. The affected
file also passed locally with 28 passed and 3 skipped.
- Greptile: 5/5 on the latest commit. Both review threads are resolved.
- No live deployment or staging task mutation has been performed.

## Risks

- Stopping the sandbox can prevent file copy-back. The result preserves
workspace restore failure evidence; termination does not imply restored
files.
- If provider termination fails, the adapter keeps ownership of its
outstanding command and does not acknowledge Stop.
- The write restriction now applies to ordinary cancelled tasks. Reads
remain allowed. Task and interaction checks add a shared run-row lock to
agent mutations. An exact server-issued receipt permits the task request
that stopped its own run to complete its handoff.
- Existing terminal tasks are not reopened automatically. An operator
must correct a historical late Done before its saved queue can continue.
- No schema migration or UI change.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and test tools. The precise backend revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 17:07:07 -07:00
Devin FoleyandPaperclip 110d176fc9 fix: preserve chat message bindings in review recovery (#13818)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - External chat messages can start work on tasks that remain in review.
> - Paperclip queues one recovery run when a task loses its review path.
> - That recovery keeps the chat source but loses the admitted message IDs.
> - The authorization check then rejects the recovery before execution starts.
> - This change retains the message IDs so the existing check can verify current access.

## Linked Issues or Issue Description

Refs #13809.

**What happened?**

A successful external-chat run can leave an active task in review without a maintained review path. Its automatic recovery then fails with `reviewed_chat_execution_binding_not_authorized`. The recovery context retains `chat:slack` but drops `wakeCommentIds`.

**Expected behavior**

An eligible recovery should retain its admitted message references and pass a new authorization check. Missing or revoked access must still prevent execution.

**Steps to reproduce**

1. Finish a chat run whose task remains in review with no maintained review path.
2. Build the bounded review recovery from that run's context.
3. Dispatch the recovery through the reviewed-chat authorization check.

The new regression tests fail before this patch. Related PR #13809 handles answered conversations that become idle. This patch handles recovery when the conversation remains active.

## What Changed

- Retain the admitted message batch for external-chat review recovery. Derive the current comment reference from that batch.
- Share the existing supported-provider selector between recovery and run-bound chat authorization. AgentMail remains on its separate email inbox path.
- Keep the existing authorization check. Do not copy prior checkout, authorization, session, or prompt state.
- Test Slack and Discord recovery, missing batches, revoked access, and changed task or company bindings.
- Document the recovery authorization contract.

## Verification

- The new tests reproduced the missing-message failure before the fix.
- Six focused suites passed: 257 tests covering review recovery, reviewed-chat authorization, issue liveness, Slack lifecycle, comment-wake batching, and external-chat waits. PostgreSQL integration tests ran against a temporary local PostgreSQL database.
- Focused test files: `review-path-recovery.test.ts`, `heartbeat-reviewed-chat-binding.integration.test.ts`, `heartbeat-issue-liveness-escalation.test.ts`, `slack-conversation-lifecycle.test.ts`, `heartbeat-comment-wake-batching.test.ts`, and `external-chat-wait.integration.test.ts`.
- Server TypeScript check passed with scratch configuration that resolves this checkout's workspace packages. Existing dependency links point to another checkout; the default check reports stale shared-type errors.
- `node scripts/check-module-boundaries.mjs` and `git diff --check` passed.
- `git diff | gitleaks stdin --redact --no-banner` passed. The diff was also checked for private identifiers and user data.
- [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35760757895) passed on the latest commit, including build, workspace typecheck, general and serialized tests, Rust checks, and all eight browser shards. The unchanged local-service readiness test and agent-chat page-load assertion passed when their failed shards were retried. The local-service suite also passed locally (6 tests). The first, superseded run lost a chat runner; its stuck browser job was cancelled to unblock the current run.
- Full local workspace typecheck, tests, and build were not run. The machine has about 2 GiB free, and those commands include Rust builds. The full PR CI checks passed before merge.

## Risks

Small context-construction change. Dispatch still checks current execution ownership and access for every admitted message. The recovery remains bounded to one attempt per consumed path. No schema changes. Existing failed runs are not retried by this patch.

## Model Used

OpenAI GPT-6 (Codex), with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-22 10:46:13 -07:00
DottaandPaperclip 8725d6ce09 fix: make answered Slack conversations idle (#13809)
Settle published, successful Slack turns as Idle; resume the same conversation on an admitted message. Preserve unfinished work, delivery errors, and explicit dispositions.

Verified through focused lifecycle/API/UI tests, full CI, and a real staging Slack conversation in the embedded browser.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-22 10:35:33 -05:00
Devin FoleyandPaperclip ded156a904 Keep interaction continuations scoped to the target task
Separate an interaction's producer from explicit resume history. Do not
import another task's comments or results. Filter newly captured foreign
origins and recover older inherited origins only when the saved producer
context and comment row prove their source. Keep missing context and
company boundaries fail-closed.

Verified 46 continuation tests, 201 related recovery tests, and server
typecheck. Added 20 database regressions for provenance and scope guards.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:06:05 -07:00
DottaandPaperclip 9d19f98b50 fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agent chat uses native runner sessions to plan, delegate, and track
that work.
> - A user can press Stop while the native session is still starting.
> - The server can acknowledge that Stop without dispatching it, then
let the session submit a turn.
> - This leaves chat recovery waiting for an execution that the user
expected to stop.
> - This PR waits for the startup handle, dispatches cancellation, and
prevents a late startup from submitting a turn.
> - New full-stack evals check the resulting records and outputs across
Claude and Codex.
> - Those evals also exposed missing ACPX readiness fields, unbounded
polling, and an old-run identity check that rejected valid warm
handoffs.

## Linked Issues or Issue Description

**What happened?**

Stop during native startup could record an acknowledged cancellation
with `dispatched: false`. The provider could then begin work. A
subsequent `/new` stayed queued. A remote Claude follow-up also
exhausted the command journal while probing warm-session readiness: ACPX
never returned the readiness fields required by the shared transport.
Once readiness worked, attachment incorrectly compared the next run
descriptor against the old run ID. The 25 ms polling loop could issue
4,800 commands during its two-minute wait, beyond the 500-command bound.
The existing chat eval treated lifecycle logs as proof of an active
provider turn, so it did not distinguish startup cancellation from
active-turn cancellation.

**Expected behavior**

A Stop during startup must reach the pending session. A late session
must not submit a prompt after Stop. Recovery must retain control when
startup exceeds the bounded wait. Chat evals must check saved task
state, document contents, worker identity, account binding, and
duplicate effects.

**Steps to reproduce**

1. Start a native Claude or Codex chat turn.
2. Press Stop after process startup is requested but before the provider
turn starts.
3. Send `/new`, then send a fresh message.
4. On the affected base, cancellation can be acknowledged without
dispatch and the reset stays queued.

**Paperclip version or commit**

The live Claude baseline reproduced this on `29d6b3509`. The branch also
includes master commit `0f5fafe16`.

Related work: #13678, #13686, #13693, #13291, #13738. A separate runner
reliability branch also contains a startup-wait fix. Its overlap must be
reconciled before merging; this branch additionally prevents prompt
submission after a late startup.

## What Changed

- Wait for a pending native startup before acknowledging a run-scoped
Stop. Preserve the existing recovery error when that wait expires.
- Keep a Stop guard on startup. Cancel a late handle before it can
submit a provider turn.
- Add regression tests for normal handle publication and publication
after the Stop deadline.
- Back off blocked warm-attachment probes. Keep the fast two-snapshot
barrier, fail closed, and record changed blockers.
- Add red/green tests for delayed readiness, persistent blockers,
alternating readiness, and readiness near the deadline.
- Publish ACPX readiness and blockers. Preserve the old authority’s
event acknowledgement barrier; only settled sessions can proceed to
attachment.
- Bind warm ACPX descriptors to the validated next authority while
retaining old-run event correlation until activation. Preserve session
identity and provider profile checks.
- Exercise two consecutive run rotations through a qualified fake
sidecar, verifying checkpointing, provider identity, pre-activation
rejection, and new-run work admission.
- Separate startup and active-turn cancellation checkpoints in the
browser eval.
- Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells
across Claude and Codex.
- Cover hiring and reuse through managed AI accounts, source-based
review, current blocked-task status, request replay after a lost HTTP
acknowledgement, server restart continuity, and Stop/reset continuity.
- Use ordinary production agent instructions. Enable API tools only for
the two coordination cases that need them.
- Calibrate the matchers with invalid records and outputs. Require
remembered context after restart and a structured status snapshot that
distinguishes the current blocker from history and task status from
active execution. Compare the public issue mutation contract and
relationships during read-only reporting. Preserve before/after source
records in failed eval evidence.
- Fix the lost-ack browser harness and verify it against a real HTTP
server. Check the chat composer after restart instead of waiting for an
unrelated document lifecycle event.
- Document the scope and limits of each case.

## Verification

- The startup regression failed on the unfixed executor and passed after
the fix.
- `pnpm test:e2e:runner:typecheck` passed.
- `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files.
- `pnpm exec vitest run
server/src/services/native-runtime/native-session-executor.test.ts`
passed: 385 tests.
- [Baseline live
campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868):
Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed.
Codex delegation was blocked by provider capacity.
- [Eval-only startup
campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786):
both providers failed as expected. Both persisted `dispatched: false`
and left `/new` queued.
- [First fixed
campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706)
on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both
providers. Failed cases exposed eval harness defects and remote
continuity failures. All attempts remain available.
- [Original workflows and stronger memory
checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649)
on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and
startup Stop passed for both providers; Codex remote restart passed.
Claude remote restart exposed the missing readiness contract. Two Codex
planning cells hit provider capacity.
- [Unchanged-model
retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548):
Codex planning and backlog creation both passed.
- [18-cell campaign with ACPX
readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963)
on `6a98ef743`: 16/18 passed, including all local/remote Stop and
committed-send cases. Claude remote continuity exposed the
next-authority check, now fixed. Codex hiring produced its checklist,
but the runner redacted the requested marker after it appeared as
“Tracking token: …”. That content-redaction policy is unchanged and
remains an explicit limitation.
- [Structured status
grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725)
on `50448c228`: both providers passed on their first attempt, including
cleanup.
- [Complete read-only state
grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011)
on `551e13892`: both providers passed.
- [Final ACPX handoff and hiring
retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456)
on `cbd637587`: all three Claude Daytona cases passed (restart
continuity, active Stop/reset, and lost-ack replay). Codex hiring
reproduced the content-redaction failure: the saved checklist contained
`Tracking token: [REDACTED]` instead of the required business marker.
All four cases completed cleanup successfully. [Published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/).
The only subsequent commit adds the qualified-sidecar integration test;
production code is identical to this live proof.
- `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without
paid models.
- Runner TypeScript typecheck passed. All 5 warm-readiness tests pass;
two failed with the prior fixed-rate loop, and the late-readiness test
failed before the pacing correction.
- ACPX readiness and warm-identity regressions each failed before their
fixes. All 292 runner-core Rust library tests passed. The
qualified-sidecar integration test passes. Rust formatting is checked.
- Status-grader regressions for misleading historical mentions and
previously unchecked mutations each failed before tightening the oracle
and pass now.
- [Latest-head
CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307)
passed on `a4093c8f1`: full build, type checks, test partitions, browser
E2E, and native runner checks. Two unrelated tests initially failed
(Sentry fixture release attribution and local-service fixture
readiness); both passed locally together (35 passed, 5 optional SDK
tests skipped) and on the failed-job retry. No changes were made to
those tests.
- Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed
and all review threads are resolved.
- The paid live suite is not fully green: the reproducible
content-redaction case remains red. This is separate from the passing PR
merge checks. No production content-redaction, prompt, model, or
completion-policy change is included.
- Managed-account hiring and review cases explicitly enable API tools;
these do not qualify default new-user onboarding.

## Risks

- Stop can wait up to 30 seconds for startup, then use the existing
pending-recovery path. This does not prove that remote cleanup has
finished.
- Blocked warm readiness adds up to 750 ms between later probes with the
two-minute remote budget, or about 32 ms with the default five-second
budget. Ready sessions retain the short second barrier.
- Paid evals can fail because of provider capacity or agent decisions.
Each failure needs evidence-based classification.
- The HTTP request replay case checks comment idempotency and duplicate
effects. It does not prove replay safety for an ambiguous provider tool
call.
- The new suite is opt-in. It does not increase the default paid
campaign.
- No production prompts or model selection change. Review-handoff
behavior and content-redaction policy remain separate product decisions.
The latter can remove harmless business content that looks like
credential syntax; the failing attempt is retained.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 10:35:38 -05:00
DottaandPaperclip 43acbcc398 fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task state to provider sessions.
> - Follow-up turns must retain provider memory and carry new user
direction.
> - Lost session IDs caused repeated context and extra input tokens.
> - Native question answers and approval races could leave valid work
blocked.
> - This pull request repairs those paths and adds regression coverage.
> - Agents can continue accepted work without repeating the conversation
or losing the user's answer.

## Linked Issues or Issue Description

Refs #13574. That merged PR shortened continuation prompts and moved
question instructions into tool documentation. This change preserves
sessions and fixes failures exposed by broader testing. Related runtime
work: #13408 and #13410.

**What happened?**

Native follow-up turns could lose the provider session ID. Completion
guidance could replace the original task with its latest comment. Claude
native questions could remain pending after the user answered. Approval
during a running tool call could suspend the run before the tool
response arrived. Onboarding and chat handoff instructions also caused
repeated planning or missing plan documents.

**Expected behavior**

Reuse a valid provider session. Send only new events when that session
already has the history. Preserve the task requirements and apply later
user direction. Store the question answer and deliver it to the waiting
run. Finish governed tool responses before suspending. Execute the
accepted plan without asking for the same approval again.

**Steps to reproduce**

Run the continuation, local-session-integrity, first-task, and
agent-chat suites with native Codex and Claude. Include
provider-question-bridge, accept-while-running, and plan-handoff.

**Paperclip version or commit**

This branch is based on master d54b75011. The active full catalog run
tests 4e75881db. Later review fixes have separate regression coverage.

**Deployment mode**

Isolated local instances and Daytona sandboxes in the existing Runner
full-stack E2E harness.

## What Changed

- Retain provider session identity across turns and late usage
snapshots. Send new continuation events on session reuse, with full
context available for a fresh session.
- Preserve task requirements and later direction in completion guidance.
Return the current contract revision after a stale completion
submission.
- Bridge native Claude questions to saved Paperclip cards. Submit
answers through the saved card and resume the same run.
- Delay governed suspension until tool results settle. Add a
deterministic test barrier for approval during an active run.
- Clarify free-text question examples, explicit onboarding plans, and
execution of accepted chat plans.
- Fix continuation readiness, verified output evidence, and declared
screenshot collection.
- Qualify the legacy Claude test CLI at 2.1.277. The old 2.1.19 CLI did
not discover mounted skills. Update the existing workflow pin and
isolated launcher together.
- Refresh the Daytona image lockfile integrity pin after reviewing
master patch updates.
- Carry continuation mode as runtime metadata instead of inferring it
from user-visible text. Install the test Claude CLI without lifecycle
scripts.

## Verification

- Targeted paid verification: 20/20 cases passed across
local-environment campaigns before the rebase.
[Report](https://pages.paperclip.ing/runner-e2e-seven-fixes-35397904249/).
- Harness checks: 379 unit tests passed; harness typecheck passed.
- Latest-head PR checks: 55 passed, two intentionally skipped. Greptile
is 5/5; the security scan passes.
- Review regressions: 350 executor tests and 204 session/driver tests
passed. A script-free Claude install was verified with the actual CLI.
- Full catalog, including the explicit-only everyday suite: [run
35417932353](https://github.com/paperclipai/paperclip/actions/runs/35417932353).
Completed: **164/205 passed; 41 failed**. [Full dashboard and failure
investigation](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/).
Includes 204 case artifacts and one pre-case GitHub authorization
timeout; missing evidence is not scored as a pass. The full run tested
`4e75881db`; Final-head metadata/CLI smoke cases both passed. In the
separate [six infrastructure
retries](https://github.com/paperclipai/paperclip/actions/runs/35419769343),
the GitHub timeout case passed and all five Docker preflight failures
repeated. [Follow-up
dashboard](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/follow-up/).
- Full local typecheck and build passed on the rebased branch. The full
local unit run completed with 657 passing files, two test timeouts and
one suite setup timeout. All three affected files passed when rerun in
isolation (84 tests). The first full local run was not clean.
- Focused regression coverage includes the live question bridge,
same-run response delivery, UI routing, stale revisions, approval
overlap, and session reuse.

## Full-catalog follow-ups

- Test infrastructure: 14 Claude everyday cells probe an absent host
CLI; six cells failed pre-task GitHub/Docker qualification (GitHub
passes on retry; all five Docker cases repeat; the workflow preflight
allowlist omits their case IDs); five ACPX Codex cells cannot create
sandbox namespaces.
- Runtime: four OpenCode completion-criteria mismatches masked by
shutdown errors, one service-approval suspension failure; three Daytona
recovery failures encounter existing skill files; one duplicate
completion wake.
- Confirmed test defects: question pagination and a noncanonical plan
document key.
- Product/behavior: mismatched visible/required question sets, an
attachment instead of the requested task document, one lone-option
onboarding question, early completion instead of review, and a Codex
Mini completion-schema failure.
- The report job itself fails on trusted master’s stale patch/lock
configuration. The linked report is rebuilt with the shared renderer
from original cell results and public fixture screenshots; it excludes
private snapshots, logs and traces.

These are investigated follow-ups, not silently regraded passes.
First-task passed 51/52. The PR checks are green independently of the
broader catalog’s behavioral/infrastructure failures.

## Risks

- Session reuse depends on a valid provider identity and context
coverage. Fresh-session fallback and reset tests cover this boundary.
- Native question delivery spans saved interaction state and a live
provider run. Tests cover duplicate events, closed runs, and same-run
answers.
- Provider behavior varies. The full paid catalog may expose failures
beyond these targeted fixes; those results will be reported without
relaxing valid approval or output checks.
- The legacy Claude version update is limited to test infrastructure. No
database migration is included.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code execution,
and browser/E2E tools. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks and all
three timeout-file reruns pass; full-run timeout caveat above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 07:42:57 -05:00
DottaandPaperclip 84fe89906d fix: complete native agent review handoffs (#13581)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Native execution uses durable runs, issue locks, wake requests, and
typed tool authority
> - A child can finish with a native agent review request while its
original assignee stays responsible for the work
> - The reviewer then needs a bounded execution path that can inspect
the child, record one decision, and finish safely
> - Before this change, assignee-only gates rejected the reviewer or
left the parent waiting after the child review ended
> - This pull request adds typed reviewer admission, scoped reviewer
tools, durable wake and recovery handling, and parent continuation
evidence
> - The benefit is that native review handoffs complete without changing
child ownership or granting broad mutation access

## Linked Issues or Issue Description

Refs: #13314
Refs: #13574

**What happened?**

A native child run could report `needs_review` for an agent reviewer.
The reviewer wake then failed assignee and execution-lock checks. The
child remained in review and the parent remained waiting.

**Expected behavior**

The named reviewer should receive one durable wake. The reviewer should
inspect the child and resolve the exact review card. The child assignee
should stay unchanged. The parent should receive the recorded review
outcome after the child reaches its terminal state.

**Steps to reproduce**

1. Run a native task with a different named agent reviewer.
2. Keep the child assigned to its original worker.
3. Let the worker finish with a native completion review request.
4. Start the durable reviewer wake.
5. Resolve the review and finish the reviewer run.
6. Observe the child and parent state.

**Paperclip version or commit**

Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification
source: `eea171aae`.

**Deployment mode**

Built from source.

**Installation method**

Built from source (pnpm build).

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Access context**

Both.

**Database mode**

Embedded PostgreSQL in the isolated live test fixtures.

## What Changed

- Add server-validated native review assignment facts.
- Admit only the exact company, issue, source run, decision, revision,
addressee, and resolver policy.
- Give reviewer runs a narrow set of Paperclip read and resolve tools.
File and shell access follow the configured agent and environment
policy, so reviewers can run tests.
- Separate server-owned reviewer instructions from untrusted persisted
review data. Escape the data boundary; retain server-enforced
authorization.
- Keep the child assignee unchanged. Atomically claim the reviewer run,
wake request, and issue execution lock. A competing lock prevents
provider startup.
- Require the exact running reviewer session and current issue lock to
resolve its assigned card. Reject missing, unrelated, or terminal
reviewer runs.
- Add durable reviewer wake, lock, stale-card, and abandoned-run
recovery handling.
- Prevent duplicate native wake dispatches during deferred admission and
recovery.
- Carry accepted or rejected child review outcomes into parent task
context and continuation evidence.
- Add focused server, runner, and native protocol coverage.
- Preserve upstream continuation rules. Add child review decisions as
separate evidence, while keeping real human answers in their own field.
- Return actionable completion validation feedback to both providers.
Permit a corrected completion after rejection. Keep strict terminal
acknowledgment validation.
- Apply exclusive shared-workspace locks to sandbox environments. Local
and SSH folders can run concurrently, including when old settings
request serialization.
- Repair test timing, native event parsing, and the review artifact
assertion. Allow a valid reject, correct, and accept review sequence.
Check the accepted card against its reviewer run and decision. Keep
polling within the existing deadline when review acceptance precedes the
parent wake projection; report a specific missing-continuation error at
timeout.
- Apply the ACPX pending-call limit to reserved finish/block calls, with
capacity-release and cancellation tests.

## Verification

- `pnpm build`: passed on `eea171aae`.
- `pnpm -r typecheck`: passed on `eea171aae`.
- `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on
`b31ad9ab8`; runner E2E typecheck also passed.
- `pnpm check:token-gates`: passed.
- Focused DB review, reviewer authority, and prompt-boundary checks: 31
tests passed. They cover invalid reviewer runs, competing locks, atomic
admission, duplicate claims, and valid resolution.
- Heartbeat, workspace, and recovery checks: 30 tests passed.
- ACPX sidecar suite: 27 tests passed. Moving the capacity guard back
below reserved handling makes both new regression cases fail.
- Four focused live continuation checks passed on their first attempt at
`f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and
question tool guidance (12/12 each). These cases do not use the reviewer
prompt path changed afterward.
- Fresh Codex and Claude review-handoff checks passed all 29 native
checks each on their first attempt at `eea171aae`. Both runs received
the expected fixed prompt and completed cleanup. Only the six selected
live flows were tested; no full paid provider catalog run.
- The final commit only extracts the existing test-harness timeout
diagnostic into a shared helper and adds positive and negative coverage.
Removing the accepted-review guard makes two regression assertions fail;
restoring it passes all six timeout tests. Production runtime code,
prompts, deadlines, and grading criteria are unchanged by this final
commit.
- Deadline regressions: a valid continuation delayed 20 seconds succeeds
within its 30-second unit-test deadline; an absent wake returns a
specific candidate-failure diagnostic at that same deadline. Both
assertions failed before the fix. Production E2E deadlines remain
unchanged.
- Historical native failures remain recorded: Docker availability
failures; a valid reject/correct/accept sequence that the first-card
grader misread; and a test that rejected the gap between accepted child
review and parent wake projection. No failed result was regraded. The
latest tests use a protected reference to the pinned Docker image and
the unchanged artifact oracle and time limits.
- Full repository verification runs in GitHub CI. Local verification
uses the focused suites above, full build, and full typecheck. An
unchanged Codex shutdown timing test failed once in CI, passed in
isolation, and its full shard passed on the final commit without changes
to that test or its causal code path. The original failure is retained
in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no
outstanding actionable findings. All review threads are resolved. All
current-head CI gates passed, including the isolated native runner
Docker build (55 successful checks; two skipped by the workflow).

## Risks

- Reviewer admission depends on exact persisted decision and interaction
bindings. A stale or changed card is rejected.
- Paperclip control-plane tools are limited to inspection and review
resolution. This is not a filesystem permission boundary; provider file
and shell access retain the configured policy.
- Deferred wake recovery changes dispatch receipt coalescing. A
scheduler regression could delay a continuation if the receipt state is
wrong.
- Parent review outcomes are evidence for the model. They do not grant
tool authority or change issue ownership.
- This change does not address legacy lease-hold handoff behavior.

> Roadmap review: native execution, review gates, and durable recovery
are existing roadmap capabilities. This PR completes a narrow
reliability path for those capabilities.

## Model Used

OpenAI `gpt-6-astra` with reasoning, tool use, and code execution.
OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and
journal work. Context window size is not exposed by this session. Live
test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the
PR authors.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 15:52:19 -05:00
Nicky LeachandPaperclip e926b13017 fix: keep sandbox termination progressing after bridge loss (#13287)
## Thinking Path

> - Paperclip must stop remote execution after losing its controller.
> - A bridge can remain blocked while the sandbox still incurs costs or
performs actions.
> - Waiting forever for that bridge prevents provider termination.
> - A temporary provider outage can also exhaust cleanup attempts
permanently.
> - This pull request bounds bridge drain and persists cleanup retries
with backoff.
> - Cleanup ends only after provider confirmation, without a user
accepting uncertain side effects.

## Linked Issues or Issue Description

Builds on merged #13285. Related #13254 added exact provider termination
receipts; merged #13272 adds explicit user retry. Merged #13352 stops
active sandbox startup before waiting for setup. This PR preserves that
immediate cancellation path and extends bounded teardown to ordinary
release and destroy. Cleanup continues automatically after repeated
provider failures. Refs #12953 for provider failures blocking execution.

**What happened?**
Daytona release waits for in-flight bridge activity before stop/delete.
A dead bridge can prevent that wait from finishing. The host also stops
cleanup after five failed attempts.

**Expected behavior**
Provider termination proceeds after a bounded bridge drain. Cleanup
retries survive service restarts and provider outages.

**Steps to reproduce**
Start a sandbox command whose bridge promise never resolves, then
release its lease. Separately, persist a pending-cleanup lease with five
failed attempts and recover the provider.

**Deployment mode**
Hosted Paperclip with a Daytona provider; rebased onto master at
`728f7185f` on September 14.

## What Changed

- Bound bridge drain and provider lifecycle calls. Prefer stop for
reusable sandboxes, with delete fallback.
- Persist cleanup attempt identity, renewable in-flight deadline, and
cooldown. Fence completion writes against superseded attempts.
- Preserve scoped explicit Retry and its activity log. Explicit Retry
can skip cooldown, but cannot take over a live cleanup attempt.
- Continue cleanup after five failures with slower retries and an
operator warning.
- Exclude leases in cooldown before paging so they do not starve due
work.
- Add hung-bridge, restart, provider-recovery, and concurrent-cleanup
regressions.

## Verification

- Rebased onto master at `728f7185f`. The outstanding diff contains only
cleanup changes; the merged controller-ownership prerequisite is
excluded.
- Daytona plugin suite: 160 passed, including immediate startup
cancellation, graceful release, hung activity, and teardown regressions.
- `pnpm exec vitest run
server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts`: 31
passed. Two added integration cases verify explicit Retry during
cooldown and while another cleanup owns the lease. They also verify run
scoping and the activity log.
- Targeted cleanup and cancellation cases in
`environment-runtime.test.ts`: 20 passed.
- Earlier live disposable Daytona test: provider stop ended background
work, resume preserved files without restarting the old process, a new
command succeeded, and the sandbox was deleted. This verifies provider
behavior; it was not repeated for this rebase.
- Latest-head CI and automated review are pending. Broad local tests,
typecheck, and build were not rerun for this focused rebase; CI supplies
those checks.

## Risks

- Timing out bridge drain permits provider termination; it never
supplies a stop receipt.
- A crashed cleanup attempt remains protected for 15 minutes, then
becomes eligible again. Repeated failures retry every 30 minutes after
escalation.
- The existing counter saturates at the escalation threshold; the new
attempt identity and deadline prevent overlapping claims.
- No schema, UI, telemetry, lockfile, or workflow change.

## Model Used

OpenAI GPT-6 through Codex, using reasoning, repository inspection, code
execution, and test tools. The precise backend revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 10:01:57 -07:00
DottaandPaperclip d0b67bfe71 feat: queue approvals and answers during active runs (#13539)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users guide running agents through messages, questions, and approval
cards.
> - Messages already wait in a queue when an agent is running.
> - Card responses did not appear in that queue. Some question answers
also steered a later run without a user click.
> - A fast approval could invalidate the agent's review handoff and
cause it to stop its own run.
> - This pull request gives card responses the same queue controls and
preserves the exact response during delivery.
> - Users can wait for completion or explicitly send the response with
Interrupt or Steer.

## Linked Issues or Issue Description

Refs #13517, which is merged. This PR targets master and adds queued
interaction responses on top of the onboarding changes. Related
continuation work: #10519 and #12866.

**What happened?**

Accepting a proposal while its source run was active left a saved
response outside the message queue. The agent could then lose its review
path, reassign the task, and cancel itself. Answers to older questions
could also steer another active turn without a click.

**Expected behavior**

Save the response immediately. Queue its continuation behind the active
run. Deliver it after completion, or when the user explicitly chooses
Interrupt or Steer. Preserve approval revisions and answer choices.

**Steps to reproduce**

1. Let an agent publish a confirmation card while its run is still
active.
2. Accept the card before the agent finishes its review handoff.
3. Inspect the message queue and the task's next run.

**Paperclip version or commit**

Reproduced on da8a3876c with the onboarding changes from #13517.

**Deployment mode**

Local development from source. The fix covers legacy adapters and native
Runner turns.

## What Changed

- Project resolved cards into the existing queue as immutable responses.
Keep answers and exact approval revisions.
- Require an explicit click to steer a response into a compatible native
turn. Use Interrupt when a fresh session is required.
- Preserve typed response context through interruption, cleanup waits,
and normal queue promotion. Keep the direct answer channel for a
provider blocked on its original question request.
- Accept the source run's review handoff after its card resolves. Reject
stale agent reassignment that would orphan a queued response.
- Add deterministic regression tests and an `accept-while-running` case
to the first-task suite. Require recorded timestamp overlap before that
case can pass.
- Keep the first-task skill name out of user-facing messages.

## Verification

- Red-green: the original route failed the queue regression; the changed
route passes it.
- Focused server/UI tests: 139 passed, including 64 queue-route tests.
- Runner harness unit tests: 314 passed.
- Server, UI, and Runner E2E typechecks passed. UI token gates passed.
- Full repository typecheck and build passed. Server typecheck passed
again after review fixes.
- Review regressions: 165 queue/reopen route tests, 53 wake admission
tests, and 18 run identity tests passed. Approval acknowledgement
recovery and both message/approval arrival orders are covered.
- Full local test run: 12,401 passed; three new admission regressions
ran against a cached pre-fix module. A fresh run of that entire suite
passed (53 tests). The complete CI suite passed on the final commit.
- Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional
Storybook checks skipped. Every server/workspace/browser shard, Runner
verification, build, typecheck/release registry, canary, policy, and
security check passed. Greptile: 5/5, no unresolved threads. Earlier
interrupted CI workers were replaced by this fresh complete run.
- After integrating the updated parent: 314 harness tests, 119
queue/admission tests, 44 onboarding/question-delivery tests, and 13
native recovery tests passed locally. Full repository typecheck and
build passed.
- Clarified the skill wording preference: routine replies describe the
action without announcing the internal skill; direct questions and
permission/security/execution disclosures remain truthful.
- The paid `accept-while-running` scenario is registered for all four
local first-task profiles. It has not been run against a model in this
change.

- Rebased onto the merged parent at `11921075a`; the resulting tree
exactly matches the locally verified integration tree. Final-head CI on
`b53054807` passed: 54 successful checks, 2 optional Storybook checks
skipped, no failed checks. Every new server/browser shard, aggregate
verify/e2e gate, Runner, typecheck, build, canary, and security check
passed on the first attempt. Greptile reviewed this exact head at 5/5
with no unresolved threads.

## Risks

- Responses now wait instead of implicitly steering another active turn.
A provider blocked on the original question still receives its answer
directly.
- Approval receipts cannot be edited, discarded, or reordered as
comments. This preserves the recorded decision.
- Interruption must still prove that the prior execution stopped. The
tests cover cleanup waits and duplicate delivery.
- The new paid overlap case can be unexercised if the model finishes
before the click lands. It cannot pass without evidence of overlap.
- No database migration is required. This repairs the existing approvals
and execution controls; it does not implement the roadmap's work-stream
queues.

## Model Used

OpenAI GPT-6 through Codex. The exact deployed model ID and
context-window size were not exposed in this session. Capabilities used:
agentic reasoning, repository inspection, code editing, terminal
commands, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 14:44:36 -05:00
DottaandPaperclip f4cdc7b231 fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The control plane prepares task workspaces before it starts an
agent.
> - Workspace preparation reads Git state so it can preserve edits and
exclude private files.
> - A failed scan was treated as a non-Git folder and lost its actual
failure code.
> - The resulting generic setup failure could not recover, even when the
cause was temporary.
> - This pull request keeps the cause and uses the existing bounded
retry schedule before provider startup.
> - Tasks can recover without human intervention, while permanent
failures and exhausted retries stop with useful guidance.

## Linked Issues or Issue Description

**What happened?**

A Git scan error during managed repository preparation became
`Configured repository folder is not a Git checkout`, followed by
generic `setup_failed`. The agent never started. Generic recovery could
not distinguish a temporary timeout from a bad workspace configuration.

**Expected behavior**

Keep the closed scan error code. Retry temporary timeouts and queue
saturation under the existing shared budget. Preserve edits, exclusions,
ownership, and pause gates. Stop permanent failures and exhausted
retries with a specific explanation. Do not replay historical generic
setup failures.

**Steps to reproduce**

1. Configure a task project with a local Git source that must be copied
into its managed repositories.
2. Make the ignored-file scan exceed its timeout before the agent
starts.
3. Before this fix, the snapshot returns null and the run ends as
non-retryable `setup_failed`.
4. Use the disposable browser fixture in
`tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and
test the full recovery path.

**Paperclip version or commit**

Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`.

**Deployment mode**

Built from source. The defect is in core workspace setup, not a specific
model provider.

Related work: Refs #13442 (managed repository preparation), Refs #11572
(bounded Git scheduler), Refs #12997 (separate adapter startup retry
work), Refs #13469 (separate terminal-workspace scan performance work).

## What Changed

- Return the non-Git fallback only for repository discovery. Propagate
failed scans of a confirmed repository.
- Replace full ignored status output with an ignored-only directory
listing. Preserve NUL-delimited paths and exclusions.
- Preserve typed, sanitized scan errors through workspace preparation
and persist pre-provider failure details.
- Retry only timeouts and queue saturation, using the existing durable
two-retry budget and issue gates. Prevent generic recovery from adding
another budget.
- Show workspace-specific failure copy and actionable exhausted-recovery
notices.
- Add red-green unit tests, real-database restart and retry-boundary
tests, and opt-in browser acceptance fixtures with real Git subprocess
timeouts.
- Document the recovery contract and browser verification procedure.

## Verification

- Red: injected scan failures returned null instead of rejecting; setup
lost the timeout code; task-thread and recovery notices had generic
copy.
- Green: 119 focused adapter/backend tests, 20 recovery-boundary tests,
and 136 task-thread tests.
- `pnpm -r typecheck` — passed.
- `pnpm build` — passed on the final production code.
- `pnpm check:token-gates` — passed.
- The initial local `pnpm test:run` overlapped source edits and was
interrupted after two late-added assertions saw pre-fix behavior; it is
not counted as a green full run. A fresh final-head run passed all 229
tests across the six affected adapter/backend/UI suites. The clean
latest-head CI full test matrix passed: all five general-server shards,
all five serialized-server shards, and all three general-workspace
shards.
- Latest-head CI also passed all three browser shards and their
aggregate gate, typecheck and release registry, build, runner
verification, canary dry run, policy, Docker context integrity, and
security gates. Greptile: 5/5, with the review thread resolved.
- Browser: created a task in a disposable instance. A real Git timeout
scheduled recovery, the next run completed through the run-scoped API
without manual Retry, and Done survived reload. The deterministic
process worker checked preserved source edits and excluded private
files; no model calls were made.
- `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec
playwright test --config
tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6
minutes). The persistent case made exactly three failed attempts, never
started the worker, showed the cause-specific notice, stayed stopped for
another scheduler tick, and retained Blocked after reload.
- Extra red-green coverage: 50 recovery tests passed after fixing an
exhausted-bootstrap classification that incorrectly implied unknown
provider actions. Missing or uncertain evidence still retains the safety
hold.
- Verified the documented Git executable override during repository
seeding.

## Risks

- A confirmed repository scan failure now fails closed instead of
falling back to directory sync. This prevents unfiltered copying but
makes previously hidden errors visible.
- Temporary host problems can create up to two additional setup
attempts, 30 seconds apart. Permanent scan errors do not auto-retry.
Generic recovery cannot reset this budget.
- The durable retry path still enforces ownership, pause, and work
eligibility. Integration tests cover restart, duplicate promotion,
pause, exhaustion, and non-retryable categories.
- No schema migration, new runtime setting, new retry budget, production
deployment, or historical task replay.

## Model Used

OpenAI Codex, GPT-5-based coding agent, with reasoning, repository
tools, shell execution, and browser testing. The exact deployment model
ID and context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 12:37:21 -05:00
cceeb0aa66 test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner must support project work, delegation, hiring, and
service access.
> - Browser tests exposed lost connection access, rejected helper
events, and stalled recovery.
> - Some eval failures also came from incorrect fixtures and decision
controls.
> - This pull request fixes those paths and adds eight everyday workflow
stories.
> - The tests retain observed failures and verify delivered files
independently.
> - The benefit is repeatable evidence for common user tasks and their
remaining gaps.

## Linked Issues or Issue Description

Related work: #13404 contains earlier workflow fixes. #13300 and #13470
changed the CI contracts used by the harness security tests. Merged
companion:
[paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22).

**What happened?**

Native ACPX sessions did not receive the assigned connection gateway.
Codex helper events could arrive before their spawn receipt and fail
thread validation. A parent continuation could take a shared workspace
before its child retried. A failed native continuation could leave the
task status without a clear recovery blocker. The eval harness also
confused tool approvals with new connection requests and could reject a
valid delegated download.

**Expected behavior**

Keep assigned gateway access and its approval checks. Verify helper
lineage before accepting helper progress. Let a waiting child proceed
before automatic parent recovery. Preserve a failed task's recovery
ownership. Grade the actual requested workflow and its delivered files.

**Steps to reproduce**

Run the everyday workflow suite with the native Codex and Claude
profiles. Exercise service approval, connection refusal, delegated
project work, and teammate reuse. The commands and case requirements are
in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm
test:runner-recovery` for controlled crash and replacement cases.

## What Changed

- Pass the scoped connection gateway binding through the native ACPX
host and sidecar.
- Recognize Codex helper lineage from parent metadata and spawn
receipts. Verify early helper events with `thread/read`. Keep helper
events separate from root completion authority.
- Guide agents to use persistent hiring, child tasks, dependency
records, and a blocked handoff while waiting for a child.
- Defer automatic parent recovery while a child has an active execution
path in the same shared workspace. Allow parent recovery when the child
needs review.
- Record Blocked status and recovery evidence when a failed native
continuation needs reconciliation, including existing active or
escalated incidents. Preserve their owner and retry budget.
- Add eight browser-driven workflow cases. Use real decision controls,
explicit child feedback delivery, managed hiring credentials, and
independent ZIP checks inside a bounded Docker sandbox. Verify sandbox
availability before task creation. Record screenshot SHA-256 at capture.
- Keep runner crash probes in controlled recovery tests. Preserve the
original failure when cleanup also fails.
- Display missing accounting and replay revisions as unavailable. Align
harness security assertions with the approved CI changes.
- Make the channel-rejection browser fixture bind its file after the
send captures its payload. This prevents live refresh from removing the
file before the simulated race.

## Verification

- Full workspace `pnpm -r typecheck` passed after merging current
master.
- Runner E2E typecheck passed. Harness unit tests passed: 216/216.
- Wake-queue database tests passed: 55/55. The two added
existing-incident tests failed before the fix and pass after it.
- Docker artifact calibration passed: 12/12. Host-file and host-loopback
isolation tests failed before the fix and pass after it. Read-only
delivery and output limits are also verified.
- Full `pnpm build` passed. Targeted recovery tests passed: 83/83.
- The channel-rejection browser test passed five consecutive runs after
fixing the fixture race found in CI.
- Local general-server (12,351 tests), UI (6,250), CLI (485), and
workspace package groups passed. The monolithic run stopped at an
unchanged lock-heartbeat fixture race; the isolated workspace group
passed on rerun (shared: 747/747). A separate local serialized run
passed 97 files before two socket errors in the unchanged issue-list
route suite; that suite passed 15/15 on isolated rerun. These local full
commands did not finish uninterrupted; the complete CI matrix below
covers the remaining suites.
- Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful
checks, 2 expected skips**, including every server/workspace shard,
browser shard, native runner verification, build, and typecheck. [Final
CI
run](https://github.com/paperclipai/paperclip/actions/runs/34989136700).
- Greptile reviewed this exact head at **5/5**; all review threads are
resolved. Both Superagent security checks are successful.
- ACPX credential-boundary tests passed: 118/118. Superagent accepted
the runner/sidecar versus provider-environment trace and cleared its
finding.
- The latest paid local campaign on source
`f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8,
Claude 7/8, Mini 7/8. These results predate the merge with current
master.
- The two remaining failures are in `hire-reuse`: Claude exceeded the
attempt deadline during final review; Mini made invalid deliverable tool
calls and remained Blocked.
- Six Daytona cases were not run because the matching immutable runner
image was unavailable. This PR does not claim new remote model results.

## Risks

The changes affect connection admission, helper identity, and recovery
scheduling. Assigned gateway grants and user approval still govern
service calls. The workspace admission gate still exists; the broader
folder-sync design is separate work. Provider behavior can still cause
the two recorded hiring failures. No database migration is required.
Paid cases are opt-in and have bounded attempt deadlines. Project
stories now require Docker and the documented pinned Python image on the
harness host.

## Model Used

OpenAI `gpt-6-astra` performed implementation, diagnosis, and
substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR
preparation, and review tracking. Both used repository tools and code
execution. Context-window sizes were not recorded.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks and
isolated reruns; full-run limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 11:04:16 -05:00
DottaandPaperclip f1d57863d2 fix: make connection checks and task handoffs reliable (#13404)
Preserve connection-probe outcomes through cleanup, reduce unrelated startup work, and report selected Claude authentication accurately. Make artifact download actions match their labels.

Route delegated feedback through its active child, retain accepted messages across completion, and avoid redundant worker runs for proven closing notes. Preserve explicit follow-ups, human input, company boundaries, source provenance, and mixed issue references.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-14 10:06:06 -05:00
DottaandPaperclip 4cd7b40255 fix: recover new messages after historical native runs stop (#13405)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native recovery must distinguish a fresh user request from replay of
failed work.
> - Older runs can lose their process fields before a local stop receipt
exists.
> - A suspended durable session can still prove that the exact runner
and provider session are idle.
> - The message admission path ignored that evidence and kept new user
messages blocked.
> - This pull request uses the existing exact-state verifier for those
historical runs.
> - The user can start one fresh turn while the old history and unknown
outcomes remain intact.

## Linked Issues or Issue Description

**What happened?**
A user sent a new message after a native run exhausted recovery.
Paperclip saved the message but said the previous run had no verified
stop record. The old runner was suspended, with no active provider turn
or pending output. Its process fields had been cleared before stop
receipts were added.

**Expected behavior**
A new user message starts a fresh turn when the exact retained session
proves it is suspended and the other execution gates pass.

**Steps to reproduce**
1. Retain a failed native run with a terminal controller, cleared
process fields, and no process receipt events.
2. Retain its exact suspended runner state and idle provider state. Keep
its recovery hold.
3. Send a new user comment. Before this fix, admission returns no
successor.

**Paperclip version or commit**
Reproduced on master at d351e08de.

**Deployment mode**
Self-hosted server. The regression uses an embedded PostgreSQL test
database and real durable state files.

Related work: Refs #13270, Refs #13338. This adds compatibility for
older stopped runs. Refs #13332 concerns separate recovery-hold scope
rules.

## What Changed

- Add exact suspended-state evidence to explicit native message
admission for runs that predate process receipts.
- Reuse the existing failed-retry verifier for run, runner, workspace,
provider identity, and pending-work checks.
- Reject this fallback if any server-authored process receipt or launch
event exists.
- Add a red/green regression with real message admission, duplicate
delivery, dry-run behavior, and blocked-state cases.
- Document the new-message recovery rule.

## Verification

- Red: the historical suspended-state regression failed on master
because admission returned null. Eight rejection cases passed.
- Green: 469 tests passed across explicit native continuation and native
session execution.
- Full repository `pnpm -r typecheck` and `pnpm build` passed locally.
- The complete Vitest CI matrix and all browser test shards passed. The
duplicate local `pnpm test:run` was stopped after these CI results; it
did not complete locally.
- The first runner verification worker received an infrastructure
shutdown signal during compilation. The single retry passed.
- Greptile reviewed commit `6836d7310` at 5/5 with no findings.
- Inspected the affected server's database and durable state read-only.
It has the historical missing-PID shape and an exact suspended runner
with no active provider turn, pending tools, or undelivered output.

## Risks

- Incorrect idle evidence could allow overlapping work. The fallback
requires an exact suspended root and rejects active or pending work,
mismatched identities, missing files, and newer process evidence.
- Normal task, controller, lease cleanup, decision, and active-run gates
remain in force.
- This does not resume old provider actions or reset recovery attempts.
Unknown outcomes remain unknown.
- No schema change or deployment is included.

## Model Used

OpenAI Codex, GPT-6. The session does not expose the exact model
snapshot or context-window size. Used reasoning, repository inspection,
code editing, shell execution, and test tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 09:45:34 -05:00
DottaandPaperclip f2c5e54dca fix(runner): preserve handoff work and publish requested files (#13355)
## Thinking Path

> - Paperclip lets people manage AI agents and their tasks.
> - A task keeps its instructions, progress, and files when its assigned
agent changes.
> - The replacement runner lost the interrupted run's context and could
overwrite an existing draft.
> - A saved message also stayed attached to the former agent and could
reopen the task after the replacement finished.
> - File tasks could report Done with only a local path that the user
could not open.
> - This pull request transfers handoff context and saved messages, and
makes requested files accessible through the existing attachment
contract.
> - Users can change agents and collect completed work without repeating
instructions or confirming bookkeeping.

## Linked Issues or Issue Description

Refs #13338. Builds on merged #13354 for queue admission and #13353 for
remote workspace retry. #10123 concerns restricted recovery-model
escalation; this change instead covers ordinary native handoff and file
completion.

**What happened?**

Codex wrote a draft before a user assigned the task to Claude. The
replacement lacked continuation context and replaced the draft. A queued
user message could later restart the former agent and reopen the
completed task. Separately, a runner could finish a requested file but
return only a machine-local path. Remote native runs had no bound file
publication tool.

**Expected behavior**

The replacement reads and preserves existing work, receives saved
messages once, and keeps each message's author. The former agent stays
stopped. A requested file has a working attachment or accessible work
product before Done. Text-only tasks do not require attachments.

**Steps to reproduce**

1. Ask Codex to save three newsletter names and then wait.
2. Queue an instruction to keep those names and expand the draft.
3. Use Interrupt and assign to select Claude.
4. Verify the original names survive, the result has a working download,
and only the source and replacement runs exist.
5. Ask either provider for a Markdown checklist and open the file from
its completed response.

## What Changed

- Carry the exact same-task interrupted run's summary, semantic
receipts, and history into handoff context. Tell the replacement to
inspect existing files before editing.
- Adopt saved ordinary task comments into the successor's receipt under
the task lock. Preserve authors and separate mention, chat, and
interaction contracts.
- Prevent a former-assignee comment wake from reopening a completed task
or starting a stale execution.
- Reject workspace-only, fabricated, and cross-task file completion
references with actionable runner feedback. New file output also needs a
matching current-run publication receipt and asset filename/size/hash,
or an accessible work product registered by the current run. Prior
output can remain context alongside a current file, or be verified and
re-registered internally. Authorized chat attachment reuse retains its
verified current-run clone receipt; older receipt shapes require an
intact matching source.
- Bind remote file reads to the active environment runner and reuse the
existing attachment and work-product publication path.
- Enforce workspace confinement, regular single-link files, stable
identity, a 10 MiB limit, and exact size and SHA-256 checks. Rotate the
native session fingerprint for the updated tool contract.
- Contain rejected remote signals and protocol-failure cleanup,
including logging failures. Preserve the original cleanup rejection for
its owner; a rejected operation never supplies stop acknowledgement or
cleanup proof.
- Allow exactly one maximum-size base64 file through the native SSH
command adapter, preserving a finite output cap.
- Document handoff and accessible file completion rules.

## Verification

- Each observed bug has a failing regression before its fix. Final
post-rebase integration passed 732 tests across 13 files before the
final receipt and signal guards; final affected results are below.
- Publication provenance and compatibility: 8 provenance regressions and
2 compatibility regressions failed before their fixes; the final four
affected suites pass 53 tests, including mixed old/new references and
real authorized chat reuse. Controls cover old attachments and work
products, filename/size/hash/origin mismatch, missing/wrong receipts,
current-run publication, same-run durable proof, internally
re-registering preserved bytes, and no-new-file follow-ups.
- Remote signal rejection: the real Node subprocess previously exited 1
when the production launcher signalled a deleted sandbox. It now stays
alive for both a failed signal and failed logging; all 349 executor
tests pass. The failed signal still provides no termination proof.
- Remote file reader and SSH command boundary: 33 tests passed,
including real Linux descriptor reads and the actual SSH adapter
subprocess output cap (network executable replaced by a deterministic
fixture). Exact 10 MiB bytes pass, one byte beyond the encoded cap
fails.
- Live local Claude and Codex Stop journeys preserve the saved file,
deliver queued instructions once, and reach Done with two total runs.
The handoff journey preserves the original names and download with
exactly two runs. Both local providers deliver exact checklist files
without a completion confirmation.
- Live combined Daytona verification passed: the original failed task's
Retry reused its sandbox; a selected Git subfolder produced an exact
downloadable file; warm and deliberately resumed Claude runs took about
33 seconds. Codex produced a 240-byte download in 33.1 seconds after 134
seconds of contention/backoff. Both cloud downloads retained exact bytes
after the two owned sandboxes were deleted.
- The sandbox-deletion retest identified a separate ignored promise in
protocol-failure cleanup. Two real Node subprocess regressions failed
under fatal unhandled-rejection policy before the fix; all 36 protocol,
lifecycle, and integrity tests now pass. The original close promise
still rejects to its owning runtime. The final live retest passed: a
normal Claude Daytona task completed in 132.352 seconds, then its
sandbox was deleted. Thirteen samples over 361 seconds confirmed the
same controller stayed healthy, the task stayed Done with unchanged run
IDs, and its attachment retained exact bytes. The post-deletion browser
download passed with zero page errors; all five owned sandboxes are
confirmed absent.
- Full repository typecheck and build passed on final commit
`de64d16f1`. Final-head Greptile is 5/5 with no unresolved threads.
[Final-head
CI](https://github.com/paperclipai/paperclip/actions/runs/34736623758)
passed: 32 successful checks and two conditional skips. The earlier
mixed-source full local test invocation was deliberately stopped before
rebase, so no pristine green full local aggregate is claimed. Its known
failures passed in later affected suites.

## Risks

- Handoff may adopt only ordinary comments from its validated former
owner. Other delivery contracts must remain independent.
- File verification fails closed if a remote file changes during
reading. The runner must retry publication or explain a blocker.
- The updated session fingerprint starts a fresh provider process where
needed to install the new tool contract.
- No schema migration or historical status reconciliation is included.

## Model Used

OpenAI `gpt-6-astra` through Codex, with reasoning, code execution,
browser testing, and tool use. The context-window size is not exposed in
this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-13 08:41:20 -05:00
DottaandPaperclip c9e3bb7ca4 fix: preserve queued work after native Stop and honor steering support (#13354)
## Thinking Path

> - Paperclip lets people manage AI agents and their tasks.
> - The runner owns execution, while the task keeps user instructions
and status.
> - Stop must stop the current response without losing instructions that
the user already sent.
> - The queue stored its original reason inside saved context. Recovery
checked the outer deferred reason and left the message waiting.
> - Claude also exposed Steer through a shared method even though its
driver did not support it. A rejected request could remove its own error
row.
> - This pull request keeps queued work until execution has stopped,
uses the driver's real capability, and preserves completion event order.
> - Users can continue work without repairing task state or repeating
messages.

## Linked Issues or Issue Description

Refs #13338. Related recovery work: #13353.

**What happened?**

A message sent during a native run stayed queued after Stop. Claude
exposed an unsupported Steer action. A steering failure could hide the
queued row and its error. A terminal event could also precede the final
provider result, and subtree Stop omitted the board actor.

**Expected behavior**

Stop ends the current execution. Once Paperclip proves that execution
has stopped, it delivers the saved instruction once through normal
admission. Pause and recovery holds still prevent dispatch. Unsupported
controls stay disabled, and a rejected action leaves an actionable error
visible. Final results precede terminal events.

**Steps to reproduce**

1. Start a Claude or Codex task that writes a file and then waits.
2. Send a follow-up instruction while it runs.
3. Press Stop. Check that the queued instruction runs once and preserves
the file.
4. Check Claude's Steer control and simulate a server rejection on the
only queued message.

## What Changed

- Recover saved native comments after acknowledged Stop using their
original wake reason.
- Require durable remote termination receipts or verified local process
termination before dispatch.
- Preserve actor identity, queued-message deduplication, Pause, and
recovery gates.
- Derive steering support from the driver descriptor and reject
unsupported calls.
- Keep the queue mounted until a steering request succeeds so its error
remains visible.
- Emit provider results before terminal events and pass the board actor
into subtree Stop.
- Document Stop and steering behavior.

## Verification

- Final-head continuation suite: 104 passed, after failing regressions
for saved wake reasons, cleanup proof, and deduplication of every queued
message. Steering UI: 121 passed. Driver capability: 26 passed. Runner
backend/transport: 205 passed; Rust library: 285 passed.
- Real Claude and Codex browser journeys both preserved the saved file,
delivered the queued instruction once after Stop, and reached Done with
exactly two total runs. The process Stop browser fixture also passed.
- Local full repository typecheck and build passed on `afaa35139`;
server typecheck and the affected 104-test suite passed after the final
queue changes. Token gates passed. Final-head CI verifies the complete
integrated source.
- Local aggregate evidence has explicit limits: the general-server
invocation overlapped the queue fixes and finished with 11,989 passed, 2
failed, and 80 skipped; both failures are covered by the final 104-test
pass. The UI and CLI then passed all 6,184 and 485 tests; the complete
145-file serialized rerun passed all 2470 tests. The shared-package lock
fixture passed unchanged on rerun, but the package phase subsequently
stopped at an embedded-Postgres bootstrap resource failure. No single
pristine green local full aggregate is claimed.
- Greptile reviewed `fa66e2bd5` at [5/5 with no unresolved
findings](https://github.com/paperclipai/paperclip/pull/13354#issuecomment-5650334242).
[Final-head CI completed
successfully](https://github.com/paperclipai/paperclip/actions/runs/34733781888/attempts/2):
33 successful checks, 2 conditional skips, including all server,
package, UI, browser, runner, typecheck, and build gates. The first
attempt hit a preview-readiness/port-collision fixture; its unchanged
local control passed 25 tests with 3 skips, and one supported unchanged
CI retry passed the affected shard and aggregate gates.

## Risks

- Queue recovery must never overlap an old execution. Unknown cleanup
state remains blocked.
- Driver descriptors are now authoritative for steering; a wrong
descriptor disables the action instead of attempting it.
- No schema migration or historical status reconciliation is included.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, browser
testing, and tool use. The exact hosted model ID and context-window size
are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 22:29:54 -05:00
DottaandPaperclip 422287eecd fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task messages, provider execution, and
task outcomes.
> - First-time user tests exposed gaps in recovery, completion
permissions, message delivery, and Stop behavior.
> - These gaps left usable output hidden, completed work waiting for
bookkeeping, or safe work unable to continue.
> - This pull request fixes the shared lifecycle and receipt paths while
preserving process ownership and action checks.
> - Users can continue work with accurate task state and durable
messages.

## Linked Issues or Issue Description

**What happened?**

A stopped local Codex execution could remain blocked even after its
processes had stopped and its complete transcript proved that no
external action needed replay. Claude under Conservative permissions
could fail to call task completion tools. Recovery could reuse an
assistant item ID and overwrite prior output. A delivered comment could
remain marked uncertain after navigation. Stop could look like Pause or
a new recovery incident. Workspace contention could look like
cancellation. A direct reply reopening Done could enter a clarification
loop.

**Expected behavior**

Recover automatically only with verified termination and complete action
receipts. Preserve answers and messages. Keep task completion available
under Conservative permissions without broad tool access. Show crashes
as Blocked, actual human decisions as In Review, and ordinary workspace
contention as waiting. Stop the current response and allow a new
direction.

**Steps to reproduce**

1. Create ordinary response tasks with local Codex and Claude Code, then
send follow-up messages through the task composer.
2. Interrupt a disposable local Codex runner during text-only work.
Verify automatic continuation and retained output.
3. Stop a response, send a new request, answer a clarification, and
reopen completed work with another message.
4. Navigate or reload while a comment submission is pending. Confirm the
exact persisted request receipt settles it without removing newer draft
text.
5. Run two tasks in a shared Daytona workspace. Confirm waiting does not
appear as failure.

**Paperclip version or commit**

Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`.
Current integration base: `6cef9743c`. Both operator-interruption and
workspace-waiting guards are preserved; native restart and legacy
permission rules remain documented.

**Deployment mode**

An isolated source-built test-drive instance, with real local Codex and
Claude Code providers and disposable Daytona environments.

Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163.
This PR addresses additional failures from ordinary task journeys,
including controller restart handoff and repeated warm sandbox setup.
Historical task status reconciliation is excluded.

## What Changed

- Persist runner ownership immediately at spawn and resume an explicitly
adopted runner even when the controller crashed before the first driver
checkpoint. Detach the controller safely across graceful restarts,
including session startup. Prevent an old finalizer from suspending or
signaling an adopted runner. Checkpoint idle warm sessions before
shutdown. Preserve the same run and queued follow-up messages.
- Scope saved legacy queue successor checks to the queue owner while
preserving ordinary task locks, operator identity, assignment gates, and
exactly-once delivery.
- Preserve managed Codex credential files when an old session is
detached for restart; normal owned cleanup still copies refreshed auth
back and removes the scoped copy.
- Reuse the bound warm shared sandbox and fully verify an existing
staged provider pack before using it. This avoids repeated uploads when
the pack is already valid.
- Add a narrow local Codex replacement path with stopped-process proof,
a closed transcript inventory, exact completion receipts, and
fresh-session lineage. Preserve no-replay holds when evidence is
incomplete. Recovery may clear only the same run's recorded Blocked
status version; manual re-blocking and dependency changes invalidate
that receipt, while queued comments do not. Later blocks stop scheduled,
queued, and final dispatch; queued/final checks re-read dependencies
even when the task status stays In Progress.
- Permit only task delivery and human-input tools through the isolated
Claude runner's exact task bridge.
- Scope assistant item identity to the provider turn and ignore only
authority-free Codex skill-change notifications during startup.
- Reconcile composer submissions by client request ID across response
loss, navigation, and reload. Retain text typed during delivery.
- Keep acknowledged run-only Stop neutral and show workspace contention
as waiting. Project exhausted native failures as Blocked.
- Restore the guarded task-page retry action for failed legacy runs,
including the server-supported explicit new-attempt path for stopped
conversation adapters. Preserve native/process recovery holds and avoid
promising Retry while a decision or execution gate hides it.
- Refresh delivered artifacts and handle direct user replies that reopen
completed work without a clarification loop.
- Check the embedded PostgreSQL PID, data directory, and actual port
before connecting or migrating.
- Document accepted behavior and add focused regressions at lifecycle,
route, transcript, and UI boundaries.

## Verification

- Final head `fece606ac2` passes the complete GitHub CI matrix: **34
green checks, two expected Storybook skips, no failures or pending
checks**, including `ci / verify`, `ci / e2e`, full runner verification,
typecheck, build, every server/workspace shard, and all browser shards.
[CI
run](https://github.com/paperclipai/paperclip/actions/runs/34727183287).
Greptile is **5/5 with no open findings**. The final two commits only
refine test fixtures; both affected suites pass 24/24 locally and in CI,
with server typecheck green.
- Complete local Vitest coverage uses the canonical groups/shards: all
635 general server suites, all 145 serialized suites, and all workspace
packages. The aggregate began on `0a8001c18` while the final queue fix
arrived: 23,903 passed, five failed, 87 skipped. The five
port/socket/timing failures passed unchanged in follow-ups (60 tests in
the exposure/file suites and 412 tests covering the serialized failures
and unrun tails). The final queue/operator-identity suites separately
passed 52/52. This is aggregate coverage plus explicit reruns, not a
pristine single-command final-head run.
- After integration with current master,
queue/operator-identity/continuation suites passed 162/162 and affected
UI suites passed 140/140. ACP Stop/continuation and legacy
task/Inbox/message browser suites passed 9/9, including both task
recovery Retry and thread Try again, automatic saved-message delivery,
exactly one new run, Done, and retained output after reload. The default
process Stop/Pause/Resume browser case passed (the native-provider case
is opt-in and skipped by default). The complete Board attachment/receipt
browser suite passed 11/11 on a disposable instance, covering both
composers, exact receipts after lost responses, no replay, bound
attachments, and newer drafts after reload.
- Blocking-intent regressions cover pre-existing Blocked, a mismatched
run/cause, an explicit manual re-block, changed dependencies, a queued
comment after failure, and a block arriving between scheduling and
provider dispatch. The negative cases reproduced before the fix. All 478
affected executor/recovery/dispatch tests passed; both database suites
ran separately after availability-probe skips in the first combined
command. The final late-dependency check passed all 143 affected
recovery/dispatch tests (zero skips) after two new negative cases
reproduced the bug.
- Focused runtime regressions cover awaited runner ownership
publication, authenticated adoption before the first checkpoint,
old-finalizer detachment, idle and busy warm-session shutdown, rejected
checkpoint propagation, provider-pack verification, and managed-Codex
credential preservation. Four managed credential detachment cases
reproduced the bug before the fix; normal owned cleanup still succeeds
exactly once.
- Live local Claude: SIGKILL 2.6 seconds into startup recovered the same
run automatically in 53 seconds, then a normal follow-up completed in 24
seconds. SIGTERM 2.5 seconds into startup preserved the same run (54
seconds) and its queued follow-up (21 seconds). Answers remained visible
and the task reached Done.
- Live Claude Daytona: a warm follow-up retained its sandbox and fell
from 121 seconds to 44 seconds. A separate cold turn took 127 seconds;
after controller shutdown and checkpointing, its follow-up completed in
33 seconds with the same sandbox, workspace, native session, and runner.
Both answers remained visible and the task was Done.
- Other live journeys covered task completion and follow-up with local
and Daytona Codex, local Codex crash recovery, Stop then new direction,
clarification response, live artifact refresh, and shared-workspace
waiting.
- Validation limits: the opt-in native composer Stop/Pause→subtree
Resume fixture exposes terminal/result ordering and subtree-cancellation
attribution bugs that can leave a child task blocked; that new finding
is assigned to a separate follow-up and is not claimed fixed here.
Default CI skips this optional native-provider fixture. Managed-Codex
credential handoff and the queue-agent integration use automated
regression evidence. Cold custom provider-pack uploads still add startup
latency.

## Risks

- Automatic replacement remains deliberately narrow: local Codex,
verified stopped identities, unchanged retained state, and a complete
text/completion-only turn. Unknown actions, partial history, or changed
ownership remain blocked.
- Claude completion permission handling changes an upstream package
patch. The exact isolated task bridge must remain pinned; unrelated
tools keep their existing permissions.
- New task failure projection changes user-visible status. No historical
status backfill or database migration is included.
- This is a broad lifecycle fix across server and UI. Live proof covers
graceful local Claude restart during startup and idle Claude Daytona
session recovery across controller shutdown. Live abrupt SIGKILL during
local Claude startup also recovered the same run. Unknown ownership or
missing action evidence still blocks reuse. Cold custom provider-pack
uploads still add startup latency; this change avoids unnecessary repeat
uploads.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, code execution, browser
automation, and tool use. The exact hosted model ID and context window
are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 19:41:15 -05:00
DottaandPaperclip 6cef9743c0 fix: deliver saved user messages after recovery stops (#13327)
Deliver saved user messages after legacy recovery stops. Validate undelivered comments and queue ownership under the task lock, preserve the operator identity checks from #13315, and prevent duplicate successors.

Add a recovery notice with Retry and inline errors, plus service, route, component, and browser coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-12 18:08:04 -05:00
DottaandPaperclip df984cbc2c fix: dispatch queued legacy messages with operator identity and task permissions (#13315)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task conversations save messages that arrive during an active turn.
> - Legacy adapters deliver these messages in a later turn.
> - A run can stop before the saved queue is delivered.
> - The Interrupt button previously required an active run, so it could
not release this queue.
> - This pull request lets a board operator send the saved queue after
the run stops and retries queues missed during finalization.
> - Manual dispatch must use the clicking operator and must not require
permission to create agents.
> - The task can continue without a duplicate message or a second
execution owner.

## Linked Issues or Issue Description

**What happened?**

A legacy task retained a queued message after its run stopped. Interrupt
was disabled because the queue had no active target. Finalization and
deferred message admission can also leave a queue without a successor.

**Expected behavior**

Interrupt sends the saved messages when no runner is active. Messages
that arrive during normal completion are delivered automatically. An
uncertain previous execution still requires proof that its process or
sandbox stopped.

**Steps to reproduce**

1. Queue a user message during a legacy conversation turn.
2. Let the turn stop or simulate a server restart before queue
promotion.
3. Open the task with a deferred queue and no active run.
4. Try Interrupt. Before this change, the button is disabled.

**Paperclip version or commit**

Reproduced against `8f40b4ad4`.

**Deployment mode**

Legacy conversation adapter. The same persisted queue state is covered
with an isolated PostgreSQL fixture.

Related public work: #13275 adds active legacy interruption. #13291
addresses automatic sandbox conversation recovery. This change handles
explicit saved-queue delivery and late queue promotion.

## What Changed

- Accept a null Interrupt target while retaining queue identity,
revision, company, and assignee checks.
- Save the operator's request on the existing queue. Reuse normal
admission after verified stop, including older messages, different
authors, and queues whose original wake came from the system.
- Strip interruption authority from caller-supplied wake payloads. Only
the board queue route can persist that authority.
- Retry durable interruption requests after restart and deferred queues
after legacy cleanup.
- Let an explicit Interrupt retry cleanup for its stopped run, including
old ephemeral leases that recorded success without a provider stop
receipt. Preserve retained resources, other lease owners, and the
automatic retry limit.
- Preserve the server's waiting explanation when normalizing and
combining queue entries.
- Revalidate the consumed board queue receipt at dispatch so a different
message author does not cause setup failure.
- Use the Interrupt user's execution identity for the new run. Preserve
original message authors. Validate the receipt independently at startup
and inherit the resulting identity on retry.
- Persist authenticated board authority for ordinary manual wakes too.
Adopting someone else's queued messages cannot switch a manual run to
that author's permissions. Strip caller-supplied authority markers and
retain private conversation ownership checks.
- Keep the clicking user when a manual wake is merged into an older
deferred receipt. Update its requester and payload in the same
transaction.
- Use the same current-queue/revision API on task details and pipeline
conversations; show Interrupt after a legacy target stops.
- Keep manual wakes out of active runs, including unscoped agent wakes.
They receive their own execution identity; a matching receipt requester
is not sufficient because an exact retry can retain a different
originating identity.
- Authorize both existing-agent wake endpoints with `agent:wake`,
available to active non-viewer company members. Keep `agents:create` for
hiring. Validate the stored task and current assignee before an exact
task retry.
- Reject viewer Interrupt requests before saving intent or stopping
execution. Keep external chat retry authorization and per-action
agent/user permission checks.
- Preserve edits and discards until dispatch. Prevent another queue
promotion when the same agent already has a successor. Keep independent
reviewer recovery available.
- Suppress cancelled/failed run toasts for intentional operator
interruption. Keep ordinary runtime error notices.
- Add UI, route, admission, restart, successor ownership, and toast
regression tests. Document the behavior.
- Reuse the existing socket reservation helper for both
credential-quorum test cases after CI exposed an ambient-port collision.
This changes test preparation only; production credential staging is
still called exactly once.

## Verification

- Failing regression tests reproduced the message-author identity bug
and an operator's `agents:create` rejection before the fixes.
- All 316 focused tests pass across eight route, queue, identity,
authorization, continuation, and responsible-user suites, including the
44-test rerun of queue admission and actual startup after the final
manual-wake restriction. Regressions reproduce cross-user merging both
with and without a task, and same-requester receipt ambiguity. The
cross-company existence guard also passes both tests.
- Startup integration tests reach adapter execution under the clicking
operator and retain that identity through follow-up. Coverage includes
mixed authors, adopted queues, system-origin queues, restarts, forged or
stale receipts, viewers, suspended memberships, changed assignees,
private conversations, and caller-supplied authority markers.
- The earlier queue/cleanup/UI regression suite passed 402 tests. The
final review corrections pass another 180 tests across queue
admission/persistence, real heartbeat startup, UI API, conversation
rendering, and pipeline suites. Regression tests reproduced both review
findings before correction. The final head has a 5/5 review with no
unresolved threads. Full CI passes on `c2002979c`, including every
general and serialized server shard, all browser shards, Paperclip
Runner verification, typecheck, build, canary dry run, and the aggregate
gates.
- Full `pnpm -r typecheck`, `pnpm build`, and UI token gates pass after
the final application changes. CI identified an outdated task-page API
mock after the shared helper extraction; the fixture now exercises the
real helper, and all 131 task-page/API tests pass. The final application
build passes with the additional manual-wake restriction.
- CI exposed a pre-existing port collision in the Codex
credential-quorum fixture. It reproduced locally; both listener cases
now use the existing bounded reservation helper. All 41 credential tests
pass on rerun. One intervening local run hit a separate ambient bind
collision in the two-occupied-port case.
- The full local `pnpm test:run` attempt was stopped after host
contention caused focused-suite timeouts. The affected focused tests
passed on rerun. An expiring trace fixture and a missing
private-conversation state were corrected. The successful full CI run is
the complete-suite verification.
- Hosted Interrupt previously cleared the original queue and produced
exactly one successor with neutral interruption feedback. It exposed the
dispatch authorization defect. Retry on that earlier build was rejected
for missing `agents:create` before creating another run.
- Deployed the final application build (`38257f391`) to the scoped
hosted instance and verified readiness. The latest PR commit changes
only the credential test fixture; application code matches that
deployment. A live Retry by the same operator without `agents:create`
created one successor attributed to that operator, passing the former
dispatch permission gate. Startup then stopped at
`configuration_incomplete` because that operator has not configured
their required personal Claude Code OAuth secret; the post-deployment
run page confirms the operator identity and no provider work started,
and the My secrets UI still shows the token as not set. Provider
execution remains unverified pending that credential. No permission
grants or credentials were changed.

## Risks

Queue admission and finalization can race. The task lock, durable queue
receipt, current comment IDs, and successor guard prevent duplicate
dispatch. Process and lease stop checks, task pauses, approvals,
ownership, and budgets remain in force. The API change only allows null
on legacy Interrupt; native steering still requires an active run. No
schema migration is required. Active non-viewer board members can now
invoke existing agents without agent-creation permission. Agent
self-invocation rules, raw provider-trace admin access, task retry
scope, external chat authorization, and action-specific user/agent
permissions remain enforced.

## Model Used

OpenAI GPT-6 through Codex. The session does not expose an exact backend
model ID or context-window size. Used reasoning, repository search, code
execution, tests, and browser tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 17:11:30 -05:00
DottaandPaperclip 0e14c61da7 fix: fence native startup against cancellation (#13316)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can pause a task while its runner prepares to start.
> - Cancellation must prevent preparation from creating new execution
authority.
> - Native runtime selection could run after cancellation and leave an
unclaimed recovery coordinator.
> - Saved user messages then waited for recovery that had no eligible
worker.
> - This pull request fences startup and lets explicit user continuation
settle verified, unclaimed startup state.
> - Tasks can continue after cleanup while keeping provider ownership
and execution safeguards.

## Linked Issues or Issue Description

Refs #13285 for startup controller leases and #13270 for explicit
continuation and saved-message recovery. Related #13293 covers retained
processes that actually started; this change covers cancellation before
the native provider claim. Related #13315 covers legacy queued-message
delivery.

**What happened?**

Pausing a task during startup could cancel its heartbeat before native
runtime selection. Stale preparation then created an observed native
coordinator on the cancelled run. The coordinator had no provider result
or eligible recovery worker. A later Continue message stayed queued
indefinitely.

**Expected behavior**

Cancellation fences native startup. After verified cleanup, a newer user
message starts one fresh conversation turn. An unverified execution
keeps its hold and a clear explanation.

**Steps to reproduce**

1. Start a task with the native runner and delay startup preparation
before runtime selection.
2. Pause the task, then release preparation.
3. Resume the task and send Continue.
4. Before this fix, native selection can persist after cancellation and
block the saved message.
5. Repeat from persisted cancelled startup state after a server restart.

**Paperclip version or commit**

Reproduced against master at `586b5ec82` with isolated PostgreSQL
regression fixtures.

**Deployment mode**

Self-hosted server built from source, with Paperclip Runner.

## What Changed

- Serialize the cancellation fence and native runtime selection on the
run row. Revalidate the startup controller lease.
- Refresh the runtime before dispatching cancellation, and reject
terminal or cancelled runs at the native provider claim.
- Recognize never-claimed coordinators only after startup and
environment cleanup are verified. Reject process, provider, owner, and
conflicting launch evidence.
- Settle that coordinator atomically with a new authenticated user turn.
Preserve history, unknown outcomes, and attempt counts.
- Reuse the saved-message worker after restart and retain pause, budget,
approval, and ownership gates.
- Add startup, restart, duplicate-admission, and negative-proof
regressions. Document the rule.

## Verification

- All 687 tests pass across the complete heartbeat recovery, explicit
continuation, and native session executor suites on the rebased branch.
- Six focused race regressions also pass: cancellation before and after
native selection, Stop racing adapter registration, process termination
during a database failure, and a run finishing during cancellation.
- `pnpm -r typecheck` and `pnpm build` pass after rebasing on master at
`ab15aff39`.
- Greptile is 5/5 on `f3dea2ab27facdf0360b56172ebd3e5219e25538`, with no
open review threads. Policy and security checks pass.
- All 32 CI checks pass on the final head, including the full
server/workspace test matrix, all three browser shards, runner
verification, typechecks, build, and canary dry run. The two optional
Storybook jobs are skipped. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/34697945514).
- Local full-suite limitation: the broad `pnpm test:run` attempt
reported two failures outside the changed area after unusually long test
durations (about 65 seconds for supporting skill-file saves and 933
seconds for setup-token login). Both cases passed isolated reruns, with
no code changes. The broad local run was stopped after CI completed
successfully; no clean full local-suite pass is claimed.

## Risks

Cancellation and startup overlap. The run and coordinator locks provide
the authority fence; cleanup and process evidence provide the
containment proof. Historical runs without sufficient evidence remain
blocked. A saved user message authorizes a fresh turn, not automatic
replay. No schema migration or dependency change.

## Model Used

OpenAI GPT-6 through Codex, using reasoning, repository inspection, code
execution, and tests. This session does not expose the exact backend
revision or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 11:39:36 -05:00
DottaandPaperclip f12b647ae8 fix: reliably interrupt and resume legacy message queues (#13275)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A task can collect more messages while its agent works.
> - Legacy runners must stop the active process before they can receive
those messages.
> - The old Interrupt action cancelled the run but could leave the queue
idle and hidden.
> - Codex could also classify a cancelled run as successful or start a
fresh process after cancellation.
> - This pull request joins cancellation, preserves the provider
session, and dispatches the current queue after cleanup.
> - The benefit is reliable interruption with the saved message order,
edits, and deletions.

## Linked Issues or Issue Description

**What happened?**

Interrupt could strand a legacy message queue. The UI could hide pending
messages after the run stopped. A Codex signal exit could race the
cancellation write. A stale session warning could also trigger a fresh
process after an interrupted resume.

**Expected behavior**

Interrupt stops the active turn and sends the remaining messages once,
in their saved order. Deleted messages stay deleted. An interrupted
Codex turn keeps its session and does not restart itself.

**Steps to reproduce**

1. Assign a task to a legacy Codex agent that runs a long command.
2. Queue three messages. Edit one, discard another, and move the last
message first.
3. Click Interrupt in the queue.
4. Repeat the interruption while the resumed session runs another
command.

Related work: Refs #13160, which moves native queue steering into the
wake-queue module. This change fixes legacy interruption and keeps
native steering unchanged.

## What Changed

- Add a revision-checked, company-scoped endpoint for legacy queue
interruption.
- Promote only the requested queue after the provider stops and releases
its lease. Retry its persisted interrupt intent from the scheduler after
a promotion error or server restart.
- Keep pending legacy queues visible after a run stops. Use server state
for the interrupt result.
- Serialize owned process cancellation before classifying the adapter
result. Preserve late session and log metadata. Acknowledge cancellation
only when an actual process or process group was owned; scheduler
placeholders retain their normal release policy.
- Send Ctrl-C to legacy Codex. Prevent missing-session fallback once the
session has started.
- Add cancellation race, multi-actor queue order, durable retry, resume
fallback, and stale request regression tests. Document the behavior.

## Verification

- Real browser tests passed with legacy Codex CLI and ACP engines, using
Codex 0.153.4 and gpt-5.6-sol.
- All three automated ACP browser scenarios passed locally: immediate
Interrupt delivery, no replay of an unfinished write, and pause
requiring Resume. Updated the old test expectation that required a
separate “go” after Interrupt.
- Browser tests covered queued edits, deletion, reordering, deleting the
final message, and repeated interruption.
- Two consecutive CLI interrupts kept one provider session. Both stopped
processes exited. The final message arrived once.
- `pnpm -r typecheck` passed.
- `pnpm check:token-gates` passed.
- All 346 post-review scheduling, recovery, queue-route,
archived-company, worktree-suppression, and stale-queue regression tests
passed.
- All 318 process-recovery and durable-chat tests passed after the final
cancellation guard.
- Codex adapter, queue UI, issue-page, and OpenAPI contract tests
passed.
- `pnpm build` passed.
- Full local suite coverage completed with
`PAPERCLIP_IN_WORKTREE=false`, using the stable runner and its CI
shards: 618 general server suites, all 145 serialized server suites, and
all workspace groups. Every failing suite passed a targeted rerun after
the fixes, rebuilding the native test fixture, correcting macOS
temporary-path setup, or retrying setup/timing failures. Existing skips
remain.
- The original monolithic run reported failures before the final fixes;
its failed suites were rerun rather than rerunning all 618 suites again.
The final process-recovery/durable-chat regression run passed all 318
tests.
- All CI checks passed for `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`:
[run 34654820774, attempt
2](https://github.com/paperclipai/paperclip/actions/runs/34654820774/attempts/2),
including typecheck, build, all test shards, E2E, and canary. The
signoff and Cursor sandbox tests each hit a timeout in the initial
attempt; both suites passed locally, and both failed shards passed their
single CI rerun. All three corrected ACP browser scenarios passed in CI.
- Greptile reviewed `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: 5/5, no
open review threads.

## Risks

Cancellation order affects local adapters. The tests cover signal exits,
graceful exits, adapter exceptions, termination errors, and cancellation
write errors. Embedded adapters keep their cancellation controls.
Ordinary run cancellation and task pause keep their distinct queue
policies. No database migration is required.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, browser testing, and code
execution. The exact serving model ID and context-window size are not
exposed in this session. The live test runner used OpenAI gpt-5.6-sol.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 18:26:59 -05:00
DottaandPaperclip a38ccf9972 fix: retry transient continuation admission locks (#13290)
## Thinking Path

> - Paperclip manages AI agents and the tasks that they execute.
> - The run scheduler checks continuation authority before it starts a
provider.
> - This check uses database locks to order execution against
conversation closure.
> - A short lock conflict could fail a valid user follow-up before the
provider started.
> - This pull request retries the admission transaction after the locks
are released.
> - Valid work can start after normal contention, while closure and
cancellation still stop execution.

## Linked Issues or Issue Description

Refs #13038. Related continuation work: #13270 and #13239.

**What happened?**

A user comment started a run through the automation queue. Its source
records and admission marker were valid. A database lock conflict at
dispatch caused `chat_control_recovery_proof_unresolved` and stopped
automatic recovery. The provider received no work.

**Expected behavior**

Retry short database lock conflicts before failing admission. Read
current ownership and conversation-close evidence on each attempt. Do
not retry provider execution.

**Steps to reproduce**

1. Queue a user follow-up through the automation transport.
2. Hold the task row lock in a separate transaction at the dispatch
boundary.
3. Release the lock after 250 ms.
4. Before this fix, the run fails before provider dispatch. With this
fix, the run passes admission once the lock is released.

**Paperclip version or commit**

Reproduced on base commit `1c4bcff2b`. Disabling the new retry
reproduces the original error in the regression test.

**Deployment mode**

Self-hosted server with PostgreSQL. Regression tests use embedded
PostgreSQL.

## What Changed

- Retry rolled-back admission transactions after lock conflicts, with up
to 50 waits of 100 ms.
- Keep queue claims nonblocking. Keep provider dispatch outside the
retried transaction.
- Recheck current run state and committed close evidence after every
conflict.
- Explain persistent database contention in the exhausted admission
error.
- Add real database contention tests and bounded retry tests. Document
the behavior.

## Verification

- `pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/services/chat-control-admission-retry.test.ts`: 273 passed.
This includes task, wake, and run locks, the native runner, close/cancel
races, unrelated failures, and retry exhaustion.
- Regression proof: disabling retries makes the user-follow-up test fail
with `chat_control_recovery_proof_unresolved`.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: stopped after all equivalent CI server/workspace
shards passed. The local run exposed a missing `fake-codex-app-server`
fixture binary in the fresh worktree; after `pnpm --filter
@paperclipai/paperclip-runner run build:rust`, the complete affected
`native-session-resume.test.ts` suite passes (37 tests).
- Greptile: 5/5, no findings, on commit `c974a496a`.
- CI: all 31 checks passed on commit `c974a496a`, including all
server/workspace test shards, browser tests, typecheck, runner
verification, build, and the release dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/34654770074).

## Risks

- A contended dispatch can wait through 50 short delays, plus
transaction time.
- Persistent contention still fails closed after the retry budget.
Invalid source evidence fails without retrying admission.
- No schema, permission, provider retry budget, or queue-claim behavior
changes.

## Model Used

OpenAI Codex, GPT-6. The exact deployed model ID and context-window size
are not exposed in this session. Used reasoning, repository inspection,
code editing, and command execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 18:11:44 -05:00
DottaandPaperclip 9031516a7e fix: recover legacy Daytona startup failures from task and inbox (#13272)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Legacy conversation adapters can run in Daytona sandboxes.
> - A server restart during provisioning can occur before the invocation
event exists.
> - Recovery then lacks the old adapter identity and leaves a hold that
ordinary user retries cannot clear.
> - A remote launch can also fail when its host relay looks for Node in
the sandbox PATH.
> - This pull request records the adapter at claim time and restores
explicit user continuation after verified cleanup.
> - Users can recover from the task or inbox while the failed run and
uncertain action history remain intact.

## Linked Issues or Issue Description

Refs #13237, #13239, #13254. Those changes cover recorded conversation
runs, native user continuation, and explicit remote Stop. This change
covers legacy failure before `adapter.invoke` and exact task/inbox
Retry.

Refs #9771 for overlapping generated-command quoting. This change also
supplies the absolute host Node executable. Refs #13163 and #13264 for
the separate native restart and retained-workspace work.

**What happened?**
A legacy Daytona run interrupted during provisioning became
`process_lost` without an invocation event. Recovery preserved an
execution hold, and Retry or a new task reply could not resume it.
Cleanup could also run before the Daytona plugin was ready. On a macOS
host, a subsequent ACP relay launch failed with `env: node: No such file
or directory` because the remote launch environment did not contain the
host Node path.

**Expected behavior**
An interrupted conversation can continue after its previous execution
stops. Explicit Retry and new user replies should start a fresh turn
with the task history. Cleanup failures must remain visible and
recoverable. The host relay must use the host Node executable.

**Steps to reproduce**
1. Use a legacy Claude adapter with a Daytona environment.
2. Interrupt the server after it acquires the sandbox lease and before
it records `adapter.invoke`.
3. Restart and inspect the task hold.
4. Retry from the task or inbox, or send a new task reply.
5. Confirm the old sandbox has stopped and one new response arrives.

**Paperclip version or commit**
Reproduced from master at `3bafac12f796fbea02e609e1074a9639f872e9c4`.
The branch is rebased on `51b0e01ea`, including #13261 and #13270.

**Deployment mode**
Built from source on macOS with a real Daytona sandbox and the legacy
Claude ACP adapter.

## What Changed

- Count new browser specs with the scheduler's median duration in the
shard-balance check. This fixes a false policy failure after new specs
arrive from both branches. The balance threshold is unchanged.

- Persist server-owned adapter identity in the queued-to-running claim
before provisioning starts.
- Wait for provider plugin startup before restart cleanup. Keep failed
cleanup leases as active ownership blockers.
- Admit exact board retries and new user comments after verified
termination. Retain the old run, task history, approvals, and unknown
action outcomes.
- Adopt repeated Retry requests. Permit one scoped cleanup attempt per
explicit user Retry after the automatic limit, with an activity record.
A later user Retry can recover after a transient provider failure;
automatic attempts remain capped.
- Resume replies deferred during cleanup, including historical legacy
startup failures.
- Launch the host ACP relay through the absolute host Node executable.
- Add a task-level Retry button and return actionable blockers when
retry admission is refused.
- Add database regressions and three browser recovery journeys. Exclude
installed third-party dependency skills from the shipped-skill audit.

## Verification

- Current head: `d23c84181`, rebased on `51b0e01ea`. Conflict resolution
retains the saved-message recovery, local stop receipts, and wait
reasons from #13270 alongside exact legacy Retry support.
- Real Daytona: interrupted the server after lease acquisition and
before adapter invocation. Restart cleanup confirmed provider
termination. Task Retry cleared a seeded historical hold and a real
Claude agent returned `Recovery verified.` in the task. Removed the
disposable sandbox and environment after testing.
- All three browser recovery journeys passed again after the final
rebase. Task Retry, Inbox Retry, and a new reply each produced one fresh
successor, completed the task, preserved the failed run, and retained
the answer after reload.
- All 29 e2e/server shard-partition tests passed. The balance check now
uses the scheduler's median fallback for unmeasured specs, with the same
balance threshold.
- Server typecheck passed after rebuilding the generated runner
dependencies. The combined recovery/route run passed 136 of 137 tests.
Its remaining route test timed out during the first cold module import
at its explicit 10-second limit; an isolated rerun reproduced that
timeout and passed the other 51 route cases. The complete CI suite
passed on this head. The same route file passed all 52 cases in CI,
including the first cold import in 7.5 seconds.
- Before the final rebase, recursive typecheck, full build, UI token
gates, 132 targeted server tests, and the complete [CI
workflow](https://github.com/paperclipai/paperclip/actions/runs/34650004085)
passed. The subsequent CI failure was the shard-balance accounting
mismatch fixed here.
- Greptile reviewed `d23c84181` at 5/5 with no outstanding actionable
findings. The complete [current CI
workflow](https://github.com/paperclipai/paperclip/actions/runs/34653327949)
passed on attempt 2. All test, typecheck, build, and canary jobs passed
on the first attempt. Docker setup timed out fetching BuildKit from
Docker Hub; retrying that job and its dependent aggregate succeeded.

## Risks

- Recovery admission changes executable authority. Company, task, agent,
user, approvals, process ownership, and provider termination checks
remain required.
- Explicit continuation starts a fresh conversation with history. It
does not certify unknown external action outcomes or rerun
non-conversation adapters automatically.
- Changing task status alone does not clear an execution hold. The task
now offers an explicit Retry action.
- Historical adapter claims and invocation events take precedence over
current agent settings. Known process or webhook runs retain their hold.
Pre-upgrade rows with no adapter evidence may receive only a new
explicit user turn after termination proof; they do not become eligible
for automatic replay.
- No schema migration or sandbox-image change is required. This branch
has not been deployed to production.

## Model Used

OpenAI GPT-6 through Codex, with repository inspection, code execution,
browser automation, and test execution. The exact deployment model ID
and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 18:08:14 -05:00
DottaandPaperclip 51b0e01ead fix: resume saved user messages after execution recovery (#13270)
Preserve verified native process-stop evidence and retry saved user messages through normal continuation admission after recovery cleanup. Show the current wait reason and serialize delivery so a saved message starts one fresh turn.

Validated with 410 focused tests, typecheck, build, token gates, all PR CI checks, and Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-11 17:12:55 -05:00
DottaandPaperclip 3bafac12f7 refactor: remove automatic productivity reviews (#13263)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its recovery loop keeps assigned work moving after execution
failures.
> - Productivity review used run counts, comment counts, and elapsed
time to create management tasks.
> - Infrastructure failures could satisfy those rules and create more
tasks without evidence that the source work needed management review.
> - This pull request removes that detector and its continuation holds.
> - Bounded recovery, budgets, explicit blockers, and normal review
stages remain in place.
> - Existing task records stay readable and unchanged.

## Linked Issues or Issue Description

Refs #5897. That request describes unwanted automatic productivity
reviews and asks to preserve existing tasks. This change retires the
feature instead of adding another configuration switch.

Related prior approaches: Refs #9191, Refs #12489. Those changes
excluded infrastructure failures or bounded review creation. This
removal replaces the detector rather than tuning its thresholds.

## What Changed

- Delete the scheduled detector, automatic task creation, evidence
refresh, and productivity continuation holds.
- Remove computed productivity fields, special attention items, badges,
and Storybook fixtures.
- Retain historical origin values, decision compatibility, and recovery
recursion exclusions. Add no migration and change no existing task data.
- Update the execution contract. Replace feature tests with regressions
for legacy task reads, ordinary attention, and bounded continuation in
the presence of an old review.

## Verification

- Targeted attention, issue-route, startup, and UI tests: 4 files and
101 tests passed.
- Updated issue-route and UI tests: 2 files and 61 tests passed.
- Bounded continuation regression: 2 cases passed, including a legacy
review plus pre-dispatch cancellation churn.
- `pnpm check:token-gates`: all four gates passed.
- `git diff --check`: passed.
- `pnpm build-storybook`: passed.
- Greptile: 5/5 on `a5a612eea`, with no actionable findings.
- Scheduler and historical recovery regressions: 2 files and 28 tests
passed.
- Repository `pnpm -r typecheck` and `pnpm build`: passed.
- The complete `pnpm test:run` suite passed across the CI server,
serialized-server, and workspace shards on `a5a612eea`. Stopped the
duplicate local monolithic run after the full CI suite passed; no
completed local full-suite result is claimed. The targeted local suites
above passed.
- CI serialized shard 5 initially hit a 10-second timeout in the first
interaction-route test. The complete file passed locally (78 tests),
then the single CI rerun passed.
- All CI gates are green, including the build and end-to-end suites.
- A local merge check against current `master` (`ce09ea40b`) completed
without conflicts.

## Risks

- API responses no longer include the computed `productivityReview`
field. Consumers must stop using it.
- The scheduler no longer creates management work from elapsed time, run
counts, or missing comments. This is the intended behavior change.
- Existing review tasks and explicit dependencies remain in place.
Historical origins still prevent recursive recovery treatment. No task
cleanup or data migration occurs.
- The native review handoff repair is separate from this removal.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact runtime model identifier and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 15:46:35 -05:00
DottaandPaperclip 663c44cb2b fix: continue conversations after confirmed remote runner stop (#13254)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can stop a run and send another message on the same task.
> - Remote runners need evidence from their sandbox provider that
execution stopped.
> - Local process checks cannot prove that a remote process exited.
> - This pull request records provider stop receipts and uses them for
conversation admission.
> - New user messages can proceed after confirmed cleanup without
repeating interrupted actions.

## Linked Issues or Issue Description

Refs #13237 and #13239. Related: #13163 covers app-restart recovery;
this change covers an explicit stop followed by a new user message.

**What happened?**

A stopped remote Claude ACP task kept its execution hold after Daytona
cleanup succeeded. Native runners also rejected remote process
identities and retained stale session cleanup gates. A message sent
during cleanup could stay deferred after the sandbox stopped.

**Expected behavior**

After the provider confirms that the old execution stopped, a new user
message starts a fresh turn. Pending user messages must not need another
message to trigger admission. Prior action outcomes remain recorded.

**Steps to reproduce**

1. Start a long-running task in Daytona with a legacy Claude ACP or
native ACP runner.
2. Cancel the run while its tool is active.
3. Send a new message immediately, or after cleanup completes.
4. Observe the execution hold despite the old sandbox having stopped.

**Paperclip version or commit**

Reproduced on master at 7b829efdf6. The
branch is rebased on current master.

## What Changed

- Add optional provider stop receipts to sandbox release and destroy
hooks. Old plugins remain compatible.
- Bind receipts to the company, run, lease, and provider resource.
Failed cleanup cannot supply stop authority.
- Acknowledge legacy remote cancellation after confirmed termination.
- Admit native user continuations using remote receipts instead of host
process checks.
- Retire only the settled cleanup owner matching the stopped company,
run, and provider resource. Isolate cleanup gates between remote
sandboxes, including two sandboxes owned by one run.
- Reconsider user messages deferred during remote cleanup through normal
admission, including successful later cleanup retries.
- Preserve receipts through cleanup retries and inline cleanup after
failed startup.
- Permit provider destruction after a terminal remote checkpoint
failure. Busy ownership still blocks destruction.
- Add regression tests and update execution semantics.
- Give the real preview fixture ten seconds for cold startup. Reuse
release-mode Rust artifacts for the filtered parity checks, avoiding a
duplicate debug test build that exhausted CI disk twice; test filters
and assertions are unchanged.

## Verification

- Recursive typecheck and build passed.
- Targeted server, Daytona plugin, and native runtime tests passed. They
cover missing or mismatched receipts, failed cleanup, local process
protection, exact cleanup ownership, and a message sent during cleanup.
- Live Daytona tests passed for legacy Claude ACP, native per-turn,
native warm, and a newly created native runner using disposable
sandboxes. Each original run was cancelled; its explicit follow-up
completed with no execution hold. The disposable case confirmed a new
sandbox after deletion.
- Native tests used the provider's Opus 5 selector, `opus[1m]`, and the
Linux runner bundle from the sandbox image. Existing checkpoint/sync
finalization warnings remained visible before the native follow-ups
reached committed success. This change does not repair those separate
warnings or guarantee recovery of uncheckpointed files.
- A direct live provider test also passed with the final delete-wait
change: the destroy hook returned its receipt only after Daytona
reported the sandbox destroyed. The final Daytona plugin suite passed
all 153 tests; its build passed.
- After rebasing on master, 284 targeted server/plugin tests and 313
native executor tests passed. The native session runtime suite passed
all 128 tests.
- The review regression passed all 145 tests across the continuation,
environment runtime, and pending-cleanup sweep suites. Recursive
typecheck and build passed again after that fix.
- The security review's exact-resource finding is fixed. Cleanup
completion requires the same provider-resource scope used during session
creation. All 128 native runtime tests, 61 server continuation/cleanup
tests, recursive typecheck, and build passed after this fix. The
two-sandboxes-in-one-run regression proves one receipt cannot retire the
other quarantine.
- Greptile reviewed the final commit at 5/5, the security scan passed,
and no review threads remain unresolved. The final native executor suite
also passed all 313 tests.
- The full local suite reached 10,616 passing tests before stopping on
four failures. All four now pass in focused reruns: the final-code
continuation/teardown suites, the attachment suite, and the native
session test after building its required fake-provider binary. The
original full run overlapped source edits and did not reach the
remaining groups; it is not counted as a full-suite pass.
- The preview exposure suite passed all 25 applicable tests (three
Linux-only tests skipped locally) after the startup allowance change.
Both release-mode Rust parity commands passed locally. All final-head PR
checks passed, including full server/workspace/browser test groups, full
native runner verification, build, typecheck, canary dry-run, and the
aggregate gates.

## Risks

- A provider must return a receipt only after confirmed termination.
Incorrect provider claims could permit overlapping execution.
- Older providers without receipts retain the existing hold. Missing
evidence, failed cleanup, active ownership, pauses, approvals, and
budgets still block admission.
- This change preserves unknown action outcomes and old checkpoints. It
does not authorize replay or alter historical runs.
- No database migration or telemetry contract change is required.

## Model Used

OpenAI GPT-6 through Codex, with repository inspection, code execution,
and browser tools. The exact backend revision and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run targeted tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 15:18:37 -05:00
DottaandPaperclip 2904a3a6cc fix(ui): hide retry countdown after execution starts (#13258)
## Thinking Path

> - Paperclip helps operators manage AI agents and their tasks.
> - Task pages show countdowns for deferred checks and automatic
retries.
> - A retry keeps its scheduled start time after it enters the queue or
starts running.
> - The countdown treated that historical time as a pending deadline and
showed an overdue warning beside active work.
> - This pull request limits retry countdowns to retries that are still
scheduled and hides waiting surfaces on terminal tasks.
> - Operators now see a warning only when the displayed retry is still
waiting to start.

## Linked Issues or Issue Description

Refs #9783, which added the monitor surfaces. Searched related PRs and
issues; no duplicate fix was found.

**What happened?**

After a service restart resumed a task through an automatic retry, the
task showed an overdue retry banner while the agent was running. The
banner also remained when the task became done.

**Expected behavior**

A queued or running retry must not show a countdown against its past
scheduled start time. Done and cancelled tasks must not show waiting
banners.

**Steps to reproduce**

1. Open a task with an automatic retry scheduled for a known time.
2. Let the retry enter the queue and start running.
3. Wait until its scheduled time is more than one minute in the past.
4. Observe the overdue banner and Check now button while the agent is
working.

**Paperclip version or commit**

Reproduced on source commit 847d00bdc3.

**Deployment mode**

Self-hosted server built from source. This is a core UI bug.

## What Changed

- Derive a retry countdown only when the retry status is
`scheduled_retry`.
- Ignore retained retry times for queued, running, and cancelled
retries.
- Hide waiting banners for done and cancelled tasks.
- Keep a separate scheduled monitor visible on an open task.
- Add state and rendered-transition regression tests. Document the
display rule.
- Stub the process start time in one restart-recovery test. CI exposed
that the fixture read real host metadata for its fake PID.

## Verification

- Focused monitor tests: 34 passed on the rebased commit.
- Token gates passed on the rebased commit.
- `pnpm -r typecheck` and `pnpm build` passed on the rebased commit. The
full local test run was stopped after embedded PostgreSQL failed to load
missing macOS library aliases. After repairing the local dependency
aliases, the 11-test native status corpus passed. CI exposed an
unrelated restart-recovery fixture that read real host process metadata.
A test-only fix removes that host dependency. All 31 CI checks passed on
the final commit; the two optional Storybook jobs were skipped.
- The runner transport test that timed out in the first CI run passed
locally: 2 tests passed. The final CI Build job passed.
- Restart-recovery fixture suite: 20 tests passed after the test-only
change.
- Greptile reviewed final commit
`b815ac97d36cfcb6d09f31c852465b6474f78fea`: 5/5, with no open review
threads.
- The rendering regression checks waiting, queued, running, rescheduled,
and done states. It also checks that hidden surfaces remove their
buttons and countdown timers.

## Risks

Low risk. This changes display state only. It does not change retry
dispatch, monitor scheduling, database records, or API contracts. A
retry that is still scheduled retains its overdue warning.

## Model Used

OpenAI GPT-6 through Codex. Used reasoning, repository inspection, code
editing, browser inspection, and test execution. The exact runtime model
identifier and context window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 15:05:07 -05:00
DottaandPaperclip 5545f6d166 feat: let user messages continue stopped native tasks (#13239)
## Thinking Path

> - Paperclip manages AI agents and their tasks.
> - A failed native run can leave a durable execution hold.
> - The hold prevents automatic replay of actions with unknown outcomes.
> - It can also prevent the agent from answering a new user message.
> - A new user message should authorize a fresh turn after the prior
execution stops.
> - This pull request adds that admission path and preserves the
existing execution gates.
> - Users can continue the conversation without certifying every past
action.

## Linked Issues or Issue Description

**Subsystem affected**

Server task wake admission and native execution recovery.

**Problem or motivation**

A native task can remain blocked after automatic recovery stops. A new
user message is saved, but its run is cancelled before the agent can
answer.

**Proposed solution**

Use a new authenticated user comment to authorize a fresh turn. Check
stopped predecessor ownership and available history. Retain uncertain
action outcomes. Commit the new run and hold retirement together.

**Roadmap alignment**

This is a focused improvement to the existing self-healing runs and
recovery behavior.

Builds on merged #13237, which covers legacy conversation continuation.
This PR adds native admission and preserves native automatic-recovery
eligibility and budgets.

## What Changed

- Admit a fresh native turn for a new user comment after every held
predecessor has stopped.
- Validate the comment author, task, timing, process ownership,
controller, and cleanup leases.
- Preserve failed runs, unknown action outcomes, and the failed
incident's attempt count.
- Record the new comment and run in the existing recovery audit history.
- Validate the saved continuation source and discard consumed user-wake
authority from later automatic replacements.
- Keep pause, approval, budget, ownership, and dependency interaction
rules.
- Add database and actual wake-path regressions. Update the execution
contract.

## Verification

- [Full CI run
34626750213](https://github.com/paperclipai/paperclip/actions/runs/34626750213)
passed on `c58e6c087e0df7530c747d80b27d491da925a9c4`: all 31 reported
checks passed, including all server/workspace suites, browser shards,
native runner verification, build, typecheck, release dry run, and
aggregate gates. The two conditional Storybook checks were skipped.
- Greptile reviewed this exact head at 5/5. All review threads are
resolved, and security checks passed.
- All 218 local targeted tests passed across explicit native
continuation, continuation history, safe replacement, durable chat
wakeups, wake queue, issue liveness, native session resume, and run
dispatch. The native implementation is unchanged by the final rebase
onto master.
- Full workspace `pnpm -r typecheck` and `pnpm build` passed on the
final head. Complete test coverage is supplied by the green CI suites;
local tests used the targeted suites above.
- Regressions cover scoped authorization, concurrent delivery, live
ownership, later admission gates, retained message receipts, and
automatic replacement after terminal-task or reviewer changes.

## Risks

- A fresh model turn can choose to repeat an action. Paperclip preserves
prior history and does not replay recorded calls.
- Missing process identity and remote ownership without a target-aware
stop proof retain the hold. A terminal database row alone does not prove
that execution stopped.
- No schema or dependency changes. Existing historical tasks are not
awakened by deployment.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and test execution. This session does not expose a more
specific model build ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 12:33:57 -05:00
DottaandPaperclip b1efd65edc fix: continue interrupted task conversations with bounded retries (#13237)
## Thinking Path

> - Paperclip manages AI agents and their tasks.
> - A task can outlive a provider process or a server restart.
> - Legacy recovery treated unknown tool outcomes as a permanent
execution hold.
> - That hold could also reject a later user message.
> - A conversation turn can use prior history without replaying prior
tool calls.
> - This pull request lets supported conversation adapters continue
within the existing retry budget.
> - Users can send a new message after automatic attempts stop.

## Linked Issues or Issue Description

**What happened?**

A server restart could interrupt a local ACP run and leave its task
behind a permanent recovery hold. A later user message could be
cancelled before the provider answered. The immediate recovery path
could also create a successor outside the durable failure counter.

**Expected behavior**

Continue with a bounded new conversation turn. Preserve a compatible
provider session or use full task context when it is unavailable. Do not
replay recorded tools. When automatic attempts stop, allow a new user
request through the normal execution gates.

**Steps to reproduce**

1. Start a task with a local conversation adapter.
2. Restart the server while the provider is working.
3. Let the previous run become interrupted.
4. Send a follow-up message and observe the recovery hold on the old
behavior.

Related work: Refs #13075 for durable task recovery. Refs #12946 for
retry-limit and checkout-lock handling. This change routes conversation
recovery through the existing bounded scheduler.

## What Changed

- Mark supported local conversation failures for continuation. Keep
native-runner and non-conversation recovery rules.
- Carry an interruption notice into the next turn. Retain stopped ACP
session history even when a write outcome is unknown.
- Clear unavailable ACP sessions so the next bounded attempt can use
full task context.
- Route immediate failure recovery through the same durable scheduler as
process-loss recovery. Release only the predecessor checkout when its
retry takes ownership.
- Retire obsolete conversation holds using immutable run evidence, in
bounded batches with an activity record. Preserve outcome evidence and
do not wake historical tasks.
- Block actual admission and Resume while a predecessor process or
environment lease is still active. Keep the original interruption notice
after a rejected wake. Preserve the upstream blocked-wake waiting
contract: bounded retry planning can happen during cleanup, while
deferred messages and execution remain gated.
- Add subprocess and database regression tests. Update the execution
contract.
- Add the current thread-status field to the native recovery provider
fixture so its damaged-journal test reaches the intended boundary.
Tolerate an already-exited fixture process during test cleanup while
still asserting both processes terminate.

## Verification

- Workspace typecheck passed: `pnpm -r typecheck`.
- Build passed: `pnpm build`.
- Module boundaries passed: `pnpm check:module-boundaries`.
- Focused tests passed: 293 recovery/session/dispatch tests, 66 retry
and response-gate tests, and 37 native-session tests. Some suites
overlap.
- Tests cover interrupted writes, missing sessions, concurrent retries,
restart persistence, pending questions and approvals, execution gates,
and historical holds.
- Built the Rust test executables with `pnpm --filter
@paperclipai/paperclip-runner build:rust` for native-runner
verification.
- Full Vitest coverage verified locally using the repository’s general
and serialized shards, with focused reruns for failures and files not
reached after a shard stopped. The ownership-gate regression is fixed
and the complete affected server shard passes (1,390 tests). Local
parallel runs also hit temporary-directory, resource, and timing
failures; those suites pass with canonical temporary paths and
sequential reruns. No test timeouts were increased.
- Final merged-branch regression run: 577 tests pass across process
recovery, retry scheduling, liveness, durable chat, wake-queue
application/adapter, dispatch, continuation, native sessions, and task
chat. Earlier focused verification also passed 19 native control tests.
Token gates and whitespace validation pass.
- Browser verification passed all three ACP Stop/continue/pause
scenarios, including a rerun after merging the upstream waiting
behavior: `PAPERCLIP_E2E_PORT=3397 pnpm test:e2e
tests/e2e/acp-stop-continuation.spec.ts`. The interrupted-write case
verifies that follow-up completes without a repeated write.

- Final-head [CI run
34625037394](https://github.com/paperclipai/paperclip/actions/runs/34625037394)
passed on `06ac4bd9d150f8b209a96e5fd609c696958794a0`: all 31 reported
checks are green, including server/workspace suites, all browser shards,
native runner verification, build, typecheck, release dry run, and
aggregate gates. The two conditional Storybook checks were skipped.
Greptile reviewed this exact commit at 5/5; all review threads are
resolved.

## Risks

- A new model turn can choose to repeat an action. Paperclip does not
replay recorded tool calls and does not certify unknown action outcomes.
- Conversation adapters now stop after their retry budget instead of
requiring action reconciliation. Explicit Stop, pause, dependency,
approval, budget, and ownership gates remain in force.
- No schema migration or dependency changes. Historical holds are folded
without changing task status or waking work.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and test execution. The session does not expose a more
specific model build ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 12:16:04 -05:00
DottaandPaperclip eb640ec129 fix(execution): keep blocked wakes waiting without repeated runs (#13236)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Wake admission decides when a task can create an execution run.
> - Recovery can prohibit replay while the previous execution needs
review.
> - Dependency reconciliation kept creating runs before dispatch
rejected that same hold.
> - Each rejected run added another startup notice without doing useful
work.
> - This change checks the hold during admission and records repeated
automatic waits once.
> - Tasks keep their messages and can resume when the current gates
permit execution.

## Linked Issues or Issue Description

Related changes: Refs #13173 (stale completed-task continuations). Refs
#12651 (dependency waits during recovery).

**What happened?**

A blocked task with completed dependencies can remain under a durable
execution reconciliation hold. Each scheduler pass created a queued run.
Dispatch then cancelled it before the adapter started. The skipped wake
did not satisfy dependency wake deduplication, so this repeated and
filled the conversation with “Couldn't start” notices.

**Expected behavior**

A known execution hold creates a waiting diagnostic without a run.
Repeated automatic observations share that diagnostic. Clearing the hold
permits a new wake only after the other gates pass. New comments remain
available for the next eligible execution.

**Steps to reproduce**

1. Assign a blocked task with a completed blocker.
2. Give the task an active reconciliation action, or a resolved action
whose automatic recovery evidence still prohibits replay.
3. Run dependency reconciliation repeatedly.
4. Observe repeated cancelled pre-start runs on the base branch. This
branch creates no runs while held and admits work after the effective
hold clears.

## What Changed

- Check effective execution holds under the issue admission lock before
inserting runs. Keep the final dispatch check for races.
- Share automatic wait diagnostics across producers, wake keys, and
service restarts. Apply the helper to reconciliation, dependencies,
pause holds, availability, budgets, and disabled heartbeats.
- Preserve ordinary comment and interaction receipts during execution
holds. Prevent release from draining them while replay is blocked. Keep
external-chat receipt authorization intact.
- Group empty pre-start reconciliation cancellations into a neutral
waiting notice. Keep started runs and the full run history.
- Document the waiting contract and add database-backed, UI, and browser
regressions.
- Stabilize two existing verification tests: allow the asynchronous chat
lease transition a bounded five-second wait, and accept either
legitimate damaged-session refusal while retaining exact
archive-evidence assertions.

## Verification

Passed targeted tests:

- `pnpm exec vitest run
server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts
server/src/modules/wake-queue/adapters/postgres.test.ts` — 28 tests.
- `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx` —
covered in the initial combined test run; UI suite passed.
- Run-dispatch adapter tests passed in the combined gate regression run.
- `pnpm exec vitest run
server/src/__tests__/durable-chat-wakeup.test.ts` — 41 tests, including
held receipt replay, promotion, and revoked access.
- `PAPERCLIP_E2E_PORT=3294 pnpm test:e2e
tests/e2e/acp-stop-continuation.spec.ts` — all 3 browser scenarios pass.
Repeated held messages create no additional runs or provider prompts and
do not replay writes.
- `pnpm check:token-gates`
- `pnpm check:module-boundaries`
- `git diff --check`

`pnpm -r typecheck` and `pnpm build` pass.

The full local `pnpm test:run` invocation did not finish green: it
encountered exhausted local PostgreSQL shared-memory slots, a missing
fresh-worktree runner test binary, and tests loaded across in-flight
edits. The affected chat/database suites passed on rerun (81 tests), and
the targeted lifecycle/recovery verification passed (3 tests). After
building the runner test binary, the full native session suite also
passed (37 tests). Final-head [CI run
34621288475](https://github.com/paperclipai/paperclip/actions/runs/34621288475)
passed on `8659618b0ed2b98df002a28f4c1bd97321b0db04`, including all
server/workspace test shards, all three browser shards, runner
verification, typecheck, build, release dry run, and the aggregate
verification gates. All 31 reported checks passed; the two conditional
Storybook checks were skipped as intended. Greptile reviewed that exact
commit at 5/5 with no unresolved review threads.

## Risks

The wait record is diagnostic only. It must never count as a delivered
wake or bypass a current gate. Tests cover repeated and concurrent
admission, resolved no-replay evidence, a remaining dependency after
hold clearance, deferred comments, and release gating. Explicit user
requests and authorized chat receipts do not share automatic
diagnostics. No migration or historical data deletion is required.
Existing provider retry budgets remain unchanged.

## Model Used

OpenAI GPT-6 through Codex. The runtime does not expose the exact hosted
snapshot ID or context-window size. Used repository inspection, code
editing, command execution, tests, and review tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 11:43:40 -05:00
DottaandPaperclip 1d26ae965e fix(ui): keep active runner status current and say Working (#13238)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task transcript shows a running agent's progress.
> - The active-run query stops polling when the live-run list has data.
> - The transcript still preferred that initial snapshot, so an old
execution-confirmation state could remain after work resumed.
> - This pull request uses the refreshed snapshot for the same run and
keeps active status text at Working.
> - Operators can see current activity without connection-state jargon.

## Linked Issues or Issue Description

**What happened?**

The task transcript said Reconnecting while the runner continued sending
messages and calling tools. The stale projection could also hide the
Thinking tail or stop the status spinner and timer.

**Expected behavior**

The selected run uses its current live snapshot. Active transcripts say
Working and show current activity. Completed and failed runs say Worked
and Stopped.

**Steps to reproduce**

1. Open a running task before its execution confirmation arrives.
2. Let the active-run query stop polling when the live-run list returns
the run.
3. Let the list refresh to working while the cached active-run snapshot
still says reconnecting.
4. Inspect the transcript status and activity tail.

**Paperclip version or commit**

Base commit: 52811c6ce.

**Deployment mode**

Built from source. The report concerns the new runner. The fix also
covers legacy transcript status text.

Searched open issues and open/closed pull requests for runner
reconnection work. No duplicate fix found.

## What Changed

- Refresh the selected active run from the polled list by matching the
task execution-run ID. Reject cached predecessors after run replacement.
- Use Working in native and legacy transcripts and active agent cards.
- Keep the active spinner, timer, and Thinking tail independent of
diagnostic execution phases. Terminal status takes precedence.
- Add stale-snapshot, timer, and terminal-state regressions. Update the
recovery story and documentation.

## Verification

- Focused Vitest suite: 148 tests passed across the run resolver, live
pill, runner turn, and task thread.
- `pnpm check:token-gates`: passed.
- `pnpm -r typecheck`: passed.
- Browser: inspected the recovery Storybook with a reconnecting
projection. It renders Working and Thinking.
- `pnpm build`: passed.
- `pnpm build-storybook`: passed.
- Full local `pnpm test:run` reported failures in unchanged server
tests; stopped the remaining general-server run after the equivalent CI
shards passed. The file-resource suite passes on rerun (35/35). Building
`build:runner-binaries` fixed a missing fake-provider binary; the
native-session-resume suite still has one continuity-reason assertion
mismatch (36/37 pass).
- Ran the remaining local test groups separately: both workspace groups
passed. All serialized suites passed except `pipelines-routes.test.ts`,
which still reports a socket hang-up on rerun (18/19 pass). The initial
access-route timeout passed on rerun. No server or runner source files
differ from the base.
- CI: all test shards, browser E2E, typecheck, runner verification, and
build passed on `22069462f30bdfcfb58f295da901a71ae5a47776`; the
aggregate verification gate passed (31 checks passed, 2 optional
Storybook checks skipped). Greptile is 5/5 with all review threads
resolved.

## Risks

- Low risk. This changes UI snapshot selection and presentation. It does
not change server recovery, leases, retry authority, or permissions.
- The live snapshot must match the task execution-run ID. Missing
matches use the cached active run only when its ID also matches.

## Model Used

OpenAI GPT-6 (Codex). The exact deployment identifier and context-window
size are not exposed in this session. Used reasoning, repository
inspection, code execution, and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused/UI and workspace
tests; full-suite limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 11:34:45 -05:00
DottaandPaperclip d10cbde815 fix(recovery): reject stale productive continuation wakes (#13173)
## Thinking Path

> - Paperclip manages agent work through tasks and runs.
> - Recovery continues assigned work when no live execution path
remains.
> - A recovery sweep can read an in-progress task before its run
completes.
> - The sweep can then observe the successful run after completion has
changed the task status.
> - This pull request checks current status and assignment under the
existing enqueue lock.
> - A stale continuation leaves a skipped wake receipt and creates no
run.
> - Task chat also omits an empty continuation cancelled before it
started because its task had become terminal.

## Linked Issues or Issue Description

Related public work: #10779 and #8419. Those older open changes address
terminal disposition across other recovery paths. This change uses the
existing scheduler guard for productive successful-run continuation and
adds real database lock contention coverage.

**What happened?**

Recovery could combine an old in-progress task snapshot with a newer
successful run. It queued an automatic continuation after the task was
done. Dispatch cancelled that run before it started, but task chat
displayed “Couldn't start” below the successful answer. This can happen
after the native runner's finish result has already been accepted. It
does not require a missing comment.

**Expected behavior**

Productive continuation must remain eligible when enqueueing acquires
the task lock. Completion, cancellation, reassignment, or a move away
from in-progress must prevent creation of the run. Actual execution
stops must remain visible.

**Steps to reproduce**

1. Let recovery select an assigned in-progress task whose latest run
succeeded with productive progress.
2. Hold the task row lock in another transaction and change the task to
done.
3. Let recovery attempt to enqueue while that transaction holds the
lock.
4. Commit completion. Before this fix, recovery creates a redundant run
from the stale snapshot.

**Paperclip version or commit**

Reproduced against master at `4042eb1c4` with deterministic integration
tests.

**Deployment mode**

Built from source with PostgreSQL. The bug is in core recovery and is
not adapter-specific.

## What Changed

- Pass the existing status-and-assignee guard for productive terminal
continuation recovery.
- Preserve a skipped wake receipt with the expected and actual task
state, without creating a run.
- Test actual PostgreSQL lock contention for native and legacy
completion, cancellation, backlog, review, blocked state, and
reassignment.
- Omit empty redundant pre-start cancellations from native and legacy
task chat. Preserve stop markers for runs that started.
- Document recovery eligibility at enqueue time.

## Verification

- All seven new race cases failed before the guard was connected.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm check:token-gates` passed.
- Task chat suite: 98 tests passed.
- Recovery integration suites: 290 tests passed, including 31
stale-queue tests.
- Full CI verification passed on `7ea71f04d`: all 31 active checks
succeeded, including all test shards, browser tests, build, typecheck,
and canary release dry run. Storybook visual regression was skipped by
its path filter.
- Greptile reviewed this commit at 5/5 with no review threads.
- The local `pnpm test:run` aggregate reported a setup failure in the
unchanged `tool-access-service.test.ts` suite. Its isolated rerun passed
all 231 tests without edits. The duplicate aggregate was stopped after
the complete CI matrix passed; it is not counted as a successful local
full-suite run.

## Risks

Low risk. The backend guard applies only to productive successful-run
recovery. It requires the task to remain in-progress with the same
agent. Other wake sources keep their current policy. The UI change only
suppresses empty redundant cancellations; run records remain available.
No schema change or migration is required.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
edits, and local test execution. The exact deployment snapshot and
context window are not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 08:51:50 -05:00
DottaandPaperclip 018ca5daaf fix: verify ACP Stop and preserve safe continuation (#13119)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task controls coordinate provider execution and queued user
messages.
> - Stop could finish before an embedded ACP provider stopped its tools.
> - A later request could be held for reconciliation without a clear
task response.
> - A restored provider could also retain the stopped run's API
credential.
> - This pull request verifies provider termination and preserves safe
session continuation.
> - Operators can continue known-safe work and see why uncertain work
cannot start.

## Linked Issues or Issue Description

**What happened?**

Stop could leave an embedded ACP provider running. A queued follow-up
followed by “go” could fail before it reached the provider. Task chat
could show a generic missing-response message. Even a restored session
could use the previous run's credential and fail its task update.

**Expected behavior**

Stop waits for confirmed provider termination. A later explicit wake
continues the same compatible session only when recorded actions have
known outcomes. It carries pending comments and the current run's
environment. Uncertain actions retain a visible reconciliation hold.
Composer Stop preserves the existing pause rule: conversation can
continue while paused, but task work requires Resume.

**Steps to reproduce**

1. Start an embedded ACP task.
2. Send a second request while the provider is running.
3. Interrupt the run, then send “go”. Also test composer Stop followed
by Resume work.
4. Check that the request is delivered once and that the provider can
complete the task through the current run's API credential.
5. Repeat with an unfinished write. Confirm that the write stops and
that further execution stays blocked with a visible reason.

**Paperclip version or commit**

Built from source on master at `3bc60dd8b` plus this branch.

**Deployment mode**

Local source build with an isolated embedded PostgreSQL instance.

Refs #11183. Refs #12552. Those changes address recovery after operator
cancellation. This change also covers embedded ACP termination, session
proof, pending-comment delivery, and task feedback.

## What Changed

- Propagate Stop into embedded ACP and wait for bounded adapter cleanup
and provider exit. Retain the actual ChildProcess object for forced
termination on all platforms; never signal a recycled numeric PID.
- Preserve interrupted checkpoints only for acknowledged, local,
persistent sessions with settled reads or no tools. Keep writes,
incomplete actions, and forced termination blocked.
- Restore the same compatible provider session with the current run's
environment. Reject fresh-session fallback for an interrupted
checkpoint.
- Adopt pending comments on the next explicit wake. Stop alone does not
dispatch them.
- Share the execution-blocker rule across dispatch, Resume, and task
detail. Show Stopped or Couldn't start with the recorded reason. Resolve
the stopped agent for the run link, including reviewer runs.
- Keep execution reconciliation holds intact when generic recovery sees
queued comments or healthy child tasks.
- Add process, service, component, and browser regression coverage. Fix
disposable database cleanup and React test settling exposed by the full
suite.

## Verification

- Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`.
- Passed all three `acp-stop-continuation.spec.ts` browser journeys.
They use an actual ACP child process and require task completion through
the agent API.
- Passed 165 adapter execution, operator-stop, and child-process control
tests, 17 queued-comment route tests, and 65 tests in the two adjusted
UI suites. Earlier focused recovery, heartbeat, and task-control tests
also passed.
- Manually used the browser to queue a request, Stop, send “go” while
paused, and Resume. The same session answered once and moved the task to
Done with the current run's credential.
- Manually interrupted an unfinished write. Its file size stayed fixed
for five seconds. “Go” showed the reconciliation reason and did not
start another provider prompt.
- Separate live Claude ACP smoke checks confirmed that Stop ended a
disposable local write and that a no-tool interruption could resume the
exact provider session. The browser fixture does not call Drive or
another external app.
- Passed all 5,615 UI tests and 3,090 other workspace tests. The CLI and
general server groups pass with targeted retries: two transient server
failures passed together on retry, and two embedded-database startup
failures passed after removing abandoned shared-memory segments from
this task's completed browser fixtures. All 144 serialized server suites
completed, with 2,189 tests passing after two transient HTTP socket
failures passed on retry.
- Passed all 135 heartbeat process/recovery tests, including a
deterministic regression that failed before the recovery-sweep fix.
- Passed 18 dispatch integration tests, including stopped-reviewer
links, company boundaries, and malformed run IDs.
- Greptile is 5/5 on `7dd170d83`, with zero unresolved review threads.
The security scan and all required CI gates pass for the same commit.

## Risks

- Safe continuation depends on complete tool reporting and a restorable
local provider session. Unknown outcomes remain blocked and require
reconciliation.
- Provider cleanup can take time. A timeout does not grant replay
permission.
- The change adds optional adapter context fields and an optional issue
projection. It does not change the database schema or require a
migration.
- Test cleanup truncates company data only in a disposable test
database.

## Model Used

OpenAI GPT-6, running as Codex with repository tools, code execution,
and browser interaction. The runtime does not expose a more specific
model deployment ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-09 22:06:06 -05:00
DottaandPaperclip 3b550c80fa fix(codex): correct startup trust, history reads, and resume usage (#13110)
## Thinking Path

> - Paperclip runs Codex locally and in remote sandboxes.
> - The runner must preserve startup configuration and session identity.
> - Missing project trust can disable repository configuration.
> - Full-history requests use deprecated provider fields.
> - Resume usage describes old work and must not become new run usage.
> - This change corrects startup trust, state reads, and usage
classification.

## Linked Issues or Issue Description

**What happened?**

Normal Codex runs could show repository-trust and history-deprecation
warnings.
Resume could report the preceding turn's token snapshot as a late-turn
warning.
The historical last-usage value could also be attributed to the new run.

**Expected behavior**

Trust the server-selected startup root in isolated configuration. Read
lightweight
provider state and paginated evidence. Use historical cumulative usage
as a
baseline without a new charge or user-facing warning.

**Steps to reproduce**

1. Start a native Codex task in a selected repository.
2. Finish the turn and resume the provider thread.
3. Inspect provider notices, history requests, and per-run usage.
4. Repeat startup and cold resume inside a Daytona sandbox.

**Paperclip version or commit**

Codex CLI 0.153.4 is the pinned runtime and reproduced baseline.
Replayed onto master at 6abeb6733. Related authority work: Refs #13092.
This PR retains its startup cleanup and protocol-integrity checks.

**Deployment mode**

Local source checkout and disposable Daytona sandbox.

## What Changed

- Classify the exact historical resume usage event before the generic
stale-turn warning.
- Persist cumulative usage baselines across recovery of the same run.
- Use excludeTurns on resume and lightweight thread reads.
- Page turn metadata and selected turn items with cursor and identity
validation.
- Reject unsupported or incomplete history instead of guessing that
execution is idle.
- Trust the startup execution root on its host, including Git worktree
trust keys.
- Start Codex in that root and retain the selected sandbox profile on
later turns.
- Keep unrelated isolated configuration and Codex's separate hook trust
policy.
- Add Rust, TypeScript, accounting, native integration, and local
run-log documentation.

## Verification

- Codex and native-transport TypeScript: 333 passed before PR replay.
- Adjacent OpenCode/ACPX driver and accounting tests: 49 passed.
- Rust library, serialized: 226 passed. Native Codex integration: 72
passed, 1 ignored, plus two pagination regressions.
- Repository typecheck and build passed. All repository test groups have
passing coverage after fixture and resource retests; the initial
monolithic command was not clean.
- Fresh real Codex native browser tasks returned correct answers without
the three targeted notices. Answers persisted after refresh and restart.
- Real same-thread TypeScript driver tests passed locally and in
Daytona, including cold resume, configuration, skills, and an approved
harmless hook.
- Local usage summed to 64,607 tokens. Daytona usage summed to 42,737
tokens. Each sum matched its final session total exactly.
- See doc/plans/2026-09-09-codex-integration-acceptance.md for the scope
and limits of the live tests.
- After replay onto current master and review fixes: 334 Codex, backend,
and live-session tests passed, including checkpoint serialization and
real-runner process restart. TypeScript checks passed.
- The native Codex integration run passed 83 tests; the large lineage
test passed separately with the release runner (its debug build exceeded
the test deadline).
- All GitHub checks passed on the final PR head. Greptile is 5/5 with no
unresolved review threads. CI regenerates the lockfile for the added
TOML dependency, per repository policy.
- The first server shard hit a timing-dependent duplicate-key failure in
the unchanged artifact-document concurrency test. Its focused 11-test
suite passed locally. One CI retry on the same head passed all 103 files
and 1,405 tests (2 skipped): [retry
result](https://github.com/paperclipai/paperclip/actions/runs/34398832930/job/102631274667).

## Risks

- Trust applies only to the server-selected startup root and isolated
configuration. Sandbox and tool permissions remain authoritative.
- Codex still requires approval of individual hook hashes. This change
does not bypass that policy.
- Providers without the required history APIs fail explicitly.
- Daytona acceptance used the production TypeScript driver. Remote
Paperclip UI and remote Rust execution were not tested.
- No new public API, database state, recovery policy, or UI control is
included.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`). Used for reasoning, code edits,
tool use,
and test execution. The exact context-window limit is not exposed in
this
session. Real-provider acceptance used Codex CLI 0.153.4 with
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-09 15:35:18 -05:00
DottaandPaperclip ca96e1eb0a fix(runner): keep streaming after task completion tools (#13108)
Keep receiving provider events after paperclip_finish, drain pending event persistence, and select the final assistant answer after the provider turn ends. Preserve cancellation, failure, and governed-wait behavior.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-09 15:22:00 -05:00