mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server's heartbeat scheduler queues, admits, and recovers agent
runs.
> - A control plane arms an instance task drain before it restarts the
process, so no new run starts mid-restart.
> - Two recovery paths treated that transient hold as a permanent
verdict: a wake that arrived during a drain was written as `skipped` and
never replayed, and a corrective disposition run that the restart
interrupted counted as an exhausted attempt, so the issue went to
`blocked`.
> - Both outcomes leave an assigned issue with no run and no path until
a person notices.
> - This pull request keeps drain-time wakes in the durable queue and
gives an interrupted corrective run the same bounded transient retry any
interrupted run gets.
> - The benefit is that a graceful restart never strands or blocks an
issue by itself.
## Linked Issues or Issue Description
No existing issue found. I searched open PRs that touch the task drain
(#13527 adds opt-in termination of active runs; #13965 handles deferred
wakes after a stale cancellation). Neither replays drain-time wakes or
changes the corrective-handoff escalation.
**What happened?**
1. An operator accepted a plan confirmation while the instance was in a
task drain (a control plane armed it before a deploy restart).
`enqueueWakeup` wrote the assignee's continuation wake with `status:
"skipped"` and `heartbeatSkip.reason: "task_drain"`. The drain is
process-local and cleared on the restart seconds later, but nothing
replays skipped wakes. The issue stayed in `todo` with no run for 20
hours until a person retried it by hand. The stranded-issue sweeper did
not help: a `todo` issue whose latest run succeeded is treated as
deliberately handed back.
2. On another issue, the watchdog correctly raised
`successful_run_missing_state` and queued the single corrective handoff
run. A graceful server shutdown (the same deploy restart) interrupted
that run with `server_shutdown_interrupted`. On the next boot,
`reconcileStrandedAssignedIssues` saw a corrective run at
`handoffAttempt >= maxHandoffAttempts` and escalated the issue to
`blocked` with a board-owned `missing_disposition` recovery action. The
agent never got to finish one corrective attempt.
**Expected behavior**
A task drain holds admission, not the request. A wake that arrives
during a drain must run once the drain lifts or the process restarts. An
interrupted corrective run is not evidence that the agent could not
choose a disposition; it should be retried like any interrupted run, and
escalate only when that retry budget is spent or a finished attempt
still leaves no disposition.
**Steps to reproduce**
1. Assign an agent an issue and `POST /api/instance/task-drain` with a
Cloud control assertion (or call `startTaskDrain({})` in a test).
2. Trigger any wake for that agent (assignment, comment, or an accepted
`request_confirmation`).
3. Observe the `agent_wakeup_requests` row: `status = skipped`, `reason
= heartbeat.scheduling_suppressed`, `payload.heartbeatSkip.reason =
task_drain`. Stop the drain or restart: no run is ever created for that
wake.
4. For the second case: let a `finish_successful_run_handoff` run be
interrupted by SIGTERM during a graceful shutdown, then boot. The issue
moves to `blocked` on the startup sweep without any retry.
**Paperclip version or commit**
master at e2f1a66aa7 (2026-09-25).
**Deployment mode**
Managed cloud instance behind a control plane that arms task drains
before deploy restarts; the code paths are the same for any operator who
uses the drain endpoint.
## What Changed
- `enqueueWakeup` no longer writes a wake as `skipped` when the only
suppression is `task_drain`. The wake and its queued run land in the
durable queue; `startNextQueuedRunForAgent` and `executeRun` already
refuse to admit work while the drain is active, and the boot-time
`resumeQueuedRuns` pass picks it up after the restart.
`worktree_instance` and `database_restore_in_progress` suppression still
write `skipped`: those holds are not transient restarts.
- `reconcileStrandedAssignedIssues`: when the latest run is a corrective
successful-run handoff at its attempt cap **and** that run is
`interrupted`, the sweeper schedules the bounded transient retry (via
`enqueueStrandedIssueRecovery` → `scheduleRecoveryRetry`) instead of
escalating. The retry keeps the handoff context, so it is still the
corrective run. If the retry budget is spent, or a finished attempt
still leaves no disposition, escalation proceeds exactly as before.
Failed corrective runs (adapter or provider failures) are unchanged:
they still escalate immediately.
- New sweep counter `successfulRunHandoffRetried`, included in the
startup and periodic recovery log lines.
- Tests: `heartbeat-task-drain-admission-release.test.ts` gains a case
that a wake during a drain is queued, held while the drain is active
(quiescence still reports true), and runs to completion once the drain
lifts. `heartbeat-process-recovery.test.ts` gains two cases: an
interrupted corrective run is retried with its handoff context and the
issue stays `in_progress`; an interrupted corrective run with a spent
transient budget escalates to `blocked` with the usual recovery action
evidence.
## Verification
```
pnpm -r --filter './packages/**' build
cd server
npx vitest run src/__tests__/heartbeat-process-recovery.test.ts # 296 passed
npx vitest run src/__tests__/heartbeat-task-drain-admission-release.test.ts \
src/__tests__/heartbeat-worktree-suppression.test.ts \
src/__tests__/heartbeat-scheduling-suppression.test.ts \
src/__tests__/heartbeat-task-drain.test.ts \
src/__tests__/instance-settings-routes.test.ts \
src/services/recovery/successful-run-handoff.test.ts \
src/__tests__/issue-recovery-actions.test.ts \
src/__tests__/heartbeat-comment-wake-batching.test.ts \
src/__tests__/attention-service.test.ts # 226 passed
npx tsc --noEmit -p tsconfig.json # clean
```
Live reproduction of the first defect: after this change was written, a
board wake issued against a draining managed instance (running current
master) was again recorded as `skipped` with `heartbeatSkip.reason =
task_drain`, which is the exact row the new test asserts no longer
appears.
## Risks
- A wake that lands during a drain now waits in the queue instead of
being dropped. `getTaskDrainStatus().pendingWakes` counts only in-flight
enqueue promises, so quiescence is unchanged and a control plane waiting
for the drain is not held longer. If a drain is stopped without a
restart, the queued run starts on the next scheduler tick.
- The interrupted-handoff retry reuses the existing transient retry
budget (two attempts) and the existing `hasActiveExecutionPath` skip, so
a sweep cannot double-schedule. Native-runtime corrective runs still
return `null` from the recovery enqueue and escalate as before.
- Self-hosted instances that never call the drain endpoint see no change
on the first path; the second path only changes behavior for corrective
runs interrupted by a graceful shutdown.
## Model Used
Claude (Anthropic) — `claude-fable-5-1`, extended thinking, tool use
(Claude Code CLI). Human-directed: diagnosis of the two production
incidents, the fix design, and review were steered by the maintainer.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>