mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The stranded-issue sweeper is what puts an assigned issue back on a
live path when its run dies.
> - A failed or interrupted run gets a bounded transient retry; when
that budget is spent, the retry scheduler queues nothing and reports the
exhaustion.
> - The sweeper treated that "nothing queued" like every other "nothing
queued" and skipped the issue, on every tick, forever: `in_progress`, no
run, no path, no notice.
> - Three server restarts in a row (a deploy storm) are enough to spend
the budget, so this is reachable in ordinary operation.
> - This pull request makes the sweeper escalate a spent budget to a
board-owned recovery action, the same visible `blocked` state its other
dead ends use.
> - The benefit is that no issue can sit assigned and silent after its
retries run out.
## Linked Issues or Issue Description
No existing issue found. Related work: #14028 (this branch's
predecessor: queue drain-time wakes, retry interrupted corrective runs)
fixed two neighbouring gaps but not this one.
**What happened?**
A routine-created issue's run was interrupted by a graceful server
shutdown, and both bounded transient retries were interrupted by the
next two shutdowns. The retry scheduler logged "Bounded retry exhausted
after 2 scheduled attempts; no further automatic retry will be queued"
and stopped. On every tick after that, `reconcileStrandedAssignedIssues`
reached the generic continuation lane, `enqueueStrandedIssueRecovery`
delegated to `scheduleRecoveryRetry`, which returned null (exhausted),
and the sweeper counted the issue as `skipped`. The issue stayed
`in_progress` with no run, no recovery action, and no comment for hours
until a person woke the agent by hand. The same hole exists in the
assigned-`todo` dispatch lane.
**Expected behavior**
When the transient retry budget for an interrupted or failed run is
spent, the sweeper should treat it as the dead end it is: escalate the
issue to `blocked` with a board-owned recovery action and a notice,
exactly as it does when a continuation retry chain or an assignment
retry chain is exhausted. Leaving the budget spent across restarts is
correct and unchanged; leaving the issue silent is not.
**Steps to reproduce**
1. Assign an agent using a conversation adapter (its interrupted runs
carry `conversationContinuation`, so legacy reconciliation does not
terminalize them) an issue and let a run start.
2. Interrupt the run with a graceful shutdown, let the transient retry
start, interrupt it, and repeat once more so `scheduledRetryAttempt`
reaches 2.
3. Run the stranded-issue sweep. Before this change: `skipped` every
tick, issue `in_progress`, no run, no recovery action. After:
`escalated`, issue `blocked`, one active board-owned
`issue_recovery_actions` row, a "No live execution path" notice.
**Paperclip version or commit**
master at bd6caf51bb (2026-09-25).
**Deployment mode**
Managed cloud instance restarted by fleet deploys; the code path is the
same for any operator whose server restarts more often than the retry
budget allows.
## What Changed
- `recoveryService` gains an optional `transientRetryBudgetSpent(run)`
dependency; `heartbeatService` wires it as
`executionFailureRetryCount(run) >=
BOUNDED_TRANSIENT_HEARTBEAT_RETRY_MAX_ATTEMPTS`, the same check
`scheduleBoundedRetryForRun` applies.
- `enqueueStrandedIssueRecovery` takes an optional `outcome`
out-parameter and sets `retryExhausted` when the failed predecessor's
retry returned nothing **because** the budget is spent. A null return
without it still means another authority owns the run (native runtime,
legacy reconciliation) and the caller leaves it alone; the deliberate
"failure recovery cannot fall through into the continuation queue" rule
is unchanged.
- The generic `in_progress` continuation lane and the assigned-`todo`
dispatch lane escalate on `retryExhausted` via
`escalateStrandedAssignedIssue` with a "No live execution path" notice
(danger tone), which creates the board-owned source-scoped recovery
action and moves the issue to `blocked`. Every other
`enqueueStrandedIssueRecovery` caller is unchanged.
- Tests (`heartbeat-process-recovery.test.ts`): an `in_progress` issue
with a spent budget escalates (blocked, one active board-owned action,
notice, no successor run, idempotent on the next sweep); an interrupted
run with budget remaining still gets its transient retry; an assigned
`todo` issue with a spent dispatch budget escalates. The existing guard
"does not reset an exhausted incident budget on server restart" keeps
its no-successor-run assertion and now expects the board escalation
instead of nothing, with a comment on why.
## Verification
```
pnpm -r --filter './packages/**' build
cd server
npx vitest run src/__tests__/heartbeat-process-recovery.test.ts \
src/__tests__/issue-recovery-actions.test.ts \
src/services/recovery/successful-run-handoff.test.ts \
src/__tests__/heartbeat-task-drain-admission-release.test.ts \
src/__tests__/attention-service.test.ts \
src/__tests__/heartbeat-comment-wake-batching.test.ts
npx tsc --noEmit -p tsconfig.json
```
Live reproduction: the exact stuck state (issue `in_progress`, latest
run `interrupted` with `scheduledRetryAttempt` 2, the "Bounded retry
exhausted" lifecycle event, sweeper `skipped` every tick) was observed
on a managed instance running current master before this change was
written.
## Risks
- Behavior change is limited to runs whose transient budget is already
spent, which previously produced no action at all. Nothing new is
retried; the change only adds the escalation, so no retry loop can be
introduced.
- The board-owned action spawns no run. Resolving it (restore to the
owner) re-dispatches through the existing recovery-action routes, the
same flow as every other stranded escalation.
- Native-runtime and legacy-reconciliation predecessors are untouched:
they return before the exhaustion check.
## Model Used
Claude (Anthropic) — `claude-fable-5-1`, extended thinking, tool use
(Claude Code CLI). The incident diagnosis and the choice to escalate
rather than re-dispatch were steered by the maintainer.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>