fix(server): restore stranded recovery continuations (#9630)

## Thinking Path

> - Paperclip is the open source control plane people use to coordinate
AI agents and their work.
> - Its server recovery layer classifies blocked issue graphs and
restores interrupted heartbeat execution.
> - A dependent issue could remain dispatch-suppressed by a cancelled
blocker without producing operator-visible attention when the dependent
still displayed as todo or backlog.
> - Separately, a monitor-triggered run that lost its process before
disposition could consume the monitor's one-shot wake without scheduling
the existing bounded continuation.
> - Both gaps strand useful work even though Paperclip already has the
relevant blocker-attention and process-loss recovery mechanisms.
> - This pull request widens the existing classification path and reuses
the single process-loss retry for monitor dispatches with no future
wake.
> - The benefit is visible, routable recovery without weakening
dependency checkout rules or introducing an unbounded retry loop.

## Linked Issues or Issue Description

No matching public GitHub issue or pull request was found.

### What happened?

Two server recovery cases could leave work stranded:

1. A non-terminal, agent-assigned issue with an unresolved cancelled
blocker remained ineligible for checkout, but blocked-chain liveness
classification only inspected issues already displaying `blocked` or
`in_review`, so the existing `blocked_by_cancelled_issue` attention was
not surfaced.
2. A one-shot issue monitor cleared its next check when dispatched. If
that monitor-triggered run ended as `process_lost` without a tracked
local child, the existing bounded retry gate rejected it and no future
monitor wake remained.

### Expected behavior

- Cancelled blockers continue to be unresolved dependencies, and their
dependents receive blocker attention regardless of whether the dependent
currently displays as backlog, todo, blocked, or in review.
- A monitor-triggered run lost before disposition receives exactly one
bounded continuation when no future monitor check exists; a second loss
follows the normal recovery-action escalation path.

### Steps to reproduce

1. Create an agent-assigned todo issue blocked by a cancelled issue and
run issue-graph liveness classification.
2. Observe that no cancelled-blocker finding appears before this change.
3. Dispatch a due issue monitor, clear its one-shot
`monitorNextCheckAt`, and mark the resulting untracked run
`process_lost`.
4. Observe that no retry is queued before this change.

### Environment

- Paperclip commit: `3e348b96b`
- Deployment: built from source / local test environment
- Adapter: not adapter-specific; core server recovery
- Database: embedded test database

## What Changed

- Inspect non-terminal, agent-assigned issues with unresolved blocker
edges during blocked-chain liveness classification.
- Include cancelled dependents in the existing blocked-inbox attention
query while preserving company-scoped relation checks.
- Allow monitor-triggered `process_lost` runs with no future monitor
wake to use the existing single bounded retry.
- Mark monitor recovery retries as continuation-needed context and
retain the existing second-loss escalation behavior.
- Document cancelled-blocker and monitor-dispatch recovery semantics.
- Add focused regressions for liveness findings, attention propagation,
one retry, and second-loss escalation.

## Verification

- `pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/__tests__/issue-blocker-attention.test.ts
server/src/__tests__/issue-liveness.test.ts` — 3 files, 126 tests
passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.

## Risks

- Low risk and server-only. The liveness scan inspects more unresolved
dependency shapes, which can produce additional existing attention
entries for previously invisible cancelled blockers.
- Monitor recovery remains bounded by `processLossRetryCount < 1`, and
the extra path only applies when the dispatch was monitor-triggered and
no future monitor check exists.
- No schema, migration, authorization, API-contract, or UI changes.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI `gpt-5.4` through Codex CLI, with reasoning, repository tool
use, command execution, and test execution capabilities.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
DottaandPaperclip authored and GitHub committed 2026-07-15 16:41:39 -05:00
1 parent bbb3e19c6f
commit 9af96461d5
7 files changed
+308 -6

No files matched your search

+4
View File
@@ -169,6 +169,8 @@ Use it for:
Blocked issues should stay idle while blockers remain unresolved. Paperclip should not create a queued heartbeat run for that issue until the final blocker is done and the `issue_blockers_resolved` wake can start real work.
`cancelled` is terminal for the blocker issue itself, but it does not satisfy the dependency. A cancelled blocker edge remains unresolved until the edge is removed or replaced, and Paperclip must surface blocker attention on the dependent regardless of whether that dependent is currently displayed as `blocked`, `todo`, `backlog`, or another non-terminal agent-owned status.
If a parent is truly waiting on a child, model that with blockers. Do not rely on the parent/child relationship alone.
## 7. Accepted-Plan Decomposition
@@ -266,6 +268,8 @@ An external wait counts as a live or waiting path only when the next move surviv
- a first-class blocker or `blocked` disposition that names the external owner and concrete action required to unblock the issue
- a delegated child issue with a responsible owner and its own healthy action path, plus a blocker edge when the source issue must wait for that child; `parentId` alone is not a dependency
A one-shot issue monitor consumes its persisted `nextCheckAt` when it dispatches the assignee wake. If that monitor-consuming run is lost before it records a new disposition or future monitor, Paperclip restores exactly one bounded continuation using the existing process-loss retry limit; if that continuation is also lost, the normal recovery-action escalation owns the next step instead of creating another monitor loop.
An unmanaged local process is not a durable action path. Shell jobs started with `&`, `nohup`, local polling loops, detached PTY sessions, adapter child processes, or similar background watchers do not keep an issue live unless Paperclip persists them as a run or pairs a managed runtime service with a monitor, scheduled wake, blocker, or delegated issue that owns the next check. A PID, session id, log file, comment, or promise to check later is evidence only. The process may be killed when the adapter invocation or heartbeat exits and cannot be assumed observable or recoverable by another worker.
Before a heartbeat finalizes, its issue disposition must therefore be evaluated from durable Paperclip state, not from processes still visible only to that heartbeat. An agent-owned issue may remain `in_progress` after the heartbeat only when another valid action-path primitive already exists. If the only claimed continuation is a local/background watcher, finalization treats the issue as having no live path even when the process has not yet been observed exiting.