mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
fix(server): restore stranded recovery continuations (#9630)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents and their work. > - Its server recovery layer classifies blocked issue graphs and restores interrupted heartbeat execution. > - A dependent issue could remain dispatch-suppressed by a cancelled blocker without producing operator-visible attention when the dependent still displayed as todo or backlog. > - Separately, a monitor-triggered run that lost its process before disposition could consume the monitor's one-shot wake without scheduling the existing bounded continuation. > - Both gaps strand useful work even though Paperclip already has the relevant blocker-attention and process-loss recovery mechanisms. > - This pull request widens the existing classification path and reuses the single process-loss retry for monitor dispatches with no future wake. > - The benefit is visible, routable recovery without weakening dependency checkout rules or introducing an unbounded retry loop. ## Linked Issues or Issue Description No matching public GitHub issue or pull request was found. ### What happened? Two server recovery cases could leave work stranded: 1. A non-terminal, agent-assigned issue with an unresolved cancelled blocker remained ineligible for checkout, but blocked-chain liveness classification only inspected issues already displaying `blocked` or `in_review`, so the existing `blocked_by_cancelled_issue` attention was not surfaced. 2. A one-shot issue monitor cleared its next check when dispatched. If that monitor-triggered run ended as `process_lost` without a tracked local child, the existing bounded retry gate rejected it and no future monitor wake remained. ### Expected behavior - Cancelled blockers continue to be unresolved dependencies, and their dependents receive blocker attention regardless of whether the dependent currently displays as backlog, todo, blocked, or in review. - A monitor-triggered run lost before disposition receives exactly one bounded continuation when no future monitor check exists; a second loss follows the normal recovery-action escalation path. ### Steps to reproduce 1. Create an agent-assigned todo issue blocked by a cancelled issue and run issue-graph liveness classification. 2. Observe that no cancelled-blocker finding appears before this change. 3. Dispatch a due issue monitor, clear its one-shot `monitorNextCheckAt`, and mark the resulting untracked run `process_lost`. 4. Observe that no retry is queued before this change. ### Environment - Paperclip commit: `3e348b96b` - Deployment: built from source / local test environment - Adapter: not adapter-specific; core server recovery - Database: embedded test database ## What Changed - Inspect non-terminal, agent-assigned issues with unresolved blocker edges during blocked-chain liveness classification. - Include cancelled dependents in the existing blocked-inbox attention query while preserving company-scoped relation checks. - Allow monitor-triggered `process_lost` runs with no future monitor wake to use the existing single bounded retry. - Mark monitor recovery retries as continuation-needed context and retain the existing second-loss escalation behavior. - Document cancelled-blocker and monitor-dispatch recovery semantics. - Add focused regressions for liveness findings, attention propagation, one retry, and second-loss escalation. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/issue-blocker-attention.test.ts server/src/__tests__/issue-liveness.test.ts` — 3 files, 126 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. ## Risks - Low risk and server-only. The liveness scan inspects more unresolved dependency shapes, which can produce additional existing attention entries for previously invisible cancelled blockers. - Monitor recovery remains bounded by `processLossRetryCount < 1`, and the extra path only applies when the dispatch was monitor-triggered and no future monitor check exists. - No schema, migration, authorization, API-contract, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.4` through Codex CLI, with reasoning, repository tool use, command execution, and test execution capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
1 parent
bbb3e19c6f
commit
9af96461d5
7 files changed
+308
-6
No files matched your search
@@ -169,6 +169,8 @@ Use it for:
|
||||
|
||||
Blocked issues should stay idle while blockers remain unresolved. Paperclip should not create a queued heartbeat run for that issue until the final blocker is done and the `issue_blockers_resolved` wake can start real work.
|
||||
|
||||
`cancelled` is terminal for the blocker issue itself, but it does not satisfy the dependency. A cancelled blocker edge remains unresolved until the edge is removed or replaced, and Paperclip must surface blocker attention on the dependent regardless of whether that dependent is currently displayed as `blocked`, `todo`, `backlog`, or another non-terminal agent-owned status.
|
||||
|
||||
If a parent is truly waiting on a child, model that with blockers. Do not rely on the parent/child relationship alone.
|
||||
|
||||
## 7. Accepted-Plan Decomposition
|
||||
@@ -266,6 +268,8 @@ An external wait counts as a live or waiting path only when the next move surviv
|
||||
- a first-class blocker or `blocked` disposition that names the external owner and concrete action required to unblock the issue
|
||||
- a delegated child issue with a responsible owner and its own healthy action path, plus a blocker edge when the source issue must wait for that child; `parentId` alone is not a dependency
|
||||
|
||||
A one-shot issue monitor consumes its persisted `nextCheckAt` when it dispatches the assignee wake. If that monitor-consuming run is lost before it records a new disposition or future monitor, Paperclip restores exactly one bounded continuation using the existing process-loss retry limit; if that continuation is also lost, the normal recovery-action escalation owns the next step instead of creating another monitor loop.
|
||||
|
||||
An unmanaged local process is not a durable action path. Shell jobs started with `&`, `nohup`, local polling loops, detached PTY sessions, adapter child processes, or similar background watchers do not keep an issue live unless Paperclip persists them as a run or pairs a managed runtime service with a monitor, scheduled wake, blocker, or delegated issue that owns the next check. A PID, session id, log file, comment, or promise to check later is evidence only. The process may be killed when the adapter invocation or heartbeat exits and cannot be assumed observable or recoverable by another worker.
|
||||
|
||||
Before a heartbeat finalizes, its issue disposition must therefore be evaluated from durable Paperclip state, not from processes still visible only to that heartbeat. An agent-owned issue may remain `in_progress` after the heartbeat only when another valid action-path primitive already exists. If the only claimed continuation is a local/background watcher, finalization treats the issue as having no live path even when the process has not yet been observed exiting.
|
||||
|
||||
Reference in new issue
Block a user