mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The heartbeat scheduler claims queued runs for each agent in `startNextQueuedRunForAgent`, and startup recovery calls it through `resumeQueuedRuns` > - `claimQueuedRun` can throw a 4xx `HttpError` that the run's own rows decide, for example the 403 from #14043 ("Queued-message interrupt authority is unavailable") > - The claim loop does not catch that error. The run stays `queued`, so every recovery pass and every restart claims it again and gets the same error > - During periodic recovery the error is logged, but it stops the claim loop, so all queued runs of every agent after it wait. During startup recovery the error is rethrown, so the server cannot boot > - We saw this on a self-hosted instance: one card-answer run from #14043 put the server in a crash loop for about 11 hours (216 failed boots) until the run was cancelled by hand in the database > - This pull request cancels a queued run when its claim fails with a 403 that cannot clear, keeps a run queued when its 4xx rejection can clear later, and in both cases continues with the rest of the queue > - The benefit is that one bad queued run can no longer stop the server from starting or stall other agents. #14510 fixes the specific identity mismatch; this change is the safety net for this whole class of error ## Linked Issues or Issue Description Refs #14043 Refs #14510, #13539, #13315 ## What Changed - `server/src/services/heartbeat.ts`: The claim loop in `startNextQueuedRunForAgent` catches errors from `claimQueuedRun` for each run. - A 403 is a permanent rejection. In the claim path it only comes from the run's persisted identity (an unverifiable interrupt receipt, a manual wake without a user), and those rows do not change. The loop cancels that run with error code `queued_run_claim_rejected`. `cancelRunInternal` also cancels the wakeup request and releases the issue execution lock. - The cancel runs after the agent start lock is released (`.finally()`), because `cancelRunInternal` promotes the next queued run under the same lock. A cancel inside the lock waits 30 seconds for the stale-lock timeout. - Any other 4xx (for example a 422 `responsible_user_unresolved`, or a 409) can clear later. The run stays queued, the loop logs a warning, and it continues with the next run. A later recovery pass tries the run again. - Non-HTTP errors and 5xx errors (for example database errors) keep propagating, as before. Startup still fails when the database or another dependency is broken. - If a cancel itself fails, the loop logs the error. The run stays queued for the next recovery pass. - `server/src/__tests__/heartbeat-queued-run-claim-isolation.test.ts`: New embedded-Postgres tests with a mocked adapter: - A queued-comment interrupt run with an unverifiable receipt is cancelled, `resumeQueuedRuns()` resolves, and a second pass stays clean. - A claimable run behind a rejected run of the same agent is claimed and executed. - A run with an unresolvable responsible user (422) stays queued, and the claimable run behind it is executed. - A rejected run for one agent does not stop the queue of another agent. ## Verification - `cd server && npx vitest run src/__tests__/heartbeat-queued-run-claim-isolation.test.ts` passes. - Without the change in `heartbeat.ts`, all four tests fail (the 403 `Queued-message interrupt authority is unavailable` and the 422 `responsible_user_unresolved` escape `resumeQueuedRuns()`). - All `heartbeat*`, `*queued*` and `run-identity` server suites pass locally (1252/1253). The one local failure, `heartbeat-process-recovery` "redacts opaque environment-bound credentials from Sentry diagnostics", fails the same way on unchanged `master` on my machine and passes in CI. - `pnpm -r typecheck`, `pnpm test:run` and `pnpm build` pass locally. ## Risks - Behavior shift: a queued run whose claim fails with a 403 is now cancelled, not left queued. Such a run could never start before this change, so no runnable work is lost. The cancel reason and the `queued_run_claim_rejected` error code make the cancel visible on the run. - A run with another 4xx rejection stays queued, as before, but it no longer stops the rest of the queue. It logs a warning on each recovery pass until the cause clears. - Low risk otherwise: no schema, API or UI change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (`claude-opus-5-5`) in Claude Code, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>