mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 16:35:27 +02:00
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents and their work > - Heartbeat admission decides when an agent should start another adapter session for an issue > - After process-loss recovery, assignment pollers and reconcilers can repeatedly request another wake while the issue remains `in_progress` > - When the preceding runs succeeded without issue-visible progress, those event-free wakes provide no new information but still pay the full cost of an adapter session > - Existing liveness evidence is too broad for this case because workspace tool calls can make a run look active without moving the issue > - This pull request adds an issue-scoped admission throttle for consecutive no-progress re-wakes while preserving every wake that carries new information or recovery intent > - The benefit is bounded recovery cost without delaying comments, operator actions, failures, or other meaningful events ## Linked Issues or Issue Description No public GitHub issue exists for this bug. **What happened?** After a process died, external wake drivers could re-wake the same agent for the same `in_progress` issue every few seconds. Each succeeded run that produced no issue-visible progress could be followed by another full adapter session despite no new issue input. In the observed recovery smoke, one recovery consumed 25 sessions and 2.4× the direct-run cost. **Expected behavior** Repeated event-free re-wakes should back off after consecutive successful runs produce no issue-visible progress. Any new information, explicit operator intent, or failed-run recovery should continue immediately. **Steps to reproduce** 1. Start an issue heartbeat and simulate process loss while the issue remains `in_progress`. 2. Allow assignment/reconciliation drivers to request repeated event-free wakes for the same agent and issue. 3. Complete each follow-up run successfully without adding a comment, issue mutation, document, work product, interaction, or continuation. 4. Observe repeated adapter sessions starting every few seconds without new issue input. **Environment** - Version: reproduced on `master` before this change - Deployment: local development, built from source - Adapter scope: core bug; not adapter-specific - Database: reproduced and tested with embedded Postgres ## What Changed - Add a pure issue re-wake throttle that detects consecutive succeeded runs without issue-visible progress and applies a 120-second exponential cooldown capped at 30 minutes. - Gate event-free `enqueueWakeup` requests and return the explicit skip reason `issue_rewake_throttled` while the cooldown is active. - Always bypass throttling for comment wakes, new issue activity, explicit resumes, `forceFreshSession`, event-shaped reasons, and post-failure recovery. - Add focused pure unit coverage and database-backed heartbeat admission coverage for throttle and bypass behavior. ## Verification - `cd server && pnpm vitest run src/__tests__/issue-rewake-throttle.test.ts` — 12 passed. - `cd server && pnpm vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6 passed with embedded Postgres. - `cd server && pnpm run typecheck` — passed. - Neighbor suites previously verified: `heartbeat-dependency-scheduling`, `heartbeat-process-recovery`, `run-continuations`, `heartbeat-issue-liveness-escalation`, `recovery-stale-issue-lock-sweep`, and `heartbeat-comment-wake-batching` — 131 tests passed. ## Risks - A progress classifier that is too narrow could defer a legitimate event-free poll; the cooldown is bounded and new issue activity bypasses it immediately. - A progress classifier that is too broad could allow the original heartbeat storm; tests intentionally distinguish issue-visible mutations from workspace-only activity. - Low compatibility risk: no schema, API contract, or migration changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent. The runtime does not expose the exact underlying model ID or context-window size; reasoning, terminal tool use, code inspection, GitHub CLI access, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public PR branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>