mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents make progress in heartbeats: the server wakes an agent session, it does a slice of work on an issue, records a disposition, and exits > - Benchmarking identical coding tasks run as Paperclip-orchestrated agent pairs vs invoking the same agent harness directly measured a 1.8–2.2× wall-clock slowdown for the Paperclip pairs, dominated by per-heartbeat orchestration overhead rather than model time > - Two contributors stood out: (1) since PF-4 (#4838) every `heartbeat_timer` wake starts a brand-new task session, so continuation work on a specific issue repays the full session-start and re-orientation cost on every heartbeat; (2) in degraded environments agents burn many tool calls retrying the same failing control-plane write before giving up > - This pull request reuses the task session for issue-scoped timer wakes (keeping the PF-4 fresh-session rule only for unscoped exploratory wakes, which were the original context-bloat case) and adds a bounded-retry rule to the wake prompt and core skill: after 2 consecutive failures of the same control-plane write, stop retrying it for the rest of the heartbeat and rely on the adapter/runtime status channel > - The benefit is materially less wall-clock and token overhead per heartbeat while preserving the context-bloat protection PF-4 was added for ## Linked Issues or Issue Description Refs #4838 (merged PF-4 change whose reset rule this refines), Refs #5287, Refs #1907 (related timer-heartbeat session work). No public GitHub issue exists for the slowdown itself; bug-report fields: - **What happened:** Agent pairs orchestrated through Paperclip heartbeats complete identical task sets 1.8–2.2× slower (wall-clock) than the same harness invoked directly. Profiling attributed the gap to per-heartbeat orchestration overhead: every timer wake discards the task session (full session start + re-orientation), and in degraded environments agents repeatedly retry the same failing control-plane write. - **Expected behavior:** Heartbeat orchestration should add minimal wall-clock overhead on top of the underlying harness; issue-scoped continuation work should not pay a fresh-session tax each interval. - **Steps to reproduce:** Run a fixed benchmark task set once through Paperclip issue heartbeats and once via direct harness invocation with the same model/config; compare wall-clock totals. - **Version/commit:** master @3d23c3b2c3, self-hosted deployment. ## What Changed - `server/src/services/heartbeat.ts`: `shouldResetTaskSessionForWake` now resets only for `heartbeat_timer` wakes with no derivable task key (unscoped exploratory wakes). Issue-scoped timer wakes reuse the issue's task session. `describeSessionResetReason` updated to stay in exact agreement. - `server/src/__tests__/heartbeat-timer-wake-session-reset-pf4.test.ts`: new cases for scoped vs unscoped timer wakes, plus the scoped case added to the reset/reason agreement invariant. - `packages/adapter-utils/src/server-utils.ts`: wake prompt template and execution contract gain a bounded-retry rule — after 2 consecutive failures of the same control-plane write, stop retrying it for the rest of the heartbeat, continue useful work, report the failure in the final response, and use the adapter/runtime status channel as the sanctioned fallback. - `packages/adapter-utils/src/server-utils.test.ts`: asserts the new prompt lines are present in both the template and the rendered wake prompt. - `skills/paperclip/SKILL.md`: documents the same bounded write-retry rule in the core Paperclip skill. ## Verification - `node_modules/.bin/vitest run packages/adapter-utils/src/server-utils.test.ts` — 1 file, 83 tests passed - `cd server && node_modules/.bin/vitest run src/__tests__/heartbeat-timer-wake-session-reset-pf4.test.ts` — 1 file, 14 tests passed - Both run on this branch rebased onto current master (3d23c3b2c3) ## Risks - Behavioral shift: issue-scoped timer wakes now reuse sessions, so a long-lived issue session can grow across heartbeats. Mitigated by keeping the PF-4 reset for unscoped wakes (the originally observed bloat case) and by existing session compaction. - Prompt/skill text changes alter agent guidance; the new rule is scoped narrowly to repeated failures of the same control-plane write. - No migrations, no API or schema changes, no dependency changes. ## Model Used - Claude (Anthropic) — `claude-fable-5` (Fable 5), extended reasoning with tool use, driven via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>