mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 03:08:10 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its web UI keeps live views (issue threads, run transcripts, dashboard) fresh via React Query polling plus a live-events websocket, coordinated across tabs with a `BroadcastChannel` layer > - Long-lived tabs viewing live agent runs grew to multi-GB memory footprints while their JS heap stayed ~60–200 MB — so the memory is off-heap (Blink/native + committed allocator arenas), not a classic JS leak > - Live profiling of a reproduced 2.7 GB / 66 MB tab showed 15–30% idle CPU, ~3 fetches/sec across overlapping poll loops, and ~8 `setInterval` create/clear cycles per second whose rate grew ~7× as the tab aged — relentless allocation churn that inflates committed memory the OS never reclaims, amplified across tabs by the cross-tab fan-out > - This pull request cuts that churn at its four largest sources (invalidation storm, per-instance 1 s timers, redundant polling, unbounded streamed-run set) > - The benefit is that idle tabs do far less periodic work, so their off-heap footprint stops ballooning over a long session ## Linked Issues or Issue Description No public GitHub issue exists; describing inline per CONTRIBUTING.md → "Link Issues or Describe Them In-PR", following the bug report template. **What happened?** Browser tabs viewing live agent runs grew to 8–16 GB memory footprint over a long session (multiple tabs open), while each tab's "live" JS heap stayed only ~150–250 MB. Every idle tab also burned 15–30% CPU. Tabs eventually approached the ~4 GB V8 heap ceiling / OS pressure and could crash. **Expected behavior** Tabs viewing live runs should hold a bounded footprint and do minimal work while idle, regardless of how long they stay open or how many tabs are open. **Steps to reproduce** Open several issue/run tabs that have agents actively streaming and leave them open for a while. Watch Chrome's Task Manager: Memory Footprint climbs into the GBs while "JavaScript Memory" stays small, and CPU stays high on idle tabs. Reproduced in ~90 minutes: a tab reached 2.7 GB footprint on a 66 MB JS heap, and the per-second timer-churn rate was ~7× higher on a 90-minute-old tab than a fresh one. **Paperclip version or commit** Branch `fix/live-updates-churn`, off `master`. **Deployment mode** Local dev (`pnpm dev`), web UI. Not adapter-specific — core UI live-updates plumbing (observed with `claude_local` / `codex_local` runs). ## What Changed - **`ui/src/lib/query-invalidation-batcher.ts` (new)** — `createInvalidationBatcher` throttles + de-dupes React Query invalidations into one flush per ~300 ms, and `createCoalescingQueryClient` wraps the client via a `Proxy` so only `invalidateQueries` is batched (optimistic `setQueryData` writes stay immediate). Wired into `LiveUpdatesProvider`, which previously invalidated synchronously on every websocket event. - **`ui/src/hooks/useSecondTick.ts` (new)** — one shared, ref-counted, page-wide 1 s ticker. `useLiveElapsed` in `IssueChatThread` now uses it instead of a per-instance `setInterval` that forced a full-thread re-render every second per live element. - **`ui/src/components/transcript/useLiveRunTranscripts.ts`** — when the realtime websocket is enabled, the recurring log poll backs off to a 30 s safety-net cadence instead of polling every 2 s on top of the live stream. Added a marker for the durable poll→push rearchitecture. - **`ui/src/lib/issueChatTranscriptRuns.ts`** — `resolveIssueChatTranscriptRuns` now caps the streamed run set (live/active runs always kept; most-recent linked runs fill up to 20) so a large run history can't open a live-transcript poll per historical run. - **`ui/src/main.tsx`** — explicit `gcTime` so cross-tab-published cache entries for unobserved resources are collected promptly. - Tests for the batcher, shared ticker, and run cap. ## Verification - `vitest`: new suites `query-invalidation-batcher.test.ts` (batcher collapses 20 invalidations → 1 flush; keeps distinct keys/variants; dispose cancels; proxy passes non-invalidate methods through), `useSecondTick.test.tsx` (single ref-counted timer, stops when idle), `issueChatTranscriptRuns.test.ts` (cap keeps newest + live). All pass. - Existing affected suites pass: `LiveUpdatesProvider` (23), `IssueChatThread` (), `useLiveRunTranscripts`, `AgentDetail.instructions` — 109 tests across affected files. - `tsc -b` clean. - Behavior confirmed by live profiling before the change (2.7 GB / 66 MB tab, ~8 interval churns/sec growing 7× with age). Runtime churn reduction should be re-measured against a rebuilt bundle with the same instrumentation. ## Risks Low-to-moderate; all changes reduce work rather than add features. - **Invalidation batching** delays live-driven refetches by up to ~300 ms. Optimistic `setQueryData` writes (e.g. the visible issue's new comment) remain immediate, so foreground updates still feel instant; only the safety-net refetch is throttled. Non-live invalidations (user actions, mutations) are unaffected — they use the real client. - **Poll back-off** relies on the websocket as the live source when realtime is enabled; a 30 s fallback poll still covers gaps/reconnects (both the transcript hook and `LiveUpdatesProvider` also auto-reconnect). - **Run cap (20)** means an issue with a very large run history streams live transcripts only for its live/active + 20 most-recent runs; older runs still open normally via their run pages. - Downstream test fallout (timing-sensitive tests around invalidation/polling) may need adjustment — flagged intentionally for follow-up. Durable follow-up (out of scope, marked in code): replace transcript/run polling with server push (SSE/websocket deltas) so idle tabs do no periodic work at all. ## Model Used - **Provider:** Anthropic, via the Claude Code CLI. - **Model:** Claude Opus 4.8 (`claude-opus-4-8`). - **Reasoning mode:** Extended thinking enabled. - **Capabilities used:** tool use (shell, file editing), sub-agent fan-out for codebase analysis, and the Chrome DevTools MCP to reproduce and profile the memory/CPU churn on a live instance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work (a bug/perf fix, not planned core feature work) - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [ ] I have updated relevant documentation to reflect my changes (N/A — no user-facing docs; rationale documented inline) - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
73 lines
2.5 KiB
TypeScript
73 lines
2.5 KiB
TypeScript
import type { ActiveRunForIssue, LiveRunForIssue } from "../api/heartbeats";
|
|
import type { RunTranscriptSource } from "../components/transcript/useLiveRunTranscripts";
|
|
import type { IssueChatLinkedRun } from "./issue-chat-messages";
|
|
|
|
/**
|
|
* Upper bound on how many runs an issue thread streams live transcripts for.
|
|
*
|
|
* Live/active runs are always included (they are the ones actually streaming);
|
|
* older linked runs are historical and only fill the remaining slots, most
|
|
* recent first. Without this cap a long-lived issue with a large run history
|
|
* would open a live-transcript poll/subscription for every run it ever had —
|
|
* multiplying the steady-state polling and re-render churn (and off-heap
|
|
* growth) that this change set is fixing.
|
|
*/
|
|
export const MAX_ISSUE_CHAT_TRANSCRIPT_RUNS = 20;
|
|
|
|
function toTimestamp(value: Date | string | null | undefined): number {
|
|
if (!value) return 0;
|
|
const ms = value instanceof Date ? value.getTime() : new Date(value).getTime();
|
|
return Number.isFinite(ms) ? ms : 0;
|
|
}
|
|
|
|
export function resolveIssueChatTranscriptRuns(args: {
|
|
linkedRuns?: readonly IssueChatLinkedRun[];
|
|
liveRuns?: readonly LiveRunForIssue[];
|
|
activeRun?: ActiveRunForIssue | null;
|
|
limit?: number;
|
|
}): RunTranscriptSource[] {
|
|
const { linkedRuns = [], liveRuns = [], activeRun = null, limit = MAX_ISSUE_CHAT_TRANSCRIPT_RUNS } = args;
|
|
const combined = new Map<string, RunTranscriptSource>();
|
|
|
|
for (const run of liveRuns) {
|
|
combined.set(run.id, {
|
|
id: run.id,
|
|
status: run.status,
|
|
adapterType: run.adapterType,
|
|
logBytes: run.logBytes,
|
|
lastOutputBytes: run.lastOutputBytes,
|
|
});
|
|
}
|
|
|
|
if (activeRun) {
|
|
combined.set(activeRun.id, {
|
|
id: activeRun.id,
|
|
status: activeRun.status,
|
|
adapterType: activeRun.adapterType,
|
|
logBytes: activeRun.logBytes,
|
|
lastOutputBytes: activeRun.lastOutputBytes,
|
|
});
|
|
}
|
|
|
|
// Live/active runs above are always retained; fill the remaining slots with
|
|
// the most recently created linked runs so the retained set is bounded.
|
|
const remainingLinked = [...linkedRuns]
|
|
.filter((run) => !combined.has(run.runId) && run.adapterType)
|
|
.sort((a, b) => toTimestamp(b.createdAt) - toTimestamp(a.createdAt));
|
|
|
|
for (const run of remainingLinked) {
|
|
if (combined.size >= limit) break;
|
|
const adapterType = run.adapterType;
|
|
if (!adapterType) continue;
|
|
combined.set(run.runId, {
|
|
id: run.runId,
|
|
status: run.status,
|
|
adapterType,
|
|
hasStoredOutput: run.hasStoredOutput,
|
|
logBytes: run.logBytes,
|
|
});
|
|
}
|
|
|
|
return [...combined.values()];
|
|
}
|