mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 05:31:46 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its web UI keeps live views (issue threads, run transcripts, dashboard) fresh via React Query polling plus a live-events websocket, coordinated across tabs with a `BroadcastChannel` layer > - Long-lived tabs viewing live agent runs grew to multi-GB memory footprints while their JS heap stayed ~60–200 MB — so the memory is off-heap (Blink/native + committed allocator arenas), not a classic JS leak > - Live profiling of a reproduced 2.7 GB / 66 MB tab showed 15–30% idle CPU, ~3 fetches/sec across overlapping poll loops, and ~8 `setInterval` create/clear cycles per second whose rate grew ~7× as the tab aged — relentless allocation churn that inflates committed memory the OS never reclaims, amplified across tabs by the cross-tab fan-out > - This pull request cuts that churn at its four largest sources (invalidation storm, per-instance 1 s timers, redundant polling, unbounded streamed-run set) > - The benefit is that idle tabs do far less periodic work, so their off-heap footprint stops ballooning over a long session ## Linked Issues or Issue Description No public GitHub issue exists; describing inline per CONTRIBUTING.md → "Link Issues or Describe Them In-PR", following the bug report template. **What happened?** Browser tabs viewing live agent runs grew to 8–16 GB memory footprint over a long session (multiple tabs open), while each tab's "live" JS heap stayed only ~150–250 MB. Every idle tab also burned 15–30% CPU. Tabs eventually approached the ~4 GB V8 heap ceiling / OS pressure and could crash. **Expected behavior** Tabs viewing live runs should hold a bounded footprint and do minimal work while idle, regardless of how long they stay open or how many tabs are open. **Steps to reproduce** Open several issue/run tabs that have agents actively streaming and leave them open for a while. Watch Chrome's Task Manager: Memory Footprint climbs into the GBs while "JavaScript Memory" stays small, and CPU stays high on idle tabs. Reproduced in ~90 minutes: a tab reached 2.7 GB footprint on a 66 MB JS heap, and the per-second timer-churn rate was ~7× higher on a 90-minute-old tab than a fresh one. **Paperclip version or commit** Branch `fix/live-updates-churn`, off `master`. **Deployment mode** Local dev (`pnpm dev`), web UI. Not adapter-specific — core UI live-updates plumbing (observed with `claude_local` / `codex_local` runs). ## What Changed - **`ui/src/lib/query-invalidation-batcher.ts` (new)** — `createInvalidationBatcher` throttles + de-dupes React Query invalidations into one flush per ~300 ms, and `createCoalescingQueryClient` wraps the client via a `Proxy` so only `invalidateQueries` is batched (optimistic `setQueryData` writes stay immediate). Wired into `LiveUpdatesProvider`, which previously invalidated synchronously on every websocket event. - **`ui/src/hooks/useSecondTick.ts` (new)** — one shared, ref-counted, page-wide 1 s ticker. `useLiveElapsed` in `IssueChatThread` now uses it instead of a per-instance `setInterval` that forced a full-thread re-render every second per live element. - **`ui/src/components/transcript/useLiveRunTranscripts.ts`** — when the realtime websocket is enabled, the recurring log poll backs off to a 30 s safety-net cadence instead of polling every 2 s on top of the live stream. Added a marker for the durable poll→push rearchitecture. - **`ui/src/lib/issueChatTranscriptRuns.ts`** — `resolveIssueChatTranscriptRuns` now caps the streamed run set (live/active runs always kept; most-recent linked runs fill up to 20) so a large run history can't open a live-transcript poll per historical run. - **`ui/src/main.tsx`** — explicit `gcTime` so cross-tab-published cache entries for unobserved resources are collected promptly. - Tests for the batcher, shared ticker, and run cap. ## Verification - `vitest`: new suites `query-invalidation-batcher.test.ts` (batcher collapses 20 invalidations → 1 flush; keeps distinct keys/variants; dispose cancels; proxy passes non-invalidate methods through), `useSecondTick.test.tsx` (single ref-counted timer, stops when idle), `issueChatTranscriptRuns.test.ts` (cap keeps newest + live). All pass. - Existing affected suites pass: `LiveUpdatesProvider` (23), `IssueChatThread` (), `useLiveRunTranscripts`, `AgentDetail.instructions` — 109 tests across affected files. - `tsc -b` clean. - Behavior confirmed by live profiling before the change (2.7 GB / 66 MB tab, ~8 interval churns/sec growing 7× with age). Runtime churn reduction should be re-measured against a rebuilt bundle with the same instrumentation. ## Risks Low-to-moderate; all changes reduce work rather than add features. - **Invalidation batching** delays live-driven refetches by up to ~300 ms. Optimistic `setQueryData` writes (e.g. the visible issue's new comment) remain immediate, so foreground updates still feel instant; only the safety-net refetch is throttled. Non-live invalidations (user actions, mutations) are unaffected — they use the real client. - **Poll back-off** relies on the websocket as the live source when realtime is enabled; a 30 s fallback poll still covers gaps/reconnects (both the transcript hook and `LiveUpdatesProvider` also auto-reconnect). - **Run cap (20)** means an issue with a very large run history streams live transcripts only for its live/active + 20 most-recent runs; older runs still open normally via their run pages. - Downstream test fallout (timing-sensitive tests around invalidation/polling) may need adjustment — flagged intentionally for follow-up. Durable follow-up (out of scope, marked in code): replace transcript/run polling with server push (SSE/websocket deltas) so idle tabs do no periodic work at all. ## Model Used - **Provider:** Anthropic, via the Claude Code CLI. - **Model:** Claude Opus 4.8 (`claude-opus-4-8`). - **Reasoning mode:** Extended thinking enabled. - **Capabilities used:** tool use (shell, file editing), sub-agent fan-out for codebase analysis, and the Chrome DevTools MCP to reproduce and profile the memory/CPU churn on a live instance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work (a bug/perf fix, not planned core feature work) - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [ ] I have updated relevant documentation to reflect my changes (N/A — no user-facing docs; rationale documented inline) - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
88 lines
2.8 KiB
TypeScript
88 lines
2.8 KiB
TypeScript
import { describe, expect, it } from "vitest";
|
|
import type { LiveRunForIssue } from "../api/heartbeats";
|
|
import type { IssueChatLinkedRun } from "./issue-chat-messages";
|
|
import { MAX_ISSUE_CHAT_TRANSCRIPT_RUNS, resolveIssueChatTranscriptRuns } from "./issueChatTranscriptRuns";
|
|
|
|
function linkedRun(n: number, isoDate: string): IssueChatLinkedRun {
|
|
return {
|
|
runId: `run-${n}`,
|
|
status: "succeeded",
|
|
agentId: "agent-1",
|
|
adapterType: "codex_local",
|
|
createdAt: isoDate,
|
|
startedAt: isoDate,
|
|
finishedAt: isoDate,
|
|
hasStoredOutput: true,
|
|
} as IssueChatLinkedRun;
|
|
}
|
|
|
|
describe("resolveIssueChatTranscriptRuns", () => {
|
|
it("uses adapterType from linked runs without requiring agent metadata", () => {
|
|
const runs = resolveIssueChatTranscriptRuns({
|
|
linkedRuns: [
|
|
{
|
|
runId: "run-1",
|
|
status: "succeeded",
|
|
agentId: "agent-1",
|
|
adapterType: "codex_local",
|
|
createdAt: "2026-04-09T12:00:00.000Z",
|
|
startedAt: "2026-04-09T12:00:00.000Z",
|
|
finishedAt: "2026-04-09T12:01:00.000Z",
|
|
hasStoredOutput: true,
|
|
},
|
|
],
|
|
});
|
|
|
|
expect(runs).toEqual([
|
|
{
|
|
id: "run-1",
|
|
status: "succeeded",
|
|
adapterType: "codex_local",
|
|
hasStoredOutput: true,
|
|
},
|
|
]);
|
|
});
|
|
|
|
it("caps linked runs to the limit, keeping the most recent by createdAt", () => {
|
|
// 30 runs, run-0 oldest … run-29 newest.
|
|
const linkedRuns = Array.from({ length: 30 }, (_, i) =>
|
|
linkedRun(i, new Date(Date.UTC(2026, 0, 1, 0, i)).toISOString()),
|
|
);
|
|
|
|
const runs = resolveIssueChatTranscriptRuns({ linkedRuns });
|
|
|
|
expect(runs.length).toBe(MAX_ISSUE_CHAT_TRANSCRIPT_RUNS);
|
|
// Newest run retained, oldest dropped.
|
|
const ids = runs.map((r) => r.id);
|
|
expect(ids).toContain("run-29");
|
|
expect(ids).not.toContain("run-0");
|
|
});
|
|
|
|
it("respects a custom limit", () => {
|
|
const linkedRuns = Array.from({ length: 10 }, (_, i) =>
|
|
linkedRun(i, new Date(Date.UTC(2026, 0, 1, 0, i)).toISOString()),
|
|
);
|
|
|
|
const runs = resolveIssueChatTranscriptRuns({ linkedRuns, limit: 3 });
|
|
|
|
expect(runs.length).toBe(3);
|
|
expect(runs.map((r) => r.id)).toEqual(["run-9", "run-8", "run-7"]);
|
|
});
|
|
|
|
it("always retains live/active runs even beyond the limit", () => {
|
|
const linkedRuns = Array.from({ length: 25 }, (_, i) =>
|
|
linkedRun(i, new Date(Date.UTC(2026, 0, 1, 0, i)).toISOString()),
|
|
);
|
|
const liveRuns = [
|
|
{ id: "live-1", status: "running", adapterType: "claude_local", logBytes: null, lastOutputBytes: 10 },
|
|
] as unknown as LiveRunForIssue[];
|
|
|
|
const runs = resolveIssueChatTranscriptRuns({ linkedRuns, liveRuns, limit: 5 });
|
|
|
|
const ids = runs.map((r) => r.id);
|
|
// The live run is always present; linked runs fill the remaining slots.
|
|
expect(ids).toContain("live-1");
|
|
expect(runs.length).toBe(5);
|
|
});
|
|
});
|