Files
DottaandPaperclip 3b03c4b9eb perf(ui): keep long task chat responsive during streaming (#13229)
## Thinking Path

> - Paperclip lets operators manage agent work through tasks.
> - Task chat keeps responses and run history together.
> - A live update rendered every historical bubble and hidden tool row
again.
> - New image callbacks also forced unchanged markdown to parse again.
> - Long conversations saturated the browser main thread.
> - This change reuses unchanged history and mounts folded tools on
first inspection.
> - Operators can read and reply while work continues.

## Linked Issues or Issue Description

**What happened?**

Chat-style tasks with substantial scrollback became almost unusable. A
deterministic browser reproduction with 200 long responses and 4,000
tools consumed 97.8% of the main thread during live updates. It
delivered only 8 updates during the sample.

**Expected behavior**

The task should remain responsive during streaming. Historical markdown
and unopened run details should not repeat expensive render work.

**Steps to reproduce**

1. Install dependencies with `pnpm install`.
2. Run `pnpm exec playwright test --config
tests/perf/task-chat/playwright.config.ts`.
3. Compare the attached performance JSON. The new test fails against the
original rendering code.

**Paperclip version or commit**

Reproduced at `a05b828bc`. The branch is rebased on current master.

**Deployment mode**

Local Chromium and Vite with deterministic fixtures. No database or
agent credentials are required.

Related work: #10463 reduces the issue-page bundle. This change
addresses repeated rendering after the page loads. No duplicate
scrollback fix was found.

## What Changed

- Keep the bubble image callback stable so unchanged markdown can skip
parsing.
- Memoize the settled history separately from the header and streaming
tail. Keep the brief renderer and default attachment array stable.
- Mount folded tool history on first expansion. Keep it mounted
afterward to preserve child state and closing motion. Runtime request
receipts remain visible.
- Add deterministic rendering tests to the normal Vitest suite and an
opt-in Chromium regression fixture.
- Document the reproduction, commands, scope, and local measurements.

## Verification

- Browser reproduction: 97.8% main-thread utilization before; 11.6%
after for tail-only updates; 31.5% after when projection recreates
history objects. Both fixed cases delivered 32 updates.
- Browser checks pass for scroll-position retention, typing, return to
latest, tool inspection, and retained expansion state.
- Focused component suite: 155 tests passed. Post-rebase
thread/performance rerun: 101 tests passed.
- Recursive typecheck, build, Storybook build, and token gates passed.
UI typecheck passed after the final edits.
- Local UI/CLI lane: 5,910 tests passed; ten files hit worker-start
timeouts, then all ten passed with two workers (20 tests).
Shared/adapter lane: 3,128 tests passed.
- Full local `pnpm test:run` was attempted and is **not green**: its
server lane recorded 10,517 passes and 11 failures plus fixture/setup
errors. Queue (31 tests), Cursor/Git-load (9 tests), and missing-binary
failures cleared on isolated reruns / building the runner test binaries.
Two native suites still cannot initialize embedded PostgreSQL on this
host.
- The remaining native-session recovery assertion was reproduced in a
clean worktree at base `a20ecce40` (1 failed, 67 passed across the
native/queue suites). It expects a settled-session error but receives a
semantic-tool-input digest error. No server or runner files changed in
this PR.
- All substantive CI jobs have passed, including build, typecheck, all
server/workspace test shards, all three e2e shards, and the canary dry
run. The unchanged Slack ordering test exhausted its one-second wait on
the first run; its shard passed on rerun. Final aggregate verification
passed: **31 passing checks**, no failures or pending checks; two
optional Storybook deployment/visual checks were skipped. Greptile is
**5/5**, with no unresolved review threads.
- The supplemental local serialized route run was stopped after CI
passed all five serialized shards; it had reported no failures.
- The browser fixture uses the real chat rendering components with a
plain textarea. It does not test the full composer or server transport.
Timing results are local samples; deterministic render-count tests
provide normal CI coverage.

## Risks

- Closed tool content becomes available to DOM search only after first
expansion. Visible run summaries and runtime request receipts remain
available immediately.
- Opened run history remains mounted to preserve child state. The first
full markdown render still scales with conversation size.
- Memo dependencies must stay current when adding render inputs. Tests
check content edits and replacement gallery callbacks.
- No API, database, or permission changes.

## Model Used

OpenAI GPT-6 (Codex). Exact serving variant and context-window size are
not exposed in this session. Used reasoning, repository tools, code
execution, and browser testing. No sub-agents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass — targeted/UI/workspace
checks pass; the full local server suite has the baseline/host failures
documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 10:15:00 -05:00

57 lines
3.3 KiB
TypeScript

import { expect, test } from "@playwright/test";
for (const reproject of [false, true]) {
test(`long scrollback stays responsive (${reproject ? "reprojected history" : "tail only"})`, async ({ page }, testInfo) => {
const errors: string[] = [];
page.on("pageerror", (error) => errors.push(error.message));
await page.goto("/tests/task-chat-perf.html");
await expect(page.getByTestId("task-chat-scroller")).toBeVisible({ timeout: 90_000 });
await expect(page.locator('[data-thread-anchor^="history-"]')).toHaveCount(200);
if (reproject) await page.getByLabel("Recreate history objects").check();
const session = await page.context().newCDPSession(page);
await session.send("Performance.enable");
const read = async () => Object.fromEntries((await session.send("Performance.getMetrics")).metrics.map(({ name, value }) => [name, value]));
await page.getByRole("button", { name: "Start streaming" }).click();
const before = await read();
await page.waitForTimeout(3000);
const after = await read();
const metrics = {
mainThreadBusyPercent: 100 * (after.TaskDuration - before.TaskDuration) / (after.Timestamp - before.Timestamp),
scriptMs: 1000 * (after.ScriptDuration - before.ScriptDuration),
layoutMs: 1000 * (after.LayoutDuration - before.LayoutDuration),
domNodes: await page.locator("*").count(),
ticks: Number(await page.getByTestId("stream-tick").textContent()),
};
console.log(JSON.stringify(metrics));
await testInfo.attach("performance.json", { body: JSON.stringify(metrics, null, 2), contentType: "application/json" });
// Read scrollback without getting pulled down by ongoing live updates.
const scroller = page.getByTestId("task-chat-scroller");
await scroller.hover();
await page.mouse.wheel(0, -700);
await expect(page.getByRole("button", { name: "Scroll to latest" })).toBeVisible();
const top = await scroller.evaluate((element) => element.scrollTop);
await page.getByRole("textbox", { name: "Reply" }).fill("Reply remains usable during streaming.");
await page.waitForTimeout(300);
expect(await scroller.evaluate((element) => element.scrollTop)).toBeCloseTo(top, 0);
await expect(page.getByRole("textbox", { name: "Reply" })).toHaveValue("Reply remains usable during streaming.");
await page.getByRole("button", { name: "Scroll to latest" }).click();
await expect.poll(() => scroller.evaluate((element) => element.scrollHeight - element.clientHeight - element.scrollTop)).toBeLessThan(2);
await page.getByRole("button", { name: "Stop streaming" }).click();
await expect(page.getByTestId("task-chat-tool-card")).toHaveCount(0);
const summary = page.getByTestId("task-chat-turn-summary").last();
await summary.click();
await expect(summary).toHaveAttribute("aria-expanded", "true");
await expect(page.getByTestId("task-chat-tool-card")).toHaveCount(20);
const tool = page.getByTestId("task-chat-tool-card").first();
await tool.getByRole("button").click();
await expect(tool).toContainText("File inspected successfully.");
await summary.click();
await summary.click();
await expect(tool).toContainText("File inspected successfully.");
expect(errors).toEqual([]);
// A broad regression ceiling, not a machine-specific benchmark target.
expect(metrics.mainThreadBusyPercent).toBeLessThan(50);
expect(metrics.ticks).toBeGreaterThanOrEqual(20);
});
}