## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks collect the files and work products that agents create. > - The Artifacts tab shows these outputs in rows, which makes videos hard to compare. > - Small text links also make artifact rows harder to open. > - This pull request adds image and video tiles with previews and makes the full artifact row clickable. > - Users can compare outputs and open the existing media viewer with one click. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task Artifacts tab and media previews in task chat. **Current behavior** Video outputs appear as file rows or icons. Users must click a small link to open a work product. A task with eight video outputs gives little visual context. **Proposed behavior** Show images and videos in a responsive gallery. Show a paused video frame as the thumbnail. Keep documents, links, and other files in rows whose entire area opens the item. **Reason and benefit** Users can compare generated media without opening each item. Larger click targets also make the sidebar easier to use. **Breaking changes** None. This uses the existing artifact URLs, media viewer, run grouping, and attachment filters. No API or database changes. Related work: #11226 added the task sidebar output surface, and #7361 added rich attachment previews. #3524 concerns a separate reviewed-assets panel. This PR improves the existing task artifact components. The duplicate search found no active PR for this change. This is polish for the shipped Artifacts & Work Products roadmap item. ## What Changed - Add a shared media tile for work products and agent attachments. - Reuse video and image previews in task artifacts and chat. Seek up to one second into videos and reset preview state when the source changes. - Use the existing task gallery for playback and downloads. Preserve grouping and attachment deduplication. - Extend native links and buttons across work-product rows, including keyboard focus indicators. - Add eight offline Storybook examples for video outputs, mixed media, clickable rows, narrow and wide panels, missing previews, empty state, and light mode. - Register the component in the design guide and document its use. ## Verification - 108 focused component tests pass, including thumbnail seeking, source changes, gallery activation, and attachment deduplication. - `pnpm build`, `pnpm -r typecheck`, Storybook build, token gates, and `git diff --check` pass locally. - Reviewed the production components in the embedded browser. Checked all eight video thumbnails, mixed media, narrow layout, light mode, blank-area row clicks, keyboard gallery activation for generic-MIME images, and playback from chat video thumbnails. - Storybook: open **Tasks / Artifact Gallery** and select **Eight Video Outputs**, **Mixed Media And Files**, or **Whole Row Clickable**. The small local clips are synthetic fixtures. - All build, typecheck, unit, runner, and end-to-end CI jobs pass for `b8ace289b5e07df5b9f2c319b159f3923ce427b4`. Greptile gives 5/5 with both review findings resolved. All 54 PR checks pass, including the external security scan. The full test suite passed in CI. The duplicate serial local test run was stopped after CI finished; the 108 focused tests, full build, and recursive typecheck passed locally. ## Risks - Video thumbnails require the browser to load metadata and a frame. A slow server or unsupported codec can leave the fallback visible; opening and downloading still use the existing viewer. - Full-row click targets change pointer interaction with work-product cards. Native link and button semantics remain in place. ## Model Used OpenAI GPT-6 in Codex, with code execution and browser tools. The exact deployment ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
7.3 KiB
Changing Paperclip's UI — a field guide
How to make visual changes now that the design system exists. Written for everyone: designers, engineers, and AI agents (AGENTS.md points here via DESIGN.md).
The one-minute mental model
Paperclip's look lives in three layers, and you almost always work in the first one:
- Tokens — every color, text size, spacing, radius, and shadow is a named value in
ui/src/index.css. Change a token, and every surface using it follows. - Components — consume tokens, never raw values. One component per job (one Button, one Card, one ToggleSwitch).
- Screenshots — 510 baseline images (255 stories × light/dark) downloaded into
tests/storybook-visual/.snapshots/from the pinned external archive intests/storybook-visual/baseline-manifest.json. They are the proof of what the UI looks like. Any visual change shows up as a screenshot diff; no visual change proves itself the same way.
The rules live in DESIGN.md (repo root). The reasoning behind past decisions lives in DECISION-SHEET.md. If you disagree with a rule, change DESIGN.md first (with review) — don't quietly diverge in code.
The three commands
pnpm check:token-gates # am I allowed to write this? (no hardcoded values)
pnpm test:storybook-visual # what does my change look like? (diff vs baseline)
pnpm test:storybook-visual:update # accept my intentional changes as the new baseline
Recipe 1 — change how something looks everywhere
"Make the corners rounder." "That amber is too loud." "Bump the smallest text size."
- Find the token in
ui/src/index.css(they're named and commented:--radius,--status-task-todo,--text-micro, …). - Change the value.
pnpm test:storybook-visual— it will "fail" on every affected story. That's the point: each failure writes a before/actual/diff image triplet intotests/storybook-visual/test-results/. Review them (ornpx playwright show-reportfromtests/storybook-visual/for a browsable version).- Happy?
pnpm test:storybook-visual:update, review the generated bundle undertests/storybook-visual/baseline-review/, publish it from a trusted maintainer environment, then commit the token edit and the updated manifest metadata together.
Notable single-knob tokens: --radius drives the entire corner ladder (sm→4xl are derived); the --status-task-* / --status-agent-* family is the app-wide status vocabulary (chips, charts, bars, live dots all follow it).
Recipe 2 — retheme the whole app
The core palette follows the shadcn token names, so a theme built at ui.shadcn.com/create applies as token values:
cd ui && pnpm dlx shadcn@latest init --preset <CODE> --force --no-reinstall
Then review the git diff and keep only the CSS-variable value changes — the CLI also tries to rewrite components.json, lib/utils.ts, and add dependencies; revert those (it once deleted 240 lines of our utils). Never use shadcn apply --preset (it overwrites component files). After the token diff is clean: Recipe 1 steps 3–4, plus a sanity pass on the Paperclip-specific tiers (agent gradients, WCAG-tuned status hues) for clashes.
Recipe 3 — build or style a component
- Values: tokens only. No hex, no
text-[11px], nop-[13px]. If no token fits, add a token — that's a feature, not a workaround. - Type: use the named ladder —
--text-nano(10px) /--text-micro(11px) / Tailwindtext-xs(12) /--text-compact(13) /text-sm(14). Letter-spacing:--tracking-label/--tracking-eyebrow/--tracking-caps. - Status: anything that means running/idle/paused/error/todo/done/blocked uses the status system (
ui/src/lib/status-colors.tshelpers or--status-*tokens). Liveness is always blue. - Primitives: check
ui/src/components/ui/anddoc/design/COMPONENT-INVENTORY.mdbefore writing a new component. Switches areToggleSwitch; badges/chips route throughbrandChipBadge. - Give it a story. New visual surface = new Storybook story = automatic screenshot coverage forever.
pnpm check:token-gatesbefore you push. If a value genuinely can't be a token (third-party config, canvas fills, intentional one-off decoration on demo pages), it goes on the allowlist with an inline comment saying why.
Recipe 4 — you changed something and snapshots failed
That's the system working. Two cases:
- You meant it → review the diffs (they're your design review), then
pnpm test:storybook-visual:update, publish the packed baseline archive, and commit the manifest update with the change. A PR with visual changes but no baseline-manifest update is incomplete; baseline changes with no explanation are a red flag. - You didn't mean it → you broke something. The diff images show you exactly where. Do not update the baseline to make it green.
Three stories are known to flake under full parallel load (they pass in isolation — see DECISION-SHEET). Re-run a single story with npx playwright test --config tests/storybook-visual/playwright.config.ts -g "<story name>" before assuming a real failure.
For AI-agent sessions
(Running a session as the human? See AGENT-SESSIONS.md — this section is instructions for the agent itself.)
This system was built to be steered by instruction. "Make all running indicators blue" or "collapse these three grays into one" should land as a token edit or a small codemod plus a snapshot diff — not a manual hunt. If a change is mechanical and touches many files, write an idempotent script in scripts/ (see codemod-*.mjs for the pattern) instead of hand-editing. DESIGN.md is loaded via AGENTS.md; follow it exactly, and record consequential choices in DECISION-SHEET.md.
What's deliberately not done yet (don't fix ad hoc)
- Tailwind palette classes (
bg-red-500, ~3,100 sites) — scheduled for a dedicated cluster-by-cluster conversion pass; piecemeal fixes will collide with it. - Hand-rolled cards/pills →
Card/Badge, sidebar agents-section unification — queued as a component-convergence pass with per-site snapshot verification. - ESLint ratchet — will eventually enforce the token rules at lint time; until then
check:token-gatesis the gate.
See DECISION-SHEET.md for the full ledger.
Task artifact galleries
The task Artifacts tab renders images and videos with MediaArtifactCard, including agent attachments that have not been promoted to work products. Its container-responsive grid preserves run grouping and chronological order; documents and other outputs span the full width. RichWorkProductCard supports gallery for media and keeps card/compact rows for other outputs. Rows use a stretched native link or button so titles, thumbnails, metadata, and padding activate the same action, with keyboard focus and normal link modifiers preserved.
ArtifactPreview is shared with the company gallery. Video previews are muted, never autoplay, and seek up to one second into the clip after metadata loads. Changing the source resets preview state. The existing task gallery handles playback and download. Storybook Tasks → Artifact Gallery covers eight generated video outputs, mixed images/files, full-row activation, narrow/expanded panels, empty state, unavailable previews, and light mode using small offline fixtures.