mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is the operator's daily surface: task lists, boards, budgets, agent status — all built on shadcn components and Tailwind > - Visual values (colors, spacing, type sizes, radii) were hardcoded at ~1,600 call sites: the same "small gray label" was 9/10/11px depending on the file, charts disagreed with chips about status colors, two toggle-switch implementations coexisted in two greens, and there was no visual regression coverage > - This made the UI drift-prone and made any restyle a hundreds-of-files project, which discourages design iteration > - This pull request extracts visual values into a single token layer in `ui/src/index.css`, adds a Storybook visual regression suite backed by external immutable baseline archives, and then applies a deliberate retune reviewed change-by-change on screenshot diffs > - The benefit is that Paperclip's look becomes a config surface: retheming is a token edit reviewed as a snapshot diff, drift is blocked by a token gate, and future UI PRs can prove exactly what changed visually without committing hundreds of PNGs ## Linked Issues or Issue Description No existing public issue covers this work (searched "design tokens", "visual regression", "design system" across issues and PRs). Related in spirit: Refs #8982 (theming a hardcoded panel — a one-off instance of the same problem class this PR addresses systematically). **Problem (feature-request form):** UI visual values are hardcoded per call site with no source of truth and no regression coverage; consistency depends on reviewer memory, and restyling requires mass file edits. **Proposed solution (this PR):** a single token layer + enforcement gate + externally stored visual snapshot suite, then an intentional restyle on top of that foundation. ## What Changed - **Token extraction (zero visual change, machine-verified during development):** committed codemods (`scripts/codemod-*.mjs`) moved ~1,600 hardcoded color/type/spacing/radius/shadow/misc values into named tokens in a non-inline `:root` block of `ui/src/index.css`. - **Visual regression suite:** `pnpm test:storybook-visual` covers 255 stories × light/dark = 510 Playwright screenshots at `maxDiffPixels: 0`, plus new primitive-coverage stories and deterministic-render fixes. - **External visual baselines:** committed PNG snapshots were removed. `tests/storybook-visual/baseline-manifest.json` pins an immutable archive URL/hash/size/count, and `scripts/storybook-visual-baseline.mjs` handles `download`, `verify`, `pack`, and trusted maintainer `upload` flows. - **Opt-in visual CI artifacts:** added a `Storybook Visual` workflow that runs on manual dispatch or PRs labeled `storybook-visual`, downloads/verifies the baseline, runs Playwright, and uploads Playwright report/test-result artifacts for review. Normal PR runs do not mutate baseline objects. - **Token gate:** `pnpm check:token-gates` — zero hex literals, zero arbitrary bracket values, zero raw font-sizes in `ui/src/components/**` and `ui/src/pages/**`, with a documented inline allowlist for legitimate opt-outs. - **Theme retune (intentional, snapshot-reviewed):** new base theme values; radius ladder derived from a single `--radius` knob; micro-type cluster collapsed to a named ladder (`--text-nano/micro/compact` + Tailwind `text-xs`/`text-sm`); letter-spacing collapsed to named steps. - **One status-color vocabulary:** charts, quota/budget bar fills, RUNNING/live chips, and liveness indicators all use the canonical `--status-*` hues. Light-mode legibility fixes for red alert surfaces that used dark-tuned text classes. - **One switch:** `ToggleSwitch` restyled to the registry capsule form, second hand-rolled implementation removed, and all call sites unified. - **Docs:** `DESIGN.md` is the design contract; `doc/design/` holds audit reports, decision logs, and updated guidance for external baseline review/update workflows. - Dead code removed (`agentStatusBadge` duplicate map), byte-identical contrast constants consolidated, semantic renames (`--project-seed`/`--project-none`, `--liveness-blue`). ## Verification - `pnpm check:token-gates` — 3/3 gates CLEAN during the design-system run - `pnpm typecheck` && `pnpm --filter @paperclipai/ui build` — green during the design-system run - `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` — pass after external-baseline rework - `pnpm exec tsc --noEmit --pretty false --module NodeNext --moduleResolution NodeNext --target ES2022 --types node,@playwright/test tests/storybook-visual/playwright.config.ts tests/storybook-visual/storybook-visual.spec.ts` — pass after external-baseline rework - `git diff --check origin/pr/9134..HEAD` — pass after external-baseline rework - `find tests/storybook-visual -type f -name '*.png' -print | wc -l` — `0` - `node scripts/storybook-visual-baseline.mjs verify` — intentionally fails closed until the first trusted maintainer publishes the baseline archive and updates `baseline-manifest.json` ## Risks - **Large but shallow:** the PR still touches many UI files due to mechanical token extraction and retune work, but committed PNG snapshot churn has been removed from the branch. - **Baseline publication required before the visual suite can pass in clean clones:** the manifest currently has placeholder archive metadata. A trusted maintainer must publish the first immutable archive, then update `baseline-manifest.json`. - **Rendering platform variance:** the external baseline should be captured in the documented Linux/Chromium environment. Future CI runs verify against the pinned archive and fail closed on checksum/count mismatch. - **Visual CI is opt-in while stabilizing:** add the `storybook-visual` label or dispatch the workflow manually to produce downloadable Playwright report/test-result artifacts. - **Scheduled follow-ups, deliberately out of scope:** Tailwind palette classes map to semantic tokens in a dedicated pass; card/pill component consolidation; ESLint ratchet. Tracked in `doc/design/DECISION-SHEET.md`. ## Model Used Claude Fable 5 (Anthropic, `claude-fable-5`, Mythos-class tier) with extended thinking, running in Claude Code with tool use; mechanical phases delegated to Claude Sonnet subagents. Follow-up external-baseline rework assisted by OpenAI Codex (`gpt-5` coding agent with repository, terminal, and GitHub tool use). All bulk rewrites executed via deterministic, idempotent scripts committed in `scripts/`; intentional visual changes were human-reviewed on screenshot contact sheets. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run targeted local verification and documented the intentional baseline-publication failure above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green *(pending new CI run after this rework)* - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(pending review)* - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) and OpenAI Codex --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
6.4 KiB
6.4 KiB
/goal Prompt — Design Language Simplification, Run 1 (v3)
Paste everything inside the code block below into Claude Code after typing /goal, from inside this worktree. Prerequisites already satisfied on this branch: DESIGN.md, PRIOR-ART.md, KNOWN-DUPLICATES.md at repo root; token-auditor + codemod-runner in .claude/agents/.
v3: condensed under the /goal 4,000-character limit (v2 was 4,648 and got rejected). The paste block now carries only mission + DONE-WHEN + guardrails; the full phase spec lives in the "Phase spec" section below, which the run reads from this file on disk.
Refactor Paperclip's UI so every visual value flows through the single
existing token layer, with provably zero visual change, working only in
this git worktree on branch design/token-extraction. Never touch master
or other working trees. DESIGN.md at the repo root is the source of
truth; follow it exactly. Read PRIOR-ART.md before auditing. Execute
Phases 0-2 exactly as specified in the "Phase spec" section of
GOAL-PROMPT.md at the repo root (Phase 0 external baseline archive,
Phase 1 audit, Phase 2 codemod extraction), delegating Phase 1 to the
token-auditor subagent and Phase 2 to the codemod-runner subagent.
Commit after each phase and in small reviewable steps.
DONE WHEN (all verified in this worktree):
1. The Storybook visual snapshot suite passes against a Phase 0
baseline that was captured BEFORE any component change and pinned in
tests/storybook-visual/baseline-manifest.json — zero visual change.
Baseline scope: primitives in ui/src/components/ui/
(add minimal stories only for those) plus the existing stories in
ui/storybook/stories/. No stories for the ~277 feature components.
2. ui/src/index.css (plus any tokens.css it imports) is the only token
source and components consume visual values only through it; no
parallel token source exists; runtime-tunable tokens live in a
NON-inline block.
3. rg gates pass: zero hex color literals, zero arbitrary px/bracket
Tailwind values, zero raw font-size declarations in
ui/src/components/** and ui/src/pages/** outside the documented
allowlist in the token source.
4. TOKEN-AUDIT.md and COMPONENT-INVENTORY.md exist at the repo root,
are current, and each contains a "Needs human decision" section.
5. pnpm build, pnpm typecheck, and pnpm build-storybook all exit 0.
GUARDRAILS
- Preserve rendered output exactly. Reuse an existing token only on
EXACT value match; otherwise mint a new token with the value
VERBATIM — no normalizing, rounding, or inventing a scale. If a
replacement cannot be made without visual change, skip it and log it
under "Needs human decision".
- All value rewrites happen via codemod scripts committed to scripts/,
never file-by-file hand edits.
- No redesign, no layout changes, no new colors/typefaces, no component
merges or deletions (consolidation and shadcn swaps are
recommendations-only in COMPONENT-INVENTORY.md), no copy renames
(issue->task is a separate later run), no new dependencies beyond
snapshot tooling, no server or app-logic changes.
- If reality conflicts with DESIGN.md, record the conflict in
TOKEN-AUDIT.md instead of guessing. If a phase cannot be completed,
stop and report rather than partially applying it.
Phase spec (referenced by the goal — the run reads this from disk)
Phase 0 — Baseline (before changing ANY component):
- Set up Storybook visual snapshot testing (Storybook test-runner with image snapshots, or equivalent already-compatible tooling; Storybook lives at
ui/storybook/, launched viapnpm storybook). - Coverage scope: the shared primitives in
ui/src/components/ui/(add a minimal story for any of the ~24 that lack one) plus all existing stories underui/storybook/stories/. Do NOT write stories for the ~277 feature components in this run. - Pack and publish the passing baseline snapshots through the external Storybook visual baseline flow, then commit the manifest metadata. Every later phase must keep snapshots matching this baseline.
Phase 1 — Audit (no code changes; delegate to token-auditor):
- Produce
TOKEN-AUDIT.mdat the repo root: every hardcoded color/spacing/radius/type/shadow value inui/src/, its frequency, file locations, and near-duplicate clusters (e.g. 13/14/15px used interchangeably). Flag clusters for human review — do NOT merge them. Cross-reference the ~80 existing tokens inui/src/index.css: for each hardcoded value, note whether it exactly matches an existing token. - Produce
COMPONENT-INVENTORY.md: all components, their variants, and suspected duplicates with evidence (similar props, similar rendered output, copy-pasted origins). Include a "shadcn candidates" section: (a) custom components duplicating an available shadcn primitive, (b) installed shadcn components drifted from the registry (npx shadcn@latest diffwhere available), (c) raw Radix/plain elements where an installed shadcn wrapper exists. For each, state the recommended replacement and expected visual impact. ALL consolidation and swap items are RECOMMENDATIONS ONLY.
Phase 2 — Extraction (mechanical, via codemod; delegate to codemod-runner):
- Token destination is
ui/src/index.css(Tailwind v4; optionally atokens.cssimported by index.css). Do NOT create a parallel token source. Tokens that must be runtime-tunable go in a NON-inline block (@theme inlinebakes literals). - Exact-match values → existing token reference; everything else → new verbatim token. Ugly values stay ugly; they are the audit.
- Codemod scripts committed to
scripts/perform the replacements; run them; no hand-edits. - Third-party overrides that cannot use tokens go on a documented allowlist in the token source, each with an inline comment saying why.
After the run (human steps)
- Eyeball pass:
pnpm storybookhere and in the master tree (-p 6007), flip between tabs. - Read TOKEN-AUDIT.md; choose the real spacing/radius scale (PRIOR-ART.md has drafted rules to start from).
- Tune tokens; snapshots now fail intentionally — the diff folders are your design-review contact sheet. This is also where the ui.shadcn.com/create preset lands, as token-value edits.
- Review COMPONENT-INVENTORY.md; approve a merge list and shadcn-swap list → those become Run 2 and Run 3, each its own /goal.
- Merge to master when satisfied (rebase first if master moved). Scrap path:
git worktree remove ../paperclip-design-simplify --force.