mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 01:57:33 +02:00
master
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7498705642 |
fix(ui): stabilize mobile task reading and document navigation (#15228)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People read tasks and send instructions from phones as well as desktop browsers. > - The mobile footer let page text show through its labels, and small text fields made Safari zoom on focus. > - Small scroll changes made the footer switch direction, and its changing page padding moved the conversation. > - Task pages also showed comments before question cards and run history arrived, so the composer and reading position moved again. > - Document links also used native navigation, which reset the reading position or reloaded a task through its UUID URL. Desktop tabs were crowded in the mobile drawer. > - This pull request keeps navigation steady, opens documents in the mounted task, and gives mobile readers a full-height panel with a vertical tab selector. > - The benefit is a stable task view and smoother scrolling on mobile. ## Linked Issues or Issue Description **What happened?** The mobile footer was translucent. Safari zoomed when a person focused a small text field. The footer switched abruptly while scrolling. A large task could show saved comments, then move the page again when a question card or run history arrived. In a local test with delayed responses, a late question moved the mobile composer by about 374 pixels. Opening a plan from the feed could reset the view or reload the task through a UUID link. The mobile document drawer left part of the feed exposed above small desktop tab controls. **Expected behavior** The footer has an opaque surface and moves smoothly after deliberate scrolling. Text fields do not cause automatic focus zoom. A task shows its initial conversation and composer together at the final scroll position. Background refreshes keep the existing conversation visible. Document links open in the mounted task with its existing cache and reading position. Mobile documents fill the viewport, show a clear close button, and offer a vertical list of open tabs. **Steps to reproduce** 1. Open a task with many long comments and a pending question in iOS Safari. 2. Delay its interactions, activity, and runs responses by different amounts. 3. Reload the page and watch the conversation and composer move as each response arrives. 4. Scroll down and back up, including small direction changes and edge bounce. 5. Focus the task composer, search field, and new-task title and description. 6. Open a plan or another task document from the feed, including a link that uses the task UUID. Close the panel and check the reading position. 7. Open several documents on a phone. Switch tabs and close both active and inactive tabs. **Paperclip version or commit** Developed from `1c07b5903` and rebased onto `59015846a`. **Deployment mode** Built from source. Tested in an isolated local test drive with iOS 26.5 Simulator Safari and Chrome. A temporary local proxy delayed independent responses for the layout test. Related work found in the duplicate search: - Refs #14727. It made saved replies appear before supporting history. This PR keeps its parallel requests and narrows the tradeoff in favor of a stable first layout. - Refs #14667. This open PR takes a different approach with per-run placeholders and retries. This PR fixes the observed question/composer movement and mobile navigation behavior. - Refs #13095 and #13597. These earlier fixes added task scroll anchors and skipped transcript waits for scheduled retries. - Refs #6550. Earlier mobile board polish. - Refs #9467. This related open PR changes list and generic tab reflow. The task-pane selector uses a separate component. ## What Changed - Give the mobile footer an opaque semantic surface. - Set a base-size floor for editable text on touch devices to prevent Safari focus zoom. Preserve larger title text. - Share mobile scroll tracking between both layouts. Accumulate scroll distance, ignore edge bounce and changed document bounds, and update once per frame. - Use shared motion tokens for the footer and composer. Keep page padding stable and honor reduced motion. - Wait for the initial question cards, attachments, work products, activity, runtime selection, plan, and relevant transcript history before the first reveal. Skip scheduled retries and older runs outside the initial comment window. - Bound the first reveal to 15 seconds. A stalled supporting request leaves saved conversation and the composer accessible with an explicit loading notice. - Keep concealed mobile history from stretching the document. Keep the composer mounted but concealed until the same reveal. Keep both visible during later refreshes. - Route first and repeated same-task document clicks in place. Recognize UUID and identifier links. Preserve the thread history entry and feed position. Keep modifier clicks, downloads, external links, and classic document behavior. - Give the mobile task panel the full viewport and safe-area padding. Use a visible X and 44-pixel touch controls. Replace the horizontal tab strip with a vertical selector that wraps titles and supports keyboard focus. - Add four interactive Storybook states for a few tabs, long names, many tabs, and the last tab. Reuse the production selector and tab controller. - Add navigation and tab regressions, update first-reveal regressions, and document the behavior in `DESIGN.md`. ## Verification - 392 tests passed across the seven focused task-loading, scroll, mobile-navigation, layout, and composer suites. After review fixes, all 339 tests across the four affected suites passed, including stalled-loading fallback on mobile and desktop and the motion-token catalog. - All 442 focused document, tab, task-thread, and scroll tests pass. The final click-propagation cleanup also passes all 136 task-detail tests. UI typecheck, UI production build, and `pnpm check:token-gates` passed. - All four cases in `artifact-tab-arrival.spec.ts` and `text-attachment-tabs.spec.ts` pass locally, covering desktop and mobile selection, composer focus, document rendering, and downloads of the original bytes. - `pnpm --filter @paperclipai/ui build-storybook` passed. Open the mobile tab stories under `Prototypes/Task detail/Mobile tabs`. - A local diagnostic proxy measured cached plan content at about 250 ms after the first click. The HTML load count and task request count did not change. Feed scroll stayed at the same position. First, repeated, and UUID document links were tested at phone and desktop widths. Task-reference links close their preview before the document reader opens. - In Chrome at desktop and phone widths, the delayed-response task showed one complete reveal. The late question no longer moved an already visible composer. - In iOS Simulator Safari, verified the large-task reload, opaque footer, navigation hide/reveal, and search/new-task/composer focus without automatic zoom. - Full workspace typecheck and build passed. The updated UI also passes typecheck, production build, and token gates. - All 54 checks pass on the final commit `bf5c9914e61833e7cc8794a69d72cf8c7057b952` (two additional checks are intentionally skipped). Greptile reviewed that commit at 5/5, and all review threads are resolved. - Full local `pnpm test:run` was attempted but stopped after server-fixture failures. Embedded PostgreSQL startup failure reproduced in an isolated native-interaction fixture after five startup attempts. The broad run also reported a rapid Slack callback ordering test failure. These server paths are unchanged by this PR, and their CI shards pass on the latest head. The full local suite is not claimed as passing. - A localhost proxy stalled the activity response for 30 seconds. Chrome revealed the available conversation after the 15-second deadline at both desktop and phone widths, kept the composer accessible, and cleared the loading notice when the response arrived. ## Risks Slow initial history requests can delay the first conversation reveal by up to 15 seconds. If that deadline expires, late data can change the available conversation while a loading notice remains visible. The reveal waits only for runs in the initial comment window, and later refreshes do not conceal an existing conversation. The larger editable text can change line wrapping on phones. Mobile navigation and composer motion use shared tokens and respect reduced-motion settings. Mobile tab selection changes the control layout. Document links retain URL history while sharing the task reading position; other tasks and external links keep their normal navigation behavior. I checked `ROADMAP.md`. This is a fix for existing UI behavior. ## Model Used OpenAI GPT-6 in Codex. The runtime does not expose a more specific model ID or context-window size. The agent used reasoning, code editing, terminal tools, and Chrome and iOS Simulator testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
59015846ae |
fix(chat): keep dismissed task questions in the feed (#15229)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask users questions in task chats and Agent Chat. > - Agent Chat keeps unanswered questions as compact entries in the feed. > - Regular task chats still show a pending composer badge after dismissal. > - A page reload can also open a dismissed question form again. > - This pull request applies the same feed behavior to both chat types and saves dismissal per person and task. > - Users can continue the chat and return to the original question later. ## Linked Issues or Issue Description **What happened?** Dismissing a question in a regular task chat leaves a pending composer badge. The question can return after a page reload. The earlier feed behavior only applied to Agent Chat. **Expected behavior** A dismissed question stays in the feed. Its form and pending badge leave the composer. Reload preserves dismissal. Opening the feed entry restores the original question and draft answer. **Steps to reproduce** 1. Open a regular task chat with a pending question. 2. Select an option, then dismiss the form. 3. Reload the page. Check that the composer stays clear. 4. Open the question in the feed. Check that the draft is restored. 5. Submit the answer. Check that the answered receipt appears. **Paperclip version or commit** Reproduced on master at `a386a599983519eb1d399f8b770bfccdb2a74762`. **Deployment mode** Browser UI in local and authenticated instances. This change does not depend on the agent adapter. Related work: Refs #14613, which added the Agent Chat feed behavior. Refs #9141, which validates real answers on the server. Refs #11434, which tracks comment-driven changes to interaction state. This PR changes question presentation only. ## What Changed - Show compact unanswered question entries in regular task chats. - Exclude durable questions from composer pending counts and navigation. - Save dismissal in local storage per person and task. Merge the latest saved IDs so dismissals from another tab survive reload. A new question can still open its form. - Keep the original question pending and answerable. Keep approval and permission controls. - Run the question-history regressions in both chat modes. Add reload, new-question, user/task scope, and stale-tab coverage. - Share the real-component Storybook fixture. Add a regular task test drive and an interactive dismissal/answer scenario. - Update the planning-mode browser test to check a dismissed question in the feed after reload and on mobile. - Update the design rules, implementation spec, and preview instructions. ## Verification - 329 focused thread, composer, and interaction-card tests pass. - `pnpm -r typecheck` passes. The UI typecheck also passes after the review fix. - `pnpm build` passes for the full repository. - UI build, Storybook build, and token gates pass. The UI build and token gates were rerun after the review fix. - Greptile gives final commit `98f71f221` a 5/5 score. The current-head check passes, and there are no unresolved review threads. - All 56 current-head checks are terminal: 54 pass and two conditional Storybook jobs skip. The CI run includes the full test shards, build, typecheck, browser tests, aggregate verification gate, and package canary. - The stale-tab regression fails in both chat modes before the review fix and passes after it. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/planning-mode-visual-verification.spec.ts` passes against a throwaway local server. It checks dismissal, reload, task navigation, and desktop/mobile planning controls. - The duplicate local `pnpm test:run` attempt was stopped after the full CI test suites passed. It did not complete locally. - Manual browser test: select Green, dismiss, reload, reopen, submit the saved answer, and inspect the answered receipt. Also send a new message while the unanswered question remains in the feed. The test drive uses real UI components with fixture response callbacks. - To repeat the browser test, run `pnpm storybook`. Open **Chat & Comments → Task Chat Unanswered Questions → Test Drive**. ## Risks - Dismissal is a browser-local preference. It does not sync to another browser or device. Clearing local storage removes it. - When local storage is unavailable, dismissal lasts for the mounted thread only. - Unanswered questions can accumulate in the feed. They stay pending until answered or resolved through the existing API. - No database migration, API change, or change to approval permissions. ## Model Used OpenAI Codex, `gpt-6.1-sol`, with xhigh reasoning, repository editing, code execution, and browser control. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
144083fd48 |
fix(interactions): wait for workspace readiness before enabling approval (#14893)
## Thinking Path > - Paperclip lets people manage AI agents and review their work. > - Task confirmations must use the work produced by their source run. > - The server blocks approval while that run still needs to sync its workspace. > - The card currently enables approval before that check can pass, so an ordinary click produces an error. > - This PR exposes the existing readiness check and shows “Preparing approval…” with acceptance disabled. > - The card refreshes itself and enables approval when the source workspace settles. ## Linked Issues or Issue Description **What happened?** A confirmation appears while its source run is still preparing or syncing its workspace. Its enabled approval button returns a conflict asking the user to retry after syncing. **Steps to reproduce** 1. Run an agent in an isolated workspace. 2. Have it create a confirmation before workspace finalization completes. 3. Click the approval button while the source workspace is still active. **Expected behavior** The card explains that approval is preparing. Acceptance becomes available automatically when the same server check permits it. Reject and revise remain available. Related work: #10770 handles this conflict after a click with retries. #9520 proposes changing the workspace acceptance barrier. This PR preserves that barrier and exposes readiness before the click, including in compact task chat. It preserves the terminal-finalize behavior from #10099. ## What Changed - Add an optional, read-only `acceptanceBlocker` to interaction responses. Readiness uses the existing source-run workspace predicate, with one check per pending source run. - Disable acceptance and show a shared preparation notice in classic and compact confirmation cards, including checkbox and secret-binding confirmations. - Refresh preparing cards every two seconds in task detail, attention, pipelines, and Skill Studio. Restore each surface's previous polling cadence when preparation clears. - Preserve live tool reviews, questions, rejection, revision, and the server acceptance barrier. No automatic acceptance occurs. - Document the preparation state and cover readiness, terminal sync outcomes, unrelated runs, historical cards, and automatic refresh. ## Verification - Focused service, card, query-refresh, and helper tests: 217 passed across five files. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm build-storybook`: passed. - `pnpm check:token-gates`: passed. - `pnpm test:run`: incomplete locally. Stopped after about 17 minutes once it reproduced seven existing environment failures: two Slack tests and two email tests lack ancestor-directory skill fixtures; three company-skills tests fail on macOS runtime-cache staging permissions. These are outside this change. The full CI test matrix passed. - CI: all 53 checks passed; two optional Storybook jobs were skipped. The branch has no conflicts with `master`. - Greptile: 5/5 on commit `8a6f216d8d`, with no review threads. - Reviewed added lines and new files for credentials, private URLs, internal task references, user paths, and run artifacts. None found. ## Risks - No database migration or change to acceptance authorization. Readiness is advisory; the server still enforces its existing gate at acceptance. - An open preparing card adds a read every two seconds. This cadence stops after readiness clears; historical cards add no workspace checks. - Failed or stale finalization retains the existing server behavior. This PR does not change recovery policy. ## Model Used OpenAI GPT-6 through Codex. The exact serving model ID and context-window size are not exposed in this session. Used reasoning, repository inspection, tool execution, and automated tests. No sub-agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (217 focused tests; full local-suite limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b2bd6563d |
fix(chat): hide obsolete execution status and system notices (#14874)
## Thinking Path > - Paperclip lets people oversee agent work through task conversations. > - Conversations should show failures and waits that still affect the task. > - Old run errors and recovery notices remained visible after later work or task completion. > - These messages looked current even when no action remained. > - This change hides obsolete execution status while preserving the responses and activity. > - The full diagnostic record stays available in run history. ## Linked Issues or Issue Description **What happened?** Task chat kept showing “Run failed”, “Stopped”, and “Waiting to resume” after the execution had been superseded or the task had finished. Stored system notices also remained in the conversation. Completed tasks disabled Retry but kept the error message. **Expected behavior** Hide system status that no longer applies. Keep the latest unresolved failure, active recovery holds, and useful recovery actions visible. Preserve messages, files, questions, session boundaries, and inspectable activity. **Steps to reproduce** 1. Let a task run fail, then record a pre-start recovery wait. 2. Complete a later attempt or mark the task done. 3. Open the task conversation. Before this change, the old error and wait remain visible. **Paperclip version or commit** Reproduced in component tests against `0f9e9be408`. **Deployment mode** Task chat in local or hosted deployments; both legacy adapters and the native runner. Related: #14857 and #14869 address execution recovery. This PR addresses the remaining conversation presentation. Searched GitHub for historical chat status, historical errors, and “Waiting to resume”; no duplicate PR found. ## What Changed - Determine status relevance from task state, attempt order, successor evidence, and recovery state. Time alone does not hide errors. - Hide obsolete run markers and stored execution notices. Require run or recovery provenance, so unrelated system updates such as child-task blockers remain visible. Keep errors from the latest failed attempt actionable. - Keep unresolved execution holds visible. A refused pre-start retry does not replace a real attempt, and another agent’s work does not resolve a run-specific error. - Show historical activity without Worked/Stopped labels. Keep historical failures out of the current turn’s summary. - Anchor activity to visible comments so removing a notice cannot remove the response or activity with it. - Document the presentation rules and cover both runner modes and both task presentation modes. ## Verification - 334 focused component and status-policy tests passed across four files, including the child-task relay regressions. - `pnpm build` passed. The UI build also passed after the final presentation changes. - `pnpm exec vitest run --project @paperclipai/ui`: 667 files and 7,157 tests passed. Subsequent focused tests cover the final activity-anchor, live-successor, and notice-provenance changes. - `pnpm -r typecheck` passed. UI typecheck and build passed again after the review fix. - `pnpm test:run` completed its general-server phase with 14,716 tests passed, 17 failed, and 87 skipped; it stopped before later phases. The failures occurred in four unchanged server suites: chat channels, email channels, company skills, and runtime skill cache. A targeted rerun reproduced missing bundled skill paths and `EACCES` during cache-directory rename on macOS. All Linux CI suites pass for the final commit, including these server suites. - Design token gates and diff checks pass. - All 53 checks pass on commit `7af9753859`; two optional Storybook checks are skipped. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/36931968751) includes build, full typecheck, all server and workspace test shards, all eight end-to-end shards, runner verification, and the canary dry run. - Greptile scores the final commit at 5/5. No review threads remain unresolved. The branch is current with `master` and has no merge conflicts. ## Risks This changes presentation only. It does not change execution, recovery, stored comments, or run history. The main risk is hiding a current diagnostic too early. Tests cover active holds, refused retries, different agents, missing timestamps and provenance, live successors, preserved responses, and the current retry target. The base branch has a dependency override/lockfile mismatch. Local installation used the same resolution fallback as CI, then restored the tracked lockfile. No dependency changes are included. ## Model Used OpenAI Codex (GPT-6). The exact runtime model identifier and context window are not exposed in this session. Used reasoning, repository inspection, code execution, and regression tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — changed-code tests pass; unrelated full-suite failures are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3c561642b4 |
fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask for decisions and optional details through cards in chat. > - A clear approval in a message can leave the matching card pending. > - An unanswered question can also block an unrelated later reply. > - Decisions need a saved source message, while optional questions need to remain answerable in history. > - This pull request records conversational decisions and lets users move on from questions and answer them later. ## Linked Issues or Issue Description **What happened?** Native Claude and Codex could act on approval in chat while the original approval card stayed pending. Pending question forms stayed above the composer, were absent from history, and could suppress later chat replies. A late native question answer could wait for a finished run to reconnect. **Expected behavior** The active agent records a clear approval or refusal against the exact card and user message. Ambiguous replies do not grant consent. Users can send another message without answering a question. The question remains pending in history and can be reopened and answered later. The saved answer reaches the agent. **Steps to reproduce** 1. Ask an agent to propose work with a confirmation card, then approve it in chat. 2. Check that the original card records that approval before work starts. 3. Ask an interactive question, send an unrelated message, and reload. 4. Open the unanswered question from history and submit an answer. Related work: #14408 added completion delivery. #14607 tests completion reporting turns. Neither records conversational answers on approval cards. ## What Changed - Add a confirmation endpoint backed by a user comment, with schema validation, OpenAPI discovery, and native Plan-mode access. Ask mode remains read-only. - Check company, active run, actor, current session, message provenance, revision, and resolver policy. Save the decision and audit in one transaction. Retries do not repeat effects. Emit resolution telemetry after commit. - Give fresh and resumed chat turns the actual pending confirmation identities. Teach agents to save clear conversational decisions before acting and to clarify ambiguity. - Keep unanswered Agent Chat questions as compact history entries. A newer user message closes the old form. Question cards never contribute to composer pending counts or navigation, including after dismissing a fresh form. The history card is the sole reminder; clicking it restores that exact form and draft. - Preserve Agent Chat questions when later messages or questions arrive. Historical ordinary inputs no longer gate later chat replies. Current-run requests, task execution, and governed approvals keep their gates. Remove the special acknowledgement-publication proof helpers that this rule replaces. - Route answers to finished native runs through durable fresh-wake delivery, with existing idempotency and source-question context. Settle late replies against contiguous completed conversation turns and freeze their history replay; failed, unhandled, and newly arriving messages remain actionable. - Add real-component Storybook scenarios, database and UI regressions, and a three-turn native Claude/Codex E2E case. Capture distinct, UI-ready screenshots and report the individual assertions. ## Verification - Focused decision/publication/UI regressions after merging master: 288 passed; subsequent UI draft, failed-send, and conversation checks: 199 passed. - Native question and durable delivery regressions: 106 passed, including all four terminal run states and exactly-once late delivery. Seven targeted regressions fail against the original implementation and pass with the fix. - Latest conversation/decision/native-delivery regressions after the master merge: 121 passed. Covers completed progress, missing or failed intervening turns, new messages during a late reply, stale sessions, and frozen retry/replay boundaries. Four new assertions fail before the ordering fix. - E2E support suite after the master merge: 792 passed. Negative controls reject expired cards, wrong questions/answers, stale or missing replies, unrelated clarification forms, and unexpected tasks. - The embedded-browser walkthrough caught one additional defect: dismissing a fresh question still showed a composer badge. Both Cancel and close-button regressions failed before the fix. The fix at `65f2ade12` passes 170 chat-thread tests and 792 E2E support tests. After merging master, 232 chat-thread/confirmation tests, server/UI typechecks, and token gates pass. The preview and two-provider live E2E pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at that commit. All 55 checks are now successful at `e5512a206` (four conditional checks skipped), including the aggregate verification gate and clean-install canary test. The first attempt was interrupted by simultaneous CI worker shutdowns; one failed-job rerun passed without code changes. - [Published Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on): nine real-component scenarios. Manually exercised move on, reopen, preserve draft, answer later, answer one of multiple questions, and a custom mobile answer in the embedded browser. Retested fresh Cancel and close-button dismissal in the updated build, then reopened and submitted the preserved Green selection and inspected its answered receipt. Static preview has no live model/backend; its callbacks are fixture responses. - [First live campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/) reproduced the late-answer completion-state defect on both providers despite correct saved answers and acknowledgements. It also exposed a valid imperative clarification rejected by the old oracle. Both issues are fixed with regression controls; this failing run is retained as evidence. - [Four-cell qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/) passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous confirmation, each on native Claude and Codex. Inspected saved state, source-message decisions, visible cards, and agent replies. Both late-answer chats settled to waiting; no unrequested tasks were created. [Final branch rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/) passed 2/2 at `142630720`: the same unanswered-question journey after merging master, plus an additional screenshot and browser assertion for the actual late-answer acknowledgement. - [Composer-reminder E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/) passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each, with explicit no-badge assertions before and after reload. Inspected saved pending/answered state, both screenshots with a clear composer, and actual Blue acknowledgements; all five behavioral matchers passed per provider and neither created tasks. Cost coverage is partial; this is bounded workflow qualification. - [Fresh-dismissal E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/) passed 2/2 at `e5512a206`: native Claude and Codex, including fresh Cancel, clear composer, reopen, unrelated message, reload, late Blue answer, and actual agent acknowledgement. All five behavioral matchers pass per provider. Inspected the fresh-dismissal screenshots and saved pending/answered identity; neither created tasks. Cost coverage is partial (4/6 runs). - Prior evidence remains available in [the earlier campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/). Its early loading screenshot and overwritten final capture prompted the UI-ready, distinct screenshot fixes. ## Risks - The model interprets intent. The server verifies permission and provenance; it does not infer consent from text. Ambiguous and unrelated replies are not approvals. - Historical questions can accumulate. They remain visible, pending, and answerable; no automatic answer or expiry is invented. - The change to completion gates is scoped to Agent Chat and ordinary historical inputs. Current-turn and governed approvals retain their existing controls. - Live qualification is limited to the selected stories. Broader native onboarding finalization remains separate work. - No database migration. Telemetry adds no fields or values; the contract and README document the commit boundary. Privacy review was requested on the PR. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser-test orchestration. The exact model ID and context-window size are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0be2afcca6 |
feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path > - Paperclip lets operators assign tasks to AI agents and review their work. > - The task composer controls the next message and its assigned agent. > - Operators needed a way to choose that agent's model and effort without leaving the composer. > - The old mode selector, upload button, and input cards made the mobile composer crowded and hid normal messaging during a pending decision. > - Harnesses publish different model and effort capabilities, so the picker must follow the selected agent. > - This pull request adds one responsive composer flow, keeps pending cards visible above it, and protects Codex ACP authentication in the local test path. > - Operators can choose run settings, send a message, and answer a pending card as separate actions. ## Linked Issues or Issue Description **Subsystem affected** Task composer UI, issue thread interactions, Codex ACP credential handling, and Storybook. **Problem or motivation** The composer did not expose model or effort for the selected agent. Mobile actions wrapped poorly. Pending questions and confirmations replaced the composer. A local Codex ACP test could also reuse host authentication after the managed key was removed. **Proposed solution** Put assignee search, model search, exact model IDs, effort, and fast mode in one picker. Use a mobile dialog. Replace the direct-upload plus action and separate mode selector with an Add menu and removable Plan or Ask chips. Place pending interaction cards above the usable composer. Keep these cards pending after an ordinary message unless their creator asks for comment superseding. Replace managed ACP auth files atomically and isolate the test key from host credentials. **Roadmap alignment** ROADMAP.md does not list an overlapping composer milestone. This change improves the existing task and review flows. ## What Changed - Added the combined assignee, model, and effort picker to both task composers. Search matches agent name, role, and harness. The server uses a curated Codex list by default and honors instance-declared models. Manual IDs remain available. - Added an effort slider for known model capabilities, a conditional Codex fast control, and reset. The picker opens in a modal on mobile. - Added the Add menu for files, supported goals, Plan mode, and Ask mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling remains available. - Adjusted mobile spacing, avatars, wrapping, and Send placement. Removed the composer divider. - Moved pending question, confirmation, review, and related cards above the composer. Ordinary comments now leave question and confirmation cards pending by default. The onboarding prompt retains explicit comment superseding. - Updated the Storybook composer group with responsive states and the production picker. Added UI, service, route, and browser regression coverage. - Isolated Codex ACP API-key authentication, skipped subscription auth merge and shared-home copy-back for remote API-key runs, and replaced the managed auth file atomically. ## Verification - `pnpm -r typecheck` — passed on the final local head. - `pnpm check:token-gates` — passed on the final local head. - `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31 tests passed, including role and harness search, declared Codex models, and filtering general OpenAI models. - `pnpm exec vitest run server/src/__tests__/issue-thread-interactions-service.test.ts` — 74 tests passed. - `pnpm exec vitest run packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed, including remote API-key copy-back isolation. - `pnpm test:run` — attempted locally; the embedded PostgreSQL test database could not initialize on macOS. The isolated `heartbeat-run-event-sequencing` suite reproduced that environment failure. GitHub CI runs the full test matrix for this head. - `pnpm build` — passed on the final head. `pnpm build-storybook` passed after the last UI change; only server code, tests, and docs changed afterward. - Live local test drive — Codex ACP ran a task with a managed API key. The test agent was restored to its default ACP configuration afterward. - Review the interactive stories under the top-level Composer group with `pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask, picker, and pending-question states. ## Risks - A pending card stays open when an ordinary comment changes the discussion. Its creator can set `supersedeOnUserComment: true` when a new comment should replace it. - Model and effort overrides persist on the task until reset or changed. An unlisted manual model ID may fail when the provider runs it. - Some harness catalogs do not report effort support. The picker hides effort for those models. - No database migration is required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID or context window to the task. The model used code execution and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI Codex <codex@openai.com> |
||
|
|
fd071748ee |
fix: stop repeated notifications for finalized run failures (#13739)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native run reconciliation repairs saved execution outcomes after
interruptions.
> - The sweep also visits runs whose final results are already
committed.
> - An unchanged failed run still received a new status delivery ID on
every sweep.
> - The browser treated each delivery as a new failure after its short
duplicate window expired.
> - This pull request makes unchanged projections a no-op and suppresses
repeated or historical run toasts.
> - Operators receive fresh failure alerts without repeated alerts for
old work.
## Linked Issues or Issue Description
**What happened?**
An old failed run repeatedly produced failure toasts while the browser
remained open. The task could already be cancelled. Reconciliation
rewrote the same failed outcome and queued another status broadcast.
**Expected behavior**
An unchanged committed run must not queue a new status notification.
Repeated deliveries must still refresh cached state without another
toast.
**Steps to reproduce**
1. Finalize a native run with a failed result and successful workspace
finalization.
2. Deliver its pending execution status and cancel its task.
3. Replay finalization and status delivery on each periodic sweep.
4. Observe another failure broadcast for every sweep before this fix.
**Paperclip version or commit**
Reproduced against source commit
|
||
|
|
924f07be8c |
feat(chat): simplify Slack onboarding and account linking (#13638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connections let people start and continue that work from Slack. > - Setup mixed app creation, credentials, URL verification, account linking, and testing on the same screens. > - People also needed a safe way to link their own Slack identity after the first operator finished setup. > - This pull request gives each step a clear place and keeps membership approval separate from identity linking. > - It also makes connection details easier to use and fixes misleading callback health behind HTTPS proxies. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: chat routes and services, shared contracts, and the Apps board UI. **Problem or motivation** Slack onboarding made users find settings without enough guidance. A second user needed operator help to link their account. Activity stopped at 100 records, and TLS termination could mark working callbacks as stale. **Proposed solution** Use six setup steps with editable app names, a generated manifest, credential guidance, URL verification, account linking, and an optional message test. Send each Slack user a private, expiring confirmation link. Require company membership or an approved access request before linking. Add cursor pagination and tolerate the internal HTTP hop in callback diagnostics. **Roadmap alignment** This improves the existing connected-app surface and supports CEO Chat without changing the task-and-comments model. The maintainer requested and reviewed the flow during a live Slack test drive. **Additional context** Related work: #7, #3349, #13000, and #13620. Those cover broader chat capabilities, older webhook paths, or plugins. This PR improves the existing native connector's setup and account-linking flow. HTTPS documentation was published separately in paperclipai/paperclip-docs#128. ## What Changed - Split Slack onboarding into six clickable sidebar steps. Keep secondary and primary actions on one row. - Generate the Slack creation link and read-only manifest from editable app, bot, and command names. Add credential prefix validation and direct instructions. - Add live account-link status and an optional mention-based message test. - Add private, single-use Slack account invitations and membership access requests. Retain cloud authentication/bootstrap checks and enforce the chat rollout flag in all identity APIs. Default new Slack connections to linked users only. - Put Settings, Access, Conversations, and Activity in the sidebar. Simplify conversation rows and remove active header badges. - Add 25-item activity pages, stable timestamp/ID cursors, and replay safety across pages. Preserve the legacy array API for clients without pagination parameters. - Fix false callback warnings when HTTPS terminates at a proxy. Keep host, port, and path drift detection. - Document the setup flow, pagination, callback diagnostics, and shared wizard footer rule. ## Verification - Passed: `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. - Passed: focused Slack callback and pagination integration tests; UI clipboard, wizard, pagination, and activity tests; OpenAPI route tests. The final access-gate fix also passes 27 focused tests covering cloud authentication/bootstrap, nonmember invitations, token validity, and the server-enforced rollout flag. - Passed: all 1,002 chat integration tests, 6,356 UI tests, and all 11 provider browser scenarios (including mobile light/dark navigation). After rebase, the identity route, sidebar, and 25 clipboard tests pass. - The full local `pnpm test:run` was attempted. The first run found 14 Slack fixtures that needed explicit guest access; those are fixed and the complete chat suite passes. Unrelated embedded PostgreSQL startup/resource failures and timeouts prevented a clean full local run. All CI checks pass on `2d858b036`, including the full chat, server, workspace, build, typecheck, and browser suites. - Live test drive: Slack app creation, credential setup, URL verification, private account confirmation, mention messages, and thread replies. Verified the callback warning clears for the existing proxied connection. - Review: create a Slack connection, follow the six steps, link a second user's account, and browse older activity with Next and Previous. ## Risks - Identity invitations carry a temporary capability. Tokens are hashed, expire after 15 minutes, work once, and require explicit confirmation by a company member. Access requests do not grant membership. - New Slack connections reject unlinked people by default. Existing connection settings remain intact. - Activity is a live ledger. Updated action rows can move forward in time. Older pages do not poll. - Proxy tolerance affects health display only. Slack signature checks and proxy authentication settings remain unchanged. - No database migration or package-lock changes. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, code execution, and browser verification. The runtime does not expose an exact model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted suites; full local-run limitations documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
422287eecd |
fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task messages, provider execution, and task outcomes. > - First-time user tests exposed gaps in recovery, completion permissions, message delivery, and Stop behavior. > - These gaps left usable output hidden, completed work waiting for bookkeeping, or safe work unable to continue. > - This pull request fixes the shared lifecycle and receipt paths while preserving process ownership and action checks. > - Users can continue work with accurate task state and durable messages. ## Linked Issues or Issue Description **What happened?** A stopped local Codex execution could remain blocked even after its processes had stopped and its complete transcript proved that no external action needed replay. Claude under Conservative permissions could fail to call task completion tools. Recovery could reuse an assistant item ID and overwrite prior output. A delivered comment could remain marked uncertain after navigation. Stop could look like Pause or a new recovery incident. Workspace contention could look like cancellation. A direct reply reopening Done could enter a clarification loop. **Expected behavior** Recover automatically only with verified termination and complete action receipts. Preserve answers and messages. Keep task completion available under Conservative permissions without broad tool access. Show crashes as Blocked, actual human decisions as In Review, and ordinary workspace contention as waiting. Stop the current response and allow a new direction. **Steps to reproduce** 1. Create ordinary response tasks with local Codex and Claude Code, then send follow-up messages through the task composer. 2. Interrupt a disposable local Codex runner during text-only work. Verify automatic continuation and retained output. 3. Stop a response, send a new request, answer a clarification, and reopen completed work with another message. 4. Navigate or reload while a comment submission is pending. Confirm the exact persisted request receipt settles it without removing newer draft text. 5. Run two tasks in a shared Daytona workspace. Confirm waiting does not appear as failure. **Paperclip version or commit** Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`. Current integration base: `6cef9743c`. Both operator-interruption and workspace-waiting guards are preserved; native restart and legacy permission rules remain documented. **Deployment mode** An isolated source-built test-drive instance, with real local Codex and Claude Code providers and disposable Daytona environments. Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163. This PR addresses additional failures from ordinary task journeys, including controller restart handoff and repeated warm sandbox setup. Historical task status reconciliation is excluded. ## What Changed - Persist runner ownership immediately at spawn and resume an explicitly adopted runner even when the controller crashed before the first driver checkpoint. Detach the controller safely across graceful restarts, including session startup. Prevent an old finalizer from suspending or signaling an adopted runner. Checkpoint idle warm sessions before shutdown. Preserve the same run and queued follow-up messages. - Scope saved legacy queue successor checks to the queue owner while preserving ordinary task locks, operator identity, assignment gates, and exactly-once delivery. - Preserve managed Codex credential files when an old session is detached for restart; normal owned cleanup still copies refreshed auth back and removes the scoped copy. - Reuse the bound warm shared sandbox and fully verify an existing staged provider pack before using it. This avoids repeated uploads when the pack is already valid. - Add a narrow local Codex replacement path with stopped-process proof, a closed transcript inventory, exact completion receipts, and fresh-session lineage. Preserve no-replay holds when evidence is incomplete. Recovery may clear only the same run's recorded Blocked status version; manual re-blocking and dependency changes invalidate that receipt, while queued comments do not. Later blocks stop scheduled, queued, and final dispatch; queued/final checks re-read dependencies even when the task status stays In Progress. - Permit only task delivery and human-input tools through the isolated Claude runner's exact task bridge. - Scope assistant item identity to the provider turn and ignore only authority-free Codex skill-change notifications during startup. - Reconcile composer submissions by client request ID across response loss, navigation, and reload. Retain text typed during delivery. - Keep acknowledged run-only Stop neutral and show workspace contention as waiting. Project exhausted native failures as Blocked. - Restore the guarded task-page retry action for failed legacy runs, including the server-supported explicit new-attempt path for stopped conversation adapters. Preserve native/process recovery holds and avoid promising Retry while a decision or execution gate hides it. - Refresh delivered artifacts and handle direct user replies that reopen completed work without a clarification loop. - Check the embedded PostgreSQL PID, data directory, and actual port before connecting or migrating. - Document accepted behavior and add focused regressions at lifecycle, route, transcript, and UI boundaries. ## Verification - Final head `fece606ac2` passes the complete GitHub CI matrix: **34 green checks, two expected Storybook skips, no failures or pending checks**, including `ci / verify`, `ci / e2e`, full runner verification, typecheck, build, every server/workspace shard, and all browser shards. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34727183287). Greptile is **5/5 with no open findings**. The final two commits only refine test fixtures; both affected suites pass 24/24 locally and in CI, with server typecheck green. - Complete local Vitest coverage uses the canonical groups/shards: all 635 general server suites, all 145 serialized suites, and all workspace packages. The aggregate began on `0a8001c18` while the final queue fix arrived: 23,903 passed, five failed, 87 skipped. The five port/socket/timing failures passed unchanged in follow-ups (60 tests in the exposure/file suites and 412 tests covering the serialized failures and unrun tails). The final queue/operator-identity suites separately passed 52/52. This is aggregate coverage plus explicit reruns, not a pristine single-command final-head run. - After integration with current master, queue/operator-identity/continuation suites passed 162/162 and affected UI suites passed 140/140. ACP Stop/continuation and legacy task/Inbox/message browser suites passed 9/9, including both task recovery Retry and thread Try again, automatic saved-message delivery, exactly one new run, Done, and retained output after reload. The default process Stop/Pause/Resume browser case passed (the native-provider case is opt-in and skipped by default). The complete Board attachment/receipt browser suite passed 11/11 on a disposable instance, covering both composers, exact receipts after lost responses, no replay, bound attachments, and newer drafts after reload. - Blocking-intent regressions cover pre-existing Blocked, a mismatched run/cause, an explicit manual re-block, changed dependencies, a queued comment after failure, and a block arriving between scheduling and provider dispatch. The negative cases reproduced before the fix. All 478 affected executor/recovery/dispatch tests passed; both database suites ran separately after availability-probe skips in the first combined command. The final late-dependency check passed all 143 affected recovery/dispatch tests (zero skips) after two new negative cases reproduced the bug. - Focused runtime regressions cover awaited runner ownership publication, authenticated adoption before the first checkpoint, old-finalizer detachment, idle and busy warm-session shutdown, rejected checkpoint propagation, provider-pack verification, and managed-Codex credential preservation. Four managed credential detachment cases reproduced the bug before the fix; normal owned cleanup still succeeds exactly once. - Live local Claude: SIGKILL 2.6 seconds into startup recovered the same run automatically in 53 seconds, then a normal follow-up completed in 24 seconds. SIGTERM 2.5 seconds into startup preserved the same run (54 seconds) and its queued follow-up (21 seconds). Answers remained visible and the task reached Done. - Live Claude Daytona: a warm follow-up retained its sandbox and fell from 121 seconds to 44 seconds. A separate cold turn took 127 seconds; after controller shutdown and checkpointing, its follow-up completed in 33 seconds with the same sandbox, workspace, native session, and runner. Both answers remained visible and the task was Done. - Other live journeys covered task completion and follow-up with local and Daytona Codex, local Codex crash recovery, Stop then new direction, clarification response, live artifact refresh, and shared-workspace waiting. - Validation limits: the opt-in native composer Stop/Pause→subtree Resume fixture exposes terminal/result ordering and subtree-cancellation attribution bugs that can leave a child task blocked; that new finding is assigned to a separate follow-up and is not claimed fixed here. Default CI skips this optional native-provider fixture. Managed-Codex credential handoff and the queue-agent integration use automated regression evidence. Cold custom provider-pack uploads still add startup latency. ## Risks - Automatic replacement remains deliberately narrow: local Codex, verified stopped identities, unchanged retained state, and a complete text/completion-only turn. Unknown actions, partial history, or changed ownership remain blocked. - Claude completion permission handling changes an upstream package patch. The exact isolated task bridge must remain pinned; unrelated tools keep their existing permissions. - New task failure projection changes user-visible status. No historical status backfill or database migration is included. - This is a broad lifecycle fix across server and UI. Live proof covers graceful local Claude restart during startup and idle Claude Daytona session recovery across controller shutdown. Live abrupt SIGKILL during local Claude startup also recovered the same run. Unknown ownership or missing action evidence still blocks reuse. Cold custom provider-pack uploads still add startup latency; this change avoids unnecessary repeat uploads. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, browser automation, and tool use. The exact hosted model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
52811c6ce6 |
fix(tasks): require resume before sending to paused tasks (#13232)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task execution controls let board users pause a task or its subtree. > - The composer still accepted messages while a pause hold was active. > - A paused task must require an explicit resume before the user can send another message. > - This pull request replaces the composer with an amber pause card and checks board comment writes on the server. > - The user keeps their draft and resumes through the existing task controls. ## Linked Issues or Issue Description Refs #13104. Refs #13119. **What existing behavior does this improve?** The task composer and existing task/subtree pause controls. **Current behavior** A paused task can still receive a board message. The pause notice sits outside the composer, which leaves the send action available. **Proposed behavior** Show an amber takeover in both task chat and the classic composer. Preserve the draft. Require the user to resume the task or the ancestor subtree before sending. Reject board comment writes through either supported write route while the pause hold is active. **Breaking changes** Board comment writes to a paused task now return HTTP 409. Agent run reports remain supported during a pause. There is no schema migration. ## What Changed - Add a shared amber composer takeover with task, subtree, saved draft, pending, and error states. - Use effective ancestor pause state in both composer interfaces. Refresh it after pause events, task updates, and rejected sends. - Preserve draft text and attachments. Hide editor, send, queued edit, and pending question controls while paused. - Check active pause holds before board comment writes can mutate tasks, store comments, or wake agents. - Connect the approved Storybook examples to the production component and update the design and behavior docs. - Add browser coverage for both composers, draft persistence, resume, inherited holds, and rejected writes. Update ACP continuation coverage for the explicit resume requirement. ## Verification - Passed: `pnpm -r typecheck`. - Passed: `pnpm build`. - Passed: `pnpm build-storybook`. - Passed: `pnpm check:token-gates` and `git diff --check`. - Passed: focused UI tests (398 tests) and server route tests (127 tests). - Passed: `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/paused-composer.spec.ts tests/e2e/acp-stop-continuation.spec.ts` (5 tests). - Passed: manual browser walkthrough in a disposable local instance. Pause with a draft, refresh while paused, resume, send, and reopen. The draft returned, and one message persisted. The amber card and resume dialog were readable with no clipping. - Full local `pnpm test:run` did not pass: the general-server stage recorded 9,072 passing tests, 6 database setup failures from macOS shared-memory exhaustion, and 4 failed tests. This stopped the script before its later groups. Latest-head CI runs those groups independently. - Local follow-up: the Git file-resource load test passed on rerun (4 tests); native finalization migration passed after clearing the abandoned browser-test database allocation. Building the native debug fixtures fixed the missing fake provider. The remaining native-session recovery assertion also reproduces on untouched base commit `87b3e5fc6` (36 pass, 1 fail on both base and PR). It expects a settled-session error but receives a semantic-input-digest error. - The final UI build, UI typecheck, token gates, both thread suites (182 tests), and all five browser tests passed after the queued-action review fix. All 31 latest-head CI checks passed, including all server, workspace, browser, build, release, and security gates. Two optional Storybook jobs were skipped by workflow policy. Greptile reviewed `32d8fb5f5` at 5/5 with no open findings. - Review the Paused Composer and Tasks / Execution Controls stories. Pause a task with a draft, verify the amber card, resume, and verify the draft can be sent once. ## Risks - Clients that used board comments to continue paused work must resume first. The response is an explicit HTTP 409. - Pause state can change while a page is open. Live updates refresh the composer, and the server rejects stale sends before their side effects. - Resume keeps the existing dialog and optional agent wake behavior. Agent reports from interrupted runs remain allowed. ## Model Used OpenAI Codex, based on GPT-6, assisted with design, implementation, code execution, and browser verification. The exact runtime model ID and context window are not exposed in this session. The agent used reasoning and tool calls. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the relevant tests locally and they pass; the full local-suite limits are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cfd30fb07 |
feat(ui): add composer Stop and simplify task controls (#13104)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task composer is where operators direct running agents. > - Operators need to stop work without leaving the conversation. > - Existing pause controls already hold task trees and interrupt both runner types. > - This pull request connects the composer to those controls and removes repeated feedback. > - Operators can pause work quickly and still queue messages while agents run. ## Linked Issues or Issue Description **What existing behavior does this improve?** Task pause, resume, and cancellation in the task page and composer. **Current behavior** The empty composer cannot stop a running task. Task controls require extra confirmation and reason text. Pause can show several notifications for the task already on screen. **Proposed behavior** Show Stop while this task runs and the composer is empty. Text or attachments switch it to Send. Stop and the menu use the same manual pause hold. Parent pauses include descendants. Keep task cancellation in the menu with a compact confirmation. Show one quiet pause row and gray cancelled-run details. **Reason and benefit** Operators can interrupt execution with one click. Drafts and queued messages keep their existing behavior. The UI waits for actual termination, including native cancellation acknowledgment. **Breaking changes** No endpoint, schema, or task-status change. Pause no longer asks for confirmation or a reason. Resume now honors the existing wake-agents option. Task notifications are suppressed for the task and subtree currently in view. Related UI work: #8228 changes navigation and composer shortcuts. This PR covers execution controls. No duplicate Stop-button PR was found. The change improves existing controls and does not duplicate a roadmap milestone. ## What Changed - Add Stop, pending feedback, duplicate-click protection, and inline errors to the composer. - Share the pause mutation across the composer, active-run controls, and menu. - Poll affected runs after a pause request. Require native cancellation acknowledgment. - Remove pause confirmation and shared reason fields. Reduce cancel confirmation to its task count and actions. - Honor wake-agents for executable tasks only. Preserve the pause when recovery review is needed; show partial wake failures inline. - Preserve explicit legacy reconciliation decisions while their continuation waits for dispatch. - Suppress notifications for visible task trees. Use quiet pause and cancellation feedback. - Add interactive stories using production controls and native/legacy end-to-end tests. ## Verification - User reviewed the running feature and revised Storybooks in the browser. - Rebased focused checks passed: 295 original targeted tests, 161 updated route/page/notification/status tests, and 26 recovery integration tests. - Both isolated runner journeys pass on the final revision (1.7 minutes). Coverage includes queueing, parent and child interruption, persisted holds, no automatic continuation, reconciled resume, cancellation, terminal exclusions, and no Stop toast. - Native coverage uses real runnerd with a deterministic provider fixture. Legacy coverage checks actual process termination. Live hosted-provider execution was not tested. - Repository typecheck and build, Storybook build, and token gates passed after rebase. The final server typecheck/build also passed. - The broad local run completed its general-server stage with 7,219 passing tests, 48 skipped, and two failures from cached pre-fix source and a stale native provider fixture. Both failed tests pass in fresh final-head reruns after rebuilding the fixture; the script did not continue to its later local stages. CI runs all test groups on the final revision. - Final revision: all 31 applicable CI checks passed; Storybook visual regression was skipped by its workflow conditions. Greptile: 5/5, zero unresolved comments. - Review `Tasks / Execution Controls` in Storybook. Type and clear a draft, stop a run, expand cancellation details, and test the menu on desktop and mobile. ## Risks - Stop pauses descendants for a parent task. This is the existing pause contract. - A held task can remain active if interruption fails. The UI shows an error instead of claiming termination. - Resume can start multiple assignees when wake-agents is selected. Backlog, blocked, and terminal tasks stay excluded. Existing execution reconciliation remains mandatory where required; Resume never invents action-outcome evidence. - Notification suppression uses the visible task and cached subtree. Notifications for unrelated work remain enabled. ## Model Used OpenAI GPT-6 through Codex. The exact runtime snapshot and context-window limit are not exposed in this session. Used reasoning, tool calls, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c185e64b77 |
feat(ui): chat-style task view behind an experimental flag (#10606)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators spend most of their time on the issue detail page. They talk to the assigned agent there through comments. > - The current page reads as a ticket form. The thread sits below properties, the composer sits mid-page, and live agent activity renders as dense transcript logs. > - Talking to an agent is a conversation. A chat-first layout matches that mental model better than a ticket form. > - A layout change this large must not disrupt current users. It needs a safe opt-in path and full parity with the existing thread features. > - This pull request adds a chat-style task view behind a new "Chat-Style Tasks" experiment toggle. The flag is off by default and the existing page is unchanged when it is off. > - The benefit is a focused, readable conversation with the agent: live tool activity folds into compact summaries, the composer stays at the bottom, and properties, plan, and artifacts move into header tabs. ## Linked Issues or Issue Description Refs #49 (chat with agents is a much-wanted feature). Related PRs found in the dedup search: - #4489 — an earlier, closed attempt to promote the conversation to the primary surface on issue detail. This PR is a fresh, flag-gated take on the same goal. - #8837 — an open PR that proposes a two-column task layout. It restructures the same page but keeps the ticket paradigm; this PR is orthogonal because it is opt-in and chat-first. **Subsystem affected** UI (issue detail page). **Problem or motivation** The issue detail page presents agent conversations as a ticket: properties first, thread below, composer in the middle of the page, and raw transcript noise during live runs. Users who mainly converse with their agents must scroll past chrome to follow the conversation, and live activity is hard to read. **Proposed solution** An opt-in chat-style view of the issue detail page, gated by a new "Chat-Style Tasks" experiment toggle in Settings → Experimental. With the flag on, the thread fills the center pane, the composer docks to the bottom of the viewport, Properties / Plan / Artifacts become header tabs, live turns show a status pill with the current tool action and elapsed time, and settled turns collapse to a "Worked · N tools" summary that expands into per-tool rows. With the flag off, nothing changes. **Alternatives considered** Restyling the existing layout in place (rejected: too disruptive without an opt-out), and a separate chat page beside the issue page (rejected: splits the task's single source of truth). A per-request lab page (`/task-chat-lab`, dev-only) was kept for design iteration instead. **Roadmap alignment** ROADMAP.md "CEO Chat" wants lighter conversations that still resolve to real work objects. This PR keeps the core task-and-comments model — it only changes presentation, opt-in — so it does not duplicate that planned work. ## What Changed - New `enableTaskChatRedesign` instance setting, exposed as a "Chat-Style Tasks" experiment card in Settings → Experimental (shared feature catalog, validators, server instance-settings service, and UI settings page). - New `ui/src/components/task-chat/` component family: chat thread with turn grouping, agent reply bubbles, live status pill, collapsible turn summaries with per-tool rows, plan tab with a sticky CTA action bar, inline interaction cards, per-request mode chips, and a bottom-docked composer. - A shared tool taxonomy (`tool-taxonomy.ts`) maps tool names to verbs and icons; the status pill, tool rows, and the classic transcript view all use it. - A transcript adapter converts stored run logs into chat turns; it dedupes tool-call updates by `toolUseId` so tool counts match the expanded rows, and it keeps a tool row's first real name when later generic updates arrive. - Composer: posts on Cmd/Ctrl+Enter, supports image paste with object-URL thumbnail previews (revoked on clear/unmount), and uploads through the issue attachments route. - `IssueDetail.tsx`: with the flag on, pane tabs move to the header bar, the header is not sticky, and the chat fills the center; with the flag off, the previous layout renders unchanged. - Motion tokens for the new animations live in `ui/src/index.css` with a `motion-tokens.ts` catalog and a test that keeps the two in sync (the catalog now also covers the shared enter/exit/swap tokens that the decision/quicklook block declares). - A dev-only `/task-chat-lab` page with fixtures and a tweak panel for motion tuning. ## Verification - `pnpm typecheck` — clean across the workspace. - `pnpm check:token-gates` — 3/3 CLEAN. - `cd ui && pnpm vitest run` — 3,344 of 3,345 tests pass locally. The one failure is `IssueProperties.test.tsx` monitor-row time formatting, which is timezone-sensitive: it also fails on unmodified `origin/master` in a non-UTC timezone and passes with `TZ=UTC`. It is not related to this change. - `cd server && pnpm vitest run src/__tests__/instance-settings-service.test.ts` — 21/21 pass (covers the new setting). - Manual: start the dev server, open Settings → Experimental, enable "Chat-Style Tasks", and open any issue. The thread fills the page, the composer docks to the bottom, and Properties / Plan / Artifacts appear as header tabs. Assign an agent and comment to watch a live run: the status pill shows the current tool action with elapsed time, and the finished turn folds into a "Worked · N tools" summary. Disable the toggle and confirm the classic page is unchanged. - Visual snapshot baselines are intentionally not updated: per `doc/design/DECISION-SHEET.md`, "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks - The flag-off path goes through the same `IssueDetail.tsx` file, so a regression there would affect current users. Mitigation: the classic markup renders through the same components as before behind explicit flag conditionals, and the full UI suite passes. - The transcript adapter interprets stored run-log formats, including legacy entries without `toolUseId`. Malformed logs degrade to generic tool rows rather than crashing. - The new view changes no server behavior other than one additive instance setting; it is additive and default-off. Overall risk with the flag off is low. ## Model Used - Claude (Anthropic), model id `claude-fable-5` (Claude Fable 5), extended thinking enabled, agentic tool use (file editing, shell, test execution) via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c07e650cd7 |
feat(ui): single-source design tokens, visual regression suite, and theme retune (#9134)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is the operator's daily surface: task lists, boards, budgets, agent status — all built on shadcn components and Tailwind > - Visual values (colors, spacing, type sizes, radii) were hardcoded at ~1,600 call sites: the same "small gray label" was 9/10/11px depending on the file, charts disagreed with chips about status colors, two toggle-switch implementations coexisted in two greens, and there was no visual regression coverage > - This made the UI drift-prone and made any restyle a hundreds-of-files project, which discourages design iteration > - This pull request extracts visual values into a single token layer in `ui/src/index.css`, adds a Storybook visual regression suite backed by external immutable baseline archives, and then applies a deliberate retune reviewed change-by-change on screenshot diffs > - The benefit is that Paperclip's look becomes a config surface: retheming is a token edit reviewed as a snapshot diff, drift is blocked by a token gate, and future UI PRs can prove exactly what changed visually without committing hundreds of PNGs ## Linked Issues or Issue Description No existing public issue covers this work (searched "design tokens", "visual regression", "design system" across issues and PRs). Related in spirit: Refs #8982 (theming a hardcoded panel — a one-off instance of the same problem class this PR addresses systematically). **Problem (feature-request form):** UI visual values are hardcoded per call site with no source of truth and no regression coverage; consistency depends on reviewer memory, and restyling requires mass file edits. **Proposed solution (this PR):** a single token layer + enforcement gate + externally stored visual snapshot suite, then an intentional restyle on top of that foundation. ## What Changed - **Token extraction (zero visual change, machine-verified during development):** committed codemods (`scripts/codemod-*.mjs`) moved ~1,600 hardcoded color/type/spacing/radius/shadow/misc values into named tokens in a non-inline `:root` block of `ui/src/index.css`. - **Visual regression suite:** `pnpm test:storybook-visual` covers 255 stories × light/dark = 510 Playwright screenshots at `maxDiffPixels: 0`, plus new primitive-coverage stories and deterministic-render fixes. - **External visual baselines:** committed PNG snapshots were removed. `tests/storybook-visual/baseline-manifest.json` pins an immutable archive URL/hash/size/count, and `scripts/storybook-visual-baseline.mjs` handles `download`, `verify`, `pack`, and trusted maintainer `upload` flows. - **Opt-in visual CI artifacts:** added a `Storybook Visual` workflow that runs on manual dispatch or PRs labeled `storybook-visual`, downloads/verifies the baseline, runs Playwright, and uploads Playwright report/test-result artifacts for review. Normal PR runs do not mutate baseline objects. - **Token gate:** `pnpm check:token-gates` — zero hex literals, zero arbitrary bracket values, zero raw font-sizes in `ui/src/components/**` and `ui/src/pages/**`, with a documented inline allowlist for legitimate opt-outs. - **Theme retune (intentional, snapshot-reviewed):** new base theme values; radius ladder derived from a single `--radius` knob; micro-type cluster collapsed to a named ladder (`--text-nano/micro/compact` + Tailwind `text-xs`/`text-sm`); letter-spacing collapsed to named steps. - **One status-color vocabulary:** charts, quota/budget bar fills, RUNNING/live chips, and liveness indicators all use the canonical `--status-*` hues. Light-mode legibility fixes for red alert surfaces that used dark-tuned text classes. - **One switch:** `ToggleSwitch` restyled to the registry capsule form, second hand-rolled implementation removed, and all call sites unified. - **Docs:** `DESIGN.md` is the design contract; `doc/design/` holds audit reports, decision logs, and updated guidance for external baseline review/update workflows. - Dead code removed (`agentStatusBadge` duplicate map), byte-identical contrast constants consolidated, semantic renames (`--project-seed`/`--project-none`, `--liveness-blue`). ## Verification - `pnpm check:token-gates` — 3/3 gates CLEAN during the design-system run - `pnpm typecheck` && `pnpm --filter @paperclipai/ui build` — green during the design-system run - `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` — pass after external-baseline rework - `pnpm exec tsc --noEmit --pretty false --module NodeNext --moduleResolution NodeNext --target ES2022 --types node,@playwright/test tests/storybook-visual/playwright.config.ts tests/storybook-visual/storybook-visual.spec.ts` — pass after external-baseline rework - `git diff --check origin/pr/9134..HEAD` — pass after external-baseline rework - `find tests/storybook-visual -type f -name '*.png' -print | wc -l` — `0` - `node scripts/storybook-visual-baseline.mjs verify` — intentionally fails closed until the first trusted maintainer publishes the baseline archive and updates `baseline-manifest.json` ## Risks - **Large but shallow:** the PR still touches many UI files due to mechanical token extraction and retune work, but committed PNG snapshot churn has been removed from the branch. - **Baseline publication required before the visual suite can pass in clean clones:** the manifest currently has placeholder archive metadata. A trusted maintainer must publish the first immutable archive, then update `baseline-manifest.json`. - **Rendering platform variance:** the external baseline should be captured in the documented Linux/Chromium environment. Future CI runs verify against the pinned archive and fail closed on checksum/count mismatch. - **Visual CI is opt-in while stabilizing:** add the `storybook-visual` label or dispatch the workflow manually to produce downloadable Playwright report/test-result artifacts. - **Scheduled follow-ups, deliberately out of scope:** Tailwind palette classes map to semantic tokens in a dedicated pass; card/pill component consolidation; ESLint ratchet. Tracked in `doc/design/DECISION-SHEET.md`. ## Model Used Claude Fable 5 (Anthropic, `claude-fable-5`, Mythos-class tier) with extended thinking, running in Claude Code with tool use; mechanical phases delegated to Claude Sonnet subagents. Follow-up external-baseline rework assisted by OpenAI Codex (`gpt-5` coding agent with repository, terminal, and GitHub tool use). All bulk rewrites executed via deterministic, idempotent scripts committed in `scripts/`; intentional visual changes were human-reviewed on screenshot contact sheets. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run targeted local verification and documented the intentional baseline-publication failure above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green *(pending new CI run after this rework)* - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(pending review)* - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) and OpenAI Codex --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |