mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
codex/plugin-task-execution
55
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2ca0d26a99 |
fix(connections): recover missing personal AI credentials in chat (#15376)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9b3fe260ba |
fix(tasks): surface Codex ChatGPT model rejection (#15299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native Runner runs report provider errors and can also save a structured failure result. > - Codex rejects a model that the user's ChatGPT account cannot use. > - Master now diagnoses this rejection, but the task's compact run data omits its message. The thread can still call it a generic run failure. > - Users need the account restriction and a clear step to repair the model selection. > - This pull request shows the existing diagnosis as “Model unavailable” on the task. ## Linked Issues or Issue Description **What happened?** Codex returns HTTP 400 with `invalid_request_error` and the message `The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account.` Master now stores an actionable diagnosis for this rejection. The task thread still labels it “Run failed” and does not receive its error text in compact run data. **Expected behavior** Show the account restriction on the task and in the run error. Tell the user to choose a supported model or clear the task's model override before retrying. **Steps to reproduce** 1. Use a Native Runner Codex agent signed in with a ChatGPT account. 2. Select a model that produces the rejection above and start a task. 3. Let the runner save its generic failed result. Inspect the task's failure marker and recovery notice. **Additional context** Refs: #15304. That merged PR diagnoses the provider failure and preserves worker and review recovery rules. This PR adds its task-facing message and guidance without changing that diagnosis or those rules. Refs: #13134. That PR improves model discovery for ChatGPT accounts. This PR exposes the rejection when a configured model still fails at execution time. ## What Changed - Return a bounded model rejection message in compact issue-run data. - Show “Model unavailable” and model-change guidance in the task thread and recovery notice. - Add database and UI regression tests, including both the original rejection text and master's fixed diagnosis. Document the new failure message. ## Verification - Focused merged-branch validation: 279 tests passed across activity service, native provider failure observation and PostgreSQL integration, TaskChatThread, and ExecutionBlockerNotice. - `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` passed on the merged source tree. - Greptile reviewed conflict-resolution commit `1db39c4e636a92e2e76ba224b71c6f1e2f556c81` at 5/5 with no findings or inline comments. - All 55 check runs completed without failure on that commit. The two optional Storybook jobs were skipped. The legacy Snyk status passed. The branch has no merge conflicts. - UI regression assertion: the task's failure marker says “Model unavailable” and retains the account restriction. It no longer says that this failure happened after a final response. ## Risks - Provider recognition and recovery are owned by the existing master implementation. This PR exposes only bounded error text for failed runs with `native_provider_model_rejected`. - Historical runs with the generic `adapter_failed` code are not reclassified. No stored run is rewritten. - Existing retry and reconciliation gates remain in place. No schema migration or model configuration change is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser inspection. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
81cd03b3b5 |
test(ui): keep initial reasoning with the saved reply (#15107)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The issue thread shows saved replies and agent run history. > - History that appears after the reply can move the page. > - Master now waits for initial run history before it shows the thread in #15228. > - This pull request checks that reasoning and the saved reply appear together in the first visible frame. > - The test protects that behavior from a later change. ## Linked Issues or Issue Description No public issue tracks this report. This is a bug regression test. **What happened?** The issue thread showed a saved answer before its reasoning and tool history loaded. The later history moved the page. **Expected behavior** The first visible thread contains the initial reasoning and the saved reply in order. **Steps to reproduce** 1. Open an issue with an agent run and a saved reply. 2. Delay initial run-history hydration. 3. Observe the saved reply before its reasoning appears. Related work: #15228 now contains the reveal gate. This PR adds a direct regression assertion for that gate. #14667 takes a different loading approach and remains open. ## What Changed - Add a saved agent reply to the initial-history test. - Assert that the native run's reasoning appears before that reply when the thread becomes visible. - Assert that the thread stays inert until history is ready. ## Verification - `vitest run ui/src/components/TaskChatThread.test.tsx ui/src/pages/IssueDetail.test.tsx` passed: 353 tests. - `git diff --check origin/master...HEAD` passed. - GitHub CI checks are in progress for the rebased head. ## Risks - Low risk. This PR changes only a test. - The production fix is already in #15228. This PR must not restore the older code from the previous branch head. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex CLI assisted with tool use and code execution. The exact model ID and context window were not exposed to this agent session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My remote branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing documentation changed) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7498705642 |
fix(ui): stabilize mobile task reading and document navigation (#15228)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People read tasks and send instructions from phones as well as desktop browsers. > - The mobile footer let page text show through its labels, and small text fields made Safari zoom on focus. > - Small scroll changes made the footer switch direction, and its changing page padding moved the conversation. > - Task pages also showed comments before question cards and run history arrived, so the composer and reading position moved again. > - Document links also used native navigation, which reset the reading position or reloaded a task through its UUID URL. Desktop tabs were crowded in the mobile drawer. > - This pull request keeps navigation steady, opens documents in the mounted task, and gives mobile readers a full-height panel with a vertical tab selector. > - The benefit is a stable task view and smoother scrolling on mobile. ## Linked Issues or Issue Description **What happened?** The mobile footer was translucent. Safari zoomed when a person focused a small text field. The footer switched abruptly while scrolling. A large task could show saved comments, then move the page again when a question card or run history arrived. In a local test with delayed responses, a late question moved the mobile composer by about 374 pixels. Opening a plan from the feed could reset the view or reload the task through a UUID link. The mobile document drawer left part of the feed exposed above small desktop tab controls. **Expected behavior** The footer has an opaque surface and moves smoothly after deliberate scrolling. Text fields do not cause automatic focus zoom. A task shows its initial conversation and composer together at the final scroll position. Background refreshes keep the existing conversation visible. Document links open in the mounted task with its existing cache and reading position. Mobile documents fill the viewport, show a clear close button, and offer a vertical list of open tabs. **Steps to reproduce** 1. Open a task with many long comments and a pending question in iOS Safari. 2. Delay its interactions, activity, and runs responses by different amounts. 3. Reload the page and watch the conversation and composer move as each response arrives. 4. Scroll down and back up, including small direction changes and edge bounce. 5. Focus the task composer, search field, and new-task title and description. 6. Open a plan or another task document from the feed, including a link that uses the task UUID. Close the panel and check the reading position. 7. Open several documents on a phone. Switch tabs and close both active and inactive tabs. **Paperclip version or commit** Developed from `1c07b5903` and rebased onto `59015846a`. **Deployment mode** Built from source. Tested in an isolated local test drive with iOS 26.5 Simulator Safari and Chrome. A temporary local proxy delayed independent responses for the layout test. Related work found in the duplicate search: - Refs #14727. It made saved replies appear before supporting history. This PR keeps its parallel requests and narrows the tradeoff in favor of a stable first layout. - Refs #14667. This open PR takes a different approach with per-run placeholders and retries. This PR fixes the observed question/composer movement and mobile navigation behavior. - Refs #13095 and #13597. These earlier fixes added task scroll anchors and skipped transcript waits for scheduled retries. - Refs #6550. Earlier mobile board polish. - Refs #9467. This related open PR changes list and generic tab reflow. The task-pane selector uses a separate component. ## What Changed - Give the mobile footer an opaque semantic surface. - Set a base-size floor for editable text on touch devices to prevent Safari focus zoom. Preserve larger title text. - Share mobile scroll tracking between both layouts. Accumulate scroll distance, ignore edge bounce and changed document bounds, and update once per frame. - Use shared motion tokens for the footer and composer. Keep page padding stable and honor reduced motion. - Wait for the initial question cards, attachments, work products, activity, runtime selection, plan, and relevant transcript history before the first reveal. Skip scheduled retries and older runs outside the initial comment window. - Bound the first reveal to 15 seconds. A stalled supporting request leaves saved conversation and the composer accessible with an explicit loading notice. - Keep concealed mobile history from stretching the document. Keep the composer mounted but concealed until the same reveal. Keep both visible during later refreshes. - Route first and repeated same-task document clicks in place. Recognize UUID and identifier links. Preserve the thread history entry and feed position. Keep modifier clicks, downloads, external links, and classic document behavior. - Give the mobile task panel the full viewport and safe-area padding. Use a visible X and 44-pixel touch controls. Replace the horizontal tab strip with a vertical selector that wraps titles and supports keyboard focus. - Add four interactive Storybook states for a few tabs, long names, many tabs, and the last tab. Reuse the production selector and tab controller. - Add navigation and tab regressions, update first-reveal regressions, and document the behavior in `DESIGN.md`. ## Verification - 392 tests passed across the seven focused task-loading, scroll, mobile-navigation, layout, and composer suites. After review fixes, all 339 tests across the four affected suites passed, including stalled-loading fallback on mobile and desktop and the motion-token catalog. - All 442 focused document, tab, task-thread, and scroll tests pass. The final click-propagation cleanup also passes all 136 task-detail tests. UI typecheck, UI production build, and `pnpm check:token-gates` passed. - All four cases in `artifact-tab-arrival.spec.ts` and `text-attachment-tabs.spec.ts` pass locally, covering desktop and mobile selection, composer focus, document rendering, and downloads of the original bytes. - `pnpm --filter @paperclipai/ui build-storybook` passed. Open the mobile tab stories under `Prototypes/Task detail/Mobile tabs`. - A local diagnostic proxy measured cached plan content at about 250 ms after the first click. The HTML load count and task request count did not change. Feed scroll stayed at the same position. First, repeated, and UUID document links were tested at phone and desktop widths. Task-reference links close their preview before the document reader opens. - In Chrome at desktop and phone widths, the delayed-response task showed one complete reveal. The late question no longer moved an already visible composer. - In iOS Simulator Safari, verified the large-task reload, opaque footer, navigation hide/reveal, and search/new-task/composer focus without automatic zoom. - Full workspace typecheck and build passed. The updated UI also passes typecheck, production build, and token gates. - All 54 checks pass on the final commit `bf5c9914e61833e7cc8794a69d72cf8c7057b952` (two additional checks are intentionally skipped). Greptile reviewed that commit at 5/5, and all review threads are resolved. - Full local `pnpm test:run` was attempted but stopped after server-fixture failures. Embedded PostgreSQL startup failure reproduced in an isolated native-interaction fixture after five startup attempts. The broad run also reported a rapid Slack callback ordering test failure. These server paths are unchanged by this PR, and their CI shards pass on the latest head. The full local suite is not claimed as passing. - A localhost proxy stalled the activity response for 30 seconds. Chrome revealed the available conversation after the 15-second deadline at both desktop and phone widths, kept the composer accessible, and cleared the loading notice when the response arrived. ## Risks Slow initial history requests can delay the first conversation reveal by up to 15 seconds. If that deadline expires, late data can change the available conversation while a loading notice remains visible. The reveal waits only for runs in the initial comment window, and later refreshes do not conceal an existing conversation. The larger editable text can change line wrapping on phones. Mobile navigation and composer motion use shared tokens and respect reduced-motion settings. Mobile tab selection changes the control layout. Document links retain URL history while sharing the task reading position; other tasks and external links keep their normal navigation behavior. I checked `ROADMAP.md`. This is a fix for existing UI behavior. ## Model Used OpenAI GPT-6 in Codex. The runtime does not expose a more specific model ID or context-window size. The agent used reasoning, code editing, terminal tools, and Chrome and iOS Simulator testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
59015846ae |
fix(chat): keep dismissed task questions in the feed (#15229)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask users questions in task chats and Agent Chat. > - Agent Chat keeps unanswered questions as compact entries in the feed. > - Regular task chats still show a pending composer badge after dismissal. > - A page reload can also open a dismissed question form again. > - This pull request applies the same feed behavior to both chat types and saves dismissal per person and task. > - Users can continue the chat and return to the original question later. ## Linked Issues or Issue Description **What happened?** Dismissing a question in a regular task chat leaves a pending composer badge. The question can return after a page reload. The earlier feed behavior only applied to Agent Chat. **Expected behavior** A dismissed question stays in the feed. Its form and pending badge leave the composer. Reload preserves dismissal. Opening the feed entry restores the original question and draft answer. **Steps to reproduce** 1. Open a regular task chat with a pending question. 2. Select an option, then dismiss the form. 3. Reload the page. Check that the composer stays clear. 4. Open the question in the feed. Check that the draft is restored. 5. Submit the answer. Check that the answered receipt appears. **Paperclip version or commit** Reproduced on master at `a386a599983519eb1d399f8b770bfccdb2a74762`. **Deployment mode** Browser UI in local and authenticated instances. This change does not depend on the agent adapter. Related work: Refs #14613, which added the Agent Chat feed behavior. Refs #9141, which validates real answers on the server. Refs #11434, which tracks comment-driven changes to interaction state. This PR changes question presentation only. ## What Changed - Show compact unanswered question entries in regular task chats. - Exclude durable questions from composer pending counts and navigation. - Save dismissal in local storage per person and task. Merge the latest saved IDs so dismissals from another tab survive reload. A new question can still open its form. - Keep the original question pending and answerable. Keep approval and permission controls. - Run the question-history regressions in both chat modes. Add reload, new-question, user/task scope, and stale-tab coverage. - Share the real-component Storybook fixture. Add a regular task test drive and an interactive dismissal/answer scenario. - Update the planning-mode browser test to check a dismissed question in the feed after reload and on mobile. - Update the design rules, implementation spec, and preview instructions. ## Verification - 329 focused thread, composer, and interaction-card tests pass. - `pnpm -r typecheck` passes. The UI typecheck also passes after the review fix. - `pnpm build` passes for the full repository. - UI build, Storybook build, and token gates pass. The UI build and token gates were rerun after the review fix. - Greptile gives final commit `98f71f221` a 5/5 score. The current-head check passes, and there are no unresolved review threads. - All 56 current-head checks are terminal: 54 pass and two conditional Storybook jobs skip. The CI run includes the full test shards, build, typecheck, browser tests, aggregate verification gate, and package canary. - The stale-tab regression fails in both chat modes before the review fix and passes after it. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/planning-mode-visual-verification.spec.ts` passes against a throwaway local server. It checks dismissal, reload, task navigation, and desktop/mobile planning controls. - The duplicate local `pnpm test:run` attempt was stopped after the full CI test suites passed. It did not complete locally. - Manual browser test: select Green, dismiss, reload, reopen, submit the saved answer, and inspect the answered receipt. Also send a new message while the unanswered question remains in the feed. The test drive uses real UI components with fixture response callbacks. - To repeat the browser test, run `pnpm storybook`. Open **Chat & Comments → Task Chat Unanswered Questions → Test Drive**. ## Risks - Dismissal is a browser-local preference. It does not sync to another browser or device. Clearing local storage removes it. - When local storage is unavailable, dismissal lasts for the mounted thread only. - Unanswered questions can accumulate in the feed. They stay pending until answered or resolved through the existing API. - No database migration, API change, or change to approval permissions. ## Model Used OpenAI Codex, `gpt-6.1-sol`, with xhigh reasoning, repository editing, code execution, and browser control. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cc67d4e1d8 |
fix: preserve steering and recover stopped task conversations (#15015)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task conversation must let a user guide a running agent and resume stopped work. > - The active run owns its input protocol, even when the user changes the next model or effort. > - Queue delivery waits for a provider receipt, which must be able to persist during the request. > - A stopped startup also needs a clear user action that passes normal task admission. > - This pull request fixes steering delivery, makes queue actions immediate, and restores explicit continuation. > - The benefit is a responsive conversation that can recover without losing saved input. ## Linked Issues or Issue Description **What happened?** A queued message could change from Steer to Interrupt while a native run prepared. A steer request could wait on its own database lock and fail to deliver. A stopped startup could then leave the conversation without a working Retry or message continuation. Interrupt also waited for the server and showed a toast. **Expected behavior** The active run keeps its input protocol. Steer delivers input to that run. Steer and Interrupt clear the submitted queue rows and show the input in the conversation immediately. Failed delivery restores the latest queue with an inline error. An eligible stopped run offers Retry, and authenticated user input can start a fresh turn through normal task admission. **Steps to reproduce** 1. Start a task with a native Paperclip Runner. 2. Change the selected model or effort while that run prepares. 3. Queue a message and press Steer. 4. Observe the provider receipt and queue state during the request. 5. Stop a startup before its provider process begins, then try Retry or send a new message. 6. Repeat queued delivery with a legacy runner and press Interrupt. **Paperclip version or commit** Reproduced on the parent of this branch, `59c07ede7`. **Deployment mode** Authenticated private deployment. The fixes also cover local task conversations. Related work: Refs #12834, Refs #13354, Refs #13275. The open refactor in #13160 moves the same queue route; it does not fix the receipt lock or stopped-run continuation addressed here. ## What Changed - Select queue behavior from the active run's immutable dispatch and runtime resolution. - Leave the run row unlocked during provider acknowledgement, then lock and read it before merging the receipt. - Retain queued input if the target run stops during that wait. Keep inline delivery errors visible after empty queue updates. - Permit exact Retry and authenticated continuation after verified native startup cancellation. Preserve pause, approval, budget, ownership, and process-stop gates. - Carry undelivered native queue input into a fresh turn once the old execution is confirmed stopped. - Show Steer and Interrupt input in the conversation and clear submitted composer rows immediately. Restore the latest queue inline on failure. Remove delivery toasts. - Keep optimistic delivery stable across stale polls, empty queues, and paginated history. Preserve classic Interrupt error handling. - Document recovery and optimistic delivery behavior. Add regression tests across server, shared queue projection, and UI boundaries. ## Verification - Red-green regression tests reproduced the queue protocol, receipt lock, stopped-startup continuation, and optimistic delivery failures. - The focused server route, continuation, queue, and runner boundary suites passed during implementation. - The queue-route suite passes with 78 tests. The three complete conversation UI suites pass with 347 tests. - UI typecheck, production build, and `pnpm check:token-gates` pass. - Workspace `pnpm -r typecheck` and `pnpm build` pass. The local monolithic `pnpm test:run` is still running; remote CI verifies the complete suite on the latest commit. - All CI gates pass on `bd9031ad56abfcde13d13a13488c1b9217c2fd3a`, including the full test shards, runner verification, browser E2E, typecheck, release registry, and canary dry run. - Greptile reports 5/5 for that commit. Both review threads are resolved. ## Risks This changes queue display and explicit continuation admission. The UI must restore rejected delivery without losing other-session edits. The server must preserve concurrent provider result updates and must not resume a process whose stop is uncertain. Focused tests cover these boundaries. This change has no database migration. ## Model Used OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and context window are not exposed in this session. Used reasoning, repository tools, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b2bd6563d |
fix(chat): hide obsolete execution status and system notices (#14874)
## Thinking Path > - Paperclip lets people oversee agent work through task conversations. > - Conversations should show failures and waits that still affect the task. > - Old run errors and recovery notices remained visible after later work or task completion. > - These messages looked current even when no action remained. > - This change hides obsolete execution status while preserving the responses and activity. > - The full diagnostic record stays available in run history. ## Linked Issues or Issue Description **What happened?** Task chat kept showing “Run failed”, “Stopped”, and “Waiting to resume” after the execution had been superseded or the task had finished. Stored system notices also remained in the conversation. Completed tasks disabled Retry but kept the error message. **Expected behavior** Hide system status that no longer applies. Keep the latest unresolved failure, active recovery holds, and useful recovery actions visible. Preserve messages, files, questions, session boundaries, and inspectable activity. **Steps to reproduce** 1. Let a task run fail, then record a pre-start recovery wait. 2. Complete a later attempt or mark the task done. 3. Open the task conversation. Before this change, the old error and wait remain visible. **Paperclip version or commit** Reproduced in component tests against `0f9e9be408`. **Deployment mode** Task chat in local or hosted deployments; both legacy adapters and the native runner. Related: #14857 and #14869 address execution recovery. This PR addresses the remaining conversation presentation. Searched GitHub for historical chat status, historical errors, and “Waiting to resume”; no duplicate PR found. ## What Changed - Determine status relevance from task state, attempt order, successor evidence, and recovery state. Time alone does not hide errors. - Hide obsolete run markers and stored execution notices. Require run or recovery provenance, so unrelated system updates such as child-task blockers remain visible. Keep errors from the latest failed attempt actionable. - Keep unresolved execution holds visible. A refused pre-start retry does not replace a real attempt, and another agent’s work does not resolve a run-specific error. - Show historical activity without Worked/Stopped labels. Keep historical failures out of the current turn’s summary. - Anchor activity to visible comments so removing a notice cannot remove the response or activity with it. - Document the presentation rules and cover both runner modes and both task presentation modes. ## Verification - 334 focused component and status-policy tests passed across four files, including the child-task relay regressions. - `pnpm build` passed. The UI build also passed after the final presentation changes. - `pnpm exec vitest run --project @paperclipai/ui`: 667 files and 7,157 tests passed. Subsequent focused tests cover the final activity-anchor, live-successor, and notice-provenance changes. - `pnpm -r typecheck` passed. UI typecheck and build passed again after the review fix. - `pnpm test:run` completed its general-server phase with 14,716 tests passed, 17 failed, and 87 skipped; it stopped before later phases. The failures occurred in four unchanged server suites: chat channels, email channels, company skills, and runtime skill cache. A targeted rerun reproduced missing bundled skill paths and `EACCES` during cache-directory rename on macOS. All Linux CI suites pass for the final commit, including these server suites. - Design token gates and diff checks pass. - All 53 checks pass on commit `7af9753859`; two optional Storybook checks are skipped. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/36931968751) includes build, full typecheck, all server and workspace test shards, all eight end-to-end shards, runner verification, and the canary dry run. - Greptile scores the final commit at 5/5. No review threads remain unresolved. The branch is current with `master` and has no merge conflicts. ## Risks This changes presentation only. It does not change execution, recovery, stored comments, or run history. The main risk is hiding a current diagnostic too early. Tests cover active holds, refused retries, different agents, missing timestamps and provenance, live successors, preserved responses, and the current retry target. The base branch has a dependency override/lockfile mismatch. Local installation used the same resolution fallback as CI, then restored the tracked lockfile. No dependency changes are included. ## Model Used OpenAI Codex (GPT-6). The exact runtime model identifier and context window are not exposed in this session. Used reasoning, repository inspection, code execution, and regression tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — changed-code tests pass; unrelated full-suite failures are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3c561642b4 |
fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask for decisions and optional details through cards in chat. > - A clear approval in a message can leave the matching card pending. > - An unanswered question can also block an unrelated later reply. > - Decisions need a saved source message, while optional questions need to remain answerable in history. > - This pull request records conversational decisions and lets users move on from questions and answer them later. ## Linked Issues or Issue Description **What happened?** Native Claude and Codex could act on approval in chat while the original approval card stayed pending. Pending question forms stayed above the composer, were absent from history, and could suppress later chat replies. A late native question answer could wait for a finished run to reconnect. **Expected behavior** The active agent records a clear approval or refusal against the exact card and user message. Ambiguous replies do not grant consent. Users can send another message without answering a question. The question remains pending in history and can be reopened and answered later. The saved answer reaches the agent. **Steps to reproduce** 1. Ask an agent to propose work with a confirmation card, then approve it in chat. 2. Check that the original card records that approval before work starts. 3. Ask an interactive question, send an unrelated message, and reload. 4. Open the unanswered question from history and submit an answer. Related work: #14408 added completion delivery. #14607 tests completion reporting turns. Neither records conversational answers on approval cards. ## What Changed - Add a confirmation endpoint backed by a user comment, with schema validation, OpenAPI discovery, and native Plan-mode access. Ask mode remains read-only. - Check company, active run, actor, current session, message provenance, revision, and resolver policy. Save the decision and audit in one transaction. Retries do not repeat effects. Emit resolution telemetry after commit. - Give fresh and resumed chat turns the actual pending confirmation identities. Teach agents to save clear conversational decisions before acting and to clarify ambiguity. - Keep unanswered Agent Chat questions as compact history entries. A newer user message closes the old form. Question cards never contribute to composer pending counts or navigation, including after dismissing a fresh form. The history card is the sole reminder; clicking it restores that exact form and draft. - Preserve Agent Chat questions when later messages or questions arrive. Historical ordinary inputs no longer gate later chat replies. Current-run requests, task execution, and governed approvals keep their gates. Remove the special acknowledgement-publication proof helpers that this rule replaces. - Route answers to finished native runs through durable fresh-wake delivery, with existing idempotency and source-question context. Settle late replies against contiguous completed conversation turns and freeze their history replay; failed, unhandled, and newly arriving messages remain actionable. - Add real-component Storybook scenarios, database and UI regressions, and a three-turn native Claude/Codex E2E case. Capture distinct, UI-ready screenshots and report the individual assertions. ## Verification - Focused decision/publication/UI regressions after merging master: 288 passed; subsequent UI draft, failed-send, and conversation checks: 199 passed. - Native question and durable delivery regressions: 106 passed, including all four terminal run states and exactly-once late delivery. Seven targeted regressions fail against the original implementation and pass with the fix. - Latest conversation/decision/native-delivery regressions after the master merge: 121 passed. Covers completed progress, missing or failed intervening turns, new messages during a late reply, stale sessions, and frozen retry/replay boundaries. Four new assertions fail before the ordering fix. - E2E support suite after the master merge: 792 passed. Negative controls reject expired cards, wrong questions/answers, stale or missing replies, unrelated clarification forms, and unexpected tasks. - The embedded-browser walkthrough caught one additional defect: dismissing a fresh question still showed a composer badge. Both Cancel and close-button regressions failed before the fix. The fix at `65f2ade12` passes 170 chat-thread tests and 792 E2E support tests. After merging master, 232 chat-thread/confirmation tests, server/UI typechecks, and token gates pass. The preview and two-provider live E2E pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at that commit. All 55 checks are now successful at `e5512a206` (four conditional checks skipped), including the aggregate verification gate and clean-install canary test. The first attempt was interrupted by simultaneous CI worker shutdowns; one failed-job rerun passed without code changes. - [Published Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on): nine real-component scenarios. Manually exercised move on, reopen, preserve draft, answer later, answer one of multiple questions, and a custom mobile answer in the embedded browser. Retested fresh Cancel and close-button dismissal in the updated build, then reopened and submitted the preserved Green selection and inspected its answered receipt. Static preview has no live model/backend; its callbacks are fixture responses. - [First live campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/) reproduced the late-answer completion-state defect on both providers despite correct saved answers and acknowledgements. It also exposed a valid imperative clarification rejected by the old oracle. Both issues are fixed with regression controls; this failing run is retained as evidence. - [Four-cell qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/) passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous confirmation, each on native Claude and Codex. Inspected saved state, source-message decisions, visible cards, and agent replies. Both late-answer chats settled to waiting; no unrequested tasks were created. [Final branch rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/) passed 2/2 at `142630720`: the same unanswered-question journey after merging master, plus an additional screenshot and browser assertion for the actual late-answer acknowledgement. - [Composer-reminder E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/) passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each, with explicit no-badge assertions before and after reload. Inspected saved pending/answered state, both screenshots with a clear composer, and actual Blue acknowledgements; all five behavioral matchers passed per provider and neither created tasks. Cost coverage is partial; this is bounded workflow qualification. - [Fresh-dismissal E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/) passed 2/2 at `e5512a206`: native Claude and Codex, including fresh Cancel, clear composer, reopen, unrelated message, reload, late Blue answer, and actual agent acknowledgement. All five behavioral matchers pass per provider. Inspected the fresh-dismissal screenshots and saved pending/answered identity; neither created tasks. Cost coverage is partial (4/6 runs). - Prior evidence remains available in [the earlier campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/). Its early loading screenshot and overwritten final capture prompted the UI-ready, distinct screenshot fixes. ## Risks - The model interprets intent. The server verifies permission and provenance; it does not infer consent from text. Ambiguous and unrelated replies are not approvals. - Historical questions can accumulate. They remain visible, pending, and answerable; no automatic answer or expiry is invented. - The change to completion gates is scoped to Agent Chat and ordinary historical inputs. Current-turn and governed approvals retain their existing controls. - Live qualification is limited to the selected stories. Broader native onboarding finalization remains separate work. - No database migration. Telemetry adds no fields or values; the contract and README document the commit boundary. Privacy review was requested on the PR. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser-test orchestration. The exact model ID and context-window size are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e5830887b |
perf: reveal task content sooner and parallelize issue reads (#14727)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - A task page must show saved replies quickly so a person can read the work. > - The title could appear while the conversation waited for unrelated metadata and transcripts. > - Thread requests also waited for enriched task details, while the server read several independent fields in sequence. > - This change shows saved content as soon as it is ready, starts thread reads earlier, and runs independent server reads together. > - Native event history takes priority over legacy log fallback, and mentioned tasks load on intent. > - Content-free timing spans make the remaining server delays visible without recording task content. ## Linked Issues or Issue Description **What happened?** Task titles and properties appeared quickly, but saved conversation content stayed hidden for several more seconds while metadata and run transcripts loaded. **Expected behavior** Saved replies and the task description should be readable without waiting for supporting history. Returning to a cached task should show content within a frame or two. **Steps to reproduce** 1. Open a task with saved comments and completed runs. 2. Delay the task activity and runs responses by five seconds in the browser. 3. Observe whether saved content remains hidden until those responses finish. 4. Navigate away and return to the task to check cached navigation. **Paperclip version or commit** The change was developed from `1b48e73e0` and rebased onto `44736c9c7`. **Deployment mode** Built from source, tested in an authenticated staging deployment and with local response replay. Related: #14667 overlaps the transcript reveal behavior and adds separate retry UX. This PR also changes navigation prefetch, parent metadata gates, native log fallback, server read scheduling, and timing spans. #12647 proposes a separate SQL predicate optimization in the runs service. #13597 and #13095 are earlier loading fixes. ## What Changed - Start activity and runs when the task page mounts, alongside task details and comments, using the route reference for shared query keys. Hover/focus prefetch does not start full history reads. - Reveal saved comments and descriptions while metadata and transcripts load. Preserve strict waits for linked-comment navigation and tasks with only runtime content. - For settled native runs, fetch legacy logs only when event history is empty or fails. Preserve live-log subscriptions for queued and running native runs. Fetch mentioned-task details on hover or focus. - Run independent issue-detail enrichment and run metadata reads in parallel while preserving recovery dependencies. - Add `Server-Timing` phases and opt-in OpenTelemetry spans, plus regression tests and observability documentation. ## Verification - **313 tests passed** across the initial seven focused component, cache, timing, and scroll suites. After review fixes, **313 tests passed** across five task-page, cache, prefetch, and live-transcript suites (`pnpm exec vitest run` with `--maxWorkers=1`; these sets overlap). - `pnpm -r typecheck` and `pnpm build` passed locally before the final UI-only review fixes. UI typecheck/build and `pnpm check:token-gates` passed after those fixes. The final commit also passes full typecheck and build in CI. - The full local Vitest run was attempted. Several unrelated embedded PostgreSQL fixtures failed to start, and parallel test workers hit timeouts. Focused reruns passed. A later local full-suite rerun was stopped after the complete CI suite passed; it is not claimed as a local full-suite pass. - Final commit `d1e147145`: **54 checks passed, 2 skipped**, including all server/workspace test shards, all eight browser E2E shards, full build/typecheck, Runner checks, release registry, and canary clean-install verification. [CI run](https://github.com/paperclipai/paperclip/actions/runs/36734642034). Greptile **5/5** after two reviews; both findings fixed and all review threads resolved. No merge conflicts. - Live browser tests on an existing task with two saved replies: median full reload to visible content fell from **1.92 s** (3 samples) to **1.40 s** (5 samples). Cached return fell from **421 ms** (1 sample) to **29 ms** (3 samples). These are observed samples, not a performance guarantee. - With activity and runs delayed by five seconds, saved content appeared in **1.38 s** on desktop and **1.33 s** on mobile. The inspected comment did not move when metadata arrived. Verified history expansion, task properties, pending-input navigation, dashboard return, and mobile layout. - A separate local replay with fixed responses reduced visible-content time from **5.12 s** to **2.15 s**. This isolates frontend behavior and is not a live-server benchmark. ## Risks Progressive history can change the thread after first paint. Existing anchor behavior is retained and covered by tests and delayed-response browser checks. Query aliases must stay aligned for invalidation. Parallel reads can increase short bursts of database work; dependent recovery operations remain ordered. Full reloads still depend on network and task-detail latency. No schema or authorization change. OpenTelemetry remains disabled without an operator endpoint. The added spans use a closed set of phase names and carry no task IDs, task content, or exception text. I checked `ROADMAP.md`; this is a performance fix within the existing task page. ## Model Used OpenAI GPT-6 in Codex. The runtime does not expose a more specific model ID or context-window size. The agent used reasoning, code editing, terminal tools, and Chrome performance profiling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d72389bee2 |
feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps gateway gives agents governed access to external tools. > - Browser Use Cloud can run browser work, but a tool result alone does not let a person watch or take over. > - A task needs a durable browser session, a visible viewer, and recorded costs. > - This pull request adds a Browser Use Cloud v4 connection and interactive browser tabs on tasks. > - People can follow the work, interact with the page, and retain the browser after the agent finishes. ## Linked Issues or Issue Description **Problem or motivation** Agents need governed access to Browser Use Cloud. People need to see and interact with the same browser from the task. A browser must remain available after a run finishes and appear at the correct point in the task feed. **Proposed solution** Add a native REST connection for the v4 API. Bind each session to its company, task, agent, and credential grant. Open its interactive viewer in the task side panel. Record provider costs as financial events. Use `browser-use-cloud` as the app and connector key. Keep its skill with the connector and deliver it only with authorized connection tools. **Alternatives considered** A v3 MCP connection would expose tools without the v4 lifecycle integration. An external viewer link would leave the task. A fixed viewer size would prevent pages from responding to changes in the task pane. **Roadmap alignment** This extends the governed Apps gateway and Connected Apps roadmap. It uses the existing task, grant, secret, approval, and financial records. The work was requested by the maintainer. A search found no duplicate Browser Use connector PR or issue. ## What Changed - Add the Browser Use Cloud app, brand asset, API-key connection, and profile settings under the `browser-use-cloud` key. - Bundle the `browser-use-cloud` skill with the connector. Keep it out of global `skills/` discovery. Deliver it only with authorized task/run connection tools. Remove retired connector skill keys from runtime overlays and preserve unrelated browser skills. - Expose seven v4 tools through the governed gateway and deliver them to native and CLI agents. - Persist sessions, browsers, runs, event and recovery cursors, shutdown leases, and cumulative cost accounting. Recover uncertain paid starts without replaying them. - Enforce task ownership, credential grants, approvals, revoked access, and budget limits. - Add interactive task browser tabs and compact chronological feed entries. Retain the viewer across tab switches and keep visible idle browsers open. - Add debounced automatic viewport fitting, standard size presets, and a viewer ownership lease. - Add lifecycle, authorization, accounting, viewport, UI, and Storybook coverage. - Add an idempotent database migration after the current master migration. Preserve deployed migration hashes. Migrate pre-release Cloud connection and financial keys without replacing grants, credentials, or browser history. - Document provider behavior, live acceptance results, and the lack of documented passkey forwarding. ## Verification - Full workspace typecheck and production build pass on the updated branch. - Token gates, brand asset validation, module boundaries, and migration ordering pass. - Cloud tests verify global skill exclusion, authorized task/run delivery, unassigned agents, disabled connections, revocation, adapter isolation, and secret exclusion. The existing AgentMail connector assignment test also passes. - Migration replay runs twice against existing browser work and financial records. It preserves the records and avoids duplicate costs. - The focused provider, app catalog, OpenAPI, connection gateway, and migration regression suites pass. Recovery coverage includes lost replies, process crashes, provider rejection, and browser arrival acknowledgement. - All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`, including the full test matrix, browser E2E shards, build, typecheck, security, and release canary. Two optional Storybook jobs are skipped. - Greptile is 5/5 on the same commit, with zero unresolved review threads. The corrected review uses the actual master-to-head diff. - The local `pnpm test:run` started and was stopped after the full CI matrix passed. It did not complete locally; the full-suite result above comes from CI. - Earlier live acceptance used an isolated company with a capped provider credential. The agent opened paperclip.ing, the embedded viewer accepted navigation, and the same browser stayed available after completion and tab switches. - The local Storybook build passes. Stories cover the panel, footer, feed entries, settings, lifecycle failures, and viewport modes with an offline viewer fixture. ## Risks - Browser Use charges for hosted work. Provider caps and local budget checks reduce exposure; reported costs can arrive after work completes. - Viewer and CDP URLs grant access to the browser. The server validates and restricts them. They are excluded from agent results and durable event data. - Runtime resizing of v4 agent browsers uses a provider option confirmed by live testing but absent from its published agent schema. Resizing during a click may invalidate coordinates. Fixed presets remain available. - Viewport ownership is process-local and resets on restart. The lifecycle and accounting records remain in the database. - The original intermittent embedded-viewer stall has not been fully diagnosed. A bounded reconnect and active-session recovery cover the observed failure paths. - Live tests did not cover every revocation, approval, rate-limit, or restart case. Deterministic integration tests cover those paths. Passkey forwarding is not claimed. - Unknown create outcomes keep the credential available for cleanup. Run-list absence cannot prove a paid POST was rejected, so recovery stays pending until it can identify provider work. ## Model Used OpenAI Codex, GPT-6. Used reasoning, repository search, code execution, browser interaction, and test tools. The exact serving model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
96bf004a79 |
fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bd6caf51bb |
fix: preserve restore failure results and stop unsafe retries (#14035)
Preserve agent output and earlier execution errors when workspace restore fails. Report the restore phase and confirmed saved-plan links. Require verified repair before retrying unsafe archives, while preserving approval states and the retry budget. Verified with full CI, 506 focused regression tests, and Greptile 5/5 with all review threads resolved. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
74a9730acb |
fix: continue native agent chats after worker loss (#13813)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat uses native workers to run Claude and Codex conversations. > - A worker crash leaves a cleanup hold because its provider did not acknowledge suspension. > - A new user message must not reuse that unverified session or repeat old tool calls. > - The existing continuation path can preserve history and start a fresh session, but local cleanup ownership remained held. > - This pull request verifies the stopped local owners and releases only their cleanup hold for a new user turn. ## Linked Issues or Issue Description Refs #13775. **What happened?** After a native worker crashed, both providers retained cleanup quarantine. A saved plan survived, but the conversation could not produce another answer. **Expected behavior** Once the old worker and provider process groups have stopped, a new user message can continue in a fresh session with the saved work and prior action history. **Steps to reproduce** Run the opt-in `agent-chat-qualification` suite with case `worker-crash-retry` on native Codex and native Claude. The fixture saves a plan, kills the exact worker through a Linux pidfd, releases a local read-only brief, and sends a new message. ## What Changed - Verify the exact local worker stop receipt, provider identity receipts, released leases, and retained state before retiring a native cleanup hold. - Recheck process liveness and state before admission. Keep the old run and durable session files intact. - Use the existing explicit conversation continuation path. Generic Retry remains blocked for cleanup quarantine, including on the old failed-run marker after a successful continuation. - Extend the live oracle to require a successful fresh session, correct predecessor context, unchanged plan, one original message, and one answer containing a reference introduced after the crash. - Add physical-proof and database-backed admission tests. Document the precise qualification scope. ## Verification - Live Product E2E: **2/2 passed**, **2/2 cleanup passed**, with real native `gpt-5.6-sol` and `claude-sonnet-5`, Chromium, server, database, and public APIs. - Core recovery proof source: `3592b04c2bc76e23795fcdf964720e38a409dc4d`. Suite definition version 8: `9867367994d81a0c726956d91f2c7fddab6417a12f41b5cef3f7e62f3be417da`. - [Core recovery campaign](https://github.com/paperclipai/paperclip/actions/runs/35741746990) · [Public evidence report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35741746990-1/). - Both cells verify the real crash boundary, blocked generic Retry, unchanged saved plan, one original prompt, one fresh successor with predecessor context, and one run-attributed answer containing the post-crash reference. Billing coverage is partial because the crashed runs did not report complete usage; missing cost is not zero cost. - [Master baseline](https://github.com/paperclipai/paperclip/actions/runs/35737364443): both providers stopped at quarantine, with cleanup passing. - [Follow-up baseline without the fix](https://github.com/paperclipai/paperclip/actions/runs/35738638866): Codex produced a complete red result. Claude reached the same error, but its artifact upload was canceled. - [Complete Claude baseline with the same version-8 definition](https://github.com/paperclipai/paperclip/actions/runs/35740555177): red at cleanup quarantine, cleanup passed, source `6479a90c5754044356b39a9278b9a3e92ce8e55e`. - Earlier candidate attempts remain retained: [first](https://github.com/paperclipai/paperclip/actions/runs/35738449214) passed Codex and found a Claude fixture wait race; [second](https://github.com/paperclipai/paperclip/actions/runs/35740408065) exposed the normalized session-open receipt mismatch. Both corrections are in the final source. - Targeted server suites: 570 passed before the final two additional receipt regression cases. The physical-proof suite, including those cases, passed 46/46. Eval oracle and catalog: 40 passed. Server and Product E2E typechecks passed. - **All 54 PR checks passed on final head `db6f775df`**, including repository typecheck, build, tests, browser shards, and canary dry run. The canary job required one retry after its runner received a shutdown signal. On the earlier core proof head, two timing-sensitive tests passed in isolation and on a single CI retry. - Final UI regression checks: 5 passed; UI typecheck and token gates passed. Updated eval oracle/catalog: 40 passed; eval typecheck passed. - Final version-9 two-provider campaign: **2/2 passed, 2/2 cleanup passed**, including the browser assertion that the quarantined historical run never regains Try again. [Final campaign](https://github.com/paperclipai/paperclip/actions/runs/35747416013), source `db6f775dfff405e1514ec02fedb0450d42c7dad2`, definition hash `bf5abf1cc45cb6dad4e082fbf818b8fa0f4d8c282776a7e98762b22919eadab7`. [Final public evidence report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35747416013-1/). Final screenshots and retained state inspected for both providers; cost coverage remains partial. ## Risks - Missing or conflicting stop evidence keeps the conversation blocked. This change does not kill an unverified process. - This qualifies new user input after local worker loss. It does not enable automatic replay, exact-session recovery, remote crash recovery, or native onboarding defaults. - Old action outcomes remain part of the continuation. A process exit is not proof that an action did not happen. - No schema migration or production prompt change. ## Model Used OpenAI Codex, GPT-6. The runtime does not expose a more specific model identifier or context-window size. Used reasoning, repository inspection, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
df66219780 |
fix(ui): retry the latest failed task attempt (#13765)
## Thinking Path > - Paperclip manages work done by AI agents. > - The task thread lets an operator retry a failed run. > - Legacy runs with transcript output did not get a failure marker. > - The thread could therefore offer Try again for an older failure. > - The server correctly reused that failure's existing retry, even when it had already failed. > - This change keeps the latest failure actionable and reports stopped retry responses to the operator. ## Linked Issues or Issue Description **What happened?** Try again could return success without starting work. An initial setup failure had an empty transcript. Its later retry produced output and failed. Only the initial failure had a retry marker, so the button kept requesting the initial failure's already-failed successor. **Expected behavior** Try again targets the latest failed attempt. An already-stopped retry response shows an error and refreshes the task's run state. **Steps to reproduce** 1. Start a legacy adapter task that fails before producing transcript output. 2. Retry it. Let this attempt produce output and a final failure comment before it fails. 3. Click Try again in the task thread. 4. Before this fix, the click targets the original failure and replays the stopped successor. **Paperclip version or commit** Reproduced against `1483bb8bcf`; the regression is also present on the branch base `8813a50105`. **Deployment mode** Authenticated server with a legacy adapter. The bug is in the shared task UI and retry API client. Related: #11650 adds a different recovery-notice action. This change fixes failed-run markers and retry response handling. Searches found no duplicate of this failure case. ## What Changed - Render legacy failure markers even when the run has a transcript or final comment. - Keep later cancelled automatic retries from replacing the failed run's retry action. - Reject already-stopped retry responses in the API client so existing error feedback appears. - Refresh task run queries after both successful and failed retry requests. - Add regression coverage for failed and timed-out attempts, execution gates with output, and retry response states. ## Verification - Before the fix, the new regression tests failed: two selected the original failure, and four accepted a stopped successor as success. - Targeted task-thread, retry API, marker, and issue-page tests: 279 passed. - `pnpm --filter @paperclipai/ui typecheck`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `pnpm check:token-gates`: passed. - Full UI suite: 6,529 tests passed across 626 files. - [CI run 35650385023](https://github.com/paperclipai/paperclip/actions/runs/35650385023): all 53 checks passed, including full workspace build, typecheck, unit/integration suites, and browser tests. The redundant local full-workspace test run was stopped after CI passed; it is not counted as a completed local pass. - Greptile: 5/5 on `608ee58c99`, with no review threads or unresolved comments. The branch is mergeable. - `pnpm -r typecheck` and `pnpm build` were attempted. Both stop at the Runner's Rust checks because this host has no `cargo`. The full workspace checks passed in CI. ## Risks Low risk. This changes UI presentation and response handling only. The server's exact-retry idempotency, authorization, execution ownership, and recovery gates remain in place. No schema changes or live task mutations. Existing documentation describes this retry action; the fix restores that behavior. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted checks; full-workspace limits described above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (existing behavior restored; no documentation change needed) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
43acbcc398 |
fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task state to provider sessions. > - Follow-up turns must retain provider memory and carry new user direction. > - Lost session IDs caused repeated context and extra input tokens. > - Native question answers and approval races could leave valid work blocked. > - This pull request repairs those paths and adds regression coverage. > - Agents can continue accepted work without repeating the conversation or losing the user's answer. ## Linked Issues or Issue Description Refs #13574. That merged PR shortened continuation prompts and moved question instructions into tool documentation. This change preserves sessions and fixes failures exposed by broader testing. Related runtime work: #13408 and #13410. **What happened?** Native follow-up turns could lose the provider session ID. Completion guidance could replace the original task with its latest comment. Claude native questions could remain pending after the user answered. Approval during a running tool call could suspend the run before the tool response arrived. Onboarding and chat handoff instructions also caused repeated planning or missing plan documents. **Expected behavior** Reuse a valid provider session. Send only new events when that session already has the history. Preserve the task requirements and apply later user direction. Store the question answer and deliver it to the waiting run. Finish governed tool responses before suspending. Execute the accepted plan without asking for the same approval again. **Steps to reproduce** Run the continuation, local-session-integrity, first-task, and agent-chat suites with native Codex and Claude. Include provider-question-bridge, accept-while-running, and plan-handoff. **Paperclip version or commit** This branch is based on master |
||
|
|
45586170e1 |
fix(ui): reveal task history while retries wait to start (#13597)
## Thinking Path > - Paperclip lets people manage AI agents and inspect their work. > - Task conversations combine comments with run transcripts. > - The first reveal waits for the relevant history to load. > - Scheduled retries have not started, so the log reader does not hydrate them. > - Waiting for those retries keeps the whole conversation hidden after its header loads. > - This change reveals available history while a retry waits to start. ## Linked Issues or Issue Description **What happened?** A task header loads, but the conversation stays behind the loading overlay while a linked run has `scheduled_retry` status. The readiness check expects that run in `hydratedRunIds`, although the log reader does not read its log. A retry delay can therefore become a task-loading delay. **Expected behavior** Show existing comments and transcripts once their data is ready. A retry that has not started must not block the conversation. **Steps to reproduce** 1. Open a task with existing comments and a linked run waiting in `scheduled_retry`. 2. Let the comments and other run transcripts finish loading. 3. Observe that the header loads but the conversation stays hidden until the retry changes status. **Paperclip version or commit** Reproduced on master at `3f1d897a7` with regression tests. **Deployment mode** Web UI, built from source. This is a client readiness bug. Searched public issues and PRs for scheduled retries, conversation loading, and history loading. No duplicate found. Related: #10255 changes the server response for active runs with no log; it does not address this scheduled-retry readiness gate. This focused bug fix does not add a roadmap feature. ## What Changed - Exclude `scheduled_retry` runs from the initial task-history readiness gate. - Extend transcript mocks to expose per-run hydration state. - Add six regression cases across legacy and native runs. Existing history loads during retry waits, while unhydrated running and completed runs still block the first reveal. - Document the exception in a code comment. No user command or configuration changes need documentation. ## Verification - All six new cases fail before the production fix. - `pnpm --filter @paperclipai/ui exec vitest run src/components/TaskChatThread.test.tsx src/components/transcript/useLiveRunTranscripts.test.tsx src/components/transcript/useNativeRunTranscripts.test.tsx`: 164 passed. - `pnpm --filter @paperclipai/ui typecheck`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `pnpm check:token-gates` and `git diff --check`: passed. - `pnpm -r typecheck` and `pnpm build`: attempted; both stop in the native runner package because the local machine has no Rust `cargo` executable. UI checks pass separately. - `pnpm test:run`: started locally, then stopped before completion after full CI passed on the same commit. The local full-suite result is incomplete. - CI on `46daa8b06`: all 53 checks passed; two optional Storybook jobs were skipped. This includes the full test matrix, build, typecheck, native runner checks, browser tests, and canary dry run. - Greptile: 5/5 on `46daa8b06`, with no review threads or actionable findings. The branch has no merge conflicts with master. ## Risks Low risk. The exception applies only to runs waiting in `scheduled_retry`. Running and completed runs retain their existing readiness checks. Retry scheduling, transcript fetching, and server behavior stay the same. No migration is needed. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and test tools. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9cfa7fd2d1 |
fix(ui): share compact live and saved activity across runners (#13421)
## Thinking Path > - Paperclip helps people manage AI agents and inspect their work. > - Task threads show live activity and saved transcripts from several runners. > - Legacy runs used a separate activity renderer that expanded tool cards as commands arrived. > - A compact live row makes ongoing work easier to follow. > - This PR shares the native runner activity group across live and saved legacy turns. > - Both paths now use the same labels, icon gutter, animation, and history controls. ## Linked Issues or Issue Description Related: Refs #13255, which introduced rolling native-runner activity groups. **What happened?** During a legacy CLI run, each new command added an expanded tool card under an activity heading. The original legacy parity story rendered a completed turn, so it did not exercise this live path. **Expected behavior** Show one current activity line per commentary group. Roll that line forward when a new activity starts. Keep tool icons aligned on the left, use friendly labels, and show history only when expanded. **Steps to reproduce** 1. Start a task with the Codex local adapter using the CLI engine. 2. Ask it to read two files in separate tool calls and run a test command. 3. Watch the activity feed while the run is active, then expand its history. **Paperclip version or commit** The live behavior was reproduced on bdee5ebb21b2d09280e49c88bc329435da1b07c8. This update includes both live and saved rendering fixes. **Deployment mode** Built from source with the local test-drive command and a real Codex CLI agent. ## What Changed - Use `TaskChatRunnerActivityGroup` for both saved legacy phases and `TaskChatLiveTail`. - Align the legacy Working spinner with the shared activity icon gutter. - Cover streamed reasoning updates, successive commands, image labels, hidden details, and explicit history expansion in tests. - Feed raw legacy transcript events through the real adapter and live renderer in Storybook. Add live, expanded, narrow, light, and completed stories with the final reply preserved. ## Verification - Passed 195 targeted tests covering the live tail, status pill, shared activity group, native turn, and full task thread. - Passed UI typechecking, token gates, UI build, and Storybook build before submission. - Observed a real Codex CLI run while it read separate files and ran tests. The current activity stayed at one 32-pixel row without an accumulated tool list; the spinner and activity icon centers aligned. - Compared native and legacy Storybooks: identical rolling animation, fixed height, persistent expansion, truncated long labels, and correct light and completed states. - Full repository `pnpm -r typecheck` and `pnpm build` passed after replaying the branch on current master, including the Rust runner. The full `pnpm test:run` suite is still running. ## Risks - Legacy activity now starts collapsed. Users can expand each group and each row to inspect the same transcript details. - Live runtime request cards must retain their timeline positions. Existing task-thread and native-runner regression tests cover this boundary. - No API, database, or transcript format changes. ## Model Used - OpenAI GPT-6 through Codex for the live-path fix, tests, and browser verification. Capabilities used: reasoning, tool use, local code execution, and image inspection. The exact deployment ID and context-window size are not exposed in this session. - The original saved-turn change recorded OpenAI Codex with GPT-5.6; its exact variant and context size were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4cdc7b231 |
fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane prepares task workspaces before it starts an agent. > - Workspace preparation reads Git state so it can preserve edits and exclude private files. > - A failed scan was treated as a non-Git folder and lost its actual failure code. > - The resulting generic setup failure could not recover, even when the cause was temporary. > - This pull request keeps the cause and uses the existing bounded retry schedule before provider startup. > - Tasks can recover without human intervention, while permanent failures and exhausted retries stop with useful guidance. ## Linked Issues or Issue Description **What happened?** A Git scan error during managed repository preparation became `Configured repository folder is not a Git checkout`, followed by generic `setup_failed`. The agent never started. Generic recovery could not distinguish a temporary timeout from a bad workspace configuration. **Expected behavior** Keep the closed scan error code. Retry temporary timeouts and queue saturation under the existing shared budget. Preserve edits, exclusions, ownership, and pause gates. Stop permanent failures and exhausted retries with a specific explanation. Do not replay historical generic setup failures. **Steps to reproduce** 1. Configure a task project with a local Git source that must be copied into its managed repositories. 2. Make the ignored-file scan exceed its timeout before the agent starts. 3. Before this fix, the snapshot returns null and the run ends as non-retryable `setup_failed`. 4. Use the disposable browser fixture in `tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and test the full recovery path. **Paperclip version or commit** Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`. **Deployment mode** Built from source. The defect is in core workspace setup, not a specific model provider. Related work: Refs #13442 (managed repository preparation), Refs #11572 (bounded Git scheduler), Refs #12997 (separate adapter startup retry work), Refs #13469 (separate terminal-workspace scan performance work). ## What Changed - Return the non-Git fallback only for repository discovery. Propagate failed scans of a confirmed repository. - Replace full ignored status output with an ignored-only directory listing. Preserve NUL-delimited paths and exclusions. - Preserve typed, sanitized scan errors through workspace preparation and persist pre-provider failure details. - Retry only timeouts and queue saturation, using the existing durable two-retry budget and issue gates. Prevent generic recovery from adding another budget. - Show workspace-specific failure copy and actionable exhausted-recovery notices. - Add red-green unit tests, real-database restart and retry-boundary tests, and opt-in browser acceptance fixtures with real Git subprocess timeouts. - Document the recovery contract and browser verification procedure. ## Verification - Red: injected scan failures returned null instead of rejecting; setup lost the timeout code; task-thread and recovery notices had generic copy. - Green: 119 focused adapter/backend tests, 20 recovery-boundary tests, and 136 task-thread tests. - `pnpm -r typecheck` — passed. - `pnpm build` — passed on the final production code. - `pnpm check:token-gates` — passed. - The initial local `pnpm test:run` overlapped source edits and was interrupted after two late-added assertions saw pre-fix behavior; it is not counted as a green full run. A fresh final-head run passed all 229 tests across the six affected adapter/backend/UI suites. The clean latest-head CI full test matrix passed: all five general-server shards, all five serialized-server shards, and all three general-workspace shards. - Latest-head CI also passed all three browser shards and their aggregate gate, typecheck and release registry, build, runner verification, canary dry run, policy, Docker context integrity, and security gates. Greptile: 5/5, with the review thread resolved. - Browser: created a task in a disposable instance. A real Git timeout scheduled recovery, the next run completed through the run-scoped API without manual Retry, and Done survived reload. The deterministic process worker checked preserved source edits and excluded private files; no model calls were made. - `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec playwright test --config tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6 minutes). The persistent case made exactly three failed attempts, never started the worker, showed the cause-specific notice, stayed stopped for another scheduler tick, and retained Blocked after reload. - Extra red-green coverage: 50 recovery tests passed after fixing an exhausted-bootstrap classification that incorrectly implied unknown provider actions. Missing or uncertain evidence still retains the safety hold. - Verified the documented Git executable override during repository seeding. ## Risks - A confirmed repository scan failure now fails closed instead of falling back to directory sync. This prevents unfiltered copying but makes previously hidden errors visible. - Temporary host problems can create up to two additional setup attempts, 30 seconds apart. Permanent scan errors do not auto-retry. Generic recovery cannot reset this budget. - The durable retry path still enforces ownership, pause, and work eligibility. Integration tests cover restart, duplicate promotion, pause, exhaustion, and non-retryable categories. - No schema migration, new runtime setting, new retry budget, production deployment, or historical task replay. ## Model Used OpenAI Codex, GPT-5-based coding agent, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
78ce96a48b |
fix(runner): restore native Claude context and read permissions (#13422)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner supplies each agent with instructions, assigned skills, and tools. > - Native Claude lost its skill snapshot before provider launch. It also changed exact model IDs to aliases. > - The read permission mode denied ordinary Paperclip reads because this runner has no interactive approval handler. > - This pull request restores the missing context and permits only assigned tools that Paperclip defines as reads. > - Other requests that need approval stop with a clear action for the operator. They do not retry automatically. > - Agents can complete read tasks. Users can see why a restricted task stopped. ## Linked Issues or Issue Description **What happened?** Native Claude could not find assigned skills. The ACP layer could replace an exact model ID with an alias and then fail identity verification. The default `approve-reads` mode denied Paperclip read tools. A denied write showed a generic transport failure. The tool descriptions also called live operations “mock” operations. The completion prompt said to call a completion tool once, although the protocol can reject a claim and require a corrected call. **Expected behavior** Load assigned skills before launch. Keep the selected model ID. Allow assigned Paperclip reads. Stop an operation that needs approval with clear instructions when no approval handler exists. **Steps to reproduce** 1. Configure a native Paperclip Runner agent with ACPX Claude and an exact model ID. 2. Assign a skill and ask the agent to use it. 3. Select the read permission mode and ask the agent to read task context and list documents. 4. Ask the agent to write a document. Check the task state and recovery message. **Paperclip version or commit** Reproduced from `d351e08deee1b49d3467a950d1a3f01131943441`. **Deployment mode** Local native runner with real Claude, Rust runnerd, the Paperclip server, and embedded PostgreSQL. Related: #13196 fixed remote skill staging in the legacy `claude_local` adapter. The native runner uses a separate path, which this pull request fixes. ## What Changed - Carry the runtime context through Rust and the ACPX sidecar. Load assigned Claude skills after the provider lifetime lease is held. Refresh the files on each open. - Write the exact requested model ID into the isolated Claude settings. - Grant exact MCP permissions for the intersection of assigned tools and Paperclip's read catalog. Provider hints cannot grant access. Existing task-control permissions stay in place. - Stop requests that need an unavailable approval handler. Preserve the typed error through the server. Show “Approval required” on the task and require operator action without automatic retry. - Label the setting “Allow Paperclip reads.” Remove “mock” from live tool descriptions and regenerate the contracts. - Change one completion-prompt sentence to require one accepted result. Add regression tests and update the runner documentation. ## Verification - Red/green regression tests cover read admission, unavailable approval handling, server recovery, and the task error message. - A real Claude read trial failed before the fix and completed after it: [red trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=95e6a326e24437f1), [green trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=85a1f347db5604d1). - Full browser tests used real Claude and the production tool authority. Reads completed. A write stopped with “Approval required.” No document was created and no automatic retry was scheduled. All nine checks passed: [events and screenshots](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=1c4ff10be0d779cc). - A separate full browser test assigned a skill, invoked it, and completed with a marker absent from the task prompt: [skill evidence](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=ba5fbcb29a60be52). - Post-rebase checks passed: 128 targeted runner tests, 46 Rust tests, 27 sidecar/protocol contract tests, `pnpm -r typecheck`, and `pnpm build`. The server/UI red-green checks and `pnpm check:token-gates` also passed. The full suite passed in CI, including all general and serialized test shards, all three browser shards, and runner `check:all`. The duplicate unsharded local full-suite run was stopped after CI passed. Braintrust links require project access. The traces contain provider/runner events and app outcomes. They do not contain raw model HTTP requests. ## Risks - `approve-reads` now allows assigned Paperclip reads. Other operations that need approval stop the turn. Users must review the operation and change permissions before retrying. - The Claude settings depend on the pinned ACP and SDK behavior. Automated tests and real Claude trials cover this boundary. - The native Codex path is unchanged. The ACPX Codex fallback shares the clearer approval failure handling. - Local execution was tested end to end. Remote execution was not run. This change has no database migration. ## Model Used OpenAI GPT-6 through Codex. The agent used reasoning, code execution, and browser tools. The exact deployment ID and context-window size were not exposed in the session. Live acceptance tests used `claude-sonnet-5`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6809314a3f |
fix: recover sandbox workspace setup and retries (#13353)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - Sandbox tasks need a workspace and a provider session before work can start. > - A selected Git subfolder is a valid workspace, but it is not a repository fetch source. > - A resumed sandbox must keep one resource identity when the provider fills in its region. > - Workspace reuse does not prove that a provider session has started. > - A recorded continuation must not leave an old failure blocking Retry for a newer attempt. > - This pull request fixes those setup and recovery boundaries while preserving ownership checks. ## Linked Issues or Issue Description **What happened?** A Daytona task failed before provider startup when its project folder was inside a parent Git repository. After the folder was repaired, Retry resumed the same sandbox and verified its workspace sentinel, but workspace preparation reported that the lease was no longer active. A later attempt could fail because the resumed workspace had no provider checkpoint. The old run also kept a recovery projection after the server recorded its explicit successor, hiding Retry for the new failure. **Expected behavior** Sync the selected folder without importing parent files or history. Keep the existing sandbox identity stable. Allow a new provider session only with proof that its exact session has never started. Let the current failed attempt retain Retry once the old recovery has a recorded successor. **Steps to reproduce** 1. Select a subfolder of a Git repository as a project workspace and start a Daytona task. 2. Leave the target region unset. Release its reusable sandbox after setup fails, then resume and realize the workspace. 3. Retry a provider setup that failed before any session directory or checkpoint was created. 4. Record an explicit successor for a native failure, fail that successor during setup, and inspect Retry in the task thread. **Paperclip version or commit** Reproduced against `8d1f0c20a` with new failing regressions before each production fix. **Deployment mode** Local controller with a Daytona environment. Automated tests use deterministic provider fixtures and real filesystem operations. Related: #13338, #13163, #13264, #13349. The open work-folder stack in #13264 includes a broader fresh-session authority change. This patch addresses the independently reproduced startup failure with existing durable bootstrap proof and atomic directory creation. It does not include the work-folder migration or credential changes from that stack. ## What Changed - Classify only the selected Git repository root as a fetch source. Sync subfolders as directories and preserve their enclosing Git ignore rules on upload and restore. Recognize the shared scheduler’s completed non-repository result so ordinary folders still sync; timeout, cancellation, and output-limit failures remain closed. - Remove the placement target from Daytona account cache identity. Keep API endpoint, credential digest, company, environment, driver, and sandbox ID boundaries. - Preserve closed-lease admission until a sentinel-verified resume reopens that same resource. - Show the already-recorded explicit successor of a resolved recovery. Keep the old failure evidence and unresolved holds. No historical status writes occur. - Permit a resumed workspace to create a new session directory only with matching durable identity, zero connections and events, untouched bootstrap commands, no backup, and an absent remote session. Claim the directory atomically. Existing, partial, or ambiguous state still blocks startup. ## Verification - Final head `c43c7403f`: [CI completed successfully](https://github.com/paperclipai/paperclip/actions/runs/34733846820/attempts/2), with 32 successful checks and 2 conditional skips. This includes every server, UI, package, serialized-route, browser, native-runner, typecheck, and build gate. Greptile reviewed the same head at [5/5 with no unresolved findings](https://github.com/paperclipai/paperclip/pull/13353#issuecomment-5650270168). - New regressions failed before each of the four production fixes. Git/archive/restore suites: 148 passed. The scheduler-wrapped non-Git regression also failed before its fix; 135 affected Git/sync/Codex tests then passed. - Daytona plugin: 230 passed, 6 opt-in live tests skipped. Native executor, projection, and TaskChatThread: 494 passed, including existing/partial state, wrong identity, prior connections or turns, backups, unavailable proof, and unresolved recovery controls. - Local full-repository typecheck, build, and token gates passed on the final head. Local CLI: 485 passed. Complete single-worker package rerun: 3,224 passed, 19 skipped. - Local verification is an aggregate with recorded retries, not one pristine green invocation: the general-server run began on `e721a920a` and finished with 12,002 passed, 3 failed, 70 skipped. Its real Codex scheduler failure is fixed above; the socket and workspace-runtime timeout failures passed unchanged in focused reruns. Both Inbox failures passed unchanged in the full 27-test Inbox file; database/shared-package failures passed in the single-worker package rerun. The supplemental local serialized-route rerun remains in progress; all five corresponding final-head CI lanes passed. - The first final-head CI attempt hit a Daytona fixture-readiness race and a signoff-browser heartbeat receipt timeout. One supported unchanged failed-job rerun passed both and the aggregate gates. Live combined user-journey verification is tracked in the related follow-up; this PR's provider tests use deterministic fixtures and real filesystem checks. ## Risks - A selected subfolder uses directory sync and does not carry parent Git history. Its ignored files stay local. - The target region remains a creation setting and part of workspace reuse policy; it does not split the account identity of an existing sandbox. - Incomplete or conflicting provider state still fails closed. This change does not erase a session, infer completed work, bypass a user decision, or replay uncertain actions. - A later remote setup failure can leave a claimed partial session directory. It remains blocked rather than being overwritten. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact hosted model ID and context-window size are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
422287eecd |
fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task messages, provider execution, and task outcomes. > - First-time user tests exposed gaps in recovery, completion permissions, message delivery, and Stop behavior. > - These gaps left usable output hidden, completed work waiting for bookkeeping, or safe work unable to continue. > - This pull request fixes the shared lifecycle and receipt paths while preserving process ownership and action checks. > - Users can continue work with accurate task state and durable messages. ## Linked Issues or Issue Description **What happened?** A stopped local Codex execution could remain blocked even after its processes had stopped and its complete transcript proved that no external action needed replay. Claude under Conservative permissions could fail to call task completion tools. Recovery could reuse an assistant item ID and overwrite prior output. A delivered comment could remain marked uncertain after navigation. Stop could look like Pause or a new recovery incident. Workspace contention could look like cancellation. A direct reply reopening Done could enter a clarification loop. **Expected behavior** Recover automatically only with verified termination and complete action receipts. Preserve answers and messages. Keep task completion available under Conservative permissions without broad tool access. Show crashes as Blocked, actual human decisions as In Review, and ordinary workspace contention as waiting. Stop the current response and allow a new direction. **Steps to reproduce** 1. Create ordinary response tasks with local Codex and Claude Code, then send follow-up messages through the task composer. 2. Interrupt a disposable local Codex runner during text-only work. Verify automatic continuation and retained output. 3. Stop a response, send a new request, answer a clarification, and reopen completed work with another message. 4. Navigate or reload while a comment submission is pending. Confirm the exact persisted request receipt settles it without removing newer draft text. 5. Run two tasks in a shared Daytona workspace. Confirm waiting does not appear as failure. **Paperclip version or commit** Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`. Current integration base: `6cef9743c`. Both operator-interruption and workspace-waiting guards are preserved; native restart and legacy permission rules remain documented. **Deployment mode** An isolated source-built test-drive instance, with real local Codex and Claude Code providers and disposable Daytona environments. Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163. This PR addresses additional failures from ordinary task journeys, including controller restart handoff and repeated warm sandbox setup. Historical task status reconciliation is excluded. ## What Changed - Persist runner ownership immediately at spawn and resume an explicitly adopted runner even when the controller crashed before the first driver checkpoint. Detach the controller safely across graceful restarts, including session startup. Prevent an old finalizer from suspending or signaling an adopted runner. Checkpoint idle warm sessions before shutdown. Preserve the same run and queued follow-up messages. - Scope saved legacy queue successor checks to the queue owner while preserving ordinary task locks, operator identity, assignment gates, and exactly-once delivery. - Preserve managed Codex credential files when an old session is detached for restart; normal owned cleanup still copies refreshed auth back and removes the scoped copy. - Reuse the bound warm shared sandbox and fully verify an existing staged provider pack before using it. This avoids repeated uploads when the pack is already valid. - Add a narrow local Codex replacement path with stopped-process proof, a closed transcript inventory, exact completion receipts, and fresh-session lineage. Preserve no-replay holds when evidence is incomplete. Recovery may clear only the same run's recorded Blocked status version; manual re-blocking and dependency changes invalidate that receipt, while queued comments do not. Later blocks stop scheduled, queued, and final dispatch; queued/final checks re-read dependencies even when the task status stays In Progress. - Permit only task delivery and human-input tools through the isolated Claude runner's exact task bridge. - Scope assistant item identity to the provider turn and ignore only authority-free Codex skill-change notifications during startup. - Reconcile composer submissions by client request ID across response loss, navigation, and reload. Retain text typed during delivery. - Keep acknowledged run-only Stop neutral and show workspace contention as waiting. Project exhausted native failures as Blocked. - Restore the guarded task-page retry action for failed legacy runs, including the server-supported explicit new-attempt path for stopped conversation adapters. Preserve native/process recovery holds and avoid promising Retry while a decision or execution gate hides it. - Refresh delivered artifacts and handle direct user replies that reopen completed work without a clarification loop. - Check the embedded PostgreSQL PID, data directory, and actual port before connecting or migrating. - Document accepted behavior and add focused regressions at lifecycle, route, transcript, and UI boundaries. ## Verification - Final head `fece606ac2` passes the complete GitHub CI matrix: **34 green checks, two expected Storybook skips, no failures or pending checks**, including `ci / verify`, `ci / e2e`, full runner verification, typecheck, build, every server/workspace shard, and all browser shards. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34727183287). Greptile is **5/5 with no open findings**. The final two commits only refine test fixtures; both affected suites pass 24/24 locally and in CI, with server typecheck green. - Complete local Vitest coverage uses the canonical groups/shards: all 635 general server suites, all 145 serialized suites, and all workspace packages. The aggregate began on `0a8001c18` while the final queue fix arrived: 23,903 passed, five failed, 87 skipped. The five port/socket/timing failures passed unchanged in follow-ups (60 tests in the exposure/file suites and 412 tests covering the serialized failures and unrun tails). The final queue/operator-identity suites separately passed 52/52. This is aggregate coverage plus explicit reruns, not a pristine single-command final-head run. - After integration with current master, queue/operator-identity/continuation suites passed 162/162 and affected UI suites passed 140/140. ACP Stop/continuation and legacy task/Inbox/message browser suites passed 9/9, including both task recovery Retry and thread Try again, automatic saved-message delivery, exactly one new run, Done, and retained output after reload. The default process Stop/Pause/Resume browser case passed (the native-provider case is opt-in and skipped by default). The complete Board attachment/receipt browser suite passed 11/11 on a disposable instance, covering both composers, exact receipts after lost responses, no replay, bound attachments, and newer drafts after reload. - Blocking-intent regressions cover pre-existing Blocked, a mismatched run/cause, an explicit manual re-block, changed dependencies, a queued comment after failure, and a block arriving between scheduling and provider dispatch. The negative cases reproduced before the fix. All 478 affected executor/recovery/dispatch tests passed; both database suites ran separately after availability-probe skips in the first combined command. The final late-dependency check passed all 143 affected recovery/dispatch tests (zero skips) after two new negative cases reproduced the bug. - Focused runtime regressions cover awaited runner ownership publication, authenticated adoption before the first checkpoint, old-finalizer detachment, idle and busy warm-session shutdown, rejected checkpoint propagation, provider-pack verification, and managed-Codex credential preservation. Four managed credential detachment cases reproduced the bug before the fix; normal owned cleanup still succeeds exactly once. - Live local Claude: SIGKILL 2.6 seconds into startup recovered the same run automatically in 53 seconds, then a normal follow-up completed in 24 seconds. SIGTERM 2.5 seconds into startup preserved the same run (54 seconds) and its queued follow-up (21 seconds). Answers remained visible and the task reached Done. - Live Claude Daytona: a warm follow-up retained its sandbox and fell from 121 seconds to 44 seconds. A separate cold turn took 127 seconds; after controller shutdown and checkpointing, its follow-up completed in 33 seconds with the same sandbox, workspace, native session, and runner. Both answers remained visible and the task was Done. - Other live journeys covered task completion and follow-up with local and Daytona Codex, local Codex crash recovery, Stop then new direction, clarification response, live artifact refresh, and shared-workspace waiting. - Validation limits: the opt-in native composer Stop/Pause→subtree Resume fixture exposes terminal/result ordering and subtree-cancellation attribution bugs that can leave a child task blocked; that new finding is assigned to a separate follow-up and is not claimed fixed here. Default CI skips this optional native-provider fixture. Managed-Codex credential handoff and the queue-agent integration use automated regression evidence. Cold custom provider-pack uploads still add startup latency. ## Risks - Automatic replacement remains deliberately narrow: local Codex, verified stopped identities, unchanged retained state, and a complete text/completion-only turn. Unknown actions, partial history, or changed ownership remain blocked. - Claude completion permission handling changes an upstream package patch. The exact isolated task bridge must remain pinned; unrelated tools keep their existing permissions. - New task failure projection changes user-visible status. No historical status backfill or database migration is included. - This is a broad lifecycle fix across server and UI. Live proof covers graceful local Claude restart during startup and idle Claude Daytona session recovery across controller shutdown. Live abrupt SIGKILL during local Claude startup also recovered the same run. Unknown ownership or missing action evidence still blocks reuse. Cold custom provider-pack uploads still add startup latency; this change avoids unnecessary repeat uploads. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, browser automation, and tool use. The exact hosted model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
df984cbc2c |
fix: dispatch queued legacy messages with operator identity and task permissions (#13315)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task conversations save messages that arrive during an active turn. > - Legacy adapters deliver these messages in a later turn. > - A run can stop before the saved queue is delivered. > - The Interrupt button previously required an active run, so it could not release this queue. > - This pull request lets a board operator send the saved queue after the run stops and retries queues missed during finalization. > - Manual dispatch must use the clicking operator and must not require permission to create agents. > - The task can continue without a duplicate message or a second execution owner. ## Linked Issues or Issue Description **What happened?** A legacy task retained a queued message after its run stopped. Interrupt was disabled because the queue had no active target. Finalization and deferred message admission can also leave a queue without a successor. **Expected behavior** Interrupt sends the saved messages when no runner is active. Messages that arrive during normal completion are delivered automatically. An uncertain previous execution still requires proof that its process or sandbox stopped. **Steps to reproduce** 1. Queue a user message during a legacy conversation turn. 2. Let the turn stop or simulate a server restart before queue promotion. 3. Open the task with a deferred queue and no active run. 4. Try Interrupt. Before this change, the button is disabled. **Paperclip version or commit** Reproduced against `8f40b4ad4`. **Deployment mode** Legacy conversation adapter. The same persisted queue state is covered with an isolated PostgreSQL fixture. Related public work: #13275 adds active legacy interruption. #13291 addresses automatic sandbox conversation recovery. This change handles explicit saved-queue delivery and late queue promotion. ## What Changed - Accept a null Interrupt target while retaining queue identity, revision, company, and assignee checks. - Save the operator's request on the existing queue. Reuse normal admission after verified stop, including older messages, different authors, and queues whose original wake came from the system. - Strip interruption authority from caller-supplied wake payloads. Only the board queue route can persist that authority. - Retry durable interruption requests after restart and deferred queues after legacy cleanup. - Let an explicit Interrupt retry cleanup for its stopped run, including old ephemeral leases that recorded success without a provider stop receipt. Preserve retained resources, other lease owners, and the automatic retry limit. - Preserve the server's waiting explanation when normalizing and combining queue entries. - Revalidate the consumed board queue receipt at dispatch so a different message author does not cause setup failure. - Use the Interrupt user's execution identity for the new run. Preserve original message authors. Validate the receipt independently at startup and inherit the resulting identity on retry. - Persist authenticated board authority for ordinary manual wakes too. Adopting someone else's queued messages cannot switch a manual run to that author's permissions. Strip caller-supplied authority markers and retain private conversation ownership checks. - Keep the clicking user when a manual wake is merged into an older deferred receipt. Update its requester and payload in the same transaction. - Use the same current-queue/revision API on task details and pipeline conversations; show Interrupt after a legacy target stops. - Keep manual wakes out of active runs, including unscoped agent wakes. They receive their own execution identity; a matching receipt requester is not sufficient because an exact retry can retain a different originating identity. - Authorize both existing-agent wake endpoints with `agent:wake`, available to active non-viewer company members. Keep `agents:create` for hiring. Validate the stored task and current assignee before an exact task retry. - Reject viewer Interrupt requests before saving intent or stopping execution. Keep external chat retry authorization and per-action agent/user permission checks. - Preserve edits and discards until dispatch. Prevent another queue promotion when the same agent already has a successor. Keep independent reviewer recovery available. - Suppress cancelled/failed run toasts for intentional operator interruption. Keep ordinary runtime error notices. - Add UI, route, admission, restart, successor ownership, and toast regression tests. Document the behavior. - Reuse the existing socket reservation helper for both credential-quorum test cases after CI exposed an ambient-port collision. This changes test preparation only; production credential staging is still called exactly once. ## Verification - Failing regression tests reproduced the message-author identity bug and an operator's `agents:create` rejection before the fixes. - All 316 focused tests pass across eight route, queue, identity, authorization, continuation, and responsible-user suites, including the 44-test rerun of queue admission and actual startup after the final manual-wake restriction. Regressions reproduce cross-user merging both with and without a task, and same-requester receipt ambiguity. The cross-company existence guard also passes both tests. - Startup integration tests reach adapter execution under the clicking operator and retain that identity through follow-up. Coverage includes mixed authors, adopted queues, system-origin queues, restarts, forged or stale receipts, viewers, suspended memberships, changed assignees, private conversations, and caller-supplied authority markers. - The earlier queue/cleanup/UI regression suite passed 402 tests. The final review corrections pass another 180 tests across queue admission/persistence, real heartbeat startup, UI API, conversation rendering, and pipeline suites. Regression tests reproduced both review findings before correction. The final head has a 5/5 review with no unresolved threads. Full CI passes on `c2002979c`, including every general and serialized server shard, all browser shards, Paperclip Runner verification, typecheck, build, canary dry run, and the aggregate gates. - Full `pnpm -r typecheck`, `pnpm build`, and UI token gates pass after the final application changes. CI identified an outdated task-page API mock after the shared helper extraction; the fixture now exercises the real helper, and all 131 task-page/API tests pass. The final application build passes with the additional manual-wake restriction. - CI exposed a pre-existing port collision in the Codex credential-quorum fixture. It reproduced locally; both listener cases now use the existing bounded reservation helper. All 41 credential tests pass on rerun. One intervening local run hit a separate ambient bind collision in the two-occupied-port case. - The full local `pnpm test:run` attempt was stopped after host contention caused focused-suite timeouts. The affected focused tests passed on rerun. An expiring trace fixture and a missing private-conversation state were corrected. The successful full CI run is the complete-suite verification. - Hosted Interrupt previously cleared the original queue and produced exactly one successor with neutral interruption feedback. It exposed the dispatch authorization defect. Retry on that earlier build was rejected for missing `agents:create` before creating another run. - Deployed the final application build (`38257f391`) to the scoped hosted instance and verified readiness. The latest PR commit changes only the credential test fixture; application code matches that deployment. A live Retry by the same operator without `agents:create` created one successor attributed to that operator, passing the former dispatch permission gate. Startup then stopped at `configuration_incomplete` because that operator has not configured their required personal Claude Code OAuth secret; the post-deployment run page confirms the operator identity and no provider work started, and the My secrets UI still shows the token as not set. Provider execution remains unverified pending that credential. No permission grants or credentials were changed. ## Risks Queue admission and finalization can race. The task lock, durable queue receipt, current comment IDs, and successor guard prevent duplicate dispatch. Process and lease stop checks, task pauses, approvals, ownership, and budgets remain in force. The API change only allows null on legacy Interrupt; native steering still requires an active run. No schema migration is required. Active non-viewer board members can now invoke existing agents without agent-creation permission. Agent self-invocation rules, raw provider-trace admin access, task retry scope, external chat authorization, and action-specific user/agent permissions remain enforced. ## Model Used OpenAI GPT-6 through Codex. The session does not expose an exact backend model ID or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
586b5ec828 |
fix(ui): improve mobile task spacing (#13304)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task detail view is the main place for task conversation. > - The mobile view used too much horizontal space for gutters and controls. > - The status control and identifier also reduced the title width. > - This pull request gives the title its own mobile row and reduces mobile padding. > - The benefit is a clearer task view with more space for useful content. ## Linked Issues or Issue Description **What existing behavior does this improve?** The mobile task detail header, conversation list, and message composer. **Current behavior** The title shares one row with status and identifier controls. The conversation and composer also use larger mobile gutters than needed. **Proposed behavior** The title uses the full mobile width. Status and metadata use the next row. The conversation and composer use smaller mobile gutters. **Reason and benefit** Long titles have more readable line lengths. The reduced padding gives task content more room on small screens. **Breaking changes** None. ## What Changed - Put the task title on a full-width mobile row. - Move the mobile status and metadata below the title. - Reduce the mobile conversation and composer padding. - Update the composer dock class test. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/task-chat/TaskChatComposer.test.tsx ui/src/pages/IssueDetail.test.tsx --reporter=dot` - `pnpm --filter @paperclipai/ui build` - `pnpm check:token-gates` - `git diff --check` ## Risks - Low risk. The layout changes apply only at the mobile breakpoint. - Desktop spacing stays unchanged. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.6, with reasoning and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb9f954bae |
fix(ui): stabilize steered chat activity presentation (#13246)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task chat shows live agent work and user steering. > - Steered messages showed a second timestamp that did not match normal chat messages. > - Folded live work also changed height as new commands arrived. > - Status carets and dots did not use the same vertical alignment. > - This pull request gives these states one stable presentation. > - The benefit is a task chat that is easier to read while an agent works. ## Linked Issues or Issue Description **What happened?** Steered chat messages showed a separate queued or steered timestamp. Folded live activity could grow and shrink. Some carets and status dots did not align vertically. **Expected behavior** Steered messages use the normal timestamp. Folded live activity shows one latest line until the user opens it. Carets and dots use the same vertical center. **Steps to reproduce** 1. Open a task chat with a running agent. 2. Steer the chat and inspect the message timestamp. 3. Keep live activity folded while new commands arrive. 4. Compare the caret with the status dot in the Reconnecting state. **Paperclip version or commit** Current `master` before this change. **Deployment mode** Local development mode. ## What Changed - Use the normal timestamp for steered chat messages. - Keep folded live activity to one line and replace it with the latest activity. - Keep the full activity timeline available after explicit expansion. - Vertically align carets and status dots, including Reconnecting. - Add focused component tests and Storybook review states. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/task-chat/TaskChatRunnerTurn.test.tsx src/components/task-chat/TaskChatStatusPill.test.tsx src/components/task-chat/task-chat-adapter.test.ts` (63 tests passed) - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui build-storybook` ## Risks - Low risk. The change is limited to task chat presentation and its tests. - A live activity item can replace the folded text. The expanded timeline keeps the full history. > This fix does not overlap with planned work in `ROADMAP.md`. ## Model Used - OpenAI Codex. The runtime did not expose the exact model ID or context window. The agent used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ce09ea40b0 |
feat(ui): show one rolling runner activity per commentary group (#13255)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task chat shows commentary and runner activity as work proceeds. > - The expanded activity list uses multiple text lines and an indented rail. > - The folded view can repeat reasoning and the current tool in separate places. > - This pull request shows one rolling activity between commentary messages. > - Users can open a group to follow its full history and inspect each item. > - The benefit is a compact feed with aligned icons and stable row height. ## Linked Issues or Issue Description **What existing behavior does this improve?** The runner activity groups in task chat, during live work and in saved history. **Current behavior** Activity is spread across a group summary, a reasoning ticker, and a current-tool row. Expanded rows have different alignment and can span several lines. **Proposed behavior** Keep one latest activity per commentary group. Roll new activities upward. Keep each expanded history row on one line. Use neutral failure text without an X icon. Keep the full details available by opening an individual row. **Additional context** Refs: #13246. That open PR covers related folded activity, steering timestamps, and status alignment. This change implements the separately reviewed activity-group design, including animated replacement, single-line expanded history, and neutral failures. It does not change steering timestamps or status pills. ## What Changed - Add a shared runner activity group with persistent expansion and reduced-motion support. - Use the group in the production runner turn and remove duplicate narration and current-tool rows. - Preserve commentary, plan cards, request receipts, and final-response placement. - Keep tool output, reasoning, research sources, file previews, and failure details accessible. - Render the production turn in desktop and mobile animated Storybook fixtures. - Add tests for animation identity, failure visibility, expansion persistence, and timeline order. - Validate resource link schemes and omit disclosure controls for items with no details. ## Verification - Focused activity, runner turn, protocol detail, and task-thread tests pass: 161 tests. - `pnpm check:token-gates` passes. - Browser checks confirm a 32-pixel compact row, centered icons, and one-line expanded rows. - `pnpm -r typecheck` and `pnpm build` pass. The final UI build also passes. - Two real native Codex runs completed in an isolated local instance. Each run created the expected file, performed separate delayed tool calls, reported a deliberate missing-file failure, recovered, and finished the task. - Browser verification confirmed that the open three-row activity group remained expanded after the second run moved into saved history. Saved error output remained accessible without red styling or an X. - The complete UI suite passes: 581 files and 5,979 tests. - `pnpm --filter @paperclipai/ui build-storybook` passes for the final production fixtures. - Greptile rates the latest commit 5/5, with zero new comments and no unresolved review threads. The security scan passes. - All repository test shards pass in CI. The duplicate local `pnpm test:run` was stopped after CI completed the full test coverage. The CI build failed twice because the worker ran out of disk space compiling the Rust runner tests (`No space left on device`, error 28). The build and its aggregate verify gate remain blocked by CI capacity. GitHub squash auto-merge is enabled and will wait for required checks to pass. ## Risks - This changes presentation and disclosure behavior. Users open a group to see earlier activity. - Animation uses stable item IDs. A provider that replaces IDs can trigger another transition. - No API, database, or runner protocol contracts change. > Reviewed `ROADMAP.md`. This improves an existing task-chat surface. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, code execution, and browser testing. The runtime does not expose a more specific model version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb640ec129 |
fix(execution): keep blocked wakes waiting without repeated runs (#13236)
## Thinking Path > - Paperclip manages AI agents and their work. > - Wake admission decides when a task can create an execution run. > - Recovery can prohibit replay while the previous execution needs review. > - Dependency reconciliation kept creating runs before dispatch rejected that same hold. > - Each rejected run added another startup notice without doing useful work. > - This change checks the hold during admission and records repeated automatic waits once. > - Tasks keep their messages and can resume when the current gates permit execution. ## Linked Issues or Issue Description Related changes: Refs #13173 (stale completed-task continuations). Refs #12651 (dependency waits during recovery). **What happened?** A blocked task with completed dependencies can remain under a durable execution reconciliation hold. Each scheduler pass created a queued run. Dispatch then cancelled it before the adapter started. The skipped wake did not satisfy dependency wake deduplication, so this repeated and filled the conversation with “Couldn't start” notices. **Expected behavior** A known execution hold creates a waiting diagnostic without a run. Repeated automatic observations share that diagnostic. Clearing the hold permits a new wake only after the other gates pass. New comments remain available for the next eligible execution. **Steps to reproduce** 1. Assign a blocked task with a completed blocker. 2. Give the task an active reconciliation action, or a resolved action whose automatic recovery evidence still prohibits replay. 3. Run dependency reconciliation repeatedly. 4. Observe repeated cancelled pre-start runs on the base branch. This branch creates no runs while held and admits work after the effective hold clears. ## What Changed - Check effective execution holds under the issue admission lock before inserting runs. Keep the final dispatch check for races. - Share automatic wait diagnostics across producers, wake keys, and service restarts. Apply the helper to reconciliation, dependencies, pause holds, availability, budgets, and disabled heartbeats. - Preserve ordinary comment and interaction receipts during execution holds. Prevent release from draining them while replay is blocked. Keep external-chat receipt authorization intact. - Group empty pre-start reconciliation cancellations into a neutral waiting notice. Keep started runs and the full run history. - Document the waiting contract and add database-backed, UI, and browser regressions. - Stabilize two existing verification tests: allow the asynchronous chat lease transition a bounded five-second wait, and accept either legitimate damaged-session refusal while retaining exact archive-evidence assertions. ## Verification Passed targeted tests: - `pnpm exec vitest run server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts server/src/modules/wake-queue/adapters/postgres.test.ts` — 28 tests. - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx` — covered in the initial combined test run; UI suite passed. - Run-dispatch adapter tests passed in the combined gate regression run. - `pnpm exec vitest run server/src/__tests__/durable-chat-wakeup.test.ts` — 41 tests, including held receipt replay, promotion, and revoked access. - `PAPERCLIP_E2E_PORT=3294 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts` — all 3 browser scenarios pass. Repeated held messages create no additional runs or provider prompts and do not replay writes. - `pnpm check:token-gates` - `pnpm check:module-boundaries` - `git diff --check` `pnpm -r typecheck` and `pnpm build` pass. The full local `pnpm test:run` invocation did not finish green: it encountered exhausted local PostgreSQL shared-memory slots, a missing fresh-worktree runner test binary, and tests loaded across in-flight edits. The affected chat/database suites passed on rerun (81 tests), and the targeted lifecycle/recovery verification passed (3 tests). After building the runner test binary, the full native session suite also passed (37 tests). Final-head [CI run 34621288475](https://github.com/paperclipai/paperclip/actions/runs/34621288475) passed on `8659618b0ed2b98df002a28f4c1bd97321b0db04`, including all server/workspace test shards, all three browser shards, runner verification, typecheck, build, release dry run, and the aggregate verification gates. All 31 reported checks passed; the two conditional Storybook checks were skipped as intended. Greptile reviewed that exact commit at 5/5 with no unresolved review threads. ## Risks The wait record is diagnostic only. It must never count as a delivered wake or bypass a current gate. Tests cover repeated and concurrent admission, resolved no-replay evidence, a remaining dependency after hold clearance, deferred comments, and release gating. Explicit user requests and authorized chat receipts do not share automatic diagnostics. No migration or historical data deletion is required. Existing provider retry budgets remain unchanged. ## Model Used OpenAI GPT-6 through Codex. The runtime does not expose the exact hosted snapshot ID or context-window size. Used repository inspection, code editing, command execution, tests, and review tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
52811c6ce6 |
fix(tasks): require resume before sending to paused tasks (#13232)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task execution controls let board users pause a task or its subtree. > - The composer still accepted messages while a pause hold was active. > - A paused task must require an explicit resume before the user can send another message. > - This pull request replaces the composer with an amber pause card and checks board comment writes on the server. > - The user keeps their draft and resumes through the existing task controls. ## Linked Issues or Issue Description Refs #13104. Refs #13119. **What existing behavior does this improve?** The task composer and existing task/subtree pause controls. **Current behavior** A paused task can still receive a board message. The pause notice sits outside the composer, which leaves the send action available. **Proposed behavior** Show an amber takeover in both task chat and the classic composer. Preserve the draft. Require the user to resume the task or the ancestor subtree before sending. Reject board comment writes through either supported write route while the pause hold is active. **Breaking changes** Board comment writes to a paused task now return HTTP 409. Agent run reports remain supported during a pause. There is no schema migration. ## What Changed - Add a shared amber composer takeover with task, subtree, saved draft, pending, and error states. - Use effective ancestor pause state in both composer interfaces. Refresh it after pause events, task updates, and rejected sends. - Preserve draft text and attachments. Hide editor, send, queued edit, and pending question controls while paused. - Check active pause holds before board comment writes can mutate tasks, store comments, or wake agents. - Connect the approved Storybook examples to the production component and update the design and behavior docs. - Add browser coverage for both composers, draft persistence, resume, inherited holds, and rejected writes. Update ACP continuation coverage for the explicit resume requirement. ## Verification - Passed: `pnpm -r typecheck`. - Passed: `pnpm build`. - Passed: `pnpm build-storybook`. - Passed: `pnpm check:token-gates` and `git diff --check`. - Passed: focused UI tests (398 tests) and server route tests (127 tests). - Passed: `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/paused-composer.spec.ts tests/e2e/acp-stop-continuation.spec.ts` (5 tests). - Passed: manual browser walkthrough in a disposable local instance. Pause with a draft, refresh while paused, resume, send, and reopen. The draft returned, and one message persisted. The amber card and resume dialog were readable with no clipping. - Full local `pnpm test:run` did not pass: the general-server stage recorded 9,072 passing tests, 6 database setup failures from macOS shared-memory exhaustion, and 4 failed tests. This stopped the script before its later groups. Latest-head CI runs those groups independently. - Local follow-up: the Git file-resource load test passed on rerun (4 tests); native finalization migration passed after clearing the abandoned browser-test database allocation. Building the native debug fixtures fixed the missing fake provider. The remaining native-session recovery assertion also reproduces on untouched base commit `87b3e5fc6` (36 pass, 1 fail on both base and PR). It expects a settled-session error but receives a semantic-input-digest error. - The final UI build, UI typecheck, token gates, both thread suites (182 tests), and all five browser tests passed after the queued-action review fix. All 31 latest-head CI checks passed, including all server, workspace, browser, build, release, and security gates. Two optional Storybook jobs were skipped by workflow policy. Greptile reviewed `32d8fb5f5` at 5/5 with no open findings. - Review the Paused Composer and Tasks / Execution Controls stories. Pause a task with a draft, verify the amber card, resume, and verify the draft can be sent once. ## Risks - Clients that used board comments to continue paused work must resume first. The response is an explicit HTTP 409. - Pause state can change while a page is open. Live updates refresh the composer, and the server rejects stale sends before their side effects. - Resume keeps the existing dialog and optional agent wake behavior. Agent reports from interrupted runs remain allowed. ## Model Used OpenAI Codex, based on GPT-6, assisted with design, implementation, code execution, and browser verification. The exact runtime model ID and context window are not exposed in this session. The agent used reasoning and tool calls. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the relevant tests locally and they pass; the full local-suite limits are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d10cbde815 |
fix(recovery): reject stale productive continuation wakes (#13173)
## Thinking Path > - Paperclip manages agent work through tasks and runs. > - Recovery continues assigned work when no live execution path remains. > - A recovery sweep can read an in-progress task before its run completes. > - The sweep can then observe the successful run after completion has changed the task status. > - This pull request checks current status and assignment under the existing enqueue lock. > - A stale continuation leaves a skipped wake receipt and creates no run. > - Task chat also omits an empty continuation cancelled before it started because its task had become terminal. ## Linked Issues or Issue Description Related public work: #10779 and #8419. Those older open changes address terminal disposition across other recovery paths. This change uses the existing scheduler guard for productive successful-run continuation and adds real database lock contention coverage. **What happened?** Recovery could combine an old in-progress task snapshot with a newer successful run. It queued an automatic continuation after the task was done. Dispatch cancelled that run before it started, but task chat displayed “Couldn't start” below the successful answer. This can happen after the native runner's finish result has already been accepted. It does not require a missing comment. **Expected behavior** Productive continuation must remain eligible when enqueueing acquires the task lock. Completion, cancellation, reassignment, or a move away from in-progress must prevent creation of the run. Actual execution stops must remain visible. **Steps to reproduce** 1. Let recovery select an assigned in-progress task whose latest run succeeded with productive progress. 2. Hold the task row lock in another transaction and change the task to done. 3. Let recovery attempt to enqueue while that transaction holds the lock. 4. Commit completion. Before this fix, recovery creates a redundant run from the stale snapshot. **Paperclip version or commit** Reproduced against master at `4042eb1c4` with deterministic integration tests. **Deployment mode** Built from source with PostgreSQL. The bug is in core recovery and is not adapter-specific. ## What Changed - Pass the existing status-and-assignee guard for productive terminal continuation recovery. - Preserve a skipped wake receipt with the expected and actual task state, without creating a run. - Test actual PostgreSQL lock contention for native and legacy completion, cancellation, backlog, review, blocked state, and reassignment. - Omit empty redundant pre-start cancellations from native and legacy task chat. Preserve stop markers for runs that started. - Document recovery eligibility at enqueue time. ## Verification - All seven new race cases failed before the guard was connected. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm check:token-gates` passed. - Task chat suite: 98 tests passed. - Recovery integration suites: 290 tests passed, including 31 stale-queue tests. - Full CI verification passed on `7ea71f04d`: all 31 active checks succeeded, including all test shards, browser tests, build, typecheck, and canary release dry run. Storybook visual regression was skipped by its path filter. - Greptile reviewed this commit at 5/5 with no review threads. - The local `pnpm test:run` aggregate reported a setup failure in the unchanged `tool-access-service.test.ts` suite. Its isolated rerun passed all 231 tests without edits. The duplicate aggregate was stopped after the complete CI matrix passed; it is not counted as a successful local full-suite run. ## Risks Low risk. The backend guard applies only to productive successful-run recovery. It requires the task to remain in-progress with the same agent. Other wake sources keep their current policy. The UI change only suppresses empty redundant cancellations; run records remain available. No schema change or migration is required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code edits, and local test execution. The exact deployment snapshot and context window are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
889947c238 |
feat: add experimental native chat connectors (#13038)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also ask agents for work in their existing chat tools. > - Each external conversation needs one task and a current authorized source. > - Retries, Stop, and provider failures must not duplicate work or expose private data. > - The first chat PR establishes the opt-in provider and data contracts. > - This PR adds experimental channel integration and its durable control plane. > - Users can request work from connected channels and inspect delivery in Paperclip. ## Linked Issues or Issue Description Refs #13100 and #13092. This is the second of exactly two chat PRs. Foundation #13100 is merged and changed 143 files. Runner prerequisite #13092 is also merged. This PR changes 400 files against master, below the 500-file review limit. It contains no wireframe images or HTML galleries. ## What Changed - Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat connections. Keep chat disabled unless the operator enables experimental chat connectors. Preserve the production GitHub tool connection and its normal setup path. - Bind each provider bot identity to one immutable Paperclip agent. Bind each admitted external conversation to one task. Paperclip owns tasks, runs, permissions, and audit records. - Add durable admission, per-conversation queues, questions, task controls, progress, final replies, images, files, and delivery receipts. Board comments remain internal unless explicitly sent to the channel. - Check current identity, provider reach, resource access, credentials, runtime generation, and exact source before provider effects. Keep private responses private. Never send raw reasoning, private logs, credentials, or tool arguments. - Hold uncertain sends for explicit audited resolution. Make Board Send-to-channel atomic and idempotent. Keep reconnect and setup credentials in Paperclip secret storage. - Preserve current native-runner authority across retries, lost acknowledgements, and recovery. Keep immutable input and completion contracts separate from newer user input. Receipt reconciliation cannot launch a provider. - Reconcile chat close/new ordering and provider-effect lock order. Audit resource access changes in the same transaction. Submit only the selected resource from each UI toggle so stale pages cannot undo unrelated access changes. - Drain Codex stdout before certifying process exit. Bound the drain with the existing shutdown grace. Preserve observed terminal authority without treating an undrained process as successful or reusable. - Incorporate master `018ca5da` with its ACP Stop, mobile task layout, runner packaging, and official lock changes. Preserve dedicated chat-answer continuations in both directions when ordinary queued comments are adopted after Stop. - Fence late adapter readiness behind an earlier Stop for the same run. Preserve verified cleanup for registered adapters. Handle single Stop, agent pause, duplicate Stops, and failure release without creating a false cancellation receipt. - Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact failed-chat retry authorization and lineage, retired question-source suppression, and the block on generic recovery that would discard the admitted source. Fresh deferred input retains its separate promotion path. - Incorporate master `2a05b5ed3` and its queue-admission extraction, simplified transaction ports, and separate runner CI job. Preserve exact durable receipts, actor separation, and dedicated-answer isolation through the new module. A failed receipt insert rolls back the accompanying deferred-wake merge. ## Verification Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are resolved. This successor fixes two test-harness boundaries exposed by CI: per-case route-module preparation and actual durable-save completion before intentional runner termination. Production code and all existing test/turn deadlines are unchanged. [Exact-head Greptile review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594) is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable findings or open review threads. [Fresh exact-head CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341) passes **all 24 jobs**, including Build and both required aggregates. Normal exact-head guarded merge was attempted and rejected by the remaining branch approval policy: CODEOWNER review is required and no human approval is present. Normal **squash auto-merge is enabled** as of September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified; no approval bypass or self-approval was used. Earlier-head results below remain historical evidence, not qualification of this successor. - Final exact-head Linux evidence: 995/995 chat integration cases; 36/36 agent-skills routes; 35/35 runner live-session cases, including real process kill/resume; 1948 runner Vitest cases with three existing benchmark/platform guards; 870/870 API-authority cases; and 104 browser cases with four existing optional skips. Rust, conformance/replay, full repository build, typecheck, canary, all server/workspace shards, and both required aggregates pass with normal CI concurrency. Earlier failed attempts remain recorded below. - Latest test-only qualification: 141/141 route/permissions/authentication cases pass in separate cold forks, with plain server types and independent review clear. The real-runner suite passes 35/35, with plain runner types and independent review clear. A controlled premature-save acknowledgement fails as expected; matching ownership/effect/process evidence, rejected saves, real turn outcome, test abort, and pre-kill liveness are covered. No local reproduction of the original CI scheduling failure is claimed. The preceding [CI run](https://github.com/paperclipai/paperclip/actions/runs/34479680858) passes 21/24 jobs, including all 995 Linux chat cases and browser aggregate (104 passed, four existing optional skips); only Build, the skills serialized shard, and the required verification aggregate fail. Its exact-head Greptile review was 5/5. Both failed job logs are retained. - Final fixture qualification: all eight focused Discord cases and all 995 chat integration cases pass. The exact modal statement/PID is observed before taking the real connection lock; the test then proves its actual blocking relationship before mutation. Original SQL execution, provider behavior, negative assertions, and 1s/15s timeouts remain unchanged. Independent review is clear and test/production hashes remain frozen. The preceding [CI attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777) passed 22 jobs, including Build/runner, typecheck, canary, all other test shards, and browser aggregate (104 passed, four existing optional skips); the two fixture failures and failed verification aggregate remain recorded, not relabeled as a pass. - Current queue-module composition: 308/308 recovery/batching/queue/Stop tests; 995/995 full chat integration; 89/89 module tests, including real PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary tests; plain server and UI types. All four actual local process/ACP browser paths pass in 1.4 minutes. Fresh databases, no skips or retries, stable reviewed source hashes. The initial boundary failure is retained; its no-op service wrapper was removed without changing recovery context or weakening the check. An exploratory standalone test-directory typecheck fails because its new upstream transformation config is not a standalone typechecking project; standard CI/build does not invoke it, and no configuration was weakened to suppress those diagnostics. - The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed [all 24 CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958) and exact-head Greptile review at 5/5. Required CODEOWNER review prevented its normal merge before master advanced again. - Final extracted-module composition: 307/307 recovery, batching, queue and Stop-control tests; 995/995 full chat integration; 49/49 module tests including eight PostgreSQL adapter cases; and 19/19 issue-update tests. Plain server types pass. All four actual local process/ACP browser paths pass in 1.3 minutes. Fresh databases, no skips or retries in these cohorts, frozen source hashes, and independent review clear. - The preceding head `3e4e1c1c` passes [all PR CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820), including Build and required `ci / verify` and `ci / e2e`. Both the original Rust failure and the previously load-sensitive lineage fixture pass with unchanged Linux concurrency. Master advanced afterward and required this reconciliation. - Final master composition: 448/448 focused UI tests, 186/186 adapter tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI, server, shared, and adapter types pass. Token gates and diff checks pass. Independent server and UI reviews are clear. - Stop-registration regression: both real-service cases fail against exact `a95` source and pass with the fix. The full corrected recovery/control suite passes 265/265. Duplicate-owner and failed-Stop controls also pass. Plain server types pass. The readiness barrier prevents provider startup without adding an acknowledgment to an already terminal run. - Final qualification strengthens terminal-field equality and repeats both affected cases successfully on a fresh database. All four actual local process/ACP browser paths pass again in 1.3 minutes, without skips or retries. The final screenshot shows Cancelled, a paused subtree, retained input, and no error toast. - Two new actual-service regressions fail before the merge fix. They prove that queued-comment adoption could consume a dedicated chat answer or add unrelated input to that answer. The fixed four-case cohort passes, including ordinary upstream continuation and adapter Stop controls. Full recovery passes 257/257. All four actual local process/ACP Stop browser flows pass in 1.4 minutes, without skips or retries, on a fresh database. - The unchanged runner artifact was qualified with 171/171 transport tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11. Six controlled reader tests prove the exit/drain repair. Its local serial Rust workspace passed 546 top-level cases plus two invoked helpers; the later passing Linux CI supplies default-concurrency evidence. - Prior exact-source full chat integration passes 995/995. Settings regressions cover concurrent stale pages, 501 destinations, pending state, rejected updates, and explicit retry. These deterministic tests do not prove live provider behavior. - Retained failed attempts and their causes are in the [qualification log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md). The first merge adapter run timed out while macOS slept for 290 seconds. Its unchanged repeat passed with a temporary sleep guard. No assertion, deadline, or CI gate was weakened. Review commands include `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh disposable databases. See the [browser runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md) for provider setup and separate live acceptance steps. ## Risks - This remains experimental. Deterministic tests and bounded live evidence do not establish every provider feature, tenant, permission layout, or media shape. Teams work-tenant qualification is still open. - Failed and uncertain provider effects remain visible and can require operator action. A transport receipt does not prove recipient visibility. - Native controller and runner artifacts must remain compatible. Preserve lease ownership, terminal authority, source binding, and quarantine during future changes. - Access and audit rows commit together, but activity notifications remain best-effort. This is not a new durable event outbox. - The PR operation does not deploy a live server, replace its runner, or change provider permissions. Remaining live qualification is documented in the [temporary handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md). ## Model Used OpenAI Codex assisted with implementation, tool execution, testing, and review. The work records `gpt-6-astra` assistance. The environment does not report a context-window size. No private reasoning traces are included. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018ca5daaf |
fix: verify ACP Stop and preserve safe continuation (#13119)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task controls coordinate provider execution and queued user messages. > - Stop could finish before an embedded ACP provider stopped its tools. > - A later request could be held for reconciliation without a clear task response. > - A restored provider could also retain the stopped run's API credential. > - This pull request verifies provider termination and preserves safe session continuation. > - Operators can continue known-safe work and see why uncertain work cannot start. ## Linked Issues or Issue Description **What happened?** Stop could leave an embedded ACP provider running. A queued follow-up followed by “go” could fail before it reached the provider. Task chat could show a generic missing-response message. Even a restored session could use the previous run's credential and fail its task update. **Expected behavior** Stop waits for confirmed provider termination. A later explicit wake continues the same compatible session only when recorded actions have known outcomes. It carries pending comments and the current run's environment. Uncertain actions retain a visible reconciliation hold. Composer Stop preserves the existing pause rule: conversation can continue while paused, but task work requires Resume. **Steps to reproduce** 1. Start an embedded ACP task. 2. Send a second request while the provider is running. 3. Interrupt the run, then send “go”. Also test composer Stop followed by Resume work. 4. Check that the request is delivered once and that the provider can complete the task through the current run's API credential. 5. Repeat with an unfinished write. Confirm that the write stops and that further execution stays blocked with a visible reason. **Paperclip version or commit** Built from source on master at `3bc60dd8b` plus this branch. **Deployment mode** Local source build with an isolated embedded PostgreSQL instance. Refs #11183. Refs #12552. Those changes address recovery after operator cancellation. This change also covers embedded ACP termination, session proof, pending-comment delivery, and task feedback. ## What Changed - Propagate Stop into embedded ACP and wait for bounded adapter cleanup and provider exit. Retain the actual ChildProcess object for forced termination on all platforms; never signal a recycled numeric PID. - Preserve interrupted checkpoints only for acknowledged, local, persistent sessions with settled reads or no tools. Keep writes, incomplete actions, and forced termination blocked. - Restore the same compatible provider session with the current run's environment. Reject fresh-session fallback for an interrupted checkpoint. - Adopt pending comments on the next explicit wake. Stop alone does not dispatch them. - Share the execution-blocker rule across dispatch, Resume, and task detail. Show Stopped or Couldn't start with the recorded reason. Resolve the stopped agent for the run link, including reviewer runs. - Keep execution reconciliation holds intact when generic recovery sees queued comments or healthy child tasks. - Add process, service, component, and browser regression coverage. Fix disposable database cleanup and React test settling exposed by the full suite. ## Verification - Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. - Passed all three `acp-stop-continuation.spec.ts` browser journeys. They use an actual ACP child process and require task completion through the agent API. - Passed 165 adapter execution, operator-stop, and child-process control tests, 17 queued-comment route tests, and 65 tests in the two adjusted UI suites. Earlier focused recovery, heartbeat, and task-control tests also passed. - Manually used the browser to queue a request, Stop, send “go” while paused, and Resume. The same session answered once and moved the task to Done with the current run's credential. - Manually interrupted an unfinished write. Its file size stayed fixed for five seconds. “Go” showed the reconciliation reason and did not start another provider prompt. - Separate live Claude ACP smoke checks confirmed that Stop ended a disposable local write and that a no-tool interruption could resume the exact provider session. The browser fixture does not call Drive or another external app. - Passed all 5,615 UI tests and 3,090 other workspace tests. The CLI and general server groups pass with targeted retries: two transient server failures passed together on retry, and two embedded-database startup failures passed after removing abandoned shared-memory segments from this task's completed browser fixtures. All 144 serialized server suites completed, with 2,189 tests passing after two transient HTTP socket failures passed on retry. - Passed all 135 heartbeat process/recovery tests, including a deterministic regression that failed before the recovery-sweep fix. - Passed 18 dispatch integration tests, including stopped-reviewer links, company boundaries, and malformed run IDs. - Greptile is 5/5 on `7dd170d83`, with zero unresolved review threads. The security scan and all required CI gates pass for the same commit. ## Risks - Safe continuation depends on complete tool reporting and a restorable local provider session. Unknown outcomes remain blocked and require reconciliation. - Provider cleanup can take time. A timeout does not grant replay permission. - The change adds optional adapter context fields and an optional issue projection. It does not change the database schema or require a migration. - Test cleanup truncates company data only in a disposable test database. ## Model Used OpenAI GPT-6, running as Codex with repository tools, code execution, and browser interaction. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2d45f42e47 | fix(ui): refine mobile task surfaces (#13122) | ||
|
|
8cfd30fb07 |
feat(ui): add composer Stop and simplify task controls (#13104)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task composer is where operators direct running agents. > - Operators need to stop work without leaving the conversation. > - Existing pause controls already hold task trees and interrupt both runner types. > - This pull request connects the composer to those controls and removes repeated feedback. > - Operators can pause work quickly and still queue messages while agents run. ## Linked Issues or Issue Description **What existing behavior does this improve?** Task pause, resume, and cancellation in the task page and composer. **Current behavior** The empty composer cannot stop a running task. Task controls require extra confirmation and reason text. Pause can show several notifications for the task already on screen. **Proposed behavior** Show Stop while this task runs and the composer is empty. Text or attachments switch it to Send. Stop and the menu use the same manual pause hold. Parent pauses include descendants. Keep task cancellation in the menu with a compact confirmation. Show one quiet pause row and gray cancelled-run details. **Reason and benefit** Operators can interrupt execution with one click. Drafts and queued messages keep their existing behavior. The UI waits for actual termination, including native cancellation acknowledgment. **Breaking changes** No endpoint, schema, or task-status change. Pause no longer asks for confirmation or a reason. Resume now honors the existing wake-agents option. Task notifications are suppressed for the task and subtree currently in view. Related UI work: #8228 changes navigation and composer shortcuts. This PR covers execution controls. No duplicate Stop-button PR was found. The change improves existing controls and does not duplicate a roadmap milestone. ## What Changed - Add Stop, pending feedback, duplicate-click protection, and inline errors to the composer. - Share the pause mutation across the composer, active-run controls, and menu. - Poll affected runs after a pause request. Require native cancellation acknowledgment. - Remove pause confirmation and shared reason fields. Reduce cancel confirmation to its task count and actions. - Honor wake-agents for executable tasks only. Preserve the pause when recovery review is needed; show partial wake failures inline. - Preserve explicit legacy reconciliation decisions while their continuation waits for dispatch. - Suppress notifications for visible task trees. Use quiet pause and cancellation feedback. - Add interactive stories using production controls and native/legacy end-to-end tests. ## Verification - User reviewed the running feature and revised Storybooks in the browser. - Rebased focused checks passed: 295 original targeted tests, 161 updated route/page/notification/status tests, and 26 recovery integration tests. - Both isolated runner journeys pass on the final revision (1.7 minutes). Coverage includes queueing, parent and child interruption, persisted holds, no automatic continuation, reconciled resume, cancellation, terminal exclusions, and no Stop toast. - Native coverage uses real runnerd with a deterministic provider fixture. Legacy coverage checks actual process termination. Live hosted-provider execution was not tested. - Repository typecheck and build, Storybook build, and token gates passed after rebase. The final server typecheck/build also passed. - The broad local run completed its general-server stage with 7,219 passing tests, 48 skipped, and two failures from cached pre-fix source and a stale native provider fixture. Both failed tests pass in fresh final-head reruns after rebuilding the fixture; the script did not continue to its later local stages. CI runs all test groups on the final revision. - Final revision: all 31 applicable CI checks passed; Storybook visual regression was skipped by its workflow conditions. Greptile: 5/5, zero unresolved comments. - Review `Tasks / Execution Controls` in Storybook. Type and clear a draft, stop a run, expand cancellation details, and test the menu on desktop and mobile. ## Risks - Stop pauses descendants for a parent task. This is the existing pause contract. - A held task can remain active if interruption fails. The UI shows an error instead of claiming termination. - Resume can start multiple assignees when wake-agents is selected. Backlog, blocked, and terminal tasks stay excluded. Existing execution reconciliation remains mandatory where required; Resume never invents action-outcome evidence. - Notification suppression uses the visible task and cached subtree. Notifications for unrelated work remain enabled. ## Model Used OpenAI GPT-6 through Codex. The exact runtime snapshot and context-window limit are not exposed in this session. Used reasoning, tool calls, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cd4c4ed205 |
fix(ui): stabilize task loading and live feeds (#13095)
Coordinate initial conversation reveal, preserve message identity and reading anchors during live updates, and bound transcript reads with recoverable retries. Cover desktop/mobile navigation and rich task loading with actual-route browser tests. Verified all Linux CI gates, 5,575 local UI tests, eight layout browser scenarios, and recorded native Codex walkthroughs. Greptile: 5/5; all review findings resolved. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
7ed122911b |
Add end-to-end session goals to Paperclip Runner
Add capability-aware slash-goal controls, durable provider goal state, PRP v2 negotiation, autonomous goal execution, and safe local session recovery. Integrate with current master, preserve provider session identity, and verify the browser goal/chat/replacement/clear workflow and unsupported-agent rejection. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e095b84dab |
feat(connections): connect services from native task feeds (#13058)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents use connections to reach external services. > - A fresh native task can have no service tools installed. > - The agent needs a way to discover services and ask the responsible person for access. > - This pull request brings the existing connection-intent flow into native task execution. > - The person can connect from the task, and the agent can continue with updated tools. ## Linked Issues or Issue Description **Subsystem affected** Native runner tool authority, connection intents, task interactions, and shared connection setup. **Problem or motivation** A task that needs an unconnected service cannot finish its work. Leaving the task to configure access also loses context. A resolved request must survive a restart and resume the correct agent once. **Proposed solution** Expose connection discovery and access requests as server-owned native tools. Render a durable task card and use the shared setup dialog. Persist outcome delivery and start a fresh provider session after access is ready. **Alternatives considered** Sending the person to the Connections page adds navigation and does not solve continuation. Polling for authorization consumes runs and can create duplicate requests. **Roadmap alignment** This extends the existing connection-intent runtime and setup experience. It reuses the shared access model and the native runner. Related: #12345, #12347. The service-slug fix in #12906 is related but separate. Companion evaluation PR: https://github.com/paperclipai/paperclip-evals/pull/21. ## What Changed - Expose `connections_search` and `connection_request` with server-bound company, task, agent, and responsible user. Preserve the legacy entry points. - Discover catalog services and authorized custom connections. Check installation, identity, health, and executable permissions before reporting ready. - Keep pending cards through ordinary messages. Reuse requests and retire stale ownership. Put Connect at the right of Not now. - Reuse the shared setup flow in a task dialog. Keep access additive and default to the requesting agent. Recover from cancelled or blocked OAuth windows with a new-tab fallback. - Persist outcome delivery with an idempotent wake key. Resume in a fresh session and recheck ownership before dispatch. - Add native browser fixtures, offline Storybook states, server contracts, and evaluation fixtures. Update guidance and documentation. ## Verification - `pnpm build`: passed after replaying the change on current master. - `pnpm -r typecheck`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed. - New continuation-policy regression cases: 16 passed. - Docker-backed PostgreSQL regressions passed for requester-only OAuth access, assignment-only expiry, terminal expiry, and credential-free setup metadata. - Shared setup and task-card UI tests: 121 passed, including configured MCP reconnect URL recovery and preserving user edits across refetch. - Storybook browser checks: all 119 passed on the latest reconnect fix. - `pnpm test:run`: 4,734 tests passed in the first server group, but embedded PostgreSQL startup failures and resulting cleanup errors prevented a complete local pass. All Linux CI lanes passed on the latest reviewed commit. One external-object route test returned an unexplained 500 on the first run; it passed twice locally and the failed shard passed on retry without code changes. - Earlier feature-checkout evidence: three deterministic native browser journeys passed, including restart delivery and an actual fixture tool result. Legacy scripted coverage also passed. All 59 added stories were inspected in light and dark themes. - Live Notion testing recorded successful provider reads. The manual test used a local-trusted instance. It does not prove authenticated/cloud deployment or every provider journey. - Native browser rerun reached the embedded PostgreSQL startup limit before bootstrap, so the latest checkout’s full native browser journey remains unverified. Both OAuth page/task regression cases passed against isolated Docker-backed PostgreSQL 17. They verify no premature task access, requester-only completion, additive retries, and reconnect preservation. - Applied both new migrations twice to isolated PostgreSQL 17. Foreign keys remained intact, duplicate active delivery keys were rejected, and failed delivery records did not block retries. Reviewer path: start a fresh test drive, enable the native runner, use an agent that can perform work directly, and ask it to summarize a Notion page. Connect from the card, then verify the resumed provider call and source-linked answer. The default test-drive CEO is instructed to delegate, so it can introduce an unrelated hiring step. ## Risks - Two additive migrations create durable deliveries and a partial unique wake index. They are idempotent. The wake index can require a maintenance window on large tables because migrations run in a transaction. - OAuth and continuation cross asynchronous boundaries. Tests cover ownership changes, retries, additive access, and restart delivery; live provider behavior still varies. - The latest requester-scope fix has not yet been exercised through live OAuth. GitHub, API-key, authenticated-user, and all recovery journeys are not claimed as verified. ## Model Used OpenAI GPT-6-based Codex assisted with implementation, tests, and review using tools and code execution. The runtime does not expose the exact model version, context window, or reasoning setting. Live evaluation used `gpt-5.6-luna`; manual native testing used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used and disclosed unavailable runtime details - [x] I have checked ROADMAP.md and confirmed this extends existing connection work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run all required tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5b56d430e9 |
feat(ui): refine core navigation and task detail (#12854)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The main navigation and task detail view are core operator surfaces. > - Several controls used different hover states, popover layouts, and spacing rules. > - Recent task actions also needed a compact menu and correct inbox archive behavior. > - These differences made the interface feel inconsistent and caused some content to look crowded or clipped. > - This pull request aligns these surfaces with the Paperclip design tokens and current interaction patterns. > - The benefit is a simpler and more consistent operator experience in light and dark modes. ## Linked Issues or Issue Description **What happened?** Profile and organization popovers used inconsistent layouts. Navigation controls used different hover and selected backgrounds. Task warnings and the composer could crowd nearby content. Archiving a recent task could also remove it from more than the inbox. **Expected behavior** Popover menus should use the same compact visual language. Navigation controls should share readable hover and selected tokens. Task detail content should keep consistent spacing. Archiving should hide a task from the inbox while keeping it in the task list. **Steps to reproduce** 1. Open the main sidebar in light or dark mode. 2. Open the profile and organization menus. 3. Hover navigation items, the organization trigger, the profile trigger, and the feedback flag. 4. Open a task with a warning banner and a long thread. 5. Use the recent task overflow menu and archive a task. **Paperclip version or commit** Reproduced on `master` before this branch. **Deployment mode** Local dev (`pnpm dev`). ## What Changed - Rebuilt the profile and organization popovers with compact token-based layouts. - Matched organization popover width and alignment to the profile popover. - Unified sidebar hover and selected states in light and dark modes. - Added a recent task overflow menu with rename, archive, and pause or restart actions. - Kept archived tasks in the task list while removing them from the inbox. - Improved warning banner and composer spacing in task detail views. - Added and updated focused UI tests for the changed behavior. ## Verification - `pnpm check:token-gates` passed. - `pnpm --filter @paperclipai/ui typecheck` passed. - The seven affected UI test files passed with 216 tests. - `pnpm --filter @paperclipai/ui build` passed. - GitHub CI passed the full build, typecheck and release registry, general test, serialized server, canary dry-run, and end-to-end matrices. - Greptile reviewed commit `6e296be85` at 5/5 with no outstanding actionable findings. ## Risks - Risk is limited to sidebar presentation, recent task actions, and task detail layout. - The recent task archive action now follows inbox-only archive semantics. - No database schema or public API contract changed. > I checked [`ROADMAP.md`](ROADMAP.md). This pull request does not duplicate planned core work. ## Model Used - OpenAI Codex, GPT-5.6. The model used high reasoning, tool use, and code execution. The context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Scott Tong <scott@scottsmbpm5max.lan> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bf95a7eae2 |
fix(ui): stabilize active-run steering queue (#12834)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue detail page shows a live agent run and accepts follow-up instructions. > - A follow-up must stay in a stable queue until the user sends, reorders, or removes it. > - Native runners can receive a steering event in the active run. > - Legacy runners must interrupt the active run and start a follow-up run. > - The current UI moved comments between the queue and the transcript and could show duplicate text or ambiguous chronology. > - This pull request makes the queue projection durable, keeps each message in one clear place, and labels when queued input was actually steered or delivered. > - The benefit is predictable steering with stable ordering, no duplicate messages, and visible causal timing. ## Linked Issues or Issue Description Refs #11374. Refs #12591. **What happened?** During an active run, a new follow-up could first appear as a transcript bubble and then move into the steering queue. After a steer or remove action, it could appear again. Progress text could also repeat the final response text. Once consumed, a queued bubble displayed only its original submission time even though it moved to its later causal slot, and a native run split by steering looked like two unrelated runs. **Expected behavior** An active-run follow-up must appear in the queue immediately. A native steer must move it once into the active run. A legacy interrupt must move it once into the follow-up run. A removed item must stay removed. Progress text that is identical to the final response must appear once. Consumed follow-ups must show both queue and steer/delivery times, and post-steer native segments must identify themselves as continuations of the same run. **Steps to reproduce** 1. Start a long-running task. 2. Send two or more follow-up messages while the agent is active. 3. Reorder the messages and remove one message. 4. Send the first queued message as steering. 5. Observe the queue and transcript during and after both runs. **Paperclip version or commit** The problem reproduced on commit `da1e40302`. **Deployment mode** Local development with the embedded database. ## What Changed - Project queued comments into the steering well for native and legacy live runners. - Send native steering to the active run and use interrupt-and-follow-up for legacy runners. - Keep optimistic queue order stable across refreshes and roll back failed actions. - Remove discarded comments from the transcript cache and keep them removed when the queue becomes empty. - Collapse only the final progress occurrence matching the durable response, including across steered transcript segments. - Show `Queued … · Steered …` for same-run input and `Queued … · Delivered …` for successor-run input at their causal positions. - Label settled and live post-steer segments `Continued after steering` and time them from the steer boundary. - Add regression tests for queue display, steering, fallback interrupt, reorder, remove, rollback, duplicate text, causal timestamps, and live/settled continuation headers. ## Verification - Ran the final focused steering/chronology UI suite with 233 passing tests. - Ran the activity-service regression suite with 5 passing tests. - Ran the broader queue-focused UI suite with 298 passing tests before the final chronology refinement. - Ran `pnpm -r typecheck` successfully. - Ran `pnpm build` successfully. - Ran `pnpm check:token-gates` successfully. - Tested native steering in a real browser with a 90-second baseline wait and a three-second steering correction. - Confirmed that the old final response did not appear before the steered response. - Tested three queued messages in a real browser. - Confirmed that reorder changed delivery order and that the removed message was never sent or shown again. - Tested a legacy runner in a real browser. - Confirmed that it used the interrupt fallback and showed the follow-up once. - Reloaded a saved mixed-steer/successor-run thread and confirmed the causal timestamps and continuation header render in the correct positions. - The complete macOS suite reaches five unrelated platform assertions in workspace-runtime tests. Two compare `/var` with `/private/var`. Three require Linux `/proc` listener data. GitHub Actions provides the authoritative Linux run. ## Risks - Low risk. The change is limited to issue-chat queue projection and transcript presentation. - The server run-history API adds only a read-only `contextIssueId` projection; the database schema does not change. - Optimistic actions restore the prior UI state when a request fails. ## Model Used - OpenAI Codex with GPT-5, extended reasoning, browser automation, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4d30efa8e3 | feat(ui): refine streamlined task experience (#12748) | ||
|
|
0f94521017 |
fix(runner): restore local session and task integrity (#12721)
## Thinking Path > - Paperclip is the control plane for agents that perform work. > - Paperclip Runner connects durable provider sessions to individual task runs through PRP. > - Provider continuity and per-run authority are different lifetimes. > - The existing implementation mixed those lifetimes and lost event metadata between provider frames, runnerd, persistence, API sanitization, and the task thread. > - That caused failed continuation, missing progress and Plans, duplicate replies, hidden failures, and unsafe recovery. > - This repair gives every heartbeat fresh authority, preserves qualified provider-session continuity, and restores one lossless presentation path without changing direct adapters. ## Linked Issues or Issue Description **What happened?** A second native heartbeat could reuse tickets, leases, command receipts, sequence state, and run identity from the first heartbeat. Provider phase and item identity could be lost before the UI read them. Redaction could corrupt protocol discriminators while still missing malformed credential tails. The task thread could fold progress into the final response, hide failures, or show more than one final answer. Native Codex also exposed approval modes that do not yet have a durable approval bridge. **Expected behavior** Each heartbeat uses a new PRP authority epoch. Codex and OpenCode preserve exact qualified provider sessions; ACPX emits an explicit continuity event when its qualified process-replacement policy is used. Every accepted provider event is presented, classified as internal, or surfaced as unsupported. The task page shows chronological progress, reasoning summaries, activity, Plans, interactions, terminal failures, and exactly one final reply. Direct adapters retain their existing path. **Steps to reproduce** 1. Enable the unified experimental Paperclip Runner setting. 2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex agent. 3. Run response, Plan, structured-question/resume, restart, cancellation, and failure scenarios. 4. Reload the task while active, waiting, failed, and settled. 5. On the old implementation, observe stale run authority, missing classifications, incomplete output, or duplicated/folded replies. **Paperclip version or commit** The repair is based directly on `master` at `87d05e194b643810d16d20612115acd01d735d43`. **Deployment mode** Local development with the embedded database. Related work: Refs #12616, #12646, #12666, #12685, and #12700. ## What Changed - Rotates PRP control-plane, outbox, ticket, lease, command, receipt, and sequence authority for each heartbeat while carrying forward only a validated provider-session identity. - Reads `control-plane-state.json`, validates both durable schemas and lifecycle values, resumes coherent current runs, archives qualified settled authority, and quarantines malformed or mismatched scoped state without moving ambiguous live legacy state. - Preserves Codex provider phase and stable item identities so commentary remains progress and only `final_answer` becomes final. - Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning lifecycle mapping. - Makes ACPX normalization lossless for visible reasoning, tool lifecycle metadata, stable bounded identities, Plan revisions, structured requests, failures, and qualified process replacement. Only the compatible terminal assistant message is promoted as final. - Applies schema-aware redaction before generic JWT-shaped detection and scans every diagnostic string leaf. Malformed raw/escaped quoted credential tails are redacted in both server and durable Rust state. - Restores snapshot-style chronological task presentation, expandable tool activity, inline Plan cards, visible waiting/resume/cancel/failure states, and exactly one final answer. - Makes `never` the only qualified native Codex permission mode and rejects unsupported persisted native modes with remediation. OpenCode and ACPX policies remain intact. - Keeps the unified experimental Runner setting as the only enablement flag. Onboarding and direct Codex, Claude, and OpenCode stay on their legacy execution/finalization paths. - Adds cross-language goldens, authority/recovery/fault coverage, exact response/count assertions, and native plus legacy acceptance scenarios. ## Verification - Pull-request GitHub Actions run Rust formatting/tests, TypeScript checks, server/UI tests, builds, protocol drift checks, browser E2E, and security scans. - A separate workflow-only validation ref is pinned directly on this PR head and runs the 35-cell paid local matrix: three core scenarios plus structured-question resume and restart/resume for native Codex, native OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode. Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315 - Acceptance requires exact single visible replies, monotonic sequences, matching envelope discriminators, one semantic terminal, one run terminal, no unresolved interaction, no duplicate mutation, no secret leakage, provider continuity, and zero native rows for direct adapters. - Per maintainer direction, tests are running in GitHub Actions rather than on the slower local host. Only formatters and static diff checks were run locally. ## Risks - Recovery from old or partial filesystem state is sensitive. The repair fails closed, preserves active or unverifiable authority, and quarantines only state whose scoped ownership is safe to move. - Provider event formats can change. Closed validators and boundary goldens turn new or malformed events into visible diagnostics instead of silent drops. - Shared task presentation could affect direct adapters. Runtime-fact gating plus the direct-adapter matrix protect the existing path. - Managed and remote providers are not qualified here. Shared code continues to compile and fail safely, but live qualification is deferred. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact deployed snapshot and context-window size are not exposed to this task. It used agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] The paid local-provider matrix is green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
72b9f92d76 |
fix(runner): restore task runtime parity (#12685)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task view shows a running agent and lets an operator guide that agent. > - The merged runner stack lost parts of the accepted task experience. > - Native event errors could hide current reasoning from the operator. > - Queued message steering had no server route on `master`. > - This pull request restores the task-runtime behavior and keeps the runner experimental gate. > - The benefit is a visible and steerable native run with durable fallback behavior. ## Linked Issues or Issue Description **What happened?** The task view could stop showing current runner reasoning. The steering action also failed because the server route was absent. Runner instruction files were not declared as supported. **Expected behavior** The task view must show current provider activity. It must use the live log when durable native events are empty or unavailable. The operator must be able to steer a queued message into the active native turn. **Steps to reproduce** 1. Enable the Paperclip Runner experimental setting. 2. Start a native runner task. 3. Open the task view while the run emits reasoning. 4. Queue a message and select the steering action. **Paperclip version or commit** The regression reproduces on `24a674f8858060e77ea1beb50689d26473e91431`. **Additional context** Related closed work: Refs #12592. ## What Changed - Restored the queued-comment steering route for active native sessions. - Added durable and queue-bound steering acknowledgements for safe retries. - Restored runner instruction bundle support. - Added live-log fallback when native events are empty or unavailable. - Restored the compact live reasoning ticker in the task view. - Added a visible temporary-unavailable state when both activity sources fail. - Kept the unified Paperclip Runner experimental gate unchanged. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm check:token-gates` - `pnpm build` - Seven focused test files passed with 140 tests. - The final steering regression file passed with 12 tests. - The broad local test run reached unrelated workspace, port, and shared database failures. The changed-area tests remained green. ## Risks - The steering route changes queue and run records in one transaction. Tests cover stale targets, unavailable sessions, lost responses, and wrong-queue acknowledgements. - Native events remain the primary transcript source. The live log is used only when event data is absent or its poll fails. - The experimental gate still hides and rejects the runner when the setting is off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — no documentation change is required for this regression repair - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
39206c0096 |
feat(ui): complete the task workspace (#12640)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task page is the main place where people guide and review agent work. > - Native runner events already project into the existing task chat on the lower stack. > - The larger task workspace must support those events without breaking direct adapters. > - Legacy questions, final replies, empty transcripts, and classic controls must keep their behavior. > - This pull request completes the provider-neutral task workspace experience. > - The benefit is one coherent task surface for native and direct execution paths. ## Linked Issues or Issue Description Refs #12639. Refs #12617. Refs #12352. **Subsystem affected** Task chat, task detail, side panels, interaction forms, and transcript presentation. **Problem or motivation** The task page does not provide one complete workspace for live activity, plans, questions, queued guidance, files, and documents. Earlier native UI work also exposed compatibility risks in legacy question and reply paths. **Proposed solution** Add the task workspace components and provider-neutral protocol presentation. Keep runner-only controls behind runtime facts. Preserve all direct-adapter composer, transcript, interaction, and finalization behavior. **Alternatives considered** A separate runner page would duplicate task behavior. Replacing legacy transcript logic would create unnecessary adapter regressions. **Roadmap alignment** This work supports the unified task experience and the Cloud and Sandbox agents milestone. ## Stack - Base PR: #12639. - Lower PR: #12638. - This PR contains only its 119-file UI delta against `runner/remote-wss-transport`. - The next stack PR adds administrator and observability controls. ## What Changed - Added a reusable task workspace side panel for files and documents. - Added provider-neutral cards for tools, plans, questions, protocol activity, and progress. - Added queued guidance and richer composer state. - Added compact and expanded interaction presentation. - Added live activity, thinking, usage, recovery, and final reply presentation. - Added document annotations and task deep links. - Added bounded question validation and response handling. - Preserved native runner event projection from master. - Preserved legacy channel-less replies and explicit final reply precedence. - Preserved direct-adapter and classic-interface controls. - Did not change server execution selection, Rust code, migrations, workflows, or `pnpm-lock.yaml`. ## Verification - GitHub Actions will run UI tests, repository tests, typecheck, build, browser tests, security, and policy gates. - Tests cover active, settled, empty-transcript, interaction, queued-message, plan, question, file, document, and classic-interface states. - Compatibility tests cover direct adapters and native runner projection together. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check runner/remote-wss-transport...HEAD` passes. - The delta contains 119 UI-only files. ## Risks - This is a large task-page change. - Most files are new focused components and tests. - Shared transcript code keeps the lower native projection and legacy direct-adapter fallbacks. - Runner-only controls use adapter and runtime facts. - No execution path or rollout flag changes in this PR. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
209680409d |
feat(ui): project native runner turns into task chat (#12617)
## Thinking Path > - Paperclip is the open source app people use to supervise AI agents and their work. > - The task page is the established place to read run progress and answer agent questions. > - Native runner events use PRP envelopes instead of the direct-adapter transcript format. > - The task page needs a narrow projection for those events without changing legacy adapter behavior. > - Unknown event versions and fields must stay hidden until the UI supports them. > - This pull request adds a runtime-gated native turn projection and preserves the classic path. > - The benefit is one task thread for native Codex runs while direct adapters keep their current UI. ## Linked Issues or Issue Description **Subsystem affected** `ui/` task chat and transcript projection. **Problem or motivation** The server can record native runner events, replies, usage, and structured interactions, but the existing task page cannot safely render those records. Reusing the native path for direct adapters would also risk the legacy question and finalization behavior. **Proposed solution** Project supported PRP v1 events into the existing transcript model only when runtime facts identify a native `paperclip_runner` run. Use an exact event and payload allowlist. Keep direct adapters on the existing transcript, composer, interaction, and finalization path. **Alternatives considered** A separate runner page was rejected because it would split task history. A universal transcript replacement was rejected because the experimental runner must not alter legacy adapters. **Roadmap alignment** This change supports activity attribution and recoverable runs. It does not replace the current task page. ## What Changed - Project supported native PRP v1 assistant, reasoning, tool, activity, usage, interaction, and result events. - Render a native runner turn only when both runtime mode and adapter type match. - Recognize the canonical `assistant_message` item kind and preserve final-reply precedence. - Fail closed for unsupported PRP versions, event types, payload schemas, and identity fields. - Keep explicit running and pending provider activities open until a real terminal state arrives. - Include channel and lifecycle identity in memoization so progress-to-final transitions rerender. - Add focused native, empty-transcript, interaction, final, memoization, and direct-adapter regressions. ## Verification - GitHub Actions is the authoritative test environment for this stack. - This top PR receives the full stack-aware CI suite. - Greptile will review this exact nine-file delta after the branch is pushed. ## Risks - The main risk is changing direct-adapter task behavior. The render gate requires both native runtime mode and the `paperclip_runner` adapter. - The next risk is exposing future provider payloads. Version, event, schema, and field allowlists fail closed. - There are no database, lockfile, workflow, build-system, or server changes in this PR. ## Stack 1. [Runner package, SDK, and developer tools](https://github.com/paperclipai/paperclip/pull/12608) 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616) 3. This PR: provider-neutral task-thread UI ## Model Used OpenAI Codex with GPT-5, extended reasoning, repository tools, and parallel review agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bc9ba7cd26 |
feat(runner): project native runs into task threads (#12321)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The experimental Paperclip Runner can execute a guarded Codex run and persist provider-neutral events. > - The task page still reads direct-adapter transcripts and cannot present those native events. > - Structured runner questions must also use the existing task interaction experience. > - Runtime selection must use the persisted run mode, not an adapter name or a current feature flag. > - This pull request projects native events and questions into the existing task thread. > - Direct adapters keep their existing transcript, composer, interaction, and finalization paths. > - The benefit is a complete native Codex task thread without a behavior change for existing adapters. ## Linked Issues or Issue Description Refs #12202. This pull request replaces that stale implementation on current `master`. **What happened?** The server persists native runner events and structured input requests. The task page only consumes direct-adapter transcripts. A native run therefore cannot present a complete transcript, usage, or question flow through the normal task experience. **Expected behavior** Native runs project persisted provider-neutral events into the existing task thread. Native structured questions use the existing interaction card. Direct adapters retain their current behavior. **Steps to reproduce** 1. Enable the experimental runner. 2. Start a native Codex run that emits progress, usage, a structured question, and a final reply. 3. Open the task page. 4. Observe that the direct-adapter transcript path cannot project the native event records. **Paperclip version or commit** `master` at `67f9867bc`. ## What Changed - Add the canonical structured-question validator and shared contract exports. - Materialize native input requests as existing task interactions. - Validate native answers and deliver them through the durable question-response receipt. - Resume the original PRP request with an idempotent `request.resolve` command. - Project native messages, tool activity, cumulative usage, and final replies into the existing transcript model. - Propagate persisted `runtimeMode` to the task page and select native handling only for `runtimeMode: "native"`. - Expire pending interactions through the shared issue service on every terminal transition, including decisions, stalled reviews, tree control, and pipeline retry cleanup. - Queue native run cancellation while a transaction is open and execute it only after the owning transaction commits. - Keep nonterminal and non-runner issue paths on their existing service call shapes and behavior. ## Verification - `pnpm --filter @paperclipai/server typecheck` — passed, including the Rust runner release build and protocol/catalog drift gates. - Focused native-thread and lifecycle suites — 18 files and 481 tests passed during review. - `issue-execution-policy-routes.test.ts` — 19/19 passed after the final transactional-queue expectation update. - `issue-agent-mutation-ownership-routes.test.ts` — 87/87 passed in the final isolated compatibility rerun. - GitHub Actions — policy, build, canary, typecheck/release registry, 5 serialized server shards, 8 general-test shards, 3 browser shards, and both aggregate gates passed on `7793f3193`. - Security — Snyk, Socket Project Report, Socket PR Alerts, and Superagent passed. - Greptile — 5/5 on `7793f3193`; all actionable review threads resolved. - `git diff --check` — passed. - Diff against `master`: 44 files. ## Compatibility Boundary - Native transcript polling only runs when the persisted run reports `runtimeMode: "native"`. - Missing or legacy runtime modes continue through `useLiveRunTranscripts`. - Legacy questions keep the existing optional free-text choice. - Native closed select sets can suppress that legacy fallback. - Terminal cleanup uses the same issue service for native and legacy interactions; only a bound native question schedules a native run cancellation. - Native cancellation happens after transaction commit, so failed or rolled-back writes do not cancel a still-valid run. - The durable delivery service checks the original native request before it considers a continuation run. - This pull request adds no migration, dependency, workflow, manifest, or lockfile change. ## Risks The main risk is routing a direct-adapter task through native handling or changing terminal issue behavior. The implementation selects the native path only from persisted runtime facts, retains the existing nonterminal call shape, and schedules native cancellation only for a validated bound native question after commit. Focused and repository-wide tests cover both paths. Native requests remain bound to the company, issue, run, and agent; answers are validated, durable, and idempotent across reconnects. ## Model Used OpenAI Codex, GPT-5 family. The client does not expose the exact deployment ID or context window. Agentic reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected local tests and they pass - [x] I have added or updated tests where applicable - [x] I have updated the compatibility notes for this change - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I addressed all Greptile and reviewer comments before requesting merge |
||
|
|
adfbe2d4b9 |
feat(environments): refer to the managed default environment by name, not the sandbox driver key (#11838)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed deployments provision a platform-managed default environment
for agent runs; the UI shows this environment in selectors, the agent
form, run details, and the environments page
> - Those surfaces append the raw driver key to the environment name, so
users see labels like "Paperclip Computer (sandbox)", "Paperclip
Computer · sandbox", and fallback copy such as "Managed sandbox" and
"The sandbox has no ready authentication"
> - "sandbox" is infrastructure vocabulary, not the product name of the
environment; showing it next to the managed environment's name is
confusing and off-brand
> - This pull request renders platform-managed environments by name
alone and rewords the sandbox-phrased copy, while user-created
environments keep the driver suffix so mixed lists stay distinguishable
> - The benefit is that the default environment reads as one clear
product name everywhere, and self-hosted users lose nothing: their own
environments still show the driver
## Linked Issues or Issue Description
**What existing behavior does this improve?**
Display of the platform-managed default environment across the UI.
**Subsystem affected**
UI (environment selectors, agent config form, environments page, agents
page, run details) and the claude-local/codex-local adapter auth checks.
**Current behavior**
The agent form labels the inherited default environment as "Name
(sandbox)". Environment selectors and the environments list render "Name
· sandbox". The agents page describes the environment as "<provider>
sandbox provider". The agent form's fallback label is "Managed sandbox".
Adapter auth checks say "The sandbox has no ready authentication for
this adapter."
**Proposed behavior**
Platform-managed environment rows (`metadata.managedByPaperclip`) render
their name alone. The fallback label is "Paperclip Computer". The agents
page describes managed environments as "Managed by Paperclip". Run
details omit the driver suffix for sandbox-driver environments (the
adjacent Provider entry already identifies the mechanism). Adapter auth
checks say "This environment has no ready authentication for this
adapter."
**Reason and benefit**
The managed environment carries a product name. Appending the raw driver
key ("sandbox") to it is noise and contradicts the product naming.
User-created environments keep the driver suffix, so mixed lists stay
distinguishable.
**Breaking changes**
None. Message text of the auth check is not read programmatically; the
UI keys off `ADAPTER_AUTH_MISSING_CHECK_CODE`. Rows without the managed
marker render exactly as before.
## What Changed
- New `environmentDisplayLabel` helper in
`ui/src/lib/managed-sandbox-environment.ts`: managed rows → name alone;
other rows → "Name · driver".
- `AgentConfigForm`: inherited-default label uses the helper; fallback
copy "Managed sandbox" → "Paperclip Computer"; environment options use
the helper.
- `ProjectProperties`, `CompanyEnvironments`: environment selector
options use the helper; the environments-list row hides the driver
suffix on managed rows; the managed detail page's fallback description
no longer says "sandbox".
- `Agents` page: managed environments are described as "Managed by
Paperclip" instead of "<provider> sandbox provider".
- `CommentThread` run details: the driver suffix is omitted for
sandbox-driver environments.
- claude-local and codex-local adapters: auth-missing check message/hint
reworded from "sandbox" to "environment" (ACP and environment-test
paths); claude-local probe/effort/login hints reworded the same way.
- Run status lines: "Syncing workspace to sandbox", "Exporting git
changes from sandbox", "Starting adapter in sandbox", and friends now
say "environment"; "Finalizing sandbox workspace" → "Finalizing
workspace". Templated transfer-progress lines map the `sandbox`
transport key to "environment" for display (`runtime-progress.ts`).
- Agent form sign-in panel: "Sign in to the sandbox" → "Sign in to the
environment"; "Authenticated. The sandbox has credentials now." → "…The
environment has credentials now."
- Feature catalog + instance settings card: "Managed Sandbox Only" →
"Managed Environment Only" (setting key unchanged; the card keeps its
alphabetical slot).
- Server agents routes: execution-target failure and test-identity copy
no longer say "sandbox"; workspace-mode label "Cloud sandbox" → "Cloud
environment".
- Tests: new `environmentDisplayLabel` unit cases; new `AgentConfigForm`
render case asserting the managed default renders without "(sandbox)" or
"· sandbox"; status-line assertions updated across adapter-utils, server
heartbeat/live-run, and UI chat suites.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `pnpm --filter @paperclipai/adapter-claude-local typecheck` and
`--filter @paperclipai/adapter-codex-local typecheck` — clean.
- `vitest run` for `managed-sandbox-environment.test.ts`,
`AgentConfigForm.render.test.tsx`, `CompanyEnvironments.test.tsx`,
`Agents.test.tsx`, `CommentThread.test.tsx`, `NewAgent.test.tsx` — all
green (118 tests across the two runs).
## Risks
Low risk. Cosmetic label changes only; no data or API changes. Rows
without `metadata.managedByPaperclip` render exactly as before, so
self-hosted deployments with their own environments see no change. The
only self-hosted-visible wording changes are the adapter auth-check
message and the driver suffix omission on sandbox-driver rows in run
details.
## Model Used
- Claude (Anthropic) — claude-fable-5 (Claude Fable 5), Claude Code CLI,
extended thinking, tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
docs reference these labels)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
de9645ab73 |
fix(ui): surface live runtime status in the task-chat tail before the first transcript token (#11802)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs on sandbox execution targets spend their first minutes in preparation phases — config seed, workspace and skills sync into the sandbox — before the agent CLI produces its first transcript token > - The engine already reports these phases through the runtime-progress mechanism (`onRuntimeProgress` → `recordCurrentHeartbeatRunRuntimeProgress` → `currentStatusMessage` on the live run), and the pre-task-chat issue view surfaced them in its live status line > - The chat-style task view's live tail dropped that affordance: with zero renderable transcript entries it shows an opaque "Waiting for transcript..." for minutes, which reads as a hang (and prompted a real is-this-broken investigation on a healthy run) > - This pull request surfaces the live run's `currentStatusMessage` as the tail's empty-state message, with the generic wait text as fallback > - The benefit is that operators watching a sandbox run see "Syncing workspace to sandbox" instead of wondering whether the run is stuck — for every adapter, with no adapter identities involved ## Linked Issues or Issue Description **What existing behavior does this improve?** The chat-style task view's live transcript tail (`TaskChatThread` → `TaskChatLiveTail`), introduced with the experimental chat-style task view. **Current behavior** While a live run has no renderable transcript entries yet — the normal state for the multi-minute sandbox preparation window — the tail shows a static "Waiting for transcript...". The run's live `currentStatusMessage` (e.g. "Syncing workspace to sandbox", emitted by the sandbox-managed runtime's progress reporting) is available on the same live-run object but unused by this surface, although the earlier issue-chat view did display it. **Proposed behavior** When the tail is streaming a live run and no transcript rows exist yet, the empty-state message prefers the run's `currentStatusMessage`; the generic wait text remains the fallback when no runtime status has been reported (e.g. local runs that produce output immediately, or the brief pre-status window). **Reason and benefit** The preparation phases are real, reportable progress that the engine already emits. Showing them turns a minutes-long apparent hang into a legible status, for every adapter and execution target, using data the view already receives. ## What Changed - `ui/src/components/TaskChatThread.tsx`: the live tail's `emptyMessage` prefers `liveRun.currentStatusMessage` (guarded to the run the tail is actually streaming) over the static "Waiting for transcript..." fallback. Queued runs keep "Waiting to start...". - `ui/src/components/TaskChatThread.test.tsx`: a test covering both branches — a live run with a runtime status shows it (and not the wait text), and a run without one keeps the generic message. ## Verification - `vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/task-chat/TaskChatLiveTail.test.tsx`: 22/22 pass (including the new test) - `pnpm --filter @paperclipai/ui typecheck`: clean - Reproduced live: a `claude_local` run on a Daytona sandbox environment showed "Waiting for transcript..." for the full sync window; with this change the same window shows the streamed preparation statuses ## Risks - Low: a one-expression change to an empty-state string, active only while a live run has produced no renderable transcript rows. The fallback path is byte-identical to today. - `currentStatusMessage` is truncated/humanized upstream by the runtime-progress reporter; this surface renders it verbatim in the same muted style as the wait text. ## Model Used - Anthropic, **Claude Fable 5** (`claude-fable-5`) via Claude Code, with repository, shell, and Git tooling. It traced the runtime-progress mechanism end-to-end (engine emitter → heartbeat recorder → live-run API → both thread views), identified the dropped affordance in the chat-style view, and wrote the fix and test. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
cbe6395cc5 |
fix(ui): task chat composer clears on send; align carets and composer with thread (#11772)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task detail view uses a chat-style thread: agent turns, activity phases, and a message composer. > - A colleague's UX review found three defects: the composer draft stayed visible after send, disclosure carets in turns and status pills were misaligned, and the composer was not horizontally aligned with the thread column. > - These defects make the chat surface feel unpolished and cause confusion about whether a message was sent. > - This pull request fixes all three defects with small, targeted UI changes and adds regression tests for each. > - The benefit is a chat surface that behaves and aligns like users expect from a messaging UI. ## Linked Issues or Issue Description No public GitHub issue exists for this. Description of the underlying problems: **What happened?** Three UI defects in the chat-style task view: 1. After a user pressed send, the composer kept the draft text until the server round-trip finished. Fast typers could see stale text and doubt the message was sent. 2. The disclosure carets on collapsed agent turns and the caret inside the status pill did not share one alignment axis. They rendered at different x-offsets and sizes. 3. The composer container had different horizontal padding than the thread column above it, so the input box did not line up with the message bubbles. **Expected behavior** The composer clears the instant a send starts. All disclosure carets sit on one vertical axis with one size. The composer's left and right edges align with the thread column. **Steps to reproduce** Open any task in the chat-style task view. Type a message and press Enter — watch the composer text. Collapse and expand agent turns — compare caret positions. Compare the composer's horizontal edges with the message bubbles above it. ## What Changed - `TaskChatComposer.tsx`: clear the draft synchronously when a send starts instead of after the request resolves; restore the draft if the send fails. - `TaskChatTurn.tsx` and `TaskChatStatusPill.tsx`: use one shared caret alignment (size, x-offset) for turn disclosure and status pill carets, with supporting utility styles in `ui/src/index.css`. - `TaskChatThread.tsx`: align the composer container with the thread column padding. - Added or updated unit tests in `TaskChatComposer.test.tsx`, `TaskChatTurn.test.tsx`, `TaskChatActivityPhase.test.tsx`, and `TaskChatThread.test.tsx`. ## Verification - Run `pnpm vitest run src/components/task-chat/TaskChatComposer.test.tsx src/components/task-chat/TaskChatActivityPhase.test.tsx src/components/task-chat/TaskChatTurn.test.tsx src/components/TaskChatThread.test.tsx` in `ui/` — 4 files, 64 tests, all pass. - Manual: open a task in the chat view, send a message, and confirm the composer clears immediately. Collapse/expand turns and confirm the carets align. Compare composer edges with the thread column; the alignment fixes were also pixel-verified with screenshots during local review. ## Risks - Low risk. All changes are render-layer only; no server or data changes. - The composer now clears optimistically. If a send fails, the draft is restored, so no user text is lost. - Caret alignment uses shared CSS utilities; visual regressions would show in the existing component tests and in any screenshot diff. ## Model Used - Implementation: OpenAI Codex CLI coding agent (Codex model family, agentic tool use) via Paperclip's Codex adapter. - Review, verification, rebase, and PR preparation: Anthropic Claude, model id `claude-fable-5` (extended thinking, tool use), via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fd472d02ba |
Show ordered live blocker work in task chat (#11487)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The task view shows operators why work cannot continue > - The redesigned task thread now shows direct and ultimate blockers > - But it does not show the ordered task queue while a blocker chain has live work > - This pull request adds a compact ordered live-work queue to the redesigned thread > - The benefit is that operators can see completed, running, and queued dependencies without opening the larger legacy notice ## Linked Issues or Issue Description **What existing behavior does this improve?** The redesigned task thread blocker summary is improved. The merged predecessor is #11456. **Current behavior** The redesigned task thread shows compact direct and ultimate blocker links. It does not show the ordered queue when the blocker tree has live work. The legacy task view shows this queue in a larger notice. **Proposed behavior** Show a compact blue live-work queue at both ends of the redesigned task thread. Order completed tasks first, then running tasks, then queued tasks. Show a live terminal leaf as `Now running`. Return to the amber blocker links when no live dependency remains. **Reason and benefit** Operators can see the active dependency order without leaving the redesigned task view. The compact presentation preserves the new thread's low-chrome layout. **Breaking changes** None. The change only adds UI for blocker data that the task view already receives. ## What Changed - Shared the live blocker ordering helper between the legacy notice and the redesigned task thread. - Added compact ordered dependency links at the top and bottom of the redesigned thread. - Added a separate `Now running` link for a live terminal blocker leaf. - Preserved the compact amber blocker rows when live work is not present. - Added component tests and a Storybook state for the new presentation. ## Verification - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/IssueBlockedNotice.test.tsx` - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - `pnpm --filter @paperclipai/ui build-storybook` - Captured and reviewed the new Storybook state in a headless browser. ## Risks - Low risk. The queue appears only for blocked tasks whose blocker attention state is `covered` and whose dependency set contains live work. - The API does not provide an explicit queue position. The UI preserves the existing legacy ordering rule: completed, running, queued, then numeric task identifier. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The runtime does not expose the exact snapshot or context-window size. Reasoning, code execution, repository tools, and browser automation were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9e9f744f58 |
Show blocker links in the task chat (#11456)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip helps operators supervise agent work through tasks and task threads. > - The redesigned task thread shows the current work and its state. > - A blocked task did not show the dependency that prevented progress. > - Operators had to leave the thread to find the direct and final blockers. > - This pull request adds compact blocker links at the top and bottom of the task thread. > - The benefit is that operators can identify and open the relevant tasks without adding a large notice to the thread. ## Linked Issues or Issue Description **What existing behavior does this improve?** The redesigned task thread did not show which task directly blocked the current task or which task ultimately blocked its dependency chain. **Subsystem affected** `server/`, `packages/shared/`, and `ui/` task-blocker presentation. **Current behavior** A blocked task can open in the redesigned thread without a visible dependency link at the top or bottom of the conversation. **Proposed behavior** Show one compact amber row for the direct blocker. Show a second row for the selected final blocker when one exists. Render the rows at both ends of the thread. **Reason and benefit** Operators can see the reason for the blocked state and open the relevant task from the conversation. The compact rows preserve thread density. **Breaking changes** None. The new blocker-attention fields are optional. Existing clients remain compatible. ## What Changed - Added a compact task-chat component for direct and selected final blocker links. - Added the blocker rows to the top and bottom of populated and empty task threads. - Added link-ready blocker-attention details so an intermediate selected task stays on its correct direct chain. - Included blocker-link changes in the thread content key so pinned threads follow a newly added bottom row. - Added component, scrolling, server contract, and Storybook coverage for the new states. ## Verification - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx server/src/__tests__/issue-blocker-attention.test.ts` (38 tests passed) - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` ## Risks - Low risk. The rows only render while the task status is `blocked` and an unresolved blocker is available. - Long titles are truncated to keep each blocker on one line. The full task label remains available in the link title. - Older server payloads keep the original leaf-selection behavior because the new sampled details are optional. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The run used tool-enabled reasoning and code execution. The context-window size was not exposed to the run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |