mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
df662197809ca7e05b4eb49479eefb474fc261de
51
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
df66219780 |
fix(ui): retry the latest failed task attempt (#13765)
## Thinking Path > - Paperclip manages work done by AI agents. > - The task thread lets an operator retry a failed run. > - Legacy runs with transcript output did not get a failure marker. > - The thread could therefore offer Try again for an older failure. > - The server correctly reused that failure's existing retry, even when it had already failed. > - This change keeps the latest failure actionable and reports stopped retry responses to the operator. ## Linked Issues or Issue Description **What happened?** Try again could return success without starting work. An initial setup failure had an empty transcript. Its later retry produced output and failed. Only the initial failure had a retry marker, so the button kept requesting the initial failure's already-failed successor. **Expected behavior** Try again targets the latest failed attempt. An already-stopped retry response shows an error and refreshes the task's run state. **Steps to reproduce** 1. Start a legacy adapter task that fails before producing transcript output. 2. Retry it. Let this attempt produce output and a final failure comment before it fails. 3. Click Try again in the task thread. 4. Before this fix, the click targets the original failure and replays the stopped successor. **Paperclip version or commit** Reproduced against `1483bb8bcf`; the regression is also present on the branch base `8813a50105`. **Deployment mode** Authenticated server with a legacy adapter. The bug is in the shared task UI and retry API client. Related: #11650 adds a different recovery-notice action. This change fixes failed-run markers and retry response handling. Searches found no duplicate of this failure case. ## What Changed - Render legacy failure markers even when the run has a transcript or final comment. - Keep later cancelled automatic retries from replacing the failed run's retry action. - Reject already-stopped retry responses in the API client so existing error feedback appears. - Refresh task run queries after both successful and failed retry requests. - Add regression coverage for failed and timed-out attempts, execution gates with output, and retry response states. ## Verification - Before the fix, the new regression tests failed: two selected the original failure, and four accepted a stopped successor as success. - Targeted task-thread, retry API, marker, and issue-page tests: 279 passed. - `pnpm --filter @paperclipai/ui typecheck`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `pnpm check:token-gates`: passed. - Full UI suite: 6,529 tests passed across 626 files. - [CI run 35650385023](https://github.com/paperclipai/paperclip/actions/runs/35650385023): all 53 checks passed, including full workspace build, typecheck, unit/integration suites, and browser tests. The redundant local full-workspace test run was stopped after CI passed; it is not counted as a completed local pass. - Greptile: 5/5 on `608ee58c99`, with no review threads or unresolved comments. The branch is mergeable. - `pnpm -r typecheck` and `pnpm build` were attempted. Both stop at the Runner's Rust checks because this host has no `cargo`. The full workspace checks passed in CI. ## Risks Low risk. This changes UI presentation and response handling only. The server's exact-retry idempotency, authorization, execution ownership, and recovery gates remain in place. No schema changes or live task mutations. Existing documentation describes this retry action; the fix restores that behavior. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted checks; full-workspace limits described above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (existing behavior restored; no documentation change needed) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1ef3b08714 |
feat(ui): integrate agent personas across the app (#13171)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A stable agent persona is useful only when the same identity appears across the app. > - Lists, task messages, selectors, and activity feeds need inexpensive static avatars. > - Onboarding and agent headers need a larger character with expressions and pointer tracking. > - This pull request connects the persona foundation to those existing views and preserves onboarding draft assignments. > - Full-page stories and Linux checks make the placements and performance contract reviewable. ## Linked Issues or Issue Description **Problem or motivation** Agents need a stable visual identity in lists, tasks, onboarding, and configuration. External tools also need an image URL for that identity. **Proposed solution** Assign each agent a permanent palette from a fixed ClipLab character library. Store the assignment on the agent. Render and cache preset PNG URLs on demand. Use static images in dense views and one animated character in larger placements. **Alternatives considered** A generated image bundle requires a separate asset build. A live renderer in every avatar adds unnecessary work in large lists. Arbitrary uploaded images do not provide the requested shared character system. **Roadmap alignment** This improves agent identity across existing control-plane views. It preserves agent permissions, company boundaries, and status labels. ROADMAP.md has no separate ClipLab persona milestone. Related approaches: #2422 adds configurable image URLs and DiceBear generation; #5578 adds optional uploaded avatars. This work uses a fixed, versioned character library and preset URLs. ## What Changed - Replace agent icons with static persona images across lists, the sidebar, org charts, tasks, comments, selectors, activity, and dashboard views. - Put one animated character in the agent header. Let it follow the pointer across the page, with reduced-motion and touch fallbacks. - Add larger padded characters to agent creation. Keep the palette stable across draft refreshes and connection retries, then reveal it after success. - Pass appearance through shared projections rather than fetching each agent separately. - Add real full-page Storybook examples for the agent list, overview, task, dashboard, new-agent dialog, and connection page. - Add Linux screenshot, clipping, density, and 500-avatar performance checks. ## Verification - `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased tree. Persona lifecycle tests pass. - The rebased feature passes 38 Linux screenshot/performance checks, including both display densities, corner pointer positions, and the no-WebGL/no-live-download contract for 500 avatars. - The final Linux persona suite passes all 38 visual, lifecycle, density, and full-page checks using the standard Storybook configuration and real on-demand avatar endpoint. - Final local focused verification: 45 avatar/native-recovery tests pass; UI identity/routine tests, typecheck/build, token gates, and Storybook build pass. - Current-head CI passes: full workspace/server tests, all serialized server groups, typecheck/release checks, build, canary validation, and end-to-end shards. The build passed after retrying a native-runner concurrency-test failure; its three targeted cases also pass locally. - Manual inspection covered stable identities in the app, header placement, full-page mouse tracking, onboarding size, and task/dashboard placements. ### Screenshots Linux captures use synthetic Storybook fixtures. Full-page captures use reduced motion. The live character, mouse tracking, and disposal are checked separately. <details> <summary>Agent overview with the character in its header</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png" width="900" alt="Agent overview with the character in its header" /> </details> <details> <summary>Task messages and assignee identity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png" width="900" alt="Task messages and assignee identity" /> </details> <details> <summary>Larger onboarding character with room for expressions</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png" width="900" alt="Larger onboarding character with room for expressions" /> </details> <details> <summary>Dashboard agent activity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png" width="900" alt="Dashboard agent activity" /> </details> ## Risks - This PR depends on #13170, the persona foundation. Merge the foundation first, then retarget this PR to master. - Many placements change from icons to character silhouettes. Human avatars and authoritative agent status labels retain their existing behavior. - Only one character can render live per view. Reduced motion, hidden/offscreen content, touch input, and renderer failures use the defined fallbacks. - The full-page stories use fixture data. They do not contact a real company or complete real provider sign-in. ## Model Used OpenAI Codex, GPT-6 family. The exact model identifier and context window are not exposed in this session. Used code editing, shell execution, browser inspection, and Linux visual testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Tonio <tonework@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
43acbcc398 |
fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task state to provider sessions. > - Follow-up turns must retain provider memory and carry new user direction. > - Lost session IDs caused repeated context and extra input tokens. > - Native question answers and approval races could leave valid work blocked. > - This pull request repairs those paths and adds regression coverage. > - Agents can continue accepted work without repeating the conversation or losing the user's answer. ## Linked Issues or Issue Description Refs #13574. That merged PR shortened continuation prompts and moved question instructions into tool documentation. This change preserves sessions and fixes failures exposed by broader testing. Related runtime work: #13408 and #13410. **What happened?** Native follow-up turns could lose the provider session ID. Completion guidance could replace the original task with its latest comment. Claude native questions could remain pending after the user answered. Approval during a running tool call could suspend the run before the tool response arrived. Onboarding and chat handoff instructions also caused repeated planning or missing plan documents. **Expected behavior** Reuse a valid provider session. Send only new events when that session already has the history. Preserve the task requirements and apply later user direction. Store the question answer and deliver it to the waiting run. Finish governed tool responses before suspending. Execute the accepted plan without asking for the same approval again. **Steps to reproduce** Run the continuation, local-session-integrity, first-task, and agent-chat suites with native Codex and Claude. Include provider-question-bridge, accept-while-running, and plan-handoff. **Paperclip version or commit** This branch is based on master |
||
|
|
45586170e1 |
fix(ui): reveal task history while retries wait to start (#13597)
## Thinking Path > - Paperclip lets people manage AI agents and inspect their work. > - Task conversations combine comments with run transcripts. > - The first reveal waits for the relevant history to load. > - Scheduled retries have not started, so the log reader does not hydrate them. > - Waiting for those retries keeps the whole conversation hidden after its header loads. > - This change reveals available history while a retry waits to start. ## Linked Issues or Issue Description **What happened?** A task header loads, but the conversation stays behind the loading overlay while a linked run has `scheduled_retry` status. The readiness check expects that run in `hydratedRunIds`, although the log reader does not read its log. A retry delay can therefore become a task-loading delay. **Expected behavior** Show existing comments and transcripts once their data is ready. A retry that has not started must not block the conversation. **Steps to reproduce** 1. Open a task with existing comments and a linked run waiting in `scheduled_retry`. 2. Let the comments and other run transcripts finish loading. 3. Observe that the header loads but the conversation stays hidden until the retry changes status. **Paperclip version or commit** Reproduced on master at `3f1d897a7` with regression tests. **Deployment mode** Web UI, built from source. This is a client readiness bug. Searched public issues and PRs for scheduled retries, conversation loading, and history loading. No duplicate found. Related: #10255 changes the server response for active runs with no log; it does not address this scheduled-retry readiness gate. This focused bug fix does not add a roadmap feature. ## What Changed - Exclude `scheduled_retry` runs from the initial task-history readiness gate. - Extend transcript mocks to expose per-run hydration state. - Add six regression cases across legacy and native runs. Existing history loads during retry waits, while unhydrated running and completed runs still block the first reveal. - Document the exception in a code comment. No user command or configuration changes need documentation. ## Verification - All six new cases fail before the production fix. - `pnpm --filter @paperclipai/ui exec vitest run src/components/TaskChatThread.test.tsx src/components/transcript/useLiveRunTranscripts.test.tsx src/components/transcript/useNativeRunTranscripts.test.tsx`: 164 passed. - `pnpm --filter @paperclipai/ui typecheck`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `pnpm check:token-gates` and `git diff --check`: passed. - `pnpm -r typecheck` and `pnpm build`: attempted; both stop in the native runner package because the local machine has no Rust `cargo` executable. UI checks pass separately. - `pnpm test:run`: started locally, then stopped before completion after full CI passed on the same commit. The local full-suite result is incomplete. - CI on `46daa8b06`: all 53 checks passed; two optional Storybook jobs were skipped. This includes the full test matrix, build, typecheck, native runner checks, browser tests, and canary dry run. - Greptile: 5/5 on `46daa8b06`, with no review threads or actionable findings. The branch has no merge conflicts with master. ## Risks Low risk. The exception applies only to runs waiting in `scheduled_retry`. Running and completed runs retain their existing readiness checks. Retry scheduling, transcript fetching, and server behavior stay the same. No migration is needed. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and test tools. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9fd2e50310 |
feat: create company skills from runner tasks (#13538)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner gives agents tools to change company resources. > - Users need agents to save reusable skills during a task. > - A saved skill needs a visible result that users can inspect and edit. > - This pull request adds `create_skill` and a task feed card linked to Skill Studio. > - Users can open the saved skill from the task and edit the same resource. ## Linked Issues or Issue Description **Subsystem affected** Runner tools, company skill storage, task feed, and Skill Studio. **Problem or motivation** The Runner has no dedicated tool to create a company skill. A user cannot follow a creation result from the task feed to the saved skill. **Proposed solution** Add a company-scoped `create_skill` tool. Save the skill with the existing company policy. Add one creation card to the task. Open a named sidebar tab from that card. Let the user open the same skill in Skill Studio. **Alternatives considered** An agent can write a local file, but that file is not a company skill. A second document copy in the task would become stale after a Studio edit. The sidebar therefore reads the saved skill directly. **Roadmap alignment** This extends the shipped Skills Manager, Skill Studio, and Skills Store milestone. The maintainer requested and approved this scope. Search found no duplicate `create_skill` PR or issue. Related UI validation work: #8715. This PR does not change that validation display. ## What Changed - Add the real Runner tool, its contract, and its mock implementation. - Validate the complete SKILL.md and derive company, task, agent, and run identity from authentication. - Apply the existing company skill policy. Do not assign the skill to an agent. - Make keyed retries return one skill and one creation event. Reject conflicting retries. - Make concurrent file creation safe. Never replace an existing published skill during creation. - Add a creation card, a named sidebar tab, and an Open in Skill Studio action. - Show saved Studio edits when the user returns to the task. - Add storage, policy, mode, retry, UI, and Product E2E tests. Document the tool. - Fix deleted-name reuse, onboarding panel persistence, immediate feed refresh, and mock validation parity from review. - Serialize Studio file edits and renames with skill deletion and recreation. Reject stale editor requests before they can change a replacement skill. - Generate the standalone mock parser and validator from the production contract. Use portable UUIDs so the browser scenario bundle builds. ## Verification - All latest-head PR checks pass on `145dd76a5`, including all server shards, browser E2E, Runner verification, build, typecheck, and release dry run. Greptile: 5/5 with no open findings. An interrupted CI runner was retried successfully. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Review regressions: 73 storage tests, 6 real API tests, 63 UI tests, and 61 semantic runtime tests passed. Parser synchronization passed. - CI exposed existing fire-and-forget Sentry test races. Reproduced the resumption race locally, then synchronized the related sweep and finalizer assertions on the actual report; all 27 tests across the three affected files pass. - Runner scenario browser build and strict content-security-policy check: passed. - Runner suite: 2,012 tests passed; 10 skipped. - `pnpm test:run`: the general-server batch had 12,416 passes and two failures. The old tool-count assertion was fixed; all 16 authority tests then passed. The chat webhook test had a socket error; it passed four isolated reruns. - Both workspace test groups passed. The isolated route suites completed. Two socket failures in the initial route batches passed on individual reruns; all remaining 61 files passed. - Product E2E `create-skill-studio`: passed with local Codex and local ACPX Claude. - Manual browser test: submit a task, observe the real tool call and creation card, open the sidebar, edit in Studio, save, and return. The task reached Done. The saved second revision and sidebar tab survived a server restart. - The new companion headless Runner Eval passed. Companion coverage PR: https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not run because no immutable runner image was configured. ## Risks - Database writes and local file writes cannot share one transaction. Recovery accepts only an exact file-for-file retry after a database rollback. Conflicting files remain untouched. - The sidebar displays the current skill. The feed card remains the historical creation receipt. - No database migration, dependency, or workflow change is included. - Remote Daytona behavior still needs a run with a configured immutable image. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and browser verification. OpenAI `gpt-5.6-luna` assisted with bounded implementation and eval work. Both used code execution and tool access. The host did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4cdc7b231 |
fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane prepares task workspaces before it starts an agent. > - Workspace preparation reads Git state so it can preserve edits and exclude private files. > - A failed scan was treated as a non-Git folder and lost its actual failure code. > - The resulting generic setup failure could not recover, even when the cause was temporary. > - This pull request keeps the cause and uses the existing bounded retry schedule before provider startup. > - Tasks can recover without human intervention, while permanent failures and exhausted retries stop with useful guidance. ## Linked Issues or Issue Description **What happened?** A Git scan error during managed repository preparation became `Configured repository folder is not a Git checkout`, followed by generic `setup_failed`. The agent never started. Generic recovery could not distinguish a temporary timeout from a bad workspace configuration. **Expected behavior** Keep the closed scan error code. Retry temporary timeouts and queue saturation under the existing shared budget. Preserve edits, exclusions, ownership, and pause gates. Stop permanent failures and exhausted retries with a specific explanation. Do not replay historical generic setup failures. **Steps to reproduce** 1. Configure a task project with a local Git source that must be copied into its managed repositories. 2. Make the ignored-file scan exceed its timeout before the agent starts. 3. Before this fix, the snapshot returns null and the run ends as non-retryable `setup_failed`. 4. Use the disposable browser fixture in `tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and test the full recovery path. **Paperclip version or commit** Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`. **Deployment mode** Built from source. The defect is in core workspace setup, not a specific model provider. Related work: Refs #13442 (managed repository preparation), Refs #11572 (bounded Git scheduler), Refs #12997 (separate adapter startup retry work), Refs #13469 (separate terminal-workspace scan performance work). ## What Changed - Return the non-Git fallback only for repository discovery. Propagate failed scans of a confirmed repository. - Replace full ignored status output with an ignored-only directory listing. Preserve NUL-delimited paths and exclusions. - Preserve typed, sanitized scan errors through workspace preparation and persist pre-provider failure details. - Retry only timeouts and queue saturation, using the existing durable two-retry budget and issue gates. Prevent generic recovery from adding another budget. - Show workspace-specific failure copy and actionable exhausted-recovery notices. - Add red-green unit tests, real-database restart and retry-boundary tests, and opt-in browser acceptance fixtures with real Git subprocess timeouts. - Document the recovery contract and browser verification procedure. ## Verification - Red: injected scan failures returned null instead of rejecting; setup lost the timeout code; task-thread and recovery notices had generic copy. - Green: 119 focused adapter/backend tests, 20 recovery-boundary tests, and 136 task-thread tests. - `pnpm -r typecheck` — passed. - `pnpm build` — passed on the final production code. - `pnpm check:token-gates` — passed. - The initial local `pnpm test:run` overlapped source edits and was interrupted after two late-added assertions saw pre-fix behavior; it is not counted as a green full run. A fresh final-head run passed all 229 tests across the six affected adapter/backend/UI suites. The clean latest-head CI full test matrix passed: all five general-server shards, all five serialized-server shards, and all three general-workspace shards. - Latest-head CI also passed all three browser shards and their aggregate gate, typecheck and release registry, build, runner verification, canary dry run, policy, Docker context integrity, and security gates. Greptile: 5/5, with the review thread resolved. - Browser: created a task in a disposable instance. A real Git timeout scheduled recovery, the next run completed through the run-scoped API without manual Retry, and Done survived reload. The deterministic process worker checked preserved source edits and excluded private files; no model calls were made. - `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec playwright test --config tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6 minutes). The persistent case made exactly three failed attempts, never started the worker, showed the cause-specific notice, stayed stopped for another scheduler tick, and retained Blocked after reload. - Extra red-green coverage: 50 recovery tests passed after fixing an exhausted-bootstrap classification that incorrectly implied unknown provider actions. Missing or uncertain evidence still retains the safety hold. - Verified the documented Git executable override during repository seeding. ## Risks - A confirmed repository scan failure now fails closed instead of falling back to directory sync. This prevents unfiltered copying but makes previously hidden errors visible. - Temporary host problems can create up to two additional setup attempts, 30 seconds apart. Permanent scan errors do not auto-retry. Generic recovery cannot reset this budget. - The durable retry path still enforces ownership, pause, and work eligibility. Integration tests cover restart, duplicate promotion, pause, exhaustion, and non-retryable categories. - No schema migration, new runtime setting, new retry budget, production deployment, or historical task replay. ## Model Used OpenAI Codex, GPT-5-based coding agent, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
78ce96a48b |
fix(runner): restore native Claude context and read permissions (#13422)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner supplies each agent with instructions, assigned skills, and tools. > - Native Claude lost its skill snapshot before provider launch. It also changed exact model IDs to aliases. > - The read permission mode denied ordinary Paperclip reads because this runner has no interactive approval handler. > - This pull request restores the missing context and permits only assigned tools that Paperclip defines as reads. > - Other requests that need approval stop with a clear action for the operator. They do not retry automatically. > - Agents can complete read tasks. Users can see why a restricted task stopped. ## Linked Issues or Issue Description **What happened?** Native Claude could not find assigned skills. The ACP layer could replace an exact model ID with an alias and then fail identity verification. The default `approve-reads` mode denied Paperclip read tools. A denied write showed a generic transport failure. The tool descriptions also called live operations “mock” operations. The completion prompt said to call a completion tool once, although the protocol can reject a claim and require a corrected call. **Expected behavior** Load assigned skills before launch. Keep the selected model ID. Allow assigned Paperclip reads. Stop an operation that needs approval with clear instructions when no approval handler exists. **Steps to reproduce** 1. Configure a native Paperclip Runner agent with ACPX Claude and an exact model ID. 2. Assign a skill and ask the agent to use it. 3. Select the read permission mode and ask the agent to read task context and list documents. 4. Ask the agent to write a document. Check the task state and recovery message. **Paperclip version or commit** Reproduced from `d351e08deee1b49d3467a950d1a3f01131943441`. **Deployment mode** Local native runner with real Claude, Rust runnerd, the Paperclip server, and embedded PostgreSQL. Related: #13196 fixed remote skill staging in the legacy `claude_local` adapter. The native runner uses a separate path, which this pull request fixes. ## What Changed - Carry the runtime context through Rust and the ACPX sidecar. Load assigned Claude skills after the provider lifetime lease is held. Refresh the files on each open. - Write the exact requested model ID into the isolated Claude settings. - Grant exact MCP permissions for the intersection of assigned tools and Paperclip's read catalog. Provider hints cannot grant access. Existing task-control permissions stay in place. - Stop requests that need an unavailable approval handler. Preserve the typed error through the server. Show “Approval required” on the task and require operator action without automatic retry. - Label the setting “Allow Paperclip reads.” Remove “mock” from live tool descriptions and regenerate the contracts. - Change one completion-prompt sentence to require one accepted result. Add regression tests and update the runner documentation. ## Verification - Red/green regression tests cover read admission, unavailable approval handling, server recovery, and the task error message. - A real Claude read trial failed before the fix and completed after it: [red trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=95e6a326e24437f1), [green trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=85a1f347db5604d1). - Full browser tests used real Claude and the production tool authority. Reads completed. A write stopped with “Approval required.” No document was created and no automatic retry was scheduled. All nine checks passed: [events and screenshots](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=1c4ff10be0d779cc). - A separate full browser test assigned a skill, invoked it, and completed with a marker absent from the task prompt: [skill evidence](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=ba5fbcb29a60be52). - Post-rebase checks passed: 128 targeted runner tests, 46 Rust tests, 27 sidecar/protocol contract tests, `pnpm -r typecheck`, and `pnpm build`. The server/UI red-green checks and `pnpm check:token-gates` also passed. The full suite passed in CI, including all general and serialized test shards, all three browser shards, and runner `check:all`. The duplicate unsharded local full-suite run was stopped after CI passed. Braintrust links require project access. The traces contain provider/runner events and app outcomes. They do not contain raw model HTTP requests. ## Risks - `approve-reads` now allows assigned Paperclip reads. Other operations that need approval stop the turn. Users must review the operation and change permissions before retrying. - The Claude settings depend on the pinned ACP and SDK behavior. Automated tests and real Claude trials cover this boundary. - The native Codex path is unchanged. The ACPX Codex fallback shares the clearer approval failure handling. - Local execution was tested end to end. Remote execution was not run. This change has no database migration. ## Model Used OpenAI GPT-6 through Codex. The agent used reasoning, code execution, and browser tools. The exact deployment ID and context-window size were not exposed in the session. Live acceptance tests used `claude-sonnet-5`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
422287eecd |
fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task messages, provider execution, and task outcomes. > - First-time user tests exposed gaps in recovery, completion permissions, message delivery, and Stop behavior. > - These gaps left usable output hidden, completed work waiting for bookkeeping, or safe work unable to continue. > - This pull request fixes the shared lifecycle and receipt paths while preserving process ownership and action checks. > - Users can continue work with accurate task state and durable messages. ## Linked Issues or Issue Description **What happened?** A stopped local Codex execution could remain blocked even after its processes had stopped and its complete transcript proved that no external action needed replay. Claude under Conservative permissions could fail to call task completion tools. Recovery could reuse an assistant item ID and overwrite prior output. A delivered comment could remain marked uncertain after navigation. Stop could look like Pause or a new recovery incident. Workspace contention could look like cancellation. A direct reply reopening Done could enter a clarification loop. **Expected behavior** Recover automatically only with verified termination and complete action receipts. Preserve answers and messages. Keep task completion available under Conservative permissions without broad tool access. Show crashes as Blocked, actual human decisions as In Review, and ordinary workspace contention as waiting. Stop the current response and allow a new direction. **Steps to reproduce** 1. Create ordinary response tasks with local Codex and Claude Code, then send follow-up messages through the task composer. 2. Interrupt a disposable local Codex runner during text-only work. Verify automatic continuation and retained output. 3. Stop a response, send a new request, answer a clarification, and reopen completed work with another message. 4. Navigate or reload while a comment submission is pending. Confirm the exact persisted request receipt settles it without removing newer draft text. 5. Run two tasks in a shared Daytona workspace. Confirm waiting does not appear as failure. **Paperclip version or commit** Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`. Current integration base: `6cef9743c`. Both operator-interruption and workspace-waiting guards are preserved; native restart and legacy permission rules remain documented. **Deployment mode** An isolated source-built test-drive instance, with real local Codex and Claude Code providers and disposable Daytona environments. Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163. This PR addresses additional failures from ordinary task journeys, including controller restart handoff and repeated warm sandbox setup. Historical task status reconciliation is excluded. ## What Changed - Persist runner ownership immediately at spawn and resume an explicitly adopted runner even when the controller crashed before the first driver checkpoint. Detach the controller safely across graceful restarts, including session startup. Prevent an old finalizer from suspending or signaling an adopted runner. Checkpoint idle warm sessions before shutdown. Preserve the same run and queued follow-up messages. - Scope saved legacy queue successor checks to the queue owner while preserving ordinary task locks, operator identity, assignment gates, and exactly-once delivery. - Preserve managed Codex credential files when an old session is detached for restart; normal owned cleanup still copies refreshed auth back and removes the scoped copy. - Reuse the bound warm shared sandbox and fully verify an existing staged provider pack before using it. This avoids repeated uploads when the pack is already valid. - Add a narrow local Codex replacement path with stopped-process proof, a closed transcript inventory, exact completion receipts, and fresh-session lineage. Preserve no-replay holds when evidence is incomplete. Recovery may clear only the same run's recorded Blocked status version; manual re-blocking and dependency changes invalidate that receipt, while queued comments do not. Later blocks stop scheduled, queued, and final dispatch; queued/final checks re-read dependencies even when the task status stays In Progress. - Permit only task delivery and human-input tools through the isolated Claude runner's exact task bridge. - Scope assistant item identity to the provider turn and ignore only authority-free Codex skill-change notifications during startup. - Reconcile composer submissions by client request ID across response loss, navigation, and reload. Retain text typed during delivery. - Keep acknowledged run-only Stop neutral and show workspace contention as waiting. Project exhausted native failures as Blocked. - Restore the guarded task-page retry action for failed legacy runs, including the server-supported explicit new-attempt path for stopped conversation adapters. Preserve native/process recovery holds and avoid promising Retry while a decision or execution gate hides it. - Refresh delivered artifacts and handle direct user replies that reopen completed work without a clarification loop. - Check the embedded PostgreSQL PID, data directory, and actual port before connecting or migrating. - Document accepted behavior and add focused regressions at lifecycle, route, transcript, and UI boundaries. ## Verification - Final head `fece606ac2` passes the complete GitHub CI matrix: **34 green checks, two expected Storybook skips, no failures or pending checks**, including `ci / verify`, `ci / e2e`, full runner verification, typecheck, build, every server/workspace shard, and all browser shards. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34727183287). Greptile is **5/5 with no open findings**. The final two commits only refine test fixtures; both affected suites pass 24/24 locally and in CI, with server typecheck green. - Complete local Vitest coverage uses the canonical groups/shards: all 635 general server suites, all 145 serialized suites, and all workspace packages. The aggregate began on `0a8001c18` while the final queue fix arrived: 23,903 passed, five failed, 87 skipped. The five port/socket/timing failures passed unchanged in follow-ups (60 tests in the exposure/file suites and 412 tests covering the serialized failures and unrun tails). The final queue/operator-identity suites separately passed 52/52. This is aggregate coverage plus explicit reruns, not a pristine single-command final-head run. - After integration with current master, queue/operator-identity/continuation suites passed 162/162 and affected UI suites passed 140/140. ACP Stop/continuation and legacy task/Inbox/message browser suites passed 9/9, including both task recovery Retry and thread Try again, automatic saved-message delivery, exactly one new run, Done, and retained output after reload. The default process Stop/Pause/Resume browser case passed (the native-provider case is opt-in and skipped by default). The complete Board attachment/receipt browser suite passed 11/11 on a disposable instance, covering both composers, exact receipts after lost responses, no replay, bound attachments, and newer drafts after reload. - Blocking-intent regressions cover pre-existing Blocked, a mismatched run/cause, an explicit manual re-block, changed dependencies, a queued comment after failure, and a block arriving between scheduling and provider dispatch. The negative cases reproduced before the fix. All 478 affected executor/recovery/dispatch tests passed; both database suites ran separately after availability-probe skips in the first combined command. The final late-dependency check passed all 143 affected recovery/dispatch tests (zero skips) after two new negative cases reproduced the bug. - Focused runtime regressions cover awaited runner ownership publication, authenticated adoption before the first checkpoint, old-finalizer detachment, idle and busy warm-session shutdown, rejected checkpoint propagation, provider-pack verification, and managed-Codex credential preservation. Four managed credential detachment cases reproduced the bug before the fix; normal owned cleanup still succeeds exactly once. - Live local Claude: SIGKILL 2.6 seconds into startup recovered the same run automatically in 53 seconds, then a normal follow-up completed in 24 seconds. SIGTERM 2.5 seconds into startup preserved the same run (54 seconds) and its queued follow-up (21 seconds). Answers remained visible and the task reached Done. - Live Claude Daytona: a warm follow-up retained its sandbox and fell from 121 seconds to 44 seconds. A separate cold turn took 127 seconds; after controller shutdown and checkpointing, its follow-up completed in 33 seconds with the same sandbox, workspace, native session, and runner. Both answers remained visible and the task was Done. - Other live journeys covered task completion and follow-up with local and Daytona Codex, local Codex crash recovery, Stop then new direction, clarification response, live artifact refresh, and shared-workspace waiting. - Validation limits: the opt-in native composer Stop/Pause→subtree Resume fixture exposes terminal/result ordering and subtree-cancellation attribution bugs that can leave a child task blocked; that new finding is assigned to a separate follow-up and is not claimed fixed here. Default CI skips this optional native-provider fixture. Managed-Codex credential handoff and the queue-agent integration use automated regression evidence. Cold custom provider-pack uploads still add startup latency. ## Risks - Automatic replacement remains deliberately narrow: local Codex, verified stopped identities, unchanged retained state, and a complete text/completion-only turn. Unknown actions, partial history, or changed ownership remain blocked. - Claude completion permission handling changes an upstream package patch. The exact isolated task bridge must remain pinned; unrelated tools keep their existing permissions. - New task failure projection changes user-visible status. No historical status backfill or database migration is included. - This is a broad lifecycle fix across server and UI. Live proof covers graceful local Claude restart during startup and idle Claude Daytona session recovery across controller shutdown. Live abrupt SIGKILL during local Claude startup also recovered the same run. Unknown ownership or missing action evidence still blocks reuse. Cold custom provider-pack uploads still add startup latency; this change avoids unnecessary repeat uploads. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, browser automation, and tool use. The exact hosted model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
df984cbc2c |
fix: dispatch queued legacy messages with operator identity and task permissions (#13315)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task conversations save messages that arrive during an active turn. > - Legacy adapters deliver these messages in a later turn. > - A run can stop before the saved queue is delivered. > - The Interrupt button previously required an active run, so it could not release this queue. > - This pull request lets a board operator send the saved queue after the run stops and retries queues missed during finalization. > - Manual dispatch must use the clicking operator and must not require permission to create agents. > - The task can continue without a duplicate message or a second execution owner. ## Linked Issues or Issue Description **What happened?** A legacy task retained a queued message after its run stopped. Interrupt was disabled because the queue had no active target. Finalization and deferred message admission can also leave a queue without a successor. **Expected behavior** Interrupt sends the saved messages when no runner is active. Messages that arrive during normal completion are delivered automatically. An uncertain previous execution still requires proof that its process or sandbox stopped. **Steps to reproduce** 1. Queue a user message during a legacy conversation turn. 2. Let the turn stop or simulate a server restart before queue promotion. 3. Open the task with a deferred queue and no active run. 4. Try Interrupt. Before this change, the button is disabled. **Paperclip version or commit** Reproduced against `8f40b4ad4`. **Deployment mode** Legacy conversation adapter. The same persisted queue state is covered with an isolated PostgreSQL fixture. Related public work: #13275 adds active legacy interruption. #13291 addresses automatic sandbox conversation recovery. This change handles explicit saved-queue delivery and late queue promotion. ## What Changed - Accept a null Interrupt target while retaining queue identity, revision, company, and assignee checks. - Save the operator's request on the existing queue. Reuse normal admission after verified stop, including older messages, different authors, and queues whose original wake came from the system. - Strip interruption authority from caller-supplied wake payloads. Only the board queue route can persist that authority. - Retry durable interruption requests after restart and deferred queues after legacy cleanup. - Let an explicit Interrupt retry cleanup for its stopped run, including old ephemeral leases that recorded success without a provider stop receipt. Preserve retained resources, other lease owners, and the automatic retry limit. - Preserve the server's waiting explanation when normalizing and combining queue entries. - Revalidate the consumed board queue receipt at dispatch so a different message author does not cause setup failure. - Use the Interrupt user's execution identity for the new run. Preserve original message authors. Validate the receipt independently at startup and inherit the resulting identity on retry. - Persist authenticated board authority for ordinary manual wakes too. Adopting someone else's queued messages cannot switch a manual run to that author's permissions. Strip caller-supplied authority markers and retain private conversation ownership checks. - Keep the clicking user when a manual wake is merged into an older deferred receipt. Update its requester and payload in the same transaction. - Use the same current-queue/revision API on task details and pipeline conversations; show Interrupt after a legacy target stops. - Keep manual wakes out of active runs, including unscoped agent wakes. They receive their own execution identity; a matching receipt requester is not sufficient because an exact retry can retain a different originating identity. - Authorize both existing-agent wake endpoints with `agent:wake`, available to active non-viewer company members. Keep `agents:create` for hiring. Validate the stored task and current assignee before an exact task retry. - Reject viewer Interrupt requests before saving intent or stopping execution. Keep external chat retry authorization and per-action agent/user permission checks. - Preserve edits and discards until dispatch. Prevent another queue promotion when the same agent already has a successor. Keep independent reviewer recovery available. - Suppress cancelled/failed run toasts for intentional operator interruption. Keep ordinary runtime error notices. - Add UI, route, admission, restart, successor ownership, and toast regression tests. Document the behavior. - Reuse the existing socket reservation helper for both credential-quorum test cases after CI exposed an ambient-port collision. This changes test preparation only; production credential staging is still called exactly once. ## Verification - Failing regression tests reproduced the message-author identity bug and an operator's `agents:create` rejection before the fixes. - All 316 focused tests pass across eight route, queue, identity, authorization, continuation, and responsible-user suites, including the 44-test rerun of queue admission and actual startup after the final manual-wake restriction. Regressions reproduce cross-user merging both with and without a task, and same-requester receipt ambiguity. The cross-company existence guard also passes both tests. - Startup integration tests reach adapter execution under the clicking operator and retain that identity through follow-up. Coverage includes mixed authors, adopted queues, system-origin queues, restarts, forged or stale receipts, viewers, suspended memberships, changed assignees, private conversations, and caller-supplied authority markers. - The earlier queue/cleanup/UI regression suite passed 402 tests. The final review corrections pass another 180 tests across queue admission/persistence, real heartbeat startup, UI API, conversation rendering, and pipeline suites. Regression tests reproduced both review findings before correction. The final head has a 5/5 review with no unresolved threads. Full CI passes on `c2002979c`, including every general and serialized server shard, all browser shards, Paperclip Runner verification, typecheck, build, canary dry run, and the aggregate gates. - Full `pnpm -r typecheck`, `pnpm build`, and UI token gates pass after the final application changes. CI identified an outdated task-page API mock after the shared helper extraction; the fixture now exercises the real helper, and all 131 task-page/API tests pass. The final application build passes with the additional manual-wake restriction. - CI exposed a pre-existing port collision in the Codex credential-quorum fixture. It reproduced locally; both listener cases now use the existing bounded reservation helper. All 41 credential tests pass on rerun. One intervening local run hit a separate ambient bind collision in the two-occupied-port case. - The full local `pnpm test:run` attempt was stopped after host contention caused focused-suite timeouts. The affected focused tests passed on rerun. An expiring trace fixture and a missing private-conversation state were corrected. The successful full CI run is the complete-suite verification. - Hosted Interrupt previously cleared the original queue and produced exactly one successor with neutral interruption feedback. It exposed the dispatch authorization defect. Retry on that earlier build was rejected for missing `agents:create` before creating another run. - Deployed the final application build (`38257f391`) to the scoped hosted instance and verified readiness. The latest PR commit changes only the credential test fixture; application code matches that deployment. A live Retry by the same operator without `agents:create` created one successor attributed to that operator, passing the former dispatch permission gate. Startup then stopped at `configuration_incomplete` because that operator has not configured their required personal Claude Code OAuth secret; the post-deployment run page confirms the operator identity and no provider work started, and the My secrets UI still shows the token as not set. Provider execution remains unverified pending that credential. No permission grants or credentials were changed. ## Risks Queue admission and finalization can race. The task lock, durable queue receipt, current comment IDs, and successor guard prevent duplicate dispatch. Process and lease stop checks, task pauses, approvals, ownership, and budgets remain in force. The API change only allows null on legacy Interrupt; native steering still requires an active run. No schema migration is required. Active non-viewer board members can now invoke existing agents without agent-creation permission. Agent self-invocation rules, raw provider-trace admin access, task retry scope, external chat authorization, and action-specific user/agent permissions remain enforced. ## Model Used OpenAI GPT-6 through Codex. The session does not expose an exact backend model ID or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d6232e7b0 |
feat: reuse provider sign-in across AI connection workflows (#13248)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect provider accounts during onboarding and agent setup. > - They should reuse and manage those accounts through the existing Connectors interface. > - A second login wizard would diverge from the established provider workflows. > - This pull request composes the existing sign-in components into Connections and agent configuration. > - Users can select accounts without changing their agent's harness or model. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add compact AI-account management to the existing Connectors pages. - Reuse AgentProviderConnection, AdapterLoginPanel, AdapterLoginChrome, and authentication controllers. - Add the shared connection picker to agent setup/settings and task requests. - Preserve onboarding's sequence and reuse existing accounts. - Add local-login recovery, retry, cancellation, and React StrictMode handling. - Add interactive Storybook scenarios, design-guide examples, and app acceptance checks. This is part 2 of the AI Connections change. The runtime foundation in #13247 is merged. This PR now targets master. ## Verification - Updated against master `47ded8bf9`, including the landed runtime foundation and upstream task-search changes. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All CI test, browser, build, packaging, and runner jobs passed on final head `dd17d3211931dd70aaa6ea619d83a7f9966dd18e`. The fresh Greptile review is 5/5, the security scan passed, and there are no unresolved review threads. The final CI aggregate gates passed. - Local general-server coverage passed 11,804 tests; three port-collision failures passed in an isolated 25-test rerun. All 6,111 UI tests passed. CLI coverage passed 484 tests; its remaining doctor test requires port 3199, which is occupied by an unrelated report server on this Mac. The complete CLI suite passed in CI. - Live browser testing verified Codex API-key reconnect inside a task card on desktop and phone. Real provider runs resumed and completed with unchanged connection/grant identity and agent routing. - Tested opening, cancelling, reopening, and completing connection creation. A regression confirms Connect another account cannot submit the new-agent form or copy provider keys into agent settings. - Added shared inline repair, automatic local sign-in checks, and responsive connection dialogs. Standalone Daytona installation ignores workspace configuration and suppresses dependency scripts. Its standalone build also passed with CI's exact pnpm 9.15.4. - Destructive live tests are excluded by default. Explicit opt-in, local deployment checks, and matching disposable fixture identities are required before any mutation. ## Risks - Local Codex/Grok creation requires the connection-specific terminal login command. - Browser sign-in uses the existing supported-environment controllers. - This update verifies live local Claude detection and Codex API-key task repair. New subscription authorization/refresh and independent-human/native-runner isolation were not reverified in this update. - No agent automatically adopts managed Connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
586b5ec828 |
fix(ui): improve mobile task spacing (#13304)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task detail view is the main place for task conversation. > - The mobile view used too much horizontal space for gutters and controls. > - The status control and identifier also reduced the title width. > - This pull request gives the title its own mobile row and reduces mobile padding. > - The benefit is a clearer task view with more space for useful content. ## Linked Issues or Issue Description **What existing behavior does this improve?** The mobile task detail header, conversation list, and message composer. **Current behavior** The title shares one row with status and identifier controls. The conversation and composer also use larger mobile gutters than needed. **Proposed behavior** The title uses the full mobile width. Status and metadata use the next row. The conversation and composer use smaller mobile gutters. **Reason and benefit** Long titles have more readable line lengths. The reduced padding gives task content more room on small screens. **Breaking changes** None. ## What Changed - Put the task title on a full-width mobile row. - Move the mobile status and metadata below the title. - Reduce the mobile conversation and composer padding. - Update the composer dock class test. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/task-chat/TaskChatComposer.test.tsx ui/src/pages/IssueDetail.test.tsx --reporter=dot` - `pnpm --filter @paperclipai/ui build` - `pnpm check:token-gates` - `git diff --check` ## Risks - Low risk. The layout changes apply only at the mobile breakpoint. - Desktop spacing stays unchanged. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.6, with reasoning and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb640ec129 |
fix(execution): keep blocked wakes waiting without repeated runs (#13236)
## Thinking Path > - Paperclip manages AI agents and their work. > - Wake admission decides when a task can create an execution run. > - Recovery can prohibit replay while the previous execution needs review. > - Dependency reconciliation kept creating runs before dispatch rejected that same hold. > - Each rejected run added another startup notice without doing useful work. > - This change checks the hold during admission and records repeated automatic waits once. > - Tasks keep their messages and can resume when the current gates permit execution. ## Linked Issues or Issue Description Related changes: Refs #13173 (stale completed-task continuations). Refs #12651 (dependency waits during recovery). **What happened?** A blocked task with completed dependencies can remain under a durable execution reconciliation hold. Each scheduler pass created a queued run. Dispatch then cancelled it before the adapter started. The skipped wake did not satisfy dependency wake deduplication, so this repeated and filled the conversation with “Couldn't start” notices. **Expected behavior** A known execution hold creates a waiting diagnostic without a run. Repeated automatic observations share that diagnostic. Clearing the hold permits a new wake only after the other gates pass. New comments remain available for the next eligible execution. **Steps to reproduce** 1. Assign a blocked task with a completed blocker. 2. Give the task an active reconciliation action, or a resolved action whose automatic recovery evidence still prohibits replay. 3. Run dependency reconciliation repeatedly. 4. Observe repeated cancelled pre-start runs on the base branch. This branch creates no runs while held and admits work after the effective hold clears. ## What Changed - Check effective execution holds under the issue admission lock before inserting runs. Keep the final dispatch check for races. - Share automatic wait diagnostics across producers, wake keys, and service restarts. Apply the helper to reconciliation, dependencies, pause holds, availability, budgets, and disabled heartbeats. - Preserve ordinary comment and interaction receipts during execution holds. Prevent release from draining them while replay is blocked. Keep external-chat receipt authorization intact. - Group empty pre-start reconciliation cancellations into a neutral waiting notice. Keep started runs and the full run history. - Document the waiting contract and add database-backed, UI, and browser regressions. - Stabilize two existing verification tests: allow the asynchronous chat lease transition a bounded five-second wait, and accept either legitimate damaged-session refusal while retaining exact archive-evidence assertions. ## Verification Passed targeted tests: - `pnpm exec vitest run server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts server/src/modules/wake-queue/adapters/postgres.test.ts` — 28 tests. - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx` — covered in the initial combined test run; UI suite passed. - Run-dispatch adapter tests passed in the combined gate regression run. - `pnpm exec vitest run server/src/__tests__/durable-chat-wakeup.test.ts` — 41 tests, including held receipt replay, promotion, and revoked access. - `PAPERCLIP_E2E_PORT=3294 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts` — all 3 browser scenarios pass. Repeated held messages create no additional runs or provider prompts and do not replay writes. - `pnpm check:token-gates` - `pnpm check:module-boundaries` - `git diff --check` `pnpm -r typecheck` and `pnpm build` pass. The full local `pnpm test:run` invocation did not finish green: it encountered exhausted local PostgreSQL shared-memory slots, a missing fresh-worktree runner test binary, and tests loaded across in-flight edits. The affected chat/database suites passed on rerun (81 tests), and the targeted lifecycle/recovery verification passed (3 tests). After building the runner test binary, the full native session suite also passed (37 tests). Final-head [CI run 34621288475](https://github.com/paperclipai/paperclip/actions/runs/34621288475) passed on `8659618b0ed2b98df002a28f4c1bd97321b0db04`, including all server/workspace test shards, all three browser shards, runner verification, typecheck, build, release dry run, and the aggregate verification gates. All 31 reported checks passed; the two conditional Storybook checks were skipped as intended. Greptile reviewed that exact commit at 5/5 with no unresolved review threads. ## Risks The wait record is diagnostic only. It must never count as a delivered wake or bypass a current gate. Tests cover repeated and concurrent admission, resolved no-replay evidence, a remaining dependency after hold clearance, deferred comments, and release gating. Explicit user requests and authorized chat receipts do not share automatic diagnostics. No migration or historical data deletion is required. Existing provider retry budgets remain unchanged. ## Model Used OpenAI GPT-6 through Codex. The runtime does not expose the exact hosted snapshot ID or context-window size. Used repository inspection, code editing, command execution, tests, and review tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
52811c6ce6 |
fix(tasks): require resume before sending to paused tasks (#13232)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task execution controls let board users pause a task or its subtree. > - The composer still accepted messages while a pause hold was active. > - A paused task must require an explicit resume before the user can send another message. > - This pull request replaces the composer with an amber pause card and checks board comment writes on the server. > - The user keeps their draft and resumes through the existing task controls. ## Linked Issues or Issue Description Refs #13104. Refs #13119. **What existing behavior does this improve?** The task composer and existing task/subtree pause controls. **Current behavior** A paused task can still receive a board message. The pause notice sits outside the composer, which leaves the send action available. **Proposed behavior** Show an amber takeover in both task chat and the classic composer. Preserve the draft. Require the user to resume the task or the ancestor subtree before sending. Reject board comment writes through either supported write route while the pause hold is active. **Breaking changes** Board comment writes to a paused task now return HTTP 409. Agent run reports remain supported during a pause. There is no schema migration. ## What Changed - Add a shared amber composer takeover with task, subtree, saved draft, pending, and error states. - Use effective ancestor pause state in both composer interfaces. Refresh it after pause events, task updates, and rejected sends. - Preserve draft text and attachments. Hide editor, send, queued edit, and pending question controls while paused. - Check active pause holds before board comment writes can mutate tasks, store comments, or wake agents. - Connect the approved Storybook examples to the production component and update the design and behavior docs. - Add browser coverage for both composers, draft persistence, resume, inherited holds, and rejected writes. Update ACP continuation coverage for the explicit resume requirement. ## Verification - Passed: `pnpm -r typecheck`. - Passed: `pnpm build`. - Passed: `pnpm build-storybook`. - Passed: `pnpm check:token-gates` and `git diff --check`. - Passed: focused UI tests (398 tests) and server route tests (127 tests). - Passed: `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/paused-composer.spec.ts tests/e2e/acp-stop-continuation.spec.ts` (5 tests). - Passed: manual browser walkthrough in a disposable local instance. Pause with a draft, refresh while paused, resume, send, and reopen. The draft returned, and one message persisted. The amber card and resume dialog were readable with no clipping. - Full local `pnpm test:run` did not pass: the general-server stage recorded 9,072 passing tests, 6 database setup failures from macOS shared-memory exhaustion, and 4 failed tests. This stopped the script before its later groups. Latest-head CI runs those groups independently. - Local follow-up: the Git file-resource load test passed on rerun (4 tests); native finalization migration passed after clearing the abandoned browser-test database allocation. Building the native debug fixtures fixed the missing fake provider. The remaining native-session recovery assertion also reproduces on untouched base commit `87b3e5fc6` (36 pass, 1 fail on both base and PR). It expects a settled-session error but receives a semantic-input-digest error. - The final UI build, UI typecheck, token gates, both thread suites (182 tests), and all five browser tests passed after the queued-action review fix. All 31 latest-head CI checks passed, including all server, workspace, browser, build, release, and security gates. Two optional Storybook jobs were skipped by workflow policy. Greptile reviewed `32d8fb5f5` at 5/5 with no open findings. - Review the Paused Composer and Tasks / Execution Controls stories. Pause a task with a draft, verify the amber card, resume, and verify the draft can be sent once. ## Risks - Clients that used board comments to continue paused work must resume first. The response is an explicit HTTP 409. - Pause state can change while a page is open. Live updates refresh the composer, and the server rejects stale sends before their side effects. - Resume keeps the existing dialog and optional agent wake behavior. Agent reports from interrupted runs remain allowed. ## Model Used OpenAI Codex, based on GPT-6, assisted with design, implementation, code execution, and browser verification. The exact runtime model ID and context window are not exposed in this session. The agent used reasoning and tool calls. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the relevant tests locally and they pass; the full local-suite limits are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3b03c4b9eb |
perf(ui): keep long task chat responsive during streaming (#13229)
## Thinking Path > - Paperclip lets operators manage agent work through tasks. > - Task chat keeps responses and run history together. > - A live update rendered every historical bubble and hidden tool row again. > - New image callbacks also forced unchanged markdown to parse again. > - Long conversations saturated the browser main thread. > - This change reuses unchanged history and mounts folded tools on first inspection. > - Operators can read and reply while work continues. ## Linked Issues or Issue Description **What happened?** Chat-style tasks with substantial scrollback became almost unusable. A deterministic browser reproduction with 200 long responses and 4,000 tools consumed 97.8% of the main thread during live updates. It delivered only 8 updates during the sample. **Expected behavior** The task should remain responsive during streaming. Historical markdown and unopened run details should not repeat expensive render work. **Steps to reproduce** 1. Install dependencies with `pnpm install`. 2. Run `pnpm exec playwright test --config tests/perf/task-chat/playwright.config.ts`. 3. Compare the attached performance JSON. The new test fails against the original rendering code. **Paperclip version or commit** Reproduced at `a05b828bc`. The branch is rebased on current master. **Deployment mode** Local Chromium and Vite with deterministic fixtures. No database or agent credentials are required. Related work: #10463 reduces the issue-page bundle. This change addresses repeated rendering after the page loads. No duplicate scrollback fix was found. ## What Changed - Keep the bubble image callback stable so unchanged markdown can skip parsing. - Memoize the settled history separately from the header and streaming tail. Keep the brief renderer and default attachment array stable. - Mount folded tool history on first expansion. Keep it mounted afterward to preserve child state and closing motion. Runtime request receipts remain visible. - Add deterministic rendering tests to the normal Vitest suite and an opt-in Chromium regression fixture. - Document the reproduction, commands, scope, and local measurements. ## Verification - Browser reproduction: 97.8% main-thread utilization before; 11.6% after for tail-only updates; 31.5% after when projection recreates history objects. Both fixed cases delivered 32 updates. - Browser checks pass for scroll-position retention, typing, return to latest, tool inspection, and retained expansion state. - Focused component suite: 155 tests passed. Post-rebase thread/performance rerun: 101 tests passed. - Recursive typecheck, build, Storybook build, and token gates passed. UI typecheck passed after the final edits. - Local UI/CLI lane: 5,910 tests passed; ten files hit worker-start timeouts, then all ten passed with two workers (20 tests). Shared/adapter lane: 3,128 tests passed. - Full local `pnpm test:run` was attempted and is **not green**: its server lane recorded 10,517 passes and 11 failures plus fixture/setup errors. Queue (31 tests), Cursor/Git-load (9 tests), and missing-binary failures cleared on isolated reruns / building the runner test binaries. Two native suites still cannot initialize embedded PostgreSQL on this host. - The remaining native-session recovery assertion was reproduced in a clean worktree at base `a20ecce40` (1 failed, 67 passed across the native/queue suites). It expects a settled-session error but receives a semantic-tool-input digest error. No server or runner files changed in this PR. - All substantive CI jobs have passed, including build, typecheck, all server/workspace test shards, all three e2e shards, and the canary dry run. The unchanged Slack ordering test exhausted its one-second wait on the first run; its shard passed on rerun. Final aggregate verification passed: **31 passing checks**, no failures or pending checks; two optional Storybook deployment/visual checks were skipped. Greptile is **5/5**, with no unresolved review threads. - The supplemental local serialized route run was stopped after CI passed all five serialized shards; it had reported no failures. - The browser fixture uses the real chat rendering components with a plain textarea. It does not test the full composer or server transport. Timing results are local samples; deterministic render-count tests provide normal CI coverage. ## Risks - Closed tool content becomes available to DOM search only after first expansion. Visible run summaries and runtime request receipts remain available immediately. - Opened run history remains mounted to preserve child state. The first full markdown render still scales with conversation size. - Memo dependencies must stay current when adding render inputs. Tests check content edits and replacement gallery callbacks. - No API, database, or permission changes. ## Model Used OpenAI GPT-6 (Codex). Exact serving variant and context-window size are not exposed in this session. Used reasoning, repository tools, code execution, and browser testing. No sub-agents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted/UI/workspace checks pass; the full local server suite has the baseline/host failures documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d10cbde815 |
fix(recovery): reject stale productive continuation wakes (#13173)
## Thinking Path > - Paperclip manages agent work through tasks and runs. > - Recovery continues assigned work when no live execution path remains. > - A recovery sweep can read an in-progress task before its run completes. > - The sweep can then observe the successful run after completion has changed the task status. > - This pull request checks current status and assignment under the existing enqueue lock. > - A stale continuation leaves a skipped wake receipt and creates no run. > - Task chat also omits an empty continuation cancelled before it started because its task had become terminal. ## Linked Issues or Issue Description Related public work: #10779 and #8419. Those older open changes address terminal disposition across other recovery paths. This change uses the existing scheduler guard for productive successful-run continuation and adds real database lock contention coverage. **What happened?** Recovery could combine an old in-progress task snapshot with a newer successful run. It queued an automatic continuation after the task was done. Dispatch cancelled that run before it started, but task chat displayed “Couldn't start” below the successful answer. This can happen after the native runner's finish result has already been accepted. It does not require a missing comment. **Expected behavior** Productive continuation must remain eligible when enqueueing acquires the task lock. Completion, cancellation, reassignment, or a move away from in-progress must prevent creation of the run. Actual execution stops must remain visible. **Steps to reproduce** 1. Let recovery select an assigned in-progress task whose latest run succeeded with productive progress. 2. Hold the task row lock in another transaction and change the task to done. 3. Let recovery attempt to enqueue while that transaction holds the lock. 4. Commit completion. Before this fix, recovery creates a redundant run from the stale snapshot. **Paperclip version or commit** Reproduced against master at `4042eb1c4` with deterministic integration tests. **Deployment mode** Built from source with PostgreSQL. The bug is in core recovery and is not adapter-specific. ## What Changed - Pass the existing status-and-assignee guard for productive terminal continuation recovery. - Preserve a skipped wake receipt with the expected and actual task state, without creating a run. - Test actual PostgreSQL lock contention for native and legacy completion, cancellation, backlog, review, blocked state, and reassignment. - Omit empty redundant pre-start cancellations from native and legacy task chat. Preserve stop markers for runs that started. - Document recovery eligibility at enqueue time. ## Verification - All seven new race cases failed before the guard was connected. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm check:token-gates` passed. - Task chat suite: 98 tests passed. - Recovery integration suites: 290 tests passed, including 31 stale-queue tests. - Full CI verification passed on `7ea71f04d`: all 31 active checks succeeded, including all test shards, browser tests, build, typecheck, and canary release dry run. Storybook visual regression was skipped by its path filter. - Greptile reviewed this commit at 5/5 with no review threads. - The local `pnpm test:run` aggregate reported a setup failure in the unchanged `tool-access-service.test.ts` suite. Its isolated rerun passed all 231 tests without edits. The duplicate aggregate was stopped after the complete CI matrix passed; it is not counted as a successful local full-suite run. ## Risks Low risk. The backend guard applies only to productive successful-run recovery. It requires the task to remain in-progress with the same agent. Other wake sources keep their current policy. The UI change only suppresses empty redundant cancellations; run records remain available. No schema change or migration is required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code edits, and local test execution. The exact deployment snapshot and context window are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
889947c238 |
feat: add experimental native chat connectors (#13038)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also ask agents for work in their existing chat tools. > - Each external conversation needs one task and a current authorized source. > - Retries, Stop, and provider failures must not duplicate work or expose private data. > - The first chat PR establishes the opt-in provider and data contracts. > - This PR adds experimental channel integration and its durable control plane. > - Users can request work from connected channels and inspect delivery in Paperclip. ## Linked Issues or Issue Description Refs #13100 and #13092. This is the second of exactly two chat PRs. Foundation #13100 is merged and changed 143 files. Runner prerequisite #13092 is also merged. This PR changes 400 files against master, below the 500-file review limit. It contains no wireframe images or HTML galleries. ## What Changed - Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat connections. Keep chat disabled unless the operator enables experimental chat connectors. Preserve the production GitHub tool connection and its normal setup path. - Bind each provider bot identity to one immutable Paperclip agent. Bind each admitted external conversation to one task. Paperclip owns tasks, runs, permissions, and audit records. - Add durable admission, per-conversation queues, questions, task controls, progress, final replies, images, files, and delivery receipts. Board comments remain internal unless explicitly sent to the channel. - Check current identity, provider reach, resource access, credentials, runtime generation, and exact source before provider effects. Keep private responses private. Never send raw reasoning, private logs, credentials, or tool arguments. - Hold uncertain sends for explicit audited resolution. Make Board Send-to-channel atomic and idempotent. Keep reconnect and setup credentials in Paperclip secret storage. - Preserve current native-runner authority across retries, lost acknowledgements, and recovery. Keep immutable input and completion contracts separate from newer user input. Receipt reconciliation cannot launch a provider. - Reconcile chat close/new ordering and provider-effect lock order. Audit resource access changes in the same transaction. Submit only the selected resource from each UI toggle so stale pages cannot undo unrelated access changes. - Drain Codex stdout before certifying process exit. Bound the drain with the existing shutdown grace. Preserve observed terminal authority without treating an undrained process as successful or reusable. - Incorporate master `018ca5da` with its ACP Stop, mobile task layout, runner packaging, and official lock changes. Preserve dedicated chat-answer continuations in both directions when ordinary queued comments are adopted after Stop. - Fence late adapter readiness behind an earlier Stop for the same run. Preserve verified cleanup for registered adapters. Handle single Stop, agent pause, duplicate Stops, and failure release without creating a false cancellation receipt. - Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact failed-chat retry authorization and lineage, retired question-source suppression, and the block on generic recovery that would discard the admitted source. Fresh deferred input retains its separate promotion path. - Incorporate master `2a05b5ed3` and its queue-admission extraction, simplified transaction ports, and separate runner CI job. Preserve exact durable receipts, actor separation, and dedicated-answer isolation through the new module. A failed receipt insert rolls back the accompanying deferred-wake merge. ## Verification Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are resolved. This successor fixes two test-harness boundaries exposed by CI: per-case route-module preparation and actual durable-save completion before intentional runner termination. Production code and all existing test/turn deadlines are unchanged. [Exact-head Greptile review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594) is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable findings or open review threads. [Fresh exact-head CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341) passes **all 24 jobs**, including Build and both required aggregates. Normal exact-head guarded merge was attempted and rejected by the remaining branch approval policy: CODEOWNER review is required and no human approval is present. Normal **squash auto-merge is enabled** as of September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified; no approval bypass or self-approval was used. Earlier-head results below remain historical evidence, not qualification of this successor. - Final exact-head Linux evidence: 995/995 chat integration cases; 36/36 agent-skills routes; 35/35 runner live-session cases, including real process kill/resume; 1948 runner Vitest cases with three existing benchmark/platform guards; 870/870 API-authority cases; and 104 browser cases with four existing optional skips. Rust, conformance/replay, full repository build, typecheck, canary, all server/workspace shards, and both required aggregates pass with normal CI concurrency. Earlier failed attempts remain recorded below. - Latest test-only qualification: 141/141 route/permissions/authentication cases pass in separate cold forks, with plain server types and independent review clear. The real-runner suite passes 35/35, with plain runner types and independent review clear. A controlled premature-save acknowledgement fails as expected; matching ownership/effect/process evidence, rejected saves, real turn outcome, test abort, and pre-kill liveness are covered. No local reproduction of the original CI scheduling failure is claimed. The preceding [CI run](https://github.com/paperclipai/paperclip/actions/runs/34479680858) passes 21/24 jobs, including all 995 Linux chat cases and browser aggregate (104 passed, four existing optional skips); only Build, the skills serialized shard, and the required verification aggregate fail. Its exact-head Greptile review was 5/5. Both failed job logs are retained. - Final fixture qualification: all eight focused Discord cases and all 995 chat integration cases pass. The exact modal statement/PID is observed before taking the real connection lock; the test then proves its actual blocking relationship before mutation. Original SQL execution, provider behavior, negative assertions, and 1s/15s timeouts remain unchanged. Independent review is clear and test/production hashes remain frozen. The preceding [CI attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777) passed 22 jobs, including Build/runner, typecheck, canary, all other test shards, and browser aggregate (104 passed, four existing optional skips); the two fixture failures and failed verification aggregate remain recorded, not relabeled as a pass. - Current queue-module composition: 308/308 recovery/batching/queue/Stop tests; 995/995 full chat integration; 89/89 module tests, including real PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary tests; plain server and UI types. All four actual local process/ACP browser paths pass in 1.4 minutes. Fresh databases, no skips or retries, stable reviewed source hashes. The initial boundary failure is retained; its no-op service wrapper was removed without changing recovery context or weakening the check. An exploratory standalone test-directory typecheck fails because its new upstream transformation config is not a standalone typechecking project; standard CI/build does not invoke it, and no configuration was weakened to suppress those diagnostics. - The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed [all 24 CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958) and exact-head Greptile review at 5/5. Required CODEOWNER review prevented its normal merge before master advanced again. - Final extracted-module composition: 307/307 recovery, batching, queue and Stop-control tests; 995/995 full chat integration; 49/49 module tests including eight PostgreSQL adapter cases; and 19/19 issue-update tests. Plain server types pass. All four actual local process/ACP browser paths pass in 1.3 minutes. Fresh databases, no skips or retries in these cohorts, frozen source hashes, and independent review clear. - The preceding head `3e4e1c1c` passes [all PR CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820), including Build and required `ci / verify` and `ci / e2e`. Both the original Rust failure and the previously load-sensitive lineage fixture pass with unchanged Linux concurrency. Master advanced afterward and required this reconciliation. - Final master composition: 448/448 focused UI tests, 186/186 adapter tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI, server, shared, and adapter types pass. Token gates and diff checks pass. Independent server and UI reviews are clear. - Stop-registration regression: both real-service cases fail against exact `a95` source and pass with the fix. The full corrected recovery/control suite passes 265/265. Duplicate-owner and failed-Stop controls also pass. Plain server types pass. The readiness barrier prevents provider startup without adding an acknowledgment to an already terminal run. - Final qualification strengthens terminal-field equality and repeats both affected cases successfully on a fresh database. All four actual local process/ACP browser paths pass again in 1.3 minutes, without skips or retries. The final screenshot shows Cancelled, a paused subtree, retained input, and no error toast. - Two new actual-service regressions fail before the merge fix. They prove that queued-comment adoption could consume a dedicated chat answer or add unrelated input to that answer. The fixed four-case cohort passes, including ordinary upstream continuation and adapter Stop controls. Full recovery passes 257/257. All four actual local process/ACP Stop browser flows pass in 1.4 minutes, without skips or retries, on a fresh database. - The unchanged runner artifact was qualified with 171/171 transport tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11. Six controlled reader tests prove the exit/drain repair. Its local serial Rust workspace passed 546 top-level cases plus two invoked helpers; the later passing Linux CI supplies default-concurrency evidence. - Prior exact-source full chat integration passes 995/995. Settings regressions cover concurrent stale pages, 501 destinations, pending state, rejected updates, and explicit retry. These deterministic tests do not prove live provider behavior. - Retained failed attempts and their causes are in the [qualification log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md). The first merge adapter run timed out while macOS slept for 290 seconds. Its unchanged repeat passed with a temporary sleep guard. No assertion, deadline, or CI gate was weakened. Review commands include `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh disposable databases. See the [browser runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md) for provider setup and separate live acceptance steps. ## Risks - This remains experimental. Deterministic tests and bounded live evidence do not establish every provider feature, tenant, permission layout, or media shape. Teams work-tenant qualification is still open. - Failed and uncertain provider effects remain visible and can require operator action. A transport receipt does not prove recipient visibility. - Native controller and runner artifacts must remain compatible. Preserve lease ownership, terminal authority, source binding, and quarantine during future changes. - Access and audit rows commit together, but activity notifications remain best-effort. This is not a new durable event outbox. - The PR operation does not deploy a live server, replace its runner, or change provider permissions. Remaining live qualification is documented in the [temporary handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md). ## Model Used OpenAI Codex assisted with implementation, tool execution, testing, and review. The work records `gpt-6-astra` assistance. The environment does not report a context-window size. No private reasoning traces are included. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018ca5daaf |
fix: verify ACP Stop and preserve safe continuation (#13119)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task controls coordinate provider execution and queued user messages. > - Stop could finish before an embedded ACP provider stopped its tools. > - A later request could be held for reconciliation without a clear task response. > - A restored provider could also retain the stopped run's API credential. > - This pull request verifies provider termination and preserves safe session continuation. > - Operators can continue known-safe work and see why uncertain work cannot start. ## Linked Issues or Issue Description **What happened?** Stop could leave an embedded ACP provider running. A queued follow-up followed by “go” could fail before it reached the provider. Task chat could show a generic missing-response message. Even a restored session could use the previous run's credential and fail its task update. **Expected behavior** Stop waits for confirmed provider termination. A later explicit wake continues the same compatible session only when recorded actions have known outcomes. It carries pending comments and the current run's environment. Uncertain actions retain a visible reconciliation hold. Composer Stop preserves the existing pause rule: conversation can continue while paused, but task work requires Resume. **Steps to reproduce** 1. Start an embedded ACP task. 2. Send a second request while the provider is running. 3. Interrupt the run, then send “go”. Also test composer Stop followed by Resume work. 4. Check that the request is delivered once and that the provider can complete the task through the current run's API credential. 5. Repeat with an unfinished write. Confirm that the write stops and that further execution stays blocked with a visible reason. **Paperclip version or commit** Built from source on master at `3bc60dd8b` plus this branch. **Deployment mode** Local source build with an isolated embedded PostgreSQL instance. Refs #11183. Refs #12552. Those changes address recovery after operator cancellation. This change also covers embedded ACP termination, session proof, pending-comment delivery, and task feedback. ## What Changed - Propagate Stop into embedded ACP and wait for bounded adapter cleanup and provider exit. Retain the actual ChildProcess object for forced termination on all platforms; never signal a recycled numeric PID. - Preserve interrupted checkpoints only for acknowledged, local, persistent sessions with settled reads or no tools. Keep writes, incomplete actions, and forced termination blocked. - Restore the same compatible provider session with the current run's environment. Reject fresh-session fallback for an interrupted checkpoint. - Adopt pending comments on the next explicit wake. Stop alone does not dispatch them. - Share the execution-blocker rule across dispatch, Resume, and task detail. Show Stopped or Couldn't start with the recorded reason. Resolve the stopped agent for the run link, including reviewer runs. - Keep execution reconciliation holds intact when generic recovery sees queued comments or healthy child tasks. - Add process, service, component, and browser regression coverage. Fix disposable database cleanup and React test settling exposed by the full suite. ## Verification - Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. - Passed all three `acp-stop-continuation.spec.ts` browser journeys. They use an actual ACP child process and require task completion through the agent API. - Passed 165 adapter execution, operator-stop, and child-process control tests, 17 queued-comment route tests, and 65 tests in the two adjusted UI suites. Earlier focused recovery, heartbeat, and task-control tests also passed. - Manually used the browser to queue a request, Stop, send “go” while paused, and Resume. The same session answered once and moved the task to Done with the current run's credential. - Manually interrupted an unfinished write. Its file size stayed fixed for five seconds. “Go” showed the reconciliation reason and did not start another provider prompt. - Separate live Claude ACP smoke checks confirmed that Stop ended a disposable local write and that a no-tool interruption could resume the exact provider session. The browser fixture does not call Drive or another external app. - Passed all 5,615 UI tests and 3,090 other workspace tests. The CLI and general server groups pass with targeted retries: two transient server failures passed together on retry, and two embedded-database startup failures passed after removing abandoned shared-memory segments from this task's completed browser fixtures. All 144 serialized server suites completed, with 2,189 tests passing after two transient HTTP socket failures passed on retry. - Passed all 135 heartbeat process/recovery tests, including a deterministic regression that failed before the recovery-sweep fix. - Passed 18 dispatch integration tests, including stopped-reviewer links, company boundaries, and malformed run IDs. - Greptile is 5/5 on `7dd170d83`, with zero unresolved review threads. The security scan and all required CI gates pass for the same commit. ## Risks - Safe continuation depends on complete tool reporting and a restorable local provider session. Unknown outcomes remain blocked and require reconciliation. - Provider cleanup can take time. A timeout does not grant replay permission. - The change adds optional adapter context fields and an optional issue projection. It does not change the database schema or require a migration. - Test cleanup truncates company data only in a disposable test database. ## Model Used OpenAI GPT-6, running as Codex with repository tools, code execution, and browser interaction. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2d45f42e47 | fix(ui): refine mobile task surfaces (#13122) | ||
|
|
8cfd30fb07 |
feat(ui): add composer Stop and simplify task controls (#13104)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task composer is where operators direct running agents. > - Operators need to stop work without leaving the conversation. > - Existing pause controls already hold task trees and interrupt both runner types. > - This pull request connects the composer to those controls and removes repeated feedback. > - Operators can pause work quickly and still queue messages while agents run. ## Linked Issues or Issue Description **What existing behavior does this improve?** Task pause, resume, and cancellation in the task page and composer. **Current behavior** The empty composer cannot stop a running task. Task controls require extra confirmation and reason text. Pause can show several notifications for the task already on screen. **Proposed behavior** Show Stop while this task runs and the composer is empty. Text or attachments switch it to Send. Stop and the menu use the same manual pause hold. Parent pauses include descendants. Keep task cancellation in the menu with a compact confirmation. Show one quiet pause row and gray cancelled-run details. **Reason and benefit** Operators can interrupt execution with one click. Drafts and queued messages keep their existing behavior. The UI waits for actual termination, including native cancellation acknowledgment. **Breaking changes** No endpoint, schema, or task-status change. Pause no longer asks for confirmation or a reason. Resume now honors the existing wake-agents option. Task notifications are suppressed for the task and subtree currently in view. Related UI work: #8228 changes navigation and composer shortcuts. This PR covers execution controls. No duplicate Stop-button PR was found. The change improves existing controls and does not duplicate a roadmap milestone. ## What Changed - Add Stop, pending feedback, duplicate-click protection, and inline errors to the composer. - Share the pause mutation across the composer, active-run controls, and menu. - Poll affected runs after a pause request. Require native cancellation acknowledgment. - Remove pause confirmation and shared reason fields. Reduce cancel confirmation to its task count and actions. - Honor wake-agents for executable tasks only. Preserve the pause when recovery review is needed; show partial wake failures inline. - Preserve explicit legacy reconciliation decisions while their continuation waits for dispatch. - Suppress notifications for visible task trees. Use quiet pause and cancellation feedback. - Add interactive stories using production controls and native/legacy end-to-end tests. ## Verification - User reviewed the running feature and revised Storybooks in the browser. - Rebased focused checks passed: 295 original targeted tests, 161 updated route/page/notification/status tests, and 26 recovery integration tests. - Both isolated runner journeys pass on the final revision (1.7 minutes). Coverage includes queueing, parent and child interruption, persisted holds, no automatic continuation, reconciled resume, cancellation, terminal exclusions, and no Stop toast. - Native coverage uses real runnerd with a deterministic provider fixture. Legacy coverage checks actual process termination. Live hosted-provider execution was not tested. - Repository typecheck and build, Storybook build, and token gates passed after rebase. The final server typecheck/build also passed. - The broad local run completed its general-server stage with 7,219 passing tests, 48 skipped, and two failures from cached pre-fix source and a stale native provider fixture. Both failed tests pass in fresh final-head reruns after rebuilding the fixture; the script did not continue to its later local stages. CI runs all test groups on the final revision. - Final revision: all 31 applicable CI checks passed; Storybook visual regression was skipped by its workflow conditions. Greptile: 5/5, zero unresolved comments. - Review `Tasks / Execution Controls` in Storybook. Type and clear a draft, stop a run, expand cancellation details, and test the menu on desktop and mobile. ## Risks - Stop pauses descendants for a parent task. This is the existing pause contract. - A held task can remain active if interruption fails. The UI shows an error instead of claiming termination. - Resume can start multiple assignees when wake-agents is selected. Backlog, blocked, and terminal tasks stay excluded. Existing execution reconciliation remains mandatory where required; Resume never invents action-outcome evidence. - Notification suppression uses the visible task and cached subtree. Notifications for unrelated work remain enabled. ## Model Used OpenAI GPT-6 through Codex. The exact runtime snapshot and context-window limit are not exposed in this session. Used reasoning, tool calls, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cd4c4ed205 |
fix(ui): stabilize task loading and live feeds (#13095)
Coordinate initial conversation reveal, preserve message identity and reading anchors during live updates, and bound transcript reads with recoverable retries. Cover desktop/mobile navigation and rich task loading with actual-route browser tests. Verified all Linux CI gates, 5,575 local UI tests, eight layout browser scenarios, and recorded native Codex walkthroughs. Greptile: 5/5; all review findings resolved. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
35fdc0c66b |
fix: make task recovery durable and preserve current requests (#13075)
Make task recovery durable and preserve the latest user request across native and legacy continuations. Keep routine recovery quiet and prevent replay when action outcomes are uncertain. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e200104727 |
feat: review connection actions from tasks (#13063)
Bring governed connection reviews into task history and composer approvals. Share resolution with Connections, add scoped remembered permissions, and resume agents through durable outcome receipts. Keep cards compact, collapse raw results, isolate untrusted provider output, bound continuation payloads, and reconcile missed live events. Add Storybook coverage, browser journeys, and service regression tests. Verification: all PR CI gates passed, Greptile 5/5, security scans passed, five connection-review browser journeys passed, and real native Codex approval/continuation was verified against the local MCP fixture. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2043e0c735 |
fix: repair runner configuration, macOS execution, and artifact galleries (#13062)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters select a provider, a model, and a runtime. > - Runner conversion rejected existing Claude agents. The model list mixed providers. > - The native Claude runner rejected custom models and could not launch on macOS. > - This pull request fixes conversion, model selection, and verified macOS execution. > - It also groups configuration fields consistently across adapters and opens artifact images in the task gallery. > - Operators can change an agent configuration and run the selected model on their Mac. ## Linked Issues or Issue Description **What happened?** Converting an existing Claude agent to Paperclip Runner failed with a Codex-only restriction. ACPX Claude showed unrelated models and required `claude-sonnet-5`. Its native runtime rejected macOS. Configuration mixed common model settings with process controls. Artifact cards labeled “Open gallery” navigated to attachment URLs instead of opening the task gallery. **Expected behavior** Conversion keeps agent identity and compatible settings. ACPX Claude uses the normal Claude catalog and accepts typed model IDs. Codex uses the native runner. The verified Claude runtime can launch on macOS ARM64 and x64. Common configuration sections place the same fields together across adapters. Artifact images open in the shared task gallery with navigation and downloads. **Steps to reproduce** 1. Open the configuration of an existing Claude agent. 2. Convert it to Paperclip Runner. 3. Select ACPX Claude and a different catalog model or a typed model ID. 4. Save the agent and run a disposable task on macOS. 5. Inspect configuration and advanced run-policy controls across adapters. **Paperclip version or commit** The bugs were reproduced on `165ca56a22adb60e5fda56045442d9c8498116a8`. This branch was rebased onto `7ed122911`. **Deployment mode** Built from source. Local test-drive instance on macOS ARM64 with an isolated database. Related work: #11798 addresses unsupported ACP session options in the existing adapter path. #13048 addresses working-folder preservation. This change fixes native runner configuration and launch behavior. ## What Changed - Remove the Codex-only conversion restriction. Preserve agent identity, instructions, directories, credentials, and compatible model settings. Reset incompatible sessions while retaining history. - Show ACPX Claude and native Codex as distinct provider choices. Remove ACPX Codex from advertised configuration. Normalize legacy configurations before fresh runs without rewriting historical run descriptors. - Select model catalogs and cache entries by provider. Support refresh and typed model IDs. Pass exact Claude IDs through session creation, model changes, and recovery. - Add verified macOS ARM64 and x64 Claude SDK snapshots. Bound executable allocation and total snapshot size. Preserve package checks, dependency isolation, process ownership, cancellation, and Linux descriptor loading. - Probe local runtime readiness. Report remote platform checks as incomplete until the remote runner verifies its runtime. - Surface actual model rejection and allow correction and retry. - Repair missing ACPX goal-capability helpers exposed by the post-rebase live test. Persist and restore the optional capability without breaking session startup. - Put Agent identity first and intentionally remove the Capabilities editor, as requested. This is removal of UI editing, not relocation: preserve existing capability metadata and API compatibility without adding another editor. Use the themed select for configurable permission modes, with normal text instead of monospace. - Put model and provider under Adapter. Give environment variables their own section. Fold command and arguments under Configuration. Fold lifecycle, timeout, and interrupt grace under Advanced Run Policy. Hide single-option permission controls. - Open image and video artifact cards in the existing task gallery, including cards in the artifacts panel. Chat attachment images use the same gallery. Preserve standalone media previews and download links. ## Verification - Rebased focused UI/API/database suites: 293 tests passed. - Rebased native runtime and ACPX suites: 242 passed, 7 skipped. - Repository typecheck, build, and token gates passed for the runner changes. Gallery follow-up UI typecheck, build, and token gates also passed. - Follow-up UI suites passed (86 tests), packaging checks passed (14 tests), and the final focused runtime suites passed (126 passed, 7 skipped). - Linux container isolation and lifecycle fixtures passed before rebase (57 passed, 2 skipped). Rust ACPX provider-session tests passed after rebase (8 tests). - Browser tests completed actual Claude and native Codex tasks on macOS ARM64. They covered conversion, catalog refresh, a non-default catalog model, a typed `haiku` ID, save/reload, cancel, follow-up session continuity, invalid-model errors, and recovery. - Final-revision live tests completed a typed Claude task, a follow-up with the same provider session, and a native Codex task on macOS ARM64. - Browser tests confirmed the moved interrupt-grace field saves and survives reload. Cross-adapter tests cover Claude, Codex, Gemini, process, gateway, and schema forms. - Full local run: 7,080 passed, 30 skipped, and two timeouts. Both timeout suites passed on isolated rerun (84 tests); the failures were the plugin login-worker exit diagnostic and the runner real-server vertical slice. - Final follow-up checks: 50 registry tests and 45 snapshot/installation tests passed (6 platform-specific skips). Oversized executable rejection is covered before allocation or reading; unsupported-platform tests invoke the real installation probe. - Runner head `ddb5101c483a297f74875ab96b3c66035b002d50`: all CI gates green, including full runner verification, repository build, typecheck, general/serialized server suites, browser tests, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34286178670). - Greptile: 5/5 on that runner head. All four review threads resolved. Superagent, Socket, and Snyk checks green. - After snapshot hardening, another real Claude task completed on this Mac using the rebuilt runtime. - Gallery follow-up: 148 focused tests passed, covering artifact selection, shared attachment collections, deduplication, image/video cards, standalone previews, downloads, and closing. Live browser verification completed on the settings follow-up: artifact selection, 6-image pagination with wrapping, download action, and closing all stayed on the same task URL. All checks passed on gallery head `96136da58ff195bf6ca00b281eb3022ad12d7bd8`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34287987536). Greptile returned 5/5 on that exact head with no unresolved threads. - Final settings polish: 96 focused tests, UI typecheck/build, and token gates passed. A real browser walkthrough verified readable permission options, identity placement, Capabilities removal, and permission save/reload. Original test-agent permission mode restored. All 31 checks passed on final head `e46540d6bf32bfb0566dca16b2f4a75ba437618c`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34292797886). Greptile returned 5/5 with no unresolved threads. ## Risks - Capabilities intentionally has no editable UI field after this change. Existing values remain readable and API-compatible; removing the field does not erase stored metadata. - macOS launch now copies verified package files into private snapshots. The implementation must retain isolation and clean up snapshots on exit. - Runtime provider or model changes reset the current session. Historical runs remain available. - The macOS x64 SDK executable digest was verified, but a live Intel Mac run was not available. Linux verification used container fixtures, not a real Claude task. - Remote environment tests report a warning when only the platform has been checked. They do not claim package readiness from the server host. ## Model Used OpenAI Codex, based on GPT-6. The exact served model identifier and context-window limit are not exposed in this session. Used reasoning, repository inspection, code execution, Rust and TypeScript tests, and browser automation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites and both timeout suites on rerun; full-run counts above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7ed122911b |
Add end-to-end session goals to Paperclip Runner
Add capability-aware slash-goal controls, durable provider goal state, PRP v2 negotiation, autonomous goal execution, and safe local session recovery. Integrate with current master, preserve provider session identity, and verify the browser goal/chat/replacement/clear workflow and unsupported-agent rejection. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e095b84dab |
feat(connections): connect services from native task feeds (#13058)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents use connections to reach external services. > - A fresh native task can have no service tools installed. > - The agent needs a way to discover services and ask the responsible person for access. > - This pull request brings the existing connection-intent flow into native task execution. > - The person can connect from the task, and the agent can continue with updated tools. ## Linked Issues or Issue Description **Subsystem affected** Native runner tool authority, connection intents, task interactions, and shared connection setup. **Problem or motivation** A task that needs an unconnected service cannot finish its work. Leaving the task to configure access also loses context. A resolved request must survive a restart and resume the correct agent once. **Proposed solution** Expose connection discovery and access requests as server-owned native tools. Render a durable task card and use the shared setup dialog. Persist outcome delivery and start a fresh provider session after access is ready. **Alternatives considered** Sending the person to the Connections page adds navigation and does not solve continuation. Polling for authorization consumes runs and can create duplicate requests. **Roadmap alignment** This extends the existing connection-intent runtime and setup experience. It reuses the shared access model and the native runner. Related: #12345, #12347. The service-slug fix in #12906 is related but separate. Companion evaluation PR: https://github.com/paperclipai/paperclip-evals/pull/21. ## What Changed - Expose `connections_search` and `connection_request` with server-bound company, task, agent, and responsible user. Preserve the legacy entry points. - Discover catalog services and authorized custom connections. Check installation, identity, health, and executable permissions before reporting ready. - Keep pending cards through ordinary messages. Reuse requests and retire stale ownership. Put Connect at the right of Not now. - Reuse the shared setup flow in a task dialog. Keep access additive and default to the requesting agent. Recover from cancelled or blocked OAuth windows with a new-tab fallback. - Persist outcome delivery with an idempotent wake key. Resume in a fresh session and recheck ownership before dispatch. - Add native browser fixtures, offline Storybook states, server contracts, and evaluation fixtures. Update guidance and documentation. ## Verification - `pnpm build`: passed after replaying the change on current master. - `pnpm -r typecheck`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed. - New continuation-policy regression cases: 16 passed. - Docker-backed PostgreSQL regressions passed for requester-only OAuth access, assignment-only expiry, terminal expiry, and credential-free setup metadata. - Shared setup and task-card UI tests: 121 passed, including configured MCP reconnect URL recovery and preserving user edits across refetch. - Storybook browser checks: all 119 passed on the latest reconnect fix. - `pnpm test:run`: 4,734 tests passed in the first server group, but embedded PostgreSQL startup failures and resulting cleanup errors prevented a complete local pass. All Linux CI lanes passed on the latest reviewed commit. One external-object route test returned an unexplained 500 on the first run; it passed twice locally and the failed shard passed on retry without code changes. - Earlier feature-checkout evidence: three deterministic native browser journeys passed, including restart delivery and an actual fixture tool result. Legacy scripted coverage also passed. All 59 added stories were inspected in light and dark themes. - Live Notion testing recorded successful provider reads. The manual test used a local-trusted instance. It does not prove authenticated/cloud deployment or every provider journey. - Native browser rerun reached the embedded PostgreSQL startup limit before bootstrap, so the latest checkout’s full native browser journey remains unverified. Both OAuth page/task regression cases passed against isolated Docker-backed PostgreSQL 17. They verify no premature task access, requester-only completion, additive retries, and reconnect preservation. - Applied both new migrations twice to isolated PostgreSQL 17. Foreign keys remained intact, duplicate active delivery keys were rejected, and failed delivery records did not block retries. Reviewer path: start a fresh test drive, enable the native runner, use an agent that can perform work directly, and ask it to summarize a Notion page. Connect from the card, then verify the resumed provider call and source-linked answer. The default test-drive CEO is instructed to delegate, so it can introduce an unrelated hiring step. ## Risks - Two additive migrations create durable deliveries and a partial unique wake index. They are idempotent. The wake index can require a maintenance window on large tables because migrations run in a transaction. - OAuth and continuation cross asynchronous boundaries. Tests cover ownership changes, retries, additive access, and restart delivery; live provider behavior still varies. - The latest requester-scope fix has not yet been exercised through live OAuth. GitHub, API-key, authenticated-user, and all recovery journeys are not claimed as verified. ## Model Used OpenAI GPT-6-based Codex assisted with implementation, tests, and review using tools and code execution. The runtime does not expose the exact model version, context window, or reasoning setting. Live evaluation used `gpt-5.6-luna`; manual native testing used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used and disclosed unavailable runtime details - [x] I have checked ROADMAP.md and confirmed this extends existing connection work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run all required tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5b56d430e9 |
feat(ui): refine core navigation and task detail (#12854)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The main navigation and task detail view are core operator surfaces. > - Several controls used different hover states, popover layouts, and spacing rules. > - Recent task actions also needed a compact menu and correct inbox archive behavior. > - These differences made the interface feel inconsistent and caused some content to look crowded or clipped. > - This pull request aligns these surfaces with the Paperclip design tokens and current interaction patterns. > - The benefit is a simpler and more consistent operator experience in light and dark modes. ## Linked Issues or Issue Description **What happened?** Profile and organization popovers used inconsistent layouts. Navigation controls used different hover and selected backgrounds. Task warnings and the composer could crowd nearby content. Archiving a recent task could also remove it from more than the inbox. **Expected behavior** Popover menus should use the same compact visual language. Navigation controls should share readable hover and selected tokens. Task detail content should keep consistent spacing. Archiving should hide a task from the inbox while keeping it in the task list. **Steps to reproduce** 1. Open the main sidebar in light or dark mode. 2. Open the profile and organization menus. 3. Hover navigation items, the organization trigger, the profile trigger, and the feedback flag. 4. Open a task with a warning banner and a long thread. 5. Use the recent task overflow menu and archive a task. **Paperclip version or commit** Reproduced on `master` before this branch. **Deployment mode** Local dev (`pnpm dev`). ## What Changed - Rebuilt the profile and organization popovers with compact token-based layouts. - Matched organization popover width and alignment to the profile popover. - Unified sidebar hover and selected states in light and dark modes. - Added a recent task overflow menu with rename, archive, and pause or restart actions. - Kept archived tasks in the task list while removing them from the inbox. - Improved warning banner and composer spacing in task detail views. - Added and updated focused UI tests for the changed behavior. ## Verification - `pnpm check:token-gates` passed. - `pnpm --filter @paperclipai/ui typecheck` passed. - The seven affected UI test files passed with 216 tests. - `pnpm --filter @paperclipai/ui build` passed. - GitHub CI passed the full build, typecheck and release registry, general test, serialized server, canary dry-run, and end-to-end matrices. - Greptile reviewed commit `6e296be85` at 5/5 with no outstanding actionable findings. ## Risks - Risk is limited to sidebar presentation, recent task actions, and task detail layout. - The recent task archive action now follows inbox-only archive semantics. - No database schema or public API contract changed. > I checked [`ROADMAP.md`](ROADMAP.md). This pull request does not duplicate planned core work. ## Model Used - OpenAI Codex, GPT-5.6. The model used high reasoning, tool use, and code execution. The context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Scott Tong <scott@scottsmbpm5max.lan> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bf95a7eae2 |
fix(ui): stabilize active-run steering queue (#12834)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue detail page shows a live agent run and accepts follow-up instructions. > - A follow-up must stay in a stable queue until the user sends, reorders, or removes it. > - Native runners can receive a steering event in the active run. > - Legacy runners must interrupt the active run and start a follow-up run. > - The current UI moved comments between the queue and the transcript and could show duplicate text or ambiguous chronology. > - This pull request makes the queue projection durable, keeps each message in one clear place, and labels when queued input was actually steered or delivered. > - The benefit is predictable steering with stable ordering, no duplicate messages, and visible causal timing. ## Linked Issues or Issue Description Refs #11374. Refs #12591. **What happened?** During an active run, a new follow-up could first appear as a transcript bubble and then move into the steering queue. After a steer or remove action, it could appear again. Progress text could also repeat the final response text. Once consumed, a queued bubble displayed only its original submission time even though it moved to its later causal slot, and a native run split by steering looked like two unrelated runs. **Expected behavior** An active-run follow-up must appear in the queue immediately. A native steer must move it once into the active run. A legacy interrupt must move it once into the follow-up run. A removed item must stay removed. Progress text that is identical to the final response must appear once. Consumed follow-ups must show both queue and steer/delivery times, and post-steer native segments must identify themselves as continuations of the same run. **Steps to reproduce** 1. Start a long-running task. 2. Send two or more follow-up messages while the agent is active. 3. Reorder the messages and remove one message. 4. Send the first queued message as steering. 5. Observe the queue and transcript during and after both runs. **Paperclip version or commit** The problem reproduced on commit `da1e40302`. **Deployment mode** Local development with the embedded database. ## What Changed - Project queued comments into the steering well for native and legacy live runners. - Send native steering to the active run and use interrupt-and-follow-up for legacy runners. - Keep optimistic queue order stable across refreshes and roll back failed actions. - Remove discarded comments from the transcript cache and keep them removed when the queue becomes empty. - Collapse only the final progress occurrence matching the durable response, including across steered transcript segments. - Show `Queued … · Steered …` for same-run input and `Queued … · Delivered …` for successor-run input at their causal positions. - Label settled and live post-steer segments `Continued after steering` and time them from the steer boundary. - Add regression tests for queue display, steering, fallback interrupt, reorder, remove, rollback, duplicate text, causal timestamps, and live/settled continuation headers. ## Verification - Ran the final focused steering/chronology UI suite with 233 passing tests. - Ran the activity-service regression suite with 5 passing tests. - Ran the broader queue-focused UI suite with 298 passing tests before the final chronology refinement. - Ran `pnpm -r typecheck` successfully. - Ran `pnpm build` successfully. - Ran `pnpm check:token-gates` successfully. - Tested native steering in a real browser with a 90-second baseline wait and a three-second steering correction. - Confirmed that the old final response did not appear before the steered response. - Tested three queued messages in a real browser. - Confirmed that reorder changed delivery order and that the removed message was never sent or shown again. - Tested a legacy runner in a real browser. - Confirmed that it used the interrupt fallback and showed the follow-up once. - Reloaded a saved mixed-steer/successor-run thread and confirmed the causal timestamps and continuation header render in the correct positions. - The complete macOS suite reaches five unrelated platform assertions in workspace-runtime tests. Two compare `/var` with `/private/var`. Three require Linux `/proc` listener data. GitHub Actions provides the authoritative Linux run. ## Risks - Low risk. The change is limited to issue-chat queue projection and transcript presentation. - The server run-history API adds only a read-only `contextIssueId` projection; the database schema does not change. - Optimistic actions restore the prior UI state when a request fails. ## Model Used - OpenAI Codex with GPT-5, extended reasoning, browser automation, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4d30efa8e3 | feat(ui): refine streamlined task experience (#12748) | ||
|
|
0f94521017 |
fix(runner): restore local session and task integrity (#12721)
## Thinking Path > - Paperclip is the control plane for agents that perform work. > - Paperclip Runner connects durable provider sessions to individual task runs through PRP. > - Provider continuity and per-run authority are different lifetimes. > - The existing implementation mixed those lifetimes and lost event metadata between provider frames, runnerd, persistence, API sanitization, and the task thread. > - That caused failed continuation, missing progress and Plans, duplicate replies, hidden failures, and unsafe recovery. > - This repair gives every heartbeat fresh authority, preserves qualified provider-session continuity, and restores one lossless presentation path without changing direct adapters. ## Linked Issues or Issue Description **What happened?** A second native heartbeat could reuse tickets, leases, command receipts, sequence state, and run identity from the first heartbeat. Provider phase and item identity could be lost before the UI read them. Redaction could corrupt protocol discriminators while still missing malformed credential tails. The task thread could fold progress into the final response, hide failures, or show more than one final answer. Native Codex also exposed approval modes that do not yet have a durable approval bridge. **Expected behavior** Each heartbeat uses a new PRP authority epoch. Codex and OpenCode preserve exact qualified provider sessions; ACPX emits an explicit continuity event when its qualified process-replacement policy is used. Every accepted provider event is presented, classified as internal, or surfaced as unsupported. The task page shows chronological progress, reasoning summaries, activity, Plans, interactions, terminal failures, and exactly one final reply. Direct adapters retain their existing path. **Steps to reproduce** 1. Enable the unified experimental Paperclip Runner setting. 2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex agent. 3. Run response, Plan, structured-question/resume, restart, cancellation, and failure scenarios. 4. Reload the task while active, waiting, failed, and settled. 5. On the old implementation, observe stale run authority, missing classifications, incomplete output, or duplicated/folded replies. **Paperclip version or commit** The repair is based directly on `master` at `87d05e194b643810d16d20612115acd01d735d43`. **Deployment mode** Local development with the embedded database. Related work: Refs #12616, #12646, #12666, #12685, and #12700. ## What Changed - Rotates PRP control-plane, outbox, ticket, lease, command, receipt, and sequence authority for each heartbeat while carrying forward only a validated provider-session identity. - Reads `control-plane-state.json`, validates both durable schemas and lifecycle values, resumes coherent current runs, archives qualified settled authority, and quarantines malformed or mismatched scoped state without moving ambiguous live legacy state. - Preserves Codex provider phase and stable item identities so commentary remains progress and only `final_answer` becomes final. - Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning lifecycle mapping. - Makes ACPX normalization lossless for visible reasoning, tool lifecycle metadata, stable bounded identities, Plan revisions, structured requests, failures, and qualified process replacement. Only the compatible terminal assistant message is promoted as final. - Applies schema-aware redaction before generic JWT-shaped detection and scans every diagnostic string leaf. Malformed raw/escaped quoted credential tails are redacted in both server and durable Rust state. - Restores snapshot-style chronological task presentation, expandable tool activity, inline Plan cards, visible waiting/resume/cancel/failure states, and exactly one final answer. - Makes `never` the only qualified native Codex permission mode and rejects unsupported persisted native modes with remediation. OpenCode and ACPX policies remain intact. - Keeps the unified experimental Runner setting as the only enablement flag. Onboarding and direct Codex, Claude, and OpenCode stay on their legacy execution/finalization paths. - Adds cross-language goldens, authority/recovery/fault coverage, exact response/count assertions, and native plus legacy acceptance scenarios. ## Verification - Pull-request GitHub Actions run Rust formatting/tests, TypeScript checks, server/UI tests, builds, protocol drift checks, browser E2E, and security scans. - A separate workflow-only validation ref is pinned directly on this PR head and runs the 35-cell paid local matrix: three core scenarios plus structured-question resume and restart/resume for native Codex, native OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode. Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315 - Acceptance requires exact single visible replies, monotonic sequences, matching envelope discriminators, one semantic terminal, one run terminal, no unresolved interaction, no duplicate mutation, no secret leakage, provider continuity, and zero native rows for direct adapters. - Per maintainer direction, tests are running in GitHub Actions rather than on the slower local host. Only formatters and static diff checks were run locally. ## Risks - Recovery from old or partial filesystem state is sensitive. The repair fails closed, preserves active or unverifiable authority, and quarantines only state whose scoped ownership is safe to move. - Provider event formats can change. Closed validators and boundary goldens turn new or malformed events into visible diagnostics instead of silent drops. - Shared task presentation could affect direct adapters. Runtime-fact gating plus the direct-adapter matrix protect the existing path. - Managed and remote providers are not qualified here. Shared code continues to compile and fail safely, but live qualification is deferred. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact deployed snapshot and context-window size are not exposed to this task. It used agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] The paid local-provider matrix is green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
87d05e194b |
feat(work-products): add rich cards and run artifact inventory (#12717)
## Thinking Path > - Paperclip is the control plane for AI-agent companies. > - Agent outputs must remain visible after a run and easy to inspect from a task. > - The thread and artifact inventory need one consistent rich-card vocabulary. > - Run uploads also need durable artifact registration and producing-run context. > - Reviewers need deterministic examples for each rich-card kind and state. > - This pull request adds the shared presentation, registration, inventory, and Storybook review coverage. > - The benefit is a complete output path that reviewers can inspect without seeded data. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change improves work-product presentation in task threads and the task Artifacts tab. **Subsystem affected** The change affects shared work-product contracts, the runner diff path, server attachment and work-product services, GitHub metadata refresh, the React board UI, and Storybook. **Current behavior** The thread used generic cards. Some files uploaded by a run existed only as message attachments. The Artifacts tab showed a flat list without run context or filters. Storybook showed only one resting card per kind. **Proposed behavior** The thread uses rich cards for supported work-product types. Each run-produced file registers one attachment-backed artifact work product. The Artifacts tab groups outputs by run and supports filters. Storybook shows every kind and requested state, PR lifecycle states, stats variants, truncation, mobile layout, and message-tail media. **Reason and benefit** Users can identify outputs quickly. Reviewers can inspect all card permutations without creating task data. **Breaking changes** None. The metadata fields and automatic artifact registration are additive. Existing attachments and work products keep their current behavior. ## What Changed - Added a shared rich work-product card with kind-specific content and a compact inventory variant. - Added pull-request and commit diff metadata plus bounded GitHub state refresh. - Added media strips and typed file chips to message-tail attachments. - Registered each run-produced attachment as an artifact work product in the same server transaction. - Grouped task artifacts by run with agent and timestamp headings. - Added type and run filters, image thumbnails, compact cards, and a company Artifacts link. - Added a Storybook kind-by-state matrix with stats variants for all eight visual kinds. - Added PR open, draft, merged, and closed examples, long-title truncation, an exact 375-pixel viewport, and message-tail overflow coverage. - Closed reconciled runtime work products when the linked runtime stops or disappears, so the card shows `Stopped` instead of `Unhealthy`. ### Screenshots Before: one resting card per kind.  After: the kind and state matrix.  After: message-tail media at 375 pixels.  [Open the Storybook evidence viewer](https://pages.paperclip.ing/rich-work-product-storybook-20260902/). The earlier artifact inventory comparison remains available in the [artifact inventory viewer](https://pages.paperclip.ing/rich-artifacts-inventory-proof-20260902/). ## Verification - `pnpm --filter @paperclipai/ui typecheck` passed. - `pnpm check:token-gates` passed. - `pnpm build-storybook` passed. - `pnpm exec vitest run server/src/__tests__/work-product-runtime-reconciliation.test.ts` passed with 5 tests. - Chromium visual checks passed at desktop and 375-pixel widths. - All 30 latest-head GitHub checks passed. One unrelated annotation test was flaky and passed on its single retry. - Greptile passed at 5/5 with zero unresolved threads. ## Risks - Low risk. The Storybook change adds review fixtures only. The runtime fix changes read-time reconciliation without database writes. - The matrix is intentionally large so every permutation stays visible in one review surface. > I checked `ROADMAP.md`. This work does not duplicate planned core work. ## Model Used - OpenAI Codex with GPT-5 and GPT-5.6-sol across this pull request. Reasoning, tool use, and code execution were enabled. The context-window size is not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public branch name describes the change and contains no internal task id - [x] I have run tests locally and the changed-path tests pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
72b9f92d76 |
fix(runner): restore task runtime parity (#12685)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task view shows a running agent and lets an operator guide that agent. > - The merged runner stack lost parts of the accepted task experience. > - Native event errors could hide current reasoning from the operator. > - Queued message steering had no server route on `master`. > - This pull request restores the task-runtime behavior and keeps the runner experimental gate. > - The benefit is a visible and steerable native run with durable fallback behavior. ## Linked Issues or Issue Description **What happened?** The task view could stop showing current runner reasoning. The steering action also failed because the server route was absent. Runner instruction files were not declared as supported. **Expected behavior** The task view must show current provider activity. It must use the live log when durable native events are empty or unavailable. The operator must be able to steer a queued message into the active native turn. **Steps to reproduce** 1. Enable the Paperclip Runner experimental setting. 2. Start a native runner task. 3. Open the task view while the run emits reasoning. 4. Queue a message and select the steering action. **Paperclip version or commit** The regression reproduces on `24a674f8858060e77ea1beb50689d26473e91431`. **Additional context** Related closed work: Refs #12592. ## What Changed - Restored the queued-comment steering route for active native sessions. - Added durable and queue-bound steering acknowledgements for safe retries. - Restored runner instruction bundle support. - Added live-log fallback when native events are empty or unavailable. - Restored the compact live reasoning ticker in the task view. - Added a visible temporary-unavailable state when both activity sources fail. - Kept the unified Paperclip Runner experimental gate unchanged. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm check:token-gates` - `pnpm build` - Seven focused test files passed with 140 tests. - The final steering regression file passed with 12 tests. - The broad local test run reached unrelated workspace, port, and shared database failures. The changed-area tests remained green. ## Risks - The steering route changes queue and run records in one transaction. Tests cover stale targets, unavailable sessions, lost responses, and wrong-queue acknowledgements. - Native events remain the primary transcript source. The live log is used only when event data is absent or its poll fails. - The experimental gate still hides and rejects the runner when the setting is off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — no documentation change is required for this regression repair - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
39206c0096 |
feat(ui): complete the task workspace (#12640)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task page is the main place where people guide and review agent work. > - Native runner events already project into the existing task chat on the lower stack. > - The larger task workspace must support those events without breaking direct adapters. > - Legacy questions, final replies, empty transcripts, and classic controls must keep their behavior. > - This pull request completes the provider-neutral task workspace experience. > - The benefit is one coherent task surface for native and direct execution paths. ## Linked Issues or Issue Description Refs #12639. Refs #12617. Refs #12352. **Subsystem affected** Task chat, task detail, side panels, interaction forms, and transcript presentation. **Problem or motivation** The task page does not provide one complete workspace for live activity, plans, questions, queued guidance, files, and documents. Earlier native UI work also exposed compatibility risks in legacy question and reply paths. **Proposed solution** Add the task workspace components and provider-neutral protocol presentation. Keep runner-only controls behind runtime facts. Preserve all direct-adapter composer, transcript, interaction, and finalization behavior. **Alternatives considered** A separate runner page would duplicate task behavior. Replacing legacy transcript logic would create unnecessary adapter regressions. **Roadmap alignment** This work supports the unified task experience and the Cloud and Sandbox agents milestone. ## Stack - Base PR: #12639. - Lower PR: #12638. - This PR contains only its 119-file UI delta against `runner/remote-wss-transport`. - The next stack PR adds administrator and observability controls. ## What Changed - Added a reusable task workspace side panel for files and documents. - Added provider-neutral cards for tools, plans, questions, protocol activity, and progress. - Added queued guidance and richer composer state. - Added compact and expanded interaction presentation. - Added live activity, thinking, usage, recovery, and final reply presentation. - Added document annotations and task deep links. - Added bounded question validation and response handling. - Preserved native runner event projection from master. - Preserved legacy channel-less replies and explicit final reply precedence. - Preserved direct-adapter and classic-interface controls. - Did not change server execution selection, Rust code, migrations, workflows, or `pnpm-lock.yaml`. ## Verification - GitHub Actions will run UI tests, repository tests, typecheck, build, browser tests, security, and policy gates. - Tests cover active, settled, empty-transcript, interaction, queued-message, plan, question, file, document, and classic-interface states. - Compatibility tests cover direct adapters and native runner projection together. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check runner/remote-wss-transport...HEAD` passes. - The delta contains 119 UI-only files. ## Risks - This is a large task-page change. - Most files are new focused components and tests. - Shared transcript code keeps the lower native projection and legacy direct-adapter fallbacks. - Runner-only controls use adapter and runtime facts. - No execution path or rollout flag changes in this PR. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
209680409d |
feat(ui): project native runner turns into task chat (#12617)
## Thinking Path > - Paperclip is the open source app people use to supervise AI agents and their work. > - The task page is the established place to read run progress and answer agent questions. > - Native runner events use PRP envelopes instead of the direct-adapter transcript format. > - The task page needs a narrow projection for those events without changing legacy adapter behavior. > - Unknown event versions and fields must stay hidden until the UI supports them. > - This pull request adds a runtime-gated native turn projection and preserves the classic path. > - The benefit is one task thread for native Codex runs while direct adapters keep their current UI. ## Linked Issues or Issue Description **Subsystem affected** `ui/` task chat and transcript projection. **Problem or motivation** The server can record native runner events, replies, usage, and structured interactions, but the existing task page cannot safely render those records. Reusing the native path for direct adapters would also risk the legacy question and finalization behavior. **Proposed solution** Project supported PRP v1 events into the existing transcript model only when runtime facts identify a native `paperclip_runner` run. Use an exact event and payload allowlist. Keep direct adapters on the existing transcript, composer, interaction, and finalization path. **Alternatives considered** A separate runner page was rejected because it would split task history. A universal transcript replacement was rejected because the experimental runner must not alter legacy adapters. **Roadmap alignment** This change supports activity attribution and recoverable runs. It does not replace the current task page. ## What Changed - Project supported native PRP v1 assistant, reasoning, tool, activity, usage, interaction, and result events. - Render a native runner turn only when both runtime mode and adapter type match. - Recognize the canonical `assistant_message` item kind and preserve final-reply precedence. - Fail closed for unsupported PRP versions, event types, payload schemas, and identity fields. - Keep explicit running and pending provider activities open until a real terminal state arrives. - Include channel and lifecycle identity in memoization so progress-to-final transitions rerender. - Add focused native, empty-transcript, interaction, final, memoization, and direct-adapter regressions. ## Verification - GitHub Actions is the authoritative test environment for this stack. - This top PR receives the full stack-aware CI suite. - Greptile will review this exact nine-file delta after the branch is pushed. ## Risks - The main risk is changing direct-adapter task behavior. The render gate requires both native runtime mode and the `paperclip_runner` adapter. - The next risk is exposing future provider payloads. Version, event, schema, and field allowlists fail closed. - There are no database, lockfile, workflow, build-system, or server changes in this PR. ## Stack 1. [Runner package, SDK, and developer tools](https://github.com/paperclipai/paperclip/pull/12608) 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616) 3. This PR: provider-neutral task-thread UI ## Model Used OpenAI Codex with GPT-5, extended reasoning, repository tools, and parallel review agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bc9ba7cd26 |
feat(runner): project native runs into task threads (#12321)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The experimental Paperclip Runner can execute a guarded Codex run and persist provider-neutral events. > - The task page still reads direct-adapter transcripts and cannot present those native events. > - Structured runner questions must also use the existing task interaction experience. > - Runtime selection must use the persisted run mode, not an adapter name or a current feature flag. > - This pull request projects native events and questions into the existing task thread. > - Direct adapters keep their existing transcript, composer, interaction, and finalization paths. > - The benefit is a complete native Codex task thread without a behavior change for existing adapters. ## Linked Issues or Issue Description Refs #12202. This pull request replaces that stale implementation on current `master`. **What happened?** The server persists native runner events and structured input requests. The task page only consumes direct-adapter transcripts. A native run therefore cannot present a complete transcript, usage, or question flow through the normal task experience. **Expected behavior** Native runs project persisted provider-neutral events into the existing task thread. Native structured questions use the existing interaction card. Direct adapters retain their current behavior. **Steps to reproduce** 1. Enable the experimental runner. 2. Start a native Codex run that emits progress, usage, a structured question, and a final reply. 3. Open the task page. 4. Observe that the direct-adapter transcript path cannot project the native event records. **Paperclip version or commit** `master` at `67f9867bc`. ## What Changed - Add the canonical structured-question validator and shared contract exports. - Materialize native input requests as existing task interactions. - Validate native answers and deliver them through the durable question-response receipt. - Resume the original PRP request with an idempotent `request.resolve` command. - Project native messages, tool activity, cumulative usage, and final replies into the existing transcript model. - Propagate persisted `runtimeMode` to the task page and select native handling only for `runtimeMode: "native"`. - Expire pending interactions through the shared issue service on every terminal transition, including decisions, stalled reviews, tree control, and pipeline retry cleanup. - Queue native run cancellation while a transaction is open and execute it only after the owning transaction commits. - Keep nonterminal and non-runner issue paths on their existing service call shapes and behavior. ## Verification - `pnpm --filter @paperclipai/server typecheck` — passed, including the Rust runner release build and protocol/catalog drift gates. - Focused native-thread and lifecycle suites — 18 files and 481 tests passed during review. - `issue-execution-policy-routes.test.ts` — 19/19 passed after the final transactional-queue expectation update. - `issue-agent-mutation-ownership-routes.test.ts` — 87/87 passed in the final isolated compatibility rerun. - GitHub Actions — policy, build, canary, typecheck/release registry, 5 serialized server shards, 8 general-test shards, 3 browser shards, and both aggregate gates passed on `7793f3193`. - Security — Snyk, Socket Project Report, Socket PR Alerts, and Superagent passed. - Greptile — 5/5 on `7793f3193`; all actionable review threads resolved. - `git diff --check` — passed. - Diff against `master`: 44 files. ## Compatibility Boundary - Native transcript polling only runs when the persisted run reports `runtimeMode: "native"`. - Missing or legacy runtime modes continue through `useLiveRunTranscripts`. - Legacy questions keep the existing optional free-text choice. - Native closed select sets can suppress that legacy fallback. - Terminal cleanup uses the same issue service for native and legacy interactions; only a bound native question schedules a native run cancellation. - Native cancellation happens after transaction commit, so failed or rolled-back writes do not cancel a still-valid run. - The durable delivery service checks the original native request before it considers a continuation run. - This pull request adds no migration, dependency, workflow, manifest, or lockfile change. ## Risks The main risk is routing a direct-adapter task through native handling or changing terminal issue behavior. The implementation selects the native path only from persisted runtime facts, retains the existing nonterminal call shape, and schedules native cancellation only for a validated bound native question after commit. Focused and repository-wide tests cover both paths. Native requests remain bound to the company, issue, run, and agent; answers are validated, durable, and idempotent across reconnects. ## Model Used OpenAI Codex, GPT-5 family. The client does not expose the exact deployment ID or context window. Agentic reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected local tests and they pass - [x] I have added or updated tests where applicable - [x] I have updated the compatibility notes for this change - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I addressed all Greptile and reviewer comments before requesting merge |
||
|
|
11f6c754c9 |
feat: dedicated import pause reason with visible paused-assignee notices (#12140)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import parks every imported agent as a safety default, and issue assignment wakes are dropped for paused agents > - The pause was recorded as the generic reason "system" and was almost invisible: the chat-style task thread showed nothing, the legacy notice had no action, and the new-task dialog gave no hint > - Users assigned tasks in an imported company, nothing ran, and there was no explanation — the imported company looked broken > - This pull request records a dedicated "import" pause reason and makes the paused state visible and fixable where the user is looking > - The benefit is that a silent no-op becomes an explained state with a one-click resume ## Linked Issues or Issue Description **What existing behavior does this improve?** Working with a company whose agents arrived paused from a company import. **Subsystem affected** Shared constants (`PAUSE_REASONS`), company import service (`server/src/services/company-portability.ts`), task thread and new-task dialog UI. **Current behavior** Imported agents get `pauseReason: "system"`, the same value plugin-managed and built-in agent pauses use. Assigning an issue to a paused agent silently drops the wake. The chat-style task thread renders no paused notice; the legacy thread's notice says "It was paused by the system." with no action and only renders when the composer is shown. **Proposed behavior** Import writes `pauseReason: "import"`. The paused-assignee notice explains the import pause, offers an inline "Resume agent" button (suppressed for budget pauses, which clear on their own), and renders for read-only viewers. The chat-style task thread shows the same notice above the composer. The new-task dialog warns when the selected assignee is paused. **Breaking changes** None. `PAUSE_REASONS` is widened, not changed; the column already stores free-text values in other paths, and every consumer is an equality check with a manual fallback, so an older client shows the generic fallback copy for the new value. ## What Changed - `packages/shared/src/constants.ts`: `"import"` added to `PAUSE_REASONS`. - `server/src/services/company-portability.ts`: the import pause patch writes `pauseReason: "import"`. - `ui/src/components/IssueChatThread.tsx`: `IssueAssigneePausedNotice` gains import copy, a Resume button, test ids, and is exported; it now renders even when the composer is hidden. New `onResumeAssignee` / `resumeAssigneePending` props. - `ui/src/components/TaskChatThread.tsx`: renders the paused-assignee notice above the composer dock (the chat-style thread previously had no paused surface at all). - `ui/src/pages/IssueDetail.tsx`: wires a resume mutation (`agentsApi.resume`) through both thread variants and invalidates the company agent list. - `ui/src/components/NewIssueDialog.tsx`: inline note when the chosen assignee is paused, with import-specific copy. ## Verification - `cd server && npx vitest run src/__tests__/company-portability.test.ts` — 82 tests pass (pause pin updated to `"import"`). - `cd ui && npx vitest run src/components/IssueChatThread.test.tsx src/components/NewIssueDialog.test.tsx src/components/TaskChatThread.test.tsx` — 120 tests pass (new: notice copy per reason, resume click, budget suppression, active-agent null render, dialog note). - `pnpm run typecheck` in `packages/shared`, `server`, and `ui` — clean. ## Risks - Low risk. The resume action calls the existing `POST /agents/:id/resume` route with its existing guards. Existing rows keep `"system"` and fall back to the current generic copy; only new imports write `"import"`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) with extended thinking and tool use, via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
adfbe2d4b9 |
feat(environments): refer to the managed default environment by name, not the sandbox driver key (#11838)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed deployments provision a platform-managed default environment
for agent runs; the UI shows this environment in selectors, the agent
form, run details, and the environments page
> - Those surfaces append the raw driver key to the environment name, so
users see labels like "Paperclip Computer (sandbox)", "Paperclip
Computer · sandbox", and fallback copy such as "Managed sandbox" and
"The sandbox has no ready authentication"
> - "sandbox" is infrastructure vocabulary, not the product name of the
environment; showing it next to the managed environment's name is
confusing and off-brand
> - This pull request renders platform-managed environments by name
alone and rewords the sandbox-phrased copy, while user-created
environments keep the driver suffix so mixed lists stay distinguishable
> - The benefit is that the default environment reads as one clear
product name everywhere, and self-hosted users lose nothing: their own
environments still show the driver
## Linked Issues or Issue Description
**What existing behavior does this improve?**
Display of the platform-managed default environment across the UI.
**Subsystem affected**
UI (environment selectors, agent config form, environments page, agents
page, run details) and the claude-local/codex-local adapter auth checks.
**Current behavior**
The agent form labels the inherited default environment as "Name
(sandbox)". Environment selectors and the environments list render "Name
· sandbox". The agents page describes the environment as "<provider>
sandbox provider". The agent form's fallback label is "Managed sandbox".
Adapter auth checks say "The sandbox has no ready authentication for
this adapter."
**Proposed behavior**
Platform-managed environment rows (`metadata.managedByPaperclip`) render
their name alone. The fallback label is "Paperclip Computer". The agents
page describes managed environments as "Managed by Paperclip". Run
details omit the driver suffix for sandbox-driver environments (the
adjacent Provider entry already identifies the mechanism). Adapter auth
checks say "This environment has no ready authentication for this
adapter."
**Reason and benefit**
The managed environment carries a product name. Appending the raw driver
key ("sandbox") to it is noise and contradicts the product naming.
User-created environments keep the driver suffix, so mixed lists stay
distinguishable.
**Breaking changes**
None. Message text of the auth check is not read programmatically; the
UI keys off `ADAPTER_AUTH_MISSING_CHECK_CODE`. Rows without the managed
marker render exactly as before.
## What Changed
- New `environmentDisplayLabel` helper in
`ui/src/lib/managed-sandbox-environment.ts`: managed rows → name alone;
other rows → "Name · driver".
- `AgentConfigForm`: inherited-default label uses the helper; fallback
copy "Managed sandbox" → "Paperclip Computer"; environment options use
the helper.
- `ProjectProperties`, `CompanyEnvironments`: environment selector
options use the helper; the environments-list row hides the driver
suffix on managed rows; the managed detail page's fallback description
no longer says "sandbox".
- `Agents` page: managed environments are described as "Managed by
Paperclip" instead of "<provider> sandbox provider".
- `CommentThread` run details: the driver suffix is omitted for
sandbox-driver environments.
- claude-local and codex-local adapters: auth-missing check message/hint
reworded from "sandbox" to "environment" (ACP and environment-test
paths); claude-local probe/effort/login hints reworded the same way.
- Run status lines: "Syncing workspace to sandbox", "Exporting git
changes from sandbox", "Starting adapter in sandbox", and friends now
say "environment"; "Finalizing sandbox workspace" → "Finalizing
workspace". Templated transfer-progress lines map the `sandbox`
transport key to "environment" for display (`runtime-progress.ts`).
- Agent form sign-in panel: "Sign in to the sandbox" → "Sign in to the
environment"; "Authenticated. The sandbox has credentials now." → "…The
environment has credentials now."
- Feature catalog + instance settings card: "Managed Sandbox Only" →
"Managed Environment Only" (setting key unchanged; the card keeps its
alphabetical slot).
- Server agents routes: execution-target failure and test-identity copy
no longer say "sandbox"; workspace-mode label "Cloud sandbox" → "Cloud
environment".
- Tests: new `environmentDisplayLabel` unit cases; new `AgentConfigForm`
render case asserting the managed default renders without "(sandbox)" or
"· sandbox"; status-line assertions updated across adapter-utils, server
heartbeat/live-run, and UI chat suites.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `pnpm --filter @paperclipai/adapter-claude-local typecheck` and
`--filter @paperclipai/adapter-codex-local typecheck` — clean.
- `vitest run` for `managed-sandbox-environment.test.ts`,
`AgentConfigForm.render.test.tsx`, `CompanyEnvironments.test.tsx`,
`Agents.test.tsx`, `CommentThread.test.tsx`, `NewAgent.test.tsx` — all
green (118 tests across the two runs).
## Risks
Low risk. Cosmetic label changes only; no data or API changes. Rows
without `metadata.managedByPaperclip` render exactly as before, so
self-hosted deployments with their own environments see no change. The
only self-hosted-visible wording changes are the adapter auth-check
message and the driver suffix omission on sandbox-driver rows in run
details.
## Model Used
- Claude (Anthropic) — claude-fable-5 (Claude Fable 5), Claude Code CLI,
extended thinking, tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
docs reference these labels)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
de9645ab73 |
fix(ui): surface live runtime status in the task-chat tail before the first transcript token (#11802)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs on sandbox execution targets spend their first minutes in preparation phases — config seed, workspace and skills sync into the sandbox — before the agent CLI produces its first transcript token > - The engine already reports these phases through the runtime-progress mechanism (`onRuntimeProgress` → `recordCurrentHeartbeatRunRuntimeProgress` → `currentStatusMessage` on the live run), and the pre-task-chat issue view surfaced them in its live status line > - The chat-style task view's live tail dropped that affordance: with zero renderable transcript entries it shows an opaque "Waiting for transcript..." for minutes, which reads as a hang (and prompted a real is-this-broken investigation on a healthy run) > - This pull request surfaces the live run's `currentStatusMessage` as the tail's empty-state message, with the generic wait text as fallback > - The benefit is that operators watching a sandbox run see "Syncing workspace to sandbox" instead of wondering whether the run is stuck — for every adapter, with no adapter identities involved ## Linked Issues or Issue Description **What existing behavior does this improve?** The chat-style task view's live transcript tail (`TaskChatThread` → `TaskChatLiveTail`), introduced with the experimental chat-style task view. **Current behavior** While a live run has no renderable transcript entries yet — the normal state for the multi-minute sandbox preparation window — the tail shows a static "Waiting for transcript...". The run's live `currentStatusMessage` (e.g. "Syncing workspace to sandbox", emitted by the sandbox-managed runtime's progress reporting) is available on the same live-run object but unused by this surface, although the earlier issue-chat view did display it. **Proposed behavior** When the tail is streaming a live run and no transcript rows exist yet, the empty-state message prefers the run's `currentStatusMessage`; the generic wait text remains the fallback when no runtime status has been reported (e.g. local runs that produce output immediately, or the brief pre-status window). **Reason and benefit** The preparation phases are real, reportable progress that the engine already emits. Showing them turns a minutes-long apparent hang into a legible status, for every adapter and execution target, using data the view already receives. ## What Changed - `ui/src/components/TaskChatThread.tsx`: the live tail's `emptyMessage` prefers `liveRun.currentStatusMessage` (guarded to the run the tail is actually streaming) over the static "Waiting for transcript..." fallback. Queued runs keep "Waiting to start...". - `ui/src/components/TaskChatThread.test.tsx`: a test covering both branches — a live run with a runtime status shows it (and not the wait text), and a run without one keeps the generic message. ## Verification - `vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/task-chat/TaskChatLiveTail.test.tsx`: 22/22 pass (including the new test) - `pnpm --filter @paperclipai/ui typecheck`: clean - Reproduced live: a `claude_local` run on a Daytona sandbox environment showed "Waiting for transcript..." for the full sync window; with this change the same window shows the streamed preparation statuses ## Risks - Low: a one-expression change to an empty-state string, active only while a live run has produced no renderable transcript rows. The fallback path is byte-identical to today. - `currentStatusMessage` is truncated/humanized upstream by the runtime-progress reporter; this surface renders it verbatim in the same muted style as the wait text. ## Model Used - Anthropic, **Claude Fable 5** (`claude-fable-5`) via Claude Code, with repository, shell, and Git tooling. It traced the runtime-progress mechanism end-to-end (engine emitter → heartbeat recorder → live-run API → both thread views), identified the dropped affordance in the chat-style view, and wrote the fix and test. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
cbe6395cc5 |
fix(ui): task chat composer clears on send; align carets and composer with thread (#11772)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task detail view uses a chat-style thread: agent turns, activity phases, and a message composer. > - A colleague's UX review found three defects: the composer draft stayed visible after send, disclosure carets in turns and status pills were misaligned, and the composer was not horizontally aligned with the thread column. > - These defects make the chat surface feel unpolished and cause confusion about whether a message was sent. > - This pull request fixes all three defects with small, targeted UI changes and adds regression tests for each. > - The benefit is a chat surface that behaves and aligns like users expect from a messaging UI. ## Linked Issues or Issue Description No public GitHub issue exists for this. Description of the underlying problems: **What happened?** Three UI defects in the chat-style task view: 1. After a user pressed send, the composer kept the draft text until the server round-trip finished. Fast typers could see stale text and doubt the message was sent. 2. The disclosure carets on collapsed agent turns and the caret inside the status pill did not share one alignment axis. They rendered at different x-offsets and sizes. 3. The composer container had different horizontal padding than the thread column above it, so the input box did not line up with the message bubbles. **Expected behavior** The composer clears the instant a send starts. All disclosure carets sit on one vertical axis with one size. The composer's left and right edges align with the thread column. **Steps to reproduce** Open any task in the chat-style task view. Type a message and press Enter — watch the composer text. Collapse and expand agent turns — compare caret positions. Compare the composer's horizontal edges with the message bubbles above it. ## What Changed - `TaskChatComposer.tsx`: clear the draft synchronously when a send starts instead of after the request resolves; restore the draft if the send fails. - `TaskChatTurn.tsx` and `TaskChatStatusPill.tsx`: use one shared caret alignment (size, x-offset) for turn disclosure and status pill carets, with supporting utility styles in `ui/src/index.css`. - `TaskChatThread.tsx`: align the composer container with the thread column padding. - Added or updated unit tests in `TaskChatComposer.test.tsx`, `TaskChatTurn.test.tsx`, `TaskChatActivityPhase.test.tsx`, and `TaskChatThread.test.tsx`. ## Verification - Run `pnpm vitest run src/components/task-chat/TaskChatComposer.test.tsx src/components/task-chat/TaskChatActivityPhase.test.tsx src/components/task-chat/TaskChatTurn.test.tsx src/components/TaskChatThread.test.tsx` in `ui/` — 4 files, 64 tests, all pass. - Manual: open a task in the chat view, send a message, and confirm the composer clears immediately. Collapse/expand turns and confirm the carets align. Compare composer edges with the thread column; the alignment fixes were also pixel-verified with screenshots during local review. ## Risks - Low risk. All changes are render-layer only; no server or data changes. - The composer now clears optimistically. If a send fails, the draft is restored, so no user text is lost. - Caret alignment uses shared CSS utilities; visual regressions would show in the existing component tests and in any screenshot diff. ## Model Used - Implementation: OpenAI Codex CLI coding agent (Codex model family, agentic tool use) via Paperclip's Codex adapter. - Review, verification, rebase, and PR preparation: Anthropic Claude, model id `claude-fable-5` (extended thinking, tool use), via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fd472d02ba |
Show ordered live blocker work in task chat (#11487)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The task view shows operators why work cannot continue > - The redesigned task thread now shows direct and ultimate blockers > - But it does not show the ordered task queue while a blocker chain has live work > - This pull request adds a compact ordered live-work queue to the redesigned thread > - The benefit is that operators can see completed, running, and queued dependencies without opening the larger legacy notice ## Linked Issues or Issue Description **What existing behavior does this improve?** The redesigned task thread blocker summary is improved. The merged predecessor is #11456. **Current behavior** The redesigned task thread shows compact direct and ultimate blocker links. It does not show the ordered queue when the blocker tree has live work. The legacy task view shows this queue in a larger notice. **Proposed behavior** Show a compact blue live-work queue at both ends of the redesigned task thread. Order completed tasks first, then running tasks, then queued tasks. Show a live terminal leaf as `Now running`. Return to the amber blocker links when no live dependency remains. **Reason and benefit** Operators can see the active dependency order without leaving the redesigned task view. The compact presentation preserves the new thread's low-chrome layout. **Breaking changes** None. The change only adds UI for blocker data that the task view already receives. ## What Changed - Shared the live blocker ordering helper between the legacy notice and the redesigned task thread. - Added compact ordered dependency links at the top and bottom of the redesigned thread. - Added a separate `Now running` link for a live terminal blocker leaf. - Preserved the compact amber blocker rows when live work is not present. - Added component tests and a Storybook state for the new presentation. ## Verification - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/IssueBlockedNotice.test.tsx` - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - `pnpm --filter @paperclipai/ui build-storybook` - Captured and reviewed the new Storybook state in a headless browser. ## Risks - Low risk. The queue appears only for blocked tasks whose blocker attention state is `covered` and whose dependency set contains live work. - The API does not provide an explicit queue position. The UI preserves the existing legacy ordering rule: completed, running, queued, then numeric task identifier. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The runtime does not expose the exact snapshot or context-window size. Reasoning, code execution, repository tools, and browser automation were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9e9f744f58 |
Show blocker links in the task chat (#11456)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip helps operators supervise agent work through tasks and task threads. > - The redesigned task thread shows the current work and its state. > - A blocked task did not show the dependency that prevented progress. > - Operators had to leave the thread to find the direct and final blockers. > - This pull request adds compact blocker links at the top and bottom of the task thread. > - The benefit is that operators can identify and open the relevant tasks without adding a large notice to the thread. ## Linked Issues or Issue Description **What existing behavior does this improve?** The redesigned task thread did not show which task directly blocked the current task or which task ultimately blocked its dependency chain. **Subsystem affected** `server/`, `packages/shared/`, and `ui/` task-blocker presentation. **Current behavior** A blocked task can open in the redesigned thread without a visible dependency link at the top or bottom of the conversation. **Proposed behavior** Show one compact amber row for the direct blocker. Show a second row for the selected final blocker when one exists. Render the rows at both ends of the thread. **Reason and benefit** Operators can see the reason for the blocked state and open the relevant task from the conversation. The compact rows preserve thread density. **Breaking changes** None. The new blocker-attention fields are optional. Existing clients remain compatible. ## What Changed - Added a compact task-chat component for direct and selected final blocker links. - Added the blocker rows to the top and bottom of populated and empty task threads. - Added link-ready blocker-attention details so an intermediate selected task stays on its correct direct chain. - Included blocker-link changes in the thread content key so pinned threads follow a newly added bottom row. - Added component, scrolling, server contract, and Storybook coverage for the new states. ## Verification - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx server/src/__tests__/issue-blocker-attention.test.ts` (38 tests passed) - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` ## Risks - Low risk. The rows only render while the task status is `blocked` and an unresolved blocker is available. - Long titles are truncated to keep each blocker on one line. The full task label remains available in the link title. - Older server payloads keep the original leaf-selection behavior because the new sampled details are optional. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The run used tool-enabled reasoning and code execution. The context-window size was not exposed to the run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cb0ce0de5 |
fix(ui): restore queued message interrupt action (#11374)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The task thread lets an operator add guidance while an agent run is active > - A new message can wait behind that active run as a queued message > - The classic task view lets the operator interrupt the target run from that queued message > - The redesigned task view did not expose the same action > - This pull request restores the action and keeps it bound to the exact target run > - The benefit is that operators can apply urgent guidance without switching task views ## Linked Issues or Issue Description **What happened?** The redesigned task view showed `Queued` for a queued operator message, but it did not show the existing interrupt action. **Expected behavior** The queued message must show `Interrupt` next to `Queued`. The action must stop the exact run that the message is waiting behind. **Steps to reproduce** 1. Open a task in the redesigned task view while an agent run is active. 2. Send a new operator message so it enters the queued state. 3. Observe that the queued message has no interrupt action. **Paperclip version or commit** `bc0b5a1642` **Deployment mode** All deployment modes that use the redesigned task view. ## What Changed - Preserve persisted queued state and the target run ID in the redesigned thread model. - Render a token-compliant `Interrupt` action beside the queued state. - Reuse the existing exact-run interrupt callback and show a disabled `Interrupting…` state during the request. - Keep an assigned queue target immutable so an in-flight comment cannot rebind its interrupt action to a replacement run. - Add regression tests for persisted queued messages, replacement-run races, and the in-progress action state. - No documentation update was required because this restores existing behavior. ## Verification - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx` - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/task-chat/task-chat-adapter.test.ts ui/src/pages/IssueDetail.test.tsx -t 'queued message actions|queues messages against a queued live run and interrupts that exact run|commentsToTaskChatItems'` - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx -t 'queued message|queues messages'` - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx` (52 passed) - `pnpm -r typecheck` - `pnpm test:run` (all server and UI groups passed; the CLI group passed after inherited static AWS credential variables were omitted) - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` - `pnpm build` - `pnpm check:token-gates` ## Risks Low risk. The change only adds an action to queued messages that have a target run and an interrupt callback. Messages without both values keep the current rendering. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The runtime does not expose a more specific deployment ID or context-window size. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d2665ff6b4 |
fix(ui): align the mobile task chat composer with the thread (#11296)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task thread is where people read work and guide agents. > - The mobile composer should use the same content width as the thread. > - The composer kept the desktop 80% width on mobile, so its edges did not align with the thread. > - Long assignee-aware placeholder text could also clip inside the mobile editor. > - Some extracted style tokens used legacy HSL wrappers around complete semantic colors, which made those declarations invalid. > - This pull request makes the composer full width on mobile, preserves the narrower desktop layout, wraps the placeholder, and repairs the invalid color compositions. > - The benefit is a stable mobile composer that aligns with the task thread and keeps its intended visual styles. ## Linked Issues or Issue Description Related work: Refs #11263. **What happened?** At mobile widths, the task chat composer used the same 80% width as the desktop composer. Its horizontal edges did not align with the full task thread. A long assignee-aware placeholder could clip on one line. The composer's extracted shadow also used a legacy `hsl(var(...))` wrapper around complete semantic color values, so the browser could reject the declaration. **Expected behavior** The composer must match the task thread width on mobile. It must stay narrower on larger screens. Long placeholder text must wrap inside the editor. Semantic color tokens must form valid shadows and gradients. **Steps to reproduce** 1. Open a task with the chat-style thread on a mobile viewport. 2. Compare the composer edges with the task thread edges. 3. Select an assignee whose placeholder text wraps to two lines. 4. Inspect the computed composer shadow and the extracted semantic color styles. **Paperclip version or commit** The change is based on `dc6fcd1ff1` from `master`. **Deployment mode** Local build from source. The behavior also applies to packaged web builds. ## What Changed - Made the task chat composer full width below the medium breakpoint and kept the 80% desktop width. - Matched the composer dock padding to the task thread padding. - Allowed long composer placeholders to wrap and reserved enough mobile editor height for two lines. - Replaced invalid legacy HSL wrappers around full semantic colors in extracted shadows, gradients, and approval styles. - Added a token gate that prevents legacy `hsl(var(--token))` wrappers from returning. - Added focused regression tests for responsive width, padding, placeholder wrapping, mobile height, and semantic shadow validity. ## Verification - `pnpm check:token-gates` — all four gates pass. - `pnpm --filter @paperclipai/ui exec vitest run src/components/TaskChatThread.test.tsx src/components/task-chat/TaskChatComposer.test.tsx src/components/task-chat/TaskChatComposerStyles.test.ts` — 37 tests pass. - `pnpm --filter @paperclipai/ui typecheck` — passes. - `pnpm --filter @paperclipai/ui build` — passes. The build prints existing CSS optimizer and bundle-size warnings. ## Risks - Low risk. The width change is limited to the mobile breakpoint. The desktop 80% layout remains in place. - The semantic token fixes can affect shadows and gradients that were previously invalid. The new gate prevents the invalid wrapper pattern from returning. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`. The context-window size is not exposed in this environment. The model used reasoning, repository tools, code execution, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0ee0543e5b |
feat(ui): refine the chat-style task workflow (#11263)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue detail view is the main place where people read work and guide agents. > - The chat-style task view needs clear messages, controls, properties, and document feedback. > - Dense metadata and disconnected controls make active work harder to scan. > - This pull request refines the existing chat-style task workflow across desktop and mobile layouts. > - The benefit is a clearer issue thread with faster access to the controls that guide work. ## Linked Issues or Issue Description **What existing behavior does this improve?** The change improves the chat-style issue detail view that was introduced in [#10606](https://github.com/paperclipai/paperclip/pull/10606) and expanded in [#10707](https://github.com/paperclipai/paperclip/pull/10707). **Current behavior** The issue thread spreads task controls across the page. Agent-turn metadata competes with the message content. Document annotation comments use inline placement that limits the document reading area. The mobile composer can overlap the bottom navigation. **Proposed behavior** The issue view keeps the thread focused on message content. It moves supporting controls into the properties area, adds searchable assignment, restores the sub-task tree, docks document comments in a side gutter, and keeps the mobile composer clear of navigation. **Reason and benefit** People can scan active work faster and find task controls without leaving the issue. The layout also gives documents and mobile conversations more usable space. **Breaking changes** None. The change updates presentation and interaction behavior in the existing issue UI. ## What Changed - Refined task-chat message spacing, metadata, agent bubbles, and composer alignment. - Added searchable assignment and restored sub-task navigation in the properties pane. - Moved document annotation comments into a right-side gutter. - Kept the mobile composer above the auto-hiding bottom navigation. - Added and updated focused component tests for the changed interactions. ## Verification - `pnpm check:token-gates` - `TZ=UTC pnpm --filter @paperclipai/ui exec vitest run src/components/InlineEntitySelector.test.tsx src/components/IssueDocumentAnnotations.test.tsx src/components/IssueProperties.test.tsx src/components/TaskChatThread.test.tsx src/pages/IssueDetail.test.tsx` - The token gates report 3/3 clean. - The focused test run passes 124 tests in 5 files. - Visual snapshot baselines were not updated. This follows the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot verification demoted to dormant (Jul 13 2026)." ## Risks - The changes affect several related issue-detail layouts. A browser review should cover desktop and mobile widths before merge. - The monitor-row test formats time in the host timezone. The verification command sets `TZ=UTC` to match CI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The context-window size is not exposed in this environment. The model used reasoning, repository tools, code execution, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
815e49bb7c | feat: make chat-style tasks the default experience (#11101) | ||
|
|
4e9a78db58 |
feat(ui): persist task chat composer drafts (#11076)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People send task instructions through the board task chat. > - A page refresh or task switch can discard an unfinished message in the redesigned composer. > - The existing task chat already supplies a task-specific draft key. > - The redesigned composer must use that key without changing attachment or send behavior. > - This pull request restores, saves, and clears text drafts in the redesigned composer. > - The benefit is that users can return to unfinished task messages without losing their text. ## Linked Issues or Issue Description Related prior work: #11070. This pull request extracts only the final composer draft behavior from that larger draft. **Subsystem affected** ui/ — React + Vite board UI. **Problem or motivation** The redesigned task chat composer does not use the draft key that the task thread already provides. A refresh, navigation, or unmount can lose an unfinished message. **Proposed solution** Persist text drafts by task key in local storage. Restore a draft when the composer mounts. Save changes after a short delay and flush pending text during unload or unmount. Clear the draft only after a successful send. **Alternatives considered** The composer could save on every keystroke. A short delay avoids unnecessary synchronous storage writes. The feature could also stay in the larger predecessor PR, but a focused PR is easier to review and verify. **Roadmap alignment** This is a focused usability improvement for the task conversation surface. It does not add or duplicate a roadmap capability. ## What Changed - Added safe draft storage helpers for load, save, and clear operations. - Connected the task-specific draft key to the redesigned task chat composer. - Preserved drafts across debounce windows, unmounts, page unloads, failed sends, and React Strict Mode probes. - Cleared drafts after successful sends without changing current attachment safeguards. - Added focused composer and thread integration tests. ## Verification - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/ui exec vitest run src/components/task-chat/TaskChatComposer.test.tsx src/components/TaskChatThread.test.tsx` ## Risks - Local storage can be unavailable or full. The helpers catch storage errors and keep the composer usable. - Only text is persisted. Attachments, work mode, and assignee selections remain session state. - Draft keys remain task-scoped, so text does not cross task boundaries. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with model `gpt-5`. The context-window size is not exposed in this environment. The model used agentic reasoning, tool use, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
01b51dc0e5 |
fix(runtime): stop Live badge and Working shimmer after task teardown (#10985)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task list and the chat views show a Live badge and a Working shimmer for an issue that has an active run. > - A finished task kept the Live badge and the Working shimmer after the run ended and the sandbox stopped. > - The user interface reads run liveness from the `heartbeat_runs.status` row. The run finalizer writes the terminal status in a step that is separate from the agent `status=done` update. When the sandbox or the run process stops between the two steps, `heartbeat_runs.status` stays `running` forever. > - A run row that stays `running` makes a finished task look perpetually Live, and the user interface has no guard for an issue that already reached a terminal status. > - This pull request closes the invariant "environment lease released implies the run is terminal" on the server, and adds a user interface guard that suppresses live state for a terminal issue. > - The benefit is that a finished task stops showing Live and Working, both at the source (the run row) and at the surface (the badge and the shimmer). ## Linked Issues or Issue Description **Bug description** - A completed task kept the Live badge and the Working shimmer after its run ended and the sandbox was torn down. **Steps to reproduce** - Run an agent task to completion. Let the sandbox tear down while the run finalizer is between the `status=done` update and the terminal run-status write. - Open the task list or the chat view for the finished task. **Expected behavior** - A finished task shows no Live badge and no Working shimmer. **Actual behavior (before this change)** - The finished task showed the Live badge and the Working shimmer because its `heartbeat_runs.status` row stayed `running`. This pull request supersedes the two separate pull requests #10954 (frontend) and #10955 (backend). It carries all of their changes for the same race. ## What Changed Server: - Run teardown terminalizes a still-running or still-queued run before it releases the environment lease. It writes `succeeded` when the issue already reached `done`, `cancelled` when the issue is `cancelled`, and `interrupted` otherwise. It never overwrites a status that another path already made terminal. - The recovery stale-lock sweep terminalizes an orphaned running run to `interrupted` after it confirms the process and the sandbox are both gone. It requires recorded process metadata, so it never terminalizes a live run, a queued run, or a scheduled retry. - Each terminal transition writes a run event. - The stale-lock sweep continues and clears the lock when the audit write fails. It logs the failure loudly. - New server tests cover both invariants. User interface: - A shared guard suppresses the Live badge and the Working shimmer when the issue status is terminal. - The guard keeps non-terminal `queued` and `running` issues live. - The guard prefers the newest issue live-status snapshot. - New user interface tests cover the guard and the snapshot preference. ## Verification Server: - `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && tsc --noEmit` in `server/` — 0 errors. - `pnpm --filter server test heartbeat-run-lease-release-terminalization.test.ts recovery-stale-issue-lock-sweep.test.ts` — 12 tests pass. User interface: - `pnpm --filter @paperclipai/ui typecheck` — 0 errors. - `pnpm exec vitest run ui/src/lib/liveIssueIds.test.ts ui/src/lib/issue-chat-messages.test.ts` — 40 tests pass. ## Risks - Low risk. The server change only forces a still-live run row to a terminal status when the lease releases or when the recovery sweep confirms the process is dead. It never overwrites an existing terminal status, and it guards the recovery path with process metadata to avoid terminalizing a live run. - The user interface change is additive. The guard only suppresses live state for a terminal issue and keeps queued and running issues live. - No database migration. No change to any external endpoint. ## Model Used - Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ea83c5c822 |
feat(task-chat): bring back copy/👍/👎 actions on the agent bubble footer (#11025)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task view shows a threaded conversation between the user and the agent, with agent replies rendered as bubbles that carry a "✓ Worked · N tools" summary line. > - The conference-room chat already offers per-message copy and thumbs-up / thumbs-down feedback, but the redesigned task thread dropped these controls from the agent bubble footer. > - Users lose a quick way to copy an agent reply or send feedback on it, and the redesign silently ignored the feedback-vote props it was already given. > - This pull request prepends a copy · thumbs-up · thumbs-down cluster to the bubble summary line and wires the existing feedback-vote props through. > - The benefit is a consistent feedback surface across both chat views, with no new API. ## Linked Issues or Issue Description No public GitHub issue exists for this change; the underlying issue is described inline below following the feature template (`.github/ISSUE_TEMPLATE/feature_request.yml`). #### Problem or motivation The redesigned task thread renders each agent reply with a "✓ Worked · N tools · <timestamp>" summary line, but it dropped the copy and thumbs-up / thumbs-down controls that the conference-room chat still shows. Users can no longer copy an agent reply or vote feedback from the task thread. The redesign component already received `feedbackVotes` and `onVote` props but ignored them. #### Proposed solution Prepend a copy · 👍 · 👎 cluster to the summary line, leading the always-visible timestamp, reusing the shared `IssueChatFeedbackButtons` so both chat views speak the same feedback language. Anchor the cluster to the turn's summary row (a sibling of the expandable tool-history fold) so it stays on the summary line whether the tool history is collapsed or expanded. #### Alternatives considered Placing the cluster inside the expandable fold — rejected because expanding the tool history then re-centered the actions to the middle of the tall fold. #### Roadmap alignment UI polish to the task thread; no core-roadmap overlap. ## What Changed - Add `TaskChatBubbleActions`: a copy · thumbs-up · thumbs-down cluster built on the shared `IssueChatFeedbackButtons`. - Render the cluster on the agent bubble's "✓ Worked · …" summary line, leading the timestamp; runless agent replies get the same cluster with the timestamp trailing. Human and system bubbles are unchanged. - Add a `leading` slot to `TaskChatTurn` so the actions sit on the summary row, a sibling of the tool-history fold, and stay anchored when the fold expands. - Wire the redesign to the `feedbackVotes` / `onVote` props it already received. - Add a demo binding in the `TaskChatLab` dev harness. ## Verification - `pnpm check:token-gates` — 3/3 gates CLEAN. - `pnpm typecheck` — clean across all packages. - `cd ui && pnpm vitest run src/components/task-chat/TaskChatBubble.test.tsx src/components/task-chat/TaskChatTurn.test.tsx` — 31/31 pass. - Manual: open a task thread, confirm the copy / 👍 / 👎 cluster shows on the agent bubble summary line before the timestamp, copy works, votes toggle, and the cluster stays on the summary line when the tool history is expanded. Visual change: snapshot baselines are intentionally not updated, per the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks Low risk. UI-only change scoped to the redesigned task-chat bubble footer. It reuses an existing shared feedback component and existing vote props; no API, schema, or server change. Human and system bubbles are untouched. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended thinking enabled, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e8ae5286eb |
feat(issues): explain cross-task agent writes with attribution, audit receipts, and actionable denials (#10843)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents write to tasks they do not own. They comment, they change
fields, and the control plane now permits this by default for
standard-trust agents on any task they can read
> - This makes a task thread ambiguous. A reader sees a comment from an
agent that is not the assignee, but no surface says whose authority that
write rode
> - The same gap applies to field edits. The activity stream named the
verb, but it did not show the before value, the after value, or the
reason the write was permitted
> - The remaining refusals are also opaque. An agent that hits a wall
receives a 403 with no boundary name, no actor who can act, and no
sanctioned path. One real incident spent a full detour to find the
workaround
> - This pull request adds the three surfaces that make open cross-task
writes legible: an attribution chip, a field-level audit receipt, and an
actionable denial contract shared by the API and the UI
> - The benefit is that a reader can answer "who did this, on whose
authority, and was it allowed?" on the task itself, and a blocked writer
is told what to do next
## Linked Issues or Issue Description
No public issue exists for this work, so the enhancement is described
here.
**What existing behavior does this improve?**
Cross-task agent writes are permitted, but they are not explained. A
task thread can hold comments from agents that are not the assignee, and
the activity stream can hold field changes made by those agents. Neither
surface names the responsible user behind the write. When a write is
refused, the error text does not name the boundary or the way forward.
**Subsystem affected**
Issue detail UI (comment thread and activity stream), the issue write
authorization responses in the server, and the shared copy contract that
both consume.
**Current behavior**
- An agent comment on a task the agent does not own looks the same as an
assignee comment.
- An `issue.updated` activity row states the verb only. It does not show
the field-level before and after values, the responsible user, or the
authorization reason.
- A refused write returns a short message such as an ownership error.
The message does not state which rule fired, who is able to perform the
action, or which alternative path is sanctioned.
**Proposed behavior**
- An agent comment on a task the agent does not own carries a chip that
reads "for {user}". The chip names the responsible user. Its tooltip
states that the author is not the assignee and cannot exceed that user's
permissions.
- Each `issue.updated` row shows a receipt: the changed fields with
before and after values, the responsible user, and the authorization
reason. This applies to board edits as well as agent edits.
- Each refusal states three things: the boundary that fired, who is able
to act, and the sanctioned path. The API error body and the in-app
notice use the same words, because both read one shared contract.
Related pull requests, found by searching this repository:
- Refs #10837 — merged. It added the default-open cross-task write rule,
the comment attribution data, and the per-run containment cap that this
pull request makes visible.
- Refs #10114 — open. It proposes a narrower authorization change in the
same area.
- Refs #7998 — open. It proposes append-only cross-assignee comments as
an alternative to opening writes.
## What Changed
- Adds `packages/shared/src/issue-write-denial.ts`. This is one copy
contract for eight ways an issue write can be refused: not visible,
responsible-user ceiling, responsible user unavailable, excluded actor
class, assignee run lock, per-run cross-task cap, missing run context,
and rejected attribution. Each entry names the boundary, who can act,
and the sanctioned path.
- Maps server authorization decisions onto that contract in
`server/src/routes/issues.ts` and
`server/src/services/cross-issue-influence-limit.ts`. The flattened
`error` string carries all three obligations, and `details.code` lets
the UI render the same words. The two cap codes keep the names they
already ship under.
- Adds `CommentAttributionChip`. It renders "for {user}" beside the
author name on agent comments where the author is not the assignee. It
renders nothing when no responsible user is recorded, so older rows stay
clean. It is wired into both `IssueChatThread` and the flagged
`TaskChatThread` redesign.
- Adds `IssueFieldChangeReceipt`. It renders the change receipt under
`issue.updated` rows in the activity stream. Ids resolve to agent and
user names where the directory is loaded. Server-truncated text is
labelled as a preview, so the receipt never implies that it shows a
whole value.
- Adds `IssueWriteDenialNotice`. It renders the shared copy in the app,
keyed off the denial events the server logs on a task.
- Adds a public `/ux-lab/cross-issue-collaboration` page. It renders all
three surfaces and their edge cases for review without a seeded thread.
This follows the existing `ux-lab` pages.
## Verification
Automated, all green:
```
pnpm --filter @paperclipai/shared exec vitest run src/issue-write-denial.test.ts # 17 tests
pnpm --filter @paperclipai/ui exec vitest run src/components/IssueWriteDenialNotice.test.tsx \
src/components/IssueFieldChangeReceipt.test.tsx src/components/CommentAttributionChip.test.tsx \
src/lib/issue-change-receipt.test.ts src/lib/comment-attribution.test.ts # 46 tests
pnpm --filter @paperclipai/server exec vitest run src/__tests__/cross-issue-influence-limit.test.ts \
src/__tests__/issue-comment-attribution-audit-routes.test.ts \
src/__tests__/issue-agent-mutation-ownership-routes.test.ts \
src/__tests__/low-trust-red-team-routes.test.ts # 98 tests
```
`tsc --noEmit` passes for the shared, ui, and server packages.
Manual, in a browser:
1. Start the UI only: `pnpm --filter @paperclipai/ui exec vite`.
2. Open `/ux-lab/cross-issue-collaboration`. No session is needed,
because `ux-lab` routes are public.
3. All three surfaces were captured at 1440x900 in light mode and dark
mode, and at 390x844. The page reported no errors.
4. The chip tooltip was opened by a hover and by a keyboard focus.
Rendering the page found defects that the tests had missed. Three copy
and contrast defects were fixed, and two of them are now pinned by a
test. A design review then found three layout defects, which are also
fixed: the denial notice orphaned its label when a value wrapped, the
receipt icon wrapped onto its own line at narrow widths, and the chip
tooltip was reachable by hover only.
## Risks
Low risk, and additive.
- Every new surface renders nothing when its data is absent. Comments
without a recorded responsible user show no chip, and activity events
without a receipt show no receipt, so existing rows do not change.
- No migration is included. The data these surfaces read already ships.
- The wire values of the two per-run cap denial codes are unchanged.
Only the human-readable text changes, plus six codes that had no
`details.code` before.
- The denial copy is read by agents as well as people. If wording must
change later, one shared module is the only place to change it.
- Roadmap check: this extends the completed "Activity log & action
attribution" area rather than duplicating planned core work.
## Model Used
Claude Opus 5 (Anthropic), model id `claude-opus-5[1m]`, 1M context
window, extended thinking, with tool use and code execution. It ran as
an agent in Claude Code and drove a real browser to capture the review
screenshots.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
9b9631b724 |
feat(ui): chat-style tasks polish — rich-text composer, attachment chips, live-turn interstitials, mobile layout (#10707)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators talk to their agents on the issue detail page. An experimental "Chat-Style Tasks" view (#10606) makes that page read as a conversation instead of a ticket form. > - The first release of that view shipped with a plain-text composer, no visible narration while an agent works, and a desktop-only layout. > - Users write formatted replies, paste screenshots, and follow long agent runs from their phones. The experimental view should support all of that before it can graduate. > - This pull request is the next iteration of the same experiment: a rich-text composer with attachments, live-turn narration on the status line, cleaner settled-turn history, and a mobile layout. > - The benefit is a chat view that feels alive while the agent works and stays readable after it finishes, on desktop and mobile, still fully behind the existing opt-in flag. ## Linked Issues or Issue Description Refs #49 (chat with agents is a much-wanted feature). Refs #10606 (the merged first release of the experimental chat-style task view; this PR iterates on it). Related PRs found in the dedup search: - #8228 — open PR that polishes the classic issue chat composer. It targets the flag-off legacy path; this PR only changes the flag-on experimental view. - #10466 — merged blockquote-recovery fix in the shared MarkdownEditor. This PR now reuses that editor inside the chat composer. **What existing behavior does this improve?** The experimental "Chat-Style Tasks" view on the issue detail page (Settings → Experimental, `enableTaskChatRedesign`, default off). **Subsystem affected** UI (issue detail page, chat-style task view). **Current behavior** With the experiment enabled, the composer is a plain textarea with no formatting, no attachment preview, and no mention support. While an agent runs, the status line shows only a static label, and the agent's narration text is hidden. Finished runs render one settled row per turn, so a run with many short turns produces a long list of near-duplicate "Worked" rows, and turns without a comment append at the bottom out of order. On mobile, the desktop bounded-height thread makes the page scroll poorly. **Proposed behavior** The composer uses the shared MarkdownEditor: markdown formatting, mentions, image paste with thumbnail previews, and non-image attachment chips. Sending posts on Cmd/Ctrl+Enter. While an agent runs, the status line rotates playful status words and surfaces the agent's own narration as short interstitial updates: each update holds for a minimum dwell, is replaced only when superseded, and slides through a one-line viewport with tokenized motion. Back-to-back settled turns coalesce into one "Worked" row with summed durations and re-derived tool counts, and comment-less settled turns insert chronologically at their run's start time. On mobile, the thread renders in the document flow with window auto-follow and a sticky safe-area composer; the desktop layout is unchanged. **Reason and benefit** The chat view is only convincing if it feels like a conversation with a working agent. Rich text and screenshots are table stakes for chat input. Live narration gives moment-to-moment feedback without opening transcripts. Coalesced history keeps long-running tasks readable. Mobile support lets operators follow runs away from their desks. **Breaking changes** None. Every change is gated behind the existing `enableTaskChatRedesign` flag, which is off by default. The flag-off page is unchanged. ## What Changed - `TaskChatComposer` swaps its textarea for the shared `MarkdownEditor`: markdown formatting, mentions, image paste with object-URL thumbnail previews (revoked on clear and unmount), and posting on Cmd/Ctrl+Enter. - Non-image attachments render as chips on a new shared `ui/attachment.tsx` primitive (adds the `@base-ui/react` dependency it builds on). - New `status-whimsy.ts`: deterministic rotation of playful status words on the live status line. - Live interstitial narration: the transcript adapter tags agent self-talk, and the live status line shows it as ephemeral one-line updates with a ~4s minimum dwell, hold-until-superseded replacement, and a slide transition driven by new `--motion-line-scroll` tokens (cataloged in `motion-tokens.ts`, which a test keeps 1:1 with `index.css`). Hover affordance applies only to the status line, with no leading icon. - Settled-turn history: `coalesceSettledTurns` merges back-to-back settled agent turns into one "Worked" row (summed per-run durations, tool counts re-derived from the merged turn); `assembleThreadItems` inserts comment-less settled turns chronologically at run start instead of appending them at the bottom; settled turns render tool rows only (the separate thinking block component is removed). - The "Worked" summary attaches to the reply timestamp row, and thread timestamps are always visible. - Mobile layout: the thread renders with `scroll={false}` in the page scroll, a new `useWindowAutoFollow` hook keeps the window pinned to new content, the composer is sticky with safe-area padding, and the editor uses 16px text so iOS does not zoom on focus. The desktop bounded chain is untouched. - New `TaskChatDescriptionBubble` renders the issue description as the first chat bubble, and `McpIcon` gives MCP tools a distinct icon. - `IssueDetail.test.tsx` stubs `TaskChatThread`: the composer's `@mdxeditor` dependency cannot load under jsdom's CSSOM, and the suite exercises the flag-off path. ## Verification - `pnpm check:token-gates` — 3/3 CLEAN. - `node scripts/check-task-chat-motion.mjs` — OK (30 files scanned, seams present). - `cd ui && npx tsc -b` — clean. - `cd ui && pnpm vitest run` — 3,482 of 3,483 tests pass locally. The one failure is the `IssueProperties.test.tsx` monitor-row time-formatting test, which is timezone-sensitive: it fails identically on unmodified `origin/master` in a non-UTC timezone and passes with `TZ=UTC`. It is not related to this change. - Manual: enable "Chat-Style Tasks" in Settings → Experimental and open an issue with an assigned agent. Comment to start a run: the status line rotates status words and shows the agent's narration as short held updates. After the run, consecutive turns fold into one "Worked" row under the reply timestamp. Paste an image into the composer to see a thumbnail chip; attach a non-image file to see a file chip; send with Cmd+Enter. Open the same issue in a narrow viewport to see the document-flow layout with the sticky composer. - Visual snapshot baselines are intentionally not updated: per `doc/design/DECISION-SHEET.md`, "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks - The composer now loads the shared MarkdownEditor inside the chat view. The editor is already used across the app (issue descriptions, comments), so its behavior is well exercised; composer-specific handling (paste, attachments, submit keys) is covered by new tests. - The transcript adapter changes how live narration and settled turns are derived from run logs. Malformed or legacy logs degrade to generic rows rather than crashing, and the adapter suites cover the merge and ordering rules. - Object URLs for paste previews are revoked on send-clear and unmount to avoid leaks; jsdom environments without `URL.createObjectURL` are guarded. - All changes are behind the default-off `enableTaskChatRedesign` flag. Overall risk with the flag off is low. ## Model Used - Claude (Anthropic), model id `claude-fable-5` (Claude Fable 5), extended thinking enabled, agentic tool use (file editing, shell, test execution) via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |