mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
e9d64ec1be3bb72b45f3a15a262b966975d7c725
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b43073d11f |
feat(connections): sync and group accounts managed by aggregators (#15254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external tools. > - Aggregator gateways can expose accounts that users already connected upstream. > - The Apps catalog did not show those accounts or their current provider status. > - Separate cards and setup tasks also made account ownership unclear. > - This pull request discovers upstream accounts and groups them under one app card. > - Users can find connected apps while each provider keeps control of its accounts. ## Linked Issues or Issue Description **Subsystem affected** Connections across the database, shared contracts, server, and board UI. **Problem or motivation** Users cannot see which apps are connected through a saved aggregator gateway. Native and upstream accounts need one app card. Discovery must preserve company, user, gateway, and credential boundaries. **Proposed solution** Sync account metadata from Composio, Arcade, and supported Executor gateways. Keep upstream account management in each provider. Use source chips and search to browse the catalog. Preserve native setup and the gateway's existing access policy. **Alternatives considered** Creating a local executable connection for each upstream account would duplicate authorization state. Using an agent task for routine Composio setup would add an unnecessary step. The board now calls the saved gateway directly for that setup. **Roadmap alignment** This extends the shipped Connected Apps and MCP Tool Gateway features in ROADMAP.md. The duplicate search found no open PR for managed account discovery. Related work: Refs #13755, Refs #13941, Refs #14725, Refs #13855. Open PR #12906 covers adjacent toolkit routing work. ## What Changed - Add provider-neutral discovery, sync, and refresh APIs. Preserve the Composio API paths. - Cache observations by company, saved gateway, viewing user, and credential version. Retain stale observations after failed or incomplete scans. - Add optional Arcade account sync credentials in the vault. Discover Executor accounts through its supported inventory interface. - Group native and upstream accounts in one app card. Imported account menus open their provider. Gateway menus own refresh and sync setup. - Add Paperclip, Composio, Arcade, Installed, and All chips. Show 50 catalog entries per page. Keep connected accounts above discovery. Keep explicit provider searches scoped. - Simplify Composio app setup and refresh its connected app list on the gateway Permissions page. - Add a compact agent access card and task creation defaults for connection setup. Preserve explicit blocks and approval policies. - Add two replay-safe migrations, service and UI tests, Storybook journeys, and acceptance stories. ## Verification - Passed the repository typecheck, full build, token gates, and migration ordering check. - Passed the focused provider adapter, connection interaction, and catalog tests after rebasing onto master. - Passed all nine database sync and migration replay tests using a disposable database on the test-drive PostgreSQL cluster. Removed that database after the run. - Verified Arcade cursor pagination against its official Go SDK and passed all eight adapter tests, including short and incomplete pages. - Passed all 45 interaction tests after making the exact requested tools and their Allowed/Ask first permissions visible before granting access. Verified the compact card in Storybook. - Passed the complete UI suite on the final code: 683 files and 7,432 tests, including the corrected Composio destination assertions. Passed 130 focused tests for the UUID, management-link, and health-status corrections. - Passed 22 Composio setup/sync tests, 23 connection-intent service tests, and the connection migration test in separate disposable databases. Database startup alone was substituted; the suites exercised their real SQL and services. - Passed all 10 OpenAPI route checks and the full-stack connection-intent browser test, including scoped consent, agent continuation, and task completion. - The local full runner encountered embedded PostgreSQL startup failures on this loaded macOS host. The earlier in-flight run also held the pre-fix Arcade transform; a fresh run of the final provider suite passes. The final-head CI is queued during GitHub’s active Actions incident: https://www.githubstatus.com/. The previous run also lost several runners simultaneously; its real catalog assertion failures are fixed and the fresh complete UI suite passes. - Tested the real test-drive server in the embedded browser with a live Composio gateway. Detected Airtable and Circleback. Verified refresh progress, account rows, source chips, search scope, and 50-entry pagination. - Arcade and Executor coverage uses provider fixtures. Live credentials were unavailable. - Storybook builds successfully and includes grouped native/provider accounts, stale and unavailable discovery, optional Arcade setup, and mobile states. The acceptance document records the simulated and live coverage separately. - Greptile reviewed final commit `217b024c27b5933e773ce9419c4e92b1032042c6` at 5/5. All six review threads are resolved, security scans pass, and the PR has no merge conflicts. The outstanding remote checks are `ci / Select trusted runner` and `review`, queued by GitHub. They need to complete before merge. ## Risks - Provider response changes can break inventory discovery. Failed scans retain observations and show stale status. - Composio scans only the supported catalog and can take time. Large inventories run in the background with progress and a bounded lease. - Arcade requires a project API key and user ID when the gateway cannot supply them. This key is used only for discovery. - Executor discovery depends on the server's exposed inventory tools. Unsupported servers report unavailable discovery. - Cached account rows do not grant access or create executable connections. Gateway policies still govern tool use. Account deletion and per-app authorization remain upstream. - The migrations add tables and one nullable column. Replay preserves existing rows and company-scoped foreign keys. ## Model Used OpenAI Codex, based on GPT-6. The session does not expose a more specific serving model ID or context limit. Used reasoning, repository tools, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7498705642 |
fix(ui): stabilize mobile task reading and document navigation (#15228)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People read tasks and send instructions from phones as well as desktop browsers. > - The mobile footer let page text show through its labels, and small text fields made Safari zoom on focus. > - Small scroll changes made the footer switch direction, and its changing page padding moved the conversation. > - Task pages also showed comments before question cards and run history arrived, so the composer and reading position moved again. > - Document links also used native navigation, which reset the reading position or reloaded a task through its UUID URL. Desktop tabs were crowded in the mobile drawer. > - This pull request keeps navigation steady, opens documents in the mounted task, and gives mobile readers a full-height panel with a vertical tab selector. > - The benefit is a stable task view and smoother scrolling on mobile. ## Linked Issues or Issue Description **What happened?** The mobile footer was translucent. Safari zoomed when a person focused a small text field. The footer switched abruptly while scrolling. A large task could show saved comments, then move the page again when a question card or run history arrived. In a local test with delayed responses, a late question moved the mobile composer by about 374 pixels. Opening a plan from the feed could reset the view or reload the task through a UUID link. The mobile document drawer left part of the feed exposed above small desktop tab controls. **Expected behavior** The footer has an opaque surface and moves smoothly after deliberate scrolling. Text fields do not cause automatic focus zoom. A task shows its initial conversation and composer together at the final scroll position. Background refreshes keep the existing conversation visible. Document links open in the mounted task with its existing cache and reading position. Mobile documents fill the viewport, show a clear close button, and offer a vertical list of open tabs. **Steps to reproduce** 1. Open a task with many long comments and a pending question in iOS Safari. 2. Delay its interactions, activity, and runs responses by different amounts. 3. Reload the page and watch the conversation and composer move as each response arrives. 4. Scroll down and back up, including small direction changes and edge bounce. 5. Focus the task composer, search field, and new-task title and description. 6. Open a plan or another task document from the feed, including a link that uses the task UUID. Close the panel and check the reading position. 7. Open several documents on a phone. Switch tabs and close both active and inactive tabs. **Paperclip version or commit** Developed from `1c07b5903` and rebased onto `59015846a`. **Deployment mode** Built from source. Tested in an isolated local test drive with iOS 26.5 Simulator Safari and Chrome. A temporary local proxy delayed independent responses for the layout test. Related work found in the duplicate search: - Refs #14727. It made saved replies appear before supporting history. This PR keeps its parallel requests and narrows the tradeoff in favor of a stable first layout. - Refs #14667. This open PR takes a different approach with per-run placeholders and retries. This PR fixes the observed question/composer movement and mobile navigation behavior. - Refs #13095 and #13597. These earlier fixes added task scroll anchors and skipped transcript waits for scheduled retries. - Refs #6550. Earlier mobile board polish. - Refs #9467. This related open PR changes list and generic tab reflow. The task-pane selector uses a separate component. ## What Changed - Give the mobile footer an opaque semantic surface. - Set a base-size floor for editable text on touch devices to prevent Safari focus zoom. Preserve larger title text. - Share mobile scroll tracking between both layouts. Accumulate scroll distance, ignore edge bounce and changed document bounds, and update once per frame. - Use shared motion tokens for the footer and composer. Keep page padding stable and honor reduced motion. - Wait for the initial question cards, attachments, work products, activity, runtime selection, plan, and relevant transcript history before the first reveal. Skip scheduled retries and older runs outside the initial comment window. - Bound the first reveal to 15 seconds. A stalled supporting request leaves saved conversation and the composer accessible with an explicit loading notice. - Keep concealed mobile history from stretching the document. Keep the composer mounted but concealed until the same reveal. Keep both visible during later refreshes. - Route first and repeated same-task document clicks in place. Recognize UUID and identifier links. Preserve the thread history entry and feed position. Keep modifier clicks, downloads, external links, and classic document behavior. - Give the mobile task panel the full viewport and safe-area padding. Use a visible X and 44-pixel touch controls. Replace the horizontal tab strip with a vertical selector that wraps titles and supports keyboard focus. - Add four interactive Storybook states for a few tabs, long names, many tabs, and the last tab. Reuse the production selector and tab controller. - Add navigation and tab regressions, update first-reveal regressions, and document the behavior in `DESIGN.md`. ## Verification - 392 tests passed across the seven focused task-loading, scroll, mobile-navigation, layout, and composer suites. After review fixes, all 339 tests across the four affected suites passed, including stalled-loading fallback on mobile and desktop and the motion-token catalog. - All 442 focused document, tab, task-thread, and scroll tests pass. The final click-propagation cleanup also passes all 136 task-detail tests. UI typecheck, UI production build, and `pnpm check:token-gates` passed. - All four cases in `artifact-tab-arrival.spec.ts` and `text-attachment-tabs.spec.ts` pass locally, covering desktop and mobile selection, composer focus, document rendering, and downloads of the original bytes. - `pnpm --filter @paperclipai/ui build-storybook` passed. Open the mobile tab stories under `Prototypes/Task detail/Mobile tabs`. - A local diagnostic proxy measured cached plan content at about 250 ms after the first click. The HTML load count and task request count did not change. Feed scroll stayed at the same position. First, repeated, and UUID document links were tested at phone and desktop widths. Task-reference links close their preview before the document reader opens. - In Chrome at desktop and phone widths, the delayed-response task showed one complete reveal. The late question no longer moved an already visible composer. - In iOS Simulator Safari, verified the large-task reload, opaque footer, navigation hide/reveal, and search/new-task/composer focus without automatic zoom. - Full workspace typecheck and build passed. The updated UI also passes typecheck, production build, and token gates. - All 54 checks pass on the final commit `bf5c9914e61833e7cc8794a69d72cf8c7057b952` (two additional checks are intentionally skipped). Greptile reviewed that commit at 5/5, and all review threads are resolved. - Full local `pnpm test:run` was attempted but stopped after server-fixture failures. Embedded PostgreSQL startup failure reproduced in an isolated native-interaction fixture after five startup attempts. The broad run also reported a rapid Slack callback ordering test failure. These server paths are unchanged by this PR, and their CI shards pass on the latest head. The full local suite is not claimed as passing. - A localhost proxy stalled the activity response for 30 seconds. Chrome revealed the available conversation after the 15-second deadline at both desktop and phone widths, kept the composer accessible, and cleared the loading notice when the response arrived. ## Risks Slow initial history requests can delay the first conversation reveal by up to 15 seconds. If that deadline expires, late data can change the available conversation while a loading notice remains visible. The reveal waits only for runs in the initial comment window, and later refreshes do not conceal an existing conversation. The larger editable text can change line wrapping on phones. Mobile navigation and composer motion use shared tokens and respect reduced-motion settings. Mobile tab selection changes the control layout. Document links retain URL history while sharing the task reading position; other tasks and external links keep their normal navigation behavior. I checked `ROADMAP.md`. This is a fix for existing UI behavior. ## Model Used OpenAI GPT-6 in Codex. The runtime does not expose a more specific model ID or context-window size. The agent used reasoning, code editing, terminal tools, and Chrome and iOS Simulator testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
59015846ae |
fix(chat): keep dismissed task questions in the feed (#15229)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask users questions in task chats and Agent Chat. > - Agent Chat keeps unanswered questions as compact entries in the feed. > - Regular task chats still show a pending composer badge after dismissal. > - A page reload can also open a dismissed question form again. > - This pull request applies the same feed behavior to both chat types and saves dismissal per person and task. > - Users can continue the chat and return to the original question later. ## Linked Issues or Issue Description **What happened?** Dismissing a question in a regular task chat leaves a pending composer badge. The question can return after a page reload. The earlier feed behavior only applied to Agent Chat. **Expected behavior** A dismissed question stays in the feed. Its form and pending badge leave the composer. Reload preserves dismissal. Opening the feed entry restores the original question and draft answer. **Steps to reproduce** 1. Open a regular task chat with a pending question. 2. Select an option, then dismiss the form. 3. Reload the page. Check that the composer stays clear. 4. Open the question in the feed. Check that the draft is restored. 5. Submit the answer. Check that the answered receipt appears. **Paperclip version or commit** Reproduced on master at `a386a599983519eb1d399f8b770bfccdb2a74762`. **Deployment mode** Browser UI in local and authenticated instances. This change does not depend on the agent adapter. Related work: Refs #14613, which added the Agent Chat feed behavior. Refs #9141, which validates real answers on the server. Refs #11434, which tracks comment-driven changes to interaction state. This PR changes question presentation only. ## What Changed - Show compact unanswered question entries in regular task chats. - Exclude durable questions from composer pending counts and navigation. - Save dismissal in local storage per person and task. Merge the latest saved IDs so dismissals from another tab survive reload. A new question can still open its form. - Keep the original question pending and answerable. Keep approval and permission controls. - Run the question-history regressions in both chat modes. Add reload, new-question, user/task scope, and stale-tab coverage. - Share the real-component Storybook fixture. Add a regular task test drive and an interactive dismissal/answer scenario. - Update the planning-mode browser test to check a dismissed question in the feed after reload and on mobile. - Update the design rules, implementation spec, and preview instructions. ## Verification - 329 focused thread, composer, and interaction-card tests pass. - `pnpm -r typecheck` passes. The UI typecheck also passes after the review fix. - `pnpm build` passes for the full repository. - UI build, Storybook build, and token gates pass. The UI build and token gates were rerun after the review fix. - Greptile gives final commit `98f71f221` a 5/5 score. The current-head check passes, and there are no unresolved review threads. - All 56 current-head checks are terminal: 54 pass and two conditional Storybook jobs skip. The CI run includes the full test shards, build, typecheck, browser tests, aggregate verification gate, and package canary. - The stale-tab regression fails in both chat modes before the review fix and passes after it. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/planning-mode-visual-verification.spec.ts` passes against a throwaway local server. It checks dismissal, reload, task navigation, and desktop/mobile planning controls. - The duplicate local `pnpm test:run` attempt was stopped after the full CI test suites passed. It did not complete locally. - Manual browser test: select Green, dismiss, reload, reopen, submit the saved answer, and inspect the answered receipt. Also send a new message while the unanswered question remains in the feed. The test drive uses real UI components with fixture response callbacks. - To repeat the browser test, run `pnpm storybook`. Open **Chat & Comments → Task Chat Unanswered Questions → Test Drive**. ## Risks - Dismissal is a browser-local preference. It does not sync to another browser or device. Clearing local storage removes it. - When local storage is unavailable, dismissal lasts for the mounted thread only. - Unanswered questions can accumulate in the feed. They stay pending until answered or resolved through the existing API. - No database migration, API change, or change to approval permissions. ## Model Used OpenAI Codex, `gpt-6.1-sol`, with xhigh reasoning, repository editing, code execution, and browser control. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a65ca09508 |
fix(runner): settle accepted results after shutdown failures (#15217)
## Thinking Path > - Paperclip manages AI agents and the tasks they perform. > - The native runner saves tool results and completion reports before it releases a session. > - Large project discovery responses can exceed the durable command limit. > - A shutdown failure can leave a saved answer waiting for workspace repair. > - Recovery reused the old assessment for a different status decision, which violated a database constraint. > - This pull request bounds discovery responses and lets recovery commit the saved result after workspace repair. > - The benefit is a task that reaches its correct final status without another provider turn. ## Linked Issues or Issue Description **What happened?** A native run can save its final answer, fail during shutdown, and leave the task In Progress after workspace repair succeeds. Reconciliation tries to reuse the failed-workspace assessment for a new decision. The one-decision-per-assessment constraint rejects the write. Replaying the old decision can also retain a fresh coordinator lease. Separately, retained session cleanup only recognizes the old `adapter_failed` label. The project-list tool returns full project records, including large descriptions and workspace configuration. A large response exceeds the runner's durable command limit. The settlement diagnostic previously recorded only a failure flag. **Expected behavior** Project discovery stays within the command limit. Recovery finishes the saved result after workspace repair, releases its lease, and preserves the original error for inspection. It does not repeat provider work or relax session ownership checks. **Steps to reproduce** 1. Return large project records from `list_projects` and observe an oversized semantic result. 2. Persist an accepted native completion result, then record a shutdown failure. 3. Finalize with a failed workspace, repeat that attempt, then record successful workspace repair. 4. Reconcile the run. Before this fix, the issue stays In Progress. **Paperclip version or commit** Reproduced on `a386a599983519eb1d399f8b770bfccdb2a74762`. **Deployment mode** Self-hosted server with Paperclip Runner. Related transport work: #12208 drains queued events; #12241 resumes interrupted semantic calls. This change addresses bounded project discovery and accepted-result finalization. ## What Changed - Read bounded project summary projections from the database and return at most 50 authorized summaries with a continuation cursor and explicit description truncation. Agent and run trust boundaries narrow the database candidates; project-specific policies still receive full authorization. Only visible projects determine continuations. The default project-list API remains unchanged. - Record bounded, content-free settlement failure causes for command limits, storage errors, and rejected dispatch. - Include workspace state in assessment identity. Preserve the initial assessment for interrupted finalization, and commit replacement assessment and decision references together. - Release the coordinator lease when an existing decision is replayed, without repeating its effects. - Clear stale errors when recovery succeeds and retain them in `recoveredExecutionFailure`. - Accept both current and legacy close-failure labels in the existing exact-state cleanup path. - Add regression tests and update the tool contracts and recovery documentation. ## Verification - Red: the new project paging, settlement diagnostic, current cleanup label, and repaired-workspace regressions failed on the original implementation. - Green: protocol/catalog/tool checks (117 tests), the full project-tool and finalizer suites (47 tests), cleanup ownership cases (70 tests), and cleanup sweep cases (4 tests) pass. Database-backed pagination covers large descriptions/configuration, complete enumeration, uppercase cursors, agent/run restrictions, project-policy scope contributions, and identical results/cursors when hidden projects are added. - The initial CI failures in interrupted Board waits, contended endpoint proof, and semantic schema validation were reproduced and fixed; all affected cases pass locally. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — started before the review corrections; it spanned several source revisions and was stopped after reporting old-behavior and timing failures. It is not claimed green. Fresh project and recovery suites pass; the recovery suite also passes all 25 cases with the broad runner’s isolated home/config. The Slack timing case passed in isolation. Latest-head CI is the authoritative complete test matrix. - Existing authorization suite — 66 tests passed. - Latest-head CI on `8e5763d915aa6f375bdab6601996899ea01496fc` — 55 checks passed, 4 intentionally skipped, no pending or failing checks. - Greptile — 5/5, zero unresolved threads on the same commit. - `git diff --check` — passed. ## Risks - `list_projects` now returns summaries. Callers must follow `nextCursor` and use the authorized project API for full records. Candidate narrowing is only an optimization: project policy and responsible-user authorization remain authoritative. - Assessment identity changes for workspace finalization. Existing evidence remains intact; no migration is needed. - Cleanup still requires matching identities, settled tool evidence, and verified process ownership. Unknown tool outcomes remain blocked from session reuse. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and GitHub tools. The exact serving model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a386a59998 |
Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native agents receive task constraints and completion tools from Paperclip. > - Completion tools already define the procedure for reporting a result. > - Repeated procedure text adds instructions to each full task turn. > - The final reply must still explain a blocker and link a saved document. > - This pull request removes repeated procedure text and keeps these visible outcome requirements explicit. > - A document receipt supplies the exact link, and stricter evals check the persisted reply and browser navigation. ## Linked Issues or Issue Description Refs: #14961. Related: #14948 and #15007. **What happened?** Native task envelopes repeat completion procedure text. A reduced envelope needs explicit final-reply requirements. The `write_document` receipt also lacks a canonical document link. **Expected behavior** Keep the completion tools as the source of procedure details. Require one accepted completion result before the final reply. A blocked reply must explain the reason, owner and unblock action. A document reply must contain a working link to the saved document. **Steps to reproduce** 1. Run the native assigned-skill document case and native blocker case. 2. Inspect the run-attributed provider final and its persisted comment. 3. Check the blocker explanation or open the final reply's document link. ## What Changed - Remove repeated completion procedure text from the native task constraints and backend instructions. - Keep explicit blocker and document-link requirements in full task turns. - Return a company/task-scoped `documentHref` from `write_document`. Preserve the link in the idempotent mutation receipt. - Repeat canonical links for this run's current saved revisions in accepted completion feedback. Give blocked providers final-response guidance for the cause, owner and unblock action. - Keep internal document/comment anchors when Markdown issue links load cached issue details. - Add a manual six-cell comparison suite with strict source, build, default-instruction and budget admission. - Capture eighteen shared runnerd RPC projections and six direct OpenCode HTTP projections across start, resume and continuation phases, using scripted local transports and no provider execution. - Apply v3 checks only to the manual instruction comparison; preserve v2 checks for the existing native completion suite. Check the actual persisted blocker reason and exact saved-document link. Click the rendered document link and check the original content marker in the classic document card or the new document tab. - Forward exact OpenCode finishing calls through the controller. Wait for acceptance, keep accepted feedback and concrete rejection text, and reject malformed responses. Preserve ordinary dynamic-tool response handling. - Settle the completion decision and tool response before mapping a racing idle/error/abort event or handling explicit close/interruption. Reject a concurrent finishing call before controller admission. - Add a provider-free regression through real runnerd, the OpenCode proxy and a fake provider. Reject the first completion, accept the corrected report in the same turn, and propose one result. - Keep all original verdicts unchanged. Treat replay under new checks as separate diagnostics. ## Verification - `pnpm -r typecheck` and `pnpm build` pass locally. - Native document-authority tests pass, including company/run authorization and idempotent replay. - Native runtime-context, backend and measurement tests pass. - Final-answer calibration, protocol scoring, source-admission and catalog tests pass. Wrong reasons, absent links and wrong link targets fail. - `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six single-attempt local cells with the declared models. - Exported `prepareNativeInstructionPreflight` then `verifyNativeInstructionPreflight` pass on this clean committed source. They build locally and make zero provider calls. - Corrective live confirmation is incomplete. Source |
||
|
|
994d6edcdd |
Fix Markdown in provisional task titles (#15047)
## Thinking Path > - Paperclip uses issues to organize work for AI agents. > - Paperclip creates a provisional issue title from the issue description. > - A description can start with a Markdown image or other Markdown syntax. > - The old title logic truncated the raw Markdown source. > - This pull request parses Markdown before it truncates the title. > - The benefit is a short and readable provisional title. ## Linked Issues or Issue Description **What happened?** When an issue description started with a Markdown image, Paperclip used the raw image syntax and URL in the provisional title. Other common Markdown markers could also appear in the title. **Expected behavior** Paperclip must remove Markdown syntax before it creates the 120-character provisional title. It must keep useful text, such as link labels and inline code content. **Steps to reproduce** 1. Create an issue without an explicit title. 2. Start the issue description with a Markdown image. 3. Add Markdown text after the image. 4. Observe that the provisional title contains raw Markdown syntax or an image URL. **Paperclip version or commit** The problem reproduces on `master` before this change. **Deployment mode** Local development with the embedded PGlite database. ## What Changed - Parse the description and remove image nodes before title generation. - Convert the remaining Markdown to plain text before the 120-character limit. - Keep link labels, inline code, and identifier punctuation. - Use image alt text, or `Image`, for an image-only description. - Fall back to a simple 120-character title if Markdown parsing or conversion fails. - Add regression tests for image and inline-code titles, parser failures, and text-conversion failures. ## Verification - `./node_modules/.bin/vitest run server/src/__tests__/task-title-routes.test.ts` - `./node_modules/.bin/tsc -p server/tsconfig.json --noEmit` - Full GitHub CI matrix passed on commit `663f3a0517c9e450f174b6c37970756457067ba5`. - Greptile reviewed the latest commit with a 5/5 confidence score and no unresolved findings. ## Risks - Low risk. This change only affects generated provisional titles. It does not change explicit titles or full issue descriptions. - The implementation uses the Markdown parser and plain-text converter that are already present in the server. > This is a focused bug fix. It does not add planned core work from `ROADMAP.md`. ## Model Used - OpenAI Codex with GPT-5. The run used reasoning, tool use, code execution, and repository access. The runtime did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db72ad4c73 |
fix(ui): explain branch artifacts without remote links (#15051)
## Thinking Path > - Paperclip helps people manage agent work and inspect its results. > - The issue artifact list presents those results as work-product cards. > - A branch can have no remote URL or structured metadata. > - The current card hides its summary and has no action in that case. > - This pull request shows the summary and provides an accessible details action. > - The card labels the absent link without making a false link. ## Linked Issues or Issue Description **What happened?** A branch work product without a URL showed a short title and an icon. It did not show the saved summary or provide an action. **Expected behavior** The card must show what the branch contains. People must be able to open the saved details. It must not imply that a remote URL exists. **Steps to reproduce** 1. Open the artifact list of an issue with a branch work product that has `url: null` and `metadata: null`. 2. Find the branch card. Its summary and details action are missing before this change. Related work: #12717 introduced rich work-product cards. ## What Changed - Keep linkless branch cards concise: show the title and no-link state. Keep the full summary behind an optional Saved description toggle. - Label a branch with no URL as having no remote link. Let a person open its details with a keyboard-accessible button. - Add a focused interaction test and two Storybook states for the no-link branch. ## Verification - `pnpm --filter @paperclipai/ui typecheck` passed. - Focused Vitest tests passed: 45 tests in two files. - `pnpm --filter @paperclipai/ui build` passed. - `pnpm --filter @paperclipai/ui build-storybook` passed. - `pnpm check:token-gates` passed. - The Storybook collapsed and expanded states rendered in Chromium. Both screenshots were captured. - Full workspace typecheck could not finish: the runner Rust package requires `cargo`, which is not installed in this environment. ## Risks - Low risk. The change affects card text and the action for records without a URL. It does not change the saved data or routes. - Long saved summaries on other linkless work products expand the card height after a person selects Details. ## Model Used - OpenAI Codex coding agent. The execution environment does not expose the exact model ID or context window to this agent. It used code execution and a browser to validate the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details (the managed execution workspace fixes this branch name) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (Storybook examples) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2a8a99e4a5 |
fix(ui): reveal completed thinking caret beside label (#15048)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task thread shows completed agent work and its reasoning. > - The completed work row places the disclosure caret at the far right. > - The caret stays visible even when the reader does not inspect the row. > - The row needs a quiet control that stays close to the work label. > - This pull request places the caret after that label and shows it on hover or keyboard focus. > - The change keeps the timestamp on the right and does not move content on hover. ## Linked Issues or Issue Description **What happened?** The completed agent work row shows a disclosure caret at the far right of the task thread. The caret stays visible when the pointer is elsewhere. **Expected behavior** The caret stays near the work label. It appears on hover or keyboard focus. The timestamp stays on the right. **Steps to reproduce** 1. Open a task with a completed agent run. 2. Look at the completed work row without hovering over it. 3. Hover over the row and then use the keyboard to focus its button. Related: #11772 addresses a different caret and composer layout. ## What Changed - Move the completed work caret next to the work label. - Keep its space reserved. Show it on hover or keyboard focus. - Add a test for caret position, visibility classes, and the open state. ## Verification - `node node_modules/vitest/vitest.mjs run ui/src/components/IssueChatThread.test.tsx` passes all 100 tests. - `node scripts/check-token-gates.mjs` passes all token gates. - `git diff origin/master...HEAD --check` passes. - CI passes on the latest head, including typecheck, build, tests, and verification. The local isolated worktree could not run typecheck because dependency installation failed while applying the existing `codex-acp` patch. ## Risks - Low risk. The same button still opens the work detail. Only the caret position and visibility change. - On a touch device, the caret is not visible without hover. The work row stays a button that users can tap. > This is a focused UI fix. It does not add a planned core feature from `ROADMAP.md`. ## Model Used - OpenAI Codex coding agent. The runner does not expose the exact model ID or context window. The agent used reasoning, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with the details available in this runner - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked public issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket ID - [x] I have run focused tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have checked documentation impact; no documentation change is needed for this visual fix - [x] I have considered and documented the risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1c3abf5075 |
fix(ui): use available space for composer labels (#15050)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task composer shows the next assignee and model before a message starts work. > - Fixed width limits shortened both labels when the composer had unused space. > - People could not read the selected agent and model even when the full text could fit. > - This pull request makes the capsule use the available composer width. > - It keeps the ellipsis when Plan mode or a narrow layout causes real space pressure. > - The benefit is clearer run settings without damage to the compact composer layout. ## Linked Issues or Issue Description **What happened?** The task composer truncated the assignee name at 6rem and the complete assignee and model capsule at 16rem. It did this even when the composer had more available space. **Expected behavior** The composer must show the complete assignee and model labels when they fit. It must use an ellipsis only when another control or a narrow viewport limits the available width. **Steps to reproduce** 1. Open a task composer with a long assignee name and a long model name. 2. Use a wide desktop layout. 3. Observe that the old capsule shortened both labels while unused space remained. **Paperclip version or commit** Current `master` before this change. **Deployment mode** Local dev (`pnpm dev`). **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific. This is a core UI bug. **Database mode** Not database-related. **Access context** Board. **Additional context** The Storybook cases cover a wide composer and a narrow composer with Plan mode. ## What Changed - Removed the fixed maximum width from the assignee and model capsule. - Removed the fixed maximum width from the assignee label. - Kept overflow ellipsis behavior when the parent row has insufficient space. - Added stable label selectors and focused component coverage. - Added Storybook cases for complete labels and Plan-mode truncation. ## Verification - `pnpm --filter @paperclipai/plugin-sdk build` - `pnpm --filter @paperclipai/ui typecheck` - `vitest run ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` - `node scripts/check-token-gates.mjs` - Storybook production build under Node.js 24.20.0 - Captured and inspected the two new Storybook cases. - The complete repository typecheck and build reach the Rust Runner step. This local environment does not have `cargo`. Hosted CI supplies the Rust toolchain. - The repository test suite reaches workspace-runtime tests. This local runner does not allow their required temporary home directories. Hosted CI supplies a writable test home. ## Risks - Low risk. The change only removes fixed width limits from one flex item. - A very narrow composer can still shorten both labels. This is the intended fallback. - The Storybook constrained case verifies that Plan mode and Send remain usable. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 through the Paperclip Codex runner. The run used reasoning, tool use, code execution, browser automation, and image inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb049aebf2 |
feat(skills): let agents update company skills safely (#15049)
## Thinking Path > - Paperclip is an open source control plane for AI-agent companies. > - Company skills give agents reusable work instructions. > - Skill Studio can edit skill files and save version history. > - Agents can create a skill, but they do not have a first-class update tool. > - An agent update needs a version check and safe retry behavior to prevent lost edits. > - This pull request adds `update_skill` through the existing company skill file API. > - The change keeps company policy, version history, and audit records in one path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: the server API, shared validation, and runner tool catalog. **Problem or motivation** An agent can create a company skill but cannot update its `SKILL.md` through a first-class tool. An unguarded retry can also create duplicate versions or overwrite a newer edit. **Proposed solution** Add `update_skill` with a required current version ID and a retry key. Route it through the existing skill file API. Reject stale versions and changed-input retries. Save the version and audit event together. **Alternatives considered** A separate write endpoint would duplicate the Skill Studio mutation path and policy checks. This PR reuses that path instead. **Roadmap alignment** This work extends Skills Manager and Skill Studio, which are listed in `ROADMAP.md`. **Additional context** The tool accepts a complete `SKILL.md`, not a partial patch. Callers must read the current version before they edit it. ## What Changed - Add optional version and retry fields to the existing skill file update contract. - Add a guarded API update with a stable retry receipt and attributed audit event. - Add `update_skill` to native and semantic runner tool catalogs, with mode and policy gates. - Add unit, integration, protocol, and semantic-tool regression coverage. - Document agent use and extend the OpenAPI request contract. ## Verification - Focused tests and direct server and runner TypeScript checks passed before this PR. - `git diff --check` passed after the rebase onto `master`. - CI passed on the latest PR head, including the full test matrix, typecheck, and build. Local full typecheck and build stopped because `cargo` is not installed. The local full test run ended without a verdict. - No dedicated end-to-end eval scenario was added or run. The protocol coverage and semantic-tool test cover the new action deterministically. - Reviewers can read a skill version, call `update_skill`, repeat the same key, then try a stale version and a changed-input key. Only the first edit must create a new version. ## Risks - File writes and database transactions must stay in sync when a write fails. The integration tests cover failed writes and retry behavior, but CI must verify them on the PR head. - Existing Skill Studio callers do not send the new optional coordination fields. Their request shape remains valid. ## Model Used - OpenAI Codex CLI assisted with this change. The runner did not expose the exact model ID or context window. The agent used code execution and repository tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details; exact model ID and context window were not exposed) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused tests; full suite is pending CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (53 pass, 4 skip on the latest head) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b17019e14d |
fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its adapters supply task context and access to Paperclip skills and tools. > - The default hire manual and shared prompts also repeat general work procedures. > - Those procedures overlap with stock provider instructions and the Paperclip skill. > - Existing E2E fixtures supply a QA manual, so they do not qualify the production default. > - This pull request reduces the generic instructions and adds real default-hire coverage. > - The benefit is less competing guidance, with inspectable evidence for preserved skills and task context. ## Linked Issues or Issue Description Refs: #14920. That merged change preserves native Codex base instructions. This PR covers the default manual, shared legacy prompts, operational skill guidance, and the narrowly approved ACP skill-discovery/session-environment repair for measured delivery and credential-persistence failures. **What existing behavior does this improve?** New non-CEO hires without a custom bundle and legacy task/chat startup and continuation prompts. **Current behavior** The shipped default manual contains 602 words. Generic task/chat prompts and ordinary resume deltas repeat work procedures already available through the harness and Paperclip skill. **Proposed behavior** The default manual contains only the eight-word company identity. Shared startup prompts retain identity and connection guidance. Ordinary resume deltas retain current work context without the generic execution contract. **Reason and benefit** Let the stock harness guide general work. Keep Paperclip-specific capabilities and independently test default hires, skills, ordered comments, and chat restart. **Breaking changes** New default hires receive less guidance. Existing saved manuals, explicit custom bundles, CEO templates, and specialized wake contracts retain their behavior. The obsolete includeExecutionContract option remains accepted for source compatibility. ## What Changed - Reduce the default hire manual to one sentence. - Reduce shared task/chat defaults and remove the generic ordinary-resume contract. - Keep connection guidance, auth, skills, custom prompts, and specialized wake context. - Add credential-free instruction-boundary gates and 26 explicit Product E2E cells across eight legacy/native profiles, including two focused Paperclip-storage cases. - Capture public hire receipts before providers run, then grade delivered prompts and independent task/chat outcomes. - Add an early legacy skill API recipe for saving a task document, checking the saved revision receipt and linking the document. Improve stock task/heartbeat skill-selection metadata and show a clickable Markdown UI-link example. Keep native tool completion separate. - Advertise bounded routing descriptions and exact successfully staged SKILL.md paths in legacy ACP Claude; keep full bodies on demand and preserve remote path rebasing. - Remove only the provider environment from copied persisted ACP session records, while loading current run credentials and preserving all other options/conversation state. - Regenerate both capability metadata inventories and reject stale manifests/inventories before provider admission. - Publish the original reduction and focused skill-repair comparisons, preserving all failures, automatic recovery, cost coverage and limitations. ## Verification **Behavioral qualification remains pending.** Original legacy ACP Claude loses the issue document only in the reduced cohort beneath an unchanged credential failure. A source-backed diagnosis finds that neither ordinary assignment reads the staged operational skill, while the runtime persists provider environment in session state. The new common repairs expose skill metadata/path and omit persisted env; strict document and credential guards stay intact. [Inspectable diagnosis and retained hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md). Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb` incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged #14961/#15007). All 222 affected adapter tests, adapter-utils/E2E typechecks, and final 96 variant/grader/retry calibrations pass. The frozen historical comparator is `c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths identical, exactly two production instruction paths plus nine declared unit expectations differ. The operational skill/discovery/environment repairs, selected model/profile/task/core grader/auth/permissions/retry policy are identical. Both actual launcher prepare→verify admissions pass with zero providers. [Immutable manifest and exact receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json). One original legacy ACP Claude cell per variant is authorized, with enforced single campaign attempts, 12-minute deadlines and company/agent 1,000-cent hard stops; every product recovery run/cost is counted. Actual live outcomes are pending. Current normal CI has one failed server shard and failed aggregate verify under diagnosis; other normal gates including typecheck/build/Rust/all eight browser shards pass. Fresh review completed successfully; the valid historical startup/resume masking finding was fixed with per-invocation task/chat checks and strict complete-snapshot capture, calibrated and resolved. Prior heads, failures and campaigns below remain historical evidence, not checks on this repair head. - Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52 current-head checks pass with two intentional Storybook skips, including repository typecheck/test/build and the browser shard. Fresh Greptile is 5/5 with zero unresolved review threads. Exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained receipt verification, zero providers/source errors. Fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck and 26-cell stock discovery pass. Canonical contract/inventory checks and the later issue-derived reference calibration are retained; that reference-only follow-up is not live-qualified by earlier frozen runs. - Prior full repository typecheck/build passed. The complete local Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three timing failures. All three affected files passed unchanged narrow reruns; original failures remain retained. Current-head CI now passes the full general checks; the original local failures remain retained. - The original 24-pair default-manual/shared-prompt comparison has two new overall classic Claude/OpenCode document-delivery failures plus an additional legacy ACP Claude document loss beneath an unchanged credential-guard failure (not closed by later runs), two newly passing OpenCode ordered cases, seven unchanged failures and 13 unchanged passes. Equal 15/24 totals do not establish behavioral equivalence. [Complete original report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md). - The skill-only repair holds the eight-word manual/shared prompts and merged #14920 fixed. All four matched profile configurations and 203 fixture/behavior files match. Candidate `abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill sources against baseline `bc83fe030234439ac51279502a28803958963e2e`. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547) and [baseline workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047) each pass 571 exact-source prerequisites before providers; all eight cells clean up successfully. Failed campaigns publish successfully and remain failed. - Repair pairs: Claude original Fail → Pass; Claude explicit Pass → Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff worsens beneath the unchanged failing UI-link grade: baseline gives a clickable API URL, candidate gives a code-formatted path without an anchor. The request's usable-link wording is narrower in the UI-only oracle. [Complete repair report and safe projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md). - The subsequent narrow stock metadata/link correction has two matched Pass → Pass cases, zero new machine failures/passes and no pending pairs. Both original-case handoff links remain deficient: candidate uses a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the preserved original oracle only requires a durable document. Both explicit clickable UI-link cases pass revision/content/link grading. All four exact-source 587-check gates, single assignment runs and cleanup pass. This does not establish fix causality because baseline also succeeds. [Candidate workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401) freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374) freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only comparison with reduced manuals/shared prompts held constant, not a repeat of the historical-manual comparison. Only original and clarified explicit classic OpenCode cases are selected, two per variant/four expected turns. 8,242 other tracked files and both profile hashes match; protected workflows admit each exact source before credentials. [Complete qualification report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md). Candidate original loads Paperclip/reference before saving publicly; baseline original loads it after writing locally, then saves publicly within the same assignment. Reported cost totals are $0.0107824490 candidate / $0.0107909015 baseline, with unmetered runtime. The later reference-only issue-derived link correction is provider-free calibrated and **not live-qualified** by these frozen runs; no further paid runs. - Retained tool calls show the repaired original OpenCode assignment loads only its assigned output skill before writing locally. Operational Paperclip is first loaded during automatic disposition recovery; its early recipe is visible then, but it never saves the missing document. Explicit candidate loads Paperclip and reads the new reference before saving successfully. All nine actual runs are counted. Reported LLM totals are $0.3802537209 baseline and $0.4918990161 candidate; local runtime is unmetered. - Initial setup, packaging, cancelled/missing-cell recovery, callback test and relative-output attempts remain retained. No completed provider failure was rerun. Frozen measurement branches are unchanged by later canonical metadata maintenance. - Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`, and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly for paid execution; it is excluded from `--all`. Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` replays this PR on merged hiring #14985 (`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four explicit custom-CEO-bundle checks, minimal generic manual boundary, and both suites. The combined fixture catalog and hiring calibrations pass 67 assertions; exact-head stock prerequisites pass 599 assertions (598 TypeScript + 1 Rust), all six gates and retained-receipt verification, zero providers/source errors, fingerprint `a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E typecheck and 26-cell stock discovery pass. Fresh current-head CI passes all 52 checks with two intentional skips, and fresh Greptile is 5/5 with zero unresolved review threads. The prior source-plan browser failure is retained: a deterministic process fixture replayed its last `fixture:plan` command on `chat_task_completed`, writing revision 2 with identical body after the approval handoff. This was not paid provider execution. Rebased current-head CI passes the same assertion without an old-head retry or a change to that browser fixture. The merged hiring change was measured separately on immutable matched unions, with this reduced/shared/operational context and native completion guidance held constant. [Complete original two-profile report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md): [candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208) / [historical baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463), 705 provider-free prerequisites each. Both pairs are unchanged Fail → Fail on the exact-five count, with six core delivery checks passing all four cells; 28 actual successful runs include eight automatic completion wakes, zero retries, four successful cleanups. Source-read coverage is uncomparable, actual model charges unknown. Separately versioned provider-free accounting remains analytical work; original verdicts are preserved. This does not rerun or qualify the completed default-manual or native campaigns. ## Risks - Legacy ACP Claude's additional delivery loss is not closed by any later matched run and blocks the no-extra-failing-behavior merge criterion. Legacy document delivery may have relied on the prior manual/shared prompts. The early skill repair improves Claude in one trial; the later OpenCode pairs pass in both variants and cannot establish causality or robust recovery. Both original-case links remain deficient beneath the storage-only grade. The later issue-derived reference correction has only provider-free validation. Native finish/block descriptions must not be supplied to legacy agents. - The comparison holds merged native Codex fix #14920 constant; it cannot measure that fix's before/after task performance. - These bounded skill/context/chat workflows do not measure general coding quality. Unrepresented providers remain unqualified. - Saved manuals and old Codex sessions are not automatically migrated. Codex through ACP still has a separate base-instruction follow-up. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution, and delegated PR/eval work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (relevant suites and all three unchanged narrow reruns pass; complete-run timing failures retained in Verification) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green on the new repair head (prior-head checks retained above) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the new repair head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
569c7203aa |
fix(ui): use latest issue runtime callbacks (#14863)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - The issue page uses an external-store adapter for comments. > - The adapter must keep one identity during an unrelated render. > - The adapter must also call the newest send and cancel functions. > - Passive effects update those functions too late for a synchronous runtime call. > - This pull request updates the function refs during render and adds a regression test. > - The benefit is a stable comment thread with no stale comment action. ## Linked Issues or Issue Description Refs #3678 **What happened?** The issue comment runtime can call the prior send function after a render. The callback refs do not update until passive effects run. **Expected behavior** The stable runtime adapter must call the newest function as soon as the render supplies it. **Steps to reproduce** 1. Render the issue runtime with one send function. 2. Render it again with a new send function and the same thread data. 3. Call the stable adapter before passive effects run. 4. Observe that the prior send function runs. **Paperclip version or commit** `6395cae072` **Deployment mode** Local development from source. **Agent adapter(s) involved** This is a core UI bug. It is not adapter-specific. ## What Changed - Update the latest send and cancel refs during render. - Keep the external-store adapter stable across callback-only renders. - Add a test for a runtime callback before passive effects run. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/hooks/usePaperclipIssueRuntime.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` ## Risks - Low risk. The change only updates two refs earlier in the same render. - The adapter identity and its data dependencies do not change. > This bug fix does not add or overlap with a roadmap feature. ## Model Used - OpenAI Codex with GPT-5. The exact deployed model ID and context window are not exposed. Reasoning, tool use, and local code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
78e0034498 |
fix(evals): account for hiring completion notifications (#15007)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Product E2E evals check real hiring and delegated task completion. > - The hiring fixture requires three requested CEO turns and two coder executions. > - The server can also wake the CEO when each delegated task completes. > - Two exact-five-run guards rejected these valid completion turns in all four retained cells. > - This pull request validates bounded completion turns in both guards. > - The benefit is accurate workflow grading while all actual runs and coverage failures remain visible. ## Linked Issues or Issue Description Refs: #14985, #14948, #14961. **What happened?** The original hiring comparison reports Codex Fail → Fail and Claude Fail → Fail. Each cell has seven successful runs. The five requested work turns are accompanied by two server task-completion notifications. All six other delivery checks pass. **Expected behavior** Require exactly three distinct user-requested CEO turns and one coder execution for each of two known tasks. Admit at most two strictly attributed server completion turns, including one turn that batches both tasks. Reject unknown, duplicate, failed, retried or extra-work runs. **Steps to reproduce** Inspect the retained four-cell report linked below. Each original result fails `five-successful-turns`. The same exact count was also enforced by the final chat-flow guard. ## What Changed - Add one typed lifecycle helper shared by the hiring scorer and the hiring-only final chat guard. - Validate public run ledgers, company/user/account identity, request attribution, task origins, completion deliveries, timing and replies. - Keep exactly five required work turns; declare seven maximum total turns for cost and timeout planning. - Count all actual runs, including notification runs and unexpected resets. Keep other chat count guards unchanged. - Version the hiring grader as v3 (turn accounting v2) and include the helper and chat guard in its definition digest. - Keep source-read, exact coder-body and all six other delivery checks unchanged. - Add 144 focused helper/scorer/settlement calibrations and separately versioned exact retained-input replay reports. - Retry complete bracketed observations, await both owed callbacks and attributed replies, and refresh the final guard consistently. - Reject unrelated completion writes and failed mutation attempts using exact canonical/native action IDs. Missing identity mapping is uncomparable action coverage. ## Verification - All 977 credential-free E2E support tests pass across 64 files, including 144 focused lifecycle/action/scorer/settlement calibrations. - E2E typecheck, ordinary plugin SDK and Runner TypeScript dependency builds, capability contract/inventory checks and the existing two-cell hiring discovery pass. - [Executable replay report](https://github.com/paperclipai/paperclip/blob/fed1729018cc100f5f4bbfb692777e49009c423b/doc/plans/2026-10-02-hiring-executable-accounting-replay.md) pins current code revision `e4077ade1818d98b9862ae79ee1d49a007dcf9c1`, v3 definition digest, exact source/input hashes and each original/new check. - The stricter replay verifies both Codex variants through both executable guards. ACPX Claude action attribution remains unresolved/uncomparable because provider execution IDs cannot be exactly joined to native request IDs; guards fail closed. No notification writes are observed. All six other outcomes and every original source/template coverage check stay unchanged. Original files and Fail → Fail machine verdicts remain preserved; zero providers are called. - Full attempts remain uncomparable in both profiles. Historical Claude also keeps its six-backtick exact-template mismatch. This grading repair does not prove model-performance equivalence. - The limited sidecar-v1 and initial executable-v2 passes checked notification-created tasks but could miss unrelated document writes. Those assessments remain preserved and do not prove harmless notifications. The stricter v3 replay is separate. - [Original measurement and separate sidecar](https://github.com/paperclipai/paperclip/blob/8eb517ca1497687237163bdef4dfc4d3332ea916/doc/plans/2026-10-02-hiring-template-live-comparison.md) retain 28 actual runs, eight automatic notifications, four successful cleanups and unknown actual model charges. No models are rerun. - The branch is replayed on master `59c07ede7`. Intervening master changes are UI-only; eval source bytes and replay verdicts match. The four-cell provider-free replay was repeated against the reachable code revision. - Initial-head normal CI retained browser failures in agent-run denial feedback and touch-picker scroll position. Those browser paths and imports were unchanged, but their cause was not established. The necessary review-fix head passes both browser checks; no blind rerun was requested. - Local full repository typecheck/test/build were not repeated. Exact-head normal CI passes the required repository gates, including typecheck, tests, build and browser shards. Fresh Greptile review completed on `fed1729018cc100f5f4bbfb692777e49009c423b` with 5/5 and zero unresolved threads. An independent rerun of the 144 focused helper/scorer/settlement tests passes on the unchanged head. **Merge readiness:** This PR repairs the evaluator. Its positive and negative calibrations pass, both guards reject missing action attribution, current-head CI and review pass, and there are no merge conflicts. The retained ACPX cells remain uncomparable because their action IDs cannot be joined. That coverage limit remains a separate follow-up; it does not require relaxing this grader or changing the old results. No model calls, production instructions, carrier changes, or historical regrades are part of this readiness update. ## Risks - Missing or inconsistent public lifecycle evidence fails the bounded helper. The focused calibrations reject plausible false positives and malformed observations. Unmatched action IDs fail closed and are reported as uncomparable rather than a model task regression. - Source-read evidence remains incomplete. This PR does not change provider event carriers or relax the coverage oracle. - The versioned count check differs from original v1 results. Reports retain both versions and exact input hashes. ## Model Used OpenAI Codex, GPT-6 family as identified by this session. The exact deployment ID and context-window size are not exposed. The assistant used reasoning, repository tools, code execution and delegated calibration work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip normal CI gates are green (exact head `fed1729018cc100f5f4bbfb692777e49009c423b`; fresh review tracked separately below) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (completed exact-head review; zero unresolved threads) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dd868ed125 |
fix(runner): share native completion tool guidance (#14961)
## Thinking Path > - Paperclip manages AI agents and their work. > - Native Runner agents report completion through finish and block tools. > - The providers receive different descriptions for those tools. > - Completion guidance belongs with the tools that enforce the result. > - This pull request shares the descriptions and refreshes retained catalogs. > - A separate native suite checks completion and blocking on production defaults. > - Legacy agents retain their separate skill and API paths. ## Linked Issues or Issue Description Refs: #14920, #14948, #14985. **Current behavior** Native Codex and MCP bridges describe finish and block differently. Retained provider sessions can keep old descriptions. **Proposed behavior** Native providers receive the same finish and block descriptions. The descriptions cover report selection, validation feedback, returned outcomes, approval gates and final-answer timing. Retained native sessions refresh from v13 to v14. **Reason and benefit** Put the completion procedure next to its native tool. Preserve stock base instructions, schemas, permissions and terminal semantics. This PR now stands alone on master. It contains no reduced manual, shared prompt or operational-skill changes from #14948. ## What Changed - Add canonical native finish and block descriptions. Use them in direct Codex and both native MCP bridges. - Advance the native tool contract to v14. Cover old-v13 refresh without replacing task identity or prior history. - Check authenticated tool catalogs, provider start/resume frames and serialized daemon catalogs. - Add an independent, explicit-only native completion suite. Preserve the original assigned-skill durable-document journey. Pair it with a concrete whole-task blocker across Codex, ACPX Claude and OpenCode. - Verify the actual public production default bundle and budgets before execution. Require independent durable disposition, native result/terminal receipts and observable provider-final ordering. - Correct the blocker browser oracle to accept the requested explanation. Keep exact owner/action/scope checks. Calibrate positive, missing and contradictory replies. - Preserve only actual `tool_call` terminal names (`paperclip_finish` / `paperclip_block`) in the native compatibility run-log projection. Require the same named call ID through its finishing result; retain all other redaction boundaries. - Admit verified hosted shallow checkout/build hydration and bind the selected runnerd to exact source/archive/binary provenance. Hosted cells truthfully reuse the existing trusted build; local admission executes Rust calibration. Forward only public source/run identifiers through both launcher preflight subprocess paths. - Enforce single attempts in the launcher for opted-in fixtures. Keep ordinary retry policy unchanged. Run exact-source, credential-free admission before credential loading. ## Verification - Frozen candidate: `d6e59e4712a3158ab4cd7d58deff1389b4578c21`, based on master `59c07ede72dc08b8aba149a01cc11e0b7a204621`; historical descriptions: `e74ed61a69fbdd8b3a8f15dd6456bc3140246e33`. Exactly the five original native production files and six unit tests differ. Both carry identical corrected fixtures, strict named finishing-call grader, closed compatibility carrier and admission. Defaults, profiles/models/auth/permissions and manifest bytes match. - Actual launcher `prepareNativeCompletionPreflight` → `verifyNativeCompletionPreflight` admission passes on both exact refs with zero providers: candidate 132 / historical 127 selected TypeScript assertions, 128 Node calibrations and one Rust normalization calibration each; E2E typecheck, manifest checks, selected binary provenance and six-cell discovery pass. Each has 257 explicitly skipped unrelated assertions, not coverage. The credential-free environment calibration exercises both real prepare/verify subprocess options with public hosted identifiers and rejects credential/ambient overrides. Complete actual launcher prepare→verify also passes on both frozen refs with explicitly synthetic hosted metadata/verified archives, separately labeled as calibration rather than a trusted GitHub run. Exact framed provenance parsing and mock source identity are calibrated without relaxing the real verifier. - [Complete matched qualification report](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-master-qualification.md), [immutable manifest](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-manifest.json) and [closed retained audit/hashes](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-results/comparison.json) are inspectable. All six candidate cells pass; historical descriptions pass five. Paired outcomes: **zero new failures, one new pass (Codex blocker), five unchanged passes, zero pending pairs**. [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37098728980) and [historical campaign](https://github.com/paperclipai/paperclip/actions/runs/37098815696) each execute six original attempt-1 native runs, with no campaign retry and successful cleanup. Their trusted workflow revision is `215586d127e97c9301d86e769a39a15c13298ca2`, separate from measured source. [Candidate public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098728980-1/index.html) and [historical public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098815696-1/index.html) retain declared screenshots. - Independent candidate evidence agrees with all original grades: 51 strict native checks, 12 served-default/budget checks and 21 original skill/document checks pass. The historical Codex blocker saves the correct whole-task blocker but omits the required marker from its actual provider final and identical saved reply. This is not semantic-summary fallback. Its original browser/matcher failure stays retained; the additional native snapshot/grade and workspace before/after digest were never written and are not fabricated by the separate API/PRP audit. Historical Codex completion has one failed finish followed by success within the same native run; the public receipt records no failure reason. All twelve runs and their usage remain counted. Reported model-cost subtotals are $0.00421482 historical/$0.00437391 candidate; Codex/Claude zero entries have unknown billing type, actual invoices are unverified and hosted execution cost is unmetered. One matched trial supports no extra failure within these six cases, not broad statistical or coding-quality equivalence. - Initial hosted `e18c2cf9` / `459455ac` and subsequent `0a9c5a7` / `00a761b` cohorts each stopped before providers in all twelve cells. The latter failed a mocked-receipt unit test under ambient hosted metadata; all source/build proofs passed. [All twelve later setup receipts](https://github.com/paperclipai/paperclip/blob/402ee94c52273ad58de355ae9a7d562dd22f8101/doc/plans/2026-10-02-native-completion-qualified-hosted-setup.json) are retained. [Exact failed setup receipts](https://github.com/paperclipai/paperclip/blob/27653eb1a8f8ce839776d760f4563f672e5a706c/doc/plans/2026-10-02-native-completion-master-hosted-setup.json) and the original manifest remain intact. Local sandbox-denied loopback and stale anchor-expectation attempts are retained separately; unchanged appropriate assertions were corrected/admitted before paid dispatch. Old anonymous OpenCode streams are not assigned inferred tool names or retroactively passed. - Full provider-free E2E support previously passed 927 tests in 67 files. Exact-head d6 normal CI run `37098409915`, attempt 1 passes full repository typecheck/build/tests, Runner Rust/static checks, all browser shards/aggregate and canary: 52 check-runs pass, four intentional skips, Snyk passes. Fresh Greptile check `111132956342` is 5/5 with zero unresolved threads. Source-specific deterministic tests do not substitute for the bounded live comparison. - Earlier native source `9138f570c341c251a5727c32d6615ce238bc8e03` is archived. Its [complete reduced-manual-context report](https://github.com/paperclipai/paperclip/blob/9138f570c341c251a5727c32d6615ce238bc8e03/doc/plans/2026-10-02-native-completion-live-comparison.md) remains intact, including original failures, grader limits and provider-free replay. It is not current-master-context qualification. ## Risks Changed tool text can change model behavior. The completed six-pair qualification shows no extra failing outcomes in this bounded trial; other tasks and repeated-run variance remain unmeasured. Observable final ordering does not prove provider feedback consumption. Public evidence can fail closed if a provider does not expose the required result sequence. This slice does not remove native fixed prompts or measure general coding quality. No database, schema, permission or legacy completion changes occur. ## Model Used OpenAI Codex, GPT-6 family, with code inspection, execution and tool use. The exact deployment ID and context-window size are not exposed in this session. They are unavailable rather than inferred from the model menu. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ffe5e9e2a8 |
fix: retain execution evidence for Retry and saved input (#15033)
## Thinking Path > - Paperclip lets people steer and recover AI-agent conversations. > - Recovery eligibility depends on retained cancellation receipts. > - Run presentation intentionally omits result JSON on SQL_ASCII databases and reduces oversized output. > - Retry and the saved-input sweep mistakenly used that presentation read for admission. > - The banner could offer Retry while the endpoint rejected the same stopped run. > - Read the narrow execution-evidence fields for internal admission and keep normal presentation unchanged. ## Linked Issues or Issue Description Refs #15024 and #15015. Searched existing recovery and redaction PRs; no duplicate fix found. **What happened?** On a SQL_ASCII instance, a verified pre-dispatch review-wait cancellation offers Retry in the recovery notice. Clicking it returns an eligibility conflict, and a saved user message remains deferred. The notice reads the retained database receipt, but the endpoint and sweep read a presentation projection where `resultJson` is null. **Expected behavior** Retry and saved user input use the recorded execution evidence and ordinary admission gates, independent of presentation redaction. Public run reads retain their existing encoding and output-size protections. **Steps to reproduce** 1. Record a cancelled, unclaimed review-wait continuation and its recovery hold. 2. Use the SQL_ASCII presentation projection, where run result JSON is omitted. 3. Click Retry or save a new user message and let the recovery sweep inspect it. 4. Verify a fresh turn starts once, with no replay of consumed input. **Paperclip version or commit** Reproduced on `9ae3d8db3`. **Deployment mode** Authenticated private self-hosted server with a SQL_ASCII database. ## What Changed - Add an explicit internal read of cancellation, startup, review-wait, tool-inventory, and Stop evidence; omit provider diagnostics. - Use that read in the Retry route, wakeup validation, and saved-input continuation checks. - Preserve the distinction between absent result JSON and an unrecognized stored result. - Cover the reproduced SQL_ASCII Retry and saved-input failures, retained public redaction, native Stop behavior, and excluded provider output. - Document the presentation and admission distinction. ## Verification - Red: four selected assertions fail before the fix, including the SQL_ASCII eligibility conflict and saved input remaining deferred. - Targeted green regressions and existing native Stop cases pass. - Workspace `pnpm -r typecheck` and `pnpm build` pass. - All 626 affected recovery, continuation, and route tests pass, including 22 focused admission and native Stop cases. Current-head CI has 54 passing gates and 2 skipped optional Storybook checks. Greptile reviewed `9b94af91e295a4e8007dfc6bff6a8d532d945e36` at 5/5 with no findings or open review threads. No complete local monolithic pass is claimed; the complete suite runs in sharded CI. ## Risks Admission still checks recorded process and controller ownership, provider events, cleanup, company scope, user authority, pending decisions, and task holds. The evidence projection must retain every field used by these eligibility predicates; existing native Stop cases guard against dropping its acknowledgement receipt. No schema, dependency, UI, or public response change. ## Model Used OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and context window are not exposed in this session. Used reasoning, repository tools, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9ae3d8db3d |
fix: keep pre-dispatch review waits out of execution recovery (#15024)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The queued-run gate can cancel a continuation that must wait for review. > - This cancellation happens before execution authority or a provider starts. > - Recovery currently treats the gate receipt as unknown provider execution. > - That mistake blocks a conversation after a successful reply and hides Retry. > - This pull request recognizes only the recorded, unclaimed review-wait state. > - Review waits keep their normal disposition path, and new user input can recover older false holds. ## Linked Issues or Issue Description Refs #15015, #15020, and #15022. Related #11614 narrows the review posture that causes cancellation; this change corrects recovery after a valid cancellation. **What happened?** After a successful agent reply, the queued-run gate cancelled an automatic continuation with `issue_continuation_waiting_on_review`. The gate retained `timeoutSource: stale_queued_run_gate` and a matching stop reason. No execution authority or provider started. Recovery still created an unknown-action hold, moved the task to Blocked, and hid Retry. **Expected behavior** Use normal review-wait disposition repair for this recorded state. Allow Retry, a new user message, or saved undelivered input to recover an older false hold after ordinary admission checks pass. Preserve the cancelled run and do not replay its input. **Steps to reproduce** 1. Finish an agent turn on an open task that has a real review target. 2. Let the automatic continuation reach the queued-run review gate. 3. Refresh after the cancelled run is checked by recovery. 4. Confirm a review wait is handled as a wait rather than unknown provider work. 5. Reproduce an older false hold for the same receipt, then request Retry or send a new message. 6. Confirm only one fresh turn starts, and contradictory execution or cleanup evidence retains the hold. **Paperclip version or commit** Reproduced on `215586d127`. **Deployment mode** Authenticated private self-hosted server, built from source. ## What Changed - Recognize the exact review-wait dispatch receipt only while all execution claims remain unset. - Exempt recovery only after checking retained launch and provider events, coordinators, and environment cleanup. Keep the synchronous classifier conservative without that database proof. - Apply the verified classification to automatic recovery, heartbeat retries, and stranded-queue release, so saved user input starts once. - Reuse guarded startup admission for Retry, new input, and saved input on older false holds. - Show a precise review-wait notice and keep continuation guidance consistent with Retry availability. - Verify provider events, launch events, coordinators, and cleanup before admission. - Add five initial red regressions, three additional red recovery evidence regressions, concurrent saved-input coverage, and negative evidence checks. - Document the review-wait contract. ## Verification - Red: five regressions fail on unchanged master. Existing review-wait and contradictory-evidence cases still pass. - Workspace `pnpm -r typecheck` passes. - All 752 affected recovery, continuation, queue, classification, and retry-scheduling tests pass. - Three additional review regressions failed before the database proof was added; all 12 focused review-wait cases then pass. - Workspace `pnpm build` passes. - Saved-input promotion and recovery notice regressions fail before their fixes and pass afterward. - Current-head CI has 54 passing checks and 2 skipped optional Storybook checks. Greptile reviewed `0bf1437110e1617ae4544afa41f4319eac763dba` at 5/5 with zero open findings. The complete suite runs in sharded CI. No complete local monolithic pass is claimed; an earlier long run was stopped, and its unrelated failing case passed in isolation. ## Risks The error code alone cannot establish that no provider started. This exception also requires the server gate receipt, matching stop reason, and null execution authority fields. User admission separately checks coordinator, launch and provider events, environment cleanup, pending decisions, ownership, task holds, budget, and active execution. No migration or dependency change. ## Model Used OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and context window are not exposed in this session. Used reasoning, repository tools, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
215586d127 |
fix: settle interrupted preparation with a retained cancellation fence (#15022)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A cancelled preparation must preserve history while allowing a new user turn. > - Older preparation can retain its cancellation fence but omit the unwind marker. > - The controller from that older boot is gone, its lease expired, and no provider started. > - Use the same narrow preparation proof for this retained-fence state. > - The benefit is working Retry and saved-message recovery after interrupted startup. ## Linked Issues or Issue Description Refs #15020 and #15015. Related #13293 addresses post-launch recovery. **What happened?** A preparation interrupted before native selection retained `startupCancellation.beforeNativeSelection: true` but no preparation-settled marker. Its owner expired across restart and it had no environment leases. The missing-receipt compatibility rule did not recognize the retained fence, leaving Retry absent and saved messages deferred. **Expected behavior** Admit one new user turn when the preparation evidence agrees, the old controller expired, and cleanup is complete. Preserve previous results and do not replay the cancelled input. **Steps to reproduce** 1. Cancel Paperclip Runner preparation before runtime selection. 2. Retain the cancellation fence without an unwind marker or environment leases. 3. Restart after its old controller lease expires. 4. Send a new message, select Retry, or allow the saved-message worker to reconsider new input. 5. Confirm one successor receives only undelivered input. Repeat with live ownership, invocation evidence, or pending cleanup and confirm execution stays held. **Paperclip version or commit** Reproduced on `94f6f3eb4`. **Deployment mode** Authenticated private self-hosted server, built from source. ## What Changed - Accept a retained before-selection cancellation fence in the expired historical preparation proof. - Retain runtime, stage, adapter, ownership expiry, process, event, and cleanup requirements. - Extend Retry, fresh-message, partial-queue, concurrent-worker, and contradictory-evidence tests to both receipt states. - Document recovery when the preparation-settled marker was not retained. ## Verification - Red: five exact-state regressions fail on the parent commit. - Green: all 203 continuation tests pass, including the five new regressions and 13 additional negative evidence cases. Process recovery and queued-comment routes add 413 passing tests. - A read-only candidate service check against the retained live run returns Retry eligibility without changing any task state. - An unrelated containment assertion failed once in CI. All 85 tests in that route suite pass locally and the CI shard passes on rerun. - Workspace `pnpm -r typecheck` and `pnpm build` pass. - Final head `48240de3d`: all 54 CI checks pass, two optional Storybook checks skip, and Greptile is 5/5 with zero unresolved threads. The complete local monolithic suite is covered by sharded CI; the earlier local run was stopped after a workspace case failed, and that case passed in isolation. - After merge, deploy the exact merged commit and verify an actual agent reply through the task composer. ## Risks A cancellation receipt alone must not certify that provider execution stopped. This path still requires unresolved preparation, no native identity or coordinator, an expired controller from another boot, no launch or provider evidence, and completed environment cleanup. Current-boot preparation remains held until it settles. No schema or dependency change. ## Model Used OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and context window are not exposed in this session. Used reasoning, repository tools, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
94f6f3eb47 |
fix: recover historical interrupted native preparation (#15020)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task conversation must accept new user input after an interrupted startup. > - Native startup begins with a legacy preparation row before runtime selection. > - Older builds did not retain the startup cancellation receipt on that row. > - The immutable adapter claim and expired controller can still prove that no provider started. > - This pull request uses that narrow proof for explicit Retry and saved user input. > - The benefit is a conversation that recovers after an upgrade without repeating old work. ## Linked Issues or Issue Description Refs #15015. Related #13293 covers retained process evidence after provider startup. This change covers interrupted preparation before native runtime selection. **What happened?** After an upgrade, a task stopped during native preparation can still show automatic recovery stopped. A new user message saves but does not start. Retry is absent because the historical run has no cancellation receipt. **Expected behavior** Offer Retry and admit new user input when immutable run evidence proves that no provider started and cleanup is complete. Preserve incomplete or contradictory evidence as a recovery hold. **Steps to reproduce** 1. Retain a cancelled run with the Paperclip Runner adapter claim, an unresolved runtime, and the preparing stage. 2. Keep its old controller boot ID and expired lease. Retain no native identity, coordinator, result, process identity, or provider events. 3. Upgrade from a build that did not save the startup cancellation receipt. 4. Send a new user message or select Retry. Confirm that one fresh turn starts. 5. Repeat with an active controller, a provider launch, or unfinished cleanup. Confirm that execution stays held. **Paperclip version or commit** Reproduced on `cc67d4e1d` with a historical interrupted preparation row. **Deployment mode** Authenticated private self-hosted server, built from source. ## What Changed - Recognize historical interrupted native preparation from immutable run evidence and an expired controller from another server boot. - Reject adapter invocation evidence in the startup proof. Keep process and environment cleanup checks. - Apply the same proof to saved user messages in the bounded recovery worker. - Recheck saved-message eligibility under the existing task and run locks. Keep normal ownership, decision, pause, and budget gates. - Add regression tests for Retry, a new message, concurrent saved-message recovery, and contradictory evidence. Document the compatibility rule. ## Verification - Red: Retry, new-message recovery, and saved-message recovery fail on the parent commit. - Green: 185 continuation tests, 335 process-recovery tests, and 78 queued-comment route tests pass. Two additional red regressions cover partially delivered saved queues and pass after the admission fix. - Workspace `pnpm -r typecheck` and `pnpm build` pass. - Final head `48db53f4f`: all 54 CI checks pass; two optional Storybook checks skip. Greptile is 5/5 with zero unresolved threads. - The local monolithic `pnpm test:run` was stopped after CI passed. One unrelated workspace case failed in that long run; all three workspace reconciliation cases pass in isolation. No complete local monolithic pass is claimed. - After merge, deploy the exact merged commit and verify recovery through the normal task composer. ## Risks Historical compatibility could grant a new turn without enough startup evidence. The proof requires an immutable native adapter claim, an unresolved preparing stage, no result or native identity, an expired controller from another server boot, and no invocation or provider evidence. Environment cleanup remains mandatory. The fix does not replay old input or change historical run results. No schema or dependency change. ## Model Used OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and context window are not exposed in this session. Used reasoning, repository tools, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cc67d4e1d8 |
fix: preserve steering and recover stopped task conversations (#15015)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task conversation must let a user guide a running agent and resume stopped work. > - The active run owns its input protocol, even when the user changes the next model or effort. > - Queue delivery waits for a provider receipt, which must be able to persist during the request. > - A stopped startup also needs a clear user action that passes normal task admission. > - This pull request fixes steering delivery, makes queue actions immediate, and restores explicit continuation. > - The benefit is a responsive conversation that can recover without losing saved input. ## Linked Issues or Issue Description **What happened?** A queued message could change from Steer to Interrupt while a native run prepared. A steer request could wait on its own database lock and fail to deliver. A stopped startup could then leave the conversation without a working Retry or message continuation. Interrupt also waited for the server and showed a toast. **Expected behavior** The active run keeps its input protocol. Steer delivers input to that run. Steer and Interrupt clear the submitted queue rows and show the input in the conversation immediately. Failed delivery restores the latest queue with an inline error. An eligible stopped run offers Retry, and authenticated user input can start a fresh turn through normal task admission. **Steps to reproduce** 1. Start a task with a native Paperclip Runner. 2. Change the selected model or effort while that run prepares. 3. Queue a message and press Steer. 4. Observe the provider receipt and queue state during the request. 5. Stop a startup before its provider process begins, then try Retry or send a new message. 6. Repeat queued delivery with a legacy runner and press Interrupt. **Paperclip version or commit** Reproduced on the parent of this branch, `59c07ede7`. **Deployment mode** Authenticated private deployment. The fixes also cover local task conversations. Related work: Refs #12834, Refs #13354, Refs #13275. The open refactor in #13160 moves the same queue route; it does not fix the receipt lock or stopped-run continuation addressed here. ## What Changed - Select queue behavior from the active run's immutable dispatch and runtime resolution. - Leave the run row unlocked during provider acknowledgement, then lock and read it before merging the receipt. - Retain queued input if the target run stops during that wait. Keep inline delivery errors visible after empty queue updates. - Permit exact Retry and authenticated continuation after verified native startup cancellation. Preserve pause, approval, budget, ownership, and process-stop gates. - Carry undelivered native queue input into a fresh turn once the old execution is confirmed stopped. - Show Steer and Interrupt input in the conversation and clear submitted composer rows immediately. Restore the latest queue inline on failure. Remove delivery toasts. - Keep optimistic delivery stable across stale polls, empty queues, and paginated history. Preserve classic Interrupt error handling. - Document recovery and optimistic delivery behavior. Add regression tests across server, shared queue projection, and UI boundaries. ## Verification - Red-green regression tests reproduced the queue protocol, receipt lock, stopped-startup continuation, and optimistic delivery failures. - The focused server route, continuation, queue, and runner boundary suites passed during implementation. - The queue-route suite passes with 78 tests. The three complete conversation UI suites pass with 347 tests. - UI typecheck, production build, and `pnpm check:token-gates` pass. - Workspace `pnpm -r typecheck` and `pnpm build` pass. The local monolithic `pnpm test:run` is still running; remote CI verifies the complete suite on the latest commit. - All CI gates pass on `bd9031ad56abfcde13d13a13488c1b9217c2fd3a`, including the full test shards, runner verification, browser E2E, typecheck, release registry, and canary dry run. - Greptile reports 5/5 for that commit. Both review threads are resolved. ## Risks This changes queue display and explicit continuation admission. The UI must restore rejected delivery without losing other-session edits. The server must preserve concurrent provider result updates and must not resume a process whose stop is uncertain. Focused tests cover these boundaries. This change has no database migration. ## Model Used OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and context window are not exposed in this session. Used reasoning, repository tools, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
862a5758ba |
fix(agents): reduce hiring templates to role descriptions (#14985)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - New agents receive role instructions from onboarding, the hiring skill, or a team package. > - These sources repeat harness procedures and impose generic work policies. > - They can crowd out the task and the harness instructions. > - This pull request reduces those sources to short role descriptions. > - It preserves configuration, skills, authentication, reporting lines, and approval controls. > - The benefit is less repeated instruction text with explicit coverage for the default hiring path. ## Linked Issues or Issue Description Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection. Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes. ## What Changed - Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets. - Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples. - Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog. - Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions. - Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents. - Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse. - Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison. Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing. ## Verification - PASS: 99 focused server tests and eight shipped-catalog tests. - PASS: catalog generation and validation for four shipped teams. - PASS: hiring skill validation. - PASS: `pnpm -r typecheck`. - PASS: `pnpm build`. - INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419. - PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests). - PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog. - PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck. - PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads. - PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge. The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence. ## Risks - New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome. - Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break. - Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review. - The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback. - Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6eaf218924 |
chore: grant Storybook publishing access through CODEOWNERS (#14984)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Storybook previews help contributors review the board UI. > - The publishing workflow uses the default branch CODEOWNERS file to authorize users. > - Tonio and Scott need access to this workflow. > - This pull request adds both accounts as release documentation owners. > - The existing workflow can then authorize both accounts without a separate user list in code. ## Linked Issues or Issue Description **What existing behavior does this improve?** Access to the Storybook build and publishing workflow. **Subsystem affected** Repository ownership and GitHub Actions authorization. **Current behavior** The workflow reads individual accounts from every CODEOWNERS rule on the default branch. Neither `tonio-alucema` nor `scotttong` is in that file. Both accounts already have repository admin access, but the workflow authorization check denies them. **Proposed behavior** Add both accounts to the `doc/RELEASING.md` ownership rule. After merge, both can start and rerun Storybook publishing workflows. A configured `storybook-deploy` environment reviewer must still approve deployment. CODEOWNERS membership does not add an account to the environment reviewer settings. **Reason and benefit** Contributors can publish UI previews through the same CODEOWNERS policy as other maintainers. The authorization code has no account-specific exceptions. **Breaking changes** None. The existing authorization and deployment approval checks remain in place. **Additional context** Related changes: #13226 added the publishing workflow. #13231 added stable branch bookmarks. A search found no open PR for these permission changes. ## What Changed - Add `@tonio-alucema` and `@scotttong` only to the release documentation ownership rule. - Test that accounts listed for release documentation can start and rerun publishing workflows. - Test that removing an account from CODEOWNERS removes its publishing access. - Explain how CODEOWNERS entries and environment reviewers affect publishing access. ## Verification - `node --test scripts/__tests__/storybook-deploy.test.mjs`: all 25 tests pass. - `actionlint .github/workflows/storybook-deploy.yml .github/workflows/storybook-visual.yml`: passes. - `git diff --check`: passes. - GitHub API checks confirm that both accounts have repository admin access. The deployment environment requires an existing reviewer and disables administrator bypass. - Repository-wide typecheck, tests, and build were attempted. They cannot complete in this worktree because workspace dependencies are not installed. Typecheck cannot find Node type definitions. Tests and build cannot load the workspace `tsx` package. - After merge, run `Storybook Deploy` from `master`, select a source branch, obtain environment approval, and check the published URLs in the run summary. ## Risks Both accounts gain permission to start and rerun Storybook publishing. Deployment still needs approval from a configured environment reviewer. No AWS permissions or live GitHub settings change in this PR. ## Model Used OpenAI GPT-6 through Codex. The exact backend model ID and context window are not exposed in this session. The agent used reasoning, repository inspection, code editing, shell execution, and GitHub API tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d7bdfc422c |
fix(ui): show agent avatar and align chat header controls (#14726)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent chat shows which agent receives each message. > - The chat header used initials instead of the agent avatar. > - The name and controls did not use the same center alignment. > - This change uses the shared avatar and centers the header content. > - Users can identify the agent and open its settings from the same row. ## Linked Issues or Issue Description **What happened?** The agent chat header showed initials instead of the agent avatar. The settings control did not align with the name and avatar. **Expected behavior** Show the agent avatar. Put the avatar, name, and settings control on the same horizontal center line. **Steps to reproduce** 1. Enable Agent Chat in Experimental settings. 2. Open a conversation with an agent that has an appearance set. 3. Check the avatar and settings control in the top bar. Related navigation work: #14706. ## What Changed - Use `AgentAvatar` in the conversation breadcrumb. - Update the breadcrumb key when the agent appearance changes. - Center breadcrumb labels that have an adjacent action. Keep task identifier baseline alignment unchanged. - Add regression checks for avatar props, appearance changes, and center alignment. ## Verification - PASS: affected Vitest suites, 129 tests. - PASS: `pnpm check:token-gates`. - PASS: `pnpm --filter @paperclipai/ui typecheck` on an isolated retry. - PASS: `pnpm --filter @paperclipai/ui build`. - PASS: browser checks at 1000 and 390 CSS pixels. Avatar, name, and settings control all have center Y = 29.5 CSS pixels. - Screenshots use the real header and avatar renderer with isolated context fixtures. No screenshot or fixture is part of this diff. - Full typecheck and build cannot complete because Cargo is not installed. - Full tests stopped with SIGKILL. The first UI typecheck also stopped with exit 137. These runs do not prove full-suite success. ## Risks - Low risk. The change only affects UI rendering. It changes no API or database contract. - Breadcrumbs with adjacent actions now use center alignment. Ordinary task identifiers keep baseline alignment. ## Model Used OpenAI Codex agent, with code editing, shell tools, and browser checks. The runtime did not expose the exact model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge No behavior or command documentation needs an update for this rendering fix. The execution environment requires the assigned branch name to stay unchanged. Pending review gates are not marked complete. Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ec82c5774 |
fix(runner): preserve task context when tool connections change (#14963)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use tools through company-scoped connections and provider sessions. > - Resolving a tool connection currently forces a fresh session even when the provider can load new tools into the existing conversation. > - A fresh provider conversation can receive too little history to continue the task. > - This pull request adds explicit tool-refresh capabilities and uses them in both runner paths. > - Fresh attempts receive bounded task history with source IDs and retrieval instructions. > - The benefit is that agents can continue the same task after a connection changes. ## Linked Issues or Issue Description **What happened?** A resolved tool connection forced a fresh provider conversation. The new conversation could lose the original goal and prior answers. Claude also rejected resume when only the MCP server set changed. **Expected behavior** Resume the provider conversation when its harness can refresh tools. When a fresh session is required, supply enough bounded history to continue the task. Preserve company, agent, task, workspace, model, instruction, and skill checks. **Steps to reproduce** 1. Start a conversation and agree on a task and its constraints. 2. Request and connect a tool needed for the task. 3. Continue the conversation after the connection resolves. 4. Check that the agent remembers the task and can use the new tool. **Paperclip version or commit** The bug was reproduced on master at `c46e41e81`. This branch is rebased on current master. **Deployment mode** Self-hosted server. Both legacy adapters and the native runner are affected. Related public work: Refs #13282 for task-backed conversations. Refs #13057 for the broader session-compaction proposal. Refs #14659 for another report about local CLI session continuity. This change fixes tool-connection continuation. Provider authentication repairs keep their existing recovery behavior. ## What Changed - Expose tool-refresh support in native harness descriptors and legacy adapter metadata. - Request tool refresh after connection resolution. Keep provider authentication repair as a fresh-session wake. - Reload current tools and credentials while retaining supported Claude, Codex, Grok, and other provider conversations. - Allow MCP-only changes during qualified native recovery. Keep all other compatibility checks. - Refresh managed-provider and ACPX tool bindings when attaching a new run. - Add a fresh-session handoff for both runner paths. Bound database reads, excerpts, and the final packet to 24,000 bytes. - Include the original request, recent messages, decisions, plans, prior answers, and source IDs. Mark omitted content. Apply reset boundaries, wake cutoffs, quarantine, and secret redaction. - Add regression tests and document the capabilities and handoff behavior. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. Rust formatting passed. - Final review fixes passed 314 server tests, 333 adapter utility tests, 139 native-session runtime tests, and 12 managed-provider Rust tests. They verify historical quarantine, raised budgets across attachment, no history reads on successful resume, and handoff delivery on fresh retry. - Broader branch verification also passed 1,401 adapter utility tests, 1,047 runner TypeScript tests, 43 Grok adapter tests, and 311 Rust core tests. - Live Claude CLI and Grok ACP probes preserved the provider session ID, recalled a prior task constraint, and called a newly added read-only MCP tool. - GitHub CI passed on `b21486d18084a7aa4cafbe8e012f7cad6585d9cc`: 55 successful checks and 4 skipped checks. This includes all test shards, all eight browser shards, runner checks, and the Grok clean public npm install canary. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37057514976). - A full local test attempt encountered a separate Git snapshot timeout. All affected local suites passed after the final edits, and the full CI test gates passed. - Review the capability matrix in `packages/paperclip-runner/README.md`. Repeat the four reproduction steps with a supported provider and with an unsupported harness. ## Risks - Provider tool refresh can fail. Existing recovery falls back to a fresh conversation where policy permits it. - A new transport can replace an old process while preserving the provider conversation. Tests cover current credentials and unchanged identity. - Long history can omit older context. Explicit markers and source IDs let the agent retrieve needed context within task scope. - Unknown and unqualified harnesses use the fresh-session path. No database migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and live provider testing. The exact model ID and context-window size are not exposed in this session. Claude and Grok also ran as test subjects. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
43f391e807 |
refactor(slack): clarify browser setup prompt from live testing (#14965)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack chat connections let people start and continue agent work from Slack. > - The setup prompt guides an agent through the Paperclip and Slack browser interfaces. > - A live setup completed, but several instructions did not match the current interfaces. > - Those gaps can send users to the wrong connection flow or leave them waiting for controls that do not appear. > - This pull request updates the prompt with the steps observed during the live setup. > - The benefit is a clearer path from a fresh instance to a verified Slack conversation. ## Linked Issues or Issue Description Refs: #13920 and #14862. The first added the Slack conversation flow. The second updated the shared setup prompt control. A search found no duplicate open PR or matching public issue. **Issue type** Outdated instructions and missing setup guidance. **Where is the issue?** `ui/src/pages/apps/chat/SlackSetupPrompt.tsx` **What's wrong?** The prompt omits the Chat connectors setting on fresh instances. It uses an old navigation label. It assumes an avatar crop dialog and a Save button always appear. It also assumes the suggested bot username matches Slack and that the Slack reply contains a task link. **Suggested fix** Use the current labels. Explain the feature prerequisite, avatar save behavior, real mention selection, clipboard recovery, and the path to task and run evidence. ## What Changed - Add the Chat connectors prerequisite and current Connectors, resume, and identity-link labels. - Explain Slack's combined Create and Install action and how to continue from its success page. - Handle avatar uploads that save immediately. Require a reload to confirm the saved icon. - Select the real bot from Slack's mention suggestions, including names with punctuation. - Recover from an empty or stale clipboard without exposing credentials. - Find the linked task through Conversations and inspect the agent's Runs page. ## Verification - `pnpm exec vitest run ui/src/pages/apps/chat/SlackSetupPrompt.test.tsx`: 10 tests passed after rebasing onto current master. - `git diff --check origin/master...HEAD`: passed. - `pnpm build`: passed. - `pnpm -r typecheck`: passed. - `pnpm check:token-gates`: passed. - Full Vitest suite: passed in CI for this commit, including all server, chat, workspace, and serialized test shards. The duplicate local `pnpm test:run` was stopped after CI finished; it did not complete locally. - Current-commit CI: all checks passed, including build, typecheck, E2E, Runner verification, and canary dry run. Greptile rated the change 5/5 with no findings or unresolved review threads. - Live browser test before the wording update: created and installed a new Slack app, verified the callback, uploaded and reopened the avatar, linked the configuring user's identity, and enabled the selected test channel. The first mention received a reply. A follow-up without another mention recalled the first message. Both agent runs succeeded. - The wording update does not repeat Slack app installation. Existing tests verify the complete copied prompt, clipboard fallback, and instance URL handling. - This is a prompt-text refactor. No runtime behavior changes, so the existing tests cover the copied result without a new test that repeats the wording. ## Risks - Low risk. This change updates prompt text in one file. - Provider interfaces can change. The prompt tells the agent to inspect the current page and handle optional controls. - The prompt handles the observed mention mismatch. This change does not alter the generated suggested username. ## Model Used - OpenAI Codex, based on GPT-6, with reasoning, repository tools, code execution, and browser automation. The session does not expose an exact runtime model ID or context window size. The live test also used an earlier model whose exact ID was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d034ba7491 |
fix(interactions): derive question storage from canonical forms (#14946)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents request human input through durable issue interactions. > - A question form has a canonical presentation and a compatibility storage format. > - The creation API required agents to write both formats. > - Tool guidance told agents to split text and choice questions across those formats. > - This pull request accepts one complete canonical form and derives storage fields on the server. > - The benefit is a complete question card with stable answer and retry behavior. ## Linked Issues or Issue Description Related work: Refs #13630 and #14430. PR #13630 addresses the display of historical partial forms. This change fixes creation and keeps the check that rejects conflicting new forms. **What happened?** A question save supplied three compatibility questions and one canonical text question. The API correctly rejected the incomplete canonical form. The Runner's tool description encouraged this split. Sending only a complete canonical form also failed because the API required compatibility questions. **Expected behavior** An agent sends one complete `payload.questionSet` with every text and choice question. Paperclip derives `payload.questions` for storage and answer compatibility. Existing legacy requests remain valid. Explicitly conflicting dual forms remain invalid. **Steps to reproduce** 1. Call `paperclip_request_human_input` with `interactionKind: "questions"`. 2. Send `payload: { version: 1, questionSet: ... }` with a required text question and a required choice question. 3. The old API rejects the missing compatibility questions. With this change, it stores both questions and preserves the canonical form. 4. Retry with the same idempotency key. Confirm that only one interaction exists. 5. Submit both answers. Confirm that the normal resolver and continuation rules apply. **Paperclip version or commit** The branch is based on `cf8ad63c8`. The problem affects the native Runner and the interaction creation API. **Deployment mode** Server deployment with the native Paperclip Runner. Integration tests use the real interaction service and an embedded test database. ## What Changed - Add one shared canonical-to-storage projection. Reuse it for native harness question requests. - Accept canonical-only question creation at the shared validator and server boundary. - Export the input type and update the plugin SDK and its RPC contract. - Advertise a typed, complete question form in the live and scenario tool schemas. - Enforce canonical text and custom-answer constraints before ordinary or native resolution. Preserve harmless display whitespace. - Run regex matching in isolated workers with a deadline and resource limits. Both answer paths await the result before persistence. Saved native delivery uses the validated answer without taking another worker slot. - Update agent guidance and generated Runner contracts. - Test mixed forms, option-ID collisions, retries, answers, legacy requests, and conflicting forms. ## Verification - Interaction service, HTTP route, native bridge, and Runner authority suites: 221 tests passed after correcting an obsolete tool-description assertion. - Shared validator, plugin SDK, CLI, and UI compatibility suites: 67 tests passed. - Runner core tool-contract suite: 20 tests passed. AJV validates live and scenario schemas. - Final review regressions: 172 shared, service, native bridge, and authority tests passed. These cover text length, pattern, numeric limits, whitespace, custom option IDs, and historical pending cards. - Runner session suites: 67 tests passed. Published example tests: 4 tests passed. - Server typecheck and the shared/server builds passed after the compatibility fixes. - Final delivery verification: 35 response-delivery tests passed. The native delivery regression proves saved answers do not enter pattern workers; server typecheck and build passed. - Pattern security and answer-flow verification: 205 tests passed after repairing the child fixture loader. These cover pathological matching, event-loop responsiveness, worker concurrency, slot cleanup, HTTP routes, native delivery, and the full helper in a child process. - `pnpm -r typecheck` passed on the bounded-worker revision. - `pnpm build` passed on the bounded-worker revision. - All 55 GitHub checks passed on `fe457af`; four optional jobs were skipped. An unchanged Cursor adapter test timed out once in CI, passed locally, and passed on one failed-job rerun. - Reviewers can send the canonical-only mixed form above and verify that the saved interaction contains both canonical and compatibility questions. ## Risks - The creation API accepts a new input shape. Stored rows and answer contracts keep the existing shape. - The shared projection must preserve synthetic free-text option IDs. Collision and native round-trip tests cover this behavior. - Historical partial rows remain readable. New conflicting dual forms, including written-answer mismatches, remain rejected. - Existing pending cards retain the written-answer paths offered by their stored options. Canonical text constraints still apply. - Ordinary answers now enforce declared canonical constraints before persistence. Invalid answers leave the card pending. - Regex validation has a one-second deadline and a four-worker capacity limit. A complex pattern or capacity error leaves the card pending with a validation error. - No database migration or change to company authorization is required. ## Model Used - OpenAI GPT-6 through Codex. The session exposes the GPT-6 model family; its exact runtime model identifier and context window size are not exposed. Used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9786f6df56 |
fix(runner): preserve credential content in document saves (#14937)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner sends authorized tool calls to the control plane. > - Agents use these calls to save plans and instruction files. > - The Runner used diagnostic secret detection to reject execution arguments. > - Ordinary credential-related prose could reject a document save before persistence. > - This pull request forwards the original arguments and leaves credential policy to the provider harness. > - The benefit is reliable saves with useful diagnostic records. ## Linked Issues or Issue Description Related foundation: Refs #12415 and #14430. No duplicate save-policy fix was found. **What happened?** A `write_document` call failed before the server saved its plan. The Runner reported `semantic tool input contains credential material; refusing to execute altered arguments`. The detector also masked ordinary phrases such as `secret manager` and `credential handling` in diagnostics. Both TypeScript dispatchers had equivalent execution gates. One dispatcher also rewrote structured approval and question payloads before execution. **Expected behavior** Paperclip forwards authorized arguments unchanged. The provider harness decides credential-content policy. Log and audit redaction does not reject or rewrite save input. **Steps to reproduce** 1. Send an authorized `write_document` call with a plan that discusses credential handling. 2. Include an intentional credential value in the body to exercise harness-owned policy. 3. The old Runner rejects the call. With this change, the document service stores the exact body. 4. Diagnostic records still mask explicit credential values. Qualified credential fields, short bearer values, opaque diagnostic pairs, and valid encoded JSON token headers have regression coverage. **Paperclip version or commit** Reproduced at `c46e41e81c03cd3c8b64cf993615b604d7fe8c62`. The branch is based on current `master`. **Deployment mode** Server deployment with the native Paperclip Runner. Local regression tests use the real document service and an embedded test database. ## What Changed - Remove credential-content vetoes from Rust admission and both TypeScript semantic dispatchers. - Preserve original structured approval and question arguments during execution. - Keep transport bounds, schema checks, authorization, idempotency, and audit masking. - Require explicit credential syntax or recognized formats for diagnostic masking. Preserve ordinary prose, metadata, and dotted identifiers. - Test exact document persistence, replay, nested argument identities, and masked audit copies. - Remove obsolete retry guidance and document harness-owned credential policy. ## Verification - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --lib --test acpx_event_payload --test acpx_provider_state --test acpx_provider_turns`: 355 tests passed. - Focused server and adapter tests: 189 tests passed after rebase. These include the real document save and the complete tool-gateway suite. - Semantic dispatcher and conformance tests: 34 tests passed. - Diagnostic redaction and MCP tests: 46 tests passed, including all six review examples. - `pnpm -r typecheck` and `pnpm build` passed on the repaired branch. - The broad local root suite was interrupted after database fixture setup failures. The focused database suites passed. CI runs the complete configured test lanes. ## Risks - Authorized tool arguments can intentionally contain credentials. The harness must enforce its content policy. - Diagnostic detection is narrower. Explicit assignments, credential fields, and recognized credential formats remain masked. - The change does not add a database migration or change company authorization. ## Model Used - OpenAI GPT-6 through Codex. The session exposes the GPT-6 model family; its exact runtime model identifier and context window size are not exposed. Used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
839cac1343 |
fix: request supported offline access for generic MCP OAuth (#14950)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use external MCP tools through the governed gateway. > - Remote MCP connections can use OAuth access tokens that expire. > - Some providers issue refresh tokens only after an offline-access request and consent. > - Resource scopes hid that identity-provider capability in the generic connection flow. > - This pull request requests supported offline access and tests expiry through a local MCP server. > - The benefit is continued tool access without another sign-in when the provider permits refresh. ## Linked Issues or Issue Description Related PR: #13447 addresses the same OAuth symptom together with managed Codex configuration. This PR focuses on generic MCP OAuth. It also covers consent, explicit scope overrides, legacy reconnects, exact scope persistence, and real HTTP expiry tests. **What happened?** A generic MCP resource can advertise only its tool scopes. Its OAuth server can separately advertise `offline_access`. Paperclip selected the resource scopes and omitted the offline-access request. A provider could then issue an access token without a refresh token. Tool access stopped after the access token expired. **Expected behavior** Paperclip adds advertised offline access to the selected tool scopes when the OAuth server does not exclude refresh tokens. It requests consent and stores the scopes sent in the authorization request. Existing connections can discover this capability when the user reconnects. Providers without this capability keep their existing scope behavior. **Steps to reproduce** 1. Run `node scripts/mcp-fixtures/servers/oauth-refresh-fixture.mjs` from the repository root. 2. Add its MCP URL as a generic connection on a local Paperclip instance. 3. Approve the test consent page and call `read_status`. 4. Let the two-minute access token expire and call the tool again. 5. Before this fix, the connection needs another sign-in. With this fix, the call refreshes the token and succeeds. **Paperclip version or commit** The integration regression reproduced the missing-refresh-token failure on the parent of this PR's fix. The same test passes with the fix. **Deployment mode** Local development. Automated tests use a loopback HTTP MCP/OAuth server and a disposable PostgreSQL database. ## What Changed - Track offline-access capability separately from MCP tool scopes. - Add supported offline access and consent for generic connections. - Preserve the actual requested scopes through callback completion and reconnect. - Discover the capability for older connections with cached OAuth endpoints. - Add a reusable MCP/OAuth fixture with PKCE, token expiry, resource binding, and refresh-token rotation. - Test shared and personal gateway calls through two refresh rotations. Cover scope selection, unsupported refresh, and legacy reconnects. - Document the behavior and local test commands. ## Verification - The two real HTTP expiry tests failed before the fix with `oauth_refresh_missing` after the first token expired. - Tests passed on the current head: 76 generic MCP regressions, all 365 tool-access service tests, and 4 fixture controls. - Full workspace `pnpm -r typecheck` and `pnpm build` passed. Server TypeScript checks also passed after the review fixes. - All remote checks passed on `d22909bb34e9d54478c0002077c498ffe105932d`. Greptile gave 5/5 with both previous findings resolved. The complete local `pnpm test:run` is still running. - Run `node --test scripts/mcp-fixtures/servers/oauth-refresh-fixture.test.mjs` for the standalone provider controls. - Run `pnpm exec vitest run server/src/__tests__/generic-mcp-connection.test.ts` for the Paperclip integration tests. ## Risks - Users can see a consent prompt when a generic provider supports offline access. - The provider can still decline to issue a refresh token. Access works until expiry, then the user must reconnect. - A provider that advertises offline access but rejects the scope produces an OAuth error. This PR does not add an automatic retry without that scope. - Existing grants without refresh tokens need another sign-in. The fix does not change them in place. - Curated Apps keep their reviewed scope and authorization-parameter allowlists. No database migration is required. - This simulation verifies the suspected failure. The reported internal MCP server has not been tested. ## Model Used - OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The runtime did not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7a52dcdc74 |
fix: repair MCP validation and cancelled execution recovery (#14951)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The tool gateway gives agents access to connected services. Recovery controls what happens when a run stops. > - Generated tool names can exceed the provider limit after the MCP client adds its prefix. > - The same invalid definition can fail each automatic retry. A cancelled run can also hold saved messages without showing its cause. > - This pull request bounds tool names, stops configuration retries, and retains cancellation evidence. > - It shows the stopped run and admits saved input only after the existing safety checks pass. > - The benefit is a clear recovery path that preserves operator Stop and prevents duplicate message delivery. ## Linked Issues or Issue Description **What happened?** A long connected MCP tool name makes the provider reject the entire request. Automatic recovery repeats the invalid request. Separately, unexpected legacy cancellations can leave saved input behind a recovery hold. The notice does not identify the stopped run or its cause. **Expected behavior** Complete MCP names fit the provider limit. Tool-definition errors require configuration repair. Cancelled runs retain their source and reason. The recovery notice shows the cause and saved-message count. Verified unexpected cancellations can start a fresh turn through the existing admission checks. **Steps to reproduce** 1. Assign an App gallery connection with a long application key and tool name to a Claude agent. 2. Start a run. The provider rejects a name over 128 characters, including its MCP prefix. 3. For cancellation recovery, stop a legacy provider turn without an operator Stop request and send a user message while the recovery hold is active. 4. Inspect the recovery notice and the deferred message queue. **Paperclip version or commit** Rebased onto master at `cf8ad63c806685bfd7c48e3ed4a919d61a7c55f1`. **Deployment mode** Hosted or self-hosted server with legacy Claude or Codex execution. Related public work: - Refs #14017. That PR caps name segments. This PR preserves existing short names and uses stable hash aliases for long complete names. It also covers classification and recovery. - Refs #4510. That PR adds a cancellation-source column. This PR records bounded evidence in the existing run result, without a migration. - Refs #12552 and #4506. Those PRs suppress recovery after operator cancellation. This PR preserves operator intent and uses the existing continuation gates. ## What Changed - Bound gateway names with the full provider prefix in the 128-character budget. Retain the original upstream tool name for dispatch and permissions. - Classify invalid tool definitions as configuration failures before diagnostic redaction. Stop automatic retries and continuation attempts for that error code. - Persist cancellation source, expectedness, initiator, reason, and time. Preserve recorded Stop intent when adapter results arrive. Report unexpected started cancellations with closed diagnostic labels. - Show the run cause, saved-message count, and Inspect run link. Offer Continue for eligible unexpected cancellations. Require verified provider stop, empty tool inventory, ownership, and the existing pause, budget, approval, and dependency gates. Use the existing queue for single delivery. - Add regression coverage and update the execution, MCP gateway, and run-log documentation. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - `pnpm check:token-gates` passed. - Ran `pnpm test:run` and completed its workspace and serialized groups. Initial resource and timing failures passed on isolated reruns. All 149 serialized route suites passed. - Reran the changed server, adapter, and UI suites after the rebase. Coverage includes long-name upstream dispatch, configuration retry suppression, cancellation evidence retention, privacy labels, oversized run projection, and concurrent saved-message delivery. - `pnpm test:e2e tests/e2e/legacy-failure-continuation.spec.ts` passed all six browser scenarios. The recovery notice shows the run cause and inspection link, and each recovery entry point reaches one new response. - Added database-backed checks for active, removed, paused, unavailable, and disabled chat connections. The final continuation and recovery-notice suites passed 167 tests. Externally bound chats hide board Continue and show a usable next action. - All 55 GitHub checks passed on `42afbf1371dcaeb72646e3d8f65c19ff7cddf8de`. Two unrelated Storybook jobs were skipped by their normal conditions. Greptile reviewed that commit at 5/5 with no findings and no open review threads. ## Risks - Long tool names change to aliases. Existing short names stay compatible. The original connection and upstream name remain the dispatch authority. - Invalid tool definitions no longer get automatic retries. An operator must repair the configuration before a new attempt. - Continuation changes apply only to positively identified unexpected legacy cancellations with complete empty tool inventory. Operator Stop, unknown historical cancellations, outstanding tools, and unverified provider termination keep their holds. - No database migration. The added projection fields are optional. Cancellation reason and initiator IDs remain local run evidence; Sentry receives only closed source and initiator-type labels and expectedness. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository editing, shell execution, and GitHub tool use. The runtime does not expose the exact model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7d59de6113 |
feat(connections): probe provider usage limits on demand (#14936)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections store the AI accounts used by legacy and native runners. > - Subscription accounts can reach session, weekly, model, or paid usage limits. > - Operators need to read these limits for a specific stored account before making a routing decision. > - This pull request adds an on-demand usage probe to the connection service and account detail. > - The result preserves provider limits, reset times, paid usage, and unknown values for later consumers. ## Linked Issues or Issue Description **Subsystem affected** Shared contracts, the connection service and API, and the account detail UI. **Problem or motivation** Managed AI accounts lack a common operation to read their current usage limits. A local harness probe can read a different login from the account selected for an agent. **Proposed solution** Add `aiConnectionService.probeUsage()` and a board-only connection usage endpoint. Probe the selected credential grant on request. Support Codex, Claude, and Grok subscriptions, plus OpenRouter API key limits. **Alternatives considered** Harness-specific automatic polling would couple the read to execution and can read ambient credentials. This change uses the managed connection credential and leaves scheduling and admission decisions to later work. **Roadmap alignment** This extends the existing Personal & Shared AI Accounts capability. It adds no routing or quota enforcement. Related: Refs #14459 for managed OpenAI quota reads; Refs #14781 and Refs #13379 for downstream pacing and budget work. This operation reads one requested account across all three subscription providers. ## What Changed - Add typed usage snapshots and a probe capability flag to managed AI connections. - Normalize Codex, Claude, Grok, and OpenRouter responses. Keep model scopes, provider admission, reset periods, and paid allowances separate. Preserve unknown values. - Enforce company membership, credential audience, grant identity, and connection lifecycle before reading the stored secret. - Add a board-only `GET /api/companies/:companyId/ai-connections/:connectionId/usage` endpoint with `no-store` responses. - Add manual **Check usage** and **Refresh** actions to account details. Show compact usage bars, resets, admission and overage status; remove repeated descriptions and account-default copy. Clear previous results during a new request or error. - Add Storybook previews using the production account components for all four providers, initial checks, loading, and permission errors. - Add provider, authorization, runner selection, API, and UI coverage. Document provider sources and live qualification. ## Verification - Initial provider, authorization, selection, API, and UI validation passed (96 focused tests): `pnpm exec vitest run server/src/services/ai-connection-usage.test.ts server/src/__tests__/ai-connections.test.ts ui/src/components/ai-connections/AiConnectionUsagePanel.test.tsx server/src/__tests__/openapi-routes.test.ts`. - `pnpm -r typecheck` passes for the initial implementation. After simplifying the UI, 9 usage-panel and date-helper tests, UI typecheck, token gates, and Storybook build pass. The initial feature module boundary check also passed. - Real Codex, Claude, and Grok credentials were saved to encrypted disposable connections. The actual usage HTTP route returned 200 with `status: ok`. Legacy and native runner selection checks passed. The tests started no model turn and exchanged no refresh token. The disposable databases and vaults were removed. - Live Claude responses added structured scoped limits. Live Grok responses omitted included-plan usage. Tests now cover both shapes and preserve the Grok omission as unknown. - The full workspace build passes. A full local test run hit a heartbeat feedback timeout. That case passes in isolation. The duplicate local run was stopped after all remote checks passed. The Slack ordering and OpenCode transport CI flakes also pass in isolation and on the CI rerun. - Current head: `ff3d479029a1c4248190323e221b2803cfb0d79d`. All 54 active checks pass. Two Storybook checks are intentionally skipped by the workflow. Greptile is 5/5 with no unresolved review findings; the branch is mergeable. ## Risks - Subscription usage endpoints can change. Credentials can lack usage-read permission. The probe returns explicit errors without fresh limits in these cases. - A successful probe can contain partial data. Missing utilization or admission remains unknown. An enabled paid-usage switch does not prove a funded balance. - This change adds no migration. It does not change runner admission or automatic provider selection. Provider requests use fixed endpoints, disabled redirects, bounded response sizes, and a 15-second deadline. ## Model Used OpenAI Codex, GPT-6, with reasoning, file editing, shell execution, and HTTP tools. The session does not expose the exact runtime model variant or context window size. Real provider credentials were used only for the authorized live checks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6c1a75da49 |
feat(connections): make AgentMail a default connection with inline setup (#14772)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents access to external services. > - AgentMail needs both a saved key and an inbox assigned to the agent. > - Chat requests offered a setup link instead of an inline card and could treat a saved key as complete. > - Inbox setup also hid address conflicts behind a generic server error and a separate review step. > - This pull request makes AgentMail a default connection, adds the inline card, reduces setup to two steps, and shows conflicts beside the address. > - Shared native dropdown styles also give every caret a consistent inset. ## Linked Issues or Issue Description **What happened?** AgentMail requests in chat did not show a usable inline connection card. Manual setup required extra screens, ignored saved account keys, and could trap new-address setup in a locked inbox dropdown. Agent selectors omitted the avatar from the selected value. A taken address could produce an HTTP 403 from AgentMail and appear as an internal server error. Native dropdown arrows also touched the right edge of their fields. **Expected behavior** Make AgentMail available as a default connection. Ask for the API key inline, with a direct link to its provider page. Default human access to the company and agent access to the requesting agent. Resume the agent only after an assigned inbox is active. Manual setup should ask for an agent and email address, then finish. Address checks should run as the user types. Taken addresses should show clickable alternatives. A domain dropdown beside the name should prefer a verified custom domain. Setup should suggest authorized saved AgentMail keys and show agent avatars in the picker and selected value. **Steps to reproduce** 1. Ask an agent to connect AgentMail when it has no assigned inbox. 2. Check that an inline API-key card appears and links to the provider's API-key page. 3. Open AgentMail setup, choose an agent, and request an address that is already taken. 4. Correct the inline error, refresh, and finish setup with the same request ID. 5. Inspect native dropdown carets in light, dark, disabled, and right-to-left states. Uses the bounded provider-error parser merged in #14768. Related work: #13256 introduced AgentMail; #14725 expanded connection search. ## What Changed - Stop recurring email queries for tasks that have no email thread. Share the query between the thread provider and activity view. Keep email-task updates and invalidation-based discovery. - Make AgentMail available without the experimental chat setting. Keep the catalog, setup and management routes, agent Channels tab, task email feed, receiving worker, and agent tools available by default. Other experimental chat providers stay gated. - Make the email address and copy icon a single clickable action with the shared Copied! confirmation. Add View inbox linking directly to the matching AgentMail console inbox, with the address encoded as one URL path segment. - Reorganize inbox Settings around the copyable email address, usage instructions, and receiving status. Move reconnect credentials into a disclosure and separate the Disconnect action. Add production Settings stories for active, paused, unassigned-address, revoked, webhook, long-address, mobile, and reconnect states. Show repair controls when the inbox has an error. Keep usage instructions tied to an active inbox with an address. - Add AgentMail channel intents and an inline key field with the direct API-key URL. - Keep setup and retry state tied to the interaction. Require an active inbox for completion. Preserve company and agent access checks. - Reduce manual setup to agent selection and email selection. Put the domain dropdown beside the address and default to a verified custom domain. Preserve explicit choices across reloads. Keep receiving settings under Advanced options. - Check the initial address and edits after a 350 ms pause. Abort superseded requests and ignore stale responses. Show clickable suggestions and retain known creation conflicts across reloads. - Add a company-scoped, manager-only address check using the saved credential. Search the visible inbox list instead of fetching an uncreated inbox: live AgentMail retains negative lookups that can break subsequent access-key creation. Unlisted addresses remain unknown; creation is authoritative. - Suggest labeled saved AgentMail keys in both manual setup and the inline card. Filter by company, provider, active credential, and current-user grants on the server. Prefer an account key and preserve the selected key or an explicit new-key choice across refresh. Use verified scope metadata and bounded concurrent checks for legacy keys. Never return secret values. - Catch an inbox-only key before the email step. Allow its existing inbox only after an explicit choice. Recover old locked drafts at the key picker. Save the replacement key before retiring an empty draft, then use a new setup URL so refresh preserves the switched account; stop if cleanup fails. Preserve already allocated addresses and their original accounts. - Use the shared AgentSelect in email setup. Show the canonical agent avatar in each option and the selected value, including other consumers of the shared component. Add regression coverage for legacy and current Lucide agent-mention icon formats. - Start each catalog Add connection with a fresh setup identity. Honor Finish setup's exact draft/account/address instead of resuming an unrelated browser draft. Return Cancel and Done to Connectors and Email settings to the inbox. Group the task/thread explanation in a How it Works card. - Route AgentMail catalog removal through the email inbox control API, including unfinished drafts. Refresh both the catalog and inbox views. - Render each inbox management tab separately. Access uses the saved account grants and agent controls; Conversations and Activity use the shared persisted email feed. Activity lifecycle actions use the email API. Reconnect returns to inbox Settings. Conversation failures show a retry instead of a false empty state. Email delivery recovery stays in the task. - Map documented provider address conflicts to a field error. Preserve actionable messages for other failures. - Preserve non-secret draft fields across refresh, scoped to the requested agent. Never save API keys in browser storage. Resume partial inbox creation with the original agent, address, and request ID. - Show an already-created address with explicit retry and new-address recovery instead of locked inputs. Preserve the original inbox and resumable draft when choosing another address. Distinguish runtime-key 404 errors and log safe provider status/operation/code. - Apply final agent access once within email setup authorization for a new account whose original installs are unchanged. Preserve later permission edits and reused account installs. Support in-place retry of progress loading. - Let a failed inline setup change keys after retiring an empty draft. Persist its replacement setup identity without storing secrets. Recover a server-saved account when refresh interrupts the save response, while preserving intentional account changes. - Render the production setup in Storybook and add error, recovery, and mobile states. - Inset native select carets in shared CSS. Preserve custom icons, listboxes, keyboard behavior, and forced-color controls. - Add browser regression coverage and an AgentMail Product E2E case with persisted-state and rendered-card evidence. ## Verification - Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `git diff --check` passed after the default-availability change. - All 485 focused tests passed. These cover setup, management, catalog and route gates, connection intents, email authorization, Cursor execution, and the OpenAPI contract. All 39 email integration tests run with the experimental chat setting off. - The shared polling change passed four behavioral tests, UI typecheck and build, and token gates. - `tests/e2e/agentmail.spec.ts` passed with the actual server setting off. This full-stack browser test uses simulated provider responses. It covers catalog entry, saved keys, editable address and domain controls, creation, conflicts, retry, all management tabs, clipboard feedback, the provider link, and task email rendering. - In the live local browser, Add connection reached the editable email step with the saved account key. The verified custom domain was selected by default. Both domain choices worked. The existing inbox Settings page remained available. Both active inboxes completed new mail checks with the setting off. No new provider inbox or email message was created for this pass. - Earlier live provider acceptance covered creation on a verified custom domain, Finish connecting on the reported draft, successful mail checks after refresh, and catalog removal of disposable draft and active connections. Clicking the email address copied the exact address and showed Copied!. View inbox opened the same inbox in AgentMail’s console. No email messages were sent. - Production setup and Settings Storybook builds and interactions passed. Settings states include active, paused, unassigned, revoked, webhook, long-address, mobile, and reconnect. Receiving and revoked-access stories had zero accessibility violations. - Full local `pnpm test:run` on an earlier revision completed with 14,709 passing, 87 skipped, and four transient failures. All four failed cases passed in focused reruns without product changes. That serial full local command was not repeated after each follow-up. The latest-head full CI suite is the final test gate. - CI found an obsolete browser assertion that hid every channel when the flag was off. Updated it to keep AgentMail and the Channels surface visible while preserving the GitHub chat route gates. All 11 provider browser tests passed locally after scoping the Channels selector to the agent sidebar. Two initial local attempts stopped at temporary Postgres initialization. The passing run used a separate disposable database on the existing local Postgres server; it was removed after the test. - Updated the remaining sidebar and aggregator discovery assertions for default AgentMail availability. Ordinary task fixtures now return no email thread. All 128 sidebar/task-page tests and all 42 aggregator tests passed locally. - Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI passed, with 54 successful checks including Snyk and two intentional Storybook skips. The CI run is https://github.com/paperclipai/paperclip/actions/runs/37020833647. A fresh Greptile review scored 5/5 with no unresolved threads. Live model evaluations and inbound/outbound email delivery were not run. ## Risks - AgentMail no longer needs experimental opt-in. Setup still requires a human to connect an account and assign an inbox. Inline setup creates an inbox after a human submits a new or saved key. Company access, agent access, inbox assignment, and completion checks remain enforced. - AgentMail read APIs cannot prove global address availability. The visible-list check is bounded to 100 entries and cannot see inboxes outside the key’s scope. The UI reports this limitation, suggests alternatives without claiming they are free, and keeps final creation conflicts inline. Lookup outages show an error without preventing the authoritative creation attempt. - Native select CSS affects the whole app. Custom-icon selects and multi-row lists are excluded. Forced-color mode keeps the browser caret. - Saved-key discovery uses stored verified scope metadata and checks authorized legacy credentials concurrently within a shared three-second deadline. Provider outages mark legacy choices unavailable; users can still enter another key. Final use rechecks authorization and provider access. - No database migration or transport default change. Live connection remains the default. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The exact served model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ec3bacc9bd |
fix(chat): hide ignored provider information (#14929)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task and agent chats show agent progress and problems that need attention. > - Codex also sends account, skill, and unrelated thread notifications. > - The runner correctly ignores that information but reports it as a warning. > - Chat then shows an internal diagnostic as an actionable provider notice. > - This pull request keeps the diagnostic in run logs and removes it from chat. > - Real provider warnings, errors, and agent replies remain visible. ## Linked Issues or Issue Description **What happened?** Chat showed “Received a provider update” and a warning with the text “ignored unrelated provider information”. Its details said “User Actionable: Yes” even though no user action was needed. Saved conversations retained the same noise. **Expected behavior** Keep ignored provider information in the run log. Do not show it as chat activity or a user warning. Preserve real warnings and errors. **Steps to reproduce** 1. Start a conversation with the native Codex runner. 2. Have the provider send an account update, skill change, or unrelated thread notification during the turn. 3. Inspect live chat and reload its saved history. The regression tests also reproduce the old stored notice without a live account. **Paperclip version or commit** Source implementation on master at `e00d10d5d`. The duplicate search found no open PR for this fix. Related prior work: #13109 improved provider-notice presentation. #12367 added Codex thread normalization. This change addresses the internal information that those paths still projected as chat warnings. **Deployment mode** Native Paperclip Runner with the Codex app-server provider. The issue was seen in hosted chat and can be reproduced with local provider fixtures. ## What Changed - Map ignored unrelated Codex information to `harness.diagnostic` in the Rust and TypeScript normalizers. - Retain a bounded allowlist of redacted provider method and thread/turn identifiers. - Use the same Unicode character limit and truncation marker in both normalizers. - Share the text redactor through a pure helper. Keep provider connection code out of the standalone demo's source closure. - Omit that diagnostic and the matching legacy notice from live chat. - Omit the matching legacy notice from saved chat history. - Test diagnostic retention, account-notification integration, live and saved chat, and continued visibility of real warnings, errors, and replies. - Document the local run-log event and historical display behavior. ## Verification - Passed: 68 tests in the two affected UI transcript suites. - Passed: 60 TypeScript tests across provider events, transport behavior, and the standalone demo boundary. - Passed: 13 Rust provider-event tests and the Codex account-notification integration test. - Passed: `pnpm check:token-gates` and Cargo formatting checks. - Passed: full `pnpm build` and `pnpm -r typecheck`. After the review fix, the provider package build, typecheck, and both provider-event suites passed again. - Full local `pnpm test:run` failed: 608 files / 10,904 tests passed, 30 server suites failed, and 104 files / 4,012 tests were skipped. Most failures were embedded PostgreSQL startup errors. Two tests timed out in `heartbeat-comment-wake-batching` and `workspace-git-snapshot-streaming`. PostgreSQL startup also failed in `heartbeat-run-event-sequencing` and `native-finalization-migration`. These server files are unchanged by this PR. Isolated heartbeat reruns were skipped locally. The stable test script stopped after this general-server group, so later groups did not run locally. - The original review thread is resolved. Greptile is 5/5 on current head `683dab7cce57187c57e84c83f5e9da4ad75c9c04`. - All current-head CI gates passed, including the full server/chat/workspace test matrix, Rust and TypeScript runner suites, browser E2E, build, typecheck, and release canary. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37021330663). - Replay the exact old warning in either transcript adapter. It must produce no chat row. A genuine provider warning or error must still produce a row. ## Risks - Low risk. The display filter matches one diagnostic code or the complete legacy warning shape. Other provider notices remain visible. - New ignored-information events use the existing harness-diagnostic event type. They retain diagnostic evidence without original account payloads. - No database migration, API permission, provider execution, or recovery behavior changes. This affects the local run log, not Telemetry or OpenTelemetry exports. ## Model Used OpenAI Codex, GPT-6. The exact backend model ID and context-window size are not exposed in this session. Used reasoning, repository inspection, code editing, tool use, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected tests locally and they pass (the broad local run has PostgreSQL startup errors and timeouts documented above; the full CI matrix passed) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e00d10d5d5 |
fix(connections): repair stale AI defaults from agent settings (#14916)
## Thinking Path > - Paperclip manages AI agents and controls the credentials used for their work. > - Managed AI connections resolve each responsible user's provider default. > - Agent settings created another account but kept the old default selected. > - A rejected provider test left the old account marked as connected. > - Claude ACP reported a typed login failure as a generic terminal-access error. > - This pull request repairs the selected account or selects the new login explicitly. > - Agents can save and run with the repaired credential, and failed logins request sign-in. ## Linked Issues or Issue Description - Fixes #14831. - Refs #13867. Environment failures remain separate from credential-health failures. ## What Changed - Add an agent-settings action to reconnect an unavailable personal default in place. Keep its connection, grant, default, and agent access. - State that a new account becomes the user's provider default. Select its returned grant before changing the agent binding. Keep the actual sign-in method. - Show default-update errors and allow retry without another provider login. - Show the agent-access choice. Connection managers start with company-wide access for their own tasks. Other members start with access for the current agent. - Use the server's connection-manager permission in the shared list response. This includes members with a custom management grant. - Mark credentials as needing attention after an explicit login rejection in Test or Save. This includes API-key 401 and 403 responses. Network, quota, and server failures keep the credential health unchanged. - Reuse the credential-generation check so an old failure cannot invalidate a newer reconnect. - Route Claude's typed provider `access` failure to the existing login-recovery flow. Replace its generic terminal-access fallback with a sign-in message. - Add regression tests and update the AI Connections documentation. ## Verification - Red: the UI tests failed on the missing reconnect action, unused returned grant, missing access choice, and lost default-update error. The server tests failed because rejected credentials stayed connected. The real ACP fixture returned `acpx_turn_failed` for typed login failures. - Green: 156 tests passed across the AI connection, hiring, agent field, and New Agent suites. All 37 environment-route tests passed. The Claude ACP authentication fixtures also passed. - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - The full local `pnpm test:run` passed 707 files and 14,503 tests, then exited with an agent-conversation timeout and embedded PostgreSQL startup failures in unchanged suites. The isolated conversation and migration tests passed on rerun. Later local test groups did not run after this failure. - [All CI gates passed](https://github.com/paperclipai/paperclip/actions/runs/37012669356) on commit `38513dfe2`. This includes the full test matrix, browser tests, typecheck, build, Runner checks, and canary dry run. - Greptile reviewed commit `38513dfe2` and returned 5/5 with no open findings. - The regression tests use a real embedded database and a real ACP fixture process. Live provider sign-in requires a valid account and was not run. ## Risks - Connecting a new account from agent settings changes the user's provider default. The dialog states this before sign-in. - The displayed access choice can allow all company agents to use the account for its owner's tasks. Reconnect keeps the existing access. Server permissions still control installs. - Claude's typed `access` category maps to the provider's `auth_required` signal. Tool and workspace request failures retain their existing classification. - No database migration or provider credential format changes are required. ## Model Used - OpenAI GPT-6 through Codex. The exact served model identifier and context window are not exposed in this session. Capabilities used: reasoning, repository tools, code editing, and command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
408f70e69f |
fix(runner): preserve stock Codex base instructions (#14920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner connects Paperclip tasks to Codex app-server. > - Paperclip passed its runtime context as `baseInstructions`. > - That field replaces the stock Codex base prompt. > - This pull request sends the same Paperclip context as additive developer instructions. > - Codex keeps its stock prompt and still receives Paperclip task instructions and tools. ## Linked Issues or Issue Description **What happened?** The native Codex driver and Rust provider sent Paperclip context through `baseInstructions` on thread start and resume. Codex used this text in place of its stock base instructions. Direct-chat resume also sent an empty replacement base. The Runner Lab session path used the same replacement field. **Expected behavior** Codex should retain its stock base prompt. Paperclip should add its runtime context through `developerInstructions`. Other provider facades should retain their current instruction handling. **Steps to reproduce** 1. Create a native Codex session through Paperclip Runner. 2. Inspect the `thread/start` request in the native provider trace. 3. Resume the session and inspect `thread/resume`. 4. Before this fix, these paths set `baseInstructions`. After this fix, the Codex paths set `developerInstructions` and omit `baseInstructions`. **Paperclip version or commit** Reproduced against master at `cad26c6bfb736039c8ed5743da650a44792a083c`. **Deployment mode** Built from source. Native Codex app-server and runnerd paths. A local protocol probe used codex-cli 0.153.4 and a localhost Responses stub. No duplicate fix or matching public issue was found in the GitHub search. ## What Changed - Send additive developer instructions on Codex start and resume in the TypeScript driver, Rust provider, and Runner Lab session path. - Carry the additive fragment through runnerd, including runtime asset path mapping. - Preserve existing instruction fields for other provider facades, including OpenCode. - Add start/resume/direct-chat regression coverage and check the actual Rust provider request. - Document the historical option and trace field names. Record progress and follow-ups in the working checklist. ## Verification - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - Targeted Codex driver lifecycle, driver, and live-session Vitest suites — 139 tests passed. - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml --locked -p paperclip-runner-core --test codex_provider` — 91 passed, 2 ignored subprocess helpers. - Real app-server probe: a localhost Responses stub captured identical 14,732-character stock base instructions on fresh start and cold resume. Both requests retained the Paperclip marker in developer input. Both stub turns completed. No paid inference was used. - Runnerd transport Vitest suite — 182 tests passed. - The initial `pnpm test:run` attempt reported local dependency-loading, embedded PostgreSQL startup, and macOS `/var` versus `/private/var` path failures. It was stopped after those failures. Loading-suite reruns passed 1,428 tests after the build; native interaction/finalization reruns passed 38 tests. A seven-suite diagnostic rerun passed 463 tests and isolated the remaining path and PostgreSQL setup failures. - With `TMPDIR=/private/tmp`, workspace, gateway, interaction, and attachment suites passed all 356 tests. The remaining environment-image and native-session-resumption suites passed all 44 tests with the same canonical temp path. All affected suites passed on rerun. The original full local command was stopped after failures and is not claimed as passing. - All 55 PR checks passed at `83281439456181396f3707eecda5d2ebc90bd14d`. Greptile scored 5/5 with no open review threads. - No paid live campaign or Product E2E browser suite was run. This change has protocol and regression coverage; it does not claim improved task quality. ## Risks - Stock Codex behavior may differ from behavior under the previous Paperclip replacement prompt. Restoring that behavior is the intended change. - Existing Codex threads retain their saved replacement base prompt. They need a provider session reset to receive the stock base. This PR does not reset active sessions or alter recovery rules. - The legacy `baseInstructions` option and trace field names remain for compatibility. They now describe the additive Paperclip fragment for Codex. - The separate Codex-through-ACP dependency patch remains a follow-up in the harness coverage checklist. This PR covers native app-server execution. ## Model Used OpenAI Codex, GPT-6. The exact runtime model variant and context window are not exposed in this session. Used reasoning, repository inspection, code editing, shell execution, and test tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c46e41e81c |
fix(heartbeat): validate native MCP gateway ownership (#14914)
## Thinking Path > - Paperclip lets people manage agents and govern their tool access. > - Native runs receive an immutable MCP tool assignment for one agent. > - The gateway must enforce that owner when it authenticates a run token. > - Older gateway rows stored the owner only in metadata. > - This change validates the gateway and profile, binds new rows, and repairs valid older rows on reuse. > - It also delivers each native assignment once. > - Other agents cannot use the assignment, and explicit shared gateways keep their configured scope. ## Linked Issues or Issue Description Builds on #14012 by @busla (Jón Levy). That PR adds agent binding and seven regressions. This PR carries that fix onto current master and adds legacy authentication, profile validation, and duplicate-delivery coverage. Related: #14864 improves discovery memory use. **What happened?** Native gateway creation stored an owner in metadata but left `agentId` null. Authentication could therefore accept another agent's run token. Managed discovery could also deliver historical native assignments again. **Expected behavior** A native assignment accepts only its owner's run token. The gateway and profile must refer to the same immutable assignment. The current assignment enters the run configuration once. **Steps to reproduce** 1. Run the database fixtures in `heartbeat-runtime-mcp-servers.test.ts` on the baseline. 2. Create a native assignment and inspect its stored gateway owner. 3. Authenticate with another agent's run token, then inspect legacy reuse and managed delivery. 4. The baseline fails six ownership and delivery cases. The fix passes all twelve cases. **Paperclip version or commit** The red baseline is `f2e0f1963`. This PR is based on `cad26c6bf`, which includes the merged discovery fix. **Deployment mode** Native Paperclip Runner execution and managed Codex MCP delivery. Reproduction uses isolated database and HTTP fixtures. ## What Changed - Store the agent owner and agent context on new native gateways. - Validate profile and gateway assignment metadata before reuse or token creation. - Bind valid legacy rows with a company-scoped, null-owner update and validate the result. - Reject mismatched run tokens before legacy repair. - Identify native assignments by gateway metadata, the reserved profile key, or profile source. Reject missing or malformed provenance, including JSON null. - Exclude historical native assignments from managed gateway delivery. Keep their rows for existing runs. - Add twelve database and HTTP regressions and document the runtime contract. ## Verification - Red baseline: six regressions fail and four controls pass before the initial fix. Two additional regressions reproduce metadata-loss admission and a JSON-null TypeError before the review fix. - All twelve ownership regressions pass on the final code, including owner admission, cross-agent rejection, metadata loss, JSON-null HTTP 401, and explicit shared-gateway admission. Policy, listing-memory, and discovery HTTP coverage also passes. - Full workspace typecheck and build pass locally. Server typecheck and compilation pass again after the review fix. The final ownership and grant patches pass 42 combined database and HTTP regressions. - [Full CI](https://github.com/paperclipai/paperclip/actions/runs/37011383657) passes for `626a08ae66361cf586105877e24d806b1a7a9c20`: all 54 checks succeed; two optional Storybook checks skip. This includes full typecheck, build, all test shards, all eight E2E shards, runner verification, and the canary dry run. - Greptile scores that exact head 5/5. No review threads remain unresolved. ## Risks Invalid historical native gateway or profile metadata now rejects authentication. Valid unbound rows are repaired only when their owner reuses the assignment. Conflicting owners are never overwritten. Historical rows are retained for existing runs. Explicit shared gateways use ordinary profiles and keep their configured scopes. The reserved native profile namespace remains agent-owned even when gateway metadata is cleared. No schema or dependency changes are included. ## Model Used Original fix and seven regressions in #14012: Anthropic Claude Opus 5.5, `claude-opus-5-5`, 1M context, as reported by @busla. Extensions and verification: OpenAI Codex (GPT-6), with reasoning, repository inspection, code execution, and tests. This session does not expose the exact serving model identifier or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
483dbc8890 |
fix(tool-access): enforce stored grant restrictions (#14915)
## Thinking Path > - Paperclip governs the tools that agents can discover and call. > - Stored grants can limit access to a tool, connection, or application. > - The grant matcher must enforce every restriction in that scope. > - A nonmatching allow list fell through to a policy selector matcher that ignores allow. > - This change requires an explicit allow match and validates additional selectors. > - Malformed and unknown restrictions deny access. > - Discovery and execution now enforce the same stored grant limits. ## Linked Issues or Issue Description Related: #14864 adds the shared database and HTTP discovery fixture used here. Searched existing public PRs for tool grant scope fixes. No duplicate scope-validation fix was found. **What happened?** A stored grant with a nonmatching `scope.allow` could authorize a tool. Empty or malformed allow lists, unknown selectors, and combined mismatching selectors could also authorize access. The fallback policy matcher does not validate stored grant JSON. **Expected behavior** An explicit allow list must match the requested tool, connection, or application. Every additional selector must also match. Unknown or malformed restrictions must deny access. Existing null and empty-object scopes keep their broad grant behavior. **Steps to reproduce** 1. Run `tool-grant-scope.test.ts` on the baseline. 2. Create a deny profile and a grant that names another tool. 3. Attempt discovery or a call for the tool outside the grant. 4. The baseline authorizes access. The fix denies it. **Paperclip version or commit** The red baseline is `f2e0f1963`. This PR is based on `cad26c6bf`, which includes the merged discovery fix. **Deployment mode** The company-scoped MCP gateway. Reproduction uses isolated database fixtures and a deterministic HTTP provider. ## What Changed - Require an explicit allow entry to match the gateway or upstream tool name, connection, or application. - Apply all additional selectors after the allow match. - Reject unknown selectors, invalid value types, empty restrictions, and non-object scopes. - Preserve null and empty-object scope compatibility. - Add sixteen regressions, including discovery, successful execution, and revocation through the HTTP gateway. - Document stored grant scope behavior. ## Verification - Red baseline: seven restricted-scope cases and three malformed-root cases fail. HTTP discovery also exposes tools outside the grant. - All 16 grant regressions and 35 adjacent policy tests pass locally. The HTTP test excludes an ungranted tool from discovery, returns 403 for its call, and verifies that no provider call occurs. It also checks successful execution and later revocation. - Server typecheck passes. The final ownership and grant patches also pass 42 combined database and HTTP regressions. - [Full CI](https://github.com/paperclipai/paperclip/actions/runs/37011177989) passes for `803fa9440111742672c94c4471e5b98f15dd3b97`: all 54 checks succeed; two optional Storybook checks skip. This includes full typecheck, build, all test shards, all eight E2E shards, runner verification, and the canary dry run. - Greptile scores that exact head 5/5. No review threads remain unresolved. ## Risks Stored scopes with unknown keys or malformed restrictions now deny access. Operators must correct those grants before they can authorize tools. Null and empty-object scopes keep their previous broad behavior. There are no schema, dependency, or API changes. ## Model Used OpenAI Codex (GPT-6), with reasoning, repository inspection, code execution, database regressions, and HTTP tests. This session does not expose the exact serving model identifier or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4039d4f06b |
fix(auth): allow scoped low-trust work and owner-chat instruction edits (#14870)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Low-trust agents must work within their assigned scope. > - Task creation currently rejects these agents before checking assignment permission or scope. > - Persistent instruction saves also reject direct requests from authorized chat owners. > - This PR checks the requested action and its recorded authority instead of denying all such work. > - Agents can organize permitted work and follow their owner's instruction-edit requests while outside work stays restricted. ## Linked Issues or Issue Description **What happened?** Low-trust agents cannot create self-assigned tasks or subtasks, even within their allowed scope. An authorized user also cannot ask an agent in their own Agent Chat to update its managed `AGENTS.md`. Agent-folder collection can hide the permission rejection behind a generic save failure. **Expected behavior** Allow task creation when assignment permissions and project or root-task scope permit it. Allow instruction self-edits during authenticated owner-chat execution, subject to the user's current edit permission. Outside tasks, subtasks, connector messages, and peer agents do not inherit that instruction authority. Explain the actual denial when a save fails. **Steps to reproduce** 1. Configure an active agent with `low_trust_review` and a project or root-task boundary. 2. Ask it to create an in-scope task assigned to itself, or a subtask of its own task. 3. As a user with permission to configure that agent, ask it in your Agent Chat to update its managed `AGENTS.md`. 4. Observe blanket permission denials rather than action-specific checks. **Paperclip version or commit** Rebased onto master at `8ec4b84e1`. This is a core authorization change, independent of adapter choice. **Deployment mode** Authenticated server. Regression coverage uses the server services, HTTP routes, native tool authority, and embedded PostgreSQL. Related work: #14775 adds human-directed task execution. #13599 concerns instruction-path configuration; this PR leaves that configuration restricted. #11988 proposes separate active-review instruction protection. #10693 reports unclear authorization denials on a different API surface. ## What Changed - Apply task-assignment checks to both HTTP creation routes and native task creation, including unassigned work. Preserve low-trust policy and source attribution on the created task and its initial plan. - Allow self-assigned decomposition within the permitted project or root-task tree. Resolve workspace-derived project scope before authorization, and reauthorize existing tasks before duplicate detection returns them. Keep cross-project and peer-assignment checks. - Derive instruction self-edit authority from the accepted run identity and authenticated owner-message wake. Recheck current permissions at save time. Bind retries to the same request and chat session. - Reject inherited instruction authority from outside tasks, subtasks, plugins, connectors, stale sessions, cancelled runs, and peer edits. - Surface permission errors in instruction and agent-folder save receipts. Tell chat agents to explain the rejected action and the specific restriction. - Update the low-trust policy and implementation documentation. ## Verification - All 297 tests in 11 focused server suites pass after the rebase. These cover owner-chat saves, private copies, warm agent directories, reset and retry boundaries, permission revocation, task creation routes, and native tool authority. - After review fixes, all 126 tests in the four affected authorization/chat suites pass. Workspace scope regressions and 146 existing creation/ownership/workspace-route tests also pass. - The final duplicate-task and CI fixes pass all 39 tests across chat-project tools, duplicate creation, environment-selection guards, and assignee-invokability routes. The duplicate-task test reproduced an unauthorized response before the fix and verifies denial plus permitted reuse afterward. - `pnpm --filter @paperclipai/server typecheck` passes after rebasing; `pnpm --filter @paperclipai/server exec tsc --noEmit` also passes after the review fixes. - `git diff --check origin/master...HEAD` passes. - Final head `7e73270b86748792649e4ae6fbc6879f73b42b73`: all 54 checks passed, with two expected skips and no pending or failed checks. This includes builds, typechecking, the full test matrix, end-to-end tests, runner verification, the canary dry run, and security scans. - Greptile is 5/5 on that exact head, with no unresolved review threads. This change has not been deployed to staging. ## Risks This changes authorization behavior. The instruction exception must not become an inherited task permission. The check uses server-owned execution records, requires the agent's own chat and instructions, and keeps normal protected-change and responsible-user checks. Saves fail closed when current provenance or permission is missing. Owner chat grants a turn-scoped capability; the server does not classify the message intent or require approval of the exact new file bytes. Prompt injection within an authorized owner-chat turn remains a model-level risk. This is the requested owner-chat trust boundary, without a new per-edit confirmation flow. No database migration or broad trust-preset change is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and test execution. The exact runtime model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8ec4b84e1c |
fix(chat): resume messages after failed runs without duplicate delivery (#14857)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A user can send a new message after a native run fails. > - The server checks that the old execution has stopped before it starts a fresh turn. > - A failed run can retain a result accepted before checkpoint or cleanup failed. > - The continuation gate treated that saved result as active recovery and held the new message forever. > - This pull request removes that false liveness signal while retaining controller, process, environment, and authorization checks. > - Live staging then exposed a second defect: chat admission created a successor without consuming the original deferred receipt, so completion delivered the message again. > - Consume that exact receipt atomically with admission, while preserving separate turns for later chat messages. ## Linked Issues or Issue Description **What happened?** A new user message stayed in the queue with `controller_settling` after the previous run had reached `terminal_failure`. The old coordinator had no lease owner but still had a `resultId`. Its remote environment had a verified stop receipt. **Expected behavior** Start one fresh turn after execution has stopped and normal admission checks pass. Preserve the failed run and its accepted result as history. **Steps to reproduce** 1. Accept a native result, then fail checkpoint or cleanup and exhaust recovery. 2. Retain the result ID on the terminal failure record and stop the execution environment. 3. Send a new user message. Before this fix, it waits forever for the finished controller. **Paperclip version or commit** Reproduced in a database-backed regression test on `26900655b`. **Deployment mode** Server with a native runner and remote sandbox. Local process stop checks also apply. Related: https://github.com/paperclipai/paperclip/pull/14775. Searched existing PRs for retained-result continuation fixes; no duplicate found. ## What Changed - Remove the retained-result veto for terminal failures. - Keep controller ownership, successor, process, environment cleanup, pending decision, and ordinary admission checks. - Add regressions for retained results, active execution, missing stop evidence, and delayed remote cleanup. - Atomically consume the resumed receipt in agent chat, even though chat does not coalesce other queued messages. - Reproduce completion-time duplicate promotion, race cleanup against periodic recovery, and prove a subsequent chat message keeps its own turn. - Document that a saved result does not make a terminal failure active. - Keep exhausted workspace export on its separate repair path, tested through the production finalizer. ## Verification - Red: retained-result admission failed with `controller_settling` before the original fix. The new chat-specific regression then reproduced duplicate promotion when the first reply finished. - Green: 406 tests across native continuation, workspace-export recovery, and the wake-queue module passed on `cbc531cc0`. - The chat regressions exercise real Postgres transactions, simultaneous recovery callbacks, successful completion, the production queue-drain use case, and repeated drain attempts. A distinct follow-up remains a separate turn. - `pnpm -r typecheck` and `pnpm build` passed on `cbc531cc0`. - The earlier full local test run encountered a timeout and follow-on failure in unchanged AI connection-adoption tests; all 50 tests passed on isolated rerun. That local run was stopped after the full CI test matrix passed on the earlier head. - All 54 CI checks passed on `cbc531cc0` (2 skipped), including the full test matrix and browser shards. One unchanged interaction-route test returned HTTP 500 on its first CI attempt; its full 84-test file passed locally, and the failed shard passed on one targeted rerun. - Greptile reviewed `cbc531cc0`: 5/5, no unresolved findings. - Live staging first verified that the original saved message resumes and receives a successful response; that test exposed the duplicate now covered above. - Deployed exact commit `cbc531cc0410e1ef6e8811c6c5c014c3528351ed` to the affected staging workspace; deployment verification, health, authentication, and startup recovery passed. - Submitted a fresh message through the browser. The agent replied in 39 seconds; server records show exactly one successful run, native phase `committed`, no error, and an empty queue. A later check more than a minute after completion found no duplicate run. ## Risks The change affects admission after native execution failure and consumption of a resumed deferred receipt. A fresh turn must never overlap the prior execution, and consuming one chat receipt must not absorb later messages. Tests retain the controller, process, and remote-stop guards. This change does not migrate data, apply an old result, or reset the old retry budget. ## Model Used OpenAI Codex (GPT-6). The exact runtime model identifier and context window are not exposed in this session. Used reasoning, repository inspection, code execution, database-backed tests, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
efc2e6810e |
fix: show each task once in dashboard agent cards (#14847)
## Thinking Path > - Paperclip helps people manage AI agents and their tasks. > - The dashboard shows recent agent activity in compact cards. > - Those cards use run records, so two runs for one task can create duplicate task cards. > - An operator needs to see each task once when scanning the dashboard. > - This pull request selects one run per linked task before it applies the card limit. > - The live runs page still shows each run for run inspection. ## Linked Issues or Issue Description **What happened?** The dashboard showed the same task in two agent cards when that task had both an active run and a completed run. **Expected behavior** The dashboard should show a linked task at most once. It should keep the active run card when one is present. **Steps to reproduce** 1. Start an agent run for a task that already has a completed run. 2. Open the company dashboard. 3. Observe two cards linked to the same task. **Paperclip version or commit** Reproduced on the pre-change master at `8b4aa0692`. **Deployment mode** Local dev, built from source. The bug is in the core dashboard UI and does not depend on an agent adapter or database mode. ## What Changed - Select distinct linked tasks from capped active and recent run samples before applying the dashboard card limit. - Keep separate cards for runs without a linked task. - Preserve the dashboard's count of additional distinct cards behind the live-runs link. - Add UI and embedded Postgres regression tests for duplicate runs and document the dashboard rule. - Give the existing multi-request cross-tenant authorization test enough time on loaded CI runners. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/ActiveAgentsPanel.test.tsx` - `pnpm --filter @paperclipai/ui exec vitest run src/api/heartbeats.test.ts` - `pnpm exec vitest run server/src/__tests__/dashboard-service.test.ts server/src/__tests__/agent-live-run-routes.test.ts` - `pnpm exec vitest run server/src/__tests__/agent-cross-tenant-authz-routes.test.ts` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui build` - `pnpm -r typecheck` - `pnpm build` - `pnpm check:token-gates` - Review the dashboard with an active and a completed run on the same task. Confirm that it shows one card. Open Live agent runs to inspect both run records. ## Risks - A very high volume of recent runs for one task can fill the capped sample and leave older tasks off the dashboard. The Live runs page remains available for full run inspection. - The dashboard may fetch up to 50 distinct run representatives to preserve its overflow count. The default run API response and persisted data are unchanged. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-6. The runtime does not expose the exact model ID or context window size to this task. The model used reasoning, tool calls, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6d654f63d1 |
feat(apps): make MCP action test results readable (#14859)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connected Apps let an operator control which MCP actions an agent can use. > - The Permissions page lets the operator run a real action as an agent. > - The Test dialog displayed the nested MCP response as escaped JSON. > - A useful result was hard to read, even when the action worked. > - This pull request renders known MCP content as a readable preview and keeps the raw response available. > - The benefit is faster validation without losing the data needed to diagnose a failure. ## Linked Issues or Issue Description **What existing behavior does this improve?** The per-action Test dialog on a connection's Permissions page. **Subsystem affected** ui/ — React board UI. **Current behavior** The dialog shows the gateway response as an escaped JSON blob. Text content that contains JSON stays inside a string. The obsolete connection Test page also keeps a separate set of stories. **Proposed behavior** The dialog uses structured MCP content when present. It parses JSON text blocks when possible. It shows compact tables, cards, fields, or plain text. It keeps the full raw response behind a control and opens that view for errors or unknown block shapes. Stories exercise the Permissions page dialog, and the obsolete Test page and stories are removed. **Reason and benefit** An operator can inspect a successful action result at a glance and still inspect the exact gateway response when a call fails or looks wrong. **Breaking changes** No API or stored data changes. The Test dialog presentation changes. The raw response stays available. **Additional context** I tested a read-only Notion search through the real Permissions page. The dialog showed three result cards and the raw response control worked. Storybook uses invented example data. No directly matching public issue or open PR was found in the GitHub search. ## What Changed - Render structured MCP output and JSON text content in the action Test dialog. - Show wide rows as cards, keep short rows as tables, and retain the raw response for diagnosis. - Remove the obsolete connection Test page and its stories. - Add focused dialog tests and Permissions page Storybook cases for success, errors, mixed blocks, and malformed blocks. - Document the Test dialog result behavior in the connection playbook. - Keep agent mention icons visible when the Lucide icon node is unavailable in server rendering, which repaired a repeatable CI failure. ## Verification - `pnpm -r typecheck` — passed. - `pnpm exec vitest run --project @paperclipai/ui` — passed (7,111 tests). - `pnpm exec vitest run ui/src/pages/apps/app-detail/ActionTestDialog.test.tsx` — passed (11 tests). - `pnpm exec vitest run --project @paperclipai/ui ui/src/components/MarkdownBody.test.tsx` — passed (53 tests). - `pnpm test:run` — started, then stopped after the review fixes changed the head; the full sharded suite passed in CI. - `pnpm build` — passed. - `pnpm check:token-gates` — passed. - Use a connected MCP app. Open Permissions, select a read action, and run Test. Inspect the preview and the raw response control. ## Risks - MCP tools can return provider-specific block shapes. Unknown blocks open the raw response so the operator can inspect the exact result. - Row and field previews limit visible data. The raw response preserves the complete result. > This is a targeted improvement to the existing Connected Apps item in `ROADMAP.md`. ## Model Used OpenAI Codex, GPT-6. The session used tool access, code execution, and browser validation. The exact deployment ID and context window were not exposed to the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
527e146980 |
feat(ui): share animated agent setup prompts (#14862)
Use the shared animated prompt-copy control across setup, invitations, webhooks, and task handoffs. Preserve first-click copying, clipboard recovery, and logo continuity. Add Storybook coverage and restore mention icon masks for the current Lucide data shape. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6395cae072 |
fix(runner): ship provider pack in the standard Docker image (#14854)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Remote OpenCode and ACPX runs need a provider pack from the application build. > - Cloud now builds its application image from the standard production image. > - The provider pack was added only to the legacy cloud image target. > - The standard image therefore cannot supply the pack to downstream Cloud images. > - This pull request adds the pack to the production image and lets the cloud target inherit it. > - Remote runs can then use the pack that matches the application source commit. ## Linked Issues or Issue Description Refs #13827. Refs #14024. The standard production image does not include the remote provider pack. Downstream Cloud images inherit that omission. Remote OpenCode and ACPX runs fail with `runner_remote_provider_artifact_incompatible` and ask for `PAPERCLIP_RUNNER_REMOTE_PROVIDER_PACK_PATH`. ## What Changed - Build and copy the provider pack into the standard production image. - Set the pack path and check that an unprivileged user can read its artifacts and execute Node. - Let the legacy cloud target inherit the pack from production. - Add regression checks for production packaging and cloud inheritance. - Document stamped image behavior and the default pack path. ## Verification - The 12 focused Docker stamp and provider-pack reuse tests pass on commit `4a11f8d52aead55f85527c8e82c6d7f2644ce0da`. - The new packaging regression failed against the old Dockerfile and passed with the fix. - On the current commit, `pnpm build` and `pnpm -r typecheck` pass. All seven standard-image contract tests also pass. - The current-head CI build, typecheck, test, browser, and native Runner checks passed. The local full suite hit one chat-channel assertion failure; that exact test passed in isolation. The remaining local run was stopped after CI completed to avoid duplicating its full suite. An earlier run on the pre-rebase base had a heartbeat comment batching timeout; the external chat wait integration suite passed all 142 tests in isolation. - [The stamped preview image build passed](https://github.com/paperclipai/paperclip/actions/runs/36885002850/job/110446106393), including the production-stage provider pack build, copy, and unprivileged artifact readability/executable check. Publication, compatibility validation, and deployment of this exact commit to a staging QA instance passed. - Reproduced the exact missing-pack error on an existing staging image with Paperclip Runner, ACPX, and Claude in a remote Daytona computer. The legacy Claude adapter succeeds with the same account and computer. After deploying this commit, the same native task succeeded: it computed `5050` with a real remote shell command, wrote a proof file, read it back in a separate call, uploaded the file as a deliverable, and completed the task. The uploaded file contents and Done state persisted after a page reload. The run trace confirms Paperclip Runner, ACPX, and Claude. The first run took 2m 59s, including approximately 97s of remote artifact preparation. A second native run read the unchanged file from the prior run and completed successfully. Its startup took about 120s; this verifies repeated execution and file persistence, not fast provider-pack reuse. ## Risks - Stamped standard images now include the provider pack and its build cost. A pack build failure now fails the production image build. - Unstamped local builds still skip pack generation. Setting the path alone does not create a pack. - No database, provider authentication, or runner verification rules change. ## Model Used OpenAI Codex, GPT-6. The exact serving model identifier and context window are not exposed in this session. Capabilities used: repository inspection, code editing, shell verification, and browser testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
33a00d2f1e |
fix(ui): reopen last visited agent chat (#14848)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent Chat keeps one conversation for each agent and board user. > - The Chat sidebar entry opens the agent chooser each time. > - A user must then find and reopen the chat they just used. > - The browser already records recent agent chat visits by company and user. > - This pull request uses that record to reopen the last available chat. > - The chooser still serves users who have no available saved chat. ## Linked Issues or Issue Description Related: #14706 added the secondary Agent Chat navigation. **What happened?** The Chat sidebar entry opened the agent chooser, even after a user opened an agent chat. **Expected behavior** The Chat entry should reopen the last agent chat visited by the current user in the current company. **Steps to reproduce** 1. Enable Agent Chat and open a chat with an agent. 2. Open another page. 3. Select Chat in the sidebar. 4. Observe the agent chooser instead of the chat. **Paperclip version or commit** Reproduced on master at `0829d94af`. **Deployment mode** Local development, browser UI. The change also uses the same browser storage path in authenticated mode. ## What Changed - Use the existing recent chat record when the Chat landing route opens. - Check saved agents against the current roster and chat history before redirecting. - Keep the chooser when no saved chat is available, and show a retry state for load errors. - Add route tests and update the Agent Chat implementation spec. ## Verification - `pnpm exec vitest run ui/src/pages/AgentChats.test.tsx ui/src/lib/recent-agent-chats.test.ts` — 16 tests passed. - `pnpm check:token-gates` — passed. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/agent-chat-sessions.spec.ts --grep 'secondary chat navigation preserves layout'` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed on the final commit. - `pnpm -r typecheck` and `pnpm build` — passed earlier in this branch; latest-head CI completed all 47 jobs successfully. - `pnpm test:run` reported an unrelated native runtime test failure before it was stopped. That test and an unrelated external object refresh test passed in isolation. CI runs the same suites on the PR. - To check in the UI: open an agent chat, leave it, and select Chat. The same chat should open. Clear the recent chat record or use another company to see the chooser. ## Risks - The recent order is stored in the browser. Clearing browser storage returns the user to the chooser. - An existing chat ID is stored with its visit. If the chat is removed, the landing route skips that visit when history loads. Cross-tab storage removal clears the identity; failed writes retain an in-tab fallback. - The landing route waits for the agent roster and validates saved issue IDs against chat history when available. If history fails, an active agent chat can still open; roster or session failures show a retry action. - No database or API contract changes are required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-6 family. The runtime did not expose an exact API model ID or context window. It used reasoning, repository tools, shell commands, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4ac374103f |
fix(connections): repair Asana MCP and add shared-app sign-in (#14756)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let agents use provider tools through the permission
gateway.
> - Asana provides an official remote MCP server, but its v2 server
requires a registered MCP OAuth app.
> - Setup can discover retired v1 endpoints and send a callback that
differs from the displayed URL.
> - This pull request repairs custom app setup and adds sign-in through
Paperclip's shared app.
> - Users can choose their own app without enrolling with Paperclip
Cloud.
> - Agents can use Asana tools after the user connects their account and
sets action permissions.
## Linked Issues or Issue Description
Related: #14739 supplies the personal credential repair used by resumed
Asana setup. No duplicate Asana authentication PR was found.
**What happened?**
Asana setup failed even with a user-created app. Root discovery metadata
still points at v1. MCP v2 uses the Asana OAuth issuer and requires an
MCP app with a client secret. Local setup also displayed a localhost
callback while an Origin header could make authorization use a numeric
loopback callback.
**Expected behavior**
Sign in with Paperclip's app when its broker profile is available. Keep
custom MCP app setup available without Cloud enrollment. Use the correct
issuer, callback, client credentials, and resource throughout setup.
**Steps to reproduce**
1. Open Asana in the connection catalog.
2. Supply an Asana MCP app's client ID and secret.
3. Start OAuth on a local instance opened with a numeric loopback
address, or resume a draft that cached v1 metadata.
4. Observe the wrong discovery endpoint or callback mismatch.
**Paperclip version or commit**
Reproduced from
|
||
|
|
0829d94af2 |
fix(auth): derive low-trust human direction from existing execution records (#14775)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Low-trust review contains work that may include hostile input. > - Its default intake boundary currently blocks direct human chat and tasks outside that boundary. > - Human direction should authorize the assigned work while preserving containment. > - Existing conversations and execution requests already identify direct human instructions. > - This pull request derives exact-task authority from those records and the current assignee. > - The agent can perform that work without gaining access to unrelated tasks or privileged tools. ## Linked Issues or Issue Description **What happened?** A low-trust agent with a project boundary rejects its owner's direct Agent Chat before provider execution. Human-assigned tasks outside that project fail the same check. **Expected behavior** An authorized human can talk to the agent or assign it a task. The exact task runs with its existing sandbox, credential, and tool restrictions. **Steps to reproduce** 1. Enable Agent Chat and isolated workspaces. Configure a sandbox agent with low-trust review scoped to an intake project. 2. Send the agent a direct board chat message, or assign it a projectless task. 3. Observe `low_trust_boundary_mismatch` before execution. Related: #14766 adds private task directories for repo-free low-trust execution. It is now merged into master and included in the branch base, so CI and staging verify the combined behavior. ## What Changed - Derive owner-chat access from existing conversation identity. - Derive exact-task access from the existing human requester and server-owned request origin, including coalesced requests. Plugin and external sender attribution do not authorize work. - Follow existing `retryOfRunId` database links for automatic continuations, checking company, agent, and task throughout; cancelled ancestors cannot grant authority. - Require a live run and current assignment. Preserve sandbox, credential, privileged-tool, responsible-user, and quarantined-output checks. - Retain board backlog assignments in existing request records without starting execution. Reassignment cancels old human requests in the common service transaction, including plugin writes; late settlement cannot revive them. - Add real database and HTTP coverage for request provenance, retry ancestry, cancelled runs, concurrent reassignment, spoofing, and containment. Document the rule. - Preserve legacy board assignment requests through their existing source, reason, and human requester. - Use the existing wrapped-error helper for concurrent chat-question idempotency; a deterministic race test reproduces the CI failure before the fix and passes after it. - No new schema, migrations, or user-identity fields. Existing requester columns hold attribution. ## Verification - Passed the focused database, policy-retention, HTTP, and reassignment tests locally. The HTTP test creates a task through the real board route and checks the resulting persisted wakeup before exercising agent reads, comments, mutations, and review handoff. - Database tests hold a reassignment transaction open to verify coherent authorization before and after commit, with a two-connection pool. They cover retries, coalesced requests, cancelled ancestry, invalid cross-company/agent/task links, cycles, and forged attribution. - Full local `pnpm -r typecheck` and `pnpm build` passed on the final commit (`d107c26df`). [Latest-head CI](https://github.com/paperclipai/paperclip/actions/runs/36815589542) passed: 54 successful checks, two expected skips, including all eight browser-test shards. Greptile is 5/5 on this exact commit with no unresolved threads. Local tests were targeted; the full test suite ran through CI’s test matrix. - The revised HTTP suite passed all 13 tests; database authorization tests passed all 11, including legacy compatibility and late watchdog settlement; the backlog route contract passed all 3 tests. Another 102 tests covering durable chat admission, wake queues, and Cursor execution passed. - All 90 interaction-service tests passed with both create calls deliberately held until their optimistic reads complete, forcing duplicate-key recovery. That forced race failed before switching to the shared wrapped-error helper. - Previous staging proof covered owner chat and projectless task persistence. The simplified revision has not been redeployed; that earlier proof is not claimed for the new implementation. ## Risks - This is an authorization change: only the live run's exact task qualifies, and normal responsible-user restrictions still apply. - Existing request and retry records are authoritative. Merely naming a responsible/originating user or an external connector sender does not qualify. - Reassignment invalidates existing human request records transactionally. A cancelled run or request cannot regain authority when the task is assigned back. - Ordinary task exceptions require server-owned origin or the legacy board assignment source/reason/actor combination. Existing owner chats use conversation identity. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository tools, code execution, and browser testing. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c8f874311c |
fix(ui): hide Google connectors only on the Connections page (#14774)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - The Connections page lists the apps and saved accounts that agents can use. > - Google Workspace verification is still pending. > - Google entries must be temporarily hidden from this page without removing their implementations. > - This PR filters the final page rows, including saved Google accounts, after the page resolves their provider. > - Definitions, direct setup routes, OAuth profiles, credentials, and runtime access stay intact. > - Review instances can keep the prior UI by staying on their pinned app release. ## Linked Issues or Issue Description **What existing behavior does this improve?** Temporary provider visibility on the Connections landing page. **Current behavior** The page can show Google Workspace catalog entries and saved accounts while verification is pending. **Proposed behavior** Hide all nine Google Workspace rows on this page. Keep every other connector and all Google integration code unchanged. Use an existing release pin for review instances instead of a hostname exception in the app. **Reason and benefit** Pause public discovery without disabling existing runtime tools or removing the implementation needed for verification and later re-enablement. **Breaking changes** Google accounts are no longer visible on this landing page. Direct setup and management routes remain available. This is not an access-control restriction. Related completed work: #13551 used catalog-level visibility. This change is deliberately limited to the landing page and also covers saved account rows. #14740 reduced Google scopes; this change leaves those scopes unchanged. No duplicate open PR or matching open issue was found. ## What Changed - Derive the Google app slugs from the existing Workspace profile registry. - Filter the combined catalog and saved-account rows only inside `Browse`. - Cover all nine Google entries, active/draft/disabled accounts, legacy connection metadata, mixed-provider rows, and independently identified non-Google connectors in regression tests. - Document the display-only hold, pinned review builds, and how to restore visibility after approval. ## Verification - Passed: `pnpm exec vitest run ui/src/pages/apps/Browse.test.tsx ui/src/pages/apps/AppsConnect.test.tsx` (199 tests, including the latest master changes). - Passed: `pnpm check:token-gates`. - Passed: `pnpm build`. - Passed: `pnpm -r typecheck` and `pnpm build` after merging the latest master. An earlier overlapping run hit a local runner codesign race; sequential checks passed. - Passed again after the final custom-provider fix: `pnpm --filter @paperclipai/ui typecheck` and `pnpm --filter @paperclipai/ui build`. - The full local `pnpm test:run` was started, then stopped after the full remote CI suite passed to avoid continuing duplicate long-running work on the developer machine. It is not claimed as a completed local pass. - All 54 latest-head CI checks passed. Two non-applicable Storybook jobs were skipped. One serialized server job lost its self-hosted runner connection; its single retry passed. - Greptile: 5/5 on `aeda167bf4494feed6ee0de2585960511fb02918`, with no unresolved review threads. - Confirmed in the existing review instance that all nine Google entries still appear after its current release was pinned. No new app release was deployed to that instance. - Reviewer steps: open Connections on this branch with Google catalog entries and saved Google accounts. None should appear. Non-Google connectors must remain. Direct Google setup routes must still load. ## Risks - Existing Google accounts cannot be found on this page during the hold. Their data and runtime access remain unchanged. - This is a UI-only filter, not an authorization gate. Direct routes and API access still work by design. - Review instances must not receive this UI build until the hold is removed. Their existing release pin excludes fleet app upgrades; an explicit targeted upgrade must still be avoided. - No migrations, backend changes, broker changes, or credential changes. ## Model Used OpenAI Codex (GPT-5-based coding agent), with reasoning, tool use, code execution, and browser inspection. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f7e36ba3e2 |
fix: isolate repository-free low-trust tasks in private directories (#14766)
## Thinking Path > - Paperclip manages work by agents within company boundaries. > - Email tasks can run under the low-trust review preset. > - These tasks must use an isolated workspace and a sandbox. > - The default workspace strategy assumed that the project had a Git repository. > - A project without a configured workspace failed before the agent could start. > - This change gives each such task a private directory and keeps the sandbox requirement. ## Linked Issues or Issue Description **What happened?** An inbound email assigned to a low-trust agent failed with `git_worktree_base_not_git_checkout` when its boundary project had no configured workspace. Setup had accepted the project and sandbox. **Expected behavior** The agent can process email without a repository. Its workspace stays isolated from other tasks and the shared agent home. **Steps to reproduce** 1. Select a low-trust agent with an active sandbox and a project boundary. 2. Leave the project without a configured workspace. 3. Receive an email through AgentMail. 4. Observe that startup fails before provider work starts. **Paperclip version or commit** Reproduced against `5edf55d73`. **Deployment mode** Hosted staging with sandbox execution. Related: #13256 added email tasks. #13636 fixed default isolation for projects without workspaces; the explicit isolation used by low-trust tasks still needed this path. ## What Changed - Select private task directories for low-trust sandbox tasks with no configured workspace or explicit workspace strategy. - Keep each directory scoped to its company and task. Retain files across turns and reassignment and reject symlink paths and mismatched workspace reuse. - Preserve Git validation for configured workspaces and explicit strategies, plus the existing authorization and remote gates for referenced projects. - Add a startup regression and directory isolation tests. Document the supported repository-free path. ## Verification - The startup regression failed before the fix with the same Git validation error. - 260 targeted email, workspace policy, heartbeat, referenced-project and directory tests pass. - Full `pnpm -r typecheck` and `pnpm build` pass on the latest commit. - All CI checks, including the complete sharded test suite and canary dry run, pass on `b4ccd9802b09b2e95499df72d48b4a3906b8c328`. - The final commit also passes the same server shard locally: 60 files, 1,024 passed / 6 skipped tests. The earlier all-groups local run was interrupted during follow-up edits; complete-suite verification comes from CI on the final commit. - Deployed the reviewed commit to staging and independently verified the full serving SHA. Two real Codex runs in Daytona succeeded and finalized the same private company/task workspace. The first wrote a 35-byte marker; the second read the existing file without modifying it and returned the independently verified SHA-256 `ce3bbeb44d07ca6822826d3a5945752a38d30b356d10829f3159a191e5aa92a6`. - Live runtime caveat: Codex reported a nested `bwrap` loopback permission error and used its configured escalated execution inside Daytona. The outer Daytona sandbox remained active for both runs. - The startup regression uses a real database, production trust checks, workspace persistence, sandbox lease acquisition and realization, and a fake provider. It checks reassignment and allows only an authorized referenced project. - The transfer regression runs production archive/sync-back/merge code against distinct filesystem roots: create output in one sandbox, restore it, then read and update it in a fresh sandbox. The provider I/O is emulated; live staging verification is separate. ## Risks - The new default applies only to low-trust sandbox tasks without workspace configuration. Standard agents and explicit Git strategies keep their existing behavior. - Task directories retain work across turns and consume instance storage. The change does not migrate or copy existing shared files. - No database migration or credential changes are required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser tools. The exact runtime model revision and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e2908fff5c |
fix: enforce terminal outcomes during recovered sandbox cleanup (#14767)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runners can keep a reusable sandbox warm after successful turns. > - Failed turns must stop their sandbox before a later retry resumes it. > - Workspace recovery can release a lease outside the executor's normal teardown. > - Successful file copy-back can then retain a sandbox even when its run failed. > - This pull request checks the durable run outcome at the shared release boundary. > - Ordinary teardown and recovery now apply the same retention rule. ## Linked Issues or Issue Description **What happened?** A real Daytona verification on `5edf55d73` produced a failed native turn. The runner process exited, but workspace recovery retained the sandbox without stopping it. The recovery callback used successful workspace copy-back to select warm retention. It bypassed the terminal-state check in normal heartbeat teardown. **Expected behavior** Keep a sandbox running only after a successful run. Failed, cancelled, timed-out, and interrupted runs must convert requested warm retention to stop-and-retain. Preserve the existing ownership hold before release. **Steps to reproduce** 1. Persist a terminal failed native run whose workspace copy-back succeeds. 2. Release its environment lease from the recovery path with a stored `keep_running` disposition. 3. Observe that the provider receives `keep_running` on the base commit. 4. With this fix, the provider receives `stop_and_retain` and the lease status follows the durable run outcome. **Paperclip version or commit** Base: `5edf55d7350c7f08c9dd132c7e0f1421fa0bf2fb`. **Deployment mode** Native runner with a reusable Daytona sandbox. Related: #14747 retires unsuccessful warm runner sessions. This change closes the separate recovery lease-release path. Related search found no duplicate fix. ## What Changed - Read the durable run status at the shared lease-release boundary. - Apply the existing terminal-outcome retention rule before calling the environment runtime. - Use that same durable status for the lease-state mapping. - Add five database-backed regressions for four unsuccessful outcomes and successful warm retention. - Document the recovery rule. ## Verification - Before the fix: the four unsuccessful-outcome regressions fail; the successful case passes. - After the fix: 163 tests pass across the lease-release, native lifecycle, and explicit continuation suites. - Repository typecheck and build pass. - `pnpm test:run` encountered the existing local `native-session-resume.test.ts:1125` assertion failure; a focused rerun reproduced the same failure. This was also recorded before this follow-up with the unmodified base executor. The full command was stopped after that confirmation, so later local groups were not completed. - All 56 latest-head checks are successful or intentionally skipped (54 passed, 2 skipped), including the complete CI test groups and browser E2E suite. Greptile is 5/5 on `e7da3b3ad`, with no unresolved review comments. - Deployed exact PR head `e7da3b3adbf7a13642c0e56f5b0f0c4666adf358` to the staging workspace and independently verified the serving commit. A real failed native turn persisted `keep_running` and completed workspace finalization, reproducing the recovery-path conditions; Paperclip automatically issued stop-and-retain, and an independent Daytona read confirmed `stopped`. No manual stop was used. - Retried that failed run through the public API after correcting its temporary API-key credential. The same stopped sandbox resumed successfully. Three successful turns retained one live runner PID/start time and the same native, runner, and provider sessions. - Verified exact canonical note contents, deletion persistence, and an unchanged 8 MiB binary after every turn. Subsequent checkpoints copied/hashed only 58 and 87 bytes. After sandbox deletion, canonical files still matched. Temporary secrets were removed and agent/configuration policy restored. - Live acceptance used temporary API-key authentication. The separate managed-subscription authentication issue and browser Retry control were not tested by this campaign. ## Risks - Recovery callers can no longer use stale success status to retain a failed run's sandbox. - The existing native ownership hold still blocks release while ownership is unresolved. - Successful warm turns and explicit destroy dispositions keep their existing behavior. - No database migration, API contract, dependency, or UI change is included. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test analysis. The exact served model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
33f2b3a159 |
fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Connectors catalog lets people give agents tools or connect agents to conversations. > - GitHub put these two uses behind one card and an extra choice. > - People should choose the connection they need from the catalog. > - This pull request keeps GitHub for tools and adds GitHub Code Review Bot as a separate card. > - Each card opens its setup directly. Both use the existing connection code. ## Linked Issues or Issue Description **What existing behavior does this improve?** GitHub connector discovery and setup. **Current behavior** With chat connectors enabled, GitHub opens a menu that asks whether to use tools or create a bot. Saved tools and bots share the same catalog entry. **Proposed behavior** GitHub opens tool account access. GitHub Code Review Bot opens agent selection. Saved bots and drafts appear under the bot card. Chat-disabled instances show only GitHub tools. **Reason and benefit** The catalog names the two uses and removes an extra setup choice. The bot keeps the existing GitHub provider, credentials, endpoint IDs, setup steps, and runtime. **Additional context** Related work: https://github.com/paperclipai/paperclip/pull/12843 and https://github.com/paperclipai/paperclip/pull/14594 established GitHub account identity. This change preserves that tool flow. No duplicate catalog split was found. ## What Changed - Split the generated app definitions into GitHub tools and GitHub Code Review Bot. Reuse the existing GitHub logo and channel method. - Open bot setup directly, including old resume and reconnect links. - Put existing bot endpoints and drafts under the bot card. Hide duplicate internal chat applications. - Keep pasted GitHub URLs mapped to the tool connection. - Add seven Storybook states for the catalog, saved connections, disabled chat, both setup paths, mobile, and light mode. - Fix narrow-screen bot rows so the label cannot overlap status and setup actions. - Update catalog, route, browser, and API tests, plus the GitHub connector guide. ## Verification - [Hosted Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog): seven states built from this branch. The deployment passed its public-file verification. - All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`. Two optional Storybook jobs skip under their normal trigger rules; the manual Storybook deployment passes. The branch has no merge conflicts. - Greptile: 5/5 on the current head, with no review comments or unresolved threads. - `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm build-storybook` passed. The final Storybook fixture also passed UI typecheck and the hosted build. - Targeted catalog, URL matching, routing, grouping, brand, and chat UI contract tests passed. - GitHub provider browser tests: 2 passed. These cover direct tool setup and the bot setup and management lifecycle with provider responses mocked. - Embedded-browser test on an isolated local instance: opened both cards, selected an agent, saved a bot draft, and resumed the same endpoint under the bot card after a reload. - Storybook Tool Setup and Bot Setup assertions pass in the published preview. Chat Disabled assertions pass locally. Inspected mobile and light mode, including the draft-row layout and official GitHub marks. - Local full-suite limitation: `pnpm test:run` was not clean. A cross-company route assertion failed in the aggregate run and passed in isolation; a workspace-runtime test reached its 30-second hook timeout. Some isolated database reruns skipped when the embedded-PostgreSQL availability probe failed. The local aggregate was stopped after CI completed. The corresponding full CI suites pass all 360 tool-access tests and all 162 workspace-runtime tests. - No live GitHub authorization or installation was performed. The isolated instance correctly stopped at the cloud enrollment or public HTTPS prerequisites. ## Risks - Low scope: catalog presentation and routing change. There is no database migration or provider credential change. - Existing GitHub bot URLs now open bot setup directly. The tool route remains `/apps/connect?source=github`. - The bot remains behind the existing chat-connectors feature flag. Existing endpoints retain `provider: github`. - Channel applications are represented by endpoint rows. Regression tests cover legacy bot applications, tools, active bots, and drafts together. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and embedded-browser tools. The exact deployed model ID, context window size, and reasoning setting are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018993140f |
feat: let agents name prompt-only tasks (#14761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5edf55d735 |
fix: retire failed warm sessions before sandbox stop (#14747)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runner sessions can stay warm between turns in a reusable sandbox. > - Heartbeat stops that sandbox when a turn fails or is cancelled. > - The native executor treated every returned terminal result as a successful warm release. > - A failed session could therefore retain a transport for a stopped sandbox. > - This pull request retires unsuccessful sessions before heartbeat stops their sandbox. > - A retry can start without inheriting that stale transport. Successful turns stay warm. ## Linked Issues or Issue Description **What happened?** A structured failed or cancelled terminal result did not throw. The host kept its native session in the warm cache even though heartbeat stopped the reusable sandbox. Later session retirement could use the stopped provider transport and reject the retry. **Expected behavior** Retire the unsuccessful session and collect its managed files before returning to heartbeat. Keep successful sessions warm. **Steps to reproduce** 1. Run a native session with a warm lifecycle and a reusable sandbox. 2. Return a structured failed or cancelled result from the provider. 3. Stop the sandbox after the executor returns. 4. Retry with a changed native session identity. 5. Verify that the previous session was closed before step 3 and is not closed again during retry. **Paperclip version or commit** Developed from `b54b2dc35`, the current `origin/master` at implementation time. **Deployment mode** Native runner with a reusable Daytona sandbox. The same warm-session release path also serves local providers. Related work: #14735 added incremental managed-file checkpoints for warm turns. #14734 preserves tool outcomes during shutdown. This change fixes the host's handling of unsuccessful terminal results. ## What Changed - Retain a warm session only when its terminal run state is `succeeded`. - Use the existing failed-session retirement path for structured failures and cancellations. - Preserve the existing checkpoint-before-retirement path, including when provider shutdown fails. Collect stopped files after successful shutdown, including edits made during shutdown. - Retire the warm owner if the checkpoint or its receipt callback rejects, then propagate the initiating error. - Add ten regression cases for failure and cancellation, with and without managed files. They check cleanup order, cache removal, retry with a new session identity, and edits preserved when close rejects. - Document the lifecycle rule. ## Verification - The failed and cancelled managed-file regressions fail on unpatched master because the provider is not closed. - All 516 native executor tests pass, including ten new regressions. - `pnpm -r typecheck` and `pnpm build` pass. `pnpm test:run` finished its general-server group with 14,537 passed, 75 skipped, and one existing failure in `native-session-resume.test.ts` (`retainedNativeCleanupJournalMatches`, line 1125). A separate run using the unmodified `origin/master` executor fails the same assertion. The command stops before later local groups; all equivalent CI groups pass on this head. - All 56 PR checks are successful or intentionally skipped. Greptile reviewed this head at 5/5 with no remaining findings. - A real Daytona public-API probe forced a Codex authentication failure, waited for a confirmed sandbox stop, corrected the credential, and retried successfully in the same resumed sandbox. It verified the agent note in canonical storage and deleted the test sandbox. - Pumpkin staging on this exact commit: forced a structured authentication failure, confirmed sandbox stop, corrected the credential, and retried successfully in that same sandbox. Three successful turns kept the same live runner PID/start ticks, native session, runner instance, and provider session. Canonical downloads verified the note, all 8 MiB binary bytes, and deletion persistence. Subsequent checkpoints hashed/copied only 52 and 78 bytes. After sandbox deletion, canonical files still matched. Temporary secrets, agent configuration, and test policy were cleaned up or restored. - Unpatched live probes with both 5-second and 90-second retry delays also succeeded. The unit regressions prove incorrect retention; the original stopped-lease exception was not reproduced in those probes. The staging recovery campaign used the public retry API after an explicit reset of the disposable task’s already-deleted old sandbox/session. The UI did not expose a Retry control on the inspected task or run detail views. - Temporary API-key authentication was used for the campaign. This change does not repair the original managed-subscription authentication 401. ## Risks - Failed and cancelled turns now close their provider sooner. Successful warm-turn behavior is unchanged. - Cleanup uses the existing failed-session path. Its existing handling of cleanup errors is unchanged. - No database migration, API contract change, or dependency change is included. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test analysis. The exact served model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (516 relevant executor tests; the full local suite has the independently reproduced baseline failure documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad55d0a281 |
fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use Apps through a gateway that checks identity, company access, and action policies. > - Personal pasted credentials can point to company secrets. Setup can show success while the gateway rejects every call. > - Several OAuth methods also omit the scopes needed for their supported write actions. > - This pull request gives setup, health checks, and invocation the same credential rules. Owners repair existing connections by reconnecting. > - New connections request reviewed permissions for their supported actions. Read-only choices remain available under Advanced. > - Agents can use the connections people give them, while existing consent, identity boundaries, and action restrictions remain enforced. ## Linked Issues or Issue Description Refs #14009 and #14008. This addresses the personal-credential defect. The separate GitHub organization-identity selection defect is outside this change. Related work: #13942 fixed part of new personal-key setup. #14200 independently fixes legacy personal reconnect and protects managed-agent profile credentials during removal. This PR covers that ownership invariant across key and secret-URL setup, reconnect, health, discovery, and invocation, and keeps owner reconnect as the repair path. #14059 tracks requested versus provider-asserted OAuth scopes; it remains separate work. I searched open PRs and issues for Zapier, Airtable scopes, connector writes, and personal credential failures. **What happened?** A Zapier secret URL saved through personal setup can become a company secret referenced by a user grant. Health checks bypass the gateway's ownership check, so the connection appears healthy but calls fail with `grant_credential_invalid`. Custom-header paths can also receive a duplicate `credentials.` prefix. Omitted OAuth scopes make write access depend on provider defaults. **Expected behavior** Personal invocation credentials belong to the selected user. Setup, health, and actual calls enforce the same rule. New connections request documented permissions for supported read and write actions. Existing tokens gain no permissions without provider consent. **Steps to reproduce** 1. Connect Zapier or a generic secret URL with the personal identity. 2. Allow an agent to use the connection and complete setup. 3. Invoke a tool through a run-scoped gateway. The legacy layout fails ownership validation despite successful setup. **Paperclip version or commit** The implementation started from `44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master at `94e8dec56`. **Deployment mode** Built from source. Regression tests use isolated PostgreSQL fixtures and controlled MCP transports. ## What Changed - Share credential writing, ownership validation, and canonical paths across initial setup, resume, reconnect, rotation, health, discovery, and gateway calls. Keep OAuth client-registration secrets separate from invocation credentials. - Existing personal connections with company-scoped credentials require owner reconnect with a fresh key or secret URL. Reconnect creates a correctly owned value and updates the existing grant and declarations. There is no automatic ownership backfill or new startup hook. - Preserve PostgreSQL timestamp precision when reconnect checks whether a grant changed. Previously, converting the timestamp to a JavaScript Date could reject reconnect with a false concurrent-change error. - Protect credentials used by other grants, connections, bindings, managed-agent profiles, routine triggers, or secret proposals from connection removal. - Review all 117 tool methods, including 84 OAuth methods. Record explicit scopes or documented provider-default exceptions with official evidence. Add Airtable's seven scopes, Hugging Face repository/job scopes, and other documented MCP permissions. - Prefer available write-capable methods. Put explicit read-only choices under Advanced. Explain pasted-key permissions and offer reconnect for missing OAuth consent. Preserve existing grants, policies, Google availability gates, and curated scope allowlists. - Reconnect generic secret URLs and custom headers using their stored credential fields. Refresh the catalog after setup, correct reconnect feedback and error guidance, and let Cancel exit invalid setup while Save & exit retains draft-saving behavior. - Apply ownership checks to the new GitHub repository/skill connection picker. Align the permission audit with the Google scope reductions merged on master. - Add run-scoped gateway, ownership, owner-reconnect, OAuth URL, insufficient-scope, UI, and catalog-wide regression coverage. Update the connector playbook and permission audit. ## Verification Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all CI/status gates (55 completed check runs, no failures or pending checks) and has a completed Greptile review at **5/5 with no outstanding findings**. GitHub reports the PR as mergeable/CLEAN. - **Embedded browser:** used the actual server and built UI from this worktree, a fresh isolated database, and local HTTP MCP fixtures. Completed personal bearer-key, secret-URL, and custom-header setup; reproduced the legacy ownership failure; reconnected through the owner’s form; and completed writes afterward. Read-back was verified for bearer-key and secret-URL connections. Public organization-wide setup appeared immediately in Browse without reload. Zapier URL validation/Cancel and Google’s enrollment gate were also exercised. - **Persistence and invocation:** verified user ownership, canonical `credentials.authorization` / `remote.url` / `headers.X-Api-Key` declarations, and unchanged connection/grant identity. The old company secrets retain their ownership. Separate HTTP calls through an actual run-scoped gateway session completed a write and read-back. - **Backend coverage:** the final gateway suite passes all 82 cases, including catalog Zapier and generic inline reconnect. It checks company/user isolation, canonical declarations, same-endpoint URL validation, fresh credentials, retained restrictions, and real gateway read/write execution using fixture transport. A timestamp with PostgreSQL microseconds covers the former false reconnect conflict. - **Local checks:** 368 catalog, gateway, repository, and UI tests passed before the final extra Zapier case; 49 GitHub skill access tests also passed. All three Apps browser regressions pass, including reconnect through the actual form and catalog visibility without reload. Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final patch, and token gates passed. Full tool-access service runs hit varying 15-second Google fixture timeouts; both affected cases and the updated reconnect assertion pass in isolation (3 tests). The complete test matrix passes in CI on this head. - **Verification limits:** no live provider account was available for Zapier/Airtable/OAuth consent or account-bound write proof. Public metadata and local fixtures do not establish provider consent. The original development database clone failed on a pre-existing missing `tool_connections_transport_check` constraint; browser acceptance used a fresh isolated database created by the normal CLI onboarding flow. ## Risks - Existing broken personal connections stay unusable until their owner reconnects. Health, discovery, and invocation return an actionable ownership error; startup does not rewrite credential ownership. - Scope changes affect new authorization requests. Providers may still require resource selection, account roles, paid plans, or app verification. Existing consent and action restrictions remain unchanged. - Shared credentials are retained rather than reassigned or revoked. Provider-default exceptions and unavailable live checks are documented in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`. - No new endpoint, database table, lockfile change, or CI workflow change is included. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell execution, web research, and browser tools. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cbd278dc03 |
fix(interactions): derive chat recipients and validate explicit users (#14742)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agents use saved questions to get human input and continue the same task. > - The standard question example recently told models to copy a user ID. > - A model can omit an identity prefix and create a question its intended recipient cannot answer. > - Agent Chat already knows the conversation owner, so the server can supply that identity. > - This pull request removes the blanket instruction and validates explicit recipients before saving. > - Ordinary questions stay simple, and explicit addressing remains available for decisions that need a particular person. ## Linked Issues or Issue Description Refs #14707, #14188. Related: #14238 handles legacy email recipients; this change prevents invalid recipients in new cards and retains exact ID matching. **What happened?** A model copied a Cloud user ID without its prefix into `addresseeUserId`. Creation succeeded. The intended user's answer then failed the exact recipient check. **Expected behavior** Ordinary chat questions use the saved conversation owner. A task may optionally name a specific recipient. The API rejects an unknown or unauthorized recipient before it creates a card. **Steps to reproduce** Create a chat question for a user whose ID is `paperclip-id:example`. Supply `example` as the addressee. Before this change, creation accepts the invalid recipient and the owner cannot answer. With this change, creation returns 422. Omitting the field saves the full owner ID and allows that owner to answer. ## What Changed - Remove `addresseeUserId` from standard question examples and remove the blanket requester-ID instruction. - Derive the recipient of ordinary chat questions from the persisted conversation owner. Reject conflicting explicit user IDs. - Keep explicit task recipients optional. Validate supplied user IDs with the existing board mutation policy, including company, viewer, and Cloud restrictions. - Preserve explicit agent routing, connector intents, confirmations, exact recipient checks, idempotent retries, and no-login local-board authority in local-trusted mode. - Update the blocker grader to accept an omitted recipient and verify the actual requester answered. - Add database and HTTP tests for prefixed identities, denied recipients, concurrent retries, saved answers, and response delivery. ## Verification - Database interaction service suite: 90 tests passed, including implicit local-board creation/answering and authenticated/Cloud denial. - Interaction HTTP route suite: 84 tests passed. - Affected interaction/native/connector/documentation suites: 231 tests passed across six files after valid-user fixtures were updated. - Resolver and interaction unit suites: 29 tests passed. - Product E2E unit/calibration suite: 793 tests passed; Product E2E typecheck and blocker catalog discovery passed. - Generated API-reference and capability contract checks passed. - `pnpm -r typecheck` and `pnpm build` passed. - Full local `pnpm test:run` did not finish green: its initial general-server pass had 14,416 passing assertions, one unrelated native-resume assertion failure on macOS, and three teardowns from an intermediate fixture cleanup fixed above. Separate broad local groups also encountered timeout/live-port failures under host load. Local UI (7,026), CLI (502), shared (817), and skills-catalog (20) tests passed; the complete final-head CI matrix is the broad verification gate. - After two CI cold-start readiness timeouts, a separate test-only commit gives the first exposure lifecycle fixture the existing normal 30-second readiness budget. Its real HTTP, ordering, and cleanup assertions remain intact; the targeted case and final Linux CI shard passed. Production deadlines are unchanged. - A separate OpenCode fixture failed twice on GitHub-hosted Ubuntu because its cached Node executable was group-writable; the same case passed on AWS runners. The fixture now qualifies its own Linux copy with mode `0500` and the actual copy digest. Host files and production security checks are unchanged. The focused macOS case passed; the new Linux-copy branch also passed on the final AWS-hosted Linux runner (1,125 passing Runner tests, 3 skipped). The final run was not on a GitHub-hosted runner. - Final-head [CI run 36762078176](https://github.com/paperclipai/paperclip/actions/runs/36762078176) passed for `116b968b24fa0a8c5724a7bf96e73a8dda5f0425`: 54 successful checks and two conditional Storybook skips, with no pending or failed checks. The 27 general/serialized test jobs reported 28,635 passing tests. Typecheck, build, Runner, browser E2E, and Canary gates passed. Greptile reviewed that exact head at 5/5; both review threads are resolved, with no open follow-ups. - No live provider replay is claimed by this PR. ## Risks - New explicitly addressed cards reject users who cannot mutate the issue, including viewers, inactive members, and invalid IDs. Callers that supplied invalid recipients must correct their request. - Existing addressed cards are not rewritten. Existing authorization checks remain strict. - Chat inference applies only to questions without an agent addressee. Connector intents and governed confirmations retain their own recipient paths. - No schema change or migration is required. ## Model Used OpenAI Codex, GPT-6 (exact serving variant and context window are not exposed in this environment). Used reasoning, tool use, code editing, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b54b2dc35c |
fix: preserve warm Codex turns with incremental managed file checkpoints (#14735)
## Thinking Path > - Paperclip manages AI agents and keeps their instructions and files durable. > - Native Codex runners can keep a process alive between compatible turns. > - Managed file collection stopped that process after each turn, which defeated warm reuse. > - Agent folders can contain large images and other files, so full copies on every turn are expensive. > - This change keeps one managed directory for the live session and saves only file changes after each turn. > - Ownership, authorization, instruction changes, and process retirement still control when reuse is safe. ## Linked Issues or Issue Description Related: #13710 introduced native warm session reuse. This fixes managed file collection that still forced those sessions to stop. No duplicate open PR or issue was found. **What happened?** With managed instructions and warm native Codex enabled, consecutive turns reused a Daytona sandbox but started a new runner process each time. The managed directory collector required process termination before saving files. **Expected behavior** Compatible turns keep the same process and managed `AGENT_HOME`. Each completed turn saves added, changed, and deleted files before the next turn starts. Unchanged large files do not transfer again. **Steps to reproduce** 1. Use a native Codex agent with managed instructions and a reusable Daytona environment. 2. Enable warm session reuse and run three turns on the same task. 3. Write a large binary on the first turn, edit a small note on each turn, and delete a file on the second turn. 4. Compare process identity across turns and read the canonical files through the public agent-files API. **Paperclip version or commit** Reproduced on `d30b03bd8c17604cdab1533eeeeb087aba30e8b1`. **Deployment mode** Local server with remote Daytona execution; cloud native runner uses the same path. ## What Changed - Retain the managed directory only for the verified owner of a live native Codex session. - Checkpoint each completed turn before releasing the session for reuse. Retry unstable captures, then stop and collect when a warm checkpoint cannot be validated. - Compare metadata and cached hashes, stream only changed file payloads, record deletions, and validate path, content, quota, and authorization before saving. - Rotate sessions when canonical files, loaded instructions, credentials, or launch policy change. Fence stale collection and cleanup callbacks from later owners. - Keep cleanup and recovery aware of the current session owner. Recheck canonical files under the writer lock at handoff, attach the successor collector before fallible bookkeeping, and emit one final save receipt on checkpoint fallback. Preserve storage warnings across unchanged checkpoints. - Add regression coverage and a three-turn Daytona test with independent public API file checks, an unchanged 8 MiB binary, deletion checks, and strict process identity checks. - Document checkpoint consistency, lifecycle behavior, and local run-log counters. - Replace a timing assumption in the Daytona teardown test with explicit transfer-arrival gates after CI exposed an unset release callback. ## Verification - Full local `pnpm -r typecheck` and `pnpm build` passed. Server checks were repeated after the final storage-warning fix. - Runner E2E typecheck and 749 runner E2E unit tests passed. - Focused file checkpoint, directory ownership, instruction collection, native session, and merge tests passed. After review fixes, the managed-directory and native-session suites passed 550 tests, including intervening canonical edits, same-run fresh restore, failed handoff collection, and one-call fallback collection. Server typecheck passed again. The Daytona plugin suite passed 218 tests. The quota-warning regression failed before the fix and passed afterward. - Three real Daytona campaigns passed before the final handoff review fixes. The latest kept PID 547 across all three turns. The first checkpoint copied 8,388,635 bytes; the next two copied 36 and 54 bytes. Public API reads verified the binary, note contents, and deletion after every turn. Test cleanup deleted the sandbox. - The final head was also deployed to an isolated cloud staging instance and passed three UI-triggered native Codex turns with managed instructions. All three retained the same process ID/start time, native session, provider session, runner instance, and Daytona sandbox. Checkpoints copied 8,388,643 bytes on turn 1, then only 52 and 78 bytes on turns 2 and 3; those warm captures also hashed only 52 and 78 bytes. Independent canonical API reads verified every byte of the unchanged 8 MiB binary and the exact note contents after every turn; the deleted file returned 404 after turns 2 and 3. After restoring the original lifecycle and agent-auth configuration, removing the temporary secret, pausing the test agent, and deleting both test sandboxes, independent canonical API reads still verified the entire binary, the final 78-byte three-line note, and the deletion. The native runner flag remained enabled and the final serving revision remained the PR head. - Two earlier staging attempts are preserved as failures and are excluded from the acceptance result: a saved ChatGPT login failed with a provider routing 401, and its subsequent stopped-sandbox retry failed before provider startup with a closed-lease admission error. The successful campaign used a fresh sandbox and a temporary encrypted API-key binding. The stopped-lease retry remains unexplained; this campaign does not establish recovery of that failed sandbox. - All [Paperclip CI gates](https://github.com/paperclipai/paperclip/actions/runs/36750397355) pass on `26ef2ef56a389259246809805c0b34a4747eb86b`, including full test partitions, build, typecheck, runner verification, E2E shards, and the Canary clean public-npm install. Greptile reviewed that exact head at 5/5 with no unresolved review threads or outstanding findings. - Full local repository coverage used the existing CI partitions, but the 40,000-file Git streaming stress test timed out and its local retry was interrupted by macOS thermal emergency sleep; this is not a green full local suite claim. The exact stress test passed on the final head in [CI server shard 2/12](https://github.com/paperclipai/paperclip/actions/runs/36750397355/job/110008294290), in 111.9 seconds. - Repeat the live test with configured credentials and a Linux runner artifact: `pnpm test:e2e:runner -- --id daytona-warm-continuity.runner-codex.daytona.warm-three-turn`. ## Risks - This is a file-level checkpoint, not an atomic snapshot of the whole folder. Background writes after a capture are saved by the next checkpoint or final stopped collection. - Metadata scans still visit all paths. Modified files transfer in full; unchanged files do not rehash or transfer. - Incorrect ownership or reuse could collect the wrong directory. Run ownership fences, current authorization, stable capture validation, and stopped collection fallbacks are covered by tests. - Warm reuse remains opt-in. No database migration or fleet default changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code editing, tool use, and test execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d432dc7fa3 |
Add GitHub-synced skill sources (#14713)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Company skills supply instructions and files to those agents. > - GitHub imports already exist, but users cannot manage repositories as skill sources. > - Repository refresh also needs caller-authorized access and complete local packages. > - This pull request adds Sources inside Skills and reuses GitHub connections from Apps. > - Installed snapshots let agents use skills without fetching GitHub during a run. > - Manual refresh preserves skill identity and leaves failed imports on their last good version. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: skills UI, server, database, shared contracts, and runtime materialization. **Problem or motivation** Users keep skills in GitHub repositories. They need a clear way to select, import, and refresh those skills. Existing imports do not expose repository management or consistently preserve supporting files. **Proposed solution** Add company-scoped skill sources. Browse repositories from all accessible GitHub connections, or paste a public repository or branch URL. Select whole skill packages, inspect included files and reference warnings, and install complete, immutable snapshots. Refresh each source manually. **Alternatives considered** Project repository settings hide the workflow from Skills. A second GitHub connector would duplicate credentials and grants. Upstream editing and PR creation are separate work. **Roadmap alignment** This implements the Skills Manager direction in ROADMAP.md. The maintainer requested this scope and reviewed the component and full-app journey stories before implementation. Related reports: Refs #10285, Refs #10949, Refs #13464. Related work: #14356, #13656, #9268. ## What Changed - Add source and entry records, an idempotent migration, company-scoped APIs, and legacy GitHub import adoption. - Reuse current caller grants and credential refresh. Combine and deduplicate repository inventories across accessible connections. Pasted public URLs also prefer the active user’s authorized connections. Tokens stay in the Git child environment, never argv or disk. - Fetch a shallow Git snapshot at one immutable commit. Scan the full local tree, including hidden and nested folders. Read Git objects without checkout or archive transformations and enforce nested package boundaries. - Bound Git downloads to 128 MiB and three minutes. Cancel active process groups and remove incomplete downloads. Preserve cancellation and deadlines while progress drains; close stalled HTTP progress streams after 30 seconds. Reuse caller-scoped temporary snapshots for preview/import after reauthorization. - Index repository package boundaries once and cap expanded work at 1,000 packages, 10,000 files, and 100 MiB, including repeated copies of shared blobs. Bound path depth and the shared path index. Discovery keeps audited manifests without retaining all package bodies. - Resolve moving branches before fetching so unchanged discovery reuses caller-scoped snapshots. Limit active scans, scan frequency, and new downloads per caller and company; quotas apply before metadata reads and across connections, and cached scans do not consume the download quota. - Store complete versions with script content, binary bytes, and executable modes. Preserve these through copies, runtime caches, and runner packaging. - Stage downloads before publication. Use source leases, revision checks, and transactional activity records. Keep installed versions after failures, upstream deletion, deselection, and disconnect. - Add the approved import flow, Sources page, selection tree, provenance, read-only Studio behavior, and saved return from GitHub setup. - Add package manifests, commit-pinned file previews, and separate runtime requirements and reference warnings. Supporting files are included together; nested skills remain independently selectable. Preview requests reauthorize the caller and re-audit package content. - Show installed skills as compact links beneath each source. Repository titles open GitHub. Keep Refresh, Select skills, and Disconnect source in a three-dot menu. Source rows omit the branch, imported count, and refresh timestamp; action alignment and repository titles work at narrow widths. - Stream discovery metadata over an opt-in NDJSON response. Show measured Git download progress and real package/file counts, animate newly checked skills, support cancellation, and require a complete scan before selection. Keep the existing JSON API. - Retain component stories and add a separate full-app journey story group. Include fixed progress states and interactive scan, large-repository, interruption, and saving stories. - Update Skills documentation and product contracts. Suppress private GitHub skill references in telemetry. Privacy review requested for the telemetry changes. ## Verification - Local repository typecheck, full build, token gates, and Storybook build passed during this work. Focused transport, authorization, scanner, persistence, route, and UI tests pass. The final UI refinement passes all eight focused UI tests, UI typecheck/build, and token gates. The scanner resource and repeated-discovery fixes pass 132 focused scanner, transport, authorization, source-service, route, and rate-limit tests, plus server typecheck/build. Full-suite verification comes from CI; the older full local Vitest run was stopped after unrelated chat failures and a font-test failure, all of which passed in fresh focused runs. At commit `1098d5996`, all 54 active checks pass; two optional Storybook jobs are skipped. CI covers repository typecheck, build, the full test suites, browser shards, and the canary dry run. Greptile is 5/5 with no open findings; the security scan also passes. - Adversarial scanner tests verify repeated-blob byte accounting with and without declared sizes, package/file/path caps, one-time repository indexing, metadata-only discovery audits, and nested package boundaries. Additional tests cover branch movement, snapshot reuse, caller/company quotas, isolation across connections, active-lease cleanup, quota recovery, and rejection before any metadata API call. - Real Git tests verify hidden paths, exact binary bytes, executable modes, export-ignore preservation, symlink/submodule reporting, pinned commits, caller-scoped cache reuse, cancellation, cleanup, and credential isolation. Regression tests hold both download slots with permanently blocked progress callbacks, verify timeout/cancellation cleanup and retry, and exercise HTTP backpressure cancellation. Access tests cover automatic public-URL connection selection and revoked grants. Database tests verify company and grant audiences. - Live isolated browser test: the public `anthropics/skills` scan now completes and discovers all 20 skills without connecting an account. Imported canvas-design with all 83 files, opened it from Sources, and verified the installed binary-font preview/download control. Package previews also expose the complete file inventory before import. Cancelled an active Git download and retried successfully to all 20 discovered skills; the browser displayed measured download progress. The current audits reject four other packages; eligible selections remain importable. - Browser checks verify the simplified source rows at desktop and narrow widths, keyboard navigation into the actions menu, Refresh from the menu, selection, and fixture disconnect with installed skills retained. Storybook includes a menu-open checkpoint and a 320px layout. - Storybook includes receiving/preparing download checkpoints and a timed full-app import journey, plus cancellation, retry, large-repository, and saving states. Streaming tests cover split UTF-8 frames, incomplete streams, late responses, cross-company requests, HTTP errors, and JSON compatibility. - Earlier live acceptance on this PR imported `stitch-skill` with `DESIGN.md`, assigned it to an agent, disconnected its source, and ran a successful Studio test that read both installed files. An editable copy changed independently. Both Skills variants, mobile selection, and return from GitHub setup were exercised. - Private access, revoked credentials, OAuth success return, binary/script preservation, concurrent refresh, transaction rollback, version pins, and legacy adoption have automated coverage. A real private-repository OAuth grant was not created during this test. ## Risks - The migration groups recognizable legacy imports without provider calls. Their first successful refresh completes the local package snapshot. - Reference checks are advisory. They cover Markdown links and explicit relative resource paths, not arbitrary runtime dependency graphs. Preview text is capped at 64 KiB; imported bytes remain complete. - Git must be installed on the server. Shallow fetches still download the branch snapshot, including files outside selected packages. Downloads have size/time/concurrency limits. Temporary caches are bounded and caller-scoped. GitHub API quota still applies to repository metadata and the connection picker; content no longer uses per-file API requests. Failed scans retain installed content. - Sources depend on the current caller's GitHub access. A saved connection does not grant access to another person's token. - GitHub script support and immediate manual refresh are explicit maintainer-approved requirements. The operator trusts the selected repository and accepts upstream script and executable-mode changes on refresh. Static audits are not a sandbox or a guarantee of safe code; agents may later invoke installed helpers under their runtime permissions. Import and refresh do not execute scripts, hooks, package installation, or builds. Raw URL and skills.sh imports keep their prior script restrictions. - Source originals remain read-only. Refresh affects subsequent unpinned runs; explicit pins and active runs retain their versions. - The telemetry change removes source-managed GitHub identifiers from skill-reference events. It introduces no event or field. Please review the privacy boundary. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7141ab68a0 |
docs: improve README quickstart, harness logos, and links (#14744)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The README introduces supported harnesses, features, and setup. > - Its logo table omits several supported harnesses. > - Feature descriptions also need direct links to their user guides. > - People interested in Paperclip Cloud need a clear waitlist link. > - The Quickstart leads with a long installer script even though npx is supported. > - This pull request adds the logos, guide links, waitlist and careers links, and a maintenance FAQ, and leads Quickstart with asking an agent to install Paperclip, followed by the npx command for manual setup. ## Linked Issues or Issue Description **Issue type** Missing documentation coverage. **Where is the issue?** The README header, “Works with” table, feature cards, system descriptions, Quickstart, deployment FAQ, and Contributing section. **What's wrong?** The icon table omits supported harnesses. Feature descriptions lack direct links to the published guides, and the header has no Cloud waitlist link. The README also lacks a careers link. Its Quickstart leads with a multi-step shell installer instead of the supported npx command. **Suggested fix** Use the brand assets already shipped with the UI. Arrange the table in two rows. Add direct documentation links in context and a Cloud waitlist callout below the embedded video and above the main heading. Link the careers page near the bottom. Lead Quickstart with asking an agent to install Paperclip, followed by npx and the Node.js prerequisite for manual setup. Link to managed-install details and format FAQ entries with Q/A labels and spacing. Refs #14741, the merged README refresh. Related: #1921, an existing proposal to replace the old Codex placeholder. This change uses the current UI Codex asset. Searched open README logo PRs before opening this follow-up. ## What Changed - Add bold Q/A labels and paragraph breaks to all six FAQ entries, with extra spacing between questions. - Lead Quickstart with “Just ask your agent to install Paperclip” and a link to the installation guide. Follow with `npx paperclipai@latest onboard --yes` and the Node.js 24.11+ prerequisite for manual setup. Replace the shell-installer blocks with the installation-guide link and use npx for related setup commands. - Add a maintenance FAQ identifying the Paperclip team, linking to paperclip.ing, and citing over 2,700 merged pull requests with a link to the merged PR list. - Add “We're hiring” beneath the Contributing paragraph, linked to the careers section. - Add a “Sign up for the Paperclip Cloud waitlist” link below the embedded video and above the main heading. - Link fifteen guide pages from features, systems, setup, and deployment text. Use published pages on docs.paperclip.ing. - Add Gemini CLI, OpenCode, Pi, Hermes, Grok Build, and Kimi Code logos. - Label Cursor Cloud and Hermes Gateway beside their shared brand marks. - Use the existing UI Codex and Cursor assets. - Remove the Bash and HTTP logo cells and arrange ten harness logos in two rows of five. Set uniform image dimensions and align icons at the top of each cell. - Select existing dark-mode assets through picture elements. - Remove the text roster that repeats the expanded table. Keep the adapter documentation link. ## Verification - Confirmed all six FAQ entries have Q/A labels and spacing, with their wording preserved. Confirmed the first Quickstart subsection is “Just ask your agent to install Paperclip.” - Checked npx support against the published installation guide, npm package metadata, and the CLI onboarding implementation. npm currently publishes `paperclipai` with the expected binary and Node.js >=24.11.0 requirement. Checked Quickstart shell syntax and confirmed all other README sections are unchanged by the Quickstart edit. No installer or onboarding command was executed. - Confirmed 2,710 merged pull requests through the GitHub search API on September 30, 2026 (`repo:paperclipai/paperclip is:pr is:merged`, `incomplete_results: false`). The FAQ uses “over 2,700” so the claim remains valid as more PRs merge. - Passed: `git diff --check`. The careers page loads and contains the Careers section. - Passed: local structure and asset checks. All sixteen SVG sources exist and parse. All ten images have labels and dimensions. Each row contains five logos. - Checked the roster against the built-in adapter registry and UI brand mapping. The table covers twelve branded adapter variants through shared harness logos. Generic process and HTTP adapters are described beneath the table. The retired ACPX adapter and Paperclip's own runner engine are not separate external harness brands. - The remaining logos were verified in the earlier GitHub dark-mode render at 32 by 32 pixels. After removing Bash and HTTP, a local HTML structure check confirmed two rows of five logos with the existing top alignment. - Passed: all fifteen new guide destinations and the waitlist URL returned HTTP 200 with matching page titles. HTML link nesting is valid. Approved four-pillars and roadmap content is preserved. - Full typecheck, runtime tests, build, and browser suites were not rerun for this Markdown-only change. No runtime source or image assets changed. - CI and automated review are pending. The PR is a draft. ## Risks Low runtime risk. Only README.md changes. The table now references existing UI asset paths, so future asset moves must update these links. No new artwork or dependencies are added. The documentation links depend on the published site paths. The primary Quickstart requires Node.js to be installed already; this prerequisite is shown immediately above the command. ## Model Used OpenAI GPT-6 via Codex. The exact model variant and context-window size were not exposed in this session. Used repository inspection, shell tools, documentation editing, and browser verification. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
25c422ba7e |
fix(apps): request minimal Google service scopes (#14740)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Google connections give agents service-specific tools through
governed credentials.
> - Each connection has a reviewed OAuth scope set.
> - Docs, Sheets, and Slides request Drive permissions in addition to
their own service scopes.
> - Google documents these permissions as alternatives, not combined
requirements.
> - This pull request removes those extra permissions and redundant
Calendar write free/busy access.
> - Users grant fewer permissions without adding tools or changing
connection access policy.
## Linked Issues or Issue Description
Refs #13820. Related #14739 changes other connector permissions; it does
not reduce these Google profiles.
Companion broker PR:
https://github.com/paperclipai/paperclip-cloud/pull/615. Ship the
matching changes together after fresh-grant validation.
**What happened?**
Seven Google profiles request redundant scopes. Docs, Sheets, and Slides
request Drive scopes. Calendar write requests free/busy even though
calendar.events authorizes its availability tool.
**Expected behavior**
Each profile requests only the scopes required for its reviewed tools.
Managed and customer-owned OAuth methods use the same set.
**Steps to reproduce**
Inspect the Google profile registry and the four app definitions on the
base commit. Compare their scope sets with Google's MCP authorization
alternatives linked in the updated documentation.
**Paperclip version or commit**
Base:
|
||
|
|
d1f3e7bb0d |
docs: refresh README capabilities and roadmap (#14741)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The README introduces the product and directs people to setup and the roadmap. > - The product now supports more harnesses, connections, skills, and team workflows than the README shows. > - Some copy still describes available features as future work or makes claims broader than the implementation. > - This pull request updates the README and matching roadmap entries from current source evidence. > - Readers can see what they can use, what requires setup, and what remains experimental. ## Linked Issues or Issue Description **Issue type** Outdated information and missing documentation. **Where is the issue?** README.md feature descriptions, four pillars, setup, FAQ, and roadmap; related entries in ROADMAP.md. **What's wrong?** The README omits supported adapters and major connection and skills workflows. It presents Connected Apps and Agent Chat as wholly future work. Some budget, audit, and approval descriptions also need more precise wording. **Suggested fix** Keep the existing structure, four-pillars picture, and completed roadmap milestones. Extend the four-pillars table without removing its existing content. Add concise descriptions of supported features. Correct capability and setup claims. Distinguish available, experimental, and planned work, and explain that the roadmap follows the default branch. Searched open README and roadmap PRs and related documentation issues. Refs #14640, which proposes a separate launch-video update; this change preserves the existing video. ## What Changed - Expand adapter coverage and describe model choice alongside durable team context. - Add six feature cards: connections, personal identities for shared agents, Skill Studio, routines, artifacts and feedback, and team templates. - Describe experimental agent conversations and external chat/email entry points. - Correct budget, audit, approval, goal, session, secret, portability, and multi-organization claims. - Qualify export portability: plain environment values and local paths can remain, so packages need review before sharing. - Retain the four-pillars image and all existing table content; add connection identities, skill history, team templates, and run history. - Correct mention wake behavior, persistent npx data, source-build prerequisites, and hosting guidance. - Retain all 18 completed README roadmap milestones and its original closing sentence. Add seven completed milestones: Connected Apps, personal/shared AI accounts, Shared Agents Use Personal GitHub Identities, skill version history, document comments and revisions, company-wide search, and mixed-model/harness teams. - Keep the README’s yellow roadmap entries to their names. Expand the new milestones in ROADMAP.md alongside the existing status updates for agent chat, memory, recovery, evaluations, and queue scope. - Preserve all existing top-level section headings and their order. No runtime files change. Research: reviewed 393 candidate change summaries from a 60-day history of 1,274 non-merge commits, then checked relevant implementation, feature defaults, contracts, and current public adapter documentation. The main gaps span adapters, connections, responsible identities, skills, routines, deliverables, teams, chat/memory status, governance claims, and setup. Suggested follow-up work: refresh the product demo; add three concrete use cases; shorten Quickstart by moving secondary setup paths into the docs. ## Verification - Passed: `git diff --check`. - Passed: local link and asset targets, Markdown anchors, and HTML table nesting in both edited files. - Passed on the earlier revision: rendered README inspection on GitHub, including the new feature cards and roadmap status labels. - Passed: exact restoration checks for the original picture and both requested passages, preservation of every original pillar-table cell, and all 18 original checked roadmap items. - Passed on this revision: seven new green milestones synchronized between README.md and ROADMAP.md, three short yellow labels, local link targets, table markup, and verification that README edits are confined to its roadmap. - Verified new milestone claims against AI Connections and managed GitHub identity contracts, Skill Studio restore behavior, document annotation/revision routes, company search, and the adapter registry. - Passed on the earlier revision: `pnpm -r typecheck`. - Passed on the earlier revision: `pnpm build`. No runtime code changed in this documentation follow-up. - `pnpm test:run` reported a failure in unchanged native-session recovery code. Stopped the broader local run after reproducing it with `pnpm exec vitest run --project @paperclipai/server server/src/services/native-runtime/native-session-resume.test.ts --no-file-parallelism --maxWorkers=1`: 40 passed, 1 failed. The failing case is “archives a damaged prior epoch before completing guarded same-task replacement with real runnerd”; `retainedNativeCleanupJournalMatches` returned false at line 1125. The full local suite did not complete. - Previous revision `4bef0af3da13b03875efacd9e1f165bc69932b25` received Greptile 5/5; the export wording finding remains addressed. A fresh review is pending for this documentation follow-up. - CI remains pending. This PR remains a draft; local runtime tests are not green. - Checked feature claims against the adapter registry, feature catalog, Connections contracts and action UI, skill and team services, task semantics, deployment docs, and source startup code. - Browser suites were not run. This change edits documentation only. ## Risks Low runtime risk: only README.md and ROADMAP.md change. The default branch can lead published packages, so the roadmap now links to release notes. Provider setup and feature flags still affect availability. Live provider integrations were not requalified for this documentation pass. ## Model Used OpenAI GPT-6 via Codex. The exact model variant and context-window size were not exposed in this session. Used repository research, reasoning, shell tools, and documentation editing. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
94e8dec56b |
fix(runner): preserve tool outcomes through shutdown and restart (#14734)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner sends authorized tool calls to the server and saves their results. > - A provider turn can stop while a server write is still running. > - The old shutdown path invented a failed result that could conflict with the real result. > - Truncated execution input and incomplete recovery records made the failure harder to diagnose. > - This pull request preserves exact inputs and actual outcomes through shutdown and restart. > - Tests force the race and crash boundaries so safe retries do not repeat writes. ## Linked Issues or Issue Description **What happened?** Stopping a turn during a server tool call could record a false failure, then reject the actual result as a conflict. The diagnostic input formatter could truncate instruction content before execution. A crash during saved-result delivery could leave that delivery permanently indeterminate. Cleanup could hide the first failure, and a retry could overwrite earlier run logs. **Expected behavior** Keep dispatched tools pending until their actual result is known. Preserve accepted input bytes. Accept identical result delivery without failing the task. Reject conflicting results with enough evidence to diagnose them. Recover saved-result delivery without repeating the business operation. **Steps to reproduce** 1. Hold an instruction update at the filesystem commit barrier. 2. Stop its provider turn before the server returns the result. 3. Release the write, deliver its result, and replay the same result. 4. Repeat with a restart before and after the delivery receipt is saved. 5. Check that there is one write and one audit row, and that the exact result survives. **Paperclip version or commit** The change was developed from `44736c9c7` and rebased onto `0e5830887`. **Deployment mode** Self-hosted server with the native runner. Tests use local runner processes, scripted providers, and PostgreSQL. Related work: #12353 added durable semantic tool receipts; #12384 added durable Codex tool recovery; #12404 bound semantic tools to ACPX sessions. #14633 covers separate native-provider cancellation and qualification work. This PR addresses server semantic-tool outcomes and their durable delivery. No duplicate fix was found. AgentMail discovery is outside this PR. ## What Changed - Close turn admission without inventing results for dispatched tools. Keep pending calls and accept late actual results. - Accept identical result replay with a diagnostic warning. Include call identity and both result hashes in real conflict errors. - Preserve exact execution arguments. Reject prohibited or oversized input before dispatch. Keep diagnostic previews redacted and bounded. - Commit instruction-attempt evidence before the filesystem write. Save completed mutation receipts so concurrent and restarted duplicates return the first result. Recheck authorization before replay. An attempt without a completed result stays unknown and cannot execute again. Definite pre-write failures save and replay their original error without another write. - Recover an interrupted saved-result delivery only for backends with durable result receipts. Never replay an ordinary business operation with an unknown outcome. - Preserve the initiating error when cleanup also fails. Record incomplete settlement evidence. Propagate typed unknown-outcome errors through the native tool wrapper without creating a false completed tool result. - Append run-log attempts and restore the durable log before appending after local file loss. Reject incomplete restores. Publish a restored prefix only if the destination is absent so concurrent attempts cannot overwrite new lines. - Add deterministic race, crash, replay, authorization, exact-content, and log-restoration tests. Document their assertions in `packages/paperclip-runner/docs/durable-recovery.md`. ## Verification - Current head: `7e088f4c7fba8ebabf98ae95485a5753b013d489`. All 55 applicable checks pass; four conditional/manual checks are skipped. This includes build, typecheck, Rust, both runner TypeScript shards, server and workspace tests, all eight browser shards, isolated runner compilation, and the clean-install release dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/36746101110). - Greptile reviewed this exact head at 5/5 with zero new findings. All three earlier review threads are resolved. - Focused local verification includes 11 instruction integration tests, 23 surrounding authority/tool tests, 26 run-log tests, and 169 controller/driver tests. The post-rebase controller/transport/runtime selection passed 415 tests. The full Rust release suite passed 617 tests with two ignored. The real-process SIGKILL recovery test passed three consecutive runs. - The fault matrix in `packages/paperclip-runner/docs/durable-recovery.md` uses explicit barriers, real PostgreSQL rollback, durable journal reloads, and killed runner processes. It covers late results, identical and conflicting replay, exact long content, concurrent log restoration, lost commit acknowledgements, and definite failure replay after the original CAS base becomes valid again. No paid model calls are needed. - Full local recursive typecheck and build passed during implementation. Server typecheck and the runner TypeScript build passed after the review fixes. The broad local repository test run was stopped after repeated database startup timeouts. Four timing/launch failures in an earlier broad runner run passed focused reruns without changed assertions or timeouts. These are local verification limitations; the complete current-head CI suite is green. An earlier CI workspace job received an infrastructure shutdown signal; its current-head replacement passed. ## Risks - A stopped turn can remain blocked when a dispatched operation has no proven result. The system does not guess its outcome or rerun its effect. - Conflicting results still fail settlement. Existing failed or conflicting journals are not repaired automatically. - Accepted semantic input is limited to 480 KiB of encoded JSON to fit the encrypted transport. Larger input fails before execution. - Instruction filesystem writes and database receipts are not one atomic storage operation. A separately committed attempt and audit record survive rollback. An attempt without a completed success or definite pre-write failure receipt remains blocked as an unknown outcome. It is not replayed or reported as success. - Run-log restoration now reads the durable object before appending when the local log is missing. Failed or incomplete reads reject the append. - No schema migration, dependency change, workflow change, or AgentMail change is included. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test analysis. The exact served model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; the broad local run limitation is recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3c561642b4 |
fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask for decisions and optional details through cards in chat. > - A clear approval in a message can leave the matching card pending. > - An unanswered question can also block an unrelated later reply. > - Decisions need a saved source message, while optional questions need to remain answerable in history. > - This pull request records conversational decisions and lets users move on from questions and answer them later. ## Linked Issues or Issue Description **What happened?** Native Claude and Codex could act on approval in chat while the original approval card stayed pending. Pending question forms stayed above the composer, were absent from history, and could suppress later chat replies. A late native question answer could wait for a finished run to reconnect. **Expected behavior** The active agent records a clear approval or refusal against the exact card and user message. Ambiguous replies do not grant consent. Users can send another message without answering a question. The question remains pending in history and can be reopened and answered later. The saved answer reaches the agent. **Steps to reproduce** 1. Ask an agent to propose work with a confirmation card, then approve it in chat. 2. Check that the original card records that approval before work starts. 3. Ask an interactive question, send an unrelated message, and reload. 4. Open the unanswered question from history and submit an answer. Related work: #14408 added completion delivery. #14607 tests completion reporting turns. Neither records conversational answers on approval cards. ## What Changed - Add a confirmation endpoint backed by a user comment, with schema validation, OpenAPI discovery, and native Plan-mode access. Ask mode remains read-only. - Check company, active run, actor, current session, message provenance, revision, and resolver policy. Save the decision and audit in one transaction. Retries do not repeat effects. Emit resolution telemetry after commit. - Give fresh and resumed chat turns the actual pending confirmation identities. Teach agents to save clear conversational decisions before acting and to clarify ambiguity. - Keep unanswered Agent Chat questions as compact history entries. A newer user message closes the old form. Question cards never contribute to composer pending counts or navigation, including after dismissing a fresh form. The history card is the sole reminder; clicking it restores that exact form and draft. - Preserve Agent Chat questions when later messages or questions arrive. Historical ordinary inputs no longer gate later chat replies. Current-run requests, task execution, and governed approvals keep their gates. Remove the special acknowledgement-publication proof helpers that this rule replaces. - Route answers to finished native runs through durable fresh-wake delivery, with existing idempotency and source-question context. Settle late replies against contiguous completed conversation turns and freeze their history replay; failed, unhandled, and newly arriving messages remain actionable. - Add real-component Storybook scenarios, database and UI regressions, and a three-turn native Claude/Codex E2E case. Capture distinct, UI-ready screenshots and report the individual assertions. ## Verification - Focused decision/publication/UI regressions after merging master: 288 passed; subsequent UI draft, failed-send, and conversation checks: 199 passed. - Native question and durable delivery regressions: 106 passed, including all four terminal run states and exactly-once late delivery. Seven targeted regressions fail against the original implementation and pass with the fix. - Latest conversation/decision/native-delivery regressions after the master merge: 121 passed. Covers completed progress, missing or failed intervening turns, new messages during a late reply, stale sessions, and frozen retry/replay boundaries. Four new assertions fail before the ordering fix. - E2E support suite after the master merge: 792 passed. Negative controls reject expired cards, wrong questions/answers, stale or missing replies, unrelated clarification forms, and unexpected tasks. - The embedded-browser walkthrough caught one additional defect: dismissing a fresh question still showed a composer badge. Both Cancel and close-button regressions failed before the fix. The fix at `65f2ade12` passes 170 chat-thread tests and 792 E2E support tests. After merging master, 232 chat-thread/confirmation tests, server/UI typechecks, and token gates pass. The preview and two-provider live E2E pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at that commit. All 55 checks are now successful at `e5512a206` (four conditional checks skipped), including the aggregate verification gate and clean-install canary test. The first attempt was interrupted by simultaneous CI worker shutdowns; one failed-job rerun passed without code changes. - [Published Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on): nine real-component scenarios. Manually exercised move on, reopen, preserve draft, answer later, answer one of multiple questions, and a custom mobile answer in the embedded browser. Retested fresh Cancel and close-button dismissal in the updated build, then reopened and submitted the preserved Green selection and inspected its answered receipt. Static preview has no live model/backend; its callbacks are fixture responses. - [First live campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/) reproduced the late-answer completion-state defect on both providers despite correct saved answers and acknowledgements. It also exposed a valid imperative clarification rejected by the old oracle. Both issues are fixed with regression controls; this failing run is retained as evidence. - [Four-cell qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/) passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous confirmation, each on native Claude and Codex. Inspected saved state, source-message decisions, visible cards, and agent replies. Both late-answer chats settled to waiting; no unrequested tasks were created. [Final branch rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/) passed 2/2 at `142630720`: the same unanswered-question journey after merging master, plus an additional screenshot and browser assertion for the actual late-answer acknowledgement. - [Composer-reminder E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/) passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each, with explicit no-badge assertions before and after reload. Inspected saved pending/answered state, both screenshots with a clear composer, and actual Blue acknowledgements; all five behavioral matchers passed per provider and neither created tasks. Cost coverage is partial; this is bounded workflow qualification. - [Fresh-dismissal E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/) passed 2/2 at `e5512a206`: native Claude and Codex, including fresh Cancel, clear composer, reopen, unrelated message, reload, late Blue answer, and actual agent acknowledgement. All five behavioral matchers pass per provider. Inspected the fresh-dismissal screenshots and saved pending/answered identity; neither created tasks. Cost coverage is partial (4/6 runs). - Prior evidence remains available in [the earlier campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/). Its early loading screenshot and overwritten final capture prompted the UI-ready, distinct screenshot fixes. ## Risks - The model interprets intent. The server verifies permission and provenance; it does not infer consent from text. Ambiguous and unrelated replies are not approvals. - Historical questions can accumulate. They remain visible, pending, and answerable; no automatic answer or expiry is invented. - The change to completion gates is scoped to Agent Chat and ordinary historical inputs. Current-turn and governed approvals retain their existing controls. - Live qualification is limited to the selected stories. Broader native onboarding finalization remains separate work. - No database migration. Telemetry adds no fields or values; the contract and README document the commit boundary. Privacy review was requested on the PR. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser-test orchestration. The exact model ID and context-window size are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e5830887b |
perf: reveal task content sooner and parallelize issue reads (#14727)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - A task page must show saved replies quickly so a person can read the work. > - The title could appear while the conversation waited for unrelated metadata and transcripts. > - Thread requests also waited for enriched task details, while the server read several independent fields in sequence. > - This change shows saved content as soon as it is ready, starts thread reads earlier, and runs independent server reads together. > - Native event history takes priority over legacy log fallback, and mentioned tasks load on intent. > - Content-free timing spans make the remaining server delays visible without recording task content. ## Linked Issues or Issue Description **What happened?** Task titles and properties appeared quickly, but saved conversation content stayed hidden for several more seconds while metadata and run transcripts loaded. **Expected behavior** Saved replies and the task description should be readable without waiting for supporting history. Returning to a cached task should show content within a frame or two. **Steps to reproduce** 1. Open a task with saved comments and completed runs. 2. Delay the task activity and runs responses by five seconds in the browser. 3. Observe whether saved content remains hidden until those responses finish. 4. Navigate away and return to the task to check cached navigation. **Paperclip version or commit** The change was developed from `1b48e73e0` and rebased onto `44736c9c7`. **Deployment mode** Built from source, tested in an authenticated staging deployment and with local response replay. Related: #14667 overlaps the transcript reveal behavior and adds separate retry UX. This PR also changes navigation prefetch, parent metadata gates, native log fallback, server read scheduling, and timing spans. #12647 proposes a separate SQL predicate optimization in the runs service. #13597 and #13095 are earlier loading fixes. ## What Changed - Start activity and runs when the task page mounts, alongside task details and comments, using the route reference for shared query keys. Hover/focus prefetch does not start full history reads. - Reveal saved comments and descriptions while metadata and transcripts load. Preserve strict waits for linked-comment navigation and tasks with only runtime content. - For settled native runs, fetch legacy logs only when event history is empty or fails. Preserve live-log subscriptions for queued and running native runs. Fetch mentioned-task details on hover or focus. - Run independent issue-detail enrichment and run metadata reads in parallel while preserving recovery dependencies. - Add `Server-Timing` phases and opt-in OpenTelemetry spans, plus regression tests and observability documentation. ## Verification - **313 tests passed** across the initial seven focused component, cache, timing, and scroll suites. After review fixes, **313 tests passed** across five task-page, cache, prefetch, and live-transcript suites (`pnpm exec vitest run` with `--maxWorkers=1`; these sets overlap). - `pnpm -r typecheck` and `pnpm build` passed locally before the final UI-only review fixes. UI typecheck/build and `pnpm check:token-gates` passed after those fixes. The final commit also passes full typecheck and build in CI. - The full local Vitest run was attempted. Several unrelated embedded PostgreSQL fixtures failed to start, and parallel test workers hit timeouts. Focused reruns passed. A later local full-suite rerun was stopped after the complete CI suite passed; it is not claimed as a local full-suite pass. - Final commit `d1e147145`: **54 checks passed, 2 skipped**, including all server/workspace test shards, all eight browser E2E shards, full build/typecheck, Runner checks, release registry, and canary clean-install verification. [CI run](https://github.com/paperclipai/paperclip/actions/runs/36734642034). Greptile **5/5** after two reviews; both findings fixed and all review threads resolved. No merge conflicts. - Live browser tests on an existing task with two saved replies: median full reload to visible content fell from **1.92 s** (3 samples) to **1.40 s** (5 samples). Cached return fell from **421 ms** (1 sample) to **29 ms** (3 samples). These are observed samples, not a performance guarantee. - With activity and runs delayed by five seconds, saved content appeared in **1.38 s** on desktop and **1.33 s** on mobile. The inspected comment did not move when metadata arrived. Verified history expansion, task properties, pending-input navigation, dashboard return, and mobile layout. - A separate local replay with fixed responses reduced visible-content time from **5.12 s** to **2.15 s**. This isolates frontend behavior and is not a live-server benchmark. ## Risks Progressive history can change the thread after first paint. Existing anchor behavior is retained and covered by tests and delayed-response browser checks. Query aliases must stay aligned for invalidation. Parallel reads can increase short bursts of database work; dependent recovery operations remain ordered. Full reloads still depend on network and task-detail latency. No schema or authorization change. OpenTelemetry remains disabled without an operator endpoint. The added spans use a closed set of phase names and carry no task IDs, task content, or exception text. I checked `ROADMAP.md`; this is a performance fix within the existing task page. ## Model Used OpenAI GPT-6 in Codex. The runtime does not expose a more specific model ID or context-window size. The agent used reasoning, code editing, terminal tools, and Chrome performance profiling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a36cbffa9e |
fix(connections): broaden natural-language and aggregator search (#14725)
## Thinking Path > - Paperclip manages agents and the services they need for work. > - Agents use connection search to discover a setup path before they request access. > - Tool-only filtering hid channel and AI methods from this search. > - Requiring every query word to match rejected normal task descriptions. > - A small aggregator index also omitted supported apps such as Circleback. > - This pull request broadens retrieval and returns purpose-specific setup guidance. > - Agents can choose a relevant result while existing access and provider-choice checks still apply. ## Linked Issues or Issue Description **What happened?** A search such as “AgentMail create an email address and manage an agent mailbox” returned no usable result. “Help me find tools for circle back” also missed Composio's supported Circleback toolkit. Queries longer than 200 characters failed validation. **Expected behavior** Return useful native and verified aggregator matches from natural-language queries. Include channel/email methods when Chat connectors is enabled. Identify each method's purpose and the correct setup path. **Steps to reproduce** Use the queries above with `connections_search` from an active task. Enable Chat connectors for the AgentMail case. The regression suite reproduces these misses before the change. Related routing work: #13941. This change does not change the runner failure path or add channel setup to tool-only connection cards. ## What Changed - Rank name and capability matches. Accept extra words, split names, small spelling errors, and queries up to 4,000 characters. - Include tool, channel/email, and AI methods. Return company-prefix setup links for channel and AI flows. - Add a dated snapshot of 1,583 official Composio toolkit names and a refresh script. Merge duplicate MCP variants for search and link each support claim to official evidence. - Find authorized indexed aggregator namespaces within longer queries. Return multiple app matches when the agent needs to choose. - Prefer exact app names over fuzzy matches for other apps; retain existing AI readiness. - Preserve native preference, company and identity boundaries, administrative denials, and saved provider consent. - Add relevance and database regressions, extend native tool-authority coverage, and document search behavior. ## Verification - Red: 15 new assertions failed against the previous implementation; the existing baseline passed. Added red-green regressions for Motion versus fuzzy Notion and existing AI access during review. A further regression covers mixed ready/unconfigured AI results and their per-result setup guidance. - Green: all 103 tests in the eight focused shared, database, runtime-tool, fixture, and route suites pass on the latest commit. - `pnpm -r typecheck` and `pnpm build` passed. - Latest-commit CI passed: 54 successful checks and two skipped checks, including the full test matrix, browser E2E, typecheck, build, and canary dry run. - The long local `pnpm test:run` attempt began before the review fixes and retained transformed pre-fix search code; it also hit an unrelated timing failure. Fresh serial reruns of the affected search suites and three timeout cases passed all 124 tests. Parallel local route shards hit two additional database setup timeouts; both suites passed all 17 tests on a fresh serial rerun. The complete corresponding CI suites also passed. Duplicate broad local runs were stopped after CI completed. The local UI suite independently passed all 7,007 tests. - Greptile: 5/5 on `cc6a0180d`; all review findings resolved. - Browser verification passed in a disposable local instance through real process-agent search requests: AgentMail opened its setup flow with the requester selected; the saved Circleback choice produced the Composio setup card; a paragraph-length Notion query produced its setup card. No provider credentials or external accounts were created. - The browser test caught an invalid UUID-based setup URL. The fix uses the company prefix and has a regression assertion. - A 3,971-character catalog query found Circleback first in a local 10 ms spot check after sharing query preparation across the catalog scan. This is a single measurement, not a performance guarantee. ## Risks - Broader retrieval can return extra candidates. Named services rank first; agents must select the relevant method. - The public support snapshot can age. It proves catalog support, not account authorization or the availability of every requested action. - Channel and AI methods use existing setup links. The tool connection card still accepts tool methods only. - No schema, migration, credential, or runner lifecycle changes. ## Model Used OpenAI Codex (GPT-6). The session does not expose a more specific model identifier or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dd7fc1f90a |
fix: raise the native journal read limit to 256 MiB (#14711)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native sessions persist control-plane state so they can resume safely. > - The state includes committed provider history needed for recovery. > - The server, runnerd recovery, and durable control plane validate this state before trusting its identity. > - Their differing 64 MiB and 192 MiB limits can reject a valid journal before recovery. > - This pull request aligns all three local state limits at 256 MiB. > - Larger files remain bounded, while recovery can read larger valid histories. ## Linked Issues or Issue Description Refs #13882 Refs #14312 ## What Changed - Raise the server and runnerd recovery limits from 64 MiB to 256 MiB, and align the durable control-plane limit from 192 MiB to 256 MiB. - Add coverage for a valid history above 64 MiB and rejection above 256 MiB. ## Verification - Matching server recovery passed with more than 64 MiB of actual committed event payloads (128 events with 512 KiB deltas). - The real runnerd exact-authority resume regression with the test Codex provider passed with 193 MiB of valid JSON whitespace appended. It crosses the former 192 MiB core limit and confirms the same provider identity. This exercises the runner process and durable control plane with a simulated provider, not a live OpenAI API call. This test used approximately 1.15 GiB peak RSS. - The actual runnerd reader accepted valid 256 MiB JSON and rejected valid 256 MiB + 1 byte. The reader call took 231 ms; the fresh process peaked at 1,244 MiB RSS. - Server tests reject mismatched identity above 64 MiB and files above 256 MiB. - `pnpm -r typecheck`, `pnpm build`, and `git diff --check` passed. - Full local `pnpm test:run`: 13,730 passed, 575 skipped, 7 failed across 6 files. All failures were embedded PostgreSQL startup errors after five attempts. They affected agent hiring, instruction revisions, environment images, reviewed chat bindings, issue monitoring, and legacy continuation authority. The focused journal tests passed; the latest pushed head passed all ordinary CI checks. Superagent is the only blocking check. ## Risks - **Open review concern:** Greptile is 5/5, but Superagent is `ACTION_REQUIRED` with two P2 findings on the server and runnerd readers. Both flag the increased synchronous parsing and memory cost. This PR keeps the requested fixed-limit change small. It does not add a worker parser or a process-wide memory budget. This resource tradeoff needs review before merge. - Large state parsing is synchronous and can consume several times the file size in memory. - Remote checkpoint archive and expanded-size limits remain 64 MiB, so this change alone does not make larger remote checkpoint transfers portable. ## Model Used - OpenAI Codex, GPT-6, with delegated assistance from `gpt-6-luna` at high reasoning effort; tool use and code execution. The GPT-6 context window is not exposed in this task runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
44736c9c7c |
fix(ui): align runner commentary with chat replies (#14716)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent chat shows run commentary and the agent's reply in one thread. > - The runner commentary text starts 4px left of the following reply. > - The runner activity group omits the gutter that the reply bubble uses. > - This pull request gives runner commentary the same token-based gutter. > - Both text blocks now share one left edge. ## Linked Issues or Issue Description **What happened?** In agent chat, the first commentary line of a completed run starts about 4px left of the following agent reply. This is visible in the Codie conversation. **Expected behavior** Runner commentary and ordinary agent reply text should have the same left edge. **Steps to reproduce** 1. Open an agent chat with a completed run that has commentary and an agent reply. 2. Compare the first commentary line with the following reply at the left edge. 3. Inspect the wrappers. Before this change, `task-chat-phase-interstitial` has zero left padding and `task-chat-agent-bubble` has 4px. ## What Changed - Add the existing `px-1` spacing token to runner commentary. - Add a regression test that compares the runner commentary gutter with the ordinary agent reply gutter. ## Verification - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui exec vitest run src/components/task-chat/TaskChatRunnerActivityGroup.test.tsx src/components/task-chat/TaskChatTurn.test.tsx src/components/task-chat/TaskChatBubble.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/ui build` - A full repository typecheck and build passed earlier on this branch. All latest-commit CI gates passed, including the server shard after a transient preview readiness failure was rerun. - Browser render: runner commentary and reply paragraphs both start at x=32px after the change. ## Risks Low risk. This changes only the horizontal padding of runner commentary. Long commentary lines may wrap 8px sooner. The full local test suite was stopped after the code changed during its run; the focused UI tests and latest-commit CI suite passed. > I checked `ROADMAP.md`. This bug fix does not duplicate planned core work. ## Model Used OpenAI GPT-6 in Codex. The runtime did not expose a more specific model ID or context window size. The agent used reasoning, shell tools, and read-only browser inspection to diagnose the layout and verify the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d30b03bd8c |
test: add persistent E2E coverage for human blocker decisions (#14707)
## Thinking Path > - Paperclip manages work for AI agents. > - Agents use the coordination skill when work needs human authority or a scope decision. > - PR #14188 replaced automatic manager escalation with direct blocker handling. > - This behavior needs real browser, server, database, and provider tests. > - The test must verify saved human input, task ownership, and resumed work. > - This pull request adds six reusable Product E2E cases and improves the skill examples that they exercise. ## Linked Issues or Issue Description Refs #14188. The merged change needs repeatable behavior coverage. The new suite tests missing administrator access, missing hiring permission, and requester scope questions. Searches found no duplicate blocker-guidance suite. This extends the existing eval system described in ROADMAP.md. ## What Changed - Add the explicit-only `blocker-guidance` Product E2E suite. It has three local scenarios on legacy Codex and legacy Claude. - Use the production UI and public APIs to create work, save a human-only question or confirmation, answer it after reload, and resume the same task. - Check requester identity, ownership history, manager activity, hiring, saved answers, and completion. Keep direct text input as a separate UX result. - Save pending and final screenshots, API checkpoints, skill hashes, provider evidence, and billing data through the existing report pipeline. - Isolate the Claude fixture home. Verify the served skill bytes before dispatch so an old installed skill cannot silently replace the evaluated skill. - Improve the coordination and hiring skill examples. Include the human-only policy, requester address, wake behavior, and handling of authorized scope changes. - Grader v5 requires the exact approved public welcome note as a new worker comment. Browser input checks reject unwritable scope cards before clicking, and confirmation direction must be saved in the resolution before the worker wakes. - Add grader calibration and browser-input tests. Update the fixture guide and generated capability inventories. ## Verification - `pnpm build`: passed after rebasing onto current master. - `pnpm -r typecheck`: passed. - `pnpm test:e2e:runner:typecheck`: passed. - `pnpm test:e2e:runner:unit`: 742 tests passed. - `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests passed. - `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells found. - Capability inventory and generated-contract checks: passed. - `pnpm exec vitest run server/src/__tests__/hiring-operational-examples.test.ts`: four tests passed after synchronizing the generated API reference and section anchor. - Full general and serialized test suites: passed in CI on `6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm test:run` was interrupted after complete CI coverage passed; it is not claimed as a completed local full-suite run. - Final GitHub checks: 54 passed, two optional Storybook checks skipped. The runtime-exposure startup test hit a 10-second readiness timeout once, passed a targeted local reproduction, and its CI shard passed the single retry without code changes. - Current-head Greptile: 5/5, clean check, zero unresolved threads. - Historical live measurement on September 29 at `4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex `gpt-5.6-sol` passed 8/9. These runs predate this rebase. - Version 5 changes the scope answer to an exact approved publication draft. The historical runs do not qualify that new requirement; the two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec` passed Codex and failed Claude. Claude posted the correct salary-free sentence but omitted its required reference line from that comment, placing the reference in a separate completion message. The `public-welcome-note` check correctly failed. An earlier Claude database-startup failure was retained separately; its fresh-instance retry reached the model. This pilot is not a six-cell qualification. - The failed Codex scope case omitted `addresseeUserId`. The strict routing check remains. All 18 attempts had clean evidence manifests and passed cleanup. - To repeat with provider credentials: `pnpm test:e2e:runner -- --suite blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is excluded from `--all`. ## Risks - The live suite measures variable model behavior. The retained 17/18 historical result and the current 1/2 scope pilot are not all-pass qualifications. These paid cases are opt-in; their observed model failures remain visible independently of deterministic CI checks. - A separate generic task-replacement diagnostic still exposed a Claude refusal. The ordinary cases use specific business decisions. The diagnostic is not a standalone catalog case in this change. - Earlier measurements included an old installed Claude skill and test defects. Their grades remain retained and are not combined with the three final repetitions. - Skill examples can affect when agents ask for human input. Downstream permission checks still apply. - Native runners, Daytona, agent-requester routing, and real external connection authorization are outside this suite. - Raw provider traces and credentials remain private. No screenshots, raw reports, secrets, workflow changes, or lockfile changes are committed. ## Model Used OpenAI GPT-6 through Codex assisted with this change. The exact deployed variant and context window size are not exposed in this session. The assistant used reasoning, repository edits, tool use, and shell execution. The evaluated models were `gpt-5.6-sol` and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1b48e73e0b |
feat(ui): add secondary navigation for agent chat (#14706)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat already provides a persistent conversation with each agent. > - Its shortcuts share the primary navigation and do not give chats a dedicated place. > - People need to find agents, start a chat, and switch conversations without moving the page layout. > - This pull request adds a secondary chat sidebar and a landing page around the existing chat surface. > - The same conversation, composer, history, and context panel remain in use. ## Linked Issues or Issue Description Refs #13283 and #13420. This extends the existing experimental Agent Chat navigation after review of the component and page stories. It supports the CEO Chat roadmap item through the existing task-backed conversation model. **Subsystem affected** The board UI and the company-scoped conversation list API. **Current behavior** Chat shortcuts sit inside the primary navigation. There is no dedicated landing page with a searchable conversation list. A separate landing header also moves the sidebar when an agent is selected. **Proposed behavior** Show a Chat entry in primary navigation. Keep a searchable agent sidebar beside the chat content. The plus button starts or reopens the current user's single conversation with that agent. Keep the header and sidebar in the same positions before and after selection. **Reason and benefit** People can find agents and return to persistent conversations without leaving the chat area or creating duplicate chats. **Breaking changes** The experimental chat navigation changes. Explicitly adding a chat now resolves its conversation immediately. Direct visits to unused agent chat URLs remain read-only. The existing per-agent routes and message contracts remain compatible. No database migration is required. ## What Changed - Add an account- and company-scoped conversation list endpoint with the existing access checks, feature gate, and OpenAPI entry. - Add the live secondary sidebar, landing page, avatars, search, loading states, errors, and retry controls. - Make the agent picker wait for chat creation and display failures. Existing agents reopen the same conversation. A dismissed selection cannot close a reopened picker or navigate over a newer choice. - Preserve recent-activity ordering and terminated agents’ chat history. Scope live list refreshes to the current user’s conversation events. A failed historical-agent lookup leaves healthy chats usable and offers a focused retry. - Keep the sidebar and header stable across chat routes. Keep mobile selection in the navigation drawer. - Use the production components in Storybook. Prepare the theme and mobile viewport before mounting the page to avoid the startup flash. - Update product documentation, the design guide, and navigation tests. Replace old browser expectations for stars and recent shortcuts with persistent conversation and layout coverage. ## Verification - `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass after rebasing onto master. - Focused UI tests pass, including 105 sidebar, picker, and live-update checks after review fixes. The 33 conversation service and route tests and 10 OpenAPI checks pass, including ownership, feature gating, and concurrent creation. - The full local test run passed 14,163 server tests before three environment or timeout failures. The embedded Postgres startup, connector socket, and native runner failures all passed direct reruns. - Browser test-drive verification covers a real provider reply, add and reopen, persisted history after reload, no-match search recovery, mobile drawer dismissal, and top-aligned context panels. - Browser measurements confirm that the sidebar has the same position and dimensions on the landing page and an agent conversation. - Storybook builds and its add-and-reopen interaction passes. - The revised browser regression passes locally against a freshly built throwaway instance. It covers stable sidebar geometry, add/reopen uniqueness, drafts, search, history, and terminated-agent history after reload. The full CI browser suite also passes. - Latest commit `b323577d9523180104df4000eaceedea2772608c`: all 54 completed checks pass, including the complete server/workspace/browser suites, aggregate verification, build/typecheck, security scans, and canary packaging. The two Storybook jobs are skipped by their workflow conditions. [CI run](https://github.com/paperclipai/paperclip/actions/runs/36714052050). - Greptile reviewed this same commit at 5/5 with no remaining actionable findings; all review threads are resolved. - Reviewer path: enable Agent Chat, click Chat, use plus to choose an agent, send a message, switch away, and reopen that agent. One conversation must remain, with its history intact. ## Risks - The new sidebar lists persistent conversations instead of starred and recent shortcuts. - Chat creation is asynchronous. Errors stay visible in the picker, and delayed responses cannot navigate into a previous company or account. - The shell adjustment is limited to chat routes and preserves the existing conversation implementation. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and browser tools. The session does not expose the exact API model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #123` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d72389bee2 |
feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps gateway gives agents governed access to external tools. > - Browser Use Cloud can run browser work, but a tool result alone does not let a person watch or take over. > - A task needs a durable browser session, a visible viewer, and recorded costs. > - This pull request adds a Browser Use Cloud v4 connection and interactive browser tabs on tasks. > - People can follow the work, interact with the page, and retain the browser after the agent finishes. ## Linked Issues or Issue Description **Problem or motivation** Agents need governed access to Browser Use Cloud. People need to see and interact with the same browser from the task. A browser must remain available after a run finishes and appear at the correct point in the task feed. **Proposed solution** Add a native REST connection for the v4 API. Bind each session to its company, task, agent, and credential grant. Open its interactive viewer in the task side panel. Record provider costs as financial events. Use `browser-use-cloud` as the app and connector key. Keep its skill with the connector and deliver it only with authorized connection tools. **Alternatives considered** A v3 MCP connection would expose tools without the v4 lifecycle integration. An external viewer link would leave the task. A fixed viewer size would prevent pages from responding to changes in the task pane. **Roadmap alignment** This extends the governed Apps gateway and Connected Apps roadmap. It uses the existing task, grant, secret, approval, and financial records. The work was requested by the maintainer. A search found no duplicate Browser Use connector PR or issue. ## What Changed - Add the Browser Use Cloud app, brand asset, API-key connection, and profile settings under the `browser-use-cloud` key. - Bundle the `browser-use-cloud` skill with the connector. Keep it out of global `skills/` discovery. Deliver it only with authorized task/run connection tools. Remove retired connector skill keys from runtime overlays and preserve unrelated browser skills. - Expose seven v4 tools through the governed gateway and deliver them to native and CLI agents. - Persist sessions, browsers, runs, event and recovery cursors, shutdown leases, and cumulative cost accounting. Recover uncertain paid starts without replaying them. - Enforce task ownership, credential grants, approvals, revoked access, and budget limits. - Add interactive task browser tabs and compact chronological feed entries. Retain the viewer across tab switches and keep visible idle browsers open. - Add debounced automatic viewport fitting, standard size presets, and a viewer ownership lease. - Add lifecycle, authorization, accounting, viewport, UI, and Storybook coverage. - Add an idempotent database migration after the current master migration. Preserve deployed migration hashes. Migrate pre-release Cloud connection and financial keys without replacing grants, credentials, or browser history. - Document provider behavior, live acceptance results, and the lack of documented passkey forwarding. ## Verification - Full workspace typecheck and production build pass on the updated branch. - Token gates, brand asset validation, module boundaries, and migration ordering pass. - Cloud tests verify global skill exclusion, authorized task/run delivery, unassigned agents, disabled connections, revocation, adapter isolation, and secret exclusion. The existing AgentMail connector assignment test also passes. - Migration replay runs twice against existing browser work and financial records. It preserves the records and avoids duplicate costs. - The focused provider, app catalog, OpenAPI, connection gateway, and migration regression suites pass. Recovery coverage includes lost replies, process crashes, provider rejection, and browser arrival acknowledgement. - All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`, including the full test matrix, browser E2E shards, build, typecheck, security, and release canary. Two optional Storybook jobs are skipped. - Greptile is 5/5 on the same commit, with zero unresolved review threads. The corrected review uses the actual master-to-head diff. - The local `pnpm test:run` started and was stopped after the full CI matrix passed. It did not complete locally; the full-suite result above comes from CI. - Earlier live acceptance used an isolated company with a capped provider credential. The agent opened paperclip.ing, the embedded viewer accepted navigation, and the same browser stayed available after completion and tab switches. - The local Storybook build passes. Stories cover the panel, footer, feed entries, settings, lifecycle failures, and viewport modes with an offline viewer fixture. ## Risks - Browser Use charges for hosted work. Provider caps and local budget checks reduce exposure; reported costs can arrive after work completes. - Viewer and CDP URLs grant access to the browser. The server validates and restricts them. They are excluded from agent results and durable event data. - Runtime resizing of v4 agent browsers uses a provider option confirmed by live testing but absent from its published agent schema. Resizing during a click may invalidate coordinates. Fixed presets remain available. - Viewport ownership is process-local and resets on restart. The lifecycle and accounting records remain in the database. - The original intermittent embedded-viewer stall has not been fully diagnosed. A bounded reconnect and active-session recovery cover the observed failure paths. - Live tests did not cover every revocation, approval, rate-limit, or restart case. Deterministic integration tests cover those paths. Passkey forwarding is not claimed. - Unknown create outcomes keep the credential available for cleanup. Run-list absence cannot prove a paid POST was rejected, so recovery stays pending until it can identify provider work. ## Model Used OpenAI Codex, GPT-6. Used reasoning, repository search, code execution, browser interaction, and test tools. The exact serving model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4f9a7c613 |
test(runner): guard continuation after journals exceed 2 MiB (#14312)
Add an actual runner resume regression above the former 2 MiB journal boundary and an explicit-only three-turn Daytona workflow that grows real execution history. Verify journal size and distinct completed tool calls without exporting private payloads. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2f6fa3b6dc |
fix: recover provider authentication inside tasks (#14629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a working model provider connection to run a task. > - A provider can reject a stored credential after the task starts. > - The failed run must ask the responsible user to repair that connection. > - This pull request adds that request directly to the task and reuses Connections sign-in. > - The user can choose an API key or subscription, then continue the task with a fresh session. ## Linked Issues or Issue Description **What happened?** A run that ended with `acpx_auth_required` or another known provider authentication error did not immediately offer an inline way to connect the provider. A repair form could also lock the user to the failed account's sign-in method. **Expected behavior** Show a provider connection card in the task as soon as the authentication failure is saved. Allow the responsible user to connect or repair the provider with any supported sign-in method. Keep the connection name automatic and resume the task after successful setup. **Steps to reproduce** 1. Run a task with a supported provider and an expired or invalid credential. 2. Let the run fail with a provider authentication error. 3. Open the task and attempt to repair the connection. Related work: Refs #13724 and #13726. This change adds the inline task repair flow and method choice. ## What Changed - Classify provider authentication failures and create one connection request for the current task. A persisted blocked classification suppresses automatic retries only after the repair card is created; unsupported providers retain their existing recovery path. - Mark only the attributed, unchanged credential as needing sign-in. Preserve credentials that were refreshed after the failed run started. - Reuse the provider sign-in controls inside the task. Allow API key and subscription choices for Claude, Codex, and Grok. Keep names hidden and generate a default from the user, provider, and method. - Keep the existing account when reconnecting with the same method. Create and select another account when the method changes. Validate updates to explicit agent bindings through the normal agent save path. - Require explicit adoption for legacy agent authentication. Validate in the agent environment, then commit the binding, connection install, audit, and card completion in one transaction. Keep failed setup and account selection visible and retryable. - Add regression tests and update the specification and Connections documentation. ## Verification - Fresh local verification: 199 tests passed across the inline form, provider method selector, default naming, authentication and recovery classifiers, run liveness, OpenAPI routes, database adoption/rollback, and Cursor execution suites. The adoption database suite also passed against disposable Docker PostgreSQL. - Full repository `pnpm build` and `pnpm -r typecheck` passed on the latest commit. Token gates are clean. - Embedded browser: opened real task cards from seeded authentication failures; switched Claude from API key to subscription and back; switched Codex from subscription to API key; confirmed the name field stays hidden. Provider sign-in was not completed with real credentials. - The broad local `pnpm test:run` started before review fixes and was interrupted after the working tree changed; it is not counted as a passing full run. Fresh focused tests passed. CI supplies the full test and browser suite results for the current commit. - CI is green on commit `4b97a4e447045ff3d7516525a187a5d1d21e0d4c`: 54 checks passed and two Storybook checks were skipped by their path rules. The workspace preview job passed on one rerun after a local-server startup timeout; its rerun passed 835 tests. - Greptile is 5/5 on the same commit with no actionable findings and no unresolved review threads. ## Risks - Incorrect authentication classification could prompt for a connection unnecessarily. Tests exclude tool authorization, quota, and unrelated runtime failures. - A method change selects the new personal provider default, which also applies to other agents that use that user's default. Explicit account bindings use the existing permission and runtime validation path. - Credential invalidation must not race with refresh or reconnect. The code compares the saved credential generation and grant update time under locks. - No database migration or new credential storage format is required. ## Model Used OpenAI GPT-6 through Codex. The exact model ID and context window size were not exposed in this session. Capabilities used: reasoning, repository editing, shell commands, database tests, and embedded-browser interaction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused tests listed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fc36d1ccd3 |
test(e2e): qualify large Daytona Git workspace continuation (#14316)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can keep one agent process and workspace across human review turns. > - Large workspaces must survive each transfer between Daytona and the host. > - Small fixtures do not cross the previous 32 MiB filename-output limit. > - Correct files alone do not prove that finalization and recovery have settled. > - This pull request adds an explicit three-turn test with 60,000 files and independent host checks. > - The test detects lost files, replaced processes, and stale finalization retries. ## Linked Issues or Issue Description Refs: #14253, #14314, #14315, #14402, #14420. This test covers the workspace Git streaming fix in #14253 and the Daytona archive validation and native finalization fixes consolidated from #14315 and #14314 into #14402. These runtime changes are merged into master. This PR adds regression coverage. ## What Changed - Add one explicit-only native Codex Daytona test. Broad matrix runs do not select it. - Create 60,000 small untracked files through ordinary provider execution. Independently check all host file contents and the 39,828,890-byte filename manifest after each turn. Each continuation changes all generated file contents, so stale host copies fail. - Check spaces, newlines, leading hyphens, Unicode, and glob characters in filenames. - Require committed native finalization, successful workspace receipts, no active transfer or runless cleanup, and no scheduled recovery before each continuation and after the final turn. - Use a fixed external instruction bundle so the same runner PID and process identity can continue across the three browser-driven turns. Managed agent folders intentionally stop the process for file collection after #14420. Set the 20-minute idle window on the environment, then verify the admitted policy on every run. Set a 25-minute Daytona auto-stop window for this large-file test. - Save public workspace-operation evidence when an E2E attempt fails. - Bound this large-file fixture to 15 minutes per turn and 50 minutes total, reserving five minutes outside the turns for setup, host verification, and cleanup. A measured CI continuation succeeded in 11 minutes 5 seconds, exceeding the ordinary warm fixture's 10-minute deadline. The ordinary fixture and all file, process, finalization, and cleanup assertions remain unchanged. ## Verification - Final PR head: `b7eed5517426507f5d7912edb8ae49163d1783b0`. - `pnpm test:e2e:runner:unit` — 56 files and 708 tests passed locally. The catalog retains all 407 existing cells and adds this one explicit-only cell. - `pnpm test:e2e:runner:typecheck` — passed locally. - [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/36640920233) — all gates passed, including repository typechecks, tests, build, browser suites, and canary dry run. The PR has 54 successful checks and two expected skips. Fresh Greptile reviewed all seven files at 5/5; all review threads are resolved. - [Live single-cell Daytona verification](https://github.com/paperclipai/paperclip/actions/runs/36637283902) — passed on the first attempt in 1,822,299 ms (30m 22s) on `227074b7b`, using native Codex `gpt-5.6-sol` and the verified Daytona image. All three turns independently verified every one of the 60,000 file contents, 39,828,890 filename bytes, and five unusual names. All runs committed with one stable runner PID/process fingerprint, native/provider sessions, runner instance, and sandbox. Final browser/download assertions, all nine matchers, explicit cleanup, and report publication passed. Results are published through the [Product E2E history](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/) and [eval hub](https://pages.paperclip.ing/evals/). - The final `b7eed5517` follow-up only extends the total allowance from 45 to 50 minutes, updates its catalog assertion/version, and documents the setup/cleanup margin. Per-turn limits and behavioral assertions are unchanged from the successful live run; that paid run was not repeated for this allowance-only follow-up. - The [first integrated-head live attempt](https://github.com/paperclipai/paperclip/actions/runs/36632127364) is retained: turn 1 passed, then turn 2 hit the old 10-minute deadline while finalizing. Its trace records successful completion after 11m 5s and successful cleanup. Process diagnostics also showed the intentional managed agent-folder stop boundary introduced in #14420. These observations motivated the larger turn budget and fixed external instructions. - The broad local `pnpm test:run` began alongside the build and encountered server setup and port-test failures. The affected setup suites and assertions passed on focused reruns after the build, using canonical macOS temporary paths where needed. The redundant broad local run was stopped; full-suite success is established by final-head CI, not by that local run. ## Risks - The live test makes provider calls and creates a billable Daytona sandbox. It runs only when explicitly selected and deletes its sandbox during cleanup. - The test creates 60,000 files and can take several minutes per transfer. Its longer retention window applies only to this fixture. - The test will fail on a runtime that does not include all three required fixes. ## Model Used OpenAI GPT-6 in Codex assisted with code, terminal tools, and test analysis. The exact serving model suffix and context-window size were not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1778075155 |
fix(server): continue unfinished tasks after status replies (#14626)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs report their task outcome through a structured finish result. > - A Board comment can permit a passive wait for the next response. > - That exception accepted reports that also admitted blocking unfinished work. > - The task then stayed In Progress without a runner, and recovery treated the wait as healthy. > - This pull request rejects that contradiction and uses the existing bounded continuation path. > - Real wait conditions and protection against obsolete requests remain in force. ## Linked Issues or Issue Description **What happened?** An agent answered a status inquiry with `yielded` and `response_wake`. The same result listed blocking remaining work. The server accepted an indefinite wait without a question, approval, dependency, or pause. No further run was queued. **Expected behavior** Unfinished ordinary tasks must continue or have a recorded reason to wait. A status reply alone must not suspend the work. **Steps to reproduce** 1. Add a Board status inquiry to an unfinished assigned task. 2. Submit a successful native result with `yielded`, `response_wake`, and `remainingWork[].blocksCompletion: true`. 3. Leave the task without any real wait condition. 4. Observe that the old policy preserves In Progress with no continuation and suppresses recovery. **Paperclip version or commit** Reproduced in the native status policy at `da887ea3e`. The branch is based on current master. **Deployment mode** Authenticated server with the native Paperclip Runner. Related work: Refs #13338 (native response waits and recovery). Refs #12071 (separate legacy recovery and retry-state work). This change fixes the native unfinished-response-wait exception. ## What Changed - Reject contradictory finish reports while the provider can still correct them. - Route accepted unfinished response waits through the existing one-follow-up continuation budget. Repeated incomplete results create a visible recovery action. - Preserve questions, approvals, dependencies, pauses, conversation lifecycles, and superseded Board requests. - Recheck the current Board source in the decision transaction before queuing repair. - Let normal recovery reconsider old committed waits that report blocking work and still have a current source. - Add policy and database regression tests. Document the rule. ## Verification - `pnpm exec vitest run server/src/services/native-runtime/status-arbiter.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts --no-file-parallelism`: 346 tests passed, no skips. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `git diff --check origin/master...HEAD`: passed. - Final head `8f0824d9bb4d631f2347f8afffe5fff6d574cb69`: all CI checks green (54 passed, 2 intentionally skipped), including all general tests, serialized server suites, runner tests, eight E2E shards, and the canary dry run. - Greptile: 5/5 on the final head, with no unresolved review threads. - The broad local `pnpm test:run` encountered an unrelated timing failure in `workspace-runtime.test.ts` (waiting for managed process-tree listeners). That test passed on an isolated rerun. The duplicate broad run was stopped; the complete CI suite passed on the final commit. - The transaction-race regression reproduced an obsolete continuation before the fix. All 12 source-change cases now pass, covering passive waits, corrective continuations, and exhausted-repair decisions. ## Risks - Agents that previously parked unfinished ordinary tasks must now continue or record an actual wait condition. - Old contradictory waits become eligible for normal recovery. Existing ownership, budget, pause, and supersession checks still apply. - The rule uses the structured blocking-work flag. It does not infer omitted work from prose. - No schema, dependency, or UI change. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and test execution. The session does not expose a more specific deployment identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e912f0df53 |
fix(ui): open text attachments in task tabs (#14297)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent work often ends with a Markdown or plain-text file. > - Task attachments currently open outside the task panel. > - Users need to inspect those files while keeping the task conversation in view. > - This pull request opens text attachments in task tabs and adds rendered, raw, and download controls. > - The same controls work in the mobile task drawer. ## Linked Issues or Issue Description **What happened?** Opening a text attachment did not put its content in a task tab. Markdown files had no in-task rendered/raw toggle. **Expected behavior** Open Markdown and text attachments in one reusable task tab. Show Markdown as rendered content or raw text. Download the original file. **Steps to reproduce** 1. Upload a Markdown file and a plain-text file to a task comment. 2. Open each attachment from the task conversation or artifact list. 3. Switch Markdown between Rendered and Raw. Download both files. 4. Repeat at a mobile viewport width. Related work: #14193 controls artifact tab arrival. This change adds text attachment content tabs. ## What Changed - Route text attachment opens from conversation and artifact cards into task tabs. - Add a text attachment panel with accessible Rendered, Raw, and Download controls. - Preserve ordinary links for other file types. - Support the selected attachment in the mobile drawer. - Keep text-tab actions on the current rich artifact cards, including CSV previews. - Render attachment image references and diagram source without loading media URLs. - Add browser regression tests, component tests, Storybook examples, and usage documentation. ## Verification - Full workspace typecheck, production build, Storybook build, and UI token gates pass locally. - All 6,960 UI tests pass. The additional media regression passes against the real Markdown renderer and fails before the fix. CSV coverage verifies direct downloads and text tabs after preview. - Both desktop and mobile browser cases pass locally. They check rendered/raw Markdown, literal plain text, reusable tabs, review controls, and exact original download bytes. The local fixture used a separate database port because an existing socket occupied the default range. - The full CI test matrix passes on `82be5efbef26927b237a031725bb3d7fa79f637f`. The duplicate local `pnpm test:run` was stopped after this CI result; it did not complete locally. - Greptile is 5/5 on the final commit. All review threads are resolved, and the security scan passes. - All 54 final-head checks pass, including the canary dry run. The two optional Storybook jobs are skipped. ## Risks - Text attachment links now open in the task panel. Other content types keep their existing link behavior. - File display still depends on the existing authenticated attachment route. There are no API or database changes. - Raw text is displayed as text, including strings that look like HTML. Rendered Markdown keeps media references inert. ## Model Used OpenAI Codex, based on GPT-6, with code execution, browser testing, and subagent tool use. The runtime does not expose an exact serving model variant or context-window size. Recovered earlier implementation changes were reviewed and tested; their exact model metadata is unavailable. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
81a52eb740 |
fix(ui): allow touch scrolling in new-task pickers (#14599)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The new-task composer lets users select an assignee and override its model. > - On phones, these selectors open as large sheets outside the task dialog DOM. > - The parent dialog's scroll lock cancels touch drags in those sheets. > - This pull request gives each mobile sheet its own modal scroll boundary. > - Users can scroll to an option, select it, and continue editing their task. ## Linked Issues or Issue Description **What happened?** An iPhone Safari user could not drag through the assignee or model list in the new-task composer. The list stayed at the top. A touch-enabled Chromium reproduction also showed canceled touchmove events and an unchanged scroll position. **Expected behavior** A finger drag scrolls the list. A tap selects an option. Closing the picker preserves the task draft and selected values. **Steps to reproduce** 1. At phone width, open New Task in a company with enough agents to overflow the picker. 2. Open Assignee and drag upward through the list. 3. Select a Codex agent, open Codex options, choose Custom, and open the model selector. 4. Repeat the drag with enough models to overflow the available viewport. **Paperclip version or commit** Reproduced on `e5bf9d49a`. The fix is rebased onto `b3eb03fcb`. **Deployment mode** Local development in an isolated test instance. The user reported iPhone Safari. Browser automation uses native Chromium touch input at phone dimensions; a physical iPhone was not available. Related: #14250 introduced the large mobile entity picker sheets. Duplicate search found no existing fix for their touch scroll boundary. ## What Changed - Use modal Radix popovers for mobile entity sheets. Their lists can scroll while the background stays locked. - Suppress opening when Radix restores trigger focus after dismissal. Escape and outside taps now close the picker without reopening it. - Add a failing-before/fixed-after regression for a mobile sheet inside a parent dialog. - Add a browser test for native touch scrolling in both lists, tap selection, draft retention, close-button/Escape/outside dismissal, and desktop mouse/keyboard selection. - Document the browser test command. ## Verification - Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. - Passed 41 focused tests for `InlineEntitySelector` and `NewIssueDialog`. - Passed the new browser test against a disposable instance running this worktree. It uses 22 real fixture agents and a fixed 24-model catalog. It covers 390×844 and 390×430 phone viewports and a 1280×900 desktop viewport. - Inspected the rendered new-task form and both selectors. Selections returned to the draft with its title intact. - `pnpm test:run` was attempted and stopped after confirming that local database suites were skipping because macOS has exhausted its system semaphore limit (`initdb`: `could not create semaphores: No space left on device`). It did not complete locally. The complete test matrix passed in Linux CI, including all 71 tool-gateway tests. - All 150 browser tests passed in CI, including the new native-touch regression. - Greptile reviewed the latest commit at 5/5. Its desktop focus finding is fixed and the review thread is resolved. All checks on `2ab22727ca768b4134c9c840bbc8e4c80d7da674` are complete: 53 passed, 2 intentionally skipped Storybook deployment checks, and the Snyk status passed. ## Risks Low risk. Mobile selectors now own focus and scroll isolation. This changes their modal behavior, so the tests cover nested dismissal and focus return. Desktop selectors remain non-modal. No schema, API, or dependency changes. ## Model Used OpenAI Codex (GPT-6), with reasoning, repository inspection, code execution, and browser testing. The exact backend deployment ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b3eb03fcba |
fix(sentry): preserve run failure stacks and diagnostic context (#14585)
Builds on merged #14575 and targets `master`. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators use optional Sentry reports to diagnose failed agent runs. > - The shared reporter currently converts each saved error message into a new exception. > - This loses the original stack and cause chain, and omits available adapter and provider diagnostics. > - This pull request selects, redacts, and bounds useful diagnostics before it sends them to Sentry. > - Operators can locate the failing operation and correlate upstream failures without copying arbitrary run data. ## Linked Issues or Issue Description Refs #14573. That change preserves ACP provider failure details in the run record. Refs #14575. That merged PR adds recorded exit codes and signals. This PR covers stacks, causes, and structured diagnostics without duplicating those fields. A thrown setup or adapter error currently appears in Sentry with the reporter's stack. A saved provider failure can contain useful details that never reach the Sentry event. Both cases use the same shared reporting path. ## What Changed - Pass caught setup and execution exceptions, and structured adapter error metadata, through successful terminal status transitions. - Snapshot a fixed set of execution, adapter, provider, and exception fields. Preserve up to four exceptions in the cause chain. - Remove registered secret values, declared runtime environment credentials, unknown inherited environment values and encoded credential forms, and credential patterns before truncation. Mark truncated fields and bound provider details and stacks. - Rebuild sanitized Sentry exceptions with the original stacks and causes. Use an adapter stack preview when available. Omit a fabricated reporting stack when the source has no stack. - Keep contexts local to each event. Preserve the optional DSN gate and existing error-code/adapter fingerprint. - Document the fields, limits, and omitted data. - Settle leftover chat fixture outbox rows only after assertions and worker shutdown, so subsequent tests cannot claim earlier cases’ pending actions or provider I/O. This fixes the CI shard contamination exposed during verification; production chat behavior and test timeouts are unchanged. ## Verification - `pnpm -r typecheck` passed. - `pnpm build` passed. - 86 focused diagnostic, Sentry, and startup tests passed after the rebase, including the real `@sentry/node@10.71.0` SDK with an in-memory transport. - Tests cover cause chains, HTTP status and request IDs, long provider details, secret redaction, failed secret resolution, cyclic causes, size limits, and context isolation. - Five targeted heartbeat integration cases passed, including thrown and returned errors containing an opaque environment-bound credential. The earlier full heartbeat integration run also passed its integration cases. - The focused suite also passed with an unknown inherited environment value set to `1`; reporting tests use a controlled environment and separately verify short-secret redaction. - Before the fixture cleanup, the full chat shard reproduced the CI Telegram timeout at `pending recovery before restart` (354 passed, 1 failed). After the cleanup, the same shard passed all 355 tests. No timeout or production behavior changed. - Latest-head CI (`0ae9e70df321d66dda025c3a7ba4787e169e0a4f`) passed typecheck, build, all 12 server shards, all 3 chat shards, workspace and serialized suites, browser shards, Runner checks, canary validation, the real Sentry SDK contract, and security checks. Required `ci / verify` and `ci / e2e` passed. - The branch has been rebased onto `master` after #14575 merged. All 55 applicable checks passed on this head (2 unrelated Storybook checks skipped). Greptile re-reviewed the current 11-file diff at 5/5 with no findings or unresolved review threads. - Local monolithic full-suite attempts were interrupted to apply fixes; full-suite success is not claimed from those runs. ## Risks - Error messages and stacks can contain credentials. The reporter uses existing redactors and the run's encrypted secret registry, reads only known fields, and skips capture if registered-secret resolution fails. - Unknown environment values remain private by default. Unrecognized short values can mask benign matches; known public settings are explicitly allowed. - Diagnostic text is bounded and can be truncated. Truncation is explicit. Data discarded upstream cannot be recovered. - These additional fields go to the operator's configured Sentry endpoint. Arbitrary request/response objects, headers, configuration, prompts, and stdout/stderr are not copied. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model ID, context window, and configured reasoning level are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2de43fc909 |
fix(issues): keep agent mentions as context and defer personal app authorization (#14577)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Each task has one assignee. Explicit assignment and review requests select who should act. > - An agent mention started another agent on a task it did not own. Native attachment staging then rejected that run. > - Allowing that run through startup could also let two agents work on the same task. > - Mentions should identify relevant context. They should not start work or forward comments to other tasks. > - A personal app installed on a shared agent must also wait until tool use to resolve the current user's grant. > - This pull request removes mention dispatch and keeps missing personal app credentials from blocking startup. ## Linked Issues or Issue Description **What happened?** A native agent mentioned on another agent's task failed with `paperclip_runner_attachment_staging_not_authorized`. The source task could already be complete. A nearby optional-app warning was a separate problem: personal app tools were excluded when their shared health state required attention. **Expected behavior** An agent mention is context only. It does not wake the agent, take ownership, or copy a comment onto another task. Normal feedback still reaches the assignee. Assignment and explicit review requests still dispatch work. An unavailable personal app does not block startup or produce a startup warning. Tool use requests the current user's authorization and never uses another user's grant. **Steps to reproduce** 1. Assign a task to agent A. Post a comment that mentions agent B, including a comment that closes A's task or references B's child task. 2. Confirm the comment retains its agent link and B receives no run or deferred wake. A can still receive normal feedback. 3. Install an active personal MCP connection on B. Give only Alice a grant and leave shared health at `error`. 4. Explicitly assign work to B for another user. Confirm it can finish without using the app. 5. Ask B to use the app. Confirm its tool call shows an inline connection request for the current user. Related work: Refs #11144. This change uses the existing execution-time personal grant resolution. ## What Changed - Remove mention dispatch from standalone comments and issue updates. Remove implicit forwarding of parent comments to a mentioned worker's child task. - Ignore new requests with the legacy mention wake reason before creating a run or deferred request. Preserve already accepted queue entries, which can combine assignments and feedback with a later mention. - Remove the native mention admission, staging, and finalization exceptions from this PR. Native task ownership checks remain intact. - Keep active, installed personal app tools available despite shared health errors. Remove optional-app startup warnings. Tool execution retains the current user's grant and policy checks. - Update agent instructions and product/API docs. Refresh generated capability source anchors. ## Verification - Red: comment-route regressions reproduced extra agent wakes and child comment forwarding. A separate regression proved that cancelling by the last coalesced reason could drop an accepted assignment. - Green: the targeted route, wake queue, heartbeat, workspace, responsible-user, MCP discovery, and HTTP gateway suites passed. The final queue and heartbeat rerun passed 104 tests, the restored queue adapter passed 56, and both comment-route suites passed 135. These include accepted assignment preservation, rejection of new mention requests, and normal assignee feedback. - `pnpm -r typecheck` and `pnpm build` passed locally. The full local `pnpm test:run` attempt was interrupted for review/CI fixes, so it is not claimed as a completed local pass. It exposed a cleanup timing race in the concurrent-mention assertion, now fixed and verified across 10 repetitions. CI also exposed an obsolete test waiting for the removed mention lookup; it was reproduced and fixed, then both comment suites passed. Final full-suite verification is through CI. - Final head `bd9ea4cb05a8f081c54e017760a8999f9ea6ef44`: 54 checks passed, 2 Storybook checks intentionally skipped; no pending or failing checks. Full CI includes general and serialized suites, all 8 browser shards, runner verification, typecheck, build, and canary dry run. Greptile is 5/5 on this exact commit, with no unresolved findings. - One unchanged Cursor adapter test hit its 10-second CI timeout. All 5 tests in that file passed locally; one retry of its CI shard passed all 674 tests (3 skipped). The aggregate verification gate then passed. No code or timeout was changed for that retry. - Live browser check: inserted a structured mention with the picker on a human-owned task. The saved link remained visible. Database checks found zero new runs and zero wake requests. - Live Codex runner check: explicitly assigned that task with the unavailable personal app attached. The run succeeded and committed completion without using the app or creating a connection card. - Live browser follow-up: asked the assignee to call PostHog and mentioned another enabled agent as context. Only the assignee ran. It succeeded and displayed the existing inline connection card. Only Alice's grant existed; the run belonged to a different user. - The HTTP regression covers tool discovery with no provider calls or connection cards, first use returning the current user's authorization request, and successful retry after that user's grant exists. - App checks use an isolated local fixture and a fake MCP provider. They do not use production app credentials. ## Risks - Intentional behavior change: workflows that used mentions to wake agents must use assignment, a bounded child task, or an explicit review request. - Already accepted queue entries retain their prior rules. An old entry can combine assignment or feedback with a later mention; its last reason cannot safely identify mention-only work. New mention requests create no run or deferred wake. - Personal apps with a shared health error remain discoverable. Actual tool use still requires the responsible user's grant and existing policy gates. - No database migration or public API schema change. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact serving model ID and context-window size are not exposed in this session. - Live native-run verification used `gpt-6-astra` through the Codex provider. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
29c8fb0b66 |
fix(composer): match effort options to adapter execution (#14576)
## Thinking Path > - Paperclip is an open source control plane for AI agents. > - The issue composer lets operators choose a model and effort for a task message. > - The picker must show only effort levels that the selected adapter can execute. > - The Codex Runner fix in #14568 exposed similar gaps in other adapters. > - Grok hid a supported control. Claude used one range for all models. Kimi ACP showed a control that it ignored. > - This pull request aligns the picker with each adapter and adds matching stories. > - Operators can see and save supported effort settings without false controls. ## Linked Issues or Issue Description **What happened?** The composer hid Grok effort. It showed an incomplete effort range for some Claude models and a slider for Claude Haiku. It showed Kimi effort on the default ACP engine even though that engine drops the value. **Expected behavior** The composer should show only effort levels supported by the selected model and execution engine. Grok effort should reach its adapter as `reasoningEffort`. **Steps to reproduce** 1. Select a Grok agent and `grok-4.7`. Observe that the picker has no effort slider. 2. Select Claude Sonnet 5 or Haiku 4.5. Observe the generic low to high slider. 3. Select a Kimi agent on ACP. Observe the slider even though ACP ignores effort. Related fix: #14568. ## What Changed - Use Claude and Grok adapter capability helpers in the production picker and Storybook preview. - Map Grok composer effort to `reasoningEffort` in per-message overrides. - Show Kimi effort only when its agent uses the CLI engine, including when it runs its default model. - Show Grok effort when it runs its default model. - Add design and production stories for Claude, Grok, and Kimi ACP and CLI cases. - Document the picker rules and add focused tests. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/task-chat/composer-run-settings.test.ts` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` passed on the first commit. - The focused test, UI typecheck, token check, and Storybook build passed again after the review fixes. - Browser smoke: production stories show effort for default Grok and CLI Kimi, and hide it for ACP Kimi and Claude Haiku. ## Risks - Existing Kimi ACP agents lose a slider that could not change execution. Their stored effort override remains unchanged until the next edit. - Claude and Grok ranges follow their adapter catalogs. A provider can change model support before the catalog is updated. ## Model Used OpenAI Codex based on GPT-6. The exact runtime model ID and context window are not exposed in this session. Reasoning, tool use, and code execution were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
83076d7e7c |
feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage. Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3b4b270650 |
fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The shared ACP adapter engine records agent failures for operators. > - ACP providers can report a failure category, title, and detailed cause. > - Our patch kept only the category in the saved error, so an operator could not diagnose a failure when tracing was off. > - This pull request preserves redacted provider diagnostics in the run error, transcript, and structured run result. > - Operators can now inspect the provider message and any supplied request ID or stack trace after the run ends. ## Linked Issues or Issue Description Refs #13889 (the diagnostic gap; this PR does not update the bundled Claude version). Refs #14484 (related model-refusal classification; this PR retains diagnostics for all terminal failure categories). **What happened?** An ACP turn failed with only `ACP agent reported a terminal service failure.` The provider's title and details were available in memory but absent from the saved error and transcript. **Expected behavior** The run retains useful provider diagnostics even when raw tracing is disabled. Credentials remain redacted. A size limit must report truncation instead of silently removing the cause. **Steps to reproduce** 1. Run an ACP agent that returns an error-severity typed session failure. 2. Include an HTTP error, request ID, and stack text in its title and details. 3. Inspect the failed run with tracing disabled. Before this change, only the category survives. ## What Changed - Both pinned ACPX patches pass complete error text to the in-memory callback, so redaction happens before truncation. - The shared engine retains the sanitized category, title, and details in `resultJson.terminalSessionFailure` and includes the text in the run error and error transcript. - Diagnostics redact configured environment values even under arbitrary names, unknown launch-environment values, connection URL passwords, run credentials, and common credential syntax. Known boolean settings remain readable, while credential values are redacted even when embedded in other text. Diagnostics remove control characters and invalid Unicode. - Title and detail limits keep escaped transcript JSON below the server's chunk limit. Truncated fields include an omission count. The safe run-result projection preserves a byte-bounded diagnostic preview when the result exceeds its byte budget, with an explicit pointer to the full adapter-bounded run error and transcript. - The existing UI and CLI display the error. Diagnostics do not become assistant output. Issue continuation summaries and session-compaction prompts receive only the generic category, preventing provider text from becoming handoff instructions. Existing quota classification, warnings, timeout precedence, and control-channel failure precedence remain in place. - Regression tests cover real ACP child processes with both pinned versions in one-shot and persistent modes, credential redaction, request IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and database retrieval of oversized multibyte diagnostics. ## Verification - Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2 intentionally skipped, no pending or failing checks**. Includes typechecking, build, all Vitest shards, Runner checks, browser E2E, and the canary packaging/public-install dry run. - Greptile: **5/5** on this commit. Superagent security scan passes. All review threads are resolved. - Local verification passed: shared ACP engine suite (395 tests); real Claude ACP child-process and diagnostic regressions across both pinned runtimes and both execution modes; run retrieval and model-handoff regressions (59 tests); ACPX patch packaging (16 tests); full typecheck and build. Affected package typechecks and focused tests were rerun after review fixes. - The broad local `pnpm test:run` was stopped after review edits made its cached imports stale. Fresh targeted runs pass, including both affected server suites. Cold-build import failures were also rerun after dependency builds: chat integration (1,063 tests) and tool access (351 tests) pass. The final commit's complete CI matrix is green. ## Risks - Provider diagnostic text is untrusted. This change retains more of it in company-scoped run records. Redaction and size bounds apply before persistence. - Diagnostics are limited to fields the provider supplies. Old runs cannot recover discarded error text. - No schema migration, recovery-policy change, or new Telemetry or OpenTelemetry export. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and test execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
da887ea3e9 |
fix(runner): honor Codex effort selected in composer (#14568)
## Thinking Path > - Paperclip manages AI agents that work on assigned tasks. > - The task composer lets a person choose an assignee, model, and effort for the next run. > - A Paperclip Runner agent can use Codex as its provider. > - The composer hid Codex effort for that agent because it checked only the older Codex adapter. > - The native Runner input also did not carry an effort choice to Codex. > - This pull request carries the chosen effort from the composer to each Codex turn. > - People can now select a supported effort and get the effort they selected. ## Linked Issues or Issue Description Refs #14322 **What happened?** The composer showed a model but no effort slider when the assignee used Paperclip Runner with the Codex provider. A task-level model override also did not reach the native Runner input. **Expected behavior** The composer shows effort choices for a known Codex model. The next native Codex turn uses the selected model and effort. **Steps to reproduce** 1. Open a task composer. 2. Select an agent that uses Paperclip Runner with the Codex provider. 3. Select a known Codex model such as `gpt-6-astra`. 4. Open the assignee and model picker. The effort slider is missing before this change. ## What Changed - Show known Codex effort levels for Paperclip Runner Codex assignees. - Save the task effort override in the native run input and send it to Codex on each turn. - Apply the task's merged model and effort overrides when the native run starts. - Apply a task model override for OpenCode Runner without changing the agent's provider. - Add Runner effort tests and desktop and mobile Storybook cases. ## Verification - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm build-storybook` passed. - `pnpm check:token-gates` passed. - Focused UI, server, Runner contract, and Codex driver tests passed. - The full CI test matrix, build, typecheck, and canary dry run passed on the latest head. ## Risks - Native Runner inputs add an optional Codex effort field to the current v5 input. Older inputs keep their previous behavior. - A known model rejects an effort that its catalog does not support. Unknown models do not show a slider. > This fixes an existing composer bug. I checked `ROADMAP.md`; it does not describe this bug as planned work. ## Model Used OpenAI Codex, GPT-6. The exact deployment ID and context window are not exposed in this session. The model used reasoning, code execution, and repository tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com> |
||
|
|
3f4f8b37ba |
fix: grade Codex clarification and refusal outcomes from evidence (#14570)
Grade clarification lists, obsolete unstarted wakes, and refusal cancellation from persisted evidence. Preserve execution and ownership assertions, add boundary regressions, and version the affected eval definitions. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
7636966452 |
fix(inbox): keep other users’ failed runs out of Mine (#14572)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Mine inbox shows work that needs the current user. > - Failed-run rows used the latest run for every agent in the company. > - A failure from another user therefore appeared in Mine and its badge. > - Run list responses also omitted the responsible user needed to filter these rows. > - This pull request uses run ownership for personal failure routing. > - Users see their own failures and can still inspect company failures in All. ## Linked Issues or Issue Description **What happened?** An agent run started for one user failed. Its row and failure badge appeared in another user's Mine inbox. **Expected behavior** Mine and its badge include failed runs for the current responsible user. Other users' failures remain available in All and run details. **Steps to reproduce** 1. Use a company with two human users. 2. Create a failed or timed-out run attributed to the first user. 3. Open Mine as the second user. Before this fix, the failed run appears there and increases the badge. **Paperclip version or commit** Reproduced in regression tests on master at `24beb0057`. **Deployment mode** Authenticated deployment with multiple users. Tests also cover the local single-user board. Related prior work: #933 addressed inbox dismissal and badge consistency. No duplicate ownership fix was found. ## What Changed - Return `responsibleUserId` in normal and summary run lists. - Share one ownership rule across both inbox versions and client/server badges. - Select the latest run per agent before applying the ownership filter. This prevents old failures from resurfacing on shared agents. - Keep unattributed historical failures in the local board's Mine view. Hide them from authenticated users with no matching owner. - Keep company health alerts outside the personal badge, consistent with the client. - Document the routing contract and add page, badge, and database regression coverage. ## Verification - Red: the new badge cases failed with three company failures instead of one personal failure; eight Mine page cases failed across both inbox versions. - Green: 113 focused tests pass in `ui/src/lib/inbox.test.ts`, `ui/src/pages/Inbox.test.tsx`, `server/src/__tests__/heartbeat-list.test.ts`, and `server/src/__tests__/inbox-dismissals.test.ts`. - `pnpm check:token-gates` passes. - Agent calls on behalf of a user have two additional red-to-green API regressions. - Full `pnpm -r typecheck` and `pnpm build` pass. Server typecheck also passes after the agent-call fix. - All CI test shards and browser tests pass on `243bfa681`. The duplicate local `pnpm test:run` was stopped after the CI test lanes completed; it did not finish locally. ## Risks - Authenticated users no longer receive unattributed legacy failures in Mine. Those failures remain visible in All. - The server badge no longer counts company health alerts, matching the existing client badge. - No migration, run state, retry behavior, or company access rules change. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, terminal execution, and browser tools. The exact deployment variant and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c9b93d7e8c |
fix: preserve terminal task owners during release (#14561)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks record an assigned owner and separate checkout and execution locks. > - A completed task must retain its owner after execution ends. > - The release endpoint currently clears that owner when it clears the locks. > - This pull request preserves the assignee of Done and Cancelled tasks during release. > - Unfinished tasks keep the existing relinquishment behavior. > - The benefit is stable task attribution without retaining execution locks. ## Linked Issues or Issue Description **What happened?** An agent completed an assigned task, then called the release endpoint. The task stayed Done, but its assignee became null. The same defect affects Cancelled tasks. It caused the legacy Claude clarification/reuse and multiple-repository handoff E2E assertions to fail. **Expected behavior** Release must clear execution locks on terminal tasks and preserve their assignee, final status, and disposition timestamps. Release of unfinished tasks must still clear the agent assignee. Only In Progress work returns to Todo. **Steps to reproduce** 1. Create an assigned task with checkout and execution locks. 2. Complete or cancel the task. 3. Call `POST /api/issues/:id/release` as the assigned agent. 4. Read the saved task. Before this fix, its assignee is null. **Paperclip version or commit** Reproduced on master commit `d172197117a14b80a1eb2d2835a0e7cce2679656`. **Deployment mode** Local tests against real PostgreSQL through the production issue routes and services. This is a core lifecycle defect, independent of the agent adapter. Refs: #11689, #6899, #7769. These are related open release proposals. This is an independent fix limited to terminal task ownership. It does not include timer scheduling changes. ## What Changed - Preserve the current assignee when releasing Done or Cancelled tasks. - Keep all execution-lock cleanup and existing unfinished-task behavior. - Cover all seven task statuses through the release API and read back saved state. - Check disposition timestamps, activity attribution, and repeated board cleanup. - Update the API contract, agent reference, and CLI help. ## Verification - Red commit `ddaabb754`: the two terminal-owner regressions failed with `assigneeAgentId: null`; 12 other route tests passed. - Green: all 14 route tests pass, plus the existing successor-checkout race test (15 selected tests total). - Command: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/issue-stale-execution-lock-routes.test.ts src/__tests__/issues-service.test.ts -t 'stale issue execution lock routes|does not let stale release clobber a successor checkout lock'`. - The local host has exhausted its SysV semaphore pool. The red/green runs used the existing test-provider hook to start disposable Docker PostgreSQL 17 instances. Routes, services, migrations, and assertions were unchanged. No database tests in the selected set were skipped. The other 134 tests were excluded by the name filter. - Capability contract and inventory drift checks pass. - `pnpm build` and `pnpm -r typecheck` pass. - The local `pnpm test:run` was interrupted after environment failures while the complete sharded CI suite ran in parallel: native PostgreSQL bootstrap fails under the host semaphore limit, and the large Git fixture hits macOS `ENAMETOOLONG`. A focused rerun confirmed these happen before the relevant assertions. The interrupted local run is not counted as a full pass. - Greptile completed on `1caeeb827e9cb658ddb71f16c2421ec20f80634e` with **5/5**, a successful check, and no review threads. - All CI gates pass on the current head: typecheck, build, general and serialized tests, Runner checks, browser E2E, release packaging, and security checks. Server shard 11 passed on one targeted retry; the first attempt had 836 passing tests but an unhandled workspace-runtime startup rejection caused by an existing timing window. All other successful jobs were reused. - No paid provider evaluations were run. ## Risks - A caller that used release to erase ownership from terminal work will now retain that owner. An explicit assignment update or the board force-release option with `clearAssignee=true` can still clear it. - No schema or migration changes. The transaction, company access, assignee/run checks, and activity log remain in place. ## Model Used - OpenAI GPT-6 through Codex, with tool use, code execution, and test debugging. The session does not expose a more specific backend model version or context-window size. - OpenAI `gpt-6-luna` assisted with read-only test discovery and review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
61b3fd57a6 |
fix(ui): recover gracefully during server restarts (#14560)
Show a clear reconnecting state during server restarts and retry safe access checks every five seconds. Preserve open drafts, wait for initial startup readiness, and keep authentication failures separate from temporary outages. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3ca196b0a6 |
feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - An agent needs personal files across tasks and sessions. > - AGENTS.md is one file in that directory. Supporting files need the same persistence. > - The Instructions Editor and agent runs must share one current directory. > - Concurrent runs should apply only the files they change. The last sync of the same file wins. > - This pull request uses existing file transport and removes temporary copies after sync. > - Old instruction-only sessions keep their restore contract. New saves do not create revision history. ## Linked Issues or Issue Description Refs #14325. This replaces its revision-oriented design with persistent agent files. Keep #14325 unmerged. Transport prerequisite #14416 merged first at `d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master and remains below 100 changed files. Related work: #4513 and #8798 cover instruction tooling. This change handles run synchronization, cross-task personal files, browser editing, and old-session restoration. ## What Changed - Keep one current directory per company and agent. Point AGENT_HOME at a temporary working copy for each active run. Keep task files and provider HOME separate. - Restore text, binary files, and nested folders through workspace transport. Exclude remote agent files from task Git snapshots with a self-ignoring file inside the reserved runtime directory; never write through repository-controlled Git metadata. - Collect after the provider and child processes have stopped. Keep resumable conversation state. - Apply changed and deleted files under the agent lock. The last sync wins for the same file. Unrelated concurrent changes survive. - Remove temporary copies after successful sync, rejected sync, and staging failure. Register ownership before copying so restart recovery can remove interrupted preparation. Retry transient synchronization up to three times. Preserve the original remote lease reference until deletion succeeds; restart cleanup never acquires a replacement sandbox. Do not create captured directories or a conflict-review queue for new runs. - Keep browser editing, stale-draft protection, and streaming binary downloads. Keep the instruction entry and text editor limited to 1 MiB. - Keep historical agent-folder sync failures on their affected runs instead of repeating them above current saved instructions. Preserve legacy candidate review and current browser-save errors. Avoid duplicate quota warnings while retaining separate sync failures when they describe a different problem. - Require target-scoped caller grants for peer instruction access, while preserving self edits, responsible-user checks, and protected-change consent. - Treat full storage as a nonblocking run warning, never an agent pause or run-admission failure. Restore already-over-quota saved folders so ordinary agent cleanup can recover; warn on each run until cleanup. The run detail view shows the warning. - Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash large files as streams. Check editor-save quotas with metadata instead of hashing unrelated files. - Preserve old native inputs, instruction-only copies, paths, digests, and pending legacy candidates. Adopt old revision heads once. New writes do not append history rows. - Add idempotent migration 0287 and verify upgrades from the preview tables and receipts. - Add nine interactive stories under **Agents / Persistent files**, including automatic incoming edits, stale browser drafts, and storage-limit diagnostics. ## Verification - Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after merging current master and the landed transport prerequisite. Integration required no manual conflict resolution; the feature remains 99 changed files. Full workspace typecheck, production build, token gates, and 715 focused tests passed on this merge candidate. Fresh Greptile review is 5/5 with no unresolved findings. All 55 checks passed, with four conditional skips, including the build, typecheck, browser E2E, and canary dry run. A single retry recovered four jobs interrupted by runner shutdowns; no source changes were required. - Historical-warning UI fix: all 6,834 UI tests across 640 files passed, including regression coverage for three old failures, legacy preserved edits, and warnings scoped to the affected run. Full workspace typecheck, production build, Storybook build, and token gates passed. Browser-verified Storybook playtests passed for Historical Failures After Successful Save, Storage Limit, and Full Storage Run Warning. - Review follow-ups at `4e20c9fb2`: all 18 focused tests passed, including external Git directories, linked worktrees, symlinks, hardlinks, and distinct I/O failures alongside storage warnings. Server and UI typechecks, token gates, and the production build passed. - Storage warning regressions at `0724f3012`: all 33 directory tests and all five heartbeat-list tests passed, with no skips in their successful runs. They cover repeated runs while full, an already-over-quota saved folder, cleanup, warnings retained after unrelated save failures, and bounded warnings in large result JSON. Server typecheck passed after the final warning fixes. - Full workspace typecheck, production build, and token gates passed during this follow-up. Product E2E harness: 631 tests passed across 52 files; harness typecheck passed. Earlier native session/context and directory/legacy collection suites passed 537 tests; Runner unit/transport suites passed 329 tests. - **Real E2E at `0724f3012` (before this follow-up):** legacy local Codex and native Daytona Codex each passed six tasks, one server restart, seven independent assertions, and cleanup verification. Both prove browser-to-agent edits, agent-to-browser edits, nested/binary restoration, per-file last-sync-wins, a successful run after an oversized save rejection, and cleanup clearing the warning. - Native local Codex also passed the six-task quota flow before the final warning-retention fixes. That pass began at `918d1ed02` while the bounded-result warning fix was being edited, so it is not claimed as exact-final-head evidence. Its final-head rerun failed during embedded PostgreSQL bootstrap before any provider run: the macOS host had 87,365 of 87,381 SysV semaphores occupied. No unrelated services or kernel limits were changed. - The final-source report intentionally records **2/3 cells passed**, preserving the blocked native-local attempt: `tests/runner-e2e/results/agent-files-quota-final-20260928-report/`. Earlier failed attempts and provenance notes remain under `tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and the original campaign directories. - Daytona used immutable image `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf` and its exact Linux runner binary. Controller source is `0724f3012`; image source is recorded separately. - Legacy-session compatibility and all three ACP Stop/resume browser regressions passed on the prior validated feature head `169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same provider session is retained and interrupted writes are not replayed. Migration upgrade tests also passed earlier. - Nine interactive stories are under **Agents / Persistent files**, including **Full Storage Run Warning**. Its playtest and visual browser inspection passed; the warning states that runs continue and the editor remains available. - Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs skipped, no failures or pending checks. All eight browser E2E shards and their aggregate passed. Fresh Greptile review is 5/5 with no findings; all review threads are resolved, the security scan passed, and GitHub reports no merge conflicts. - The broad local follow-up test run was interrupted after host semaphore exhaustion affected isolated PostgreSQL instances. It also encountered the existing macOS long-path fixture failure and two timeout failures. This is not a claim that the full local suite passed. Logs are retained; focused storage/warning tests passed. ## Risks - A later sync can overwrite an earlier edit to the same file, including a saved browser edit. There is no text merge or retained version. This is the intended last-sync-wins policy. - A save that exceeds a storage limit is rejected and its temporary copy is discarded. The run itself continues normally, and later runs restore the last saved files with a warning until cleanup. Transient sync failures get bounded retries. An I/O failure partway through a sync can leave some files updated; a failed receipt does not claim whole-folder success. - Larger folders increase copy time, network traffic, and temporary disk usage. Active runs still need working copies. Terminal runs do not accumulate archives. Operators must provision disk for agents and configured concurrency; these limits are not company-wide quotas. - A restored old native session remains instruction-only until a fresh session starts. Its original conflict fence and existing pending candidates remain compatible. - Provider processes close at the collection boundary. Conversation resume remains available, but warm process reuse is lost. - Backups must include the instance filesystem and database. External bundles keep their existing behavior until explicitly moved to managed storage. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and browser tools assisted this change. Real provider E2E uses `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing> |
||
|
|
d172197117 |
feat(storage): add plain directory sync with conflict preflight (#14416)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs use workspace transport to restore and collect files. > - Some directories belong to the agent across tasks. > - Those directories need plain file transport without task Git state. > - A concurrent file edit must be detected before a merge changes any file. > - This pull request adds optional plain-directory sync and conflict preflight. > - Existing task workspace sync keeps its defaults. ## Linked Issues or Issue Description Refs #14325. This is the transport prerequisite for a replacement of its instruction revision design with current agent files. ## What Changed - Add an opt-in plain-directory transport mode to command and sandbox runtimes. - Add strict merge preflight for file edits, deletions, and directory changes. - Accept identical replay after an interrupted merge. Preserve competing changes. - Set the compiled OpenCode test executable to 0755, independent of the CI host’s file-creation mask. Preserve the original startup error if cleanup also fails. ## Verification - At `ced53ae532ce6966cad5a83575d83db61af98126`, all 212 targeted transport tests passed across workspace restore, remote managed runtime, SSH fixture, and execution-target sandbox suites. Adapter-utils typecheck passed. - Real isolated SSH retry fixture previously passed with `PAPERCLIP_ENABLE_DARWIN_SSH_ENV_LAB=1`; stale deleted files remain absent while gitignored binary bytes survive. - Dependent PR #14420 passed real native local, legacy local, and native Daytona persistence E2E at `169fab46d5af21caa2269b4c1b29b69c933a6951`, which includes all production transport changes through `ced53ae53`; the subsequent two commits only fix the OpenCode test fixture. Nine tasks, three server restarts, and all cleanup checks passed. - A hosted OpenCode fixture failed twice at `ced53ae53`. Reproduced the failure locally and in Linux with `umask 0002`: the compiler created a group-writable executable, correctly rejected by the qualified launch boundary. Explicit 0755 permissions fix the test without weakening the production guard. The focused test and non-root Linux reproduction now pass under that same mask. - Before rebase, head `69e97de0475d34aac5d532e559a405eaf015fd2b` includes the deterministic fixture permission fix and preserves original bootstrap diagnostics. All production transport code is unchanged since the 212-test validation. Fresh Greptile review is 5/5 on this exact head with no unresolved findings. All 54 current-head checks passed, with two conditional skips. The full CI run completed successfully, including the previously failing OpenCode runner shard. - Merge validation on rebased head `c509d79dd190c5cb00dc65edfde209097ff21465`: all five commits are patch-identical to the reviewed branch. All 54 checks passed with two conditional skips, and fresh Greptile review is 5/5 with no findings. One retry cleared an npm archive 404 and a Cursor fixture timeout. ## Risks - New behavior is opt-in. Existing task snapshot behavior retains its defaults. - Generic strict merge preflight remains opt-in. The dependent agent-folder feature rebases changed paths before applying them to provide per-file last-sync-wins; it does not create a conflict-review queue. - This change adds no database migration, dependency, or UI. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and tool use assisted this change. Live provider validation used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0be2afcca6 |
feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path > - Paperclip lets operators assign tasks to AI agents and review their work. > - The task composer controls the next message and its assigned agent. > - Operators needed a way to choose that agent's model and effort without leaving the composer. > - The old mode selector, upload button, and input cards made the mobile composer crowded and hid normal messaging during a pending decision. > - Harnesses publish different model and effort capabilities, so the picker must follow the selected agent. > - This pull request adds one responsive composer flow, keeps pending cards visible above it, and protects Codex ACP authentication in the local test path. > - Operators can choose run settings, send a message, and answer a pending card as separate actions. ## Linked Issues or Issue Description **Subsystem affected** Task composer UI, issue thread interactions, Codex ACP credential handling, and Storybook. **Problem or motivation** The composer did not expose model or effort for the selected agent. Mobile actions wrapped poorly. Pending questions and confirmations replaced the composer. A local Codex ACP test could also reuse host authentication after the managed key was removed. **Proposed solution** Put assignee search, model search, exact model IDs, effort, and fast mode in one picker. Use a mobile dialog. Replace the direct-upload plus action and separate mode selector with an Add menu and removable Plan or Ask chips. Place pending interaction cards above the usable composer. Keep these cards pending after an ordinary message unless their creator asks for comment superseding. Replace managed ACP auth files atomically and isolate the test key from host credentials. **Roadmap alignment** ROADMAP.md does not list an overlapping composer milestone. This change improves the existing task and review flows. ## What Changed - Added the combined assignee, model, and effort picker to both task composers. Search matches agent name, role, and harness. The server uses a curated Codex list by default and honors instance-declared models. Manual IDs remain available. - Added an effort slider for known model capabilities, a conditional Codex fast control, and reset. The picker opens in a modal on mobile. - Added the Add menu for files, supported goals, Plan mode, and Ask mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling remains available. - Adjusted mobile spacing, avatars, wrapping, and Send placement. Removed the composer divider. - Moved pending question, confirmation, review, and related cards above the composer. Ordinary comments now leave question and confirmation cards pending by default. The onboarding prompt retains explicit comment superseding. - Updated the Storybook composer group with responsive states and the production picker. Added UI, service, route, and browser regression coverage. - Isolated Codex ACP API-key authentication, skipped subscription auth merge and shared-home copy-back for remote API-key runs, and replaced the managed auth file atomically. ## Verification - `pnpm -r typecheck` — passed on the final local head. - `pnpm check:token-gates` — passed on the final local head. - `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31 tests passed, including role and harness search, declared Codex models, and filtering general OpenAI models. - `pnpm exec vitest run server/src/__tests__/issue-thread-interactions-service.test.ts` — 74 tests passed. - `pnpm exec vitest run packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed, including remote API-key copy-back isolation. - `pnpm test:run` — attempted locally; the embedded PostgreSQL test database could not initialize on macOS. The isolated `heartbeat-run-event-sequencing` suite reproduced that environment failure. GitHub CI runs the full test matrix for this head. - `pnpm build` — passed on the final head. `pnpm build-storybook` passed after the last UI change; only server code, tests, and docs changed afterward. - Live local test drive — Codex ACP ran a task with a managed API key. The test agent was restored to its default ACP configuration afterward. - Review the interactive stories under the top-level Composer group with `pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask, picker, and pending-question states. ## Risks - A pending card stays open when an ordinary comment changes the discussion. Its creator can set `supersedeOnUserComment: true` when a new comment should replace it. - Model and effort overrides persist on the task until reset or changed. An unlisted manual model ID may fail when the provider runs it. - Some harness catalogs do not report effort support. The picker hides effort for those models. - No database migration is required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID or context window to the task. The model used code execution and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI Codex <codex@openai.com> |
||
|
|
24c58e479a |
Improve task artifacts with rich cards and editable stories (#14469)
Render eight artifact card types from real task records and share them with editable Storybook stories. Preserve document review and media/file actions, and load bounded CSV previews on request. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
18e8c121d9 |
fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path > - Paperclip manages agents through a shared native runner. > - Built-in harness support should ship with Paperclip's public distribution. > - Grok already speaks ACP; it does not require a new public bridge package. > - Sandbox provisioning owns the native executable and its pinned version. > - The runner must verify that prerequisite without downloading it during npm installation. > - This change separates built-in launcher identity from external runtime identity. > - Clean npm installation and live staging checks verify the distribution boundary. ## Linked Issues or Issue Description Refs #13882, #13973, #13977, #13979. This follow-up now targets master after #13882 was squash-merged. It replaces the private `@paperclipai/grok-acp` workspace package with runner-owned assets. Current master is included so the branch also contains the merged scheduler, complete-event capture, and durable cleanup fixes. ## What Changed - Ship Grok launcher and qualification metadata inside the runner's compiled output and the public server's vendored runner tree. - Remove the separate Grok npm package and all package-manager install hooks for this runtime. - Require the checksum-verified Grok Build 1.0.13 binary at `/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution environment. Provision it explicitly in the Daytona image and CI setup. - Keep native binaries outside the provider pack. Bind the built-in launcher into the pack manifest. - Preserve executable leases, descriptor-backed startup, credential fences, permissions, and exact ACP model admission. - Use `builtin:grok-acp` and `native:grok` as profile identities. Historical package-profile sessions fail closed on resume rather than being silently reinterpreted. - Resolve built-in assets from the authenticated sidecar location, including public server npm layouts. Keep the controller path out of provider environments. - Add clean npm tarball installation verification to the existing trusted canary CI job and the admitted manual EC2 verification path. It stages a unified release version and runs npm lifecycle scripts, then verifies missing-prerequisite rejection and admission after separate provisioning without credentials or inference. - Include the controller-owned provider pack in stamped Cloud images. Unstamped local images omit the pack and remain usable; remote ACPX requires full source provenance. - Correct CLI approval-page metadata for an already authenticated Cloud board user; approval authorization remains unchanged. - Honor explicit native-runner enablement in the Cloud agent picker and direct setup page, keeping the flag disabled by default. - Allow selecting the execution environment before connecting credentials. Include Grok in the existing authenticated hello-probe flow, targeting its pinned native prerequisite for runner setup. - Recover an existing subscription sign-in conflict through an explicit cancel-and-retry action, serialized after cancellation succeeds. - Preserve the selected ACPX harness before normalizing config fields, so new Grok agents use the Grok default model. - Keep the credential-free Cloud provider pack root-owned and readable after runtime UID remapping; verify manifest and referenced asset access under an unrelated unprivileged UID during image builds. - Archive prior failover backups alongside explicitly replaced harness state, preserving evidence while preventing stale backups from blocking a fresh replacement. - Update Daytona image content inputs and contract tests for the built-in assets and explicit provisioner. - Document and regression-test the shared `approve-all` default for Grok setup, saved configuration, and native execution. Explicitly saved restrictions remain unchanged. ## Verification Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05` incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the base PR was squash-merged. All 12 conflicts came from incoming files identical to the tested pre-squash base. The final tree exactly matches a three-way merge using that original base, preserving built-in Grok distribution and removal of the obsolete private package. All 252 focused runner/UI tests, six npm-isolation tests, and token gates pass. Fresh exact-head Greptile review is 5/5 with no outstanding findings; security scans and EC2 native compilation pass. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([run 36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)). The repository owner explicitly authorized bypassing code-owner approval after all checks passed; no CI checks or repository protection settings are bypassed or changed. The only remaining PR was removed from the completed stack metadata to permit native auto-merge. Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)). The initial attempt lost two EC2 runners to shutdown signals and stalled a third shard during dependency preparation; all three passed the same-commit failed-job-only retry. Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. [Final public npm verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764) passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public packages, an executed offline lifecycle sentinel, unchanged consumer lock, built-in launcher, missing-prerequisite rejection, and verified separately provisioned binary/command lease. Provisioning and cleanup require no host privilege elevation; only the positive probe mounts the temporary native binary read-only. The verifier is unchanged by the final master merge. All six isolation tests and an offline npm smoke test pass. The prior head had 56 green CI checks and a 5/5 review after two unchanged tests timed out and passed a failed-job-only retry ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)). All 56 recovery-display/lineage tests pass; re-review cleared the already-covered missed-retry concern. Earlier EC2 failures remain retained: [npm lockfile rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203), [missing compiler in the slim image](https://github.com/paperclipai/paperclip/actions/runs/36440210984), and the aggregate 15-minute test timeouts in those broad runs. Both broad attempts passed typecheck, token gates, Product E2E type/unit checks and build. The focused EC2 lane preserves the existing trusted-actor and immutable-source gates. Earlier documentation/test checkpoint `ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior unchanged. 154 focused tests pass across configuration building, native provider resolution, permission policy, credentials, UI configuration, and new-agent setup (including both Grok auth modes); token gates pass. All fresh CI is green for this head: 56 successful checks/statuses and two intentional skips ([run 36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)). Greptile is 5/5 with no new findings. Grok already inherits the shared `approve-all` default, so unattended setup requires no manual permission change. Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final staging continuation failure before provider startup: explicit replacement archived the old harness but left its failover backups active, which caused `runner_harness_state_mismatch`. The regression fails before the fix and passes after it; all eight adjacent recovery-safety cases also pass. Old backups remain inspectable inside the continuity archive. All fresh CI is green at this head ([run 36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)), with a 5/5 review. One unrelated Cursor test timed out in the initial server shard; the same-commit failed-job rerun passed, and both attempts are retained. Staging deployment is confirmed healthy on this revision. The controller image is `ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`. The final browser-created staging task passed on this exact revision with API authentication: context read → structured human question → controller restart → answer submission → same native provider session resumed → document saved → task Done. The two turns took approximately 119s and 77s. The actual write receipt was applied, and the saved document has exactly one revision containing the selected answer and requested marker. Usage and cost were not reported. [Controller image build](https://github.com/paperclipai/paperclip/actions/runs/36360889243). - Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`: all CI green (53 successful checks/statuses, two intentional skips), including repository typecheck/build/tests, native Runner tests, browser shards, and canary installation checks. [CI run 36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529). Greptile is 5/5 with no unresolved findings. - Focused checks cover Grok credentials, executable admission, launcher assets, provider-pack paths/permissions, workflow contracts, setup defaults, CLI authorization, and subscription conflict recovery. All 39 protocol definitions validate. Final integration checks pass 124 catalog/evidence/cache tests and nine project-form tests; token gates pass. Some local dependency checks could not load the stale installed dependency tree; the corresponding fresh EC2 checks pass. - Clean public npm installation passed on EC2 at `8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run 36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)): 17 unified-version packages, lifecycle scripts enabled, built-in launcher present, no separate Grok package or npm-downloaded binary, missing prerequisite rejected, separately provisioned native executable and command lease verified. No credentials or inference were used. Subsequent changes preserve this npm asset layout. - The immutable Daytona prerequisite image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`, built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous Cloud controller image was `ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`, built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded by the latest image above. Its EC2 build verified provider-pack access under an unrelated unprivileged UID. - Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed full Grok onboarding with the correct `grok-4.7` model, saved credential delivery, and pinned Daytona execution. A browser-created task read context and asked the structured human question. After a controller restart, answering the persisted question resumed the same native provider session, saved the requested document, and completed the task. Actual tool outcomes and durable state agree: one question and one document revision. The two successful turns took 42.7s and 63.1s; usage and cost were not reported. - Restricted policy returned the expected `approval_required` outcome. Functional staging tests explicitly selected `approve-all`; controller authorization and governed approvals remain enforced. Temporary board CLI access was revoked and verified rejected (HTTP 401), and the disposable onboarding agent was paused. Failures remain retained: the pre-fix continuation failure (its task remains blocked; the passing final task is fresh), the original Cloud provider-pack permission failure, the expected restricted-policy denial, the superseded npm staging failure, and an earlier monolithic CI infrastructure timeout. Browser CI exposed a project alias/form race; the final stack uses master's stronger draft-preservation fix and all browser shards pass. Historical full subscription/API protocol and Product rosters retain their original source revisions and do not qualify this packaging revision. No local Docker or Rust build was used. ## Risks The branch includes master’s draft-preservation fix for project URL aliases. It keeps the same project’s edit form mounted and clears prior data when the project or company changes. Custom sandboxes and local execution hosts must provision the pinned binary before Grok starts. Missing, changed, unsupported-platform, and symlinked executables fail admission. The new builtin profile cannot resume sessions created with the former private-package profile. Existing Claude/Codex npm bridge profiles retain their package pins. Grok restricted modes preserve the selected policy but cannot automatically admit Paperclip calls: ACP permission metadata does not independently bind tool authority, so those calls stop with `approval_required`. New Grok configurations default to `approve-all`, including API configurations that omit the mode. Existing explicitly restricted configurations remain restricted; controller authorization and governed approvals remain enforced. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
992f720262 |
fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task descriptions, comments, continuation data, skills, and execution rules enter several agent adapters. > - The same source can be rendered by more than one automatic input carrier. > - Failed resumes can also rebuild input from stale or compact context. > - This pull request gives each Paperclip-owned source one delivery owner and preserves the required transport boundaries. > - It adds deterministic adapter, interaction, runner, and browser tests for these boundaries. > - The benefit is more predictable context delivery with explicit evidence for later live qualification. ## Linked Issues or Issue Description Related: #13144 removes a duplicate environment payload and bounds wake lists. Related: #11360 addresses Hermes resume behavior. This pull request preserves compatible active-session formats while repairing context ownership and stale question creation. **What happened?** Task descriptions and comments could enter more than one automatic context block. Native transports could wrap a complete model input in a second task envelope. Some legacy and gateway adapters could omit the owned assignment on ordinary tasks or rebuild a failed resume with stale compact context. A continuation could also request a question after newer human comments had arrived. **Expected behavior** Each task or comment source has one automatic model-facing owner. Distinct comment IDs and repeated wording remain distinct. Fresh fallback attempts rebuild the required full context. A question request is rejected when newer queued human direction makes it stale. Harness access policy remains owned by execution configuration. **Steps to reproduce** 1. Build a task with a description and current comments. 2. Capture the actual adapter or runner input. 3. Compare source ownership and task-envelope nesting. 4. Queue a human comment before a continuation requests a question. 5. Trigger a failed resume and inspect the fresh retry input. 6. Run the focused adapter, interaction, runner, and browser checks. ## What Changed - Add shared prompt-section selection at the provider-attempt boundary. - Deliver owned assignment context through native, legacy CLI, ACP, gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and Hermes paths. - Rebuild full or compact context after resume recovery changes the attempt. Add native and Claude ACP tests of actual recovery requests. - Preserve custom templates, loaded instruction files, execution policies, and older active-session formats. - Record continuation source metadata and reject stale question creation under the issue-row lock. - Add explicit Product E2E context-integrity profiles, prerequisite gates, credential-isolation checks, and report fixtures. - Bypass service-worker forwarding for same-origin Vite development modules. A real Chromium test fails with resource exhaustion before the repair and passes after it. Production asset caching keeps its existing policy. - Add browser diagnostics and service-worker module-loading regressions. - Add an explicit zero-retry eval option. The default retry behavior remains unchanged. Each campaign records its effective policy. - Remove the model-facing working-directory sentence from four prompt builders. Existing workspace, sandbox, permission, and custom-template configuration remains unchanged. - Align the everyday workflow assertion with the current 47-entry catalog. Compared with current upstream master, the branch carries the context-ownership implementation and its tests, the explicit context-integrity catalog and evidence harness, and the focused browser regression checks. ## Verification **Merge assessment:** focused regression evidence supports merge. This is not full completion of the original broad qualification matrix. The maintainer has authorized merge after fresh verification of the master integration. - Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`. All 14 conflicts are resolved. Cancellation checks, workspace finalization, native Grok support, and both sets of tests are retained. - Current-head Greptile: **5/5**, with no blocking findings. The review names this exact commit. All **59 reported checks are terminal: 55 successful, 4 skipped, zero pending or failing**. This includes the full root general and serialized suites, separate runner checks, typecheck, build, canary, browser E2E, Docker, and security checks. The successful legacy security status is included in that total. - After integration: workspace typecheck and full build passed. Separate runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust tests, and 39 preparation checks**. Other passing checks include 621 Product E2E harness units, 376 focused shared/adapter tests, 160 real-database/API tests, 86 Hermes tests, 18 browser-support checks, and Product E2E typechecking. The complete root suite passed in CI. The duplicate local monolithic root run was stopped after that CI result; it is not counted as a completed local pass. - New native recovery coverage retains full assignment, completion contract, and explicit skill selection after safe replacement, for old and prepared input formats. Full native session test file: **136/136 passed**. - New Claude ACP coverage captures actual fresh, resumed, and missing-session fallback requests. It verifies one assignment copy, comment order, identical text under distinct comment IDs, and full fallback context. Full file: **33/33 passed**. Both affected TypeScript checks passed. - Existing deterministic tests cover source revisions, approval and trust boundaries, completion validation, custom templates, compatible sessions, standalone driver wrapping, and maintained adapter transport requests. - Provider-free browser support: **17/17 passed** after the master merge. Service-worker unit tests: **33/33 passed**. The module-overload regression failed before the repair and passed after it in real Chromium. ### Fresh live comparisons The new batch ran exactly four Product E2E attempts. **All four passed on the first attempt; no retries.** Each has six terminal matchers plus the existing browser lifecycle and invariant checks. | Exact case ID | Control | Candidate | |---|---|---| | `core-compatibility.runner-codex.local.plan-revise-accept` | Passed | Passed | | `local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume` | Passed | Passed | The plan case checks a revised canonical plan and revision-bound approval before completion. The question case restarts the server before submitting the answer, then verifies the continuation completes. Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical frozen definitions and provider versions: Codex `0.156.0` with `gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with `claude-sonnet-5`. The September 24 head added master browser recovery and test-only changes. The September 28 head also integrates newer master changes, including cancellation, workspace finalization, and native Grok. These are frozen-source live results, not exact-head live runs. The candidate received one description copy where the control initially received three. The submitted initial plan envelopes were 7,969 versus 19,097 characters. Question envelopes were 7,592 versus 18,919. These are structural measurements, not whole-provider token or dollar savings. ### Earlier evidence and failed attempts - The preceding fresh batch has four effective passing pairs: OpenCode comment continuation and assigned skill, native Codex comment continuation, and native Claude comment continuation. It retains **11 attempts: eight passed and three failed**. - Original failures remain recorded: missing local PostgreSQL library links before task creation; host-sleep cleanup after task/page checks passed; and a Claude **control** session-open rejection before a model turn. Setup was repaired identically on both worktrees. The permitted unchanged infrastructure retries passed. The underlying Claude provider startup error was not retained and remains unknown. - Older R2 retains **17 passes and one failure** across 18 attempts, including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode blank-page failure led to the service-worker repair. R2 is historical evidence: master changed the native fixed prompt and removed duplicate wake environment data afterward. - The September 24 CI run initially failed one unrelated preview readiness test (`ECONNREFUSED` on its local fixture). Its test and production code match master. Isolated local verification passed **28 tests, 3 skipped**. One unchanged CI retry passed the full shard: **831 passed, 1 skipped**, including all **31 preview-exposure tests**. The aggregate CI gate passed afterward. The precise startup cause remains unknown; a port race is a hypothesis, not a proved cause. ### Limits The original wider profile/workflow matrix, repeated trials, and remote Daytona qualification are incomplete. These results support a focused merge recommendation, not statistical equivalence or universal harness qualification. Some usage receipts are missing in both variants, so no token or dollar savings are claimed. The $500 ceiling was preserved using conservative allowances; failed attempts and unknown charges remain in the ledger. Reproduce the focused additions with `pnpm exec vitest run packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/native-session-runtime.test.ts`. Full checks use `pnpm -r typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner checks. Paid evals require the frozen definitions, profiles, and credentials; do not use `--all` as a substitute for the selected cases. ## Risks - Context placement changes can affect model behavior. Deterministic checks cover the selected paths, but live qualification remains incomplete. - The stale-question guard can reject a request when queued human comments arrived during the run. This is intended. - New stored inputs and model envelopes retain compatibility readers for older active sessions. - Custom templates may intentionally repeat content. - Removing a model-facing working-directory sentence does not change filesystem, command, sandbox, or permission configuration. - The worker bypass applies only to same-origin development module paths. Cache-policy tests preserve private-response handling and production asset caching. Mounted HTTP fixture changes remain test-only. - This PR does not claim measured token savings or statistical equivalence across every harness. ## Model Used OpenAI Codex, exact model gpt-6-astra, with repository tools and code execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The serving context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR using the required issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused local checks and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect these changes - [x] I have considered and documented risks above - [x] All current-head Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups for the current head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2f585ef26a |
fix(ui): preserve newer drafts after repeated receipt cleanup (#14332)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task composer keeps unsent text across page reloads. > - A server receipt confirms a submitted message after a reload. > - Replayed cleanup can apply the old text offset to a newer draft twice. > - This pull request reconciles each confirmed attempt once and preserves both tabs’ unsent intent. > - The full newer draft stays available for the next send. ## Linked Issues or Issue Description **What happened?** Reloading the classic composer while a save was pending could truncate a newer draft. Replayed receipt cleanup changed `A newer draft written while delivery was pending.` into `nding.`. The storage helper rejected the duplicate settlement, but the effect still changed the editor and its body reference. **Expected behavior** A receipt settles its retained submission once. Later cleanup must preserve the newer draft in both the editor and browser storage. **Steps to reproduce** 1. Send a comment and hold its HTTP response after the server accepts it. 2. Type a newer draft, then reload the page. 3. Restore the matching receipt under React StrictMode. 4. Inspect the newer draft after effect replay and unmount. The deterministic regression reproduces this on current master. The existing browser test exposed the issue during #14329 verification. Related draft persistence code came from #13338. A search found no separate fix for this duplicate settlement. ## What Changed - Reconcile each confirmed attempt once in memory so duplicate effects cannot trim its newer draft again. - Unlock confirmed drafts when storage writes fail. Never acquire another tab’s pending receipt or assume its text is newer. Keep a conflicting local draft and its attachments in tab-scoped session storage while preserving the shared draft unchanged. - Add component regressions for StrictMode replay, exact editor and stored bytes, foreign receipts, attachment-only differences, reload recovery, unavailable storage, and task navigation before autosave. A brief notice explains when this tab has a separate draft. ## Verification - RED: the new test received `nding.` instead of the full newer draft before the fix. - Review RED: three cases reproduced a locked composer after failed storage writes or another tab's settlement. A further negative case preserves a newer retained attempt and its attachments. - Cross-tab review RED: four deterministic cases reproduced lost local text, foreign receipt takeover, attachment loss when text matched, and overwriting newer stored text. - GREEN: 118 tests across `IssueChatThread`, `composer-draft`, and `comment-submit-draft`, including same-mounted A → B → A navigation and a full recovery-storage failure. Two further lifecycle RED tests verify finishing a recovery returns to the shared draft on a later visit while continued typing and attachments retain recovery. - Both existing browser reload cases passed against a fresh server and database on exact head `ceb80aca77fc8cc0f813c328ba87025b3e1a2222` (42.5 seconds), with classic mode enabled and disabled. The browser flow verifies one original comment, a preserved draft, and a successful second send. - UI typecheck, token gates, and the shipped static UI build passed. The initial typecheck required the fresh worktree's plugin SDK build; the retry passed after that dependency built. - The session-only failure mock also preserves localStorage on platforms where both share the Storage prototype; this fixes the Linux workspace test failure. - All 56 checks passed on `4c30dcccc4fe5b90dcd08dd1d90475ea89e4a376`; Greptile is 5/5 with zero open review threads. An unrelated runner baseline scan hit its existing 100 ms deadline once; its isolated test and the single failed CI job passed on retry without source changes. This four-file UI fix does not change server or database code. - Final alternate-staging acceptance passed on deployed source `d884e1ab046cc76004e35e6091e9e6e2c918c9eb`, with exact health checked before and after. Three real browser cases used shipped static UI and real API saves: reload during an accepted-but-unacknowledged save in both composer variants, plus two classic tabs with different drafts. The test deliberately held only its own accepted POST acknowledgement and temporarily withheld its own GET receipt from one tab to reproduce settlement ordering. Both drafts survived independent reloads, the shared draft remained intact after the other tab sent, and finishing recovery returned to the shared draft on a later visit. Exact request IDs/counts and server comment bytes passed; no provider was invoked. Screenshots preserve the live drafts before fixture cleanup. - An ordinary authenticated Chrome/CUA visit independently showed the exact original and newer comments; an additional unsent draft survived reload and was then cleared. All three isolated fixtures are complete, browser contexts are closed, and original user tasks were untouched. ## Risks - Reconciliation uses the exact request ID and draft key. Conflicting drafts stay separate; the tab recovery survives same-tab reload through session storage and does not outlive the browser tab. If recovery storage is unavailable, the editor keeps its in-memory text and explicitly warns the user to copy it before leaving. - Attachment selections remain bounded to 20 per draft; if combining equal-text snapshots would exceed that limit, each original selection stays in its respective draft. - No API, schema, or migration changes. The status notice uses existing design tokens. ## Model Used OpenAI Codex, GPT-6, with tool use and code execution. The runtime does not expose the exact deployment variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1a394bd30 |
feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d9d2147171 |
fix(auth): keep Cloud tenants on the Cloud sign-in flow (#14407)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud owns human identity and passes a verified identity to each tenant. > - The tenant can report no session while the Cloud session is still valid. > - The access gate and direct `/auth` route then show the instance password form. > - This pull request sends those users through the configured Cloud entry endpoint. > - Cloud can renew the tenant session or show its login page, then return to the original task. ## Linked Issues or Issue Description **What happened?** A Cloud tenant can display the self-hosted email/password form after an instance session check returns no session. This gives Cloud users the wrong login method. **Expected behavior** An active Cloud session renews tenant access automatically. A signed-out user signs in through Cloud. Staging and production use their own configured Cloud origins. Self-hosted instances keep their instance login form. **Steps to reproduce** 1. Open a Cloud tenant task or an `/auth?next=...` link. 2. Keep the Cloud session active but make the instance session check return 401. 3. Observe the instance password form instead of Cloud session recovery. **Deployment mode** Cloud-managed authenticated instances. No database or server API changes. Searched related authentication PRs. Native self-hosted OIDC support in #10411 is a separate feature; this change uses the existing Cloud entry contract. ## What Changed - Wait for deployment metadata before showing an instance login form. - Use the health response's Cloud origin and stack slug for session recovery. - Preserve the tenant path, query, and fragment. Reject external and recursive login return targets. - Limit automatic recovery per tab. Show a manual Cloud retry after failed recovery. Show service failures as errors. - Keep self-hosted login and local trusted access. Add focused tests, browser regressions, deployment documentation, and an unavailable-state design example. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - UI typecheck and `pnpm check:token-gates` passed after the final UI edits. - All UI tests passed: 639 files, 6,777 tests. - 63 focused Vitest tests passed across Auth, CloudAccessGate, Cloud links, and recovery coordination. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/cloud-auth.spec.ts`: 5 passed. The tests use the real tenant UI and database with a simulated Cloud HTTP endpoint. They cover both Cloud origins, direct auth/task links, no password-form flash, preserved URLs, reload, and self-hosted login. - Hands-on browser test used the real Cloud gateway and a fresh tenant build with disposable local data. Active Cloud session plus a forced missing instance session returned to the task. An expired tenant cookie also renewed automatically and returned to the task. Removing both sessions reached the real Cloud email/social login UI. Persistent failure stopped at the retry screen; retry succeeded after removing the injected fault. The fixture used a loopback transport adapter and a simulated signed-out OIDC issuer. No production session or deployment was changed. - The default local browser startup hit the host's embedded PostgreSQL resource limit. The passing run used a separate disposable database on the test PostgreSQL process. - The full local `pnpm test:run` sweep was stopped after about 31 minutes once CI completed the full suite. It had reported 33 failures in the unchanged runner API unit/integration files; both files pass in isolation (1,749 + 28 tests). The local sweep did not reach the later workspace/serialized groups. CI completed all of those groups successfully. - Greptile reviewed commit `d603fd4e39455de44da9dae81b72197096c0e1e8` at 5/5 with its only thread resolved. All CI gates are green on this commit, including all general/serialized server groups, workspace tests, Runner checks, typecheck, build, and all eight browser shards ([run](https://github.com/paperclipai/paperclip/actions/runs/36446230697)). ## Risks - Recovery depends on valid Cloud origin and stack metadata. Incomplete metadata shows an unavailable message instead of a password form. - Browsers with session storage disabled use the manual Cloud link, since automatic retries cannot be bounded across documents. - The external identity provider's email/social login was not completed in this local test. Existing Cloud authentication owns that flow. - No migration, credential format, membership rule, or production deployment changes. ## Model Used OpenAI GPT-6 through Codex. The exact served variant and context-window limit are not exposed in this session. Used reasoning, repository tools, code execution, and browser testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
270afd2fb8 |
feat(ui): show running commit in staging account menu (#14410)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The account menu shows the signed-in user's identity. > - Staging users need to know which server commit is running after a deploy. > - The health endpoint already returns that commit, but the menu does not show it. > - This pull request adds the short commit below the email on staging hosts. > - Users can open the menu to check a deploy without opening deployment tools. ## Linked Issues or Issue Description **What existing behavior does this improve?** The account menu on staging instances. **Current behavior** The menu shows the user's name and email. It does not show the running server commit. **Proposed behavior** On `*.staging.paperclip.app`, show `SHA 8751e2d` below the email. Use the current `/api/health` commit. Show the full SHA on hover and link to the commit on GitHub. Refresh the health query when the menu opens. Hide the label on other hosts and when commit metadata is unavailable. **Reason and benefit** A user can confirm which commit a staging instance runs after an automatic deploy. **Breaking changes** None. The server already returns the commit field. Refs #14060 for related account-menu work. This change adds deployment information only. ## What Changed - Add the existing health response commit field to the UI type. - Share the staging host check and a separate health query across both account-menu variants. A menu refresh failure leaves the access gate health state unchanged. - Show a short SHA below the email. Link to the full commit on GitHub and include the full SHA in its accessible name and hover title. - Document the staging label and cover staging hosts, other hosts, missing metadata, refresh on reopen, and request failure isolation. ## Verification - `pnpm exec vitest run ui/src/components/SidebarAccountMenu.test.tsx` — 24 tests passed. - `pnpm check:token-gates` — passed. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. The UI build and typecheck also passed again after review fixes. - Full test matrix — passed on the latest commit in [CI](https://github.com/paperclipai/paperclip/actions/runs/36447810742). The local `pnpm test:run` was stopped before completion after CI finished the same suites. The focused local tests, full local typecheck, and full local build passed. - Browser shard 6 passed on one rerun. Its first attempt lost part of the draft text in the existing attachment-receipt reload test. No code changed for the rerun. - Rendered the real account menu in a local browser fixture with a staging hostname condition and mocked health response. Confirmed the SHA fits below the email in the dark menu. - Greptile review — 5/5, all review threads resolved. - Manual check after deployment: open the menu on a staging host and compare the SHA with `/api/health`. Open the menu again after a deploy to refresh it. Confirm the label is absent on production and localhost. ## Risks - Low risk. Each menu opening on staging can make one additional health request. - The host check applies to `*.staging.paperclip.app`. Other staging domains will need an explicit update. - The label identifies the running server commit. It can briefly show cached data while the request completes. ## Model Used OpenAI Codex, GPT-6. The exact model variant and context window are not exposed in this session. Used code execution, repository inspection, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #123` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
14795136f5 |
fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider. Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3447609d22 |
fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.
## Linked Issues or Issue Description
**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.
**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.
**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.
Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.
## What Changed
- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.
## Verification
- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.
## Risks
- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.
## Model Used
OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
cea8dda472 |
test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users can delegate work through onboarding and Agent Chat. > - A completed task does not prove that its result reached the original conversation. > - Existing tests do not isolate completion after the source chat becomes idle. > - This pull request adds four explicit native-runner probes across Claude and Codex. > - The probes preserve the result and reply so we can separate delivery failures from inaccurate answers. ## Linked Issues or Issue Description Refs #13775. Refs #13813. These evals extend native-runner qualification. They measure completion updates before we choose a product change. ## What Changed - Add the opt-in `completion-updates` suite with two stories for each native provider. - Test completion in the existing onboarding task flow and after an Agent Chat handoff becomes idle. - Gate the chat worker on a brief inside its managed project workspace. Prove the source is idle before releasing the worker. - Check durable task completion, saved output, a subsequent source reply, and rendered access to the result. - Preserve replies, task state, screenshots, run events, and a separate semantic review rubric. - Add grader regression tests and update the documented eval contract. - Preserve the suites added on master and include four completion cases in the 306-cell catalog. Production behavior and prompts are unchanged. ## Verification - Passed all 565 eval support tests across 45 files after merging current master: `node node_modules/vitest/vitest.mjs run --config tests/runner-e2e/vitest.config.ts`. - Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p tests/runner-e2e/tsconfig.json`. - Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs tests/runner-e2e/launch.ts --list --suite completion-updates`. - Four-cell behavior campaign on source `ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`: https://github.com/paperclipai/paperclip/actions/runs/36072337485. - A screenshot-only follow-up waits for the restored source reply to render after result-link navigation. Its one-cell Claude onboarding verification passed on final head: https://github.com/paperclipai/paperclip/actions/runs/36075716141. Corrected report: https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/. The original four-cell onboarding screenshots caught navigation loading; its saved reply evidence remains valid. The follow-up again found stale wording: "That work will run next" was posted 38 seconds after the child was Done. The four-cell campaign keeps its original source and measurements. - Published evidence: https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/. - Suite definition: `afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`, version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local execution, one attempt per cell. All four cleanup checks passed. Onboarding billing coverage is partial; reported zero cost must not be read as a free run. | Story | Automated delivery/access | Separate semantic review | | --- | --- | --- | | Codex onboarding | Pass | Pass: accurate completion reply with an accessible result | | Claude onboarding | Pass | Fail: reply says it will save the note once the task runs, after the note is already saved and the task is Done | | Codex idle chat handoff | Fail | Worker completed and saved the note; no completion reply during the full observation window | | Claude idle chat handoff | Fail | Worker completed and saved the note; no completion reply during the full observation window | Both chat cases positively recorded the source waiting and the worker at the brief gate before release. Both saved outputs include the brief-only start time. The opt-in campaign is red because it exposes current behavior. It is not a required merge gate. The PR does not fix that product behavior. Semantic review is a recorded human/agent assessment of retained evidence; it is not an automated prose-quality judge. - Second campaign: https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex chat reached the idle boundary and completed its task, then received no completion reply during the full window. Claude onboarding again returned a stale handoff answer. Claude chat exceeded the prior 110-second handoff setup budget; this revision raises that bounded setup window to 180 seconds. - Retained baseline: https://github.com/paperclipai/paperclip/actions/runs/36069427676. Onboarding passed delivery/access for both providers, but Claude gave a stale handoff answer. Chat cases stopped at fixture problems; they do not establish a completion-delivery failure. This revision fixes the workspace path and competing reference requirements. - On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR checks passed and two were skipped, including typecheck, tests, and build. Broad checks ran in CI, not locally. That head received Greptile 5/5 with no unresolved findings. The unchanged mobile repository-settings browser test passed on one targeted retry after a detached/disabled Save-button timeout. - Merged current master in `9b4491e1f` and resolved the catalog-count conflict. Eval support tests and eval TypeScript pass locally. All individual CI jobs passed on this merge commit, including build, typecheck, server tests, runner checks, and browser shards. The final aggregate check also passed: 54 checks passed and two were skipped. Greptile reviewed this exact commit at 5/5 with no unresolved findings. ## Risks - These explicit probes can expose current product failures. They do not change the default paid test selection. - Mechanical delivery and result access do not establish answer accuracy. The preserved reply still requires semantic review. - A fixture failure before the idle boundary or worker completion cannot establish a completion-update failure. - The handoff setup window lasts three minutes. The worker brief wait is bounded at four minutes. The observation window lasts two minutes after worker completion. It retains later replies without erasing earlier accessible delivery. ## Model Used OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository inspection, code execution, and GitHub tool use. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8751e2de46 |
fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task board shows which recovery actions are active. > - A native run can stop while a person must repair its workspace. > - The board previously called that state “Recovery in progress.” > - The label implied that work would continue without operator action. > - This pull request derives the label from the recovery owner and live continuation. > - Operators can distinguish scheduled recovery from a repair that needs attention. ## Linked Issues or Issue Description **What happened?** A blocked task showed “Recovery in progress” after native finalization stopped and no automatic continuation remained. **Expected behavior** Show “Recovery needed” for an idle board repair or when no live recovery path exists. Show “Recovery in progress” while the recorded continuation can run, including an explicitly admitted export whose exact callback is executing even if the old recovery action remains board-owned. **Steps to reproduce** 1. Complete a native run whose workspace export cannot be recovered automatically. 2. Inspect the task recovery action and badge. 3. Compare the board-owned action with the old “Recovery in progress” label. Related work: #14314 serializes native workspace finalization and fences stale recovery outcomes. This change reports the recovery action that currently owns the task. ## What Changed - Show “Recovery needed” for board-owned active-run recovery unless the exact native export callback is positively verified as executing. - Require the native continuation run and a live or future continuation before showing progress. - Project native activity from the exact company, source issue, and run. Include active workspace export while the original heartbeat remains failed. - Require an executing callback before using a running export row as evidence. Preserve activity for long exports and clear it when the callback joins. - Give native resume its own card explanation. Preserve ordinary watchdog observation behavior. - Remove the redundant ownership sentence from all six recovery-card explanations that used it. - Document the labels and add regression cases for stopped, scheduled, and active recovery. ## Verification - Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests and token gates pass. No UI occurrence of the removed sentence remains. `pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54 successful checks, two optional Storybook checks skipped). Greptile is 5/5 with no open review threads. The duplicate local `pnpm test:run` was stopped after the full CI suite passed; it did not complete locally. - Original regressions: seven failures before the change, then 52 focused cases pass. - Review regressions: seven UI failures and nine database failures before the follow-up. All 75 database/API recovery tests, 167 UI tests, and 18 workspace lifecycle/finalizer tests pass. A further three RED cases cover explicit board retry activity; one RED case rejects orphaned running export rows after controller loss. Wrong company, issue, run, service, phase, and completed-operation cases remain inactive. - Final recursive typecheck, production build, token gates, and complete local suite coverage pass. Embedded PostgreSQL startup/socket failures passed in isolated retries with the canonical test environment; no expected behavior was weakened. The recovery regression added during the earlier full run passed in its final complete 75-case file. - Before the copy-only follow-up, all 56 CI checks passed on `c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no unresolved review threads. The final native-activity staging repeat passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`. - Verified on a separate staging instance: a real failed Daytona workspace export retains its board-owned repair action and displays “Recovery needed” in the task list. The repair card remains actionable without starting another provider turn. - Real staged export-only repair: the actual task list showed “Recovery in progress” while the original run had a positively identified running export operation, then Done after exact copyback of all 20,000 nonce-bound files. The accepted result and full provider session/turn/terminal envelopes remained unchanged. All 17 independent final checks passed; the browser downloaded the exact 19-byte result. The separate fixture was cleaned up with independent provider-absence verification. A control transport process restarted during repair; it did not submit another provider turn. - Final deployed-source repeat on `d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral Daytona allocation, actual browser export repair, and a saved full activity projection referencing the exact executing export operation with no scheduled retry. The task list showed recovery in progress, then Done; all 21 final checks passed, including 20,000 exact host files, unchanged provider provenance, and provider deletion only after committed copyback. The downloaded 19-byte result matched independently. This integrates #14334; no source change was required here. ## Risks - The label depends on the persisted recovery action. A separate runtime defect can still stop work; this change makes that condition visible. - Future recovery kinds must supply a valid continuation path before they can display progress. ## Model Used OpenAI Codex, based on GPT-6, with code execution, browser testing, and subagent tool use. The runtime does not expose an exact serving model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cbc5132e6c |
fix(adapters): expose a verified provider stop before workspace restoration (#14311)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters own local or remote provider processes. > - Run cleanup must know when the final provider process has stopped. > - A remote timeout or a lost transport does not prove that the provider stopped. > - Retried provider invocations also make an earlier stop signal stale. > - This pull request adds a verified final-invocation stop callback before workspace restoration. > - The dependent instruction revision change uses that boundary to preserve private instruction edits safely. ## Linked Issues or Issue Description **What happened?** Adapter completion did not expose a reliable point between provider shutdown and workspace restoration. Cleanup could lose provider-written files, or treat a remote timeout as proof that a process stopped. **Expected behavior** Cleanup runs once after the final provider invocation has a verified stop receipt and before workspace restoration. An incomplete remote command keeps collection pending. **Steps to reproduce** 1. Run a remote provider that returns a timeout without a numeric exit code. 2. Let the adapter return or retry the provider. 3. Attempt to collect provider-written files during cleanup. The adapter has no verified final-process boundary to use. This is the prerequisite for the stacked canonical instruction revision pull request. It has no database or UI dependency. Related #13291 verifies remote termination for later recovery; this change exposes the earlier adapter-owned stop boundary before workspace restoration. The scopes do not duplicate each other. ## What Changed - Add a stop callback to adapter execution context and a final-invocation fence. - Confirm local child closure and complete remote exit receipts. Reject remote timeouts, missing exits, transport failures, and SSH exit 255 as stop proof. - Invoke collection once before workspace restoration in eight CLI adapters and at the confirmed ACP stop boundary. - Preserve the stop observation when a later log flush fails. - Run bridge and workspace cleanup in `finally` even when collection rejects. ACP records a safe error without exposing a raw filesystem path. Grok keeps collection errors separate from workspace restore failures and preserves completed provider results when both cleanup steps fail. - Correct the existing Cursor test shell fixture so bounded remote file reads run against real fixture files. ## Verification - Three stop-boundary regressions failed before the callback implementation and passed after it. - Independent prerequisite branch: 377 tests passed across 24 adapter, process-target, and ACP suites. Two added collector-rejection tests failed before the cleanup fix and passed after it. - A third regression reproduced Grok misclassifying a collection failure as failed workspace restoration. Two additional cases covered completed and failed provider turns when collection and restore both fail. The Grok and restore-classifier suites passed 48 tests. - All nine affected package typechecks and affected package builds passed; Grok checks passed again after its classification fix. - Integrated instruction branch: native local, legacy local, and native Daytona each passed three browser tasks with exact persisted bytes, fresh-task readback, history/restore, and explicit conflict resolution. Unchanged warm Daytona passed three turns. Legacy Codex passed all five checks again after the exception-safe cleanup fix. - Alternate staging passed the same three-task native Daytona flow: exact stopped-run save, independent downloaded readback, browser history/restore, and explicit resolution of a real concurrent edit. The deployed source was `14c3d810c9e05625121b3d27767aea9317b03125`, which covers the initial adapter callback. Later Grok cleanup failures are qualified by the adapter tests above. - The final combined native Daytona flow passed again on deployed `2bedd0f23bf4698b1f8b818f6796900647030427`: three fresh tasks proved ordinary instruction edits, independent readback, History/Restore, concurrent board conflict, and explicit candidate resolution. This native staging flow does not claim to exercise the Grok adapter. ## Risks - An unverified remote stop intentionally does not trigger collection. A later controller with verified stop evidence must recover it or report the copy unavailable. - The callback is optional. Callers that do not register it retain their existing behavior. - The callback runs before workspace restoration and can delay cleanup if its caller does not bound its own work. The dependent instruction collector uses bounded reads and retries. A rejected callback still permits bridge and workspace cleanup; it cannot claim an instruction save. ## Model Used OpenAI Codex, GPT-6, with tool use and code execution. The runtime does not expose the exact deployment variant or context-window size. Multiple Codex agents implemented and verified the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b38db9622 |
fix(runtime): stream workspace Git snapshots through disk manifests (#14253)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed runs copy a selected workspace to an execution environment and restore its changes. > - Git snapshots select the files for that copy and for later recovery. > - A fixed output limit stops large generated trees before the run can start. > - Increasing the limit still keeps the complete filename lists in memory. > - This pull request stores those lists and merge baselines in disk manifests. > - Large snapshots can now complete with bounded filename buffers and explicit failure handling. ## Linked Issues or Issue Description Refs #14194. This is the streaming follow-up to the merged 32 MiB limit fix. Related: #13619 and #11621 cover workspace scan admission and demand. This change keeps the shared scheduler and changes the snapshot data path. ## What Changed - Stream changed, untracked, deleted, and ignored paths through the shared scheduler and the standalone adapter path. - Use SQLite manifests for file selection, duplicate removal, ignored-path lookup, baseline capture, and merge lookup. - Set a configurable 30-minute snapshot deadline. Keep the existing interactive scan deadlines. - Wait for each child process and pending sink write before removing temporary storage after failure or cancellation. - Use NUL archive lists and bounded deletion batches. Preserve unusual names, explicit selection, nested repositories, and source root checks. - Store manifest references in native recovery format v2. Check their location and digest before recovery reads. Keep v1 descriptors readable. - Remove temporary manifests at lifecycle completion. Use fixed-size temporary copy names for long basenames. - Admit each manifest with a SQLite page allowance based on current disk capacity. Keep a configurable free-space reserve and fail explicitly when either limit is reached. - Preserve a host file that replaces a directory deleted by the sandbox, and continue the rest of the restore. ## Verification - Current head `5ab622ec43cd16d35429d79dedee6a5d8e3d2df2` has 54 successful checks/statuses and two skipped Storybook jobs. No checks failed or remain pending. - [CI passed](https://github.com/paperclipai/paperclip/actions/runs/36318966368): typecheck, build, all test shards, E2E, Rust checks, and the aggregate verify job. - [Greptile is 5/5](https://github.com/paperclipai/paperclip/pull/14253#issuecomment-5855670338) on the current head. All four review threads are resolved. Security checks passed. - 229 focused tests passed across Git sync, runtime staging, merge, manifest integrity, native recovery, and the scheduler (214 adapter/runtime tests and 15 scheduler tests). - A real 40,000-file fixture produces 43,428,890 filename bytes. The original standalone and scheduled scans fail. The new test passes all four filename paths, complete staging, exclusion of late files, unusual names, and deletion replay. - Recovery tests reject changed bytes, symlinks, and paths outside the controller state directory. Adapter-utils typecheck passed. - A test executor returned buffered output and caused two retry integration failures. The fixture now uses the shared streaming scheduler. All 13 tests passed with `corepack pnpm exec vitest run server/src/__tests__/heartbeat-project-repositories.test.ts`. The same CI shard now passes. - Ran `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` locally. Each full local command hit SIGKILL/exit 137 in the 4 GiB container. These local commands did not pass. The current-head CI gates above provide the full verification. - A real-Git disk-capacity regression confirms a typed failure and removal of the incomplete manifest. Repeated writer attempts cannot exceed the permitted page count. - Follow-up real Daytona and separate staging qualification passed with the related archive validator (#14315) and exact-owner finalization fix (#14314). Three successive turns copied back all 60,000 files with 39,828,890 filename bytes and five unusual names. Independent host inventories verified every file and the pinned Git HEAD. Native, provider, session, and process identities stayed fixed; no retry remained. The task reached Done, and its browser-downloaded final proof matched exactly. The reusable regression is #14316, including an assertion of the effective environment idle policy. ## Risks - SQLite manifests use disk space. Each receives one quarter of the available capacity above the host reserve at creation. The reserve defaults to 256 MiB and has a 64 MiB configuration minimum. Disk capacity, filesystem quotas, per-path limits, Git resource use, and execution deadlines remain limits. - Each path and sink chunk has a 64 KiB limit. SQLite connections use a 1 MiB page cache. Invalid or incomplete records fail explicitly. - Restore transport keeps fixed and configured archive exclusions. A remotely created Git-ignored file can be transferred, but the host merge excludes it through the manifest. - Provider archive buffers, Git and tar memory, repository metadata, legacy v1 arrays, and the separate referenced-source resolver retain their own limits. Existing provider safety validators still buffer textual tar listings: Daytona allows 32 MiB and Kubernetes allows 64 MiB. These separate transport limits can stop a sufficiently large restore before merge. This change does not claim bounded total process memory or unlimited transport size. - New descriptors use v2. Existing v1 recovery remains supported; a downgrade cannot read v2 descriptors. ## Model Used OpenAI GPT-6 through Codex. The exact deployment ID and context limit are not exposed in this run. The agent used code editing, terminal execution, tests, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
890d11137f |
fix(ui): register artifact tabs without opening the panel (#14193)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks keep agent outputs in the Artifacts tab.
> - An output can arrive while the user writes a message or reads a
document.
> - Opening the side panel on arrival interrupts that work, especially
on mobile.
> - This pull request adds the Artifacts tab without opening the panel
or changing the selected tab.
> - Users can open their outputs when they choose.
## Linked Issues or Issue Description
**What happened?**
New agent outputs opened the task side panel or mobile drawer. An
arrival could also replace the selected document or workspace file.
Existing outputs did not always register an Artifacts tab.
**Expected behavior**
Register one Artifacts tab for existing and new outputs. Keep a closed
panel closed. Preserve composer focus, the selected tab, and document or
file links.
**Steps to reproduce**
Open a task from the inbox. Close its side panel. Enter a message draft.
Create an agent output in that task. The panel must stay closed and the
draft must keep focus. Open the panel to see the Artifacts tab. Repeat
on a mobile viewport.
**Paperclip version or commit**
Base commit:
|