mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
f56aee54236f653ad191d82e70072a3ebc8894af
4579
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f56aee5423 |
fix(evals): capture complete durable run event streams (#13977)
## Thinking Path > - Paperclip manages AI agents and records their durable work outcomes. > - Product E2E checks those outcomes through the browser and public API. > - A successful long run can emit more than 1,000 durable events. > - The harness read one page and missed the later completion evidence. > - This pull request reads every page before it checks runtime invariants. > - Invalid or incomplete capture still fails. A missing page cannot produce a pass. ## Linked Issues or Issue Description Related: #13882 (Grok qualification). A duplicate search found no existing event-pagination fix. **What happened?** The local structured-question case in [campaign 36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537) completed the task and passed its six outcome matchers. It failed native runtime invariants because the capture contained exactly 1,000 events. The last captured event preceded the run's completion by more than a minute. The API caps each response at 1,000 rows. The harness did not request the next page. **Expected behavior** Read the complete durable event stream through the public API before checking semantic-result and terminal-event counts. Reject incomplete or malformed evidence. **Steps to reproduce** 1. Complete a native task that emits more than 1,000 durable events. 2. Place the semantic-result and terminal events after row 1,000. 3. Capture the run with the Product E2E harness. 4. Before this fix, the invariant checker sees only the first page. **Paperclip version or commit** Observed at `4196a4cd76db434854b679035e4146c7f69689ce`. The same single-page capture exists on master. The original failed result remains unchanged; missing historical tail evidence is not reconstructed or graded as a pass. ## What Changed - Add a bounded event collector that advances through the public `afterSeq` cursor. - Use it for task success/failure evidence and shared chat run evidence. - Reject invalid pages, missing or non-increasing sequence numbers, repeated cursors, failed later requests, and an exhausted page limit. - Test completion events beyond the first page, exact page boundaries, and malformed evidence. - Correct the existing Everyday catalog test from 38 to the maintained 47 cells. The suite stays explicit-only. - Document the complete-capture requirement and its bound. ## Verification - Product harness typecheck passes. - The full credential-free harness suite passed 477 tests. After adding the chat integration regression, all 34 chat evidence tests pass. - Seventeen pagination tests cover the valid tail and malformed-evidence cases. - `git diff --check` passes. - [Full repository CI](https://github.com/paperclipai/paperclip/actions/runs/36080683422) passes at `aed6f79c089226be79e75dcf390969561ec4f787`: 52 successful checks and two intentional skips, including Rust, typecheck, build, server tests, and browser shards. Greptile gives this exact head 5/5 with no findings. No local Docker, Rust build, browser suite, or model invocation was used for this change. - [Live Grok requalification](https://github.com/paperclipai/paperclip/actions/runs/36080870743) is running on combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, with the runner, scheduler, and evidence fixes. Its full credential-free harness passes all 493 tests and typecheck. Live results are pending; the original campaign remains a failed measurement. ## Risks Long runs need more read-only API requests and larger private evidence files. Capture stops with an explicit error after 100 full pages. This changes neither production APIs nor provider behavior. It does not relax an invariant or change a historical grade. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c96c751f39 |
fix(heartbeat): claim task ownership with queued runs (#13973)
## Thinking Path > - Paperclip manages agents and their tasks. > - The scheduler can start several runs for one agent. > - Each task still needs one execution owner. > - Task creation and recovery can queue assignments for the same task. > - The scheduler previously marked both runs running before starting either executor. > - This change claims the run and task ownership in one transaction. > - Tasks with different IDs retain concurrent execution. ## Linked Issues or Issue Description Refs #13882. Related queue work: #11830 and #13965; those address different admission and cancellation paths. **What happened?** The Grok subscription Product campaign exposed two assignment runs for one task. Task creation and periodic recovery started runs within 200 ms. Both acquired Daytona sandboxes. One failed before provider work with `paperclip_runner_attachment_staging_not_authorized` because the other owned the task. Fixture cleanup then found an active session. The original failed result is retained in [campaign 36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537). **Expected behavior** Only the task execution owner may start provider setup. Competing queued work must wait. Different tasks may use the agent's available slots. **Steps to reproduce** 1. Set an agent's concurrent run limit to two. 2. Queue task-creation and recovery assignment runs for the same task. 3. Resume the queue while holding adapter execution open. 4. Before this fix, both runs become running. Only one has the task execution lock. **Paperclip version or commit** Reproduced on master `8781f06a8` and in the Grok campaign at `4196a4cd`. ## What Changed - Use the same company-scoped task ownership gate for assignments, direct comments, and queued comments. - Leave competing work queued while another live run owns the task. Transfer a terminal pointer only after the tracked executor and durable environment leases/finalization settle, including cleanup on another controller. - Commit the running state and task execution owner together. - Preserve the dedicated review path and the native replacement checkout guard. - Add sixteen database regressions for assignment/comment orderings, unrelated concurrent tasks, and cross-controller cleanup fences. ## Verification - Live subscription Product E2E planning passed on its first attempt in both Daytona (351,197 ms) and local execution (268,883 ms), with 6/6 matchers and cleanup passing in each, on combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e` ([campaign](https://github.com/paperclipai/paperclip/actions/runs/36080870743)). The local screenshots and persisted state show the same plan revised, the first approval rejected, the revised approval accepted, and one completion. The full campaign remains unqualified because of a separate startup retry and EC2 Spot interruption. - The new same-task regression failed before the fix: two running rows instead of one. - All seven focused concurrency cases passed. Three comment cases failed before the shared gate was added. Five cross-controller cleanup cases failed before the durable gate was added. A warm-retention regression also failed before its successful release receipt was admitted; missing/failed receipts and a different retention policy stay blocked. - `node ../node_modules/vitest/vitest.mjs run src/__tests__/heartbeat-stale-queue-invalidation.test.ts` from `server`: 48 passed. Native cleanup admission adds one passing test; task-drain release adds two. Total focused checks: 51 passed. - `git diff --check`: passed. - Current-head [CI](https://github.com/paperclipai/paperclip/actions/runs/36078495577) passes repository typecheck, tests, Rust checks, build, and browser suites. The unrelated repository-form browser shard passed one bounded rerun without assertion or code changes; its original navigation failure remains retained. No local Docker or Rust build was used. - The failed test's exact two sandboxes were checked: one was already absent; the other was deleted after its company, environment, and run labels matched. ## Risks The claim transaction adds a task row lock. It follows task-before-run lock order. A competing run stays queued until the current owner releases the task. The change does not alter attachment authorization, provider permissions, schemas, or public APIs. Review is 5/5 with all threads resolved on `23c0d13b6`. All CI checks pass on that exact head. Live Grok requalification will use the combined runner and scheduler source. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ### Live regression verification The Daytona plan/revise/accept lifecycle in [Grok campaign 36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743) passed on its first attempt at combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, in 351,197 ms, with all six outcome matchers and cleanup passing. The earlier competing-run failure remains retained. This is one live planning result; the complete Grok matrix and repeated qualification are separate gates, and this campaign also has separately retained infrastructure failures. --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.925.0-canary.1 |
||
|
|
efce9356b5 |
fix(ui): offer recovery when the app fails before React starts (#13970)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The browser must load its JavaScript before React can render a task. > - A failed import can stop that process before the React error boundary exists. > - The HTML entry then leaves an empty page with no recovery action. > - This PR adds a small recovery screen that works without React. > - The user can retry the same page and return to saved task content. ## Linked Issues or Issue Description Refs #13824 and #13895. This is a follow-up to their browser startup investigation. **What happened?** Interrupting the app bundle or a required import leaves an empty React root. A startup exception has the same effect. A React error boundary cannot handle these failures because React has not started. **Expected behavior** The page must explain the startup failure and offer a manual retry. A late successful load must dismiss the recovery message without a reload. **Steps to reproduce** 1. Open a saved task in the browser. 2. Abort the application bundle request, or make a required module return HTTP 503. 3. Observe the empty page before this change. With this change, use Reload page after the fault clears and verify the saved task and comment. **Paperclip version or commit** The failing regression baseline used master at `8781f06a8`. **Deployment mode** Local source build and compiled UI. Tests cover both initial navigation and a page controlled by the production service worker. The exact cause of the older intermittent Vite stall remains unconfirmed. Forty app loads and thirty replays of retained responses did not reproduce it. This PR fixes the missing recovery path; it does not claim to remove that historical cause. A normal HTTP 304 response is not a failure. ## What Changed - Add an inline startup guard and recovery screen in the HTML entry. It does not depend on the app module graph. - Show a manual reload action after a startup error or after 30 seconds without rendered root content. - Remove the notice, timer, observer, and error listeners when the app starts. Never reload automatically. - Keep the recovery screen outside the React root so it cannot satisfy app-readiness checks. - Add browser tests for interrupted imports, a stalled import, an evaluation error, service-worker-controlled retry, repeated offline retry, and cleanup after successful startup. - Return a static, uncached HTML retry screen when a service-worker-controlled navigation fails offline. It contains no task content. - Add a full-app test that retries an interrupted compiled bundle and checks the saved task, comment, composer, route, and absence of agent runs. - Document the coverage and the limits of the historical diagnosis. ## Verification - Red baseline: four recovery cases failed; the normal-startup case passed. After the change, all five recovery cases passed. The review found an offline retry gap; that additional case failed before the worker fix and passed afterward. - Full provider-free browser-support suite: 16 passed. - Compiled-app browser tests: four passed, including saved-task reload, interrupted-bundle recovery, slow-CPU service-worker reload, and sidebar navigation. - Expanded service-worker, offline response, PWA, and worker build-ID unit tests: 37 passed. The two old plain-text offline expectations were reproduced as failures and updated for the HTML retry contract. - UI production build, full local repository typecheck (`pnpm -r typecheck`), runner-E2E typecheck, and design token checks passed. - Manual browser check: a temporary server failed the compiled bundle once. The recovery screen appeared. Clicking Reload page restored the same saved task, comment, and composer. - Full local `pnpm build` passed. - Full local `pnpm test:run` was attempted with a bounded deadline and stopped after it timed out. Workspace runtime/cleanup tests reported timeouts on this host. The monolithic local run is not a pass. The focused tests above and the complete Linux CI run provide the successful verification. - Final-head [CI run](https://github.com/paperclipai/paperclip/actions/runs/36072201966) passed. All 53 check runs succeeded; the two Storybook jobs were intentionally skipped. The legacy security status also passed. - Greptile reviewed `f83e0f51fb760541d83353f2c1df4e182f3948f9`: 5/5. Both review findings are fixed and resolved. ## Risks - The guard only handles startup before React renders root content. Existing React boundaries handle later rendering errors. - A slow startup can show the message after 30 seconds. A later successful render removes it; the page does not reload by itself. - The fallback uses native HTML when the app stylesheet is unavailable. - The worker changes only its offline navigation response. It returns static HTML with a reload button and `Cache-Control: no-store`. Its cache allowlist, private-response protections, task state, provider prompts, and grading rules stay unchanged. - This does not establish or fix the unknown cause of the historical intermittent Vite stall. ## Model Used OpenAI GPT-6 through Codex. The session exposes the GPT-6 family but not an exact served model ID or context window size. Used reasoning, code editing, shell tools, and browser testing. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.925.0-canary.0 nightly/v2026.925.0-nightly.0 |
||
|
|
8781f06a87 |
feat(connections): enable MCP aggregators by default (#13964)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Connections let those agents use external services with explicit access rules. > - Zapier, Arcade, Composio Connect, and Executor already have setup and runtime support. > - Their experimental switch still blocks discovery and setup by default. > - This pull request removes those gates and the Settings toggle. > - Users can connect these providers without enabling an experiment. ## Linked Issues or Issue Description Refs #13755. Refs #13941. **What existing behavior does this improve?** Apps browsing, inline setup, and agent connection search for the four MCP aggregators. **Current behavior** An instance must enable the MCP aggregators experiment before users or agents can start setup. **Proposed behavior** All four providers are available by default on local and managed instances. Old stored and managed values still parse but cannot disable them. ## What Changed - Remove the aggregator gates from Apps, inline setup, server setup, and agent search. - Remove the Settings toggle and its UI hook. - Retain the old setting key only for upgrade compatibility. Normalize it to true and ignore managed overrides, as Apps already does. - Replace opt-in fixtures with default-on coverage. Test old false values, all four setup flows, provider choice, and the removed toggle. - Update current connector guidance and remove the opt-in from the runner acceptance fixture. ## Verification - 306 focused tests passed across eight files: shared remote MCP contracts; server remote MCP lifecycle, aggregator fallback, settings normalization, and managed overlay; UI Apps browsing, setup, and experimental settings. - Server and UI TypeScript checks passed. - UI token gates and `git diff --check` passed. - The full local suite was not run, per the maintainer's instruction. All 54 CI checks passed; two checks were skipped. One unrelated workspace-preview readiness timeout passed on one failed-shard retry. - The setup fixtures use simulated MCP responses. This change does not claim new live provider acceptance. ## Risks - Existing instances now show all four providers, even if the old flag was false. This is intentional. - External provider choice, credentials, company isolation, agent grants, and tool policies still apply. Showing a connector does not authorize an external account. - No data migration is required. The compatibility key keeps old managed configuration documents valid. - Historical Zapier live acceptance remains incomplete in the existing evidence report. The maintainer explicitly requested the default-on rollout for all four existing providers; the report records that scoped exception. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, repository tools, and test execution. The context window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.5 |
||
|
|
aa8fc86331 |
feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external services. > - Native connections should remain the first choice for a supported app. > - Other apps may be available through an external MCP provider. > - The user must know which external provider handles the connection and choose it before setup. > - This pull request adds ranked alternatives and server-authored instructions to connection search. > - Agents can follow the returned instructions while Paperclip validates saved choices and access. ## Linked Issues or Issue Description Related: #13879, which fixed inline MCP provider setup. This PR adds discovery and provider selection on top of that work. **Subsystem affected** Cross-cutting: shared connection contracts, server search and intent services, native runtime, CLI, inline setup UI, and evals. **Problem or motivation** An agent cannot offer a clear external-provider choice when Paperclip has no native connection for an app. Adding provider-specific branches to the core prompt would make those instructions harder to maintain. **Proposed solution** Prefer a native connection. Otherwise return verified alternatives in Composio, Arcade, Executor, Zapier order. Include an external-service disclosure, a question with None, and the next instruction in the search result. Validate the saved human choice before creating a selected fallback setup card. Reuse existing provider accounts and verify underlying app access separately. **Alternatives considered** Do not silently choose a provider. Do not claim that broad execution tools prove support for every app. Reuse existing questions and connection intents rather than add another connection model. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback capabilities. The MCP aggregators experiment remains the gate. No duplicate provider-routing PR was found in the public search. ## What Changed - Add a dated support index and authorized cached-tool evidence for external routes. - Return provider questions and next-step instructions from `connections_search`. - Preserve pending choices and declines across continuation. Validate company, task, agent, human, app, and current route eligibility. - Carry the selected app into new setup and account reuse, validate explicit provider requests against persisted human messages, and distinguish provider readiness from app authorization. - Sync native, MCP, REST, and CLI contracts. Keep core agent instructions provider-neutral. - Add production-component Storybooks, focused database tests, and three real-agent browser eval cases. - Record the plan, observed failures, fixes, passing evidence, and acceptance limits. ## Verification - Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and all review threads resolved. - After rebasing on master `18dac1e1e`: 64 focused shared, validator, route-contract, and database tests passed; server typecheck passed. - Embedded-browser test drive on the rebased head: native Jira card, HubSpot external-provider question, Arcade account reuse, one actual MCP read against a local synthetic fixture, reload persistence, and None preventing further calls. A real OpenAI-backed agent performed discovery and continuation. - UX observation: the agent initially combined mutually exclusive request fields; the server rejected it and the agent recovered without changing access. This extra retry remains visible in the transcript. - After rebase: 23 focused eval grader/catalog tests, affected TypeScript checks, token gates, production UI build, and Storybook build passed. Full local tests are intentionally excluded at the maintainer's request. - Before rebase: four browser/real-agent attempts passed: native Jira, None, and reuse of the second provider on two Codex profiles. - Browser evals used an isolated deterministic MCP fixture through the real Paperclip gateway. They do not prove production compatibility with all four providers. - Review `Apps / Connections / Provider choice` in Storybook. Choose Arcade, continue through Access, and verify the app name, external-service disclosure, and URL configuration. - The detailed verification report is `doc/connections/2026-09-23-aggregator-routing-verification.md`. ## Risks - The public support index is finite and can age. Account capability and app authorization still require verification after selection. - Existing installed-tool permissions remain in effect. Provider choice is not a new execution permission boundary. An early Mini attempt skipped search; clearer provider-neutral instructions made the targeted rerun pass. This is not a measured reliability rate. - Explicit requests skip provider confirmation only when a clear persisted human message or saved provider choice supports them. Other phrasing falls back to confirmation; the agent query alone is not consent. These routes do not add tool permissions. - No database migration or legacy Composio broker is added. Real-provider acceptance remains separate from fixture proof. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser tools. The exact deployment variant and context-window size are not exposed in this session. Product evals separately used the repository's primary Codex and Codex Mini profiles; those agents supplied test behavior, not independent provider compatibility proof. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.4 |
||
|
|
18dac1e1ef |
feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path > - Paperclip manages AI agents and their work. > - Connections give agents governed access to external tools. > - Agents need durable memory across tasks and execution environments. > - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory tools. > - This pull request adds their setup flows behind an experimental toggle. > - It also delivers assigned MCP tools through the native remote Codex runner. > - Operators can connect a provider once and use the same governed tools locally or in Daytona. ## Linked Issues or Issue Description **Problem or motivation** The Apps catalog lacks a complete set of memory providers. Remote native Codex agents also need access to assigned managed MCP tools without receiving provider credentials. **Proposed solution** Add five memory connectors behind the disabled-by-default Experimental memory connectors setting. Use the existing connection setup and permissions UI. Default their tools to Allowed. Preserve company boundaries, operator permission changes, provider scopes, and audit attribution. **Alternatives considered** Direct provider credentials in each sandbox would duplicate setup and bypass the managed gateway. The remote runner instead uses its existing protocol channel to call the gateway on the server. **Roadmap alignment** This maintainer-requested experiment supports the Memory / Knowledge and Connected Apps roadmap areas. It adds provider connections without introducing a separate memory UI. Related connector authoring documentation is tracked in #13692; no duplicate memory-provider implementation was found. ## What Changed - Add provider definitions, official branding, and the experimental setting for all five providers. - Use OAuth for Zep and Supermemory, API credentials for Mem0 and Honcho, and a bundled Cloud API bridge for Cognee with no runtime downloads or subprocesses. - Default memory tools to Allowed and classify destructive actions explicitly. - Fix personal remote credential resolution and propagate provider tool errors. - Relay assigned managed MCP tools to remote native Codex through the runner protocol. Recheck current authority for each call and rotate stale tool contracts. - Add 21 Storybook states and complete OAuth walkthrough fixtures. - Document provider research, sanitized tool inventory, and live local and Daytona proof. ## Verification - Passed workspace typecheck: `pnpm -r typecheck`. - Passed production build: `pnpm build`. - Full local `pnpm test:run`: 13,273 passed, with failures from process/readiness timeouts under parallel load. Reran all 16 affected suites with one worker: 456 passed, leaving two macOS `/var` versus `/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`. No test failures remain unverified. Latest-head remote CI passes all 54 checks (two optional Storybook jobs skipped). Greptile is 5/5 with no unresolved findings. - Latest Cognee gateway regression: 71 passed, including public deployment without a runtime host and immediate recovery after a provider error. Bundled bridge tests: 25 passed. - Browser setup and real agent tasks exercised all five providers. Mem0, Cognee, Zep, and Honcho have successful store/retrieve proof. - Supermemory now has scoped read/write consent. Local storage and real Daytona write, document read, and semantic recall passed; indexing completion was verified before claiming success. - All five providers were exercised through a real Daytona sandbox and its native runner MCP relay. Provider credentials remained on the server. The final bundled Cognee bridge also passed a fresh Daytona store/recall run. All disposable sandboxes were removed and verified absent after testing. - Storybook is rebuilt and contains the experimental toggle, catalog, setup, permissions, and error states. The Zep and Supermemory access steps advance correctly. ## Risks - Provider OAuth scopes and plan limits remain independent of Paperclip tool permissions. An Allowed tool can still be rejected by the provider. - Providers can queue memory indexing; save acceptance does not prove that semantic recall is ready. - Remote tool contracts must stay synchronized with current connection authority. Regression tests cover revocation and stale contracts. - This change has no database migration. Existing connections remain usable when the experimental catalog toggle is disabled. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository edits, code execution, and browser automation. The exact context window size is not exposed in this session. Live acceptance agents used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e006c18f20 |
test(paperclip-runner): stabilize the assigned-skills ACPX runtime host test (#13453)
Release sealed skill homes before test cleanup and reuse one prepared sandbox across the assigned-skills test opens. Preserve the current skill-refresh and prompt assertions without increasing test timeouts. Co-Authored-By: Devin Foley <devinfoley@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.3 |
||
|
|
2914501c08 |
test(chat): retire fixture leases before recovery sweeps (#13952)
Retire verifying chat fixtures and their synthetic endpoint leases after service shutdown. Add a regression that recovers a stale receipt while preserving another company's lease. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f3a214fe77 |
Fix relative symlinks in secondary sandbox repositories (#13953)
## Thinking Path > - Paperclip manages agents and their task workspaces. > - A task can use several independent Git repositories. > - Sandbox staging copies secondary repositories from temporary clones. > - The copy changed relative symlinks into absolute host paths. > - Those links broke skill discovery and made workspace restore fail. > - This change preserves link targets while keeping extraction checks intact. ## Linked Issues or Issue Description **What happened?** Staging a secondary repository rewrites a link such as `.claude/skills/demo -> ../../skills/demo` to an absolute path in a temporary Git clone. That clone is then removed. The link is broken in the sandbox, and Daytona refuses the outbound archive during workspace restore. An agent can finish its turn but still have its run fail during restore. **Expected behavior** Repository links keep their original targets after staging. Links within a repository remain usable, and changes return to the local checkout. Unsafe outbound archive links still fail before extraction. **Steps to reproduce** 1. Create a project with a primary repository and a secondary repository. 2. Commit a relative skill directory link in the secondary repository. 3. Stage the workspace for sandbox execution and inspect the copied link. 4. Restore that repository through Daytona. Before this fix, the link points at a removed host temporary directory and restore rejects it. **Paperclip version/commit** Reproduced on `0f8750627f11d855552abce9a837d7f3b67c9ddf` in the multi-repository sandbox path. Related: #13442 introduced multi-repository provisioning. #13882 adds native Grok but keeps legacy adapters; #12991 addresses Grok instruction isolation and leaves skill staging unchanged. Searches of open/closed PRs and open issues found no direct fix for this copy behavior. ## What Changed - Set `verbatimSymlinks: true` when copying secondary Git clones. This [Node option](https://nodejs.org/api/fs.html#fspromisescpsrc-dest-options) preserves the stored link target instead of resolving it against the temporary source. - Test directory, file, chained and dangling links after temporary-clone cleanup. Assert the copied Git checkout remains clean. - Cover skill-link reads, edits and restore in fresh, warm-adoption and durable-seed workspace modes. - Extend Daytona checks for valid relative directory links and rejected absolute targets. - Document the staging behavior. ## Verification - Four regression cases fail without the source fix: one clone test and three staging modes. - `pnpm exec vitest run packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts`: 301 passed. - `pnpm -r typecheck` and `pnpm build`: passed. - `pnpm test:run` was started locally, then stopped after the full Linux CI suite passed. It did not finish locally; this is not a full local-suite pass. - Full PR CI: 53 checks passed, two conditional skips, on `07b30ff298baab327a4e60708ade899935688a3f`. Three server shards were interrupted by runner shutdowns; the unchanged mobile repository test timed out waiting for a disabled Save changes button. One same-commit failed-job rerun passed. Original attempts remain in [run 36036538369](https://github.com/paperclipai/paperclip/actions/runs/36036538369). - Greptile: 5/5 on the same head, no review threads or actionable findings. - Diff scanned for secrets and private identifiers; no matches. ## Risks Low risk: the production change is one copy option. It preserves symlinks instead of following or materializing their targets. Daytona extraction guards, workspace exclusions, authentication and database behavior do not change. This prevents corruption in newly staged snapshots. It does not rewrite an already corrupted warm workspace or durable seed; those need fresh staging from the source checkout. No live provider run or customer-task replay was performed. The tests use real Git, filesystem and tar operations with mocked provider transport. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.2 |
||
|
|
0f8750627f |
fix(codex): preserve typed ACP quota classification and reset time (#13945)
## Thinking Path > - Paperclip runs agents and recovers failed tasks. > - Codex ACP can report usage exhaustion as a typed terminal failure. > - The shared engine removes provider text before it stores the result. > - Codex did not use the existing terminal classifier hook, so a quota failure became a generic turn failure. > - This change classifies explicit usage exhaustion before that text is removed. > - Recovery can then use quota backoff and the supported reset clock. ## Linked Issues or Issue Description **What happened?** A Codex ACP `limit` failure with explicit usage-exhaustion text produced `acpx_turn_failed`. Recovery could schedule the ordinary short retries because the quota classification and reset time were lost. **What did you expect to happen?** Keep the failure visible and use the existing provider-quota wait. Preserve a supported reset timestamp without storing provider text. **Steps to reproduce** Return a terminal ACP session failure with category `limit` and title `You've hit your usage limit for GPT-5. Switch to another model now, or try again at 4:30 PM (America/Chicago).` The real child-process regression tests exercise both pinned ACPX versions in persistent and oneshot modes. **Paperclip version** Base commit: `32573876d4`. Related work: #13651 and #13831 provide the shared hook and Claude classification. #11854 includes quota handling as part of optional credential rotation, but reads the already-sanitized result; this patch handles the typed terminal boundary without adding rotation. #13549 reads recovery text after this boundary and cannot recover discarded provider text. #9011 concerns the Codex CLI backoff. This change leaves those other mechanisms in place. ## What Changed - Register a Codex terminal-failure classifier with the existing ACP engine hook. - Recognize explicit usage exhaustion only in a typed `limit` failure. - Reuse the Codex reset-time parser and existing provider-quota recovery fields. - Leave context, turn, rate, budget, storage-capacity, and unknown failures on their existing paths. - Test real ACP children, both dependency patches, privacy, and recovery classification. Document the boundary. ## Verification - 88 focused tests passed across Codex ACP, parsing, and server recovery classification. - An initial cross-adapter run passed 53 tests, including the existing Claude quota suite. - Removing only the classifier registration makes five new integration tests fail; restoring it passes all 19 new tests. - `pnpm -r typecheck` and `pnpm build` passed. - Full Linux CI passed on `b10d60e002`, including all unit/integration shards, browser shards, typecheck/build, runner checks, and release/package gates. - The duplicate local `pnpm test:run` reported three skill-cache failures and two runner-suite failures. It is not claimed as a full local pass. - The three skill-cache failures reproduce in a clean worktree at the unchanged base commit (3 failed, 107 passed, 3 skipped across the skills and runner files). The cache rename reports `EACCES` on macOS. - An isolated runner-suite run reports two embedded PostgreSQL startup failures before its assertions (37 tests pass). No runner source is changed; the Linux CI runner and server suites passed. - Greptile reviewed the current head at 5/5 with no actionable findings or unresolved threads. ## Risks - Explicit usage-exhaustion failures now wait for quota recovery instead of short generic retries. - Unknown wording retains the existing behavior. The generic historical terminal-limit message alone cannot establish quota exhaustion. - Reset parsing keeps the existing Codex clock formats. A missing or unsupported reset uses the existing quota backoff. - No schema, credential, UI, dependency, or deployment changes. Provider text stays in memory and is absent from results and logs. ## Model Used OpenAI GPT-6 via Codex. Exact runtime model variant and context-window size were not exposed. Used reasoning, repository inspection, editing, and local tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.1 |
||
|
|
32573876d4 |
chore(lockfile): refresh pnpm-lock.yaml (#13840)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
b0155a681a |
feat(slack): connect Paperclip conversations and scheduled messages (#13920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack conversations use the same tasks and agents as the Paperclip board. > - A board reply must reach that Slack conversation and let the agent continue the work. > - An assigned agent also needs its Slack tools during normal tasks and scheduled routines. > - Both paths must keep the linked user's authority, delivery rules, and conversation history. > - This pull request adds those paths and reduces setup friction for Slack bots. ## Linked Issues or Issue Description **Subsystem affected** Server orchestration, Slack connector tools, shared contracts, and chat setup UI. **Problem or motivation** Replies entered in Paperclip did not provide a complete round trip to the linked Slack thread. Slack and board wakeups could select different model sessions for the same task. Agents also lacked their assigned Slack tools outside Slack-origin work, which prevented a routine from sending its responsible user a briefing. Inviting a bot could leave the new channel disabled. **Proposed solution** Mirror human board messages with author attribution and route the agent result to the same thread. Use the same session key across both entry points. Supply Slack tools to the connection's assigned agent in normal tasks and routines, using the current responsible user's verified link. Enable newly invited channels while preserving explicit disabled choices. Add a browser-agent setup prompt to the Slack wizard. **Alternatives considered** A separate Slack scheduler or task dispatcher would duplicate existing Paperclip workflows. Reusing the connection owner's identity would grant the wrong authority. Replaying old channel history could start unintended work. This change uses ordinary task wakeups, routine dispatch, and Slack's original invitation mention event instead. **Roadmap alignment** Extends the shipped Scheduled Routines and governed Apps capabilities. It does not add a separate task lifecycle. Related work: #13828 and #13809. Related test stabilization: #13877. The existing plugin Slack-control proposals are separate from this built-in connector change. ## What Changed - Queue human Paperclip messages for the original Slack thread with display-name attribution and stable delivery identities. Require the author’s current linked Slack identity and recheck access before delivering messages or agent replies. - Apply pause, dependency, cancellation, and closed-workspace guards before explicit Board sends request work and again when the durable outbox dispatches it. - Route agent results back to Slack and preserve model-session continuity, including replies that reopen completed tasks. - Resolve assigned Slack connections for normal agent tasks and routines. Recheck the responsible user's link, membership, and permissions at execution. - Add `slack_open_dm` for the responsible user's bot DM and request the `im:write` scope. - Enable newly discovered invited channels. Keep explicit OFF choices and normal admission and deduplication rules. - Add a copyable Slack setup prompt for a computer-use agent, with Storybook coverage. Share the prompt-button component with GitHub. - Update Slack tool documentation and runtime instructions. - Stabilize the mobile project browser test by waiting for the final canonical route before editing, preserving all persistence assertions. ## Verification - Live staging: invited the bot after the first mention. The channel became enabled and the bot answered that original mention. - Live staging: a normal Paperclip reply appeared in Slack with author attribution. The agent completed the calculation and replied once in the original thread and in Paperclip. - Live staging: a codeword entered in Slack was recalled from Paperclip. A following Slack calculation used the result from the Paperclip turn. Run metadata confirmed the same model session for both entry points. - Live staging: a scheduled routine used `slack_open_dm` and `slack_post_message` to deliver one DM. The existing app was reinstalled with `im:write`. The test routine was paused after verification. - Before the master merge: 397 focused feature tests passed. The continuity fix passed all 76 issue comment/update route tests and six focused route/integration cases. Typecheck, build, and token gates passed. - Review fixes: 47 focused integration cases passed, covering link revocation/replacement, private membership removal, guarded outbox dispatch, concurrent workers, lost scheduler responses, a real one-connection pool, exact reply provenance, and attachment retries. All 99 issue-comment route tests and the Slack catalog browser test passed. - Full local typecheck, production build, token gates, and module-boundary checks passed. The full local test command passed 25,915 tests before a 15-second timeout in `issue-thread-interaction-routes.test.ts`; that entire suite passed on isolated rerun (81 tests). Remaining serialized coverage is provided by the current-head CI shards. - An unchanged Cursor adapter test hit its 10-second limit in CI; all five tests in that file passed on a local rerun in 3.11 seconds, and the failed CI shard passed on its single retry. - The preview-server readiness test passed a local rerun (28 tests). The mobile-project readiness fix passed three repetitions of both browser tests (6/6). - Final commit `64ac0d9897f4353375996f1b1b38e5040bdeb0a0`: all CI gates passed, including all eight browser shards, all server/chat suites, typecheck, build, runner checks, and security checks. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - Human messages on a Slack-linked task now publish to its Slack thread. The task banner states this behavior. Incoming Slack messages and internal agent bookkeeping must not echo back. - Normal tasks and routines can now use the assigned bot. Authority remains bound to the current responsible user's link; it does not fall back to the connection owner. Revocation, private-context limits, and queued-write checks still apply. - Existing Slack apps need `im:write` and a reinstall to open DMs. Other existing capabilities remain available without that scope. - New invited channels default to enabled. Explicit disabled choices remain disabled. Channels created by bot tools still require a person to enable responses. - No database migration or new provider credentials are required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, GitHub CLI, and browser tools. The runtime does not expose a more specific model build identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c341588bdb |
fix(ui): add Cloud invitations to the Members page (#13922)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People manage collaborators from the Members page. > - Cloud manages invitations outside the tenant's local invitation system. > - The local Invites tab is hidden on Cloud, so this page has no way to invite a person. > - This pull request adds an Invite people action for the current Cloud stack's owner or admin. > - The action opens the existing Cloud People settings for that stack. ## Linked Issues or Issue Description **What happened?** A Cloud owner opens Organization Settings → Members and finds no invitation action. The tenant-local Invites tab is hidden, and the page does not link to Cloud's invitation flow. **Expected behavior** Cloud owners and admins can start an invitation from Members. **Steps to reproduce** 1. Sign in to a Cloud-managed instance as the current stack's owner or admin. 2. Open Organization Settings → Members with `company.invites` hidden. 3. Look for an invitation action beside the page heading. **Paperclip version or commit** `7b7c4d4172d6aac14919e2682b702ae87bc17653`. **Deployment mode** Cloud-managed, authenticated. Related search: #2388 proposes broader member-management UI. This change only connects the existing Members page to Cloud invitations. No duplicate Cloud invitation action PR was found. ## What Changed - Add **Invite people** beside the Members heading for the current Cloud stack's owner/admin. - Read the role from the authenticated Cloud portfolio. Ownership of another stack does not enable the action. - Navigate to the current stack's People settings on the configured Cloud origin. Keep the local Invites tab hidden when configured. - Cover allowed roles, denied roles, loading, failed refresh, missing configuration, current-stack selection, and self-hosted behavior. Document the navigation contract. ## Verification - Focused Members and Cloud link tests: 23 passed. - Full UI suite: 6,669 passed across 634 files. - UI typecheck and `pnpm check:token-gates`: passed. - `pnpm build`: passed. - `pnpm -r typecheck`: passed. - All PR CI checks passed, including the full general, serialized, browser, and runner test jobs. The duplicate local repository-wide `pnpm test:run` was stopped after CI completed; it is not reported as a local pass. - Manual acceptance after tenant rollout: an owner/admin opens Members, selects **Invite people**, and reaches the same stack's Cloud People settings. A member does not see the action. ## Risks - The action needs a tenant app update before it appears on an existing stack. - The portfolio request must identify the current stack and its role. The action stays hidden when that information is unavailable or the request fails. - Cloud rechecks invitation authorization at the destination. No schema, API, or invitation-acceptance behavior changes. ## Model Used - OpenAI Codex, GPT-6, with reasoning, code execution, and repository tools. The session does not expose a more specific model identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7b7c4d4172 |
fix(ui): recover an archived Cloud stack from its own health probe (#13913)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - On Paperclip Cloud, a stack whose organization is archived is served through the Cloud harness router. > - An archived stack answers every tenant request — including the SPA's own health probe — with a 423 archived status page. > - If the SPA is already loaded (a stale tab, or a load that raced the stack's archival), that 423 surfaces as a dead "Failed to load health" screen with no way out. > - This pull request recovers it the same way an expired tenant session already does: a top-level reload re-enters the Cloud harness, which redirects a fresh browser navigation on an archived host to the portfolio. > - The benefit is that an archived stack never traps the user on a broken health screen. ## Linked Issues or Issue Description No public issue exists. Description follows the bug template: **What happened?** A Cloud stack that becomes archived while its SPA is open (or is opened in a stale tab) shows a "Failed to load health" screen. The SPA's health probe receives the harness's 423 archived status page and has no recovery path, so the browser is stuck on a dead screen. **Expected behavior** The browser should leave the archived stack for the Cloud portfolio, as it already does for an expired tenant session — no dead "Failed to load health" screen. **Steps to reproduce** 1. On a Cloud-managed instance, open an organization's SPA in the browser. 2. Archive that organization (or open a stale tab for one that was archived). 3. The SPA's health probe returns 423 (archived) and the page shows "Failed to load health" with no way out. **Paperclip version or commit** `master` (Paperclip Cloud managed instances). **Deployment mode** Cloud-managed (`authenticated`). Self-hosted instances are unaffected: the 423 archived status page only originates from the Cloud harness router that fronts managed stacks; a self-hosted server serves its own health. **Subsystem affected** UI: tenant document recovery (`ui/src/lib/tenant-session-recovery.ts`), which the health/client/heartbeats/audit API surfaces already route error responses through. ## What Changed - `isArchivedStackRecoveryError(status, body)`: true for a 423 whose body carries a `statusPage.code === "archived"`. - `isTenantDocumentRecoveryError` unifies the existing 401 tenant-session codes with the new archived case; the recovery coordinator now fires on either. - The recovery action is unchanged — a single top-level reload — so an archived stack re-enters the harness and is redirected to the portfolio. The existing single-reload guard prevents any loop, and only the exact `archived` code triggers it (suspended/deleted/other 423s pass through as normal errors). ## Verification - `pnpm exec vitest run src/lib/tenant-session-recovery.test.ts src/api/auth.test.ts src/api/audit.test.ts` — pass. - New tests: the predicate accepts a 423 archived body and rejects suspended/deleted/error-shaped/503/null; the coordinator triggers the same single top-level reload for an archived body and stays inert for unrelated responses. - `pnpm exec tsc --noEmit` in `ui/` is clean. - Pairs with the harness-side redirect (paperclipai/paperclip-cloud#545): the harness redirects fresh navigations on an archived host to `/orgs`; this makes an already-loaded SPA trigger that navigation on its own health 423. ## Risks - Low. The recovery path and its single-reload guard are unchanged; only the set of conditions that trigger it grows by one exact code. Non-archived 423s and all other statuses behave as before. ## Model Used - Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended thinking and tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergenightly/v2026.924.0-nightly.0 canary/v2026.924.0-canary.0 |
||
|
|
94a0aa7726 |
fix(ui): hide workspace isolation controls for managed hosts (#13907)
## Thinking Path > - Paperclip manages AI agents and their work. > - Workspace isolation keeps task checkouts separate. > - Managed hosts can enable isolation and hide its experimental toggles. > - Project, task, and routine forms still expose choices that override that policy. > - This pull request adds an operator visibility key for those controls. > - Workspace access stays available, and execution keeps its existing policy. ## Linked Issues or Issue Description **What existing behavior does this improve?** Operator control over workspace isolation settings in the UI. This follows the settings list cleanup in #13905. **Current behavior** Hiding the experimental isolation toggles leaves project policy editors, task selectors, routine and pipeline overrides, recovery actions, and workspace configuration visible. **Proposed behavior** Set `PAPERCLIP_HIDDEN_SETTINGS=workspaces.isolation` to hide these controls. Keep workspace navigation, status, files, and runtime access. Hide experimental toggles separately. Instances that do not set this key keep their controls. **Reason and benefit** Users on managed hosts should use the host's isolation default. A hidden form must not submit a stale draft that overrides it. ## What Changed - Add the UI-only `workspaces.isolation` key to the shared visibility registry. - Hide project workspace policy, task and subtask selectors, routine and pipeline overrides, and isolated re-issue actions. - Hide the workspace Configuration tab and redirect direct links to workspace issues. - Omit hidden new-task and routine overrides. Keep explicit task/subtask workspace launch context, saved policies, and automatic branch values for workspace routine runs. - Wait for the health visibility policy before showing controls. Keep workspace access and all execution APIs available. - Document the key and test visibility, form payloads, deep links, and unchanged workspace access. ## Verification - All 52 CI checks pass on `50cc770d33` (two expected skips). The branch is mergeable. Greptile is 5/5 with no unresolved comments. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm check:token-gates` passed. - Targeted UI checks passed: 346 tests across 14 suites, including hidden project/task controls, stale task drafts, routine branch defaults, recovery actions, configuration deep links, and workspace access. - Shared settings-visibility tests passed: 11 tests. - `pnpm test:run` was run and stopped after reproducing five failures in unchanged server tests: two `chat-channels.integration` cases (linked-request provenance and direct external-chat finals) and three `company-skills-service` cases (runtime refresh, concurrent download, and explicit update). The earlier local run for #13905 showed the same failures. The full local suite is not claimed green. Targeted UI/shared checks pass. All PR CI shards, including the affected chat and skills suites, pass. - Reviewed the diff for secrets, private links, and run artifacts. ## Risks This key changes UI visibility only. It does not reject API calls or change feature values. Operators must enable isolation and its default through their existing policy mechanism. Older app versions ignore the new key until upgraded. Removing the key restores the controls. No schema changes. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally; targeted checks pass (full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8ee8f1fd6e |
ci: retire recurring public cloud image builds (#13827)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Core publishes standard images and source verification for downstream services. > - Managed services can now compose private images from the signed standard image. > - Core still builds a second public cloud image on every master push and release. > - That duplicate producer consumes build capacity and retains an obsolete readiness contract. > - This pull request retires recurring cloud publication while preserving the standard producer and rollback artifacts. ## Linked Issues or Issue Description Refs #13797 and #13789. Related: #12856 changes image dependency packaging; it does not retire this producer. **What existing behavior does this improve?** Core's recurring Docker publication and Cloud readiness workflow. **Current behavior** Master pushes call the legacy cloud publisher from Cloud readiness. Release tags and manual Docker runs call it too. Canary promotion also requires the legacy image. **Proposed behavior** Publish standard Core images and retain `Cloud source verified v1`. Let downstream services build their managed image. Keep explicit commit previews and existing images available. ## What Changed - Remove `docker-cloud.yml`, its master and release callers, and its unused cache selector. - Remove the legacy image/migrator wait and `Cloud deployable v1` job. Keep the full source verification workflow and exact source-proof name. - Make canary promotion inspect and promote the standard image only. - Preserve signed standard-image publication, direct migrator publication, and explicit `release.yml` previews. The preview path still uses the Dockerfile `cloud` target. - Update workflow, preview, build-stamp, and packaging tests. Exercise the promotion shell with mocked registry commands, including missing-image and missing-tag cases. - Document frozen legacy aliases, consumer requirements, preview compatibility, and rollback retention. ## Verification - All 377 workflow tests pass: `node --test .github/scripts/tests/*.test.mjs`. - All 129 release-registry tests pass: `pnpm test:release-registry`. - Focused source-proof, standard-image, preview, and workflow tests pass: 256 tests. - Focused image packaging/build-stamp tests pass: 16 tests. - Actionlint passes on all three changed workflow files. `git diff --check` passes. - Full local `pnpm build` and `pnpm -r typecheck` pass. - The policy follow-up updates an old assertion that required the removed readiness job. All 37 source-proof/release-workflow tests pass locally. - Full local `pnpm test:run` did not complete successfully while the Mac ran out of disk space. No full-suite pass is claimed. Removed 1.2 GiB of generated Cargo output from this isolated worktree with `cargo clean`. GitHub CI passed on the final head: 52 successful checks and 2 optional skips. - Fresh Greptile review for `4f5fe1951f0bd7f7739cf6655d395ff78f1ed944`: **5/5**, successful current-head check, zero review threads. - September 23 refresh: the unchanged PR head merges cleanly with current master `db8f8fe5b73a2697684a30261b0d306a9c631aba`. In an isolated temporary worktree, all 377 workflow tests and 29 release/preview tests pass on the combined tree. `git diff --cached --check` passes. - Refreshed Actionlint workflow validation passes with ShellCheck disabled. Full Actionlint reports the same 10 existing ShellCheck diagnostics as master, with no added diagnostics. No source changes or new PR commits were needed. - The full local build/typecheck and current-head Linux CI results above remain the verification for the unchanged PR head. They were not rerun for this metadata-only refresh. No image publication or tenant deployment was initiated for this refresh. ## Risks **Deployment prerequisite satisfied (September 23):** The combined cleanup release is deployed to staging and production, and production Support is verified. Active managed-fleet automation uses standard-image composition. Explicit immutable previews remain supported by the retained preview publisher. This PR is ready for maintainer review; keep auto-merge disabled and wait for explicit merge authorization. - A consumer still selecting `Cloud deployable v1` will stop advancing at the last legacy-ready commit. Confirm active automatic consumers use the standard-image composition contract before merge. - Legacy cloud release-channel aliases stop advancing. Standard self-hosted aliases continue. - This PR deletes no registry images, cache tags, migrators, credentials, or runner infrastructure. Existing immutable releases remain usable for rollback. - Explicit legacy previews remain for commit-specific operator deployments. Retiring that compatibility path requires a separate consumer migration. - These changes affect CI publication, not database schema or application behavior. ## Model Used OpenAI Codex, GPT-6. The runtime does not expose a more specific model identifier or context-window size. Used repository inspection, reasoning, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db8f8fe5b7 |
fix(evals): select Grok subscription protocol credentials explicitly (#13901)
## Thinking Path > - Paperclip manages agent work through shared runner contracts. > - Direct protocol evals qualify provider behavior against a mock control plane. > - Grok supports API keys and company subscription credentials. > - The hosted protocol workflow selected an API key for every Grok cell. > - Product subscription support did not enable subscription protocol runs. > - This change adds explicit subscription selection and checks the recorded authentication mode. ## Linked Issues or Issue Description Refs #13878, #13882, #12618. The direct Grok protocol roster cannot run with subscription authentication through the trusted default-branch workflow. Add an explicit selector while keeping API-key dispatches compatible. Keep the actor allowlist, protected environment, immutable source revisions, and publication gates. ## What Changed - Add `grok_authentication` with `api_key` and `subscription` choices. Keep `api_key` as the compatibility default. - Deliver the protected `GROK_AUTH_JSON` secret only to a subscription-selected Grok cell. Do not provide an API key to that cell. - Read authentication mode from the pinned eval program's actual roster summary. Retain it in the cell, catalog, campaign roster, and result. - Reject missing or mismatched authentication evidence during aggregation. Preserve cell metadata and an allowlisted failure reason before failing a cell, so malformed evidence cannot hide the retained attempt. - Document credential setup, source separation, and temporary-secret cleanup. ## Verification - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-campaign.test.mjs packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`: 34 tests passed. - Validated all 39 Grok cells at eval revision `3213dbec7e8ca1865ea95e6db7e7d34b095eb47a`; every selected cell requests only the subscription credential. Validation made zero provider calls. - Negative coverage rejects invalid selectors and missing or API authentication evidence in an otherwise passing subscription attempt. - `git diff --check` passed. All 53 current-head checks passed; the unchanged callback-drain timing test passed its bounded rerun, and the failed attempt is retained. Greptile reviewed `ecd3dcc0e998f07cf56fcb1f087946f50f388bec` at 5/5 with no remaining findings. - No Docker or broad builds ran on the developer machine. CI performs repository checks on the configured fleet. ## Risks Grok runs require an eval revision that records `authenticationMode` in the roster summary. Missing evidence fails closed. The credential contains account access and refresh tokens; an owner must approve its delivery to the protected environment before a live run. The change adds no PR trigger or authorization bypass. Live subscription protocol qualification remains pending this workflow reaching master. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
794b09f834 |
fix(ui): alphabetize experimental settings and hide empty groups (#13905)
## Thinking Path Paperclip operators use Experimental settings to find and manage optional controls. New cards have accumulated outside alphabetical order, and operator-hidden cards leave empty developer and legacy sections. Sort the displayed controls within their existing groups and remove groups with no visible controls. ## Related Issues **What existing behavior does this improve?** The Instance Settings → Experimental page. **Current behavior** Agent Chat follows Cases, MCP aggregators precedes External Objects, and hidden developer/legacy controls leave empty headings. The worktree execution card also ignores its operator visibility key. **Proposed behavior** Cards appear alphabetically by their displayed title within each section. Empty developer and legacy sections disappear. The worktree execution card follows the same operator visibility policy as other experimental controls. **Reason and benefit** Operators can scan the list predictably. Hosted installations show only the controls their operator permits, without empty sections or a worktree-only exception. No matching sorting PR was found in the duplicate search. This is a small improvement to an existing settings page and does not add a roadmap feature. ## What Changed - Reorder existing cards without changing their values or mutation handlers. - Hide empty developer/legacy sections and respect the hidden worktree-execution key. - Cover alphabetical ordering, conditional isolated-workspace controls, a restricted three-control policy, and partly visible sections. - Document sorting and operator visibility. ## Verification - `pnpm --dir ui exec vitest run src/pages/InstanceExperimentalSettings.test.tsx`: 45 tests passed. - `pnpm check:token-gates`: passed. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm --filter @paperclipai/ui typecheck`: passed. - All PR CI checks passed on `88011049be`, including the chat integration shards, full typecheck, build, browser tests, and policy checks. Greptile is 5/5 with no review threads. - Local `pnpm test:run` reported eight failures in unchanged chat, company-skills, email-channel, and workspace exposure suites. Stopped the remaining local run after CI completed successfully. Isolated chat rechecks were skipped by the host database support gate and do not count as passes. All matching CI shards passed; the full local suite is not claimed as green. - Frozen installation is blocked on the base branch by existing overrides/patch configuration drift from the lockfile. Local validation uses `pnpm@9.15.4 install --no-frozen-lockfile`; the original lockfile and manifests are unchanged in this PR. - Diff reviewed for secrets, private references, and run artifacts. ## Risks Settings only change position or visibility. Existing values, managed locks, API contracts, and feature dependencies are unchanged. Each section keeps its own alphabetical list. Reverting this change restores the previous presentation. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the enhancement template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the relevant tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fff410dfe7 |
fix(ui): retain bounded context for browser render errors (#13904)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its browser error boundaries recover from failed renders and report the exception when error monitoring is enabled. > - The boundaries already receive the React component stack, but discard it before reporting the error. > - A minified DOM insertion error can therefore lack enough context to identify the affected component. > - Browser translation can replace text nodes that React still uses as insertion anchors. > - This pull request preserves bounded component names and browser state for the next failure. > - Maintainers can locate the failed component without collecting page content or customer URLs. ## Linked Issues or Issue Description Related: #13719 attributes browser errors to the loaded release. #13784 supplies the browser environment. This change adds context to error-boundary reports. **What happened?** A browser error boundary reports a DOM `NotFoundError` with only a minified JavaScript stack. The React component trace is logged to the console, which production builds remove. Browser translation is a plausible cause, but the report cannot identify the component or confirm the translation marker. **Expected behavior** An error-boundary report includes a bounded component trace and limited browser state. It excludes component props, text, HTML, element IDs, arbitrary CSS classes, page URLs, and query strings. **Steps to reproduce** 1. Render a component with conditional content before a text node. 2. Replace that text node with a translation element outside React. 3. Enable the preceding conditional content. React tries to insert before the detached text node and throws `NotFoundError`. 4. The boundary shows its recovery UI, but the old report loses the component trace. **Paperclip version or commit** Based on `6681c71b40`. The regression test reproduces the DOM mutation with the installed React version. **Deployment mode** Built browser UI with optional Sentry monitoring enabled for the signed-in session. ## What Changed - Pass the component stack and boundary kind from both error boundaries. - Keep at most 40 component names from at most 16 KiB of stack input. Drop locations and unrecognized lines. - Snapshot document readiness, visibility, and the browser translation root-class marker before the asynchronous reporting queue runs. - Attach diagnostics to that event only. Preserve the existing monitoring gate, sign-out behavior, and original exception if diagnostics fail. - Preserve function names in production bundles. Test the actual Vite production pipeline. - Document the fields, privacy limits, and translation-marker limitations. ## Verification - Focused diagnostics, boundary, real Sentry SDK, and production-build tests: 49 passed. - The translation DOM-mutation test reproduces `NotFoundError`, verifies the failed component trace, and keeps the recovery UI usable. - The real SDK test checks emitted events, private fixture exclusion, event isolation, and no capture after sign-out. Its transport stays in-process. - `pnpm check:token-gates`: passed. - [Greptile review](https://github.com/paperclipai/paperclip/pull/13904#issuecomment-5804220318): 5/5 on `92a2a23d9c`, with no review threads. - `pnpm -r typecheck` and `pnpm build`: passed. - Full CI unit and integration test matrix: passed, including all server, chat, workspace, runner, and serialized suites. - Local `pnpm test:run` was started, then stopped after the equivalent full CI matrix passed. The unsharded local run was not completed and is not claimed as a pass. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/project-repositories.spec.ts --repeat-each=3 --trace=on --reporter=line`: six tests passed against the production build. - Initial CI browser shard 8 timed out after the repository form replaced an enabled Save button with a disabled one before the click. The retained page snapshot shows the original repository selection. No error-boundary fallback appeared. [CI attempt 2](https://github.com/paperclipai/paperclip/actions/runs/35929990870/attempts/2) passed the failed jobs on the same commit. The full PR check matrix is green. This confirms an intermittent failure, but does not establish the cause of the first failure. - Full browser builds with and without name preservation passed. The initial JavaScript chunk grows from 1,571,919 to 1,668,453 gzip bytes (+6.1%). Total JavaScript across all chunks grows by 219,610 gzip bytes (+5.3%). ## Risks - This adds diagnostics. It reproduces a translation failure mode but does not identify or repair the specific application component from a past report. - The root-class marker is a hint. Other translation tools may omit it, and its presence does not prove causation. - Name preservation increases bundle size as measured above. No source maps are published by this change. - The error is still reported. DOM operations, browser translation, and recovery behavior are unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code editing, terminal tools, and test execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.14 |
||
|
|
a6448cd060 |
fix(runner): tolerate Codex account notifications during active turns (#13902)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Rust runner connects task runs to Codex app-server sessions. > - Codex sends account notifications on the connection during authentication updates. > - These notifications have no task, thread, or turn identity. > - The runner treated a valid account update as an invalid task event and stopped the run. > - This pull request classifies account updates as connection information while preserving task identity checks. > - Sandbox images can also contain an older runner with compatible metadata. The server must compare its bytes before reuse. > - Slack conversations use the deployed runner and can complete when Codex refreshes account state. ## Linked Issues or Issue Description Related: #13853. That change fixes the supported Codex version range. This fixes a separate notification failure after the version check succeeds. **What happened?** An active Codex task stopped with `thread_binding_mismatch` when the provider emitted `account/updated` without a thread ID. Slack showed that the agent stopped before completing its turn. **Expected behavior** Connection-level account updates must not stop a task or acquire task authority. Account details must not appear in task output. **Steps to reproduce** 1. Start a task with the native Codex runner. 2. Emit `account/updated` during the turn, with `authMode` and `planType` but no thread ID. 3. Observe that the old runner rejects the notification and fails the turn. **Paperclip version or commit** Reproduced after #13853. The fix is based on `b648d8cdd`. **Deployment mode** Cloud staging with a remote Codex runner and a Slack chat connection. ## What Changed - Classify `account/updated` and `account/login/completed` as connection information. - Reject account notifications that contain execution identity fields. - Reuse the existing bounded diagnostic path without publishing account payloads. - Test both notification types during two consecutive turns. Preserve existing identity rejection tests. - Compare preinstalled sandbox runner bytes with the controller artifact. Stage the deployed binary after a mismatch, failed checksum, or timeout. Keep exact retained artifact reuse. ## Verification - `cargo test --release --locked -p paperclip-runner-core`: passed, 582 test executions; two existing ignored cases. - `cargo test --release --locked -p paperclip-runner-core --test codex_provider`: 89 passed; two existing ignored cases. - `cargo fmt --all` and `git diff --check`: passed. - Repository-wide `pnpm -r typecheck`: passed. Server typecheck also passed after the artifact-selection change. - Server executor suite: 456 tests passed, including preinstalled digest match, mismatch, checksum failure, and timeout. - Repository-wide Vitest and build are running. - Deployed `5648d90d5dd762c5a8b697face269f6eea9cd59d` to one staging stack and verified its serving SHA. - Retried the failed conversation through the Slack UI. The agent replied and the task completed. - Sent a fresh Slack mention asking which bots belong to the channel. Verified successful `slack_members` and `slack_user` calls, a successful run, and the delivered Slack answer. - Sent a follow-up in the same Slack thread without another mention. The agent returned the requested names from context, with one final reply. ## Risks - Only two known connection-level notification methods change behavior. Unknown authoritative methods and malformed or conflicting identities still fail closed. - An older sandbox runner now requires one upload when its bytes differ from the controller artifact. Remote OS and architecture checks still apply. - No schema, credential, permission, Codex version-range, or UI changes. The Codex minimum remains 0.149.0. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code editing, terminal tools, and browser verification. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.13 |
||
|
|
30182c704c |
feat(telemetry): propose connector telemetry events behind proposal markers (#13582)
Adds connection.created / connection.updated / connection.invoked as proposed-telemetry wrappers (unregistered; the TelemetryClient cannot queue or send them) plus the server-side emitter wiring and tests. No telemetry is transmitted by this change. Proposal: #13578. Coordinated release plan and review package: PAP-24 (Step A.1).canary/v2026.923.0-canary.12 |
||
|
|
6681c71b40 |
fix(ci): stabilize chat startup and close the initial live-update gap (#13895)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Browser tests check chat state across navigation, reload, and agent runs. > - Runtime tests check that a service is ready before Paperclip publishes its address. > - CI for #13891 failed in these paths, then passed attempt 3 with the same code. > - The browser traces stopped during development asset startup. A separate sidebar assertion used an unstable focus path through the rich editor. > - This PR gives browser tests fresh built assets and a direct keyboard path to the star button. It adds service-worker reload coverage under CPU throttling. > - A later browser failure exposed a real reload race: comments can change between the first query and the first live subscription. The UI now refreshes active queries when that subscription opens. > - Runtime fixtures now have separate registry state and better failure evidence. The readiness deadline remains unchanged. > - A later CI run exposed a wall-clock backoff assertion and three authorization cases sharing one test lifecycle. The PR anchors the assertion to transport time and separates the cases. ## Linked Issues or Issue Description Refs #13891. Related runtime ownership and cleanup work: #11791, #11389, #11278. Evidence: [original CI run, attempt 3](https://github.com/paperclipai/paperclip/actions/runs/35910335089/attempts/3). Attempt 3 passed both affected shards without code changes. This PR is separate from the wake-payload change, which has since merged. | Failure | Diagnosis and classification | | --- | --- | | Sidebar blank page and retry text missing after reload in CI | Both traces show a blank document before React startup. Vite connects, but the failing page makes no application API requests. The service worker forwards the unbundled development module graph. The retry response is already stored and its process-adapter run succeeded before reload. This places the failure in browser bootstrap, not wake payload or reply persistence. The exact reason the development module graph stopped is not established by the retained trace. The harness now serves built assets, and the new test covers startup and controlled reload at 4x CPU throttling. | | Star opacity remains zero on macOS | Reproduced on unchanged master `f55759942b`. The old test clicked the rich editor, focused the star, then used Tab and Shift+Tab. An instrumented baseline run captured the sequence: Tab moved from the star to the next sidebar link, then the editor bundle called `focus()` on its contenteditable before Shift+Tab. That key reached the composer, where it is a work-mode shortcut. This confirms a pending editor selection update stole focus; it was not the browser skipping sidebar buttons. The assertion did not prove the star retained focus. The test now moves the pointer away, focuses the preceding sidebar link, presses Tab once, and asserts both actual focus and opacity. This is test synchronization and keyboard traversal, not a demonstrated CSS defect. | | Runtime readiness exceeds 10 seconds | The old error only says `fetch failed`. There is no child startup output in that failure, so it cannot distinguish slow process startup from a refused or stalled probe. It passed unchanged locally and took 3.416 seconds in attempt 3. Resource contention is plausible but unproved. This PR does not claim a proven historical runtime root cause: it isolates fixture registry/log files, checks ports before spawn and after stop, checks live backends at publication, and preserves transport errors, probe count, elapsed time, and fixture startup timestamps for the next occurrence. | | Later CI: recovery reply disappears after reload | Product synchronization bug, distinct from the blank-page bootstrap failure. Reproduced locally with a trace: the comments request started at `21:56:07.443`, the server saved the reply at `.520`, and the first live subscription started at `.541`. The reload fetched a successful run and its complete log, but missed the comment event. The provider refreshed after reconnects only. It now refreshes active queries on the first connection too. A deterministic regression test fails before the fix. No browser assertion was changed. | | Later CI: Slack backoff and authorization tests | The 30-second backoff assertion required more than 25 seconds to remain when it read the saved action. CI spent 10.311 seconds in the test, exceeding that five-second allowance. A local six-second read delay reproduces the failure; the new transport-anchored lower and upper bounds pass the same fault injection. The neighbouring authorization test ran three independent fixtures in one test and timed out at 15 seconds. Each case now has its own fixture cleanup and the normal per-test deadline, so earlier cases do not remain active during later global worker sweeps. No specific production slow call was established. | | Self-hosted runner loses communication or shuts down | The original lost-communication failure has no assertion. On final-head [attempt 1](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/1), Build, Runner Vitest 1/2, and chat 2/3 ran on three separate fleet instances. All received a runner shutdown signal at `22:04:55 UTC`, within 35 milliseconds, then cancellation. Server shard 6/12 received the same shutdown signal one minute later. All four were Spot `m7i-flex.xlarge` instances in `us-east-1a`. No test assertion or build error preceded those stops. This is infrastructure interruption; the reason the fleet stopped the runners is not available in job logs. | ## What Changed - Build the browser fixture UI into the static server's preferred directory, `server/ui-dist`, and disable Vite middleware. Ignore these generated assets. Start the source CLI directly from the repository root, as required by the CLI invocation safety contract. - Keep all existing chat assertions. Add first takeover and three service-worker-controlled reloads under CPU throttling. Assert that the page uses built module assets and that opening it creates no chat task. - Use forward keyboard traversal from the Zeta link to its star. Check focus before checking the reveal style. - Give each runtime exposure test a temporary Paperclip home and restore environment state after process cleanup. - Check that reserved ports are free before spawn, serve the fixture response before exposure, and become free after stop. - Include the nested fetch error, probe count, and elapsed time in readiness failures. Add a unit test for this diagnostic contract. - Refresh active queries when the first live connection opens, covering events missed during initial page loading. Keep reconnect toast suppression unchanged. - Measure Slack retry timing from the transport attempt and recovery completion. This checks the full provider-requested delay without spending a small wall-clock allowance on unrelated processing. - Run each Slack authorization-revocation scenario as a separate test, with cleanup between cases. All assertions remain. - Document the browser fixture's build and serving mode. No configured assertion deadline, readiness deadline, retry count, or skip was added. ## Verification - Final-head [Linux CI, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/2): **green**. All 52 check runs passed; two conditional checks were skipped. The legacy Snyk status also passed. No pending or failed checks remain. - `pnpm -r typecheck` passed. - `pnpm build` passed again after the final UI fix. - Focused runtime suites: 38 passed, 3 existing platform skips. - CLI invocation safety suite: 39 passed. - Slack timing negative control: inserting a six-second delay before reading the saved action fails the old assertion. All four revised authorization/backoff cases pass with that same delay. The diagnostic delay is not committed. - Live update suites: 92 passed, including the new regression. UI typecheck and token gates passed. - Recovery browser negative control: the unchanged tests reproduced the missing reply (1 failed, 9 passed). After the first-connection fix, all six recovery paths passed twice (12 passed). Browser assertions and deadlines are unchanged. - Server typecheck passed after the Slack test adjustment. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-chat --shard-index 0 --shard-count 3`: 342 passed. The other 682 tests belong to the remaining shards; collection verified exact coverage. - Final direct-source CLI launch: all five chat session tests passed. - `PAPERCLIP_E2E_PORT=32993 PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/agent-chat-sessions.spec.ts --repeat-each=3`: 15 passed, including nine controlled reloads at 4x CPU throttling. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-server-without-chat --shard-index 10 --shard-count 12`: 55 files passed, 867 tests passed, 21 existing skips. An earlier run could not initialize PostgreSQL because this Mac exhausted its System V shared-memory slots. After reclaiming the orphaned segment from this task's stopped browser server, the full shard passed. - On final commit `1f1fafc08d`, [browser 4/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400837947) passed all 16 tests; [browser 8/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838155) passed all 23 tests; [server 11/12](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838328) passed 887 tests with one existing skip. The readiness lifecycle case took 1.842 seconds. Chat 1/3 also passed. All four jobs interrupted by runner shutdowns passed unchanged on their single rerun. - Greptile reviewed `1f1fafc08d`: **5/5**, with no unresolved comments. - Full local `pnpm test:run` was attempted and stopped after confirming failures outside this patch: a sibling `skills` directory shadows bundled Slack/AgentMail skills; macOS rejects rename of read-only skill-cache directories (`EACCES`, also reproduced in an isolated filesystem probe); and the host exhausts PostgreSQL System V shared-memory slots. Focused reruns confirmed these limits. The complete Linux CI run is the repository-wide verification; the full local run is not green. - Baseline evidence: the original sidebar test failed on unchanged master; a development-mode run at 4x CPU throttling passed six selected cases, so CPU pressure alone did not reproduce the CI bootstrap stall. ## Risks - Default browser tests now exercise the shipped static UI. They no longer implicitly cover Vite middleware or HMR; use the development server for those checks. - The new browser startup test uses Chromium CDP, matching the only configured browser project. - Opening a live subscription now causes one active-query refresh to close the initial event gap. This adds startup API reads but no recurring poll. - Runtime behavior and deadlines are unchanged except for error details. The historical readiness stall remains unconfirmed; a green rerun alone cannot establish its cause. - Local verification runs on macOS. The runtime lifecycle checks also passed on Linux CI. ## Model Used - OpenAI GPT-6 through Codex. The session identifies the model family as GPT-6; an exact served model ID and context window size are not exposed. Used reasoning, repository inspection, shell tools, code editing, and test execution. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #123` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted suites passed; full local host limits are listed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24429024e7 |
feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps gives agents governed access to external tools through stored credentials. > - Routines start work when an external service sends an event. > - Fireflies provides meeting transcripts and summaries through an official hosted MCP server. > - This PR adds that connection and accepts signed meeting events through the shared app webhook flow. > - Agents can review completed meetings with the same permissions and audit records as other work. ## Linked Issues or Issue Description **Problem or motivation** Operators need agents to read Fireflies meetings and start follow-up work when a summary is ready. The Apps catalog lacks Fireflies. The shared app webhook flow needs to accept its signed deliveries. **Proposed solution** Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer API key. Extend the existing Another app or script flow with signed webhook support. Verify the raw-body signature and pass the JSON payload as external data. Select Meeting Summarized in Fireflies. Deduplicate identical signed deliveries, including setup deliveries. **Alternatives considered** A separate REST connector would duplicate the governed MCP path. Polling, legacy V1 payloads, and automatic provider-side webhook registration are outside this change. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps and Scheduled Routines surfaces. It adds a provider to those systems. It does not introduce a second integration framework. **Additional context** A GitHub search found no existing Fireflies issues or PRs. Provider references and verification limits are in `doc/connections/FIREFLIES.md`. ## What Changed - Add the official Fireflies catalog definition, generated registry, provider evidence, and branded artwork. - Reuse Access → Connect, dynamic discovery, Permissions, vault storage, policy, and audit behavior. - Classify Fireflies sharing, movement, and access revocation as writes. - Preserve Off and Ask first restrictions during OAuth reauthorization and API-key replacement. New actions retain normal defaults. - Add `app_webhook` authentication to the shared Another app or script flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier `fireflies_hmac` triggers and revision snapshots for compatibility. Existing text columns need no migration. - Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact request body. Preserve generic event payloads and deduplicate identical signed requests. - Keep the routine wizard generic. Show one webhook URL and secret in Another app or script. Keep all new app webhook event names provider-neutral. Keep provider setup instructions in the connector documentation. - Pass generic webhook JSON to the task in an explicit external-data block, capped at 16,384 characters. Keep strict meeting validation for existing legacy Fireflies triggers. ## Verification - Feature implementation commit `0882dc8a1`: all 54 CI checks passed; two conditional Storybook checks skipped. This includes full tests, typecheck, build, browser E2E, canary dry run, and security checks. Greptile rated this commit 5/5; all review threads are resolved. - Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed on the final code. Targeted connector, gateway, webhook, revision, and UI suites passed during implementation. After the provider-neutral follow-up, all 84 app-webhook and routine-service tests passed; the final payload-to-task assertion also passed in the 72-test routine suite and a clean-config rerun. - The long local `pnpm test:run` invocation started before the final edits and was stopped after the final-commit CI suites passed. It reported one generic webhook test failure while those files were changing; that test and the entire routine suite passed on the final source, including a clean-config reproduction. The interrupted local run is not counted as a full-suite pass. - In the embedded browser, completed official OAuth consent and discovered 20 live actions. Real meeting listing, transcript retrieval, and summary/action-item retrieval succeeded as the selected agent. Turning a live read Off blocked its test; catalog refresh preserved the restriction. - Embedded-browser Another app or script setup, back/save/resume, narrow layout, and a signed synthetic Fireflies delivery succeeded. The UI reported authentication passed without creating a task. Fixtures cover signature tampering, malformed requests, ordinary app event names, duplicate/setup deliveries, rotation, revisions, pause/archive, and company isolation. - Existing MCP browser suite: 8 passed and 2 provider-dependent cases skipped. Branding checks passed; connector artwork and webhook setup were checked at desktop/mobile widths and in light/dark modes. - An unauthenticated POST to a correctly formatted public webhook URL reached the staging tenant verifier through the existing Cloud gateway. - A real Fireflies webhook delivery remains unverified. A staging callback is available for the operator walkthrough. Live API-key authorization, credential expiry, and a new meeting's summary completion were not tested against the provider. Fixtures cover these protocol and lifecycle paths where applicable. - Storybook follow-up `c54174faa`: 27 production-component stories cover every UI change, with a source-to-story map in the connector documentation. Static Storybook build, UI typecheck, token gates, and Playwright checks for all stories and the mobile footer pass. All PR checks passed for this Storybook follow-up; Greptile reviewed `c54174faa` at 5/5. ## Risks - Fireflies may change its hosted MCP tools or OAuth behavior. Tool discovery stays dynamic. Experimental search/fetch tools are not required. - Public webhook setup requires HTTPS and a separate signing secret. Fireflies normally emits events for meetings owned by the configuring account. - Reauthorization touches shared MCP permission code. Regression tests cover existing restrictions, new actions, connection removal, and other gateway callers. - Webhook receipt grants no tool access. The routine agent still needs an authorized Fireflies connection. - New generic triggers rely on provider event subscriptions. Without a sender-supplied idempotency key, changed request bytes count as a new event. Existing legacy Fireflies triggers retain summary-only filtering and per-meeting deduplication. ## Model Used OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing, code execution, and embedded-browser testing. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b648d8cdda |
fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path > - Paperclip manages AI agents and their provider connections. > - Product E2E checks real tasks through the browser, server, and runner. > - Grok qualification needs separate API-key and subscription evidence. > - Product subscription tests and direct Grok protocol evals need explicit credential delivery. > - This change supplies each credential only to its selected profile and prepares the pinned binary. > - Maintainer authorization and protected-environment gates remain required. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13882. The Grok feature branch has a manual subscription qualification profile. The trusted master workflow must admit its selected credential and prepare the same verified binary and artifact verifier as the API profile. Direct protocol evals also need the selected xAI key and pinned Grok binary. These prerequisites do not register or schedule the new profiles on master. ## What Changed - Deliver `GROK_AUTH_JSON` from the protected paid environment only when the selected profile requests that credential. - Install the checksum-verified Grok binary for the local subscription profile. - Prepare the pinned artifact verifier for the manual subscription suite. - Extend workflow security assertions to cover the new credential and profile. - Add the ACPX Grok credential mapping to the trusted-master catalog, then deliver only the selected `XAI_API_KEY` to direct protocol cells and install the target’s checksum-verified Grok binary before packaging. - Allow a direct-protocol concurrency override from two cases up to the existing configured ceiling; it can only lower concurrency. - Document the Grok protocol workflow and its API-only credential boundary. - Render missing LLM usage and cost as Unavailable, and label partial observations with coverage. Preserve raw records, grades, and actual zero costs. - Preserve measured campaign source metadata during report regeneration instead of inheriting the renderer checkout or CI event; skip empty legacy source records when recovering older provenance. ## Verification - Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54 reported checks successful, two intentional skips, Greptile 5/5, and zero unresolved review threads. [CI run](https://github.com/paperclipai/paperclip/actions/runs/35890978288). - After merging current master, all 17 workflow security/image tests and 23 catalog/workflow policy tests passed. The trusted catalog also generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per shard. The new policy tests execute the concurrency guard against valid, out-of-range, and malformed values. - The Grok branch separately passed 450 Product harness unit tests, including private company credential staging, cleanup, and token-fragment redaction. - The fresh-login native subscription smoke passed three repetitions of tool execution, session resume, restrictive permissions, and cleanup. These are setup evidence; full subscription Product qualification remains pending. - All 72 focused report/billing/history/catalog tests and the Product harness typecheck passed for the report-display change. The initial sandbox run could not open the tsx IPC socket; the permitted rerun passed. A zero-provider-call replay of the actual 16-cell Grok campaign preserved all result records, grades, timing, and source provenance while correcting missing usage labels. - Review the thirteen-file diff. Provider credentials still enter only the selected paid-test step; default-branch, numeric-actor, and environment restrictions are unchanged. ## Risks This admits a refreshable subscription credential to explicitly selected trusted tests. Store it only in `runner-e2e-paid`, use a test login, and remove it after qualification. Unselected profiles receive an empty value. Pull requests cannot trigger the paid workflow. This PR changes no fleet admission or actor allowlist. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.11 |
||
|
|
4b8ec588f3 |
Stop duplicating wake context in adapter environments (#13891)
## Thinking Path
> - Paperclip manages agent work and preserves task context.
> - Built-in adapters already include wake context in the agent prompt.
> - They also copy the full wake JSON into a process environment
variable.
> - A large environment entry can prevent the agent from starting with
`spawn E2BIG`.
> - This change removes the duplicate environment entry and uses the
existing prompt delivery.
> - The agent keeps its context without extra file transport or new
history limits.
## Linked Issues or Issue Description
Refs #13144, #13860, #13872, #13793.
Large wake payloads can exceed the operating system limit for one
environment entry. The launch-envelope fix in #13793 handles the outer
transport but leaves that child environment entry intact.
Credit to @nickyleach for the prompt-only approach in #13144. This PR
applies that part on current master. It does not include that PR's
30-item history limits or recovery-history endpoint. Those behavior
changes can be reviewed separately from the process launch fix.
## What Changed
- Stop exporting `PAPERCLIP_WAKE_PAYLOAD_JSON` in the shared ACP engine
and all ten built-in adapter writers.
- Ignore configured values of the retired variable so saved adapter
settings cannot restore the oversized entry. Also drop inherited copies
in Hermes, which builds its environment directly.
- Keep scalar runtime variables, existing prompt rendering, continuation
history, resume deltas, gateway bodies, and Hermes JSON template
variables.
- Document the prompt delivery contract and the migration for custom
instructions that read the retired variable.
- Test large local and sandbox child-process launches, fresh and resumed
ACP turns, SDK delivery, and configured-variable filtering.
## Verification
- `pnpm -r typecheck` passed.
- Focused adapter utility, ACP, Codex child-process, and Cursor Cloud
suites: 340 tests passed.
- Hermes execution and prompt tests: 18 tests passed using its package
Vitest configuration.
- The child-process tests deliver over 128 KB of context through stdin
and check the complete text. The ACP test retains 50 complete messages
and 50 completed actions, then checks the resumed delta.
- `pnpm build` passed.
- `pnpm test:run` was attempted, then stopped after it reproduced ten
macOS runtime-skill-cache permission failures (also reproduced on
unchanged master) and one HTTPS backfill test failure. The HTTPS test
passed when rerun unchanged on this branch and master. The complete
local suite was not completed; Linux CI provides the full-suite gate.
- CI is green on
canary/v2026.923.0-canary.10
|
||
|
|
f55759942b |
fix(runner): accept compatible Codex releases from 0.149.0 (#13853)
Keep the minimum fixed at 0.149.0 until a deliberate maintainer change. Accept stable versions below 0.157.0 separately from the reproducible install pin. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.9 |
||
|
|
b41ccf097f |
fix(apps): configure MCP aggregators from inline task cards (#13879)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agents request connections through cards in task threads. > - MCP aggregators need provider-specific URLs and authentication. > - The task dialog used the generic setup form and omitted these fields. > - This pull request uses the same provider setup controller in tasks and Apps. > - Users can configure a connection without leaving the task. ## Linked Issues or Issue Description Related: #13755, #13855. **What happened?** An inline Executor request opened a very wide dialog with an empty credential step. Connect failed because the MCP URL was missing. The other aggregator cards also bypassed their provider setup. **Expected behavior** Each card shows its provider instructions, URL field, and authentication options in a bounded dialog. Completing setup grants access only to the requesting agent. **Steps to reproduce** 1. Enable experimental MCP aggregators. 2. Have an agent request Zapier, Arcade, Composio, or Executor from a task. 3. Open the card and continue past Access. ## What Changed - Route page and task setup through the same provider controller. - Bound the task dialog width and preserve the requesting agent's access scope. - Support existing accounts, saved drafts, URL/token setup, and task-bound OAuth. - Keep a sign-in link available when the browser cannot open a popup. Verify completion through the existing durable callback path. - Add inline Access, configuration, and narrow Storybooks for all four providers. - Document the shared setup requirement and correct Executor's URL instructions. ## Verification - Focused Vitest selection: 24 passed. Covers all four inline forms, requester access, existing accounts, saved drafts, OAuth retry, callback validation, popup cleanup, and generic reconnect endpoint preservation. - UI typecheck, UI build, design-token gates, and Storybook build passed. - Live local browser: new Executor, Arcade, and Composio connections completed provider consent from task cards. Each appeared Connected with the requester selected. - Real Test calls returned Executor output `4`, an Arcade public GitHub star count, and Composio tool-discovery results. An ungranted agent was denied access. Real Paperclip process-agent runs discovered each provider catalog with only the requester’s connection installed. - Deployed implementation commit `5d722e89c` to the isolated staging tenant. The original failing Executor card now completes, discovers seven actions, limits access to its requesting agent, and resumes that agent. Its continuation completed real Executor calls and the provider resume flow, then returned an upstream Airtable authorization link. A real staging Test call returned `4` in 1.5 seconds. Later PR commits add regression coverage and popup-unmount cleanup. - Zapier fresh-token browser test remains pending a provider clipboard handoff. Its URL and token flow passes focused tests. - Latest-head CI (`d14f4c73c`): 53 checks passed; optional Storybook deployment and visual regression jobs skipped. Greptile 5/5, both review threads resolved. Canceled runners and unrelated chat timeouts passed the single retry on unchanged code. - The full local test suite was not run, as requested by the maintainer. CI runs the repository gates. ## Risks - OAuth popup behavior differs by browser. The explicit sign-in link and durable server completion checks provide recovery. - Saved task drafts store only a connection ID in browser storage. Credentials remain in the existing server vault. - No database or server protocol changes. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code execution, and browser tools. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.8 |
||
|
|
60c7c9cd1a |
fix(runner-e2e): pass verified lock digest to Daytona image build (#13876)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Product E2E campaigns test the native runner in local and Daytona environments. > - Each campaign resolves one target lockfile and verifies its downloaded artifact. > - The Daytona image job did not pass that artifact digest to Docker. > - Docker used an older default digest and stopped before any selected task ran. > - This pull request passes and validates the campaign digest at the image build boundary. > - The image keeps its checksum check and frozen package installation. ## Linked Issues or Issue Description **What happened?** The merged-master [qualification campaign](https://github.com/paperclipai/paperclip/actions/runs/35863582409) stopped in the Daytona image build. The resolved target lock digest was `e0c928a494f90ddad3c00791e83f09315ee8c82df2a0418a809dcf93649a8ab3`. Docker used its default digest, `57b298aceebc48bb94ea0593347348256475da7b2fddb77025d9e57cc8759420`. The checksum check rejected the mismatch. All 12 selected model cells were skipped. This follows the image provenance work in #13814. **Expected behavior** The image build must check the same lockfile artifact that the campaign restored and verified. A changed lockfile must still fail the checksum check. **Steps to reproduce** 1. Start a Product E2E campaign with a Daytona cell on master `7944ed3d976d1a7cc26a2d0cee51f227f3542084`. 2. Resolve a target lockfile whose digest differs from the Dockerfile default. 3. Observe the provider-pack image stage reject the lockfile before model execution. **Paperclip version or commit** `7944ed3d976d1a7cc26a2d0cee51f227f3542084`. **Deployment mode** GitHub Actions Product E2E campaign with a Daytona image build. **Install method** Built from source with the campaign lockfile artifact. **Agent adapter(s) involved** Native Codex and ACPX Claude cells were selected. No model cell ran in this failed campaign. **Database mode** Not involved. The failure occurs during image creation. **Access context** The authorized default-branch paid workflow. The build receives no provider credentials. ## What Changed - Read the image checksum from the existing target-lock job output. - Require a 64-character lowercase hexadecimal digest before image inspection or build. - Pass the digest as the existing Docker build argument. - Add regression checks and document the campaign checksum handoff. ## Verification - The Daytona image regression fails with the original workflow and passes with the fix. - All six Daytona image contract tests pass. - All 450 Product E2E unit tests pass. - Product E2E typecheck passes. - Actionlint passes for the changed workflow. - A context-shaped resolution probe preserves the downloaded lockfile bytes and digest. - All latest-head CI checks passed on `67d41fd9439b2a9a809ddb05765f8617585072c5` ([run](https://github.com/paperclipai/paperclip/actions/runs/35865739359)). - Greptile gave 5/5 on this head; its test-scoping comment is addressed and resolved. - A hosted Daytona image rebuild and the three remote qualification cells remain pending after merge. ## Risks The campaign digest comes from the existing trusted target-lock job. The restored artifact checks, Docker checksum check, frozen install, content identity, image signing, and verification remain in place. The standalone Docker default remains available. This change does not alter task behavior, prompts, credentials, or dependency versions. ## Model Used OpenAI `gpt-6-astra` through Codex performed diagnosis and review with code execution tools. OpenAI `gpt-5.6-luna` assisted with investigation, implementation, and verification. Context window limits are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.7 |
||
|
|
4721f55803 |
fix(apps): recover MCP OAuth setup after consent errors (#13855)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external tools. > - MCP aggregator setup can return from provider consent to a saved draft. > - The branded setup did not explain failed or cancelled authorization. > - It also offered identity changes that the server does not apply when a saved connection resumes. > - This pull request explains OAuth return outcomes and keeps the displayed identity consistent with the saved policy. > - Users can understand the outcome and retry the same connection. ## Linked Issues or Issue Description Related: #13755 introduced the MCP aggregators. #13758 retired the legacy Composio broker. #13584 proposes changes to the provider handoff window; this fix retains the current handoff behavior. **What happened?** Cancelling Composio consent returned to the setup form with no explanation. The same controller ignored failed OAuth callback outcomes. Returning to Access on a saved draft also offered personal/shared choices, although the server retains the saved identity. This could make the OAuth request disagree with that identity. **Expected behavior** Explain cancellation or failure, preserve the draft, and offer Try again. Display the retained credential identity and start OAuth with the policy returned by the server. **Steps to reproduce** 1. Enable MCP aggregators and start a Composio connection. 2. Continue to provider consent and cancel it. 3. Observe the return screen. Before this change, it showed the form without cancellation feedback. 4. Go back to Access. Before this change, the form offered personal/shared choices even though resume retains the original identity. **Paperclip version or commit** Observed before the fix on `8c6cc7dccf91523e0720bd86f95487e66b4b0e63`. The original report of successful consent leaving setup unfinished did not reproduce. This PR addresses the recovery defects observed during that investigation. **Deployment mode** Isolated local development instance and authenticated staging deployment, with real Composio consent and provider calls. ## What Changed - Read the OAuth callback outcome in branded MCP setup and show cancellation or failure feedback. - Retry the same saved draft without displaying untrusted callback error text. - Keep saved personal/shared identity fixed in Access and select OAuth identity from the returned credential policy. - Add six focused regression cases and authorization-failure Storybook states for Arcade, Composio, and Executor. - Document return-screen recovery and retained identity in the connector playbook. ## Verification - `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx -t 'OAuth return' --maxWorkers=1`: 6 passed, 143 skipped, including after integration with current master. - `pnpm --filter @paperclipai/ui exec tsc --noEmit`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed for the implementation commit. - Browser recovery: cancelled real Composio consent, observed the new feedback, returned to Access, retried the same personal draft, completed consent, and ran a real tool call. - Staging on implementation commit `84bb40aa70662e0c8955bf692c5714661b4bea93`: fresh shared and personal connections each completed on the first consent attempt and loaded 11 tools. Real discovery, execution, and schema calls succeeded. Both connections stayed Connected after reload. A real agent used the shared connection through the Paperclip gateway and returned the public repository documentation hierarchy with one success and zero errors. - All current-head CI gates passed on `662f84a67e867a52a2e5526026adbed00f6b59bf`. The Cursor execution and agent-chat browser shards each had an initial timeout; both passed on one targeted rerun without code changes. Greptile reviewed this exact head at 5/5, with no open review threads. - No full local suite was run, as requested. CI provides the broader checks. The PR adds a master merge and documentation after the live-tested implementation commit. ## Risks - This shared setup controller also serves Arcade and Executor. Their callback rendering and retry behavior have focused test coverage; this investigation used Composio for live provider testing. - Saved identity remains fixed during resume. A different identity requires a new connection, consistent with server behavior. - No database, protocol, credential storage, or gateway policy changes. ## Model Used OpenAI GPT-6 via Codex, with code editing, shell tools, and browser testing. The exact runtime model ID and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.6 |
||
|
|
7944ed3d97 |
fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Native runner agents need governed tools, durable runtime state, and useful execution evidence. > - A first activity trace waited 53.467 seconds even though tool activity took 6.274 seconds; provider input arrived before the server API call executed. > - Native agents also need a safe way to hire teammates without asking the model to rebuild runtime configuration. > - This pull request separates the observed ACP input-stream window from the actual server `tool.execute` span and adds a server-owned native hire contract. > - The benefit is clearer latency evidence and safer native teammates with existing approval, auth, and company boundaries preserved. ## Linked Issues or Issue Description Related Daytona provenance work is in [#13814](https://github.com/paperclipai/paperclip/pull/13814). No duplicate public PR was found for this combined timing and native-hire change. **What existing behavior does this improve?** Native runner agents can use governed tools and request hires. The server did not expose a safe native hire operation that reused the caller's validated runtime settings. First-activity traces also mixed provider input timing with server tool execution timing. **Current behavior** A native hire must construct a separate runner configuration. Full configuration copying could expose paths, instructions, secrets, or sessions. Timing evidence could make a provider or MCP identity join appear proven when the trace did not contain that join. **Proposed behavior** The native `hire_agent` operation accepts identity and persona inputs. The server sends `adapterType: "paperclip_runner"` with `inheritRuntimeFrom: "caller"`, then copies only validated provider, model, permission, lifecycle, and bounded execution settings. It inherits and validates the default environment, derives the managed AI binding through existing normalization, preserves approval and permissions, and creates fresh child instructions. Caller secrets, paths, prompts, and sessions are excluded. Provider events now include the optional boolean `inputUpdated`, with Rust forwarding support. Timing evidence separately records the ACP input-stream window and the actual server `tool.execute` activity. It does not claim a provider or MCP join without matching evidence. **Reason and benefit** Native agents can hire teammates that start with the caller's approved execution policy. Operators retain company boundaries, auth rules, approval gates, and requalification. Reviewers can distinguish provider streaming time from server API execution time when diagnosing first-activity delays. **Breaking changes** None for existing hires or tool calls. `inheritRuntimeFrom` is optional and only applies to same-company native agent callers. Conflicting explicit runtime settings are rejected. The provider event field is optional for existing producers. ## What Changed - Added the native `hire_agent` protocol action, catalog entry, API contract, and runner authority checks. - Added `inheritRuntimeFrom: "caller"` validation and a closed native runtime inheritance allowlist. - Preserved managed AI binding normalization, default-environment validation, approval snapshots, permissions, requalification, and fresh child instructions. - Added provider `inputUpdated` schema support and Rust forwarding. - Added first-activity and server tool timing evidence with conservative identity-join handling. - Added route, authority, provider-event, sidecar, API, catalog, and Rust-focused tests. - Kept private Honeycomb links, raw traces, and local result paths out of this description. ## Verification Focused checks passed: - 458 timing/session checks. - 61 native hire inheritance checks. - 20 hire authority checks. - 1,741 API checks. - 106 catalog checks. - 54 provider sidecar checks. - 12 Rust provider checks. Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2). R1 stopped at missing Docker image setup. The final trace is available at https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces. Latest-head CI passed all required build, typecheck, Rust, static, Vitest, serialized-server, workspace, chat, and E2E jobs. The focused local checks listed above passed; the broad local suite was not run before the live evaluation, while CI provides the full repository verification. ## Risks - Timing fields describe separate observed windows. They do not prove a provider or MCP owner without a valid trace join. - The inheritance allowlist must stay synchronized with native runner configuration fields. - Approval snapshots include resolved safe inherited settings and should be reviewed when native configuration fields change. - The focused local suite is narrower than the full repository suite; latest-head CI covers the broader repository checks. > Roadmap review: `ROADMAP.md` places this work within Paperclip's bring-your-own-agent direction. It extends existing native runner hiring and observability behavior. ## Model Used OpenAI GPT-6 (exact serving model ID is not exposed), with extended reasoning and repository tool use; GPT-5.6 Luna assisted with focused implementation and verification work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.5 nightly/v2026.923.0-nightly.0 |
||
|
|
64faf0ae90 |
fix(ui): leave the archived company's settings after archiving it (#13846)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A person can archive a company from that company's settings page. > - After the archive, the page stays on the archived company. The only visible change is the button text "Already archived". > - It is unclear that anything happened, and the user has no next step on screen. > - This pull request navigates away after a successful archive: to another active company, to the Cloud portfolio when no active company remains on a managed instance, or to the companies list when self-hosted. > - The benefit is a clear outcome: the user sees where they are now and a toast that names what happened. ## Linked Issues or Issue Description No public issue exists. Description follows the enhancement template: **What existing behavior does this improve?** The archive action on the company settings page. The mutation works, but the view stays on the archived company and gives no feedback beyond a disabled button. **Subsystem affected** UI: company settings page, company selection helpers, cloud links. **Current behavior** Archive succeeds. The button changes to "Already archived". The user stays on the archived company's settings. On a single-company instance nothing else changes. **Proposed behavior** After a successful archive, the app departs: another active company's dashboard with a toast; the Cloud portfolio's manage view when no active company remains on a managed instance; the companies list when self-hosted with nothing active, because that list shows the archived state and owns the unarchive action. **Reason and benefit** Staying on the archived company reads as "nothing happened". Leaving to a live surface makes the outcome clear and gives the user their next step. ## What Changed - `ui/src/lib/company-selection.ts`: new pure `resolveCompanyArchiveDeparture` helper. Another active company wins; the just-archived company is excluded by id, so a stale cached status cannot select it. Cloud portfolio is the fallback on managed instances; the companies list is the final fallback. - `ui/src/lib/cloudLinks.ts`: new `cloudPortfolioManageUrl` (`/orgs?manage=1`). The manage view matters: the plain launchpad auto-forwards a solo user back into their one openable stack — the page this navigation is escaping. - `ui/src/pages/CompanySettings.tsx`: the archive mutation now departs on success — toast + `navigate` for in-app destinations, a top-level navigation for the Cloud portfolio — instead of only setting the selected company id. - `ui/src/pages/CompanySettingsRenameHint.test.tsx`: render harness wraps the page in a `MemoryRouter` for the new router dependency. ## Verification - `pnpm exec vitest run src/lib/company-selection.test.ts src/lib/cloudLinks.test.ts src/pages/CompanySettingsRenameHint.test.tsx src/pages/CompanySettings.test.tsx` — 22 tests pass. - New unit tests cover: active sibling wins (also on cloud), the just-archived company is never the destination even with a stale cached status, cloud portfolio fallback, self-hosted companies-list fallback, and the portfolio URL helper (null base included). - `pnpm exec tsc --noEmit` in `ui/` is clean after the plugin SDK build. ## Risks - Low risk. The archive API call is unchanged; only post-success navigation is new. The toast uses the optional actions hook, so surfaces without a ToastProvider stay safe. - On Cloud, the portfolio navigation is a full top-level navigation, so the query-cache invalidation for the departed document is skipped intentionally. ## Model Used - Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended thinking and tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.923.0-canary.4 |
||
|
|
e4237c45f3 |
fix(ui): prevent organization title flicker during plugin loading (#13854)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The sidebar identifies the current organization. > - An optional plugin can replace this navigation surface. > - The built-in title appears before discovery and module loading finish. > - This pull request reserves the trigger until its owner is known. > - The organization name appears once, while failures retain built-in navigation. ## Linked Issues or Issue Description Refs #13832. Searched related pull requests and issues; no duplicate fix found. **What happened?** The organization switcher renders a provisional built-in title before an installed replacement loads. Unrelated plugin imports can also affect its loading state. **Expected behavior** Reserve the trigger with a neutral placeholder, then show the resolved navigation surface. Keep the built-in menu on failed or absent contributions. **Steps to reproduce** Install an organization-switcher contribution. Delay session, company, contribution, and module responses. Reload the page and watch the trigger through each stage. ## What Changed - Reserve the trigger through account, company selection, slot discovery, and module loading. - Distinguish failed session lookup from pending lookup so errors retain usable navigation. - Load and await only contributions matching the requested slots. Observe completion of imports started by another consumer. - Document loading behavior and add regression coverage for loading, failures, unrelated modules, and identity transitions. ## Verification - `pnpm -r typecheck` passed, including Rust checks. - `pnpm build` passed. - All 629 UI test files passed: 6,593 tests. The 42 focused UI/API/plugin tests also passed. - `pnpm check:token-gates` and `git diff --check` passed. - `pnpm test:run` was also attempted. The broad local server run was stopped after recording skill-cache/channel fixture failures outside this diff (for example, runtime skill source status `missing` instead of `available`). The original cause is not established. All latest-head Linux CI gates pass; the complete UI suite and affected local checks pass. - Desktop (1440px) and mobile (390px) Chromium checks passed with real host components, dynamic module loading, and the built Account bundle. Delayed fixture responses produced exactly two title states: empty placeholder, then the resolved name. A slow refresh preserved the title and trigger dimensions; absent/failed plugin fallback and Escape dismissal passed, with zero uncaught browser errors. This is browser component integration, not a live signed-in tenant test. ## Risks A cold load displays a neutral placeholder until discovery completes. Absent, ambiguous, failed, and invalid contributions still use the built-in menu. No migrations or authorization changes. Scoped module loading changes when an unrelated contribution is imported; each surface loads its own matching modules. ## Model Used OpenAI Codex, GPT-6, with reasoning, code execution, and browser verification. The exact deployment ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
106d89314f |
fix(evals): keep raw diagnostics out of public history (#13850)
## Thinking Path > - Paperclip manages AI agents and their work. > - Product E2E campaigns retain evidence from live provider runs. > - The public publisher copied per-attempt files based on their extension. > - Credential redaction does not remove hidden reasoning or provider session IDs. > - This pull request keeps those raw diagnostics out of public evidence bundles. > - Public reports retain normalized grades and declared fixture screenshots. ## Linked Issues or Issue Description **What happened?** The Product history publisher accepted all JSON, log, Markdown, and text files under an attempt directory. API snapshots and process logs can contain provider reasoning even after credential redaction. **Expected behavior** Keep raw attempt diagnostics in retained Actions artifacts. Publish normalized results and declared fixture screenshots. Preserve the original grades and attempt evidence. **Steps to reproduce** Place a credential-redacted API snapshot with a reasoning event under an attempt's snapshots directory. The previous public path check admitted it because it ended in .json. ## What Changed - Remove extension-based admission for per-attempt text files in both S3 and Pages staging. - Require PNG paths to appear in the existing screenshot declaration allowlist. - Preserve root normalized results, grading, billing, provenance, and reviewed screenshot behavior. - Add seven denial cases for raw diagnostics, renamed files, malformed JSON, and false screenshot declarations. - Update report copy and the publication security contract. ## Verification - Focused history and report tests: 47 passed. - Replayed the filter against a copy of retained Grok evidence: six diagnostic files removed, one declared screenshot retained, zero provider calls. Original artifacts remain unchanged. - git diff --check passed. - Product typecheck in the reused local dependency tree reports an unrelated plugin type mismatch for organizationSwitcher in server/src/services/plugin-capability-validator.ts. CI will verify a fresh install. - No local browser or Docker execution. ## Risks Public reports no longer link raw per-attempt logs or API snapshots. Those files remain in the original Actions artifact. This change affects future publication only; it does not withdraw or rewrite existing immutable public campaigns. Normalized result content and marked screenshot review retain their existing boundaries. No workflow authorization, provider credentials, application behavior, or database changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool execution. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.3 |
||
|
|
ac854a9b59 |
ci: build Daytona eval images on the authorized fleet (#13847)
## Thinking Path > - Paperclip manages AI agents and their work. > - Product E2E campaigns test the browser, server, and runner together. > - The trusted workflow selects an authorized execution fleet. > - Daytona image builds still use a fixed GitHub-hosted runner. > - This pull request applies the existing fleet selection to image builds. > - Campaign builds and tests then use the configured EC2 fleet. ## Linked Issues or Issue Description Refs: https://github.com/paperclipai/paperclip/pull/13845 **What existing behavior does this improve?** The location of Daytona Product E2E image builds. **Current behavior** The image job uses ubuntu-latest even when RUNNER_E2E_AWS_ENABLED selects EC2 for the rest of the campaign. **Proposed behavior** Use the existing authorized runner output for the image job. Preserve the GitHub-hosted fallback when the EC2 switch is off. ## What Changed - Route Daytona image builds through the existing authorized runner selection. - Assert that the image job depends on authorization and receives no provider credentials. - Document the build routing and credential boundary. ## Verification - Product E2E workflow security: 11 tests passed. - git diff --check passed. - Reviewed actor checks, workflow triggers, target commit selection, signing, and package permissions. They are unchanged. - Required CI checks pass on the latest head. The first CI attempt had failures in unchanged chat tests; one diagnostic rerun passed. Both attempts remain in Actions. - Greptile reviewed the latest head at 5/5 with no inline findings. - Live EC2 image verification follows after this trusted workflow change is merged. No local Docker execution. ## Risks EC2 image builds depend on the fleet having working Docker and sufficient disk space. Existing authorization, image digest checks, signing, and the GitHub-hosted fallback remain in place. No provider secrets are added to the image job. No application or database changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool execution. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
950ccb8eef |
ci: enable Grok qualification in the trusted paid workflow (#13845)
## Thinking Path > - Paperclip manages AI agents and their work. > - Product E2E tests verify tasks through the browser, server, and runner. > - These paid tests use a trusted workflow from master and an isolated target commit. > - The Grok target branch selects XAI_API_KEY, but the trusted workflow does not deliver that credential. > - Its artifact test also needs the pinned Python verifier before execution. > - This pull request adds both bindings inside the existing paid boundary. > - The tests can then run on the configured EC2 fleet without laptop Docker. ## Linked Issues or Issue Description Related evaluation infrastructure: https://github.com/paperclipai/paperclip/pull/11297. No duplicate Grok paid-workflow change was found. **What existing behavior does this improve?** Branch-targeted Grok Product E2E qualification in the existing paid workflow. **Current behavior** Grok cells cannot receive their selected API credential. The Grok build-revise case also misses the artifact-verifier setup step. **Proposed behavior** Deliver XAI_API_KEY only when the selected matrix credential is XAI_API_KEY. Prepare the existing pinned verifier for the Grok qualification suite. **Reason and benefit** Run the controller, browser, runner and artifact checks on the EC2 fleet. Preserve default-branch workflow authorization and protected environment secret access. ## What Changed - Bind the selected XAI credential only in the paid test step. - Install the checksum-verified Grok binary for local cells before provider access. - Include Grok qualification in the existing pinned artifact-verifier preparation. - Add an optional max_parallel input that can only lower the configured campaign concurrency. Use 1 for the Grok test key. - Extend security assertions and document setup. ## Verification - Ran the Product E2E workflow-security tests: 11 passed. - Checked the diff for whitespace errors. - Reviewed credential selection, setup ordering, numeric actor gates, target commit pinning, and trusted report checkout. - Live Grok execution follows after this workflow is available on master. This PR does not claim completed Grok qualification. ## Risks The paid test step can use the selected XAI credential and incur provider charges. The credential remains in runner-e2e-paid and is absent from setup, build, and reporting jobs. The default-branch gate and existing environment restrictions remain in place. No database migration or product behavior changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool execution. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.2 |
||
|
|
8c6cc7dccf |
feat(cloud): sync primary-company archive state with the Cloud control plane (#13837)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud runs managed instances and controls their lifecycle
from a control plane.
> - Inside the app, a person can archive a company. On a managed
instance, the Cloud-pinned primary company is the whole organization.
> - Today that archive stays local. The control plane does not learn
about it, so it keeps the instance running and shows the organization as
live.
> - This pull request notifies the control plane when the primary
company crosses the archived boundary, and adds two control endpoints so
the control plane can verify the state and undo the archive during a
restore.
> - The benefit is that an archived organization stops running and shows
as archived, and a restore brings the company back without manual steps.
## Linked Issues or Issue Description
No public issue exists. Description follows the enhancement template:
**What existing behavior does this improve?**
Company archive on Cloud-managed instances. Archiving the primary
company pauses agents and cancels runs, but the hosting control plane
never learns about it.
**Subsystem affected**
Server: company service, cloud-control middleware, instance routes.
**Current behavior**
A person archives the primary company. The instance keeps running. The
Cloud portfolio still shows the organization as live. Unarchive after a
Cloud restore requires manual steps inside the product.
**Proposed behavior**
The company service rings a Cloud lifecycle doorbell when the primary
company is archived or unarchived. The control plane verifies the state
through `GET /api/instance/lifecycle` before it acts. During a restore,
the control plane calls `POST /api/instance/lifecycle/unarchive-primary`
to bring the company back. Self-hosted instances are not affected.
**Reason and benefit**
An archived organization should not keep running, and its hosting
console should show it as archived. The verified read-back keeps the
doorbell a hint: a forged or duplicated ring cannot change state.
## What Changed
- New `services/cloud-lifecycle-sync.ts`: fire-and-forget doorbell `POST
{cloudOrigin}/v1/tenant/lifecycle-changed` with bounded retries. It runs
only on cloud-managed instances, and only for the derived primary
company id. It never blocks or fails the company mutation.
- `services/companies.ts`: ring the doorbell after a committed archive
or unarchive transition, in both the `update()` status-patch path and
`archive()`.
- New `GET /api/instance/lifecycle`: reports the primary company id, its
status, and how many other companies are not archived. Bound to a
`lifecycle:read` Cloud control assertion; board members can also read
it.
- New `POST /api/instance/lifecycle/unarchive-primary`: idempotent
unarchive of the primary company as a system actor. Bound to
`lifecycle:unarchive-primary`; requires instance admin otherwise.
- `middleware/cloud-control.ts`: a closed endpoint→method→action table
replaces the single hardcoded endpoint. Each assertion still authorizes
exactly one action on one endpoint.
- `services/cloud-instance.ts`: the primary-company id derivation moves
here as the single definition; `middleware/auth.ts` delegates to it. The
derivation itself is unchanged.
## Verification
- `pnpm exec tsc --noEmit` in `server/` is clean.
- `pnpm exec vitest run src/__tests__/cloud-lifecycle-sync.test.ts
src/__tests__/cloud-control-task-drain.test.ts
src/__tests__/instance-settings-routes.test.ts
src/__tests__/companies-service.test.ts
src/__tests__/cloud-tenant-company-provisioning.test.ts` — all pass.
- New tests cover: doorbell env-gating, retry bounds, no-throw contract;
the read-back status and sibling count; idempotent unarchive and the
admin gate; non-cloud 404s; control-assertion action binding for the new
endpoints, including cross-action and wrong-method rejection; and a
companies-service test that proves the doorbell rings exactly on
archived-boundary transitions.
## Risks
- Self-hosted instances see no behavior change: without a Cloud signal
the doorbell is a no-op and the endpoints answer 404.
- The doorbell is advisory by design. The control plane verifies through
the read-back before it acts, so a lost or duplicated ring cannot
corrupt state.
- The unarchive endpoint reuses the existing `companyService.update`
path, so agent reactivation and activity logging behave exactly like an
in-product unarchive.
## Model Used
- Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended
thinking and tool use enabled.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
canary/v2026.923.0-canary.1
|
||
|
|
be6f49a425 |
feat(runner): refresh shared coding harness runtimes (#13838)
## Thinking Path > - Paperclip runs agents through local adapters and the native runner. > - Both paths must use the same installed provider CLI. > - New models require current harness releases. > - The runner still pins Codex 0.153.4, Claude SDK 0.3.263, and OpenCode 1.18.29. > - Changing the image alone would fail the runner's exact version and executable checks. > - This pull request updates those dependencies, integrity checks, controller checks, and image pins together. > - Shared installations can then run the current models without a task-time download. ## Linked Issues or Issue Description Refs #13829, which updates model choices and reasoning controls. Searches found no open PR that updates these runtime pins. **Current behavior** The shared provider pack ships old CLIs. Claude Code 2.1.263 cannot run Opus 5.5, which requires 2.1.280. Remote controllers reject provider packs whose versions differ from their declared pins. **Proposed behavior** Use Codex 0.156.0, Claude Agent SDK 0.3.280 / Claude Code 2.1.280, and OpenCode 1.18.32 throughout the runner. Keep the reviewed ACP bridge patches and one shared CLI installation per provider. **Reason and benefit** Current harnesses support the new model IDs while preserving executable verification and remote provider-pack compatibility checks. ## What Changed - Update dependency overrides, the Codex ACP package patch, runtime profiles, and remote controller pins. - Verify the new Claude Linux x64 and macOS arm64/x64 executables and Codex Linux x64 executable against integrity-verified npm archives. - Refresh OpenCode version checks, fixtures, and the runner configuration label. - Refresh the eval image's Grok, Gemini, Kimi, Cursor, and GitHub CLI pins and archive hashes. Hermes remains current at 0.19.0. - Refresh the build-time lock digest from clean pnpm 9.15.4 resolution. Leave lockfile commits to repository automation. - Document model compatibility and the separation between CLI runtimes and patched ACP bridges. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - Rust workspace release tests passed. - Package/patch and OpenCode binary-materialization contract tests: 11 passed. - Real Codex 0.156.0 startup-ownership and paginated session-resume probes passed with isolated synthetic homes and no model turn. - Codex app-server `thread/start` preserved `gpt-6-sol` and `gpt-6-luna`; no `turn/start` was sent. An unauthenticated built-in catalog does not include those account-served entries. - Installed Claude integrity probes passed for `claude-opus-5-5` and `claude-fable-5-1`. - `pnpm --filter @paperclipai/paperclip-runner test:opencode:qualification` passed with the actual OpenCode 1.18.32 executable under Node 24 and Node 25. The loopback provider exercise covers health/version, session creation/read/delete, SSE, and a completed async prompt. - `pnpm check:token-gates` passed. - The targeted runner suite passed 130 tests. Three macOS failures in snapshot module lookup and OpenCode final-message selection also reproduce on the unchanged base; Linux CI will provide the platform check. - [Final Linux CI](https://github.com/paperclipai/paperclip/actions/runs/35798076399): all gates passed. Four jobs needed one retry after their CI workers received shutdown signals. The PR has 55 successful checks, two skipped checks, Greptile 5/5, and no unresolved review threads. - Changed runner configuration UI tests: 5 passed. - Full macOS `pnpm test:run` reached 13,094 passing server tests, 84 skipped, and 18 failures before the wrapper stopped. Failures involved skill-cache publication permissions, missing bundled connector skills in the worktree, and a conversation-reset timing case. The 10 cache permission failures reproduce on the unchanged base; both conversation-reset cases passed on a targeted retry. The wrapper did not reach its later workspace/serialized groups locally; Linux CI covers those groups. - The local Docker daemon did not respond, so no local Docker build was run. No billable model requests were made. ## Risks - Deploy the matching controller and provider pack together. Older controllers enforce their previous exact pins. - Current upstream CLIs can change behavior. Existing protocol tests and isolated real Codex probes cover the integration boundaries; authenticated model inference is not part of these checks. - ACP bridge package versions and executable digests stay unchanged because their executable bytes are unchanged. Only the underlying CLI/SDK dependencies move. - No schema migration. Revert the runtime and image pins together to roll back. ## Model Used OpenAI GPT-6 via Codex, with repository tools, code execution, and web research. The exact serving model ID and context window were not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass for the changed surfaces and real-executable probes; full macOS-suite limitations are listed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.0 |
||
|
|
cdf04a33fa |
feat(adapters): refresh current coding models and reasoning controls (#13829)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its adapters supply model catalogs and reasoning controls to agent setup. > - Several provider releases are missing from the fallback catalogs. > - Some newer models also have effort levels that the UI does not offer. > - Operators need the exact supported IDs and controls when discovery is unavailable. > - This pull request updates the existing adapters from current provider documentation. > - Operators can select current coding models without entering custom IDs. ## Linked Issues or Issue Description **What existing behavior does this improve?** Model selection and reasoning controls across the existing coding-agent adapters. **Subsystem affected** Claude, Codex, Grok, Gemini, Cursor, Kimi, and OpenCode adapters; model discovery tests; agent creation and editing. **Current behavior** The catalogs omit Opus 5.5, GPT-6 Sol/Luna, Grok 4.7/4.6/4.5, current Gemini Flash models, and several Cursor/Kimi choices. Bedrock has obsolete IDs. The UI omits supported effort levels and saves Grok effort under a key the runtime does not read. **Proposed behavior** Offer verified current model IDs and model-specific efforts. Remove retired Gemini 2.0 choices. Keep configured defaults and saved model IDs. Keep runtime discovery for account-specific choices. **Reason and benefit** Catch up with provider releases through September 22, 2026. Correct the picker and runtime controls together. **Breaking changes** No database or API change. Gemini 2.0 options leave the picker after their June 1 shutdown. Existing saved IDs remain unchanged. Corrected Bedrock catalog IDs do not rewrite saved configuration. **Additional context** Supersedes the separate GPT-6 Sol PR #13830. Fable 5.1 was already merged in #12730, and GPT-6 Astra in #12851. The Grok 4.6/4.5 proposal #11324 was closed and parked by its author. This change retains the default-sentinel fix from #12062. Related discovery proposals #13127 and #13565 do not supply these catalog and effort updates. Searches found no open PR for the additional model IDs. See [the dated audit](https://github.com/paperclipai/paperclip/blob/feat/claude-opus-5-5/doc/adapter-model-audit-2026-09-22.md) for exact scope, primary sources, runtime observations, and account-specific limits. This updates existing adapters and does not duplicate planned core work. ## What Changed - Add Opus 5.5 for direct Claude and Bedrock, with a Claude Code 2.1.280 gate. Correct and extend Bedrock model IDs. - Add GPT-6 Sol/Luna and Fast mode. Offer Ultra for Astra/Sol and GPT-5.6 Sol/Terra, and Max for both Luna generations. - Add Grok 4.7/4.6/4.5, expose supported Extra High effort, and save Grok edits under `reasoningEffort`. - Add Gemini Flash 3.8/3.7/3.6/3.5, Flash Lite 3.5/3.1, and 3 Flash Preview. Remove retired 2.0 choices. - Add the current documented Cursor fallback models, including Fable 5.1, Composer 2.5, and Muse Spark 1.3. - Refresh OpenCode fallback IDs used in remote environments from its installed provider registry. - Add Kimi K3 256K. Update the existing coding alias to K2.8 Preview and enable its CLI effort settings. - Use model-specific Claude/Grok efforts in creation and editing. Clear unsupported effort when switching models. - Add catalog, CLI/ACP forwarding, compatibility, and UI persistence coverage. Record the audit and sources. ## Verification - Latest head `6e63c9ef53b54ba869cd4fb431a8570bebe289f4`: 53 CI checks passed, 2 skipped. This includes full workspace typecheck, build, and all test shards. Greptile is 5/5 with zero unresolved threads. GitHub reports no merge conflicts. - 340 focused tests passed across adapter metadata, CLI/ACP arguments, Claude version checks, Kimi effort, Grok execution, server model discovery, and UI effort selection/persistence. - `pnpm --filter @paperclipai/adapter-claude-local --filter @paperclipai/adapter-codex-local --filter @paperclipai/adapter-grok-local --filter @paperclipai/adapter-gemini-local --filter @paperclipai/adapter-kimi-local --filter @paperclipai/adapter-cursor-local --filter @paperclipai/adapter-opencode-local typecheck` — passed. The same filters with `build` passed. - `pnpm check:token-gates` and `git diff --check` — passed. - Full workspace and UI typechecks were attempted locally. They stop on existing missing `three` dependencies in `packages/shared/src/cliplab`. - Full `pnpm test:run` and `pnpm build` were not run locally. Worktree creation exhausted disk space, so a clean dependency install is not feasible on this host. Focused checks reuse existing dependencies. CI supplies full workspace verification. - No provider inference was run. Account-specific runtime model lists were inspected where available. - Manual check: select the new models in agent setup and editing. Confirm Luna has Max but no Ultra, Grok 4.7 has Extra High, and Fable 5.1 has Extra High/Max. Save Grok effort and confirm `adapterConfig.reasoningEffort` contains the selection. ## Risks - Catalog presence does not grant account access. Older CLIs and restricted accounts can reject a model. Opus 5.5 has an explicit upgrade check. - Higher effort can increase cost and latency. Existing agent defaults are unchanged. - Cursor fallback IDs come from public model documentation; the local account exposed no live catalog. Runtime discovery still adds account-specific variants. - Kimi effort remains supported only on its explicit CLI engine. This does not add effort support to its default ACP engine. - Saved obsolete Bedrock or retired Gemini IDs are not migrated automatically. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7badae6981 |
feat(plugins): add an optional organization switcher slot (#13832)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Plugins can add UI surfaces to the board. > - Organization navigation is still fixed in the host sidebar. > - A distribution needs a supported way to supply its own organization menu. > - This pull request adds one optional React slot with host-owned navigation controls. > - The built-in menu stays available when the optional contribution cannot render. ## Linked Issues or Issue Description **Subsystem affected** Plugin SDK, server capability validation, and sidebar UI. **Problem or motivation** An installed plugin cannot replace the organization switcher without editing the host menu. Existing sidebar and overlay slots do not provide this replacement surface. **Proposed solution** Add an `organizationSwitcher` slot that requires `ui.sidebar.register`. Pass display state, an icon renderer, and navigation/logout callbacks. Keep the built-in menu for absent, ambiguous, missing, failed, or unsupported contributions. **Roadmap alignment** This keeps distribution UI in plugins and adds a small host contract. It does not add an account system or change company authorization. The maintainer requested this extension. Related prior menu changes: #12788, #10917, and #10850. No duplicate replacement-slot PR was found. ## What Changed - Add the slot to shared validation, SDK types, and server capability checks. - Wrap both sidebar menu variants with the optional replacement. - Resolve selection against the current account query before mounting, and reset replacement state on account or company changes. - Add host-specific component props and a fallback to `PluginSlotMount`. - Document the React-only contract and its trust boundary. - Report runtime-supervisor fixture startup details when CI readiness fails. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - `pnpm check:token-gates` passed. - `pnpm exec vitest run --project @paperclipai/ui`: 6,580 tests passed. - Targeted manifest, replacement, and built-in menu tests: 26 passed, including the incoming-account selection regression. - Installed a local test contribution into an isolated server. CLI inspection reported `ready`. Browser checks covered the loaded production UI, keyboard dismissal, current-organization selection, and an expired remote session. Remote account responses were fixtures. - `pnpm test:run` was attempted. Its first server phase passed 8,323 tests but failed in 36 files due to embedded PostgreSQL startup and filesystem permission errors on this Mac. Later phases did not run. The current Linux CI run is green: 54 checks passed and two were skipped, including build, typecheck, and browser gates. See https://github.com/paperclipai/paperclip/actions/runs/35796722770. - The earlier runtime-supervisor readiness failure did not reproduce locally. The complete affected shard passed locally: 58 files and 843 tests. The six supervisor tests also passed on Node 24.21.0 with CI flags. Added fixture startup diagnostics for the selected Node executable and listener port. The affected Linux shard then passed all 843 tests. The original root cause remains unconfirmed; no production runtime behavior or timeout was changed. ## Risks - Plugin UI remains trusted same-origin code. Display props do not authorize account requests. - A replacement can change navigation behavior. The host retains the built-in menu when discovery or rendering fails and keeps logout/session cleanup host-owned. - No database migration. Existing menus and portfolio behavior remain available. - Local full-suite verification is limited by the environment failures listed above. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser testing. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.14 |
||
|
|
f447990d77 |
fix(claude): recognize ACP quota fallback errors (#13831)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Claude adapter reports failures to the recovery system. > - Recovery must distinguish account quota from other provider limits. > - The Claude ACP bridge has a default message for an account with no quota. > - The current classifier misses that message and returns `acpx_turn_failed`. > - This pull request recognizes that exact typed message and uses the existing quota wait. ## Linked Issues or Issue Description Refs #13651. This is a narrow follow-up to its typed quota classification. Related PRs #13549 and #10276 address broader quota classification and reset handling. This change covers the bridge's exact default quota message. **What happened?** A typed Claude ACP `limit` failure with the title `The Claude account has no available quota.` returns `acpx_turn_failed`. The recovery result has no quota label. **Expected behavior** Return `provider_quota`. Use the existing one-hour quota backoff when the provider gives no reset time. Keep context, turn, rate, and configured budget limits out of the quota path. **Steps to reproduce** 1. Use the local ACP fixture to return the exact title above with category `limit` and severity `error`. 2. Run the Claude adapter in oneshot or persistent mode. 3. Before this change, the new regression cases receive `acpx_turn_failed` instead of `provider_quota` on ACPX 0.12.0 and 0.13.1. **Paperclip version or commit** Reproduced on `a959e4750`. Rebased onto current `master` before submission. **Deployment mode** Local source checkout with isolated ACP child-process fixtures. No live provider calls or customer-stack changes. ## What Changed - Recognize the exact Claude bridge quota fallback only for typed `limit` failures. - Add real-process regression cases for both pinned ACPX versions and both session modes. - Verify that unrelated categories, positive quota wording, and historical generic limit errors do not imply quota exhaustion. - Document the fallback and its existing recovery backoff. ## Verification - Red/green reproduction: all four new fallback cases failed before the classifier change and passed afterward. - 108 targeted tests passed across Claude ACP, quota, parser, and server recovery suites. - Claude adapter typecheck and build passed. - Real-process tests verify that provider text stays out of results and logs. - Full local `pnpm -r typecheck` and `pnpm build` passed. - The full local `pnpm test:run` attempt stopped after embedded PostgreSQL could not load a missing library symlink. The dependency setup was repaired in the worktree. The isolated database test then passed. The complete test suites passed in GitHub CI. - All 54 latest-head checks passed on `3d4d4e65386c5b6023ba34e6a2abcd94cff9a9bc`. Two optional Storybook jobs were skipped. - Greptile: 5/5, with no inline review threads or requested changes. ## Risks - Low risk. An exact match is required inside an existing typed `limit` failure. - If upstream changes this wording, this fallback can stop matching. Existing quota-message detection remains in place. - No schema, credential, or permission changes. Historical generic errors remain ambiguous and are not reclassified. ## Model Used OpenAI Codex, GPT-6. The session identifies the model family as GPT-6 but does not expose a more specific model ID or context-window size. Used reasoning, repository inspection, code editing, and terminal test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.13 |
||
|
|
a10702a878 |
feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people start and continue agent tasks from other services. > - A Slack conversation needs access to its surrounding discussion and Slack collaboration tools. > - The agent must use the linked requester's access and keep private material within its permitted audience. > - This pull request adds Slack tools through the existing connector contribution and approval framework. > - People can ask an invited bot to read a discussion, create follow-up tasks, and collaborate in Slack. ## Linked Issues or Issue Description **Subsystem affected** Chat connectors, connector runtime, tool gateway, and connection Settings/Access. **Problem or motivation** Slack-origin tasks can receive messages but cannot inspect the rest of a channel or act through the originating bot. People must paste context or configure a separate integration. **Proposed solution** Supply typed Slack tools and a bundled skill only to the originating task and assigned agent. Resolve the linked requester on the server. Check bot and requester access before reads and writes. Use existing durable actions and approvals. Retrieved messages remain source material. **Alternatives considered** Slack's user-OAuth MCP server does not replace the customer-created chat bot. An unrestricted Web API proxy would not provide suitable permission or publication boundaries. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps work with a provider contribution. It does not add a task dispatcher or a separate Slack task lifecycle. Related: #11144 covers generic per-user MCP grant execution; this change binds Slack bot operations to chat-origin tasks. ## What Changed - Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and shared native/HTTP execution. - Bind tools to company, endpoint, task, run, assigned agent, and admitted linked requester. Check membership and revocation on each call and before queued writes. - Add paginated reads, bounded history search, source links, messages, file uploads, reactions, pins, bookmarks, topics, canvases, lists, and approved channel operations. - Restrict private-source publication, including automatic replies and uploaded deliverables. Keep other people's bot DMs inaccessible. - Reuse action receipts, idempotency, approvals, and reconciliation. Suppress an identical explicit-send/final-reply duplicate. Return governed results through their verified originating conversation. - Add endpoint-bound personal search OAuth storage and lifecycle. Keep native real-time search disabled until a runtime meets Slack's transient-result requirements. Current runtimes use bounded history search. - Show capabilities, scope upgrades, and personal search authorization in Settings/Access and Storybook. Document provider and runtime limits. ## Verification - Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no unresolved findings. One unchanged rapid-callback timing test passed on a single CI retry. - Approval presentation regressions cover board-comment precedence and exact Slack publication; the expanded database assertion passed in CI. The local PostgreSQL startup probe later became unavailable, so that final assertion was verified in CI. Slack setup and failed-run retry browser tests also passed locally. - Full workspace typecheck and build passed. Server typecheck/build passed again after the approval routing fix. - Broad local suites passed in separate groups: server 12,958 tests, UI 6,555, shared 770, skills catalog 20, and other workspace packages 2,652. CLI and serialized server checks passed after environment/timeout retries. These are composite results, not one uninterrupted green full-suite invocation. - PostgreSQL authority regression covers admitted identity, cross-company/task/agent rejection, recovery, retained-session revocation, OAuth refresh/disconnect races, approval execution, exact publication lineage, retries, uncertain sends, and duplicate suppression. - Gateway/response regressions cover separate-origin approval batches and durable continuation. Focused provider, access, search, native runtime, route, and AgentMail regressions pass. - Storybook capability, missing-scope, OAuth configuration, authorization, and disconnect states were inspected in the browser. - Live staging: read a channel decision and full thread, create exactly two assigned backlog tasks, add a reaction, paginate discovery to exhaustion, and return bounded search matches with source links and coverage. - Live staging: create/edit/read a canvas and list, inspect the canvas in Slack, post/edit one message, and create a channel only after approval. New channels remain disabled for responses. - Live staging: read a response-disabled channel from the requester's DM; writes to that channel were denied. The test setting was restored. - Final live retest passed: explicit file upload and exact content read-back; approved deletion of only the disposable bot message; continuation confirmation returned to the original Slack thread without repeating the action. - Optional OAuth, private multi-user boundaries, native RTS, and CLI provider execution are not fully live-qualified. The staging agent initially supplied malformed tool arguments; valid arguments succeeded, and the tool/skill descriptions now emphasize UUID write keys. ## Risks - Existing Slack apps must add scopes and reinstall for new capabilities. Provider plans and document permissions can still restrict operations. - Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET` for governed tool actions. The staging instance was configured with explicit operator approval; fleet provisioning is a separate gap. - Native RTS is not exposed on current transcript-retaining runtimes. Bounded history scans are deliberately reported as incomplete. Inline file reads support text/canvas content up to 256 KiB; other types return metadata. - Private document edits fail closed when the full audience cannot be verified. Uncertain effects other than posts/uploads require inspection instead of blind retries. - Shared approval-delivery code now separates outcomes by source run to preserve origin boundaries. No database migration is required. - A separate completion-validator gap remains when the agent cites a prior run's registered artifact during finalization. It asked for registration again even though Slack delivery was confirmed. This change does not add a connector-specific task-completion policy. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and browser testing. The exact deployed model identifier and context-window size were not exposed in the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.12 |
||
|
|
a959e47508 |
fix(apps): reduce Google Chat scopes and block unread filters (#13820)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents controlled access to external services. > - Google Chat uses OAuth profiles and reviewed MCP tools. > - Those profiles request membership and read-state access that the supported feature set does not need. > - Removing read-state access also requires us to block unread search filters, including on existing connections. > - This pull request reduces both OAuth methods and enforces the reduced search contract before dispatch. > - Users retain conversation lookup, message history, ordinary search, and approved message sending. ## Linked Issues or Issue Description **What happened?** Both Google Chat profiles request membership and read-state scopes. The supported tool set does not include membership listing or read-state updates. Message search still advertises an unread filter. Related work: Refs #12619. **Expected behavior** Managed and customer-owned OAuth request only the scopes needed for supported features. Unsupported unread filters fail clearly before any provider call. Existing cached catalogs and broader grants must not bypass that policy. **Steps to reproduce** Start Google Chat OAuth from either connection method and inspect the requested scopes. Inspect the message-search tool schema, then submit a search with `searchParameters.isUnread` set to true or false. **Paperclip version or commit** The scope change is based on master at `110d176fc`. **Deployment mode** Managed Cloud and self-hosted instances with Google Chat Apps enabled. ## What Changed - Remove `chat.memberships.readonly` and `chat.users.readstate.readonly` from shared profiles and all four Chat connection methods. - Hide unsupported read-state fields and instructions in agent and board Test tool schemas. - Reject explicit unread filters, including false, null, snake-case fields, and encoded filter objects, before provider dispatch. - Recheck previously approved calls and support existing profile-bound and URL-only Chat connections. - Add scope, signed broker request, OAuth URL, allowlist, schema, and dispatch regression tests. - Document coordinated app/broker rollout, existing-grant reconnects, and the remaining deployment checks. ## Verification - All seven focused OAuth and Chat gateway test files pass: 498 tests on the rebased branch. - `pnpm -r typecheck` and `pnpm build` pass, using pinned pnpm 9.15.4. - The complete sharded CI test matrix passes, including general server, Chat, workspace, serialized server, Runner, and all eight browser e2e shards. The duplicate unsharded local `pnpm test:run` was stopped after CI passed; it did not complete locally. - `git diff --check` passes. - `node scripts/ingest-app-definitions.mjs` succeeds and leaves the branch unchanged. Google Workspace JSON is the durable reviewed input used by the generator. - All current-head CI checks pass at `7ee755371714dc036fff1c7da844776fee3f2ec1`, including typecheck, build, and canary dry run. Greptile is 5/5 with zero unresolved threads. The generator concern was withdrawn after review of the source and regeneration evidence. - No production deployment or live Google consent test was performed. After coordinated deployment, verify reduced consent scopes, normal search/history, message sending, and rejection of unread filters. ## Risks - Coordinate deployment with the companion Cloud broker scope change. Mixed versions can reject exact-scope requests. - Existing tokens are not narrowed or revoked. Grants with old scopes need new consent. Do not revoke a shared Google client to migrate one profile. - Explicit unread filters now return an error instead of being sent to Google. Ordinary search and the approved send tool remain available. - No database, UI, lockfile, or workflow changes. ## Model Used OpenAI Codex, a GPT-5-based coding agent, with tool use and code execution. The exact runtime model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9a76390cfa |
test(runner): improve blank-page diagnostics and infrastructure coverage (#13824)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner E2E tests verify real tasks and retain evidence for failures.
> - Some exposure tests assumed that port 42000 was free.
> - A blank task page could also fail without enough browser startup
evidence.
> - This pull request tests occupied ports and task reloads, and records
private startup diagnostics.
> - These changes make test failures easier to reproduce and explain.
## Linked Issues or Issue Description
Refs #13815. Report publication already received its production fix in
#13750; this PR adds a regression check for that workflow.
**What happened?**
Three exposure tests assumed that the allocator would select port 42000.
The synthetic failure disappeared when another process occupied that
port. A prior Daytona run also retained an empty task page after
navigation, but its evidence did not record pending modules or
service-worker control.
**Expected behavior**
Exposure tests must exercise the intended failure on the actual assigned
port. Browser failure evidence must distinguish an empty root from
loaded content. A saved task must remain usable after navigation and
reload.
**Steps to reproduce**
Run the exposure regression with the base port pair marked unavailable.
Run the browser-support tests with an unresolved entry module. Open a
saved task under the service worker, navigate to the same URL, and
reload it.
**Paperclip version or commit**
Based on master at
|
||
|
|
1ccae464c5 |
feat(ui): add task artifact media gallery and full-row links (#13825)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks collect the files and work products that agents create. > - The Artifacts tab shows these outputs in rows, which makes videos hard to compare. > - Small text links also make artifact rows harder to open. > - This pull request adds image and video tiles with previews and makes the full artifact row clickable. > - Users can compare outputs and open the existing media viewer with one click. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task Artifacts tab and media previews in task chat. **Current behavior** Video outputs appear as file rows or icons. Users must click a small link to open a work product. A task with eight video outputs gives little visual context. **Proposed behavior** Show images and videos in a responsive gallery. Show a paused video frame as the thumbnail. Keep documents, links, and other files in rows whose entire area opens the item. **Reason and benefit** Users can compare generated media without opening each item. Larger click targets also make the sidebar easier to use. **Breaking changes** None. This uses the existing artifact URLs, media viewer, run grouping, and attachment filters. No API or database changes. Related work: #11226 added the task sidebar output surface, and #7361 added rich attachment previews. #3524 concerns a separate reviewed-assets panel. This PR improves the existing task artifact components. The duplicate search found no active PR for this change. This is polish for the shipped Artifacts & Work Products roadmap item. ## What Changed - Add a shared media tile for work products and agent attachments. - Reuse video and image previews in task artifacts and chat. Seek up to one second into videos and reset preview state when the source changes. - Use the existing task gallery for playback and downloads. Preserve grouping and attachment deduplication. - Extend native links and buttons across work-product rows, including keyboard focus indicators. - Add eight offline Storybook examples for video outputs, mixed media, clickable rows, narrow and wide panels, missing previews, empty state, and light mode. - Register the component in the design guide and document its use. ## Verification - 108 focused component tests pass, including thumbnail seeking, source changes, gallery activation, and attachment deduplication. - `pnpm build`, `pnpm -r typecheck`, Storybook build, token gates, and `git diff --check` pass locally. - Reviewed the production components in the embedded browser. Checked all eight video thumbnails, mixed media, narrow layout, light mode, blank-area row clicks, keyboard gallery activation for generic-MIME images, and playback from chat video thumbnails. - Storybook: open **Tasks / Artifact Gallery** and select **Eight Video Outputs**, **Mixed Media And Files**, or **Whole Row Clickable**. The small local clips are synthetic fixtures. - All build, typecheck, unit, runner, and end-to-end CI jobs pass for `b8ace289b5e07df5b9f2c319b159f3923ce427b4`. Greptile gives 5/5 with both review findings resolved. All 54 PR checks pass, including the external security scan. The full test suite passed in CI. The duplicate serial local test run was stopped after CI finished; the 108 focused tests, full build, and recursive typecheck passed locally. ## Risks - Video thumbnails require the browser to load metadata and a frame. A slow server or unsupported codec can leave the fallback visible; opening and downloading still use the existing viewer. - Full-row click targets change pointer interaction with work-product cards. Native link and button semantics remain in place. ## Model Used OpenAI GPT-6 in Codex, with code execution and browser tools. The exact deployment ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
92d4868e79 |
fix(server): isolate run errors and redact runtime capability headers (#13826)
## Thinking Path
> - Paperclip manages AI agents and their work.
> - Operators use Sentry to investigate failed runs and server errors.
> - Run reports attach a task ID, run ID, error code, and adapter
fingerprint.
> - The server skips Sentry's OpenTelemetry setup to preserve its
separate tracing and privacy settings.
> - Without an async context manager, a scope mutation can attach old
run data to later errors.
> - The HTTP logger also retains a runtime credential capability header.
> - This change isolates run metadata and redacts that header so
diagnostics identify failures without leaking credentials.
## Linked Issues or Issue Description
Refs #13446 and #13719.
**What happened?**
After a terminal run failure, an unrelated server exception can inherit
that run's tags, context, and fingerprint. Sentry then groups a database
error with an earlier adapter failure. The real SDK reproduces this with
the application's `skipOpenTelemetrySetup: true` setting. HTTP request
logs also retain the `x-paperclip-github-capability` header, which must
be treated as a credential.
**Expected behavior**
Run metadata belongs to the terminal run event. Later exceptions must
not inherit it. Every genuine error must still be captured. Runtime
capability headers must be redacted on success and failure logs.
**Steps to reproduce**
1. Initialize the optional Sentry SDK with the application's options and
an in-memory transport.
2. Capture a terminal run failure.
3. Capture an unrelated exception.
4. Inspect the second event. Before this fix, it contains the first
run's identity and fingerprint.
5. Send a request with a fixture runtime GitHub capability header.
Before this fix, HTTP logs retain the fixture value.
## What Changed
- Pass tags, context, and fingerprint directly to `captureException`
instead of mutating the ambient scope.
- Preserve the existing run fields, grouping keys, ordinary exception
capture, and privacy settings.
- Test two run identities interleaved with unrelated exceptions against
the real optional SDK.
- Update the capture contract tests and document event-local run
metadata.
- Redact the runtime GitHub capability header through the existing HTTP
logger policy. Test successful, denied, and failed requests.
- Add a dedicated GitHub-hosted CI check that installs the exact
optional SDK version declared in `server/package.json`. It fails if the
real-SDK regression would be skipped. The SDK stays outside the
workspace and production dependency graph.
## Verification
- The real-SDK regression failed before the fix because the unrelated
event contained `contexts.run_failure`.
- Five focused suites passed: 123 tests, including all optional SDK
tests. Suites: `run-failure-sentry-real-sdk.test.ts`,
`run-failure-sentry.test.ts`, `sentry.test.ts`,
`run-failure-report.test.ts`, and `http-log-redaction.test.ts`. A custom
in-memory transport prevented outbound Sentry delivery.
- All three new header-redaction cases failed before the policy fix and
passed afterward.
- The dedicated CI command passed locally with
`PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1` and the audited SDK available
through `NODE_PATH`.
- Server TypeScript check passed with a scratch configuration that
resolves this checkout's workspace packages. The existing dependency
links point to another checkout.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- Gitleaks and a separate private-data scan passed before push.
- Full local workspace typecheck, test, and build were not run. The
machine has less than 2 GiB free and those commands include Rust builds.
Full PR CI must pass before merge.
- The dedicated real-SDK GitHub check passed with 1 test executed and no
skips: https://github.com/paperclipai/paperclip/actions/runs/35774449002
- Greptile reviewed
canary/v2026.922.0-canary.11
|
||
|
|
a68f3d8e35 |
fix(runner-e2e): align Daytona image and provider pack provenance (#13814)
## Thinking Path > - Paperclip is an open source app people use to manage AI agents for work. > - The Daytona runner image provides the native runner and its provider package. > - The controller also sends a provider package to native Daytona cells when the image package does not match. > - A stale image and a package from another source revision caused a 1.8 GB upload before Claude could run. > - This pull request refreshes the reviewed lock checksum and documents how to reuse the exact package from an immutable image. > - The benefit is a reproducible setup path and clear evidence when image and package provenance do not match. ## Linked Issues or Issue Description **What happened?** A Daytona native Claude run used image source revision `45c99a0d06cbd5b04982b06b79de321149930ac5` with a controller provider package from revision `294853dc...`. Runtime verification rejected the image package and staged a large package upload before model execution. A hosted campaign also failed during image setup because the Dockerfile expected lock checksum `d7d96cf0...` while the resolved lockfile checksum was `4b796c312833ebf2be4c38228babc0292c76774fb40d43bd18b54bd6b205753d`. Related public work reviewed: [#12795](https://github.com/paperclipai/paperclip/pull/12795), [#12862](https://github.com/paperclipai/paperclip/pull/12862), and [#12887](https://github.com/paperclipai/paperclip/pull/12887). **Expected behavior** The Daytona image and controller provider package must come from the same verified build. A package extracted from the immutable image must pass the existing manifest, source revision, lockfile, binary, bridge, and artifact checks before a native Claude run starts. **Steps to reproduce** 1. Set `PAPERCLIP_E2E_DAYTONA_IMAGE` to the old immutable image digest. 2. Set `PAPERCLIP_RUNNER_REMOTE_PROVIDER_PACK_PATH` to a package built from a different source revision. 3. Run a native Claude Daytona cell. 4. Observe provider package verification failure followed by the large staging upload. 5. Build the image with the stale Dockerfile lock checksum and observe the checksum failure. **Paperclip version or commit** `3b8df3dcd6f99e99277faa45f3351989edb5c239`. **Deployment mode** Daytona native runner E2E. **Install method** Built from source. **Agent adapter(s) involved** Claude Code through the native ACPX runner. **Database mode** Not database-related. **Access context** Not applicable to the setup failure. ## What Changed - Refreshed `PAPERCLIP_RUNNER_LOCK_SHA256` in `docker/daytona-runner/Dockerfile` to the resolved lockfile checksum. - Added a local guide for extracting the provider package from an immutable verified image with Docker. - Documented the required provenance checks and the expected manifest-matched runtime log. - Documented that cold package upload coverage must remain separate from recovery coverage. - Pinned the preview-service test guest to the test runner’s Node executable and logged guest startup and bound ports for readiness diagnostics. ## Verification - 442 focused E2E tests passed. - Typecheck passed. - `pnpm build` passed. - Five provider-pack reuse tests passed. - Six Daytona image contract tests passed. - Fixed Daytona campaign [35740613581](https://github.com/paperclipai/paperclip/actions/runs/35740613581) passed. - The overall hiring campaign [35739993219](https://github.com/paperclipai/paperclip/actions/runs/35739993219) failed because of an unrelated Mini metadata failure; its Claude cell passed. - Full local `pnpm test:run` was attempted but did not complete. The isolated Postgres install was repaired and its 15-test probe passed. - After the fixture change, all seven preview reservation tests passed locally and CI server shard 6/12 passed on `20234f75f`. This removes login-shell Node resolution variance; the exact cause of the earlier CI-only timeout is not established. - CI run [35766034635](https://github.com/paperclipai/paperclip/actions/runs/35766034635) passed on `20234f75f`. All server, runner, browser, build, and typecheck gates passed. - The signoff browser case initially failed waiting for an approver run. All five signoff tests passed locally without changes; the one allowed CI retry passed all 19 shard tests. This is recorded as an intermittent failure, not a demonstrated product fix. - Greptile reviewed `20234f75f`: 5/5, no actionable findings. - [Follow-up report](https://pages.paperclip.ing/runner-daytona-hiring-20260922/) includes timings, evidence links, and the remaining hiring configuration failure. - Review the immutable image source revision and extracted `provider-pack.json` before another paid recovery run. ## Risks - The Dockerfile checksum gate intentionally fails when the resolved lockfile changes. A future dependency change must refresh the reviewed checksum with the image change. - The local extraction guide requires Docker and a pullable immutable image. - An image built from an older source revision can still fail runtime manifest verification. The guide does not bypass that check. - The change does not alter runner prompts, approval policy, or recovery behavior. ## Model Used OpenAI Codex using the primary GPT-6 backend; the exact backend deployment ID is not exposed. Repository analysis and code execution used tool access. Assistance also came from OpenAI gpt-5.6-luna. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.922.0-canary.10 |
||
|
|
2788f20fc0 |
fix(ui): show ancestors in the task detail Tasks panel (#13823)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks form a hierarchy that explains why each piece of work exists. > - The streamlined Tasks panel shows child tasks and work created by the current task. > - It does not show ancestors, although the task response already includes them. > - This pull request adds linked ancestors above Subtasks in root-to-parent order. > - Users can now move up the task hierarchy from the same panel. ## Linked Issues or Issue Description Related implementation: #13241. A search found no duplicate ancestor-panel PR. **What happened?** The Tasks panel omitted the current task's ancestors. The streamlined header also hides hierarchy breadcrumbs. **Expected behavior** The Tasks panel should show the ancestor chain and let users open each ancestor. **Steps to reproduce** 1. Open a task with a parent and grandparent in the streamlined UI. 2. Open the Tasks tab in the side panel. 3. Observe that it shows child and created tasks but no ancestors. **Paperclip version or commit** Reproduced on master atcanary/v2026.922.0-canary.9 |
||
|
|
5f1100e3b3 |
refactor(server): remove retired operator UI snippet injection (#13789)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators can extend its interface through trusted plugin UI contributions. > - The server also accepts executable HTML through two legacy environment settings. > - This older path bypasses the plugin installation and lifecycle model. > - This pull request removes snippet injection from static and development pages. > - Operators must migrate existing integrations before upgrading. ## Linked Issues or Issue Description **What existing behavior does this improve?** Retire the operator HTML injection path from the server. Related public changes: #13168, #13245 and #13496 introduced the legacy settings; #13646 supplies the generic plugin host contract. **Current behavior** Managed instances append operator-supplied HTML or a decoded script body to every page. The same integration can use supported trusted plugin UI slots. **Proposed behavior** Ignore both retired settings. Serve normal branded static and development HTML, and keep plugin contributions unchanged. **Breaking changes** Installations that rely on `PAPERCLIP_CLOUD_UI_SNIPPET` or `PAPERCLIP_CLOUD_UI_SNIPPET_B64` must migrate before upgrading. The maintainer-owned staging and production deployments have completed the migration prerequisite. Other operators must migrate their integrations before adopting this change. ## What Changed - Remove the snippet injector and its static/dev rendering integration. - Remove injector-specific tests and retain a regression that old settings no longer change served HTML. - Replace setup instructions with a retirement and plugin-migration note. ## Verification - Rebased onto current master; `pnpm exec vitest run server/src/__tests__/static-index-html.test.ts server/src/__tests__/vite-html-renderer.test.ts`: 5 tests passed. - With repository-pinned Rust/Cargo installed, full `pnpm -r typecheck` and `pnpm build` pass after rebase. - Previous full local `pnpm test:run` encountered unrelated macOS runtime-cache rename `EACCES` errors and a missing AgentMail skill path; a focused reproduction confirmed 5 failures / 95 passes. That is not a passing full-suite result. The unchanged focused suites, build and typecheck were repeated after rebase; the full local suite was not repeated. All 54 refreshed GitHub checks/contexts passed on `bd62bff63f9d7980bfd10e54cb0102693d27ecfb`; fresh Greptile is 5/5 with no unresolved threads. - No UI component styling, database or API contract changed. - Maintainer approved the remaining rollout and cleanup. Staging snippet retirement and sleep/wake verification are complete. Production migration and removal of the legacy settings are complete for serving tenant instances. Remaining old warm inventory is excluded from new signups until configuration reconciliation completes. The maintainer authorized upgrading the remaining old deployments and clearing their pins. ## Risks - Removing the settings disables integrations that still depend on them; operators outside the completed maintainer rollout must migrate before upgrading. This is an intentional behavior change, documented at the existing setup-doc path. - Keep a previous image and its configuration for rollback. Existing browser tabs need a refresh to unload already-injected code. - Plugin UI remains trusted same-origin code. This does not add a security sandbox or change ordinary branding. ## Model Used - OpenAI GPT-6 (Codex; exact deployment variant and context-window size are not exposed in this session). Reasoning, repository inspection, local code execution and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused rendering tests; unrelated local full-suite failures are disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All refreshed checks pass on `bd62bff63f9d7980bfd10e54cb0102693d27ecfb` - [x] Fresh Greptile is 5/5 on the current head, with no unresolved findings - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3d78e3a4ec |
fix(runner): keep warm sessions alive with managed GitHub access (#13815)
## Thinking Path > - Paperclip manages AI agents and their work. > - The native Runner keeps a live provider process between task turns. > - Managed GitHub access used a token tied to one run. > - A new run forced Paperclip to replace that process to replace its token. > - This PR gives the session a stable credential transport and binds each operation to the active run. > - The agent can keep its process while Paperclip checks current identity and grants. ## Linked Issues or Issue Description Follow-up to #13738. Related credential-rotation work: #11770 and #8208 use process replacement for other adapter credentials; this change applies to managed GitHub access in the native Runner. **What happened?** A configured GitHub connection forced a warm native provider process to close at each new run. The saved conversation survived, but the live process did not. **Expected behavior** Keep the warm provider process. Resolve GitHub access for the current run when each command starts. Deny access while idle or after the run ends. **Steps to reproduce** 1. Configure managed GitHub access for a native Runner agent with a warm session. 2. Complete a turn, then send another message to the same task. 3. Observe the provider process close with the reason `warm native session configuration changed`. **Paperclip version or commit** Reproduced on master `8326e33ad`. Rebased onto `e3d8fb087` before submission. **Deployment mode** Local and remote native execution, including the sandbox callback bridge. ## What Changed - Move configured native GitHub transport and launcher ownership from the run to the provider session. - Bind the broker only after the executor acquires session ownership. Clear that binding when the run exits. - Keep the shared live-run, identity, grant, and trust-policy checks for each credential request. - Reject wrong scopes, idle requests, and credential responses that arrive after their run binding changes. - Retire transport and launcher files with the provider session. Keep anonymous commands available if bridge startup fails. - Add red/green executor tests, real subprocess and callback-bridge tests, and database checks. Update the runtime documentation. ## Verification - Before the fix, both new local and remote warm-session reuse tests failed. - After the fix, 435 targeted tests passed across the executor, broker, launcher, token, and database suites. - A real long-lived test process kept the same PID and original environment across two runs, including through the production callback bridge on local test processes. - Server typecheck and TypeScript compilation passed. - Full workspace typecheck and build passed. Server typecheck passed again after the review fix. - The fallback-logging regression failed before the fix; all 9 broker tests pass afterward. - The exact chat sidebar browser scenario passed locally. The initial CI timeout showed failed Vite module downloads; all eight browser shards pass on the latest commit. - All 53 latest-head checks passed, including the full CI test matrix and security checks (two unrelated conditional checks skipped). - The duplicate full local test run was stopped after CI passed; it is not claimed as a completed local pass. Targeted local tests, workspace typecheck/build, and the browser scenario passed. - Greptile reviewed the latest commit at 5/5 with no unresolved findings. - No fresh paid provider or Daytona campaign has run for this change. ## Risks - The broker now lives as long as the provider session. Tests cover idle denial, late cleanup, late responses, shutdown, and failed startup. - Its in-memory authority does not survive a controller restart. Existing checkpoint and process-recovery rules still apply. - Raw GitHub credentials remain confined to individual command processes. The session transport token cannot select a different task, agent, company, or run. - No database migration or public API change. ## Model Used OpenAI Codex, GPT-6, with reasoning, terminal tools, and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
74a9730acb |
fix: continue native agent chats after worker loss (#13813)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat uses native workers to run Claude and Codex conversations. > - A worker crash leaves a cleanup hold because its provider did not acknowledge suspension. > - A new user message must not reuse that unverified session or repeat old tool calls. > - The existing continuation path can preserve history and start a fresh session, but local cleanup ownership remained held. > - This pull request verifies the stopped local owners and releases only their cleanup hold for a new user turn. ## Linked Issues or Issue Description Refs #13775. **What happened?** After a native worker crashed, both providers retained cleanup quarantine. A saved plan survived, but the conversation could not produce another answer. **Expected behavior** Once the old worker and provider process groups have stopped, a new user message can continue in a fresh session with the saved work and prior action history. **Steps to reproduce** Run the opt-in `agent-chat-qualification` suite with case `worker-crash-retry` on native Codex and native Claude. The fixture saves a plan, kills the exact worker through a Linux pidfd, releases a local read-only brief, and sends a new message. ## What Changed - Verify the exact local worker stop receipt, provider identity receipts, released leases, and retained state before retiring a native cleanup hold. - Recheck process liveness and state before admission. Keep the old run and durable session files intact. - Use the existing explicit conversation continuation path. Generic Retry remains blocked for cleanup quarantine, including on the old failed-run marker after a successful continuation. - Extend the live oracle to require a successful fresh session, correct predecessor context, unchanged plan, one original message, and one answer containing a reference introduced after the crash. - Add physical-proof and database-backed admission tests. Document the precise qualification scope. ## Verification - Live Product E2E: **2/2 passed**, **2/2 cleanup passed**, with real native `gpt-5.6-sol` and `claude-sonnet-5`, Chromium, server, database, and public APIs. - Core recovery proof source: `3592b04c2bc76e23795fcdf964720e38a409dc4d`. Suite definition version 8: `9867367994d81a0c726956d91f2c7fddab6417a12f41b5cef3f7e62f3be417da`. - [Core recovery campaign](https://github.com/paperclipai/paperclip/actions/runs/35741746990) · [Public evidence report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35741746990-1/). - Both cells verify the real crash boundary, blocked generic Retry, unchanged saved plan, one original prompt, one fresh successor with predecessor context, and one run-attributed answer containing the post-crash reference. Billing coverage is partial because the crashed runs did not report complete usage; missing cost is not zero cost. - [Master baseline](https://github.com/paperclipai/paperclip/actions/runs/35737364443): both providers stopped at quarantine, with cleanup passing. - [Follow-up baseline without the fix](https://github.com/paperclipai/paperclip/actions/runs/35738638866): Codex produced a complete red result. Claude reached the same error, but its artifact upload was canceled. - [Complete Claude baseline with the same version-8 definition](https://github.com/paperclipai/paperclip/actions/runs/35740555177): red at cleanup quarantine, cleanup passed, source `6479a90c5754044356b39a9278b9a3e92ce8e55e`. - Earlier candidate attempts remain retained: [first](https://github.com/paperclipai/paperclip/actions/runs/35738449214) passed Codex and found a Claude fixture wait race; [second](https://github.com/paperclipai/paperclip/actions/runs/35740408065) exposed the normalized session-open receipt mismatch. Both corrections are in the final source. - Targeted server suites: 570 passed before the final two additional receipt regression cases. The physical-proof suite, including those cases, passed 46/46. Eval oracle and catalog: 40 passed. Server and Product E2E typechecks passed. - **All 54 PR checks passed on final head `db6f775df`**, including repository typecheck, build, tests, browser shards, and canary dry run. The canary job required one retry after its runner received a shutdown signal. On the earlier core proof head, two timing-sensitive tests passed in isolation and on a single CI retry. - Final UI regression checks: 5 passed; UI typecheck and token gates passed. Updated eval oracle/catalog: 40 passed; eval typecheck passed. - Final version-9 two-provider campaign: **2/2 passed, 2/2 cleanup passed**, including the browser assertion that the quarantined historical run never regains Try again. [Final campaign](https://github.com/paperclipai/paperclip/actions/runs/35747416013), source `db6f775dfff405e1514ec02fedb0450d42c7dad2`, definition hash `bf5abf1cc45cb6dad4e082fbf818b8fa0f4d8c282776a7e98762b22919eadab7`. [Final public evidence report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35747416013-1/). Final screenshots and retained state inspected for both providers; cost coverage remains partial. ## Risks - Missing or conflicting stop evidence keeps the conversation blocked. This change does not kill an unverified process. - This qualifies new user input after local worker loss. It does not enable automatic replay, exact-session recovery, remote crash recovery, or native onboarding defaults. - Old action outcomes remain part of the continuation. A process exit is not proof that an action did not happen. - No schema migration or production prompt change. ## Model Used OpenAI Codex, GPT-6. The runtime does not expose a more specific model identifier or context-window size. Used reasoning, repository inspection, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |