mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
master
338
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0e0b63e5a5 |
feat(connections): add experimental task-pinned AI routing (#14967)
## Thinking Path > - Paperclip manages AI agents and their work. > - AI Connections separate account access from models and harnesses. > - A pool must act as one connection while retaining each task’s account. > - Core must enforce member access and preserve session and recovery rules. > - A plugin supplies rotation policy without receiving credentials. > - This change adds durable routing and native connector setup and management. ## Linked Issues or Issue Description **Subsystem affected** AI Connections, Connectors, plugins, run dispatch, and session compatibility. **Problem or motivation** Operators need to rotate new tasks across saved accounts while each task keeps its account and session. Pool setup must fit the existing connector catalog and account workflow. **Proposed solution** Add an experimental router binding, a capability-gated plugin hook, and transactional task pins. Plugins declare native pooled connectors through `aiConnectionRouter`. Core hosts the existing-account picker, ordering step, and account settings. Related usage contract: #14936. Companion private plugin: https://github.com/paperclipai/paperclip-cloud/pull/643. **Roadmap alignment** This extends Apps and AI Connections. Core supplies generic enforcement and native connector UI; the private plugin owns rotation and quota policy. The prior duplicate search found no matching router implementation. ## What Changed - Add a router binding without changing existing concrete bindings. Keep the instance flag and new pools disabled by default. Require manual operator configuration. Show no routing toggle in Experimental settings on either open-source or Cloud installs, even after routing is enabled. - Persist company-scoped pools, one shared cursor per pool, and pins keyed by company, pool, agent, and task. Commit pins and cursor advances together with revision checks and bounded retries. Persist run-ID affinity before allocation. - Pass only authorized metadata and normalized usage to plugins. Core retains credential handling, member access checks, runtime qualification, and recovery evidence. Probe outside locks with a shared 15-second budget and freshness cache. - Resolve routing before credential preparation and backend selection. Preserve pins through turns, session resets, removed members, and quota waits. Retain admitted recovery after disable or uninstall. - Separate credential session epochs from token generations. Verified refresh preserves the epoch; reconnect and manual replacement change it. Include the credential slot ID in session and usage-cache identity, so reconnecting an indexed legacy account invalidates its old session even when both epochs are zero. - Validate pool member installations before accepting saved-agent bindings and recheck compatibility when the harness changes. Install only authorized members in the new-agent transaction and record their IDs in local activity. Pool membership cannot install a restricted shared connection. - Preserve pool bindings when agents hire teammates through either creation API or native caller runtime inheritance. Block stale manager credential references; retain explicit child authentication precedence and reject incompatible inherited pools. - Add native connector registration through plugin metadata. Reuse the Connectors catalog, setup header, account header, sidebar, dialogs, and usage display. Setup selects and orders saved connections. Advanced settings hold usage rules and member runtime defaults. New-account setup opens in another tab. - Use revision-checked pool archival from the Connectors catalog and account page. Keep task pins, cursors, recovery evidence, and underlying connections. Reject ordinary connection updates or removals that bypass pool revisions. - Add pool selectors, composer models, override notes, quota status, run details, activity records, and local run-log records. Keep session-adoption copy minimal. - Show **Used by** below the pool connections. List current company agents with shared avatars and profile links. Include paused agents; exclude terminated agents and agents using another pool. - Add Core stories for the generic connector workflow and runtime surfaces. Cloud stories reuse these production routes and tokens through a preview-only alias. ## Verification - Final head `73cb953bca30ed83e4505dd820edd9b5edffd28b`: full workspace `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass locally. - All 422 focused connector/settings/shared-contract/migration tests and all 156 database-backed AI connection, hiring, reconnect, and durable-routing cases pass (69 hiring cases rerun after the final auth-precedence fix). The merged shared contract retains connection instructions and pool metadata. The pool migration is generated at sequence 0299 after the latest upstream migrations; this PR makes no lockfile changes. - All four full-app Playwright tests pass on the final head after a cold restart and migration, against the installed private plugin and isolated database, with no route or pool-API mocks. They cover hidden routing controls after manual opt-in, native pool creation, ordering, membership edits, rename, paused defaults, enabling/save/refresh persistence, stale edits, cancellation/removal, preserved underlying accounts, unavailable routers, and Used by avatars and profile links. Exact command: `PAPERCLIP_CONNECTION_POOL_E2E=1 AI_CONNECTIONS_TEST_COMPANY_ID=a37b9625-5ecf-4e29-8081-04df3d6e7d6f AI_CONNECTIONS_TEST_URL=http://127.0.0.1:3108 pnpm exec playwright test --config tests/ai-connections-app/playwright.config.ts connection-pools.spec.ts`. - [Native setup, ordering, and management screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-6006976278) address the review follow-up. [Earlier selector, quota, and run-detail screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-5971537316) show the runtime surfaces. Core previews: `pnpm --filter @paperclipai/ui storybook`, then **Connectors / Pool host** or **AI Connections / Connection pools**. Cloud owns its host-backed plugin stories; both repositories’ Operator Setup Required story assertions pass. - Live acceptance used OpenAI/Codex and Anthropic/Claude ACPX, resumed both exact sessions after restart, preserved pinned accounts through explicit reset and controlled quota deferral/recovery, and committed only two allocations across fourteen runs. A later UI-created task test again rotated OpenAI then Anthropic and resumed OpenAI through follow-up/restart/quota recovery. That later Anthropic execution was blocked by its saved OAuth token expiring (provider 401). No live usage probes ran. - The full local `pnpm test:run` was attempted earlier and did not complete because of macOS embedded PostgreSQL bootstrap/shared-memory failures and the 40,000-file Git fixture timeout. The focused database suites above now pass; full-suite verification is provided by the split CI lanes. The preceding CI run had one runtime readiness timeout; it passes locally both alone and inside the larger runtime suite. That larger local suite also encountered an embedded PostgreSQL setup failure and two macOS temporary-path alias assertions; those two assertions pass with canonical TMPDIR=/private/tmp. All final-head CI checks are terminal green, including full general/serialized server suites, Runner checks, browser E2E shards, canary verification, build, and typecheck. Greptile is 5/5 on that exact head with no unresolved threads. ## Risks - The migration adds routing tables and a credential epoch column. Install the private plugin only with the compatible Core contract. - Routing and each pool require opt-in. Production distribution and fleet defaults remain unchanged. - Unknown usage stays eligible. Known pinned exhaustion waits; revoked access requires operator repair. - Legacy adapters require compatible members. Runner model and effort overrides remain limited by qualified backend support. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and browser testing. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes:` / `Refs:` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted suites; full-suite limitations are reported above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4857799a88 |
feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes. Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
984f092ddf |
Add project and agent resource lifecycle hooks (#15280)
## Thinking Path > - Paperclip manages agents and projects for work. > - Plugins need reliable lifecycle hooks for these resources. > - In-process notifications can disappear during a restart. > - A lifecycle record must commit with the resource change. > - This pull request adds a generic lifecycle journal without provider calls. > - A later plugin delivery layer can use these records without losing lifecycle changes. ## Linked Issues or Issue Description **Problem or motivation** Resource provisioning needs durable hooks for agent hiring, pause, resume, termination, and project creation. Hooks must cover shared service callers, budget actions, and approval paths. Capture should work for all plugins and deployment modes. **Proposed solution** Record content-free events in the resource transaction. Deduplicate creation and unchanged status. Preserve each pause/resume cycle. Commit approval, activation, and creation together. Commit termination and API-key revocation together. **Alternatives considered** Event subscriptions alone cannot survive process failure. Cloud-only capture would exclude other plugins. Provider calls inside tenant transactions would couple resource creation to external services. **Roadmap alignment** This is lifecycle infrastructure for the existing plugin system. It adds no provider, adapter, UI, or plugin read API. Repository mutations and backfill remain separate work. Searches found no duplicate lifecycle-journal PR. ## What Changed - Add a company-scoped lifecycle journal and an additive migration. - Record hired-agent and project creation from shared services. - Record agent pause, resume, and termination, including generic updates and budget actions. - Lock agent state changes to suppress concurrent duplicate hooks. - Make hire approval, rejection, and termination transactions atomic with their lifecycle records. - Document capture, ordering, future per-plugin acknowledgments, and migration scope. ## Verification - Full workspace typecheck: `pnpm -r typecheck` passed. - Production build: `pnpm build` passed. - Focused agent, project, approval, budget, and built-in regressions: 102 tests passed across 9 suites. - Final lifecycle journal check: 14 tests passed. It covers repeated/concurrent transitions, rollback, key revocation, self-hosted capture, and company boundaries. - Review regression: 32 lifecycle and approval tests passed, including rejected-hire rollback and retry; server typecheck passed after the fix. - The initial local `pnpm test:run` overlapped the rejection fix and reported the new rollback regression against the earlier service code. A fresh lifecycle run passed all 14 tests. Fresh full local shards were stopped once the complete CI test matrix passed on the final commit. - `git diff --check` and a local secret/PII scan passed. - Greptile: 5/5 on `3ee3903f8e1a171ea3ba2bea9caa2b5766a3f807`, with the review thread resolved. - [Complete CI passed](https://github.com/paperclipai/paperclip/actions/runs/37374227222) on `3ee3903f8e1a171ea3ba2bea9caa2b5766a3f807`: general server/chat/workspace tests, serialized server suites, runner checks, build, typecheck, canary, and all end-to-end shards. The requester waived CI during the Actions outage, but the workflow subsequently completed successfully. ## Risks - The migration creates an empty table. It does not scan or backfill existing resources. - Apply the normal database migration before running this server version. Event-write failure intentionally rolls back the resource change. - Pending hires cannot bypass approval through pause or resume. - Capture works on all deployments. Plugin delivery, retention, retries, and provider actions remain separate work. No plugin can read this journal through a new API in this PR. - Future delivery must enforce company scope, track acknowledgments per plugin, and preserve resource order. A global sequence cursor can skip uncommitted transactions. - A termination hook does not authorize deleting persistent volumes. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code execution, and tool use. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b43073d11f |
feat(connections): sync and group accounts managed by aggregators (#15254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external tools. > - Aggregator gateways can expose accounts that users already connected upstream. > - The Apps catalog did not show those accounts or their current provider status. > - Separate cards and setup tasks also made account ownership unclear. > - This pull request discovers upstream accounts and groups them under one app card. > - Users can find connected apps while each provider keeps control of its accounts. ## Linked Issues or Issue Description **Subsystem affected** Connections across the database, shared contracts, server, and board UI. **Problem or motivation** Users cannot see which apps are connected through a saved aggregator gateway. Native and upstream accounts need one app card. Discovery must preserve company, user, gateway, and credential boundaries. **Proposed solution** Sync account metadata from Composio, Arcade, and supported Executor gateways. Keep upstream account management in each provider. Use source chips and search to browse the catalog. Preserve native setup and the gateway's existing access policy. **Alternatives considered** Creating a local executable connection for each upstream account would duplicate authorization state. Using an agent task for routine Composio setup would add an unnecessary step. The board now calls the saved gateway directly for that setup. **Roadmap alignment** This extends the shipped Connected Apps and MCP Tool Gateway features in ROADMAP.md. The duplicate search found no open PR for managed account discovery. Related work: Refs #13755, Refs #13941, Refs #14725, Refs #13855. Open PR #12906 covers adjacent toolkit routing work. ## What Changed - Add provider-neutral discovery, sync, and refresh APIs. Preserve the Composio API paths. - Cache observations by company, saved gateway, viewing user, and credential version. Retain stale observations after failed or incomplete scans. - Add optional Arcade account sync credentials in the vault. Discover Executor accounts through its supported inventory interface. - Group native and upstream accounts in one app card. Imported account menus open their provider. Gateway menus own refresh and sync setup. - Add Paperclip, Composio, Arcade, Installed, and All chips. Show 50 catalog entries per page. Keep connected accounts above discovery. Keep explicit provider searches scoped. - Simplify Composio app setup and refresh its connected app list on the gateway Permissions page. - Add a compact agent access card and task creation defaults for connection setup. Preserve explicit blocks and approval policies. - Add two replay-safe migrations, service and UI tests, Storybook journeys, and acceptance stories. ## Verification - Passed the repository typecheck, full build, token gates, and migration ordering check. - Passed the focused provider adapter, connection interaction, and catalog tests after rebasing onto master. - Passed all nine database sync and migration replay tests using a disposable database on the test-drive PostgreSQL cluster. Removed that database after the run. - Verified Arcade cursor pagination against its official Go SDK and passed all eight adapter tests, including short and incomplete pages. - Passed all 45 interaction tests after making the exact requested tools and their Allowed/Ask first permissions visible before granting access. Verified the compact card in Storybook. - Passed the complete UI suite on the final code: 683 files and 7,432 tests, including the corrected Composio destination assertions. Passed 130 focused tests for the UUID, management-link, and health-status corrections. - Passed 22 Composio setup/sync tests, 23 connection-intent service tests, and the connection migration test in separate disposable databases. Database startup alone was substituted; the suites exercised their real SQL and services. - Passed all 10 OpenAPI route checks and the full-stack connection-intent browser test, including scoped consent, agent continuation, and task completion. - The local full runner encountered embedded PostgreSQL startup failures on this loaded macOS host. The earlier in-flight run also held the pre-fix Arcade transform; a fresh run of the final provider suite passes. The final-head CI is queued during GitHub’s active Actions incident: https://www.githubstatus.com/. The previous run also lost several runners simultaneously; its real catalog assertion failures are fixed and the fresh complete UI suite passes. - Tested the real test-drive server in the embedded browser with a live Composio gateway. Detected Airtable and Circleback. Verified refresh progress, account rows, source chips, search scope, and 50-entry pagination. - Arcade and Executor coverage uses provider fixtures. Live credentials were unavailable. - Storybook builds successfully and includes grouped native/provider accounts, stale and unavailable discovery, optional Arcade setup, and mobile states. The acceptance document records the simulated and live coverage separately. - Greptile reviewed final commit `217b024c27b5933e773ce9419c4e92b1032042c6` at 5/5. All six review threads are resolved, security scans pass, and the PR has no merge conflicts. The outstanding remote checks are `ci / Select trusted runner` and `review`, queued by GitHub. They need to complete before merge. ## Risks - Provider response changes can break inventory discovery. Failed scans retain observations and show stale status. - Composio scans only the supported catalog and can take time. Large inventories run in the background with progress and a bounded lease. - Arcade requires a project API key and user ID when the gateway cannot supply them. This key is used only for discovery. - Executor discovery depends on the server's exposed inventory tools. Unsupported servers report unavailable discovery. - Cached account rows do not grant access or create executable connections. Gateway policies still govern tool use. Account deletion and per-app authorization remain upstream. - The migrations add tables and one nullable column. Replay preserves existing rows and company-scoped foreign keys. ## Model Used OpenAI Codex, based on GPT-6. The session does not expose a more specific serving model ID or context limit. Used reasoning, repository tools, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cf8ad63c80 |
build(deps-dev): bump tsx from 4.23.12 to 4.23.15 (#12965)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.23.12 to 4.23.15. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/privatenumber/tsx/releases">tsx's releases</a>.</em></p> <blockquote> <h2>v4.23.15</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.14...v4.23.15">4.23.15</a> (2026-09-20)</h2> <h3>Bug Fixes</h3> <ul> <li>exclude bare builtins from namespace inheritance (<a href="https://github.com/privatenumber/tsx/commit/38e158857e50bca311be7c232a5057e5a2e5347a">38e1588</a>)</li> <li>expose require.cache and require.extensions to tsImport CommonJS modules (<a href="https://github.com/privatenumber/tsx/commit/2da34075afaed43e2b7fd0aca5fbebaaf337ff3a">2da3407</a>)</li> <li>make namespaced register() overloads portable for declaration emit (<a href="https://github.com/privatenumber/tsx/commit/562c434a5c8695e74327bbeb51cfeb9b86fc7e15">562c434</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.15"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.14</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.13...v4.23.14">4.23.14</a> (2026-09-20)</h2> <h3>Bug Fixes</h3> <ul> <li>restore the CJS bridge namespace for Node 24 require(esm) under tsImport() (<a href="https://redirect.github.com/privatenumber/tsx/issues/802">#802</a>) (<a href="https://github.com/privatenumber/tsx/commit/6e5236b065738d3687a06396d064774cfede390f">6e5236b</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.14"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.13</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.12...v4.23.13">4.23.13</a> (2026-08-30)</h2> <h3>Bug Fixes</h3> <ul> <li><strong>cache:</strong> bound shared transform cache memory (<a href="https://redirect.github.com/privatenumber/tsx/issues/835">#835</a>) (<a href="https://github.com/privatenumber/tsx/commit/28e1f12d04cd2afe1db17f8555b14fe5fb567c6e">28e1f12</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.13"><code>npm package (@latest dist-tag)</code></a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/privatenumber/tsx/commit/ca66105a17a2a4c6503fe3a12b5b9ec408286011"><code>ca66105</code></a> test: fix drive-less file URLs in ESM resolver fixtures</li> <li><a href="https://github.com/privatenumber/tsx/commit/2da34075afaed43e2b7fd0aca5fbebaaf337ff3a"><code>2da3407</code></a> fix: expose require.cache and require.extensions to tsImport CommonJS modules</li> <li><a href="https://github.com/privatenumber/tsx/commit/38e158857e50bca311be7c232a5057e5a2e5347a"><code>38e1588</code></a> fix: exclude bare builtins from namespace inheritance</li> <li><a href="https://github.com/privatenumber/tsx/commit/562c434a5c8695e74327bbeb51cfeb9b86fc7e15"><code>562c434</code></a> fix: make namespaced register() overloads portable for declaration emit</li> <li><a href="https://github.com/privatenumber/tsx/commit/edfb1f05a3f40b879a41a03a0801c2abd3a3ecf9"><code>edfb1f0</code></a> build: upgrade pkgroll and externalize CJS loader reference</li> <li><a href="https://github.com/privatenumber/tsx/commit/70e78284837c859f09b96cd10cd71d007aa4b795"><code>70e7828</code></a> test: upgrade tinyspy for disposable API</li> <li><a href="https://github.com/privatenumber/tsx/commit/9ed2022dfa9ea1be9511fe6abcde8110c25055a7"><code>9ed2022</code></a> ci: avoid duplicate release notifications</li> <li><a href="https://github.com/privatenumber/tsx/commit/872e77ffc5e96ca5c4727e74c0694debcb26219b"><code>872e77f</code></a> refactor: use disposables for cleanup</li> <li><a href="https://github.com/privatenumber/tsx/commit/6e5236b065738d3687a06396d064774cfede390f"><code>6e5236b</code></a> fix: restore the CJS bridge namespace for Node 24 require(esm) under tsImport...</li> <li><a href="https://github.com/privatenumber/tsx/commit/28e1f12d04cd2afe1db17f8555b14fe5fb567c6e"><code>28e1f12</code></a> fix(cache): bound shared transform cache memory (<a href="https://redirect.github.com/privatenumber/tsx/issues/835">#835</a>)</li> <li>See full diff in <a href="https://github.com/privatenumber/tsx/compare/v4.23.12...v4.23.15">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
cad26c6bfb |
fix(tool-gateway): bound MCP discovery memory and concurrency (#14864)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents discover governed tools through the MCP gateway. > - A listing repeated policy and full run-row reads for each catalog tool. > - Parallel listings multiplied those allocations during run startup. > - One 900-tool baseline listing used 11,489 queries and about 1.7 GiB of extra heap in a fixture. > - This pull request shares reads within a listing and bounds whole listings across the process. > - The benefit is lower discovery memory use while execution still checks current policy. ## Linked Issues or Issue Description Refs #13115. Its on-demand target change affects the same listing loop. **What happened?** MCP discovery repeated roughly 13 reads per tool. Full run snapshots and repeated connection configurations caused large allocations. Per-listing bounds alone did not limit concurrent listings across gateways. **Expected behavior** Discovery reads shared inputs once per listing. The process bounds active and queued listings. Catalog payload size and policy evaluation still grow with the catalog. Tool execution checks current access rules. **Steps to reproduce** Create a company with a remote MCP connection, 900 catalog tools, large schemas, and a large run snapshot. Send concurrent tools/list requests using a run-bound gateway token. Run the committed benchmark for a deterministic reproduction. **Paperclip version or commit** Baseline: |
||
|
|
98d8a6ccac |
Stop replaying ambiguous database disconnects (#14773)
## Thinking Path > - Paperclip stores agent work and control state in PostgreSQL. > - Its database client must not repeat a mutation after an uncertain result. > - The global retry wrapper treated `write CONNECTION_CLOSED` as proof that PostgreSQL never received a statement. > - postgres.js also uses that message when the connection closes after statement delivery. > - This pull request removes that global replay and tests the actual driver over a local wire connection. > - Callers retain control of retries when they can prove the complete operation is idempotent. ## Linked Issues or Issue Description Follow-up to #13417. Preserve the transaction disconnect handling from #13643 and the explicit actor synchronization retries introduced in #12773. Searched open and closed issues and PRs for database retries, disconnects, and `CONNECTION_CLOSED`. The open circuit-breaker proposal #11142 addresses outage queue growth; it does not establish whether an already-sent statement can be replayed. **What happened?** The database wrapper replayed an arbitrary statement up to three times after `write CONNECTION_CLOSED`. The driver adds `write ` to connection-close errors even after the peer receives the statement. A local protocol peer receives the same submitted INSERT three times when it drops each response. A committed write could therefore execute more than once. **Expected behavior** An ambiguous statement result must fail without automatic replay. A subsequent operation must be able to reconnect. **Steps to reproduce** Run the new wire regression against the parent commit. The six Simple Query cases and the parameterized Drizzle case receive three executions instead of one. The named prepared-client case was already safe and stays covered. The peer reads the entire statement and then closes the connection. This demonstrates repeated delivery with the real driver; it does not claim that a historical incident duplicated a committed write. **Paperclip version or commit** Reproduced on source commit `018993140f` with the patched postgres.js 3.4.9 dependency. **Deployment mode** Built from source with a local PostgreSQL protocol peer. No live provider or customer database is used. ## What Changed - Pass the original postgres.js client to Drizzle and remove the global statement replay wrapper. - Add eight wire regressions: six Simple Query cases for INSERT, side-effect-capable SELECT, and a data-changing CTE, plus parameterized Drizzle and named prepared-client cases. The extended peer processes Parse, Describe, Bind, and Execute, verifies bound parameters, and drops the response only after Execute. Each case checks one delivery and recovery on a fresh query. - Document ambiguous outcomes and the retry compatibility tradeoff. Keep explicit idempotent actor-sync retries and disconnected-transaction handling unchanged. ## Verification - Before the fix: the six Simple Query cases and the parameterized Drizzle case failed with three executions instead of one. The named prepared-client case was already safe. All eight wire cases pass on this branch. - Final focused client, pool teardown, configuration, and actor-sync retry checks: 28 tests passed. `pnpm --filter @paperclipai/db typecheck` also passed after the test-only follow-up. - First implementation head, `pnpm exec vitest run --project @paperclipai/db`: all 158 tests passed across 45 files, including real PostgreSQL transaction/reserved-connection recovery. The local embedded dependency's symlinks were hydrated before this run. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - The complete local `pnpm test:run` did not finish; no complete local suite pass is claimed. All CI test, typecheck, and build gates passed on the first implementation head `11f8b22d90`. Final-head CI is pending after the test-only follow-up. - `git diff --check`: passed. Reviewed the diff for secrets, personal data, generated output, and run artifacts. ## Risks Some transient statement failures that the global wrapper previously replayed now reach the caller. Operation owners must retry only when they have an idempotency guarantee or a durable receipt that prevents duplicate effects. A connection error is not proof that a write failed to commit. There is no SQL-text retry heuristic, new suppression, schema change, or migration. This change prevents unsafe replay; it does not prevent network disconnects. ## Model Used OpenAI Codex / GPT-6, with reasoning, repository inspection, code execution, and local protocol tests. The exact backend model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Final verification (September30): every final-head CI check passed at `3004c5bda39c985c3557547ec45e33870ce5d010`. Greptile scored5/5 on this head, all review threads are resolved, and the branch is mergeable. Eight real-wire regressions cover simple, parameterized Drizzle, and named prepared queries. Full local suite did not produce a completed result; the complete CI matrix passed. This public PR remains open for maintainer merge. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018993140f |
feat: let agents name prompt-only tasks (#14761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d432dc7fa3 |
Add GitHub-synced skill sources (#14713)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Company skills supply instructions and files to those agents. > - GitHub imports already exist, but users cannot manage repositories as skill sources. > - Repository refresh also needs caller-authorized access and complete local packages. > - This pull request adds Sources inside Skills and reuses GitHub connections from Apps. > - Installed snapshots let agents use skills without fetching GitHub during a run. > - Manual refresh preserves skill identity and leaves failed imports on their last good version. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: skills UI, server, database, shared contracts, and runtime materialization. **Problem or motivation** Users keep skills in GitHub repositories. They need a clear way to select, import, and refresh those skills. Existing imports do not expose repository management or consistently preserve supporting files. **Proposed solution** Add company-scoped skill sources. Browse repositories from all accessible GitHub connections, or paste a public repository or branch URL. Select whole skill packages, inspect included files and reference warnings, and install complete, immutable snapshots. Refresh each source manually. **Alternatives considered** Project repository settings hide the workflow from Skills. A second GitHub connector would duplicate credentials and grants. Upstream editing and PR creation are separate work. **Roadmap alignment** This implements the Skills Manager direction in ROADMAP.md. The maintainer requested this scope and reviewed the component and full-app journey stories before implementation. Related reports: Refs #10285, Refs #10949, Refs #13464. Related work: #14356, #13656, #9268. ## What Changed - Add source and entry records, an idempotent migration, company-scoped APIs, and legacy GitHub import adoption. - Reuse current caller grants and credential refresh. Combine and deduplicate repository inventories across accessible connections. Pasted public URLs also prefer the active user’s authorized connections. Tokens stay in the Git child environment, never argv or disk. - Fetch a shallow Git snapshot at one immutable commit. Scan the full local tree, including hidden and nested folders. Read Git objects without checkout or archive transformations and enforce nested package boundaries. - Bound Git downloads to 128 MiB and three minutes. Cancel active process groups and remove incomplete downloads. Preserve cancellation and deadlines while progress drains; close stalled HTTP progress streams after 30 seconds. Reuse caller-scoped temporary snapshots for preview/import after reauthorization. - Index repository package boundaries once and cap expanded work at 1,000 packages, 10,000 files, and 100 MiB, including repeated copies of shared blobs. Bound path depth and the shared path index. Discovery keeps audited manifests without retaining all package bodies. - Resolve moving branches before fetching so unchanged discovery reuses caller-scoped snapshots. Limit active scans, scan frequency, and new downloads per caller and company; quotas apply before metadata reads and across connections, and cached scans do not consume the download quota. - Store complete versions with script content, binary bytes, and executable modes. Preserve these through copies, runtime caches, and runner packaging. - Stage downloads before publication. Use source leases, revision checks, and transactional activity records. Keep installed versions after failures, upstream deletion, deselection, and disconnect. - Add the approved import flow, Sources page, selection tree, provenance, read-only Studio behavior, and saved return from GitHub setup. - Add package manifests, commit-pinned file previews, and separate runtime requirements and reference warnings. Supporting files are included together; nested skills remain independently selectable. Preview requests reauthorize the caller and re-audit package content. - Show installed skills as compact links beneath each source. Repository titles open GitHub. Keep Refresh, Select skills, and Disconnect source in a three-dot menu. Source rows omit the branch, imported count, and refresh timestamp; action alignment and repository titles work at narrow widths. - Stream discovery metadata over an opt-in NDJSON response. Show measured Git download progress and real package/file counts, animate newly checked skills, support cancellation, and require a complete scan before selection. Keep the existing JSON API. - Retain component stories and add a separate full-app journey story group. Include fixed progress states and interactive scan, large-repository, interruption, and saving stories. - Update Skills documentation and product contracts. Suppress private GitHub skill references in telemetry. Privacy review requested for the telemetry changes. ## Verification - Local repository typecheck, full build, token gates, and Storybook build passed during this work. Focused transport, authorization, scanner, persistence, route, and UI tests pass. The final UI refinement passes all eight focused UI tests, UI typecheck/build, and token gates. The scanner resource and repeated-discovery fixes pass 132 focused scanner, transport, authorization, source-service, route, and rate-limit tests, plus server typecheck/build. Full-suite verification comes from CI; the older full local Vitest run was stopped after unrelated chat failures and a font-test failure, all of which passed in fresh focused runs. At commit `1098d5996`, all 54 active checks pass; two optional Storybook jobs are skipped. CI covers repository typecheck, build, the full test suites, browser shards, and the canary dry run. Greptile is 5/5 with no open findings; the security scan also passes. - Adversarial scanner tests verify repeated-blob byte accounting with and without declared sizes, package/file/path caps, one-time repository indexing, metadata-only discovery audits, and nested package boundaries. Additional tests cover branch movement, snapshot reuse, caller/company quotas, isolation across connections, active-lease cleanup, quota recovery, and rejection before any metadata API call. - Real Git tests verify hidden paths, exact binary bytes, executable modes, export-ignore preservation, symlink/submodule reporting, pinned commits, caller-scoped cache reuse, cancellation, cleanup, and credential isolation. Regression tests hold both download slots with permanently blocked progress callbacks, verify timeout/cancellation cleanup and retry, and exercise HTTP backpressure cancellation. Access tests cover automatic public-URL connection selection and revoked grants. Database tests verify company and grant audiences. - Live isolated browser test: the public `anthropics/skills` scan now completes and discovers all 20 skills without connecting an account. Imported canvas-design with all 83 files, opened it from Sources, and verified the installed binary-font preview/download control. Package previews also expose the complete file inventory before import. Cancelled an active Git download and retried successfully to all 20 discovered skills; the browser displayed measured download progress. The current audits reject four other packages; eligible selections remain importable. - Browser checks verify the simplified source rows at desktop and narrow widths, keyboard navigation into the actions menu, Refresh from the menu, selection, and fixture disconnect with installed skills retained. Storybook includes a menu-open checkpoint and a 320px layout. - Storybook includes receiving/preparing download checkpoints and a timed full-app import journey, plus cancellation, retry, large-repository, and saving states. Streaming tests cover split UTF-8 frames, incomplete streams, late responses, cross-company requests, HTTP errors, and JSON compatibility. - Earlier live acceptance on this PR imported `stitch-skill` with `DESIGN.md`, assigned it to an agent, disconnected its source, and ran a successful Studio test that read both installed files. An editable copy changed independently. Both Skills variants, mobile selection, and return from GitHub setup were exercised. - Private access, revoked credentials, OAuth success return, binary/script preservation, concurrent refresh, transaction rollback, version pins, and legacy adoption have automated coverage. A real private-repository OAuth grant was not created during this test. ## Risks - The migration groups recognizable legacy imports without provider calls. Their first successful refresh completes the local package snapshot. - Reference checks are advisory. They cover Markdown links and explicit relative resource paths, not arbitrary runtime dependency graphs. Preview text is capped at 64 KiB; imported bytes remain complete. - Git must be installed on the server. Shallow fetches still download the branch snapshot, including files outside selected packages. Downloads have size/time/concurrency limits. Temporary caches are bounded and caller-scoped. GitHub API quota still applies to repository metadata and the connection picker; content no longer uses per-file API requests. Failed scans retain installed content. - Sources depend on the current caller's GitHub access. A saved connection does not grant access to another person's token. - GitHub script support and immediate manual refresh are explicit maintainer-approved requirements. The operator trusts the selected repository and accepts upstream script and executable-mode changes on refresh. Static audits are not a sandbox or a guarantee of safe code; agents may later invoke installed helpers under their runtime permissions. Import and refresh do not execute scripts, hooks, package installation, or builds. Raw URL and skills.sh imports keep their prior script restrictions. - Source originals remain read-only. Refresh affects subsequent unpinned runs; explicit pins and active runs retain their versions. - The telemetry change removes source-managed GitHub identifiers from skill-reference events. It introduces no event or field. Please review the privacy boundary. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d72389bee2 |
feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps gateway gives agents governed access to external tools. > - Browser Use Cloud can run browser work, but a tool result alone does not let a person watch or take over. > - A task needs a durable browser session, a visible viewer, and recorded costs. > - This pull request adds a Browser Use Cloud v4 connection and interactive browser tabs on tasks. > - People can follow the work, interact with the page, and retain the browser after the agent finishes. ## Linked Issues or Issue Description **Problem or motivation** Agents need governed access to Browser Use Cloud. People need to see and interact with the same browser from the task. A browser must remain available after a run finishes and appear at the correct point in the task feed. **Proposed solution** Add a native REST connection for the v4 API. Bind each session to its company, task, agent, and credential grant. Open its interactive viewer in the task side panel. Record provider costs as financial events. Use `browser-use-cloud` as the app and connector key. Keep its skill with the connector and deliver it only with authorized connection tools. **Alternatives considered** A v3 MCP connection would expose tools without the v4 lifecycle integration. An external viewer link would leave the task. A fixed viewer size would prevent pages from responding to changes in the task pane. **Roadmap alignment** This extends the governed Apps gateway and Connected Apps roadmap. It uses the existing task, grant, secret, approval, and financial records. The work was requested by the maintainer. A search found no duplicate Browser Use connector PR or issue. ## What Changed - Add the Browser Use Cloud app, brand asset, API-key connection, and profile settings under the `browser-use-cloud` key. - Bundle the `browser-use-cloud` skill with the connector. Keep it out of global `skills/` discovery. Deliver it only with authorized task/run connection tools. Remove retired connector skill keys from runtime overlays and preserve unrelated browser skills. - Expose seven v4 tools through the governed gateway and deliver them to native and CLI agents. - Persist sessions, browsers, runs, event and recovery cursors, shutdown leases, and cumulative cost accounting. Recover uncertain paid starts without replaying them. - Enforce task ownership, credential grants, approvals, revoked access, and budget limits. - Add interactive task browser tabs and compact chronological feed entries. Retain the viewer across tab switches and keep visible idle browsers open. - Add debounced automatic viewport fitting, standard size presets, and a viewer ownership lease. - Add lifecycle, authorization, accounting, viewport, UI, and Storybook coverage. - Add an idempotent database migration after the current master migration. Preserve deployed migration hashes. Migrate pre-release Cloud connection and financial keys without replacing grants, credentials, or browser history. - Document provider behavior, live acceptance results, and the lack of documented passkey forwarding. ## Verification - Full workspace typecheck and production build pass on the updated branch. - Token gates, brand asset validation, module boundaries, and migration ordering pass. - Cloud tests verify global skill exclusion, authorized task/run delivery, unassigned agents, disabled connections, revocation, adapter isolation, and secret exclusion. The existing AgentMail connector assignment test also passes. - Migration replay runs twice against existing browser work and financial records. It preserves the records and avoids duplicate costs. - The focused provider, app catalog, OpenAPI, connection gateway, and migration regression suites pass. Recovery coverage includes lost replies, process crashes, provider rejection, and browser arrival acknowledgement. - All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`, including the full test matrix, browser E2E shards, build, typecheck, security, and release canary. Two optional Storybook jobs are skipped. - Greptile is 5/5 on the same commit, with zero unresolved review threads. The corrected review uses the actual master-to-head diff. - The local `pnpm test:run` started and was stopped after the full CI matrix passed. It did not complete locally; the full-suite result above comes from CI. - Earlier live acceptance used an isolated company with a capped provider credential. The agent opened paperclip.ing, the embedded viewer accepted navigation, and the same browser stayed available after completion and tab switches. - The local Storybook build passes. Stories cover the panel, footer, feed entries, settings, lifecycle failures, and viewport modes with an offline viewer fixture. ## Risks - Browser Use charges for hosted work. Provider caps and local budget checks reduce exposure; reported costs can arrive after work completes. - Viewer and CDP URLs grant access to the browser. The server validates and restricts them. They are excluded from agent results and durable event data. - Runtime resizing of v4 agent browsers uses a provider option confirmed by live testing but absent from its published agent schema. Resizing during a click may invalidate coordinates. Fixed presets remain available. - Viewport ownership is process-local and resets on restart. The lifecycle and accounting records remain in the database. - The original intermittent embedded-viewer stall has not been fully diagnosed. A bounded reconnect and active-session recovery cover the observed failure paths. - Live tests did not cover every revocation, approval, rate-limit, or restart case. Deterministic integration tests cover those paths. Passkey forwarding is not claimed. - Unknown create outcomes keep the credential available for cleanup. Run-list absence cannot prove a paid POST was rejected, so recovery stays pending until it can identify provider work. ## Model Used OpenAI Codex, GPT-6. Used reasoning, repository search, code execution, browser interaction, and test tools. The exact serving model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f38b5693f6 |
fix: always enable keyboard shortcuts (#14643)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The web UI has keyboard shortcuts for the inbox, task lists, cases, and task detail, plus global shortcuts such as `c`, `/`, `?`, `[`, and `]` > - Shortcut enablement was an instance-wide General setting until #14141 moved it to a per-user preference that defaults to off > - The move did not carry the old instance value over, so every existing user lost shortcuts on upgrade and had to find a new toggle under Profile settings > - A toggle that only turns off a standard, input-safe feature costs a setting, a database column, two API routes, and a React context for little benefit > - This pull request removes both the instance setting and the personal preference and enables keyboard shortcuts for every signed-in user > - The benefit is one less thing to configure, no silent loss of shortcuts on upgrade, and less code to maintain ## Linked Issues or Issue Description Refs #14141 (the change that introduced the personal preference). **What existing behavior does this improve?** Keyboard shortcuts in the web UI stay off unless each user turns them on in Profile settings. **Subsystem affected** Web UI shortcuts, Profile settings, instance general settings, the `/api/auth/preferences` routes, and the `user` table. **Current behavior** Shortcuts default to off per user. #14141 moved the toggle from Instance settings → General to Profile settings and did not carry the old instance value over. Users who had shortcuts on lost them after the upgrade and had to find the new toggle. **Proposed behavior** Keyboard shortcuts are always enabled for every signed-in user. There is no instance setting and no personal preference. Shortcuts already ignore key presses inside text inputs and modal dialogs, so an opt-out is not needed. **Reason and benefit** Fewer settings, no silent loss of shortcuts on upgrade, and removal of a database column, two API routes, a query hook, and a React context that existed only to gate this feature. **Breaking changes** `GET` and `PATCH /api/auth/preferences` are removed. `PATCH /api/instance/settings/general` no longer accepts `keyboardShortcuts`; that schema is strict, so the key now returns 400. `instance.general.keyboardShortcuts` is no longer a valid `PAPERCLIP_HIDDEN_SETTINGS` key; the parser ignores unknown keys with a warning. ## What Changed - Removed the Keyboard shortcuts section from Profile settings, the `useUserPreferences` hook, `queryKeys.auth.preferences`, and `authApi.getPreferences` / `authApi.updatePreferences`. - Removed `GeneralSettingsContext`. The inbox, legacy inbox, task list, legacy task list, cases, and task detail pages no longer gate their key handlers. - Removed the `enabled` option from `useKeyboardShortcuts`. The app shell always registers the global shortcuts. - Removed `GET` and `PATCH /api/auth/preferences`, their OpenAPI entries, and the `currentUserPreferencesSchema` / `updateCurrentUserPreferencesSchema` validators. - Removed `keyboardShortcuts` from `InstanceGeneralSettings`, the general settings zod schema, the settings service defaults, and `HIDEABLE_GENERAL_SECTIONS`. - Added migration `0289_drop_user_keyboard_shortcuts`, which drops `user.keyboard_shortcuts`. - Updated `AGENTS.md`, `doc/SPEC.md`, `doc/SPEC-implementation.md`, and `docs/deploy/environment-variables.md`. - Parsed the stored general settings row with `instanceGeneralSettingsSchema.strip()` in the feedback vote path, so a retired key left in the row cannot reset the sharing preference to `prompt` and overwrite the stored choice. - Kept every bare global shortcut (`c`, `?`, `[`, `]`, `/`) out of open modal dialogs in `useKeyboardShortcuts`; only `/` had that guard before. - Updated the affected tests and added a Profile settings test that asserts the toggle is gone, a hook test for the modal dialog guard, and a feedback service regression test for the retired-key case. ## Verification - Typecheck passes for `@paperclipai/shared`, `@paperclipai/db` (including the migration numbering and safety checks), `@paperclipai/server`, and `ui`. - `pnpm exec vitest run server/src/__tests__/instance-settings-routes.test.ts server/src/__tests__/openapi-routes.test.ts server/src/__tests__/auth-routes.test.ts server/src/__tests__/sentry.test.ts` → 119 passed. - `pnpm exec vitest run ui/src/components/Layout.test.tsx ui/src/pages/ProfileSettings.test.tsx ui/src/pages/IssueDetail.test.tsx ui/src/pages/Inbox.test.tsx ui/src/pages/Cases.test.tsx ui/src/hooks/useKeyboardShortcuts.test.tsx ui/src/pages/Agents.test.tsx ui/src/pages/InstanceGeneralSettings.test.tsx` → 286 passed. - `pnpm exec vitest run packages/shared/src/settings-visibility.test.ts` → 16 passed. - `pnpm exec vitest run ui/src/hooks/useKeyboardShortcuts.test.tsx` → 7 passed. - `pnpm exec vitest run server/src/__tests__/feedback-service.test.ts` (embedded Postgres) → the new retired-key test passes with the fix and fails without it. - Manual: sign in with no settings changed, open the inbox, press `j` and `k` to move the selection, press `?` to open the cheatsheet. Open Settings → Profile and confirm there is no Keyboard shortcuts section. ## Risks - The migration drops a column. It uses `DROP COLUMN IF EXISTS`, and the column has no readers after this change. If you roll back to a build from before this PR after the migration has run, re-add the column first: `ALTER TABLE "user" ADD COLUMN "keyboard_shortcuts" boolean DEFAULT false NOT NULL;`. The older build's ORM selects that column when it loads users. - Any external client that still sends `keyboardShortcuts` to `PATCH /api/instance/settings/general` receives a 400. No in-repo client does. - Stored `instance_settings.general.keyboardShortcuts` values are stripped on read and ignored. - Users who never turned the toggle on now get shortcuts. The handlers skip text inputs, contenteditable regions, and modal dialogs, so typing is unaffected. ## Model Used Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
83076d7e7c |
feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage. Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3ca196b0a6 |
feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - An agent needs personal files across tasks and sessions. > - AGENTS.md is one file in that directory. Supporting files need the same persistence. > - The Instructions Editor and agent runs must share one current directory. > - Concurrent runs should apply only the files they change. The last sync of the same file wins. > - This pull request uses existing file transport and removes temporary copies after sync. > - Old instruction-only sessions keep their restore contract. New saves do not create revision history. ## Linked Issues or Issue Description Refs #14325. This replaces its revision-oriented design with persistent agent files. Keep #14325 unmerged. Transport prerequisite #14416 merged first at `d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master and remains below 100 changed files. Related work: #4513 and #8798 cover instruction tooling. This change handles run synchronization, cross-task personal files, browser editing, and old-session restoration. ## What Changed - Keep one current directory per company and agent. Point AGENT_HOME at a temporary working copy for each active run. Keep task files and provider HOME separate. - Restore text, binary files, and nested folders through workspace transport. Exclude remote agent files from task Git snapshots with a self-ignoring file inside the reserved runtime directory; never write through repository-controlled Git metadata. - Collect after the provider and child processes have stopped. Keep resumable conversation state. - Apply changed and deleted files under the agent lock. The last sync wins for the same file. Unrelated concurrent changes survive. - Remove temporary copies after successful sync, rejected sync, and staging failure. Register ownership before copying so restart recovery can remove interrupted preparation. Retry transient synchronization up to three times. Preserve the original remote lease reference until deletion succeeds; restart cleanup never acquires a replacement sandbox. Do not create captured directories or a conflict-review queue for new runs. - Keep browser editing, stale-draft protection, and streaming binary downloads. Keep the instruction entry and text editor limited to 1 MiB. - Keep historical agent-folder sync failures on their affected runs instead of repeating them above current saved instructions. Preserve legacy candidate review and current browser-save errors. Avoid duplicate quota warnings while retaining separate sync failures when they describe a different problem. - Require target-scoped caller grants for peer instruction access, while preserving self edits, responsible-user checks, and protected-change consent. - Treat full storage as a nonblocking run warning, never an agent pause or run-admission failure. Restore already-over-quota saved folders so ordinary agent cleanup can recover; warn on each run until cleanup. The run detail view shows the warning. - Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash large files as streams. Check editor-save quotas with metadata instead of hashing unrelated files. - Preserve old native inputs, instruction-only copies, paths, digests, and pending legacy candidates. Adopt old revision heads once. New writes do not append history rows. - Add idempotent migration 0287 and verify upgrades from the preview tables and receipts. - Add nine interactive stories under **Agents / Persistent files**, including automatic incoming edits, stale browser drafts, and storage-limit diagnostics. ## Verification - Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after merging current master and the landed transport prerequisite. Integration required no manual conflict resolution; the feature remains 99 changed files. Full workspace typecheck, production build, token gates, and 715 focused tests passed on this merge candidate. Fresh Greptile review is 5/5 with no unresolved findings. All 55 checks passed, with four conditional skips, including the build, typecheck, browser E2E, and canary dry run. A single retry recovered four jobs interrupted by runner shutdowns; no source changes were required. - Historical-warning UI fix: all 6,834 UI tests across 640 files passed, including regression coverage for three old failures, legacy preserved edits, and warnings scoped to the affected run. Full workspace typecheck, production build, Storybook build, and token gates passed. Browser-verified Storybook playtests passed for Historical Failures After Successful Save, Storage Limit, and Full Storage Run Warning. - Review follow-ups at `4e20c9fb2`: all 18 focused tests passed, including external Git directories, linked worktrees, symlinks, hardlinks, and distinct I/O failures alongside storage warnings. Server and UI typechecks, token gates, and the production build passed. - Storage warning regressions at `0724f3012`: all 33 directory tests and all five heartbeat-list tests passed, with no skips in their successful runs. They cover repeated runs while full, an already-over-quota saved folder, cleanup, warnings retained after unrelated save failures, and bounded warnings in large result JSON. Server typecheck passed after the final warning fixes. - Full workspace typecheck, production build, and token gates passed during this follow-up. Product E2E harness: 631 tests passed across 52 files; harness typecheck passed. Earlier native session/context and directory/legacy collection suites passed 537 tests; Runner unit/transport suites passed 329 tests. - **Real E2E at `0724f3012` (before this follow-up):** legacy local Codex and native Daytona Codex each passed six tasks, one server restart, seven independent assertions, and cleanup verification. Both prove browser-to-agent edits, agent-to-browser edits, nested/binary restoration, per-file last-sync-wins, a successful run after an oversized save rejection, and cleanup clearing the warning. - Native local Codex also passed the six-task quota flow before the final warning-retention fixes. That pass began at `918d1ed02` while the bounded-result warning fix was being edited, so it is not claimed as exact-final-head evidence. Its final-head rerun failed during embedded PostgreSQL bootstrap before any provider run: the macOS host had 87,365 of 87,381 SysV semaphores occupied. No unrelated services or kernel limits were changed. - The final-source report intentionally records **2/3 cells passed**, preserving the blocked native-local attempt: `tests/runner-e2e/results/agent-files-quota-final-20260928-report/`. Earlier failed attempts and provenance notes remain under `tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and the original campaign directories. - Daytona used immutable image `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf` and its exact Linux runner binary. Controller source is `0724f3012`; image source is recorded separately. - Legacy-session compatibility and all three ACP Stop/resume browser regressions passed on the prior validated feature head `169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same provider session is retained and interrupted writes are not replayed. Migration upgrade tests also passed earlier. - Nine interactive stories are under **Agents / Persistent files**, including **Full Storage Run Warning**. Its playtest and visual browser inspection passed; the warning states that runs continue and the editor remains available. - Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs skipped, no failures or pending checks. All eight browser E2E shards and their aggregate passed. Fresh Greptile review is 5/5 with no findings; all review threads are resolved, the security scan passed, and GitHub reports no merge conflicts. - The broad local follow-up test run was interrupted after host semaphore exhaustion affected isolated PostgreSQL instances. It also encountered the existing macOS long-path fixture failure and two timeout failures. This is not a claim that the full local suite passed. Logs are retained; focused storage/warning tests passed. ## Risks - A later sync can overwrite an earlier edit to the same file, including a saved browser edit. There is no text merge or retained version. This is the intended last-sync-wins policy. - A save that exceeds a storage limit is rejected and its temporary copy is discarded. The run itself continues normally, and later runs restore the last saved files with a warning until cleanup. Transient sync failures get bounded retries. An I/O failure partway through a sync can leave some files updated; a failed receipt does not claim whole-folder success. - Larger folders increase copy time, network traffic, and temporary disk usage. Active runs still need working copies. Terminal runs do not accumulate archives. Operators must provision disk for agents and configured concurrency; these limits are not company-wide quotas. - A restored old native session remains instruction-only until a fresh session starts. Its original conflict fence and existing pending candidates remain compatible. - Provider processes close at the collection boundary. Conversation resume remains available, but warm process reuse is lost. - Backups must include the instance filesystem and database. External bundles keep their existing behavior until explicitly moved to managed storage. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and browser tools assisted this change. Real provider E2E uses `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing> |
||
|
|
14795136f5 |
fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider. Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3447609d22 |
fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.
## Linked Issues or Issue Description
**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.
**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.
**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.
Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.
## What Changed
- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.
## Verification
- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.
## Risks
- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.
## Model Used
OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
01d9a12185 |
fix: make keyboard shortcut enablement a personal preference (#14141)
Store keyboard shortcut enablement per user and expose it in Profile settings. Co-Authored-By: Codie <Codie@users.noreply.github.com> |
||
|
|
a10702a878 |
feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people start and continue agent tasks from other services. > - A Slack conversation needs access to its surrounding discussion and Slack collaboration tools. > - The agent must use the linked requester's access and keep private material within its permitted audience. > - This pull request adds Slack tools through the existing connector contribution and approval framework. > - People can ask an invited bot to read a discussion, create follow-up tasks, and collaborate in Slack. ## Linked Issues or Issue Description **Subsystem affected** Chat connectors, connector runtime, tool gateway, and connection Settings/Access. **Problem or motivation** Slack-origin tasks can receive messages but cannot inspect the rest of a channel or act through the originating bot. People must paste context or configure a separate integration. **Proposed solution** Supply typed Slack tools and a bundled skill only to the originating task and assigned agent. Resolve the linked requester on the server. Check bot and requester access before reads and writes. Use existing durable actions and approvals. Retrieved messages remain source material. **Alternatives considered** Slack's user-OAuth MCP server does not replace the customer-created chat bot. An unrestricted Web API proxy would not provide suitable permission or publication boundaries. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps work with a provider contribution. It does not add a task dispatcher or a separate Slack task lifecycle. Related: #11144 covers generic per-user MCP grant execution; this change binds Slack bot operations to chat-origin tasks. ## What Changed - Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and shared native/HTTP execution. - Bind tools to company, endpoint, task, run, assigned agent, and admitted linked requester. Check membership and revocation on each call and before queued writes. - Add paginated reads, bounded history search, source links, messages, file uploads, reactions, pins, bookmarks, topics, canvases, lists, and approved channel operations. - Restrict private-source publication, including automatic replies and uploaded deliverables. Keep other people's bot DMs inaccessible. - Reuse action receipts, idempotency, approvals, and reconciliation. Suppress an identical explicit-send/final-reply duplicate. Return governed results through their verified originating conversation. - Add endpoint-bound personal search OAuth storage and lifecycle. Keep native real-time search disabled until a runtime meets Slack's transient-result requirements. Current runtimes use bounded history search. - Show capabilities, scope upgrades, and personal search authorization in Settings/Access and Storybook. Document provider and runtime limits. ## Verification - Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no unresolved findings. One unchanged rapid-callback timing test passed on a single CI retry. - Approval presentation regressions cover board-comment precedence and exact Slack publication; the expanded database assertion passed in CI. The local PostgreSQL startup probe later became unavailable, so that final assertion was verified in CI. Slack setup and failed-run retry browser tests also passed locally. - Full workspace typecheck and build passed. Server typecheck/build passed again after the approval routing fix. - Broad local suites passed in separate groups: server 12,958 tests, UI 6,555, shared 770, skills catalog 20, and other workspace packages 2,652. CLI and serialized server checks passed after environment/timeout retries. These are composite results, not one uninterrupted green full-suite invocation. - PostgreSQL authority regression covers admitted identity, cross-company/task/agent rejection, recovery, retained-session revocation, OAuth refresh/disconnect races, approval execution, exact publication lineage, retries, uncertain sends, and duplicate suppression. - Gateway/response regressions cover separate-origin approval batches and durable continuation. Focused provider, access, search, native runtime, route, and AgentMail regressions pass. - Storybook capability, missing-scope, OAuth configuration, authorization, and disconnect states were inspected in the browser. - Live staging: read a channel decision and full thread, create exactly two assigned backlog tasks, add a reaction, paginate discovery to exhaustion, and return bounded search matches with source links and coverage. - Live staging: create/edit/read a canvas and list, inspect the canvas in Slack, post/edit one message, and create a channel only after approval. New channels remain disabled for responses. - Live staging: read a response-disabled channel from the requester's DM; writes to that channel were denied. The test setting was restored. - Final live retest passed: explicit file upload and exact content read-back; approved deletion of only the disposable bot message; continuation confirmation returned to the original Slack thread without repeating the action. - Optional OAuth, private multi-user boundaries, native RTS, and CLI provider execution are not fully live-qualified. The staging agent initially supplied malformed tool arguments; valid arguments succeeded, and the tool/skill descriptions now emphasize UUID write keys. ## Risks - Existing Slack apps must add scopes and reinstall for new capabilities. Provider plans and document permissions can still restrict operations. - Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET` for governed tool actions. The staging instance was configured with explicit operator approval; fleet provisioning is a separate gap. - Native RTS is not exposed on current transcript-retaining runtimes. Bounded history scans are deliberately reported as incomplete. Inline file reads support text/canvas content up to 256 KiB; other types return metadata. - Private document edits fail closed when the full audience cannot be verified. Uncertain effects other than posts/uploads require inspection instead of blind retries. - Shared approval-delivery code now separates outcomes by source run to preserve origin boundaries. No database migration is required. - A separate completion-validator gap remains when the agent cites a prior run's registered artifact during finalization. It asked for registration again even though Slack delivery was confirmed. This change does not add a connector-specific task-completion policy. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and browser testing. The exact deployed model identifier and context-window size were not exposed in the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8813a50105 |
feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path > - Paperclip manages agent work as tasks and runs. > - GitHub chat brings repository conversations into those tasks. > - A review bot needs the assigned agent, its authority, and governed provider tools. > - The existing channel connection did not supply that review workflow or a complete setup journey. > - This pull request adds GitHub App setup, account access, event prompts, task-bound review tools, and exact-commit checks. > - Operators can inspect each review through the same task, run, and activity systems. ## Linked Issues or Issue Description **Subsystem affected** GitHub chat, governed connection tools, task execution, shared/database contracts, and connector setup UI. **Problem or motivation** Operators need a GitHub review bot that runs their assigned Paperclip agent. Mentions and PR events must preserve task ownership and requester authority. Provider publication must use the bot App identity and enforce the configured permissions. **Proposed solution** Extend the existing GitHub chat connector with resumable App onboarding, linked-member and sponsored-guest access, editable event prompts, and governed review operations. Validate structured assessments on the server and compute a stable Paperclip Review check for the exact head commit. **Alternatives considered** A separate review scheduler would duplicate Paperclip execution and permissions. Reusing personal GitHub credentials would change the bot identity and credential boundary. **Roadmap alignment** This extends the existing Connected Apps and governed-tool infrastructure. The project owner requested and approved this design. Related PR #8645 imports external Codex review feedback; this change runs an assigned Paperclip agent and publishes its results through the existing chat connector. ## What Changed - Include the current Paperclip instance origin in the copied setup prompt. Storybook uses its configured Paperclip origin; callback parameters and URL credentials are excluded. - Add a Claude/Codex copy button in the real setup and Storybook opening step. Its detailed prompt asks four setup questions and guides embedded-browser setup, verification, and optional required checks. Clipboard failure exposes selectable instructions. - Add a tutorial that explains why App installation, review scheduling, and required checks are separate choices. - Add manifest registration, an existing-App path, separate installation and repository selection, repository refresh, and explicit account confirmation. - Add low-trust agent guidance, effective capability verification, member selection, and explicit restricted guests with a sponsor. - Add configurable PR events, prompts, repository overrides, rating thresholds, and separate formal-review permissions. - Give the assigned agent governed App tools to read PRs, comment, begin an assessment, submit findings, and optionally submit a formal review. - Bind review history, root PR events, and inline replies to ordinary tasks. Deduplicate deliveries/findings and reject stale publication. - Link check Details to the underlying task on the current trusted hostname, or to Reviews before task creation. - Add schema migration 0283, API contracts, production UI, and 49 interactive Storybook states. - Repair local lease recovery. Keep the Cloud Dockerfile identical to master; no provider-pack layer or runtime-default environment variable is added. - Retry only rolled-back wake-admission transactions after transient endpoint-lock contention. A deterministic held-lock regression proves one accepted wake. ## Verification - Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero diff against master. Final workspace typecheck and build passed. The new PostgreSQL migration regression passed and preserves existing relation and constraint identities after replay. - Greptile reviewed this exact head at 5/5. There are zero unresolved review threads and no merge conflicts. - All current-head checks are green: 54 passed and two conditional Storybook jobs skipped. This includes complete server/workspace test suites, build, typechecks, policy checks, Runner suites, browser suites, and security status. One timing-sensitive callback-ordering test passed in isolation and its CI shard passed one retry. The duplicate local full-suite run was stopped after CI completed; it is not counted as a local full-suite pass. - Before the final Slack rebase and migration renumbering, 186 focused GitHub tests, 14 native bootstrap cases, token gates, and Storybook build passed. The final rebase retained the new Slack communication guidance. - The embedded-browser setup test copied the full detailed prompt, including the configured Paperclip instance URL. Desktop and narrow layouts were checked. Component tests cover successful copying and clipboard failure with selectable text and retry. - Live local and hosted GitHub acceptance evidence refers to application revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks exercised issue mentions, automatic PR reviews, inline findings, repeated mentions, task continuation, and failing-to-passing checks after a push. The Storybook agent generated, built, and browser-rendered pages; missing acceptance text failed, matching text passed, and broken JSX produced an incomplete result. - Live cases also covered independently disabled push events, prompt injection, duplicate signed deliveries, rapid pushes, stale-result rejection, finding deduplication, and restart recovery. Formal reviews were denied while disabled and published only after explicit enablement. Check Details links pointed to the underlying task on the trusted hostname. - Those hosted native Claude runs used the provider-pack layer now removed from this PR. They do not prove native Claude works on the standard Cloud image. A replacement hosted native Codex run is not yet verified: the disposable QA tenant has only an Anthropic AI connection. No new staging or production deployment was made for the packaging removal. - Required-check merge enforcement could not be tested because the private disposable repository's GitHub plan rejected the rules configuration. Published success/failure/incomplete check states were verified directly. ## Risks - Latest master allocated migration 0282 to Slack. The GitHub migration is regenerated as 0283 with replay-safe table/index/constraint creation; a PostgreSQL regression verifies existing relations and constraints are preserved. Existing preview tenants remain subject to the fleet migration-history compatibility preflight; no bypass is introduced. - Migration 0283 adds company-scoped configuration, registration, review, and publication records. Existing connections retain their behavior until reviews/tools are enabled. - Signed webhooks and expiring registration state remain required. Hosted installations also need the companion narrow Cloud gateway exemptions. - Agent assessments can be incomplete or wrong. The server enforces coverage/result structure, current-head publication, rating policy, and separate formal-review permission; it does not replace code-review judgment. - No Cloud image packaging changes are included. Remote native ACPX/Claude and OpenCode retain their existing operator-supplied provider-pack prerequisite. Native Codex and Codex with managed MCP tools do not require that pack. Earlier staging deployment evidence refers to its stated revision, not this packaging-removal head. Production rollout and merging remain outside this change. ## Model Used OpenAI GPT-6 through Codex, with repository, code execution, API, and embedded-browser tools. The exact serving model ID and context-window size were not exposed by the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d9b3a5653e |
feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people use the same tasks and agent tools from external conversations. > - Agents need communication guidance that fits the conversation medium. > - That guidance belongs in the original task context, without repeated instructions on each turn. > - Connection owners also need clear settings and a consistent way to remove a connection. > - This pull request adds initial Slack guidance, optional connection instructions, and chat connection menus. > - The benefit is clearer Slack replies with the existing Paperclip workflow and permissions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent replies in Slack and chat connection management in the Apps catalog. **Current behavior** Slack tasks do not carry a saved communication profile. The catalog shows a separate Manage button and does not offer removal on every chat connection row. **Proposed behavior** Save Slack guidance when a new conversation creates a task. Restore that original guidance when a model session is rebuilt. Do not append it to ordinary follow-ups. Expose optional additional instructions in Slack Settings. Put Manage and Remove connection in a three-dot menu for all chat providers. Keep Finish setup visible for drafts. **Reason and benefit** Small answers fit in Slack. Substantial deliverables use ordinary document or artifact tools with a useful Slack summary. Connection settings apply to new tasks and cannot change permissions. Users can remove both active and unfinished chat connections from the catalog. **Breaking changes** Two additive database columns store endpoint preferences and the initial conversation snapshot. Existing endpoints default to empty preferences. Existing conversations keep their original behavior. Non-Slack guidance is unchanged. Related public context: https://github.com/paperclipai/paperclip/pull/13741 improves native chat recovery. This change adds communication context to those existing execution paths. A search found no duplicate communication-guidance PR. ## What Changed - Add a provider-guidance registry, enabled for Slack first. - Persist optional endpoint communication instructions and capture an immutable snapshot when a conversation creates a task. - Resolve guidance from the verified company-scoped connection. Restore it for fresh native and legacy sessions without per-turn reminders, extra model calls, or extra context queries. - Add the Slack Settings field, validation, audit coverage, and Storybook save/error states. - Add Manage and Remove connection menus for all seven chat providers. Keep the draft setup button. Require removal confirmation and allow retry after failure. - Add regression coverage, an active/draft menu story, and connector documentation. ## Verification All CI checks are green for |
||
|
|
aeef493f4a |
chore(db): keep only the newest 5 drizzle snapshots and stop shipping them (#13687)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - `@paperclipai/db` owns the Drizzle schema and the migration history > - Drizzle writes a full copy of the schema as a snapshot for each generated migration. Each snapshot is now about 1.3 MB. > - The `meta/` folder is about 103 MB. That is about half of each checkout and each worktree. The build also copies it into `dist`, so the published `@paperclipai/db` package is 112.8 MB unpacked. > - `drizzle-kit generate` reads only the newest snapshot. The runtime migrator reads only the `.sql` files and `_journal.json`. > - This pull request keeps the newest 5 snapshots and removes snapshots from `dist`. > - The benefit is a checkout that is about 100 MB smaller, and a published package that is about 1.5 MB instead of 113 MB. ## Linked Issues or Issue Description Refs #11240, #11254, #12333 (earlier snapshot work: diff collapse, binary diffs, drift repair) **What existing behavior does this improve?** The size of the Drizzle migration snapshots in the repository and in the published `@paperclipai/db` package. **Subsystem affected** `packages/db`: migrations and the build. **Current behavior** `packages/db/src/migrations/meta/` holds 141 snapshots (102.8 MB). The size grows faster than the number of migrations, because each snapshot is a full copy of the schema. `build` runs `cp -r src/migrations dist/migrations`. `@paperclipai/db@2026.916.0` contains 147 snapshot files. It is 112.8 MB unpacked and 5.3 MB as a tarball. **Proposed behavior** Keep the newest 5 snapshots. `generate` deletes older snapshots after it runs. `dist` gets only `*.sql` and `meta/_journal.json`. **Reason and benefit** - In drizzle-kit 0.31.10, `generate` sorts `meta/*` and diffs against the last snapshot only (`bin.cjs`, `preparePrevSnapshot`). The [generate docs](https://orm.drizzle.team/docs/drizzle-kit-generate) also say it compares against "the most recent" snapshot. - Gaps in the snapshot history already work. 138 of the 280 migrations never had a snapshot, because they were written by hand before `doc/DATABASE.md` required `generate`. - We keep 5 snapshots instead of 1. This lets a developer undo the latest generated migration, and it keeps the `prevId` chain for recent branches. - Git keeps the history cheaply. The 631 snapshot versions use only 2.2 MB of the pack, because git stores each version as a delta of the previous one. The cost is in the checked-out files, not the clone download. Therefore this change does not use git-lfs and does not rewrite history. Old snapshots stay available with `git show <rev>:<path>`. ## What Changed - `packages/db/package.json`: a new `prune:snapshots` script keeps the newest 5 `*_snapshot.json` files. It is `ls | sort -r | tail -n +6 | xargs rm -f`, which works with the BSD tools on macOS and the GNU tools on Linux. `generate` runs this script after `drizzle-kit generate`. - `packages/db/package.json`: `build` copies only `src/migrations/*.sql` and `meta/_journal.json` into `dist/migrations`. - Deleted 137 older snapshots. `0277`–`0281` remain. - `chat-identity-migration-reconciliation.test.ts`: removed the walk over snapshots `0254`–`0268`. Those files do not change after merge, and the walk would fail after pruning. The journal-order assertions in the same test remain. `migration-snapshot-drift.test.ts` still makes sure that the newest snapshot matches the schema. - `doc/DATABASE.md`: documented the retention rule. - The snapshots were already marked `linguist-generated=true -diff -merge` by `packages/db/.gitattributes` (#11240). No change there. `git check-attr` confirms it. ## Verification - `pnpm --filter @paperclipai/db exec vitest run`: 43 files, 160 tests pass. - `pnpm --filter @paperclipai/db build`: `dist/migrations` contains 280 `.sql` files and `meta/_journal.json`. It is 1.5 MB, compared with about 105 MB before. - `prune:snapshots` was run on macOS (BSD) and in `debian:stable-slim` (GNU findutils 4.10). With 7 fixture snapshots, it keeps the newest 5. When 5 or fewer are present, it deletes nothing and exits 0 on both. ## Risks - Low risk. Runtime migration does not read snapshots. The only commands that read older snapshots are `drizzle-kit check` and `drizzle-kit drop`. No script or CI job calls them, and they still have the newest 5 snapshots. - A branch that is open now can still add its own snapshot. If the branch conflicts, the rule is the same as today: renumber the migration and run `generate` again. - When we upgrade to drizzle-kit v1 (folder per migration), check whether its new cross-branch "commutativity" checks need a longer snapshot history. ## Model Used - Claude Opus 5 (`claude-opus-5`) in Claude Code, with tool use (shell, file edits, web fetch). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f589660ec0 |
feat(routines): add safe webhook setup and in-routine run management (#13637)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Routines turn scheduled work and external events into tasks for an assigned agent. > - Webhook setup was disabled, and actor authentication rejected valid webhook bearer keys. > - Operators need to connect and test a sending app before events can start work. > - This pull request adds a guided setup with durable connection tests that cannot dispatch a task. > - It keeps trigger management, execution tasks, and activity within the routine. > - The benefit is a webhook that can be configured, verified, and operated from one place. ## Linked Issues or Issue Description Fixes #11937. Related: #13216 adds provider-specific Sentry support. This PR addresses general routine setup and ingress. #6841 addresses legacy secret bindings; this PR retains the existing secret service. **Current behavior** Webhook creation is disabled. Bearer deliveries can fail in agent authentication before the routine checks its key. Setup has no safe connection test. Runs and Activity send the operator away from the routine. **Proposed behavior** Choose a schedule or a webhook. Follow the setup steps, copy credentials or complete agent instructions, and test delivery without creating work. Finish setup to allow future events to start tasks. Edit or remove compact trigger cards, undo removal, and inspect tasks and activity inside the routine. **Reason and benefit** An operator can verify credentials and delivery before enabling automatic work. Durable setup state survives refreshes and restarts. Retry receipts prevent an old test event from starting work after activation. ## What Changed - Add a production trigger wizard using reusable Slack setup navigation and footer components. - Add schedule and webhook choices, one-time credentials, agent instructions, and live connection feedback. - Persist pending setup, test delivery receipts, connection status, and reversible trigger removal. - Keep setup checks free of routine runs, tasks, and agent wakeups. Preserve delivery idempotency after activation. - Add compact trigger cards, inline editing, key rotation, pause controls, removal, and Undo. - Keep Runs and Activity in the routine. Use the shared task list and compact activity rows. - Permit only exact public delivery POSTs through actor authentication. Retain webhook authentication, JSON-object validation, and log redaction. - Add production-backed Storybook states and focused server, database, and UI coverage. - Document signing modes, setup checks, retries, rotation, HTTPS ingress, and navigation. ## Verification - Full workspace typecheck, build, and token gates passed on the rebased branch. Storybook also builds. - Focused routine, middleware, logging, shared wizard, and UI coverage passes on the rebased branch: 195 tests across 14 files. The migration passed on a fresh PostgreSQL database and on two repeated applications. - Browser testing used the real app, database, and a deterministic process worker through Tailscale HTTPS and the current Cloud proxy code. - Verified rejected keys, safe setup deliveries, persisted state after restart, activation, retry deduplication, key rotation, schedule editing, removal, and Undo. - Fresh bearer and GitHub-signed deliveries created tasks that the worker checked out and completed. Runs and Activity stayed within the routine. - Current Cloud ingress tests passed. Public delivery POSTs passed through without a browser session; management routes remained gated. - All 54 current-head PR checks pass, including general and serialized tests, all eight browser E2E shards, typecheck, build, runner checks, security checks, and the canary dry run. Two optional Storybook jobs are skipped by workflow conditions. - Greptile is 5/5 on commit `7ea63a61e`, with no unresolved review threads. The stale connection-status finding is fixed and covered by a regression test. - No production deployment was performed. ## Risks - Migration 0281 adds three trigger columns and a test-receipt table. It is additive and safe to reapply. Apply it before running the new server. Existing triggers remain live by default. - Requests without delivery IDs are new events after activation. Senders must reuse an event's delivery ID for retries. - Completed webhooks keep normal dispatch behavior. Their management connection check can start work; the UI states this. - Removing a trigger archives it. Undo restores the URL and credentials. Permanent deletion remains available through the existing API. - Public ingress must remain restricted to the delivery POST route. The tenant verifies credentials. Cloud sleeping-stack behavior is unchanged. - Shared setup components also serve Slack. Existing setup contracts and navigation tests cover that integration. - Senders must use application/json with an object. Other media types receive 415. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1ef3b08714 |
feat(ui): integrate agent personas across the app (#13171)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A stable agent persona is useful only when the same identity appears across the app. > - Lists, task messages, selectors, and activity feeds need inexpensive static avatars. > - Onboarding and agent headers need a larger character with expressions and pointer tracking. > - This pull request connects the persona foundation to those existing views and preserves onboarding draft assignments. > - Full-page stories and Linux checks make the placements and performance contract reviewable. ## Linked Issues or Issue Description **Problem or motivation** Agents need a stable visual identity in lists, tasks, onboarding, and configuration. External tools also need an image URL for that identity. **Proposed solution** Assign each agent a permanent palette from a fixed ClipLab character library. Store the assignment on the agent. Render and cache preset PNG URLs on demand. Use static images in dense views and one animated character in larger placements. **Alternatives considered** A generated image bundle requires a separate asset build. A live renderer in every avatar adds unnecessary work in large lists. Arbitrary uploaded images do not provide the requested shared character system. **Roadmap alignment** This improves agent identity across existing control-plane views. It preserves agent permissions, company boundaries, and status labels. ROADMAP.md has no separate ClipLab persona milestone. Related approaches: #2422 adds configurable image URLs and DiceBear generation; #5578 adds optional uploaded avatars. This work uses a fixed, versioned character library and preset URLs. ## What Changed - Replace agent icons with static persona images across lists, the sidebar, org charts, tasks, comments, selectors, activity, and dashboard views. - Put one animated character in the agent header. Let it follow the pointer across the page, with reduced-motion and touch fallbacks. - Add larger padded characters to agent creation. Keep the palette stable across draft refreshes and connection retries, then reveal it after success. - Pass appearance through shared projections rather than fetching each agent separately. - Add real full-page Storybook examples for the agent list, overview, task, dashboard, new-agent dialog, and connection page. - Add Linux screenshot, clipping, density, and 500-avatar performance checks. ## Verification - `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased tree. Persona lifecycle tests pass. - The rebased feature passes 38 Linux screenshot/performance checks, including both display densities, corner pointer positions, and the no-WebGL/no-live-download contract for 500 avatars. - The final Linux persona suite passes all 38 visual, lifecycle, density, and full-page checks using the standard Storybook configuration and real on-demand avatar endpoint. - Final local focused verification: 45 avatar/native-recovery tests pass; UI identity/routine tests, typecheck/build, token gates, and Storybook build pass. - Current-head CI passes: full workspace/server tests, all serialized server groups, typecheck/release checks, build, canary validation, and end-to-end shards. The build passed after retrying a native-runner concurrency-test failure; its three targeted cases also pass locally. - Manual inspection covered stable identities in the app, header placement, full-page mouse tracking, onboarding size, and task/dashboard placements. ### Screenshots Linux captures use synthetic Storybook fixtures. Full-page captures use reduced motion. The live character, mouse tracking, and disposal are checked separately. <details> <summary>Agent overview with the character in its header</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png" width="900" alt="Agent overview with the character in its header" /> </details> <details> <summary>Task messages and assignee identity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png" width="900" alt="Task messages and assignee identity" /> </details> <details> <summary>Larger onboarding character with room for expressions</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png" width="900" alt="Larger onboarding character with room for expressions" /> </details> <details> <summary>Dashboard agent activity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png" width="900" alt="Dashboard agent activity" /> </details> ## Risks - This PR depends on #13170, the persona foundation. Merge the foundation first, then retarget this PR to master. - Many placements change from icons to character silhouettes. Human avatars and authoritative agent status labels retain their existing behavior. - Only one character can render live per view. Reduced motion, hidden/offscreen content, touch input, and renderer failures use the defined fallbacks. - The full-page stories use fixture data. They do not contact a real company or complete real provider sign-in. ## Model Used OpenAI Codex, GPT-6 family. The exact model identifier and context window are not exposed in this session. Used code editing, shell execution, browser inspection, and Linux visual testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Tonio <tonework@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
685d4faba3 |
Fix PostgreSQL recovery after a transaction connection closes (#13643)
Reject queued and late work from disconnected transaction and reservation scopes. Keep closed reservations out of the open pool, and clear old connection buffers and responses so new requests can reconnect safely. Twelve real-PostgreSQL regression cases cover crash prevention, recovery, and transaction isolation in both ESM and CommonJS. Database checks and all PR CI checks pass. Greptile: 5/5, no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e1d2e279a3 |
fix(db): replay a query whose socket write failed on a recycled pooled connection (#13417)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - A hosted instance keeps its data in Postgres, and it reaches that
database through a connection pooler
> - A pooler recycles the server side of a connection that sat idle. The
client still holds the socket and believes it is open
> - The next query fails while the driver writes it to that dead socket,
with `write CONNECTION_CLOSED host:5432`
> - The request that happened to draw the recycled connection fails,
although nothing is wrong with the query or the database
> - Only one code path guards against this today, so the same error
keeps reaching users and error reporting from every other path
> - This pull request replays a query when the write itself failed,
because those bytes never reached the server
> - The benefit is that a recycled connection costs one retry instead of
one failed request
## Linked Issues or Issue Description
No existing issue. The problem, in the bug report format:
**What happened**
Requests fail with `write CONNECTION_CLOSED <host>:5432` (driver code
`CONNECTION_CLOSED`). It is most visible on the first queries after an
idle period, when the pool holds connections the pooler already
recycled.
**Expected behavior**
A connection that the pooler recycled while it was idle is not a
user-visible failure. The client notices the dead socket and uses a live
connection.
**Steps to reproduce**
1. Run the server against a pooled Postgres endpoint (for example a Neon
`-pooler` host).
2. Leave the instance idle until the pooler recycles the server side of
the pooled connections.
3. Issue any request that queries the database.
**Paperclip version or commit**
Present on master.
## What Changed
- `packages/db/src/transient-write-retry.ts` (new) wraps the root
`postgres.js` client. When a query fails while the driver writes it to
the socket, the wrapper runs it again on a fresh connection. Three
attempts, 50 ms then 100 ms backoff.
- `packages/db/src/client.ts` gives Drizzle the wrapped client. The
teardown registry keeps the real client, because shutdown must end the
actual pool.
- The wrapper is deliberately narrow:
- It matches only the write phase (`code === "CONNECTION_CLOSED"` and a
message that starts with `write CONNECTION_CLOSED`). The write failed,
so the server never saw the query, and a replay cannot run anything
twice. That makes it safe for reads and writes alike.
- An error after the write propagates untouched, because the server may
have acted on the query.
- `CONNECTION_ENDED` and `CONNECTION_DESTROYED` propagate untouched,
because they mean a deliberate shutdown.
- Queries inside `db.transaction()` run on the scoped client that
`sql.begin()` returns, which the wrapper does not touch. A transaction
that loses its connection must abort, not replay.
- A query runs once however many handlers attach to it, and it still
executes lazily, like `postgres.js` itself.
## Verification
```
npx vitest run packages/db/src/transient-write-retry.test.ts # 7 passed
npx vitest run packages/db/src/client-teardown-registry.test.ts \
packages/db/src/client-options.test.ts # existing db suites pass
npx vitest run server/src/__tests__/cloud-tenant-transient-db-retry.test.ts # 6 passed
cd packages/db && npx tsc --noEmit # clean
```
New tests cover the replay, the `.values()` form Drizzle uses, the
attempt budget, an error that must not replay, and one execution per
pending query. A Drizzle round trip runs through a fake wire-protocol
server, which shows the wrapper is transparent to ordinary queries.
## Risks
Low risk.
- The retry only fires for a failure during the socket write, where the
server never received the query. A replay therefore cannot duplicate an
effect.
- A permanently unreachable database costs two extra attempts and 150 ms
before the same error surfaces.
- Transactions keep exactly their current behavior.
- Existing `retryOnTransientDbConnectionError` in the auth middleware
stays. It wraps a broader set of codes for one path, and it is
unaffected.
- No migration. No configuration change. No API change.
## Model Used
- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documented behavior or configuration changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — pending first run
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending first review
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
728f7185f6 |
feat: add native in-app announcements with persistent dismissal (#13403)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Self-hosted boards need a way to show occasional product announcements. > - An app release should not be required to publish or withdraw a card. > - Native card controls keep publishing consistent; the hero can use a static image or isolated HTML/CSS animation. > - This pull request renders a validated JSON feed with native components. > - It stores dismissals per account on each instance, so a closed card stays closed across companies and browsers. > - Named staging feeds let authors test content before production publication. ## Linked Issues or Issue Description **Subsystem affected** Board application shell, announcement delivery, and user preferences. **Problem or motivation** Operators need a small, optional announcement card. Users need reliable dismissal state. Authors need to test remote content without changing the production feed. **Proposed solution** Add one non-modal AnnouncementWell. Fetch validated JSON and content-addressed media through the instance server. Keep card controls native, with optional sandboxed HTML/CSS animation in the hero. Use stable announcement IDs for dismissal, an explicit empty manifest and quiet 404 handling. Provide a staged publishing helper and isolated test-drive guide. **Alternatives considered** Hosting the entire card as a page would move navigation and dismissal into remote content. This change limits HTML to a scriptless, isolated visual hero and keeps controls native. Browser-only storage would lose dismissals across browsers, so the instance stores account preferences. **Roadmap alignment** ROADMAP.md has no overlapping announcement feature. A GitHub title search found no related announcement pull requests. This work implements a maintainer-requested feature. ## What Changed - Add shared feed types, strict validation of every object, supported routes, expiration and version checks. - Add a board-only current-feed API, constrained media proxy, and idempotent dismissal API. Store the first dismissal and its company audit entry in one transaction. - Cache upstream data for one hour. Use conditional requests, request deduplication, response limits, public destination checks, and a three-second deadline. Treat a remote 404 as an empty feed with a fifteen-minute retry cooldown. - Keep announcement visibility stable when focus moves to browser chrome or another app pane; only tab visibility starts a return check. - Add a responsive native announcement card. Respect onboarding, dialogs and toast placement. Sync pending dismissals across tabs and retry after reconnect or return. - Add idempotent migrations for dismissals and validated publication IDs, design-guide examples, static and animated Storybook examples, and focused tests. The publication registry supports offline retries without accepting caller-invented IDs. - Add HTML/CSS animated heroes with static posters, automatic playback, reduced-motion handling, strict DOMPurify validation, an empty iframe sandbox and CSP that blocks scripts/network resources. - Add validated staging publication, content-addressed assets, an empty production manifest, preview fixtures, and authoring/operator documentation. ## Verification - The preceding implementation passed 98 targeted shared/server/publisher/route/OpenAPI/UI tests and 127 tests including the master rebase. The playback-control removal passes all 21 announcement UI tests, covering the rendered sandbox, fallback, reduced motion, dismissal and slow/stale state lookups. The preceding shared/server tests cover HTML validation and response sandbox headers. - The playback-control removal passes UI typecheck, production UI build, Storybook build and token gates locally. Browser verification confirms the animated card has only its dismiss button and two links, with no page errors. The full canonical CI matrix passed on current head `00e416431edb610861599d50490270bbd0f3c6b6`: 32 successful checks and two optional Storybook deployment checks skipped. This run needed no retries. Greptile reviewed this same head at 5/5 with no outstanding findings. - The local canonical general-server run passed 12,063 tests before reporting embedded-PostgreSQL startup failures in an unrelated fixture. All 31 tests in that fixture passed across isolated retries. The UI group passed 6,219 tests and other workspace groups passed 3,201; two CLI database-startup failures also passed individually. Serialized server suites were verified by the full CI matrix rather than repeating them locally. No source changes were needed for these environment failures. - The real S3/CloudFront staging manifest and both media asset headers were verified. Production remains empty/unpublished. The guide distinguishes the preview host's disabled edge cache from production cache requirements. - In the isolated test-drive, the animation visibly moves without playback controls. A 390×844 browser viewport keeps the card above navigation. Reduced motion makes no animation request. Both themes render correctly and browser page errors are empty. Browser fault injection verified that scripts cannot execute and CSS cannot make network requests; a missing animation leaves its poster and controls. - Refresh leaves the animated card visible. Closing it persists after reload and the API returns null. Earlier live checks verified dismissal across browsers, company-relative CTA navigation, modal deferral/restoration, and new-ID eligibility after restarting the same database. - The deployed empty feed and a real remote 404 return HTTP 200 with null from the board API, with a usable dashboard and no announcement popup or browser warnings. - Authoring documentation covers staging, animated HTML constraints, test-drive, withdrawal, ID reuse and cache-refresh steps. ## Risks - Animation supports self-contained visual HTML/CSS and inline SVG, without JavaScript or external resources. A static image is required. Older builds that do not recognize the optional animation field quietly hide that unsupported feed. - The default feed makes an outbound request from an instance when a board is used. Operators can disable it. Requests contain no account IDs, company data, cookies or interaction events. - Feed publication and withdrawal can take about 65 minutes to reach returning users because of CDN and instance caches. Expiration also removes visible cards locally. - Dismissals follow an account within one instance. No-login instances share the existing local-board identity. Separate installations do not share state. - Both tables are additive. A unique key prevents duplicate dismissals; the transaction prevents duplicate first-dismissal audit entries. The publication registry retains only validated IDs. AGENTS.md and the implementation spec document the required exception to company scope for these instance-level records. - Publication was limited to separate public staging prefixes on the existing preview host. Production remains empty/unpublished. No AWS policies or infrastructure were changed. ## Model Used OpenAI GPT-6 through Codex. The exact runtime model ID and context-window size are not exposed in this session. Capabilities used: reasoning, code editing, shell execution, tests, browser interaction, and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d1f0c20af |
fix: let responsible users choose either AI subscription or API key (#13351)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - AI Connections select the account used for each run. > - A responsible-user binding must follow the person whose work the agent performs. > - The saved sign-in method currently blocks users with another method for the same provider. > - This pull request resolves a personal default by company, user, and provider. > - Each user can use a subscription or API key with the same bot and model. ## Linked Issues or Issue Description Refs #13247, #13248, #13346, #13347. **What happened?** A bot configured with a Claude subscription rejects another responsible user’s Claude API key. Inline repair also limits that person to the original sign-in method. **Expected behavior** The same bot uses each responsible user’s default Claude account, whether it is a subscription or API key. Explicit shared account selections remain fixed. **Steps to reproduce** 1. User A connects a Claude subscription and creates a bot using the responsible user’s connection. 2. User C connects a personal Claude API key. 3. User C runs the same bot. Before this fix, credential resolution fails. ## What Changed - Add a personal provider-default table. Preserve legacy per-method preferences and backfill the most recently updated preference, including unavailable defaults. Repeated migration does not replace a selection. A database trigger propagates old-server default updates without treating new accounts as replacement defaults. - Resolve responsible-user bindings by provider. Retain the method as a wire compatibility hint for old servers. Explicit selections still require the exact method and grant. - Use the selected account’s method for credential isolation, refresh locking, and run attribution. - Update onboarding, agent setup, the picker, and inline task repair. Keep existing authentication components and harness/model settings. - Add mixed-method runtime, migration, repair, and Storybook coverage. Include upstream’s duplicate Anthropic option fix through the base branch. ## Verification - Focused resolver, migration, connection-intent, onboarding, agent setup, model, and connector UI suites: 331 tests passed. - Onboarding and new-agent regression suites passed during the initial focused run. - UI typecheck, token gates, and Storybook build passed. - Live browser checks passed for Claude and Codex API-default execution, switching both back to subscriptions, and both existing shared-account bots. Bot configuration remained unchanged. - One Daytona startup command stalled before Claude launched. The test run was cancelled, its sandbox stopped, and the same account/task passed on retry. Startup cancellation remains a separate environment finding; this PR does not change that command transport. - Browser review: all eight assertions passed in the new mixed-method story, including shared selection, return to responsible-user selection, and unchanged harness/model. - Repository build and typecheck passed after refreshing upstream dependencies. Final resolver and historical rollback verification: 38 tests passed. All latest-head CI gates passed, including server/workspace/serialized suites, browser E2E, build, typecheck, and runner verification. The extra serial local full-suite run was stopped after equivalent CI passed; focused local checks completed. ## Risks - Users with both historical method defaults get their most recently updated preference as the initial provider default. They can change it explicitly in Connections. - A revoked or unavailable default blocks. Connecting an additional account does not silently replace it. - Existing legacy authentication is unchanged. Managed responsible-user bindings intentionally stop pinning a method. - Live staging: the same Claude and Codex bots completed real API-key runs after changing only the personal default, then completed subscription runs after restoring the original defaults. Read-only database verification confirms unchanged bot configuration and actual method attribution. Distinct-user concurrency is covered by automated real-database tests with synthetic credentials, not two live human logins. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, code execution, and browser interaction. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
47ded8bf97 |
feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need credentials for a specific provider and sign-in method. > - Connections already owns accounts, grants, and access permissions. > - AI authentication should use those same boundaries. > - This pull request adds the storage, API, adoption, and runtime foundation. > - Legacy agents keep their authentication until they explicitly adopt a managed connection. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add AI-purpose/runtime-auth contracts and an additive, idempotent migration. - Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and catalog entries. - Store credentials on grants. Resolve responsible-user defaults or explicit permitted grants. - Isolate managed credentials and provider sessions across accounts. Block missing credentials without ambient fallback. - Keep imported legacy secrets unchanged during reconnect. Use independent local Codex/Grok sign-in attempts for rotating credentials. - Add authorization, migration, concurrent refresh, retry, cancellation, and legacy-compatibility tests. This is part 1 of a two-PR stack. The app UI follows in #13248. Merge the foundation first. ## Verification - Updated against master `04e364236`, preserving upstream provider login and connector workflows. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All current-head CI checks passed on `2a996560a`, including all server/workspace tests, browser shards, runner verification, typecheck, build, and canary dry run. Greptile reviewed that commit at 5/5 with no unresolved threads. Earlier local full-suite attempts hit the Mac PostgreSQL shared-memory limit; the complete suites passed in CI. - Renumbered the additive AI migration to `0276` after upstream migrations and regenerated its snapshot. Existing legacy agents retain their configuration. - Added local login status checks, owner-scoped retry, managed OpenCode remote homes, credential-aware model discovery, and task connection-repair delivery. ## Risks - Managed credential failures intentionally block execution. They do not restore legacy fallback. - Preview-era copied Codex/Grok subscriptions require independent reconnect. - The integrated branch has live provider acceptance coverage. This update verifies local Claude detection and Codex API-key task repair; it does not add a new subscription authorization/refresh or Daytona stress pass. - Runtime-auth connections must stay excluded from tool and channel handling. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4d317274ce |
feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Channels connect external conversations to company tasks and agent execution. > - Slack, Discord, and AgentMail already provide durable delivery and access controls. > - People also need to reach an agent from Apple Messages and send photos. > - Photon provides shared Pro DMs, dedicated numbers, and authenticated event recovery. > - This pull request connects Photon to the existing channel services. > - People can message an agent while Paperclip retains task ownership and approval authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: channel services, shared contracts, database constraints, Apps, and agent Channels UI. **Problem or motivation** Paperclip has no iMessage channel. A person cannot use Apple Messages to start a task, send a photo, or answer an agent's pending question. **Proposed solution** Add experimental **iMessage Photon** with Pro-compatible shared DMs or a dedicated Photon Cloud number per agent channel. Reuse channel admission, identity links, task generations, publication, and interaction continuation. Keep groups disabled for shared allocation. Dedicated lines support groups that an operator explicitly enables. Require a fresh linked message and a published agent response before setup completes. **Alternatives considered** Shared allocation has no owned phone number, so it reserves one project and allows DMs only. Dedicated allocation reserves one stable number. Local Mac access needs a separate deployment model. The upstream Photon Chat SDK adapter does not persist the poll mappings and send receipts required here. This change uses the lower-level SDK without adding another agent runtime. **Roadmap alignment** This extends Connected Apps and agent communication through the existing channel subsystem. It does not add a parallel tool connection or agent loop. GitHub searches for Photon and iMessage found no matching provider implementation. **Additional context** This ships behind the existing experimental channel gate. Dedicated-line release qualification remains incomplete. Real Photon Pro DMs passed task/reply, native poll, text answers, confirmation rejection, media, restart, pause, reconnect, revocation, and removal tests. An operator-supplied iPhone camera HEIC also passed the full round trip. Dedicated groups remain unqualified. See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the implementation plan](doc/plans/2026-09-11-imessage-photon.md). ## What Changed - Add the provider catalog entry, shared setup contracts, and a forward migration. A global partial index reserves the dedicated number or shared project until its endpoint is archived. - Add Cloud project inspection, vaulted project credentials, selected-line token renewal, and a leased receiver. Persist checkpoint updates under the receiver lease. Shared project replay accepts sparse increasing sequences only after a complete recovery barrier. - Connect DMs and enabled groups to existing task generations, sender authorization, ordered delivery, and publication services. Keep each iMessage conversation on its task after completion; only explicit `/new` or `/close` releases the binding. Publish committed inbound comments live and label their human bubbles “Sent from iMessage” in both task-chat renderers. - Persist immutable text/file send identities, upload receipts, poll IDs, option IDs, per-person drafts, and canonical interaction continuation proofs. - Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG previews, and related Live Photo companion video retention. - Add the three-step setup flow and channel management surfaces with official branding. Preserve the experimental gate and existing pause/disconnect behavior. - Add interactive production-component Storybooks for setup, access, recovery, and ongoing conversations. Add provider, integration, catalog, and browser regression coverage. Document setup, recovery, supported boundaries, and qualification gaps. ## Verification - Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and receive native Codex replies in Apple Messages. Unlinked senders cannot start work. - Three real follow-ups each reopened the same completed task. Incoming bubbles appeared on its open page without reload and showed “Sent from iMessage.” The third follow-up ran after restarting the server on `4d7222110`; the agent correctly repeated its previous reply from before the restart. - Native polls after restart, sequential text drafts, required-field correction, explicit submission, approval rejection with a required reason, and native continuation passed against Photon. - PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC passed in both directions. The camera photo produced a 3024×4032 JPEG preview. The native agent described it and returned the received HEIC byte-for-byte. - Pause/resume, reconnect, identity revocation, removal, `/status`, `/new`, `/close`, and stale answers after close passed live. Messages suppressed by pause did not become work on resume. Removal stopped intake and removed credential bindings. - All 304 focused tests passed on `4d7222110`. These cover Photon unit/integration behavior, both task-chat renderers, live comment hydration, completed-task continuity after restart, enabled groups, duplicate delivery, and explicit reset/close. The selected Teams completion-boundary regression also passed. Full workspace typecheck/build and token gates passed for the conversation fix; the final UI changes passed their affected typecheck/build and tests. - All 26 new Photon Storybook Playwright cases passed in light and dark themes, including the complete shared-DM setup journey and 390px mobile follow-ups. UI typecheck and the Storybook build passed. These stories use simulated Photon responses and do not replace the live evidence above. - The full chat-adapters browser suite previously passed all 39 cases. Migration checks passed, and migration 0275 applied to the isolated live instance with the earlier Photon migration already applied. - The local full Vitest run was previously interrupted by the host's embedded-Postgres shared-memory limit; it is not a full-suite pass. All 30 applicable CI checks passed on preceding head `7a5419cac`, with two skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit required-story discovery guard to the 26 passing Storybook cases. Greptile rates this final head 5/5 with no unresolved review threads. All 30 applicable CI checks passed, with two optional checks skipped. - A repeated live send key suppressed the duplicate but returned gRPC 6 / SDK `internalError` without an original receipt. Paperclip keeps unknown delivery unresolved. This provider behavior is covered by a regression test. - See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package versions, redacted live evidence, deterministic coverage, and remaining qualification gaps. ## Risks - Dedicated group qualification remains unrun; groups are disabled for the approved Pro scope. Real iPhone camera HEIC passed transport, preview generation, agent inspection, and return. Keep the channel experimental; the dedicated-line release matrix remains incomplete. - Shared recovery and attachment aliases were verified against the live gateway. Duplicate writes currently return an error without the original receipt; unresolved sends require operator resolution. The implementation fails visibly on invalid replay ordering, a reset cursor, or changed identity. - The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF binaries have not been executed in this work. Linux musl has no packaged converter. Unsupported conversion retains the original and reports the missing preview. - The migration adds a global reservation across companies for Photon numbers and shared projects. Paused and revoked endpoints keep that reservation until removal. - Integration touches shared channel services. Existing provider browser coverage passes; broad repository verification is recorded above. - `pnpm-lock.yaml` is intentionally excluded under repository policy. The repository bot owns lockfile updates. The additional Superagent supply-chain scan is neutral/inconclusive because these new dependencies are not yet in the committed lockfile. Its security scan passed; all required CI checks pass. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code execution, browser testing, and tool use. The exact served model identifier and context-window size are not exposed in this session. No sub-agents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed50a39c3f |
fix: preserve NUL characters in run-event payloads (#13325)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dbf5ea432d |
fix: protect starting runs during overlapping deployments (#13285)
## Thinking Path > - Paperclip controls agent work across service deployments. > - A run can provision a remote sandbox before a process or invocation event exists. > - Each container previously treated its own missing process handle as proof that the run was orphaned. > - Overlapping deployments could therefore fail a run owned by another container. > - This pull request records and renews a controller lease before provisioning. > - A recovery worker must revoke an expired owner before it finalizes the run. ## Linked Issues or Issue Description Merged PR #13272 records startup adapter identity and restores explicit user continuation. This PR adds controller ownership on top of current master. Refs #7997 and #10442 for related replica and ownership problems. Related #13138 addresses silence and detached local processes; this change does not infer death from silence. **What happened?** During an overlapping hosted service deployment, a new container reaped a legacy conversation run that another container was provisioning. The run had no PID or adapter invocation yet. **Expected behavior** A live controller keeps its run. After controller loss, one recovery worker takes cleanup authority and the old controller cannot dispatch further work. **Steps to reproduce** Claim a legacy run in controller A. Start controller B against the same database before A finishes provisioning. Run the startup reaper in B. **Paperclip version or commit** Observed on `663c44cb2b9c28336d38d0b4a6971f4f1964bce6` in a hosted Railway deployment with a Daytona environment. ## What Changed - Add nullable controller boot ID, lease deadline, and execution stage columns. Claim ownership in the queued-to-running update. - Renew ownership independently of run output. Abort and reject dispatch if renewal fails. - Serialize reaper revocation against renewal. Let unfinished recovery claims expire after a restart. - Restrict graceful shutdown to legacy runs owned by the current controller. - Hand ownership back to the existing native coordinator when runtime selection becomes native. - Add twelve database regressions and document the lease contract. Update the task-drain regression to require controller expiry before reaping. ## Verification - `pnpm exec vitest run server/src/services/legacy-controller-lease.test.ts server/src/__tests__/heartbeat-task-drain-admission-release.test.ts`: 14 passed after rebasing onto master (`f12b647ae`). - Queue-interruption regressions in `heartbeat-process-recovery.test.ts`: 2 passed after preserving the new cleanup promotion from #13275. - `pnpm --filter @paperclipai/server exec tsc --noEmit`: passed after rebuilding runner TypeScript outputs for the updated master. Broad local tests are omitted at the maintainer’s request; CI owns broad coverage. - Latest-head CI passed on `f255e8e4d5ab2b24b12638a02434e6aa8a2285c5`: [run 34658248569](https://github.com/paperclipai/paperclip/actions/runs/34658248569). All test shards, browser suites, typecheck, build, canary, and security checks passed. Greptile is 5/5 with no unresolved review threads. ## Risks - Additive, idempotent migration; historical rows retain the previous recovery behavior. - Database unavailability aborts new dispatch rather than permitting an unfenced controller to continue. - Lease expiry is permission to clean up, not evidence that remote inference stopped. Follow-up PRs add persistent cleanup and automatic continuation. - Mixed-version deployment still includes old binaries whose reapers do not understand controller leases. ## Model Used OpenAI GPT-6 through Codex, using reasoning, repository inspection, code execution, and test tools. The precise backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2083bf6f9a |
feat(connections): add AgentMail inboxes and email tasks (#13256)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents controlled access to external services. > - Experimental channels already map conversations to tasks and durable work queues. > - Email needs inbox ownership, recipient envelopes, delivery records, and explicit sends. > - This pull request adds AgentMail to that infrastructure and keeps the provider key in the server vault. > - Agents can receive and send email from local or sandbox execution while the board follows each conversation in its task. ## Linked Issues or Issue Description **Problem or motivation** Agents need dedicated email addresses. Incoming email should become assigned work. Internal task comments and progress must never become outgoing email by accident. **Proposed solution** Add experimental AgentMail connections, an inbox assignment wizard, durable email intake and publication, task email cards, and authenticated API, CLI, and native runtime actions. Agents use Paperclip credentials to request sends. Paperclip owns the provider key and enforces access and task authority. **Alternatives considered** A general mailbox MCP connector does not provide durable task binding or publication boundaries. A separate mailbox application duplicates task collaboration. The board instead directs the agent through the normal task conversation. **Roadmap alignment** This extends the existing experimental connections and task infrastructure. Product scope and interaction design were reviewed with the maintainer. Related connection authority work: #11831 and #11818. The duplicate search found no competing task-based AgentMail integration. ## What Changed - Add AgentMail catalog data, shared contracts, company-scoped email records, and an additive migration. - Add vaulted setup, inbox assignment, access grants, trust guidance, and provider-side allowlist guidance. - Support WebSocket and signed-webhook intake through a shared durable pipeline, deduplication, catch-up, and task wakeups. - Queue explicit new conversations and replies with immutable send intents, idempotency, delivery state, and uncertain-send resolution. - Show inbound and outbound email cards in normal task conversations. Keep internal messages internal. - Add task-scoped CLI actions and the sandbox callback routes required for Daytona execution. - Provide a dedicated AgentMail skill automatically only to agents with active authorized inbox assignments. Keep email instructions out of the universal Paperclip skill. - Advertise connector-owned `agentmail_inboxes`, `agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery` tools only in eligible native sessions. Recheck live authority on execution. - Isolate Codex CLI connector skills by agent and skill revision. Deliver the assigned skill in the run prompt for adapters that use shared skill directories, including resumed turns. Keep automatic skills out of manual persistent sync. Show them as read-only and document the pattern in the connector playbook. - Fix AgentMail health checks that entered local-stdio validation and optional missing Codex credential cleanup in sandboxes. - Add API, pipeline, authorization, sandbox, browser, and Storybook coverage. ## Verification - Live AgentMail testing covered WebSocket intake, signed webhooks, restart catch-up, and a full receive → task → Daytona Codex CLI → explicit reply → Delivered round trip. The reply was verified in the other inbox. The normal task composer also initiated an outgoing email child task. - The connector-skill change was verified in the browser: AgentMail appears once as an automatic, read-only skill with its assigned address. Disabling experimental chat connections removes it; re-enabling restores it. A regression test covers assignment data arriving after library data. - Connector regression coverage passed 178 runtime utility, email integration, skill-route, and heartbeat tests. All 17 Codex execution tests passed, including per-agent skill isolation, model identity, revision changes, removal, and prompt delivery without shared skill files. - After rebasing onto master, all 44 focused email, heartbeat, and native-authority tests passed. All 313 native-session executor tests passed. The UI regression suite passed all 3 tests. These test sets overlap earlier focused runs. - Full workspace typecheck and build passed after the rebase. Token gates passed. Earlier focused Playwright task/setup coverage and the Storybook build also passed. - Native connector tool execution uses deterministic integration tests. Live Daytona qualification used the Codex CLI adapter; the new shared-home prompt fallback has deterministic coverage. - The full repository suite is run by CI. The earlier unsharded local full-suite attempt was stopped after the equivalent CI suites passed and is not reported as a completed local run. Greptile reviewed `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved threads. All server, workspace, serialized server, and browser suites passed in CI. The build job hit a five-second timeout in a runner transport test; both variants and the full 80-test file passed locally with unchanged timeouts. The build passed on retry on the same commit without code or timeout changes. All required CI gates, including the final `ci / verify` and `ci / e2e` summaries, are green on `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`. ## Risks - Email from external senders can start normal agent work. Setup recommends a low-trust agent and AgentMail sender controls. Sender addresses never grant board membership. - Provider timeouts can leave uncertain sends. Retries retain their idempotency key; expired windows require reconciliation or operator resolution. - Connector skills and native tools are assignment-dependent and require current access. Revocation denies retained calls; assignment changes select a new runtime context. - Activation remains behind the experimental-channel setting. The native runner path has deterministic coverage; live Daytona qualification used the Codex CLI adapter. - Schema changes are additive. Inbox ownership is unique across companies. Disconnect preserves provider inboxes and task history. ## Model Used OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bce976d60d |
feat: bind an agent to a Codex login whose account differs from the company default (#13067)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `codex_local` adapter signs agents in to OpenAI, and a company keeps one default Codex identity in its shared company home > - A login with a DIFFERENT account than the company default is deliberately kept out of the shared home — one agent's sign-in must not switch every unbound agent's credentials — but that left the cross-account login inert: nothing connected the agent the operator was configuring to the credential the login stored > - The stored credential and its company secret already exist; only the last mile — an agent actually using them — was missing > - This pull request reports a non-secret binding claim on the authenticated login and lets the agent page bind that one agent's `CODEX_HOME` to the account's secret, exactly and only when the identities differ > - The benefit is that multi-account Codex becomes one click on the agent that needs it, with company-wide identity untouched ## Linked Issues or Issue Description **What happened?** On an agent's detail page, "Sign in with Codex" using a different OpenAI account than the company default succeeds but changes nothing for that agent. The credential lands in the per-identity store and a company secret names it, but the agent keeps using the company default. The Test keeps reporting that authentication is needed, and no repeat login helps. **Expected behavior** When the operator deliberately signs an agent's page in with a different account, that agent starts using that account. Agents that were not part of the action keep the company default. A same-account login keeps working through the shared company home with no per-agent pinning. **Steps to reproduce** 1. Configure a company whose Codex home holds account A. 2. Open a `codex_local` agent's detail page with a sandbox environment and complete "Sign in with Codex" using account B. 3. Press Test. Before this change the agent still resolves account A and the authentication-needed check returns. ## What Changed - `packages/adapters/codex-local` — the prerequisite shield: `isCodexAuthCachePath` recognizes per-identity credential-store entries, and `seedManagedCodexHome` refuses to symlink, heal, or API-key-overwrite an entry's `auth.json`. The seeding pass runs before every probe and execute; without the shield, an agent bound to an entry would have its stored login silently swapped for the host credential. Static shared config files still copy in. Rotation already survives binding: the sandbox copy-back writes rotated credentials into the identity-keyed store slot. - `server` — the promotion records whether the company default home ended on a different account than the login (any read failure degrades to `false`, so the client can never be told to bind wrongly). After the terminal commit, the routes layer remembers a non-secret claim — the opaque account-home secret id plus that verdict — in a bounded in-memory map, and merges it into the owner read of an `authenticated` `codex_local` session. A restart drops the claim; the panel then shows plain success. - `packages/shared` — `CodexAccountBindingClaim` on the owner session response. It carries no account identifier and no credential byte. - `ui` — the login panel reports the claim upward once. The edit-mode form binds the agent's `CODEX_HOME` to the secret and saves in one step, only when `companyIdentityDiffers` is true. Same-account logins bind nothing on purpose: the company-home refresh already carried them, and an unbound agent keeps following the company default across rotations. Create mode is unchanged. ## Verification - Adapter suite: 381 passed, 1 skipped (includes the new store-entry shield tests and the path-predicate cases). - Server suites (8 files): 130 passed, 15 skipped — including two new route tests that drive a login to `authenticated` and assert the claim with both identity verdicts. - UI render suite: 85 passed — including a panel test that the claim is reported upward exactly once. - `tsc --noEmit` clean in `packages/shared`, the adapter package, and `ui`; `server` clean for the touched file. ## Risks - The bind changes one agent's configuration through the normal agent-update patch, initiated by the operator's own login on that agent's page. The failure direction of every fallback is "offer nothing": a missing claim, a restart, or an unreadable company home all degrade to no bind. - The seed shield narrows what the seeding pass may touch; homes outside the credential store behave exactly as before, covered by the existing seed tests. - Builds on the sign-in credential-resolution fix (#13064), now merged; this branch is rebased onto master and the diff contains only the binding feature. Supersedes #13066, which GitHub auto-closed when its stacked base branch was deleted on merge. ## Model Used Claude (Anthropic) — Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use in Claude Code (terminal). **Related PRs (searched; no duplicates found):** #12740, #12082, and #9621 touch adjacent Codex credential sync paths; #8495 is the standing hardening effort for probe auth seeding. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no standalone docs cover this flow; the behavioral contracts are documented in-line at each changed site) - [x] I have considered and documented any risks above |
||
|
|
6abeb67334 |
feat: add opt-in chat provider and data foundation (#13100)
Add dormant provider contracts, qualified patched adapters, tenant-scoped persistence and lifecycle ownership without activating chat routes. Preserve the experimental integration as dependent PR #13038. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
82f662656a |
fix(runner): restore legacy Git access and independent networking (#13094)
## Thinking Path
> - Paperclip runs agents for people with different GitHub accounts.
> - Managed operations must use the intended person's eligible
connection.
> - A failed duplicate connection must not hide a healthy grant for the
same account.
> - Legacy hosts also need their existing Git configuration when managed
access is not configured.
> - Runner networking and local Git operations must not depend on GitHub
broker availability.
> - This pull request separates those policies and improves failure
diagnostics.
## Linked Issues or Issue Description
**What happened?** New runs always cleared host Git credentials and
installed managed launchers. Network permission depended on GitHub
environment variables. A launcher failure could stop even local `git
status`. A newer unhealthy duplicate could take precedence over a
healthy connection, and generic health errors were shown as reconnect
requirements.
**Expected behavior:** Use a healthy eligible managed connection for the
intended account. Preserve host authentication only for unconfigured
standard-trust local or SSH execution. Permit local Git during broker
failures and keep network permission independent of GitHub credentials.
**Steps to reproduce:** Configure healthy and unhealthy grants for one
GitHub account, dispatch an agent, and execute Git commands. Separately
run an unconfigured legacy host with existing GitHub CLI authentication.
Stop the broker and run local `git status`.
**Paperclip version or commit:** Master at
|
||
|
|
35fdc0c66b |
fix: make task recovery durable and preserve current requests (#13075)
Make task recovery durable and preserve the latest user request across native and legacy continuations. Keep routine recovery quiet and prevent replay when action outcomes are uncertain. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e200104727 |
feat: review connection actions from tasks (#13063)
Bring governed connection reviews into task history and composer approvals. Share resolution with Connections, add scoped remembered permissions, and resume agents through durable outcome receipts. Keep cards compact, collapse raw results, isolate untrusted provider output, bound continuation payloads, and reconcile missed live events. Add Storybook coverage, browser journeys, and service regression tests. Verification: all PR CI gates passed, Greptile 5/5, security scans passed, five connection-review browser journeys passed, and real native Codex approval/continuation was verified against the local MCP fixture. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
7ed122911b |
Add end-to-end session goals to Paperclip Runner
Add capability-aware slash-goal controls, durable provider goal state, PRP v2 negotiation, autonomous goal execution, and safe local session recovery. Integrate with current master, preserve provider session identity, and verify the browser goal/chat/replacement/clear workflow and unsupported-agent rejection. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e095b84dab |
feat(connections): connect services from native task feeds (#13058)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents use connections to reach external services. > - A fresh native task can have no service tools installed. > - The agent needs a way to discover services and ask the responsible person for access. > - This pull request brings the existing connection-intent flow into native task execution. > - The person can connect from the task, and the agent can continue with updated tools. ## Linked Issues or Issue Description **Subsystem affected** Native runner tool authority, connection intents, task interactions, and shared connection setup. **Problem or motivation** A task that needs an unconnected service cannot finish its work. Leaving the task to configure access also loses context. A resolved request must survive a restart and resume the correct agent once. **Proposed solution** Expose connection discovery and access requests as server-owned native tools. Render a durable task card and use the shared setup dialog. Persist outcome delivery and start a fresh provider session after access is ready. **Alternatives considered** Sending the person to the Connections page adds navigation and does not solve continuation. Polling for authorization consumes runs and can create duplicate requests. **Roadmap alignment** This extends the existing connection-intent runtime and setup experience. It reuses the shared access model and the native runner. Related: #12345, #12347. The service-slug fix in #12906 is related but separate. Companion evaluation PR: https://github.com/paperclipai/paperclip-evals/pull/21. ## What Changed - Expose `connections_search` and `connection_request` with server-bound company, task, agent, and responsible user. Preserve the legacy entry points. - Discover catalog services and authorized custom connections. Check installation, identity, health, and executable permissions before reporting ready. - Keep pending cards through ordinary messages. Reuse requests and retire stale ownership. Put Connect at the right of Not now. - Reuse the shared setup flow in a task dialog. Keep access additive and default to the requesting agent. Recover from cancelled or blocked OAuth windows with a new-tab fallback. - Persist outcome delivery with an idempotent wake key. Resume in a fresh session and recheck ownership before dispatch. - Add native browser fixtures, offline Storybook states, server contracts, and evaluation fixtures. Update guidance and documentation. ## Verification - `pnpm build`: passed after replaying the change on current master. - `pnpm -r typecheck`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed. - New continuation-policy regression cases: 16 passed. - Docker-backed PostgreSQL regressions passed for requester-only OAuth access, assignment-only expiry, terminal expiry, and credential-free setup metadata. - Shared setup and task-card UI tests: 121 passed, including configured MCP reconnect URL recovery and preserving user edits across refetch. - Storybook browser checks: all 119 passed on the latest reconnect fix. - `pnpm test:run`: 4,734 tests passed in the first server group, but embedded PostgreSQL startup failures and resulting cleanup errors prevented a complete local pass. All Linux CI lanes passed on the latest reviewed commit. One external-object route test returned an unexplained 500 on the first run; it passed twice locally and the failed shard passed on retry without code changes. - Earlier feature-checkout evidence: three deterministic native browser journeys passed, including restart delivery and an actual fixture tool result. Legacy scripted coverage also passed. All 59 added stories were inspected in light and dark themes. - Live Notion testing recorded successful provider reads. The manual test used a local-trusted instance. It does not prove authenticated/cloud deployment or every provider journey. - Native browser rerun reached the embedded PostgreSQL startup limit before bootstrap, so the latest checkout’s full native browser journey remains unverified. Both OAuth page/task regression cases passed against isolated Docker-backed PostgreSQL 17. They verify no premature task access, requester-only completion, additive retries, and reconnect preservation. - Applied both new migrations twice to isolated PostgreSQL 17. Foreign keys remained intact, duplicate active delivery keys were rejected, and failed delivery records did not block retries. Reviewer path: start a fresh test drive, enable the native runner, use an agent that can perform work directly, and ask it to summarize a Notion page. Connect from the card, then verify the resumed provider call and source-linked answer. The default test-drive CEO is instructed to delegate, so it can introduce an unrelated hiring step. ## Risks - Two additive migrations create durable deliveries and a partial unique wake index. They are idempotent. The wake index can require a maintenance window on large tables because migrations run in a transaction. - OAuth and continuation cross asynchronous boundaries. Tests cover ownership changes, retries, additive access, and restart delivery; live provider behavior still varies. - The latest requester-scope fix has not yet been exercised through live OAuth. GitHub, API-key, authenticated-user, and all recovery journeys are not claimed as verified. ## Model Used OpenAI GPT-6-based Codex assisted with implementation, tests, and review using tools and code execution. The runtime does not expose the exact model version, context window, or reasoning setting. Live evaluation used `gpt-5.6-luna`; manual native testing used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used and disclosed unavailable runtime details - [x] I have checked ROADMAP.md and confirmed this extends existing connection work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run all required tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
023e640a7e |
fix(db): reap idle pool connections, name the pool, and end it on shutdown (#12956)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server keeps one postgres.js pool (`packages/db/src/client.ts`, `createDb`) for every query it runs. #10795 made the pool tunable from the environment, but the defaults stayed at the driver defaults: an idle connection never closes, the pool reports itself as `postgres.js`, and no code path ever calls `sql.end()`. > - On a hosted Paperclip deployment the server entered a restart loop (a bundled plugin failure that #12953 describes made every run fail, and the pool saturated). Each generation opened its ten connections, died, and left the backends open on the PostgreSQL side until TCP keepalive reaped them hours later. After about 20 generations the backends exceeded `max_connections`, and every later boot died on its first bootstrap query with `sorry, too many clients already`, before `server.listen()`. The loop could not heal itself. #9555 describes the same shape on a launchd-supervised self-hosted install. > - Three properties of the pool combine to make this possible: idle connections are never reaped, the pool is never ended on any exit path, and an operator cannot even find the leaked backends in `pg_stat_activity` because they carry the generic driver name. > - This pull request gives the pool a 60 second idle timeout and the `paperclip` application name by default, exposes `max_lifetime` and `application_name` through the same `DATABASE_*` environment contract that #10795 introduced, and ends the pool on the orderly SIGINT/SIGTERM path and on the fail-loud startup path. > - The benefit is that a restarting or crash-looping server releases its backends instead of accumulating them, and an operator can see and count Paperclip's connections. ## Linked Issues or Issue Description - Refs #9555 — database connection pool leak causes an infinite restart loop under load. This PR closes the "pool never ends, idle connections never close" part of that report. - Refs #12953 — hosted outage report. The pool exhaustion is the second half of that incident; the first half (a stuck sandbox provider plugin) has its own PR. - Related prior PRs: #9597 and #8780 both propose hard-coded `idle_timeout` / `max_lifetime` values in `createDb`. Both predate #10795 (merged), which made these options environment-driven; this PR builds on the merged shape and adds the shutdown `end()` that neither covers. #4006 and #7481 are closed earlier attempts in the same area. ## What Changed - `packages/db/src/client.ts` - New `resolveDatabaseClientOptions()` applies Paperclip defaults on top of the environment: `idleTimeoutSeconds` defaults to 60 (`DEFAULT_DATABASE_IDLE_TIMEOUT_SECONDS`) and `applicationName` to `paperclip` (`DEFAULT_DATABASE_APPLICATION_NAME`). `createDb` uses it for both the environment path and explicit options. - `DATABASE_IDLE_TIMEOUT_SECONDS` now accepts `0` to restore the driver default (keep idle connections open). Negative or non-integer values still throw. - New environment variables: `DATABASE_MAX_LIFETIME_SECONDS` (positive integer, maps to `max_lifetime`) and `DATABASE_APPLICATION_NAME` (non-empty string, maps to `connection.application_name`). - `postgresJsOptions()` maps the two new options. - `server/src/shutdown.ts` - `finalizeServerShutdown` gains two optional ordered steps: `closeHttpListener` runs first, before the application services stop; `closeDatabase` runs after the application services and before the embedded PostgreSQL stop. A failure in either is logged and does not stop the teardown. Final order: listener → application services → database pool → embedded PostgreSQL → instrumentation → Sentry. - New `closeHttpListenerForShutdown()`: stops accepting requests, closes idle keep-alive sockets, waits up to 5 s for open connections, then closes whatever is left. Requests still in flight are drained while every service is available, and none can reach a route after `sql.end()`, on the signal path and the programmatic path alike (the programmatic path's later `server.close` finds the listener closed and skips). - `server/src/app.ts`: the app shutdown hook (`shutdownAppServices`) now stops the plugin job scheduler, whose tick queries the database, so a programmatic `shutdown()` leaves no timer running against the ended pool. - `server/src/index.ts` - `startServer()` is now a thin wrapper around the boot sequence. When the boot sequence throws after the pool exists, the wrapper ends the pool (and the separate migration pool, when configured) before it rethrows. This covers the `process.exit(1)` path in the main module and the CLI `paperclip run` path alike. - The orderly shutdown passes the same `closeDatabaseClients` to `finalizeServerShutdown`. - `endDatabaseClient` tolerates a client without `$client` (test doubles) and uses a 5 second end timeout. - Docs: `docs/deploy/database.md` gets a "Connection Pool Settings" table with every `DATABASE_*` pool variable, its default and its effect; `doc/DATABASE.md` lists the two new variables. - Tests - `packages/db/src/client-options.test.ts`: parsing of the new variables, `0` for the idle timeout, rejection of malformed values, driver option mapping, and the `resolveDatabaseClientOptions` defaults. - `packages/db/src/client.test.ts` (embedded PostgreSQL): `createDb(url)` reports `application_name = paperclip` for its own backend, and a pool with `idleTimeoutSeconds: 1` has zero backends in `pg_stat_activity` after the timeout. - `server/src/shutdown.test.ts`: the listener closes before the application services, and the database close runs between the application services and the embedded PostgreSQL stop; a failing database close is logged while the teardown still finishes; `closeHttpListenerForShutdown` closes idle sockets and resolves on close, force-closes after the grace period, and is a no-op when the listener was never bound. ## Verification - `pnpm --filter @paperclipai/db typecheck` — passes (`check:migrations` + `tsc --noEmit`). - `cd server && pnpm typecheck` — passes. - `cd packages/db && pnpm exec vitest run src/client-options.test.ts src/client.test.ts src/client-teardown-registry.test.ts` — 9 + 18 + 3 tests pass (the `client.test.ts` cases need embedded PostgreSQL; the new one waits up to 10 s for the idle reap and passed in about 3 s). - `cd server && pnpm exec vitest run src/shutdown.test.ts src/__tests__/server-startup-feedback-export.test.ts src/__tests__/bootstrap-claim-routes.test.ts` — 34 + 11 tests pass. The startup-feedback suite exercises `startServer()` with a mocked `createDb`, which is why `endDatabaseClient` tolerates a client without `$client`. - Manual check for a reviewer: start the server against any PostgreSQL, then run `SELECT application_name, state, count(*) FROM pg_stat_activity GROUP BY 1, 2;`. Paperclip's backends now show `paperclip`. Leave the server idle for more than 60 s and the idle backends disappear. Send SIGTERM and the backends close before the process exits. ## Risks - Behavior change with no environment set: idle pooled connections now close after 60 s. The next query after an idle period pays a reconnect (single-digit milliseconds on a local socket). postgres.js reconnects transparently. Set `DATABASE_IDLE_TIMEOUT_SECONDS=0` to keep the previous behavior. - `application_name` changes from `postgres.js` to `paperclip`. Anything that filtered `pg_stat_activity` on the old name would need an update; nothing in this repo does. - The HTTP listener now closes at the start of the final teardown (after the heartbeat run drain, which still needs the API for running agents). The pool close runs after the application services. A late query from a timer that survived the service shutdown would fail with a driver "connection ended" error instead of running; the known database-backed timer (the plugin job scheduler) is now stopped in the service shutdown. - The listener drain adds at most 5 s to a shutdown while long-lived connections (for example WebSocket clients) are open; after that they are closed forcibly. - `startServer()` is split into a wrapper and the boot sequence. The exported signature and return type are unchanged. - No migration, no schema change. ## Model Used - Claude Fable 5.1 (`claude-fable-5-1`) via Claude Code, extended thinking, tool use (file edits, shell, test runs). The change was produced with the model and reviewed by the submitting human. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_014t3bi2beVNVVHAxK36dmXm --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1cc45086d3 |
feat: use the responsible person's GitHub for shared agent operations (#13005)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Several people can send instructions to the same agent and task. > - A fixed GitHub token in the provider process can keep the first person's access after another person's message is accepted. > - Task ownership cannot select credentials for each accepted instruction or preserve the identity of an operation already in progress. > - This pull request records ordered execution identity contexts and resolves credentials when managed Git, gh, or GitHub tools start. > - The benefit is automatic personal GitHub access for shared agents, with durable continuation rules and no teammate credential fallback. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: orchestration, connection grants, database, runtime adapters, native runners, and run details. **Problem or motivation** A shared agent must use the person whose instructions it has accepted. A queued message must retain its author. A retry or approval without new instructions must retain the originating identity. GitHub must remain optional for ordinary work. **Proposed solution** Persist execution identity separately from task ownership. Give new processes a run-scoped broker capability and token-free managed launchers. Capture identity at operation start. Keep an explicit dedicated-agent grant as an override. Show redacted diagnostics in run details. **Alternatives considered** Per-task ownership, fixed provider tokens, and mutable repository author configuration do not handle accepted steering or concurrent operations. A manual account-selection action would add unnecessary setup to each turn. **Roadmap alignment** This completes the existing Multiple Human Users, MCP Tool Gateway & Apps, Secrets Manager, and Self-healing Runs capabilities. The implementation follows the maintainer-approved plan. Related work: Refs #12843, Refs #12907. Existing proposals #4618 and #8945 cover per-agent or per-worktree author configuration. This change instead follows the accepted human instruction across runtime types. Refs #11831 for governed personal connection delegation; this change preserves connection audience checks and does not use standing delegation as a personal credential fallback. ## What Changed - Add durable, ordered identity contexts and active run references. Preserve message authors through consolidation, steering, retries, delegation, approvals, routines, and restart. - Add an authenticated operation-time GitHub credential broker and local/remote managed git and gh launchers. Keep personal tokens out of the long-lived provider process. - Resolve GitHub gateway and server-side Git operations through the same responsible-person or dedicated-grant selection rules. - Make absent and unavailable GitHub credentials non-blocking at generic startup. Clear host and prior-person credentials. Keep anonymous Git access where supported. - Add run-detail identity history and the dedicated-account warning. Keep task ownership and queue-versus-steer decisions unchanged. - Preserve personal OAuth declarations through connection edits. Retain exact selected grants in the gateway. - Fix continuation races found during real acceptance: verify a warm owner before credential rotation, and wait for bounded durable runner suspension before the next run starts. - Make migrations replay-safe. Retain identity through agent/run deletion, remove it with its company, and clean terminal launcher directories before releasing execution environments. Document coordinated release and rollback. ## Verification - Full workspace typecheck, build, and token gates passed. The complete local suite passed in its normal test groups: 17,120 passing tests, including all 143 serialized server suites. After integrating the newly merged runner API work, full local typecheck and build passed again, along with 890 focused integration tests. All 31 checks on the integrated revision passed, including build, browser E2E, release registry, canary dry run, typecheck, security and all test suites. Greptile is 5/5 with all review threads resolved. - Current focused checks passed: 142 native executor tests, 67 runtime lifecycle tests, 9 durable identity tests, 75 credential/routine tests, 19 low-trust/resumption tests, and the executable migration replay test. - Authenticated browser acceptance with two Paperclip users and two GitHub accounts on one shared native agent passed. Real commits and pushes followed A → B accepted steering → queued A continuation in the same saved conversation. GitHub commit author and committer identities matched all three operations. Both runs succeeded and task ownership stayed unchanged. - Real GitHub MCP calls switched from A to B after accepted steering. A delegated subtask retained its originating identity across a server restart. - Disabling B's GitHub connection left ordinary work successful. Managed gh was unauthenticated and the provider had no inherited GH_TOKEN or GITHUB_TOKEN. - The browser displayed run-detail diagnostics and the exact dedicated-account warning. A final controller-restart check followed by another-person continuation retained the conversation, selected the correct GitHub login and Git author, and removed each terminal launcher directory. - Company-lifetime migration and all five previously failing CI suites passed locally (167 tests). Same-token gateway A → B → A and six broker/launcher boundary tests passed. - Remote callback, launcher, sandbox, and runtime contract tests passed. Both native and legacy Codex completed actual Daytona executions on the integrated revision ([campaign results](https://github.com/paperclipai/paperclip/actions/runs/34155056509)). The remote package-manager shim staging regression also passed locally. ## Risks - Deploy the migrations, server broker, launchers, and runner artifacts together. Existing processes finish with their original contract. New managed processes need the broker endpoint for GitHub operations. - Finish or stop new managed executions before rolling application code back. Keep the additive schema and identity history during rollback. - Scripts that require a persistent raw GH_TOKEN must use managed git, gh, or GitHub gateway tools. Run capabilities authorize code executing within that run to acquire its current identity; this is not hostile-code isolation within one execution principal. Managed commands prevent automatic credential carryover; arbitrary code deliberately copying a credential is outside that boundary. - Uncertain steering acknowledgement deliberately holds new credential acquisition until reconciliation. Already-started operations retain their captured identity. - GitHub private access and provider outages can still fail the specific operation that needs them. Dedicated grant failure does not fall back to personal access. ## Model Used OpenAI GPT-6 through Codex assisted implementation, review, shell execution, and browser acceptance. The exact model variant and context-window size are not exposed in this session. Tool use included TypeScript and Rust tests, database integration tests, GitHub CLI, and authenticated browser control. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
932ddb7b37 |
feat: browse GitHub repository access across organizations (#12998)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub connections give agents access to approved repositories. > - One GitHub identity can use installations across several organizations. > - The permissions page linked to one installation and showed an unfiltered list. > - Users could not easily find another organization or inspect a large selection. > - This change adds account filtering, search, and access configuration links. > - Users can inspect repository access in one compact view. ## Linked Issues or Issue Description Refs #12993. Related repository-catalog work in #11228 and #11234 was checked. This change only improves the existing GitHub connection permissions page. **What existing behavior does this improve?** The GitHub connection permissions page and its repository display metadata. **Current behavior** The page links directly to an existing installation. The repository list has no account filter, search, height limit, or private-repository marker. Refresh access occupies a separate section. **Proposed behavior** Show all authorized repositories by default. Filter by account or organization and search by name. Open GitHub's account chooser to configure access across organizations. Show GitHub icons and private-repository locks. Keep refresh beside configuration and limit the visible list to about ten rows. **Reason and benefit** Users can find repositories across organizations and configure missing access without creating another GitHub identity. Large repository lists no longer fill the page. **Breaking changes** None. Repository display metadata gains an optional private flag. Older snapshots remain valid and gain the flag after access refresh. No SQL migration is required. ## What Changed - Add an All accounts view, account filter, search, and empty states. - Link both configuration controls to GitHub's app account chooser. - Place an accessible refresh icon beside the configuration button. - Keep the repository heading and list in one section. - Add GitHub icons and private-repository locks. - Cap the scrollable list at ten rows using a design token. - Persist GitHub's private flag only when the provider returns a boolean. - Recover missing legacy app configuration from GitHub installation metadata. - Update tests and the GitHub connection runbook. ## Verification - Focused tests passed: 54 permissions-page tests and four GitHub metadata tests. - UI and server typechecks passed before submission. Token gates passed. - Browser checks verified account filtering, search, empty results, and the configuration destination. - The live list contained 40 repositories. Its final height was 272 pixels, which fits ten single-line rows with gaps. Scrolling retained all rows. - A live access refresh populated 30 private-repository lock icons from GitHub metadata. - Full workspace typecheck and build passed. The broad local suite stopped in the general-server group with 18 failed files. Failures include macOS temporary-path handling and embedded PostgreSQL startup. That run also overlapped the legacy fix and retained a stale GitHub module; the final focused run passed all 58 tests. Clean-runner CI is tracked separately. - Latest-head review is 5/5 with the legacy chooser finding resolved. All CI checks passed on commit `0ff2b63f348f5c87d8b7df6e43388f60f5d872d9`, including build, typecheck, all test shards, browser tests, and canary dry run. ## Risks - Older repository snapshots lack visibility metadata until refreshed. Unknown visibility does not display a lock. - The account filter lists authorized installation owners. Users add other organizations through GitHub's chooser. - Filtering changes only the displayed list. GitHub remains authoritative for repository access. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) via Codex. Reasoning, code execution, and browser tools were used. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bac60d9d31 |
fix: preserve GitHub sign-in and show connected repository access (#12993)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub connections give agents an account with selected repository access. > - Fresh local instances enroll with production Paperclip Cloud. > - Enrollment could finish while the GitHub OAuth profile remained disabled. > - Setup then switched to a personal access token form without explanation. > - This change preserves sign-in intent and shows the connected account and repositories. ## Linked Issues or Issue Description Related: #12907, #12943, #12947. Existing open GitHub connection work was checked. No duplicate was found. **What happened?** After Cloud enrollment, a fresh test-drive asked for a GitHub key. Production did not advertise the managed GitHub profile. Staging did. The permissions page also omitted the authenticated username and repository names. **Expected behavior** Continue with GitHub OAuth when available. Explain unavailable sign-in and allow retry otherwise. Show the GitHub username and complete accessible repository list. **Steps to reproduce** Start a fresh test-drive. Choose GitHub and complete instance enrollment while the Cloud GitHub profile is disabled. Open an existing GitHub connection's permissions page. ## What Changed - Preserve managed sign-in intent when the gallery omits its profile. - Refresh the selected gallery entry on retry without resetting the audience. - Fetch all pages of GitHub installations and repositories. - Store only repository IDs, full names, and installation IDs in grant metadata. - Show the GitHub username, repository list, management link, and refresh action. - Discard the repository snapshot after newer installation lifecycle events. Preserve snapshots verified after delayed events. - Lock and re-read grant metadata when applying installation events or saving refreshed access. Patch only webhook fields for other events. Reject snapshots if access changed during the external fetch, using unique access revisions even when timestamps collide. - Show repository installation recovery for managed OAuth even when the app also offers an advanced PAT method. - Update tests and the GitHub connection runbook. No SQL migration is required. ## Verification - Local typecheck, build, and token gates passed. All latest-head CI gates passed, including the complete test matrix and browser suites. Greptile is 5/5 with no unresolved findings. - All 382 focused setup, permissions, metadata, service, and webhook tests passed across final runs. One socket-hang-up test passed on rerun with the full service suite. Final service, metadata, and webhook checks passed all 230 tests. - The broad local suite was stopped after failures. Seven workspace-runtime exposure and control-conflict failures reproduce on base commit `54a99d884`. The broad run also overlapped local iteration; final focused tests and clean-checkout CI are tracked separately. - Browser: a fresh production-backed instance completed enrollment, retried after profile enablement, reached GitHub consent, recovered from a missing installation, and completed OAuth. - Browser: the permissions page showed the authenticated username and the selected private test repository. A real `get_me` call returned the same account. Reading the selected repository passed; reading an unselected private repository failed with 404. - Browser: a second fresh instance completed enrollment and OAuth without a PAT form or unavailable state. Its username and repository list survived reload and refresh. A real get_me call on the final code returned the displayed account. ## Risks - Repository names are now stored in company-scoped grant metadata and shown with that credential. They are display data, not authorization data. - Large selections require more GitHub API calls. A failed later page rejects the refresh rather than reporting a partial list. - Older grants and webhook-invalidated snapshots require Refresh access to load the list. - Cloud profile enablement is separate deployment configuration. This PR does not change OAuth scopes or GitHub App permissions. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) via Codex. Reasoning, code execution, and browser tools were used. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0ffc091473 |
feat(connections): add durable GitHub identities and webhooks (#12843)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents need source control access for repository work > - A shared token cannot preserve the responsible person's identity or an agent's dedicated identity > - GitHub App tokens also need durable refresh, repository access checks, and webhook delivery > - Paperclip already has managed connections, encrypted grants, run secret leases, and merge-confirmation behavior > - This pull request extends those systems with GitHub identities instead of adding a parallel credential system > - The benefit is durable GitHub access with explicit identity, repository, runtime, and webhook boundaries ## Linked Issues or Issue Description No public GitHub issue describes this connection change. This description follows the feature request template. **Subsystem affected** Connected Apps, connection grants, secret resolution, native Git runtime setup, webhook processing, and the Apps UI. **Problem or motivation** Users need to connect GitHub once and let agents use the correct GitHub identity. A run should use a dedicated agent account when one exists. Otherwise, it should use the responsible person's account. The connection must survive token expiry, repository access changes, and temporary instance downtime. **Proposed solution** Add user-owned and agent-owned GitHub grants to the existing connection model. Resolve one identity for MCP, Git, `gh`, health checks, and webhook bindings. Store provider tokens in the existing encrypted secret system. Refresh expiring token pairs under the existing lease and compare-and-swap path. Register signed Cloud webhook bindings and process normalized pull request and installation events through a durable local inbox. **Alternatives considered** An organization-wide GitHub token would lose person and agent attribution. Environment variables alone would bypass the managed connection and grant model. A new GitHub-only credential store would duplicate the existing secret and access systems. GitHub App installation tokens and private-key custody remain outside this first version. **Roadmap alignment** This change implements the Connected Apps direction. It also extends the shipped MCP Tool Gateway, per-agent secret access, and action-attribution systems. It does not add a repository catalog. The open repository catalog work in [#11234](https://github.com/paperclipai/paperclip/pull/11234) is related and complementary. ## What Changed - Added agent-owned connection grants and a per-agent credential policy with company and subject constraints. - Added a managed GitHub App method while keeping the personal access token method as an advanced fallback. - Added durable access-token and refresh-token handling with proactive rotation and one automatic recovery after a provider `401`. - Added GitHub identity and installation summaries without storing repository-name lists. - Added signed Cloud webhook binding, event lease, acknowledgement, local idempotency, pull request merge processing, and installation access handling. - Added one identity resolver for MCP, native Git, `gh`, checkout, health checks, and webhook bindings. - Added a class-3 run projection for `GH_TOKEN`, `GITHUB_TOKEN`, a `github.com`-only credential helper, SSH-to-HTTPS rewrite, and GitHub noreply commit attribution. - Added personal and dedicated-agent setup choices plus identity, repository, continuity, and webhook status in the Apps UI. - Added schema migrations, tests, and connection documentation. ## Verification - The current head is fully green in GitHub CI, including build, typecheck, all serialized/general server shards, all browser shards, policy, canary dry run, review, and security checks. - Live staging proof completed with a non-expiring GitHub App user token, selected-repository installation, repository add/remove refresh, managed MCP, native `gh`, HTTPS clone/push/delete, GitHub noreply commit attribution, signed merged-PR webhook acceptance, durable Cloud-to-instance delivery, and installation-access event processing. Temporary branches and temporary repository access were removed afterward. - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed before and after the rebase onto `origin/master`. - `pnpm build` passed. - The focused connector suite passed 285 tests after the rebase. - The full stable suite passed 5,790 tests and failed 22 tests across 8 general server files. The failures reproduced as shared-runner environment issues. They included `/tmp` versus `/private/tmp`, closed database connections, and invalid high ephemeral ports. The focused connection tests pass in isolation. ## Risks - Migrations add agent grant subjects and a durable connection-event inbox. Migration numbering and safety checks pass. - A raw GitHub user token enters the agent process for Git and `gh`. Per-tool Ask-first controls cannot limit those shell operations. The UI warns users about this boundary. - GitHub App user tokens can be non-expiring. Paperclip performs a continuity check every 30 days, but provider revocation still requires a reconnect. - The webhook path accepts only signed and bounded payloads. It stores a minimal normalized record and no raw provider payload. - GitHub repository permissions remain authoritative. Removed access can make a cached repository count temporarily stale, but runtime access fails immediately. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, browser control, and multi-file repository editing. The context window size was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7b094724e6 |
fix(runner): recover native sessions across restarts (#12845)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner keeps durable run and provider state outside one server process. > - A server restart can leave that runner alive or can interrupt it after a provider checkpoint. > - The old startup path used handoff intent and PID evidence, but it did not reconstruct native ownership. > - That gap could block the issue, create a replacement run, or start duplicate provider work. > - This pull request adds durable same-run recovery for coordinated and uncoordinated restarts. > - The benefit is exact recovery of the run, runner, session, provider, steering, and finalization state. ## Linked Issues or Issue Description Refs #9628. That pull request added earlier local-adapter hot-restart work. This change adds native PRP authority reconstruction and same-run provider resume. Refs #10935. That pull request handles missing hot-restart snapshots. This change also supports hard restarts with no snapshot. Refs #11624. That pull request prevents unsafe retry after an adopted legacy process exits. This change reconciles native terminal evidence before provider recovery. Refs #12070. That pull request improves process liveness checks. This change also binds recovery to a process-start fingerprint and fails closed on ambiguity. **What happened?** The server could record hot-restart intent, but startup did not rebuild native runner ownership. A live runner could not re-register its PRP authority. A dead runner could not resume the exact native and provider session on the same heartbeat run. Generic recovery could then block the issue or create replacement work. **Expected behavior** A live native runner must reconnect with the same PID and logical identities. A dead runner must resume the same durable session and heartbeat run with only a new operating-system PID. A proposed or terminal result must finalize once before any provider turn starts. Ambiguous process or session evidence must stay blocked without a signal or duplicate spawn. **Steps to reproduce** 1. Start a Paperclip Runner heartbeat and wait for an active provider turn. 2. Restart only the Paperclip server, with or without a hot-restart marker. 3. Observe that the old startup path does not reconstruct the native control-plane authority. 4. Kill both the server and runner after a provider checkpoint. 5. Observe that the old path cannot resume the exact native session on the original heartbeat run. **Paperclip version or commit** The defect was reproduced from commit `1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto the current `master`. **Deployment mode** Local development and self-hosted server deployments that use the local Paperclip Runner. ## What Changed - Added correlated hot-restart requests and version-compatible native handoff fields. - Added controller boot identity, process-start identity, controller generation, recovery state, request id, and bounded history to the native finalization ledger. - Added transactional recovery claims for live-runner reattach, dead-runner resume, and incomplete bootstrap. - Added fail-closed ownership takeover rules and process identity validation. - Added live runner adoption to the local runner transport without a duplicate spawn. - Added same-run provider checkpoint resume and legacy retry-row compatibility. - Reconciled proposed and terminal results before runner or provider recovery. - Bound the HTTP and PRP listener before startup recovery and delayed scheduling and generic reapers until classification completes. - Added restart-aware health diagnostics, run-log recovery transitions, durable runner diagnostics, and bounded shutdown finalizer draining. - Moved restart-survivable diagnostics into runner-owned, pre-redacted bounded writes; raw stdout and stderr are never persisted. - Added process-start fencing for controller, runner, and provider PIDs; startup classifies every candidate without an implicit cap. - Added crash-recoverable, contention-safe development restart-request coordination and failed-startup listener cleanup. - Added a credential-free real-process restart suite for eight restart, scale, and identity scenarios. - Documented native restart operation, persistence, diagnostics, and verification. ## Verification - The documented native restart commands passed. They ran eight real-process/database recovery scenarios and the live runner adoption transport test. - Native executor tests passed: 111 tests. - Heartbeat recovery tests passed: 124 tests. - Hot restart, health, and shutdown tests passed: 52 tests. - The broader affected server suite passed: 350 tests. - Focused native recovery and startup tests passed: 49 tests. - Runner transport and control-plane tests passed: 63 tests. - Runner-owned diagnostic tests passed for write-time bounding, credential redaction, private file modes, and raw stream non-persistence. - Development restart coordination tests passed: 11 tests. - Database migration checks and the partial-application/replay regression test passed. - Server, database, and Paperclip Runner typechecks passed. - `git diff --check` passed. - Full Paperclip PR CI passed, including build, canary, all five general server shards, all five serialized server shards, all three browser E2E shards, workspace suites, and release-registry verification. - Greptile completed at 5/5 with no outstanding findings, recommendations, follow-ups, or open review threads. ## Risks - Moderate risk. This changes startup ordering and ownership transfer for active native runs. - The migration adds nullable columns and does not rewrite existing rows. - Recovery fails closed when process or durable session identity is incomplete or contradictory. - The first implementation supports the local Paperclip Runner. Remote targets keep their existing behavior. - The real-process suite covers cleanup and asserts that no runner or provider process survives each test. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime did not expose a more specific model revision or context-window size. Repository editing, shell execution, database tests, and real-process test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a0028d7e1b |
chore(deps-dev): bump vitest from 4.1.10 to 4.1.11 (#12262)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.10 to 4.1.11. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v4.1.11</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Revive global concurrency limit for test lifecycle [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> and <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10992">vitest-dev/vitest#10992</a> <a href="https://github.com/vitest-dev/vitest/commit/5146df80b"><!-- raw HTML omitted -->(5146d)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Encode iframeId in tester iframe URL [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a>, <strong>Pduhard</strong> and <strong>Claude Opus 4.8</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10955">vitest-dev/vitest#10955</a> <a href="https://github.com/vitest-dev/vitest/commit/10b2cd201"><!-- raw HTML omitted -->(10b2c)<!-- raw HTML omitted --></a></li> <li>Trigger playwright/chromium gc on lower disk availability [backport to v4] - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>OpenCode</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10951">vitest-dev/vitest#10951</a> <a href="https://github.com/vitest-dev/vitest/commit/9851dbc41"><!-- raw HTML omitted -->(9851d)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>mocker</strong>: <ul> <li>Restrict redirect mocks to the fs allowlist [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10974">vitest-dev/vitest#10974</a> <a href="https://github.com/vitest-dev/vitest/commit/fe5a11d3c"><!-- raw HTML omitted -->(fe5a1)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v4.1.10...v4.1.11">View changes on GitHub</a></h5> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/9bd8d464e6328c567c2dbcd8fdd977d57a9425c2"><code>9bd8d46</code></a> chore: release v4.1.11 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10995">#10995</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/9851dbc41c286a30abfb6b29cce65f3e5b7b40a1"><code>9851dbc</code></a> fix(browser): trigger playwright/chromium gc on lower disk availability [back...</li> <li>See full diff in <a href="https://github.com/vitest-dev/vitest/commits/v4.1.11/packages/vitest">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
fdf8c8464d |
feat(runner): add managed provider backends (#12699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner provides durable, provider-neutral agent execution. > - The current stack supports qualified local providers but omits the managed provider paths from the integration branch. > - Claude Managed Agents and AWS AgentCore need explicit profile qualification, durable recovery, usage accounting, and cleanup controls. > - This pull request adds those managed backends as the third part of the Runner parity stack. > - The benefit is managed execution without weakening the default-off Runner rollout gate. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: Runner, server orchestration, database profiles, CLI, and adapter configuration UI. **Problem or motivation** The current Runner stack cannot select or execute the managed Claude Agents API or AWS Bedrock AgentCore Harness backends. It also lacks qualified profile storage and recovery checks for those remote resources. **Proposed solution** Add qualified managed and remote profiles, API and CLI management, exact provider selection, durable lifecycle handling, cumulative usage accounting, bounded cleanup, and retention acknowledgement. Keep `enableNativeRunner` default-off. **Alternatives considered** A direct copy of the old integration branch was rejected because its provider contracts, model values, credential flow, and migration history no longer match the current base. A single large parity pull request was also rejected because stacked review keeps each subsystem bounded. **Roadmap alignment** This continues the existing Runner architecture and rollout work. It does not introduce a separate execution system. **Additional context** This pull request is based on the merged #12691 and #12685 stack. It also closes the delayed security-review findings reported on #12691 by binding qualified ACPX and OpenCode launch artifacts to the bytes actually executed. A GitHub search for managed agent, AgentCore, and Claude managed work found no duplicate public issue or pull request. ## What Changed - Add Claude Managed Agents and AWS AgentCore provider executors to runnerd. - Add qualified managed and remote profile storage, routes, OpenAPI contracts, CLI commands, and migration 0237. - Validate profile ownership, enabled state, exact qualified revision, model, agent version, and secret binding before persistence and recovery. - Persist durable provider session and owned skill state for restart-safe cleanup. - Reconcile uncertain create responses and delete remote sessions before owned skills. - Track cumulative provider usage and enforce positive session spend caps. - Recover interrupted AgentCore usage at the next turn boundary by charging the prior invocation ceiling exactly once; keep the session gated until an explicit monotonic budget raise. - Isolate AgentCore AWS configuration from host profiles and credential-process/SSO configuration while preserving workload identity. - Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact paths; remove the ambient executable override. - Snapshot and content-verify ACPX and OpenCode commands, scripts, and provider executables before launch. Linux executes sealed inherited descriptors; macOS uses authenticated private snapshots with retry-safe rematerialization at the spawn boundary. - Persist canonical ACPX and OpenCode launch-profile digests, reject drift across fresh recovery, and make recovery failures sticky. - Close and journal unsafe ACPX active-turn recovery before any provider bootstrap or reconnect. - Add managed provider fields to the Runner configuration UI and permission projection. - Preserve the default-off `enableNativeRunner` experimental flag. ## Verification - `pnpm -r typecheck` - `pnpm build` - Focused managed server, database, CLI, Runner TypeScript, Rust, Claude, AgentCore, ACPX, OpenCode, process-supervisor, and durable-recovery tests passed. - `cargo test -p paperclip-runner-core --lib --locked` (160 tests) - `cargo check --workspace --all-targets --locked` - Native Codex integration tests passed (60 tests); native provider tests passed (7 tests); server native-runtime tests passed (87 tests). - Verified-launch replacement, nested-spawn retry, exact-version, profile-drift, sticky-failure, and no-bootstrap active-recovery tests passed. - `git diff --check` - The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust workspace lockfile adds the approved `rustix` dependency used for safe descriptor handling while `#![forbid(unsafe_code)]` remains enabled. ## Risks - The provider APIs can change while they are in beta. Exact qualification and fail-closed recovery checks limit drift. - Remote cleanup can fail after a partial create. Durable ownership inventories and retry-safe deletion preserve recovery state. - Migration 0237 adds profile tables. The generated migration and snapshot pass the repository migration checks. - Managed execution can incur provider cost. Positive default spend caps and explicit retention acknowledgement limit accidental use. - An interrupted AgentCore invocation without final metadata is conservatively charged to its active session ceiling. This can overstate cost, but cannot undercount it; later work requires an explicit budget increase. - Linux qualified launches use sealed memory descriptors. macOS lacks executable-descriptor APIs, so the runner uses owner-only private snapshots and minimizes linked-path lifetime; hostile same-UID processes remain outside the documented local-host trust boundary. - The global Runner feature remains default-off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4b6de5327e |
Remove cheap model profiles (#12683)
## Thinking Path > - Paperclip manages agents that use different model providers and adapters. > - Paperclip must keep agent execution rules clear and predictable. > - The cheap-model profile added a second execution mode across adapters, task recovery, APIs, and the UI. > - That mode increased configuration and recovery complexity. > - This pull request removes the cheap-model profile as a product feature. > - The benefit is one model-selection path for normal work and recovery work. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change simplifies model selection across agent configuration, task execution, recovery, and adapter capabilities. **Current behavior** Paperclip exposes cheap-model profiles in adapter metadata, agent runtime configuration, task overrides, recovery rules, APIs, and the board UI. Recovery work can select a different model profile from the agent's configured model. **Proposed behavior** Paperclip uses the agent's configured model for normal work and recovery work. Status-only recovery stays limited to coordination work. The API rejects legacy model-profile configuration. A migration removes stored model-profile values from existing agent, issue, and historical revision records. **Reason and benefit** One model path reduces configuration, API, UI, and recovery complexity. It also prevents status recovery from becoming a separate product-level model-routing feature. **Breaking changes** This change removes model-profile fields and adapter capability metadata. Existing stored model-profile values are removed by an idempotent migration. The validators reject new legacy profile values with clear errors. ## What Changed - Removed model-profile types, adapter capabilities, API fields, and model selection logic. - Removed cheap-model controls from agent and task UI surfaces. - Kept status-only recovery limited to coordination context while normal continuations use the configured agent model. - Added an idempotent migration that removes stored model-profile values from agents, issues, and configuration revisions without changing issue update timestamps. - Updated tests and product documentation for the single-model behavior. ## Verification - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` completed with 5,607 passing tests and 8 environment-sensitive failures in unrelated fixed-port and database-deadlock suites. The same failures repeated in an isolated rerun. CI is the final clean-room result. ## Risks - This is an intentional breaking change for clients that send model-profile fields. - The migration changes legacy agent, issue, and configuration-revision JSON. It is idempotent and preserves unrelated fields and issue update timestamps. - The change is cross-cutting because the removed feature existed in adapters, shared contracts, the server, plugins, and the UI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. Reasoning and tool use were enabled. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
25cf079ec5 |
feat(runner): add Codex-native application integration (#12591)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package is useful only when the application can start, observe, and recover a native Codex run safely. > - Existing direct adapters must keep their current execution and finalization paths. > - The application boundary therefore needs additive persistence, authorization, coordination, and recovery behind an explicit experimental adapter. > - This pull request adds that Codex-only boundary without activating generalized providers, remote environments, or the later task/SDK surfaces. ## Linked Issues or Issue Description **Subsystem affected** Shared contracts, database persistence, adapter utilities, server native-runtime services, and the experimental Paperclip Runner adapter. **Problem or motivation** The already-landed runner package has a qualified Codex path, but the application needs durable native-run state, guarded runtime selection, authenticated coordination, tool security, finalization, and recovery before the experimental adapter can be exercised safely. **Proposed solution** Add a Codex-only `paperclip_runner` application path behind the existing default-off native-runner setting. Bind native state and coordination to company/run identity, preserve persisted-run recovery, and leave every direct adapter on its existing legacy execution path. **Alternatives considered** The earlier stack boundary introduced a generalized executor and remote-environment lifecycle here. That made this PR depend on implementations in higher PRs and changed reusable sandbox behavior globally. Those pieces are now deferred together to #12592. **Roadmap alignment** ROADMAP.md does not list a conflicting native-runner integration project. This change adds the application boundary for the existing Runner architecture. ## What Changed - Added native run/result/finalization/provider-trace persistence, shared validators, and idempotent migration/replay coverage. - Added guarded Codex-only runtime selection, authenticated PRP coordination, recovery, finalization, and interaction services. - Added run/company-bound tool-gateway authorization, credential redaction, SSRF protections, and replay-safe behavior. - Added the explicit `paperclip_runner` adapter behind the default-off rollout setting. - Preserved legacy answered-question wake projection and direct-adapter execution/finalization paths. - Hardened cancellation so only owned in-memory child processes are signaled; persisted recycled PIDs/process groups are never trusted. - Retained the narrow Claude ACPX isolated-context security follow-up discovered after #12590. - Deferred the generalized executor, provider ingress, remote lifecycle, SDK/lab/eval work, release-process changes, and lockfile. ## Verification - Changed-file delta against `master`: 133 files. - GitHub Actions is the authoritative verification environment for this PR. - Full CI, security, and Greptile review will run on this lowest unmerged stack PR. - Local tests/build/typecheck were not run because this checkout is resource constrained. - Static diff/reference checks pass, and `pnpm-lock.yaml` is unchanged. ## Risks - This touches central heartbeat and agent-route code, so legacy compatibility is the primary risk. - Runtime selection remains Codex-only and explicit; direct Codex, Claude, OpenCode, process, HTTP, and plugin adapters remain on their existing paths. - Fresh native starts fail closed while the rollout flag is off; persisted native records remain readable and recoverable. - Cancellation, company/run binding, tool calls, status decisions, and completion writes are guarded or replay-safe. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI and security gates are green - [ ] Greptile is 5/5 with no open actionable findings - [x] I will address all Greptile and reviewer comments before merge ## Stack - Position: 3 of 5 overall; lowest of 3 currently unmerged - Base: `master` - Previous: [#12590](https://github.com/paperclipai/paperclip/pull/12590), qualified Claude ACPX runtime — merged - Next: [#12592](https://github.com/paperclipai/paperclip/pull/12592), generalized Codex executor, task experience, and developer SDKs --------- Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
d387cc0ff0 |
feat(connections): add managed external MCP connectors (#12346)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connection intents need secure provider implementations to complete setup. > - Some providers use managed OAuth or external credential brokers. > - Those tokens must stay out of durable Paperclip state and fail closed when refresh fails. > - This pull request adds managed connector backends and the required storage contract. > - The benefit is safer provider setup with governed credential lifecycles. ## Linked Issues or Issue Description Refs #11965 This is stack 8 of 11. It depends on stack 7 and replaces another reviewable part of #11965. ## What Changed - Add managed Google Workspace and external connector backends. - Add Vercel Connect support without storing provider bearer tokens. - Add replay-safe migration 0232 and its generated snapshot. - Fail closed and clear stale token bindings when organization OAuth refresh needs reauthorization. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 194 tests passed. - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` - `pnpm exec vitest run --project @paperclipai/server server/src/services/remote-url-credentials.test.ts` (5 passed, including URL userinfo vault extraction) ## Risks - Broker metadata errors can block provider setup. - OAuth refresh failure disables the shared organization connection until reauthorization. - Migration 0232 is generated, ordered after 0231, and safe to replay. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the public source pull request with `Refs #` - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |