mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 16:35:27 +02:00
fcff93abdd915eff5590c86a8d4324c7f2a9f7cd
4873
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fcff93abdd |
Merge provider-routing core from master before UI qualification
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0d43238bae |
fix(setup): keep sign-in environment selection and cleanup evidence reachable
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9f7057e122 |
feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use a harness, a model, and a credential to run tasks. > - Connections already store credentials and control who can use them. > - Custom providers also need an endpoint and a supported API format. > - A per-agent endpoint would duplicate credentials and access rules. > - This pull request stores routing on the connection and projects it into the harness. > - Isolated credentials and protocol checks keep the selected connection authoritative. ## Linked Issues or Issue Description Refs #37, #13083, #14104, #14565, #12692. #14016 is a reference only. This PR has its own schema, vault persistence, routing validation, runtime projection, and tests. None of the five commits in #14016 is an ancestor of this branch. We do not depend on or plan to merge it. #14967 addresses task-pinned account pools. #14422 addresses another provider integration. This is the first of two linked PRs. Merge this connection change before #15341, which refines agent setup and adds the qualification harness. The split keeps each review below 100 changed files. Provider catalog entries have usable setup forms in this PR. Local browser subscription sign-in is included. ## What Changed - Store non-secret routing metadata on AI connections. Vault provider API keys, including Bedrock bearer API keys. Reject general AWS access keys. - Enforce company, owner, human audience, agent access, connection status, and protocol checks before resolving credentials. Keep reconnect destinations immutable and retain connection identity during key rotation. - Project OpenRouter and compatible custom endpoints into Codex, Claude, OpenCode, and local Hermes. Carry these settings through both legacy and native runner transports. Clear conflicting host credentials and redact keys from diagnostics. - Preserve older OpenRouter accounts and native personal defaults. Add Google API-key accounts and migration `0306` for the two provider-default constraints. - Run local Claude and Codex subscription sign-in behind the existing browser sign-in card. Use private attempt homes and owner-bound completion instead of a copied terminal command. - Seed isolated Gemini authentication and preserve OpenCode workspace permissions. Keep the selected connection authoritative. The independent Gemini and Grok workflow fixes are in #15341. - Keep native OpenCode custom gateway keys in a runner-owned selected-model proxy; the harness config contains only a session-scoped capability. Honor runtime outgoing proxy and certificate settings. Preserve streamed responses and revoke the proxy on close or startup failure. - Allow ordinary members to connect native personal accounts before an agent exists. - Repair routed accounts from task cards using the saved provider destination, protocol, model aliases, and connection identity. - Add provider catalog definitions, model discovery, pinned logos, and complete native and routed setup forms. Allow a personal routed connection before a new agent exists. Keep endpoint authentication keys out of Hermes terminal children. - Recover cancelled or restarted browser sign-in with a clear restart action. Support no-auth endpoints without a vault credential. Add isolation and recovery regressions and runtime documentation. ## Verification - Updated with `origin/master` at `22a3ea341`. Migration `0306` follows the new master migration and passes migration and snapshot checks. - The integrated connection regressions passed 152 tests and 50 native OpenCode driver tests, including key-free child-shell configuration reads, authenticated/no-auth forwarding, streaming, model/path restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass. Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34 executable also completed a turn through the proxy against a local synthetic provider; the reusable key was absent from its config. A second real-executable smoke passed with an HTTPS CONNECT proxy and runtime-specific synthetic certificate trust. Certificate-file and certificate-directory regressions pass. - Task-card repair passed 48 tests, including OpenRouter, Bedrock, and custom gateway reconnect cases. UI typecheck and token gates passed. - The prior core regression set passed 133 tests across new-agent setup, provider forms, browser sign-in, routing projection, and connection authorization. Token gates and UI typecheck passed. - Full workspace typecheck and production build passed again after the latest integration and credential-proxy fix. The merged deterministic runner E2E suite passed 1,400 Vitest tests and 128 Node tests. - Full workspace typecheck passed on the prior linked combined implementation. Production build, Storybook build, 1,316 browser-harness Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the latest master integration and regenerated migration. This exact head passed 54 remote checks with four expected skips and Greptile 5/5; no review threads remain open. An unchanged server fixture had a random six-character issue-prefix collision on its first attempt. All 245 tests passed locally and the single CI retry passed. - A provider-free terminal check used the cited supported Hermes source and dummy keys. Gateway and OpenRouter terminal children could not read the selected key. - The broad local Vitest attempt passed 15,442 tests but was not green. It had an embedded-Postgres startup failure, an HTTP logger timeout, an origin socket error, and a browser cancellation wait timeout. The cancellation wait was corrected. The relevant connection tests and the full origin test file passed separately. Latest-head CI must pass before merge. - Prior credential-backed acceptance exercised task creation, tool use, artifact delivery, completion, and context-dependent follow-up. Claude legacy and native runners passed Bedrock with `us-east-1` and `us.anthropic.claude-sonnet-4-6`. - Historical local qualification retained 43 passing API/gateway cells out of 46. Those attempts span earlier builds. They do not qualify this exact commit or staging. All subscription combinations and staging remain unqualified. - Verify native subscription and API-key setup. Connect a regular provider catalog row. Verify an incompatible harness and a changed reconnect URL are rejected. Use #15341 for the complete browser campaign. ## Risks - Migration `0306` changes two check constraints. It preserves rows and is safe to reapply. It takes normal constraint-change locks. - Credential projection touches several harnesses. CLI upgrades can change provider configuration and session behavior. - The native OpenCode proxy adds a loopback hop, pins requests to the selected model, limits request bodies to 16 MiB, rejects redirects, and expires at session close. It prevents reusable keys in the child configuration; it is not an OS isolation boundary against a process debugger running as the same user. - Custom endpoints must be reachable from the agent environment. Saving a connection does not prove connectivity. Bedrock keys require rotation before expiry. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Provider overloads and an unresolved follow-up timeout also affect live Gemini qualification. We have not patched the installed CLI or marked those cases as passing. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS identity, arbitrary auth headers, and custom routing for other harnesses are excluded. - These PRs do not establish production or staging qualification for every provider and login method. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7be51f30e1 |
docs(release): canonicalize stable notes for v2026.1005.0 (#15357)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release process keeps stable notes under a beta-keyed path during the soak and renames them to the versioned path after the stable ships > - Stable v2026.1005.0 is published. The `canonicalize_stable_notes` job pushed the rename branch but it does not open a pull request > - Until the rename merges, `releases/v2026.1005.0.md` does not exist on master and the announcement links do not resolve > - This pull request merges the workflow's rename commit. It is a pure rename with no content changes > - The benefit is that the repository returns to the canonical release-notes layout and the notes link in the announcements resolves ## Linked Issues or Issue Description **Issue type** Missing content **Where is the issue?** `releases/` — the stable notes for v2026.1005.0 still live at the beta-keyed path `releases/beta/v2026.1002.0-beta.0.md`. **What's wrong?** The `canonicalize_stable_notes` job in the stable release run pushed branch `release-notes/v2026.1005.0-canonicalize` with the rename, but it does not open a pull request. The versioned path `releases/v2026.1005.0.md` does not exist on master until this merges. **Suggested fix** Merge the workflow's rename commit. The notes were already corrected before the stable dispatch in #15251, so no content change is needed here. ## What Changed - Renamed `releases/beta/v2026.1002.0-beta.0.md` to `releases/v2026.1005.0.md` (workflow commit `ecc10236`, authored by `github-actions[bot]`) - No content changed. This is a pure rename. The file already carries the correct header, commit count, and Contributors section from #14928 and #15251 ## Verification - The compare view for this branch against master shows one commit and one file with status `renamed`, zero line changes - The GitHub Release body for v2026.1005.0 is identical to this file (one trailing newline differs) - `git rev-list --count v2026.1001.0..v2026.1005.0` returns 179, which matches the Contributors section - After merge, https://github.com/paperclipai/paperclip/blob/master/releases/v2026.1005.0.md returns 200 ## Risks - Low risk: a documentation-only rename with no content changes. ## Model Used - Claude (Anthropic), model ID `claude-fable-5-1` (Claude Fable 5.1), extended thinking enabled, tool use via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>canary/v2026.1006.0-canary.20 |
||
|
|
f485863c21 |
Merge verified provider-routing task-card repair
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9d964c8d91 |
fix(connections): preserve provider routing during task-card repair
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c582338692 |
merge: preserve provider setup alongside master assistant connections
Combine the provider and public MCP qualification matrices and retain both connection catalog regression groups. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a97c39874f |
merge: align provider migration with latest master schema
Regenerate the provider-default constraints as migration 0306 after master public MCP invitation migrations; preserve replay safety and generated metadata. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a96b0722e1 |
merge: include proxy and certificate trust regression coverage
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
bb11e5779a |
test(opencode): cover selected certificate trust and proxy bypass
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0adf8e5fd8 |
merge: include custom certificate trust correction
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e76affdd11 |
fix(opencode): copy default trust before extending runtime certificates
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a150e1bcfe |
merge: include runtime proxy preservation fix
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5b6945401c |
fix(opencode): retain runtime proxy and certificate settings
Use isolated Node HTTP and HTTPS agents for authenticated gateway requests, preserving proxy fallbacks, bypasses, and selected certificate trust. Keep harness loopback traffic outside outbound proxies. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
22a3ea3414 |
Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Public MCP lets people use their organization from an external assistant. > - Operators need a visible control for this experimental access. > - Hosted users should select an organization once and then approve its permissions. > - This pull request adds the setting, invitation-first setup, and browser or device consent. > - Connections provides a copyable invitation with public instructions that grant no access. > - Users reach browser consent from their assistant and return to inspect or revoke access. ## Linked Issues or Issue Description Builds on merged foundation #14846. This PR now targets master. Related settings convention: #13905. **Current behavior** The preview uses an environment variable to enable MCP. Hosted consent repeats organization selection. Assistant access has no entry in Connections, so users must already know the endpoint and how to reach consent. **Proposed behavior** An administrator enables Settings → Experimental → Assistant connections (MCP). A hosted connection shows the selected organization and its icon, then asks for permissions. Requested write access starts checked when the user’s role permits it; the user can opt out before connecting. Direct instance connections show an organization picker with the first available organization selected. The selection stays fixed across refetches and still requires an explicit Connect action. Connections includes Assistant Connection (MCP). Its setup page explains the canonical endpoint, client configuration, browser authentication, and connected access. It connects as the current person and does not select or impersonate an agent. **Reason and benefit** Operators manage access with the other experiments. Users select one organization, and both the UI and server enforce that choice. **Breaking changes** The old enable variable has no effect. Preview operators must enable the setting once. Apply the additive consent-request migration before deploying the tenant, then deploy the compatible Cloud broker. Existing direct requests and grants keep their behavior. ## What Changed - Simplify OAuth and device consent: show the Paperclip logo beside “Connect {client} to Paperclip”, fall back to “your assistant”, and show the identifying origin plus its favicon below, with the callback URL also visible when different. Remove the hosted-organization creation action. Default to the first available organization without silently changing it on refetch; preserve company restrictions and write opt-outs. Keep the button row contained on narrow screens. - Make Copy invitation the primary action, using the shared animated AgentSetupPrompt and a collapsed manual setup section with icon-labeled line tabs. Remove redundant link actions, copy-status text, the extra first-prompt well and revocation explanation from the setup page. Serve shared version-aware HTML and Markdown instructions without private organization data. - Support guarded Client ID Metadata Documents alongside dynamic registration, and include authorization response issuer identification. - Add RFC 8628 device authorization with separately hashed codes, expiry, shared request quotas, persistent polling backoff and atomic redemption. Reuse human consent, role checks, scoped grants, audit and revocation. - Add CLI device login and a local stdio bridge. Store credentials separately with private permissions and serialize rotating refreshes. - Add device consent stories and five cold-start paid Product E2E cases with independent grant, configuration and durable-work assertions. - Add `enablePublicMcp` to the settings validator, normalizer, feature catalog, and toggle UI. Check it live for OAuth, tools, subscriptions, and event delivery. Keep connection management and revocation available while disabled. - Default the MCP origin to the existing auth public URL, with strict validation and an explicit override. - Persist the optional OAuth `company_id` restriction. Describe only that company and reject approval for any other company, even if the person belongs to both. Keep active-membership and role checks. - Show the Paperclip icon and a large organization icon during consent. Return the saved company logo through the company-scoped request response and reuse the standard fallback icon. Use the requested concise permission labels: “Read all of your Paperclip data” and “Allow write access and creating tasks as me”. Use concise permission copy, retain a compact client and callback-origin disclosure, and remove the footer link. - Default requested write access on for eligible roles. Preserve opt-out across organization changes and refetch, reset defaults for a new request, and submit read-only access when the request or role does not allow writes. Align the shared checkbox with its label. - Use organization wording in consent, management, settings, and walkthroughs. Keep the organization fixed for hosted requests and retain direct-instance choice. - Add an Assistant Connection (MCP) card to the Connectors catalog, a setup page in the app shell, and a return link from Experimental settings. Include Codex, Claude Code, OpenCode, and generic remote MCP instructions. - Read the live gate and canonical server URL through authenticated setup metadata. Show only the current person’s grants for the selected organization, refresh after consent, and support revocation. Surface catalog status failures with an explicit retry action; do not present them as an empty connection list. Opening setup grants no authority. - Start the eight guided chapters in Connections. Keep presenter notes and chapter controls around real product pages in the app shell. Explain the terminal, consent, delegation, retrieval, and revocation handoffs. Mark conversation examples as illustrative. Cover first use, client setup, connected, loading, and error states. Keep the existing consent and management stories. - Keep the paid-eval setup and browser helper aligned with the setting and consent button. ## Verification - Warm-standby integration fix `ec64ea05e`: public MCP ingress now follows the Cloud claim guard; MCP and discovery paths return 503 instead of SPA HTML while unclaimed. Event polling checks the in-memory claim before reading the persisted experimental setting. All 97 focused OAuth/Cloud tests and server typecheck pass, including new request and timer regressions for idle-before-claim and resume-after-claim behavior. Fresh review is 5/5 with no unresolved threads, and all security scans pass on this final head. All browser shards, typecheck, build, canary installation and other test groups passed on the first attempt. The unchanged Cursor sandbox default-command test timed out at 10 seconds; the exact test passed locally without edits in 587 ms. The single failed-job retry passed, with the original failure retained in workflow 37500711895. All 54 final-head checks pass on `ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs are intentionally skipped). - Final master integration `8457828fc`: merged foundation #14846 and current master, preserving the invitation changes and all 33 files from the two newer upstream changes. No migration renumbering was required. All 95 focused OAuth/Cloud integration tests, full recursive typecheck and token gates pass. All CI gates passed on that integration head; review identified the warm-standby issue fixed above. - Security-review fix `8c1d0b696`: commit shared global/per-source admission before outbound CIMD work, preserve failed-attempt receipts, and validate resource/scope before fetching. Added migration `0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive typecheck and production build pass. The security scanner passed that commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18 authorization requests sharing one proxy across two service instances use just two fetches. All 70 OAuth/metadata tests and server typecheck pass after that refinement. Final follow-up `8ebeae84c` reports admission-storage failures as retryable HTTP 503 instead of invalid client metadata. Its regression proves no outbound request before admission and successful retry after storage recovers. All 71 OAuth/metadata tests and server typecheck pass. Final-head security scanning passes; Greptile is 5/5 with no unresolved findings. CI passed all browser shards, typecheck, build, token gates and canary installation. One unchanged adapter-utils bridge test raced a response-file write (expected a JSON error, received the safe file-changed error). The exact test passed locally without edits. The single failed-job retry passed; the original failure is retained in workflow 37490609192. All 54 checks now pass on final head `8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final integration above now targets master. - Integration with current master: preserved the new Connections source filters and pagination, kept all eval suites, and regenerated the consent/device snapshots as migrations 0303/0304. All four MCP migration SQL hashes are unchanged from the staging versions. Full recursive typecheck and production build, 132 focused UI tests (including catalog filtering), 89 server authorization/settings tests, 26 migration tests, 120 eval calibration tests and token gates pass. Review follow-up `4820ce74c` also keeps active assistant grants in Installed, with pending/error recovery and revocation/company-isolation coverage. All 76 setup/catalog tests, UI typecheck and token gates pass after that fix. The unchanged signoff browser test timed out waiting for a heartbeat in CI at `4820ce74c`; the exact test passed locally without code changes, and the preceding CI head passed that shard. That same unchanged test failed at the reviewer stage in the next CI run. All five signoff tests passed three times locally (15/15), without test changes. All eight browser shards pass at final head `8ebeae84c`; no browser-test edits or failed-browser-job retries were needed. - Setup-page refinement at `9ab009178`: all 17 focused setup/consent tests pass, along with UI typecheck, production build, Storybook build and token gates. Browser exercised the shared prompt preview and client tab switching, and the updated InvitationCopied Storybook interaction checks its clipboard fixture. All final-head CI checks pass at `9ab009178`, with no unresolved review findings. Deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524. Verified the actual page, tab switching and line styling, removed actions/copy, and successful native copy/paste of the complete Butter invitation into a local-only test field. The existing Claude grant was left intact. - Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates pass. UI typecheck and production build passed again at `4e4d5e4d9`; Storybook build and eval-helper typecheck passed for `28101cf91`. Follow-ups let the primary button wrap on narrow screens, preserve a distinct callback URL, and use only bundled icons to avoid pre-consent requests to client-selected sites. Browser-verified the real consent component in desktop and 320px mobile stories, including default selection, write access and preserved opt-out. Updated E2E heading/default-selection helpers. All CI checks passed at `dc8e9fd11`, with review 5/5 and no unresolved threads. The Butter preview publication needed a retry because npm initially accepted the DB package before making it visible; the retry succeeded and `dc8e9fd11` deployed. Verified a fresh, unapproved native Codex CIMD request on Butter: default organization/write selection, known-client heading and icon, distinct callback origin, and removed creation action. No grant was approved for this UI check. Prior paid runs below retain their exact source provenance; this UI-only follow-up did not rerun paid qualification. - Source-pinned paid matrix at `2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign `local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config, unavailable host, denied consent and reconnect/later retrieval, with independent configuration/grant/task/run/document assertions. Original failures, transcripts, source fingerprints and billing remain retained. - Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`. Latest `0637b9f1c` shares that same guidance across HTML, Markdown and manual UI after review; generated Markdown is verified byte-identical to the paid-evaluated version. Shared build, server/UI typechecks, token gates and 63 auth/metadata tests passed again. Every CI gate passed at prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads. - Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval calibration tests, server/UI/eval typechecks, token gates and Storybook build passed. Full recursive typecheck and production build passed during implementation; CI also passed them at `2992ef271`. - Local full-suite limitations: a large-file Git streaming test times out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh MCP reruns passed, and the corresponding CI groups passed. No claim that the local full suite is green. - Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD consent; device CLI reaches verification/consent. New grants await human approval. Existing local OpenCode retrieved a saved result in a fresh conversation through its previously approved grant. - Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received the exact copied invitation, read public setup, configured its server and started PKCE consent. Its shell command timed out; background retry reached the client's own callback deadline while approval remained pending. Latest instructions cover that handoff. **No completed Butter read/delegation/result retrieval is claimed.** - Cloud companion https://github.com/paperclipai/paperclip-cloud/pull/672 passes checks/review and deployed. Anonymous setup and device-protocol routing verified. Core `2992ef271` deployed successfully and the actual Claude web flow now reaches consent. Its extra JWT-bearer metadata is filtered to implemented grants; unsupported token grants remain rejected. Final `0637b9f1c` deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300; live HTML and Markdown both contain the final guidance. The superseded instruction-only build was canceled before deployment. This is a core-only staging preview; private Cloud plugins are omitted. ChatGPT web is signed out, so browser connector use is unverified. - Screenshot gallery begins at Butter's dashboard and distinguishes real setup/pending consent from local reuse and fixtures. It records the timeout finding. New persistent access needs human confirmation before the remaining actual-client acceptance work. - Manual path: Connectors → Assistant Connection (MCP) → Copy invitation → paste into assistant → configure and start authorization → sign in and approve → verify `paperclip_connection` → delegate → retrieve the saved report later. - Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md` and `doc/public-mcp.md`. ## Risks - Apply additive, replay-safe migration `0304_curvy_shadow_king.sql` before using device authorization. The public setup link carries no credential. Device codes and tokens stay private; neither sharing instructions nor installing a plugin authorizes access. - Apply additive migration `0305_chubby_vin_gonzales.sql` before deploying the shared metadata admission gate. It retains at most 60 short-lived, hashed-source receipts per instance and rejects excess attempts with 429. - CIMD metadata fetching is a new external-input boundary. It requires HTTPS, exact client ID and redirect validation, bounded responses and guarded DNS/network access. Client names remain self-reported. - Device support is per-instance. The central Cloud broker retains its existing grant support. Host installation and tool reload capabilities vary by client; instructions describe manual settings and restart requirements. - Consent names the registered client in its heading and displays its identifying origin below. Known-origin icons are bundled; all other origins show a neutral site icon without contacting client-selected sites. Client names are self-reported; the callback origin is the recipient check. The Cloud chooser also displays the original client and receiving origin before tenant handoff. - A user who accepts the preselected write permission can create tasks and comments. Task creation and comments can start or wake agents and use execution budget; the consent label uses the concise wording explicitly requested by the maintainer. Scope requests, role checks, and the final Connect action still apply. - Migration `0303_supreme_garia.sql` adds one nullable UUID column with `IF NOT EXISTS`. Requests without a company restriction keep the direct-instance picker. The binding stays recorded if its company is deleted; consent then fails closed. - Deploy tenant support before the Cloud broker sends `company_id`. Unknown or inaccessible organizations must never fall back to a different company. - The setting defaults off. Disabling access does not cancel work already delegated. Existing tokens and unexpired subscriptions can resume when enabled again; revocation remains separate. - The catalog entry is visible for discovery while the feature is off. Setup instructions, OAuth, and tool execution remain gated. No access is granted by viewing the entry. - Assistant sign-in starts in the external client so it owns PKCE and callback state. Client command syntax can change and links to official setup documentation are included. - An authenticated instance and valid public URL are required. Hosting, paid execution, and store publication remain separate rollout steps. ## Model Used OpenAI GPT-6 in Codex, with tool use and code execution. The exact serving model version and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks pass; unrelated local full-suite timeouts are explicitly recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>nightly/v2026.1006.0-nightly.0 beta/v2026.1006.0-beta.0 canary/v2026.1006.0-canary.19 |
||
|
|
b7babda93a |
merge: include reviewed provider credential and setup fixes
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
7fed41a66d |
fix(connections): protect native gateway keys and allow personal provider setup
Keep custom OpenCode authentication in a session-scoped selected-model proxy rather than an agent-readable config. Permit ordinary members to create native personal accounts before an agent exists. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e2891ef538 |
merge: integrate current master with provider setup qualification
Preserve provider-connection and public MCP suites together, including their independent fixture documentation and matrix expectations. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
07ebd89ac2 |
merge: refresh provider connections on current master
Regenerate provider constraints as migration 0303 after the new master migrations, retaining replay safety and synchronized schema metadata. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0fe47882cf |
Allow configurable Runner listening ports (#15353)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner supports authenticated provider ingress. > - Each listening Runner currently binds port 43127. > - Concurrent Runner processes on one computer need distinct listening ports. > - This pull request permits an explicit listening port and retains the default. > - Providers can route each run without changing the Runner protocol. ## Linked Issues or Issue Description **Subsystem affected** Paperclip Runner launch and durable transport. **Problem or motivation** Two listening Runner processes in one network namespace cannot bind the same fixed port. The CLI accepts a port flag but rejects every value except 43127. **Proposed solution** Accept `--listen-port` values from 1 through 65535. Default to 43127 when the flag is omitted. Preserve the wildcard bind address, exact run path, PRP authentication, and secure frames. A warm attachment retains its existing listening port. **Alternatives considered** Separate network namespaces or a shared Runner daemon need more changes. Configurable launch ports preserve the existing process model. **Roadmap alignment** This extends existing Cloud / Sandbox agent support. The duplicate search found no matching Runner listener-port change. ## What Changed - Default an omitted listener port to 43127 and reject invalid values. - Validate configurable ports in the durable transport. - Reuse the selected port during warm attachment and reject port changes. - Cover default and explicit ports, invalid input, concurrent listeners, and warm attachment. - Update transport documentation. Daytona still uses its existing default port. ## Verification - Native `cargo test --locked --workspace` passed (two existing tests ignored). - Targeted listener and CLI tests passed, including executable launches on two concurrent ports and warm listener retention. - `pnpm -r typecheck` and `pnpm build` passed. - Complete GitHub CI is green, including general and serialized test suites, all browser shards, Runner Rust/Vitest lanes, builds, and release checks. - The additional full local `pnpm test:run` is still running; no final local result is claimed. - `git diff --check` passes. ## Risks An explicit port can already be occupied. Runner fails its bind without choosing a different port. Port allocation and ingress authorization remain provider responsibilities. No schema or PRP wire format changes. Existing explicit port 43127 callers continue to work. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository inspection, and code execution. The session does not expose the exact model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1006.0-canary.18 |
||
|
|
e34abee670 |
feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path > - Paperclip gives teams durable tasks, agent execution, budgets, and approvals. > - People also use assistants in Codex, Claude, and other MCP clients. > - Those assistants need a scoped connection that preserves the person’s permissions and attribution. > - Delegating a task must not turn the assistant into the assigned agent. > - This PR adds opt-in user OAuth, ten first-party tools, browser consent, and workflow packages. > - Paid product evals verify the resulting tasks, documents, attribution, retries, and access boundaries. > - The team keeps working after the assistant conversation ends. ## Linked Issues or Issue Description **Problem or motivation** A person cannot connect an external assistant to an existing team through browser consent and safely delegate durable work as themselves. **Proposed solution** Expose an opt-in `/mcp/paperclip` endpoint with individually described first-party operations. Bind every connection to a person, client, company, resource, and scopes. Reuse domain authorization and scheduling. Package shared team-review, delegation, and follow-up workflows for OpenAI/Codex and Claude. **Alternatives considered** Related PRs #9393 and #12549 cover earlier remote MCP and board-operator approaches. This change uses user OAuth and a bounded public catalog. It does not expose a generic executor, operator administration, static shared board credentials, or external agent execution. Registry listing work in #9851 is a separate distribution step. **Roadmap alignment** This maintainer-requested implementation extends the governed MCP gateway, activity attribution, durable work products, and hosted deployment direction in `ROADMAP.md`. It implements the first release of the saved design plan; external agent participation and granted third-party tools remain later releases. ## What Changed - Add MCP 2.0 discovery and task status/comment/document Events on the same authenticated endpoint. Persist subscriptions and delivery receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt callback material, recheck permissions/Cloud membership, and bound retries/expiry. Older MCP clients keep their existing tools. - Add discovery, dynamic client registration, S256 PKCE, resource validation, rotating refresh tokens, revocation, and company consent. Store credentials as hashes and recheck membership at execution. - Add tools for connection identity, agents/projects, task search/read/create, human comments, documents/deliverables, and pending-approval links. Preserve current domain permissions and scheduling. - Add durable mutation receipts across reconnects. Matching retries replay results; uncertain outcomes keep the same request ID and require inspection. - Add consent and connection-management pages, OAuth log redaction, shared plugin workflows, and separate OpenAI/Codex and Claude package outputs. - Add eight paid Product E2E cases across three models, independent durable-state grading, usage evidence, cleanup, and report integration. Add task-document guidance and regenerate the runner capability inventories. - Add migrations 0301 and 0302, the dated implementation plan, result notes, and direct-client setup instructions in `doc/public-mcp.md`. ## Verification - Merge integration `e180b1948`: resolved conflicts with current master, preserved both eval registries, regenerated capability catalogs, and regenerated migrations as 0301/0302 while keeping the original replay-safe SQL byte-identical. Local migration safety/snapshot tests (26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval catalog/grading tests (198) pass. Token and capability gates pass. Full recursive typecheck passed. Fresh Greptile review is 5/5 with no unresolved findings. CI is green on this exact head (55 successes, two intentional skips, one neutral result): one unchanged Cursor sandbox test timed out at 10 seconds, then passed locally in 856 ms. A single retry of that failed shard and the aggregate workflow passed. Merge remains blocked on the repository code-owner approval rule. Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55 successes, two intentional skips and one neutral result. [The earlier CI run](https://github.com/paperclipai/paperclip/actions/runs/36901592350) includes all test shards, browser tests, typecheck, build and canary dry run. Greptile was 5/5 on that commit with no unresolved review threads. GitHub still requires code-owner review under the repository merge rules; passing checks do not bypass that approval. Paid source fingerprints remain separate below and in the dated result note. - Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku 4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback, signature verification and report retrieval in a fresh conversation. A final Mini regression passes after the quota/status fixes. All evidence validates. Bounded tunnel startup retries occur before provider calls and remain visible; failed earlier attempts retain their original grades. - The earlier complete seven-case matrix passes **21/21**, with a separate **3/3** delegation regression. Two preceding matrices also passed 21/21 each. A complete 24-cell matrix including Events has not been run. [The dated results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain exact source fingerprints, failures, model IDs and partial costs. - Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass after merging master. Server typecheck passes after the final quota/status changes. Eval typecheck and all 892 eval-support tests pass. - All 33 real MCP/OAuth tests pass. The preceding combined MCP, redaction, private-address and DNS-rebinding run passed 129 tests; two later MCP regressions cover quota reuse and unchanged-status suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI then found a null checkout result in the existing concurrent-workspace path; logging now uses optional status access. All 12 closed-workspace tests and all 33 MCP tests pass after that correction. The exact-start event calibration exposed a timestamp gap; scanning now includes the subscription start, with all 33 MCP tests and server typecheck passing. These two narrow corrections follow the paid regression. - A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0 subscription/delivery, current membership loss, unsubscribe, legacy SDK tools, refresh and revocation. Its callback transport is a fixture with independent HMAC verification. The paid Events campaigns separately prove public HTTPS delivery. - Earlier component qualification passed UI 7,117 tests, CLI 502, shared 832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module boundaries, migration order and plugin regeneration passed. CI covers general/serialized suites, eight browser shards, runner checks, typecheck, build and canary dry run. - **Local full-suite limitation:** the earlier monolithic run was not clean. It encountered overlapping schema rebuilding, Mac database shared-memory limits and isolated CLI/fixture failures. Targeted reruns passed. The existing >32 MiB Git filename stress test still hit its 300-second Mac timeout. The additional serialized sweep stopped after 62 passing suites once CI passed. Original failures and partial logs remain; this PR does not claim a wholly green local monolithic run. - Local Codex CLI and Claude Code OAuth login and MCP SDK interoperability were verified. Public-store installation, actual ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted newcomer provisioning remain release gates. Enablement is moving to **Settings → Experimental → Assistant connections (MCP)** in the stacked follow-up [#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge both for the intended setup experience. This foundation branch alone still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set `PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and connect to `/mcp/paperclip`. Select a team and allow writes in browser consent. Configure an available agent and budget, then delegate and retrieve results later. For Events, rescan the deployed plugin catalog in ChatGPT Work Cloud; the host supplies its webhook credentials when the user asks to watch a task. See [the setup runbook](doc/public-mcp.md). ## Risks - Events are at-least-once and may arrive out of order. No replay cursor is advertised. Clients must refresh finite subscriptions, read current state and avoid comment feedback loops. Callback material uses the instance secrets master key; hosted subscriptions require the updated Cloud broker and are bounded to five minutes/the access proof expiry. - ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging subscription remain deployment gates. Local signed-webhook and paid model evidence does not claim those surfaces have been exercised. - Disabled by default. Merging adds schema and opt-in code; it does not deploy a public endpoint, publish a store listing, create a team, or start paid agents. - Migrations 0301 and 0302 are additive and idempotent. Their SQL is unchanged from the earlier preview numbers, so hash-aware upgrade reconciliation preserves prior staging applications. Normal instance upgrades must apply it before enabling MCP. - Task creation and comments can schedule paid agent work. Consent and tool descriptions disclose that effect. Revocation blocks future calls but does not undo delegated work. - Public deployments need edge rate limits and credential-safe logging. Internal dispatch is restricted to the closed catalog and carries a request-local verified actor. - Hosted onboarding requires the companion Cloud broker, encryption-key configuration, and tenant rollout. Self-hosted direct connections can use this PR alone. - Store acceptance and agent-mode participation are not claimed. Checked-in plugin endpoints are development defaults; rebuild packages for a real deployment before installation. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A more specific serving version and context-window size were not exposed by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`, `claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted/component checks; full local-run limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3a4939ba81 |
fix: stop provider qualification campaigns across their entire lifecycle
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
31da0558d1 |
fix: preserve Grok session metadata when history exceeds retention limits
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
797e578452 |
fix: restore provider workflow fixes and persist ACP resource errors
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
202c2d307e |
fix(server): leave unclaimed warm Cloud databases idle (#15314)
## Thinking Path > - Paperclip manages work performed by AI agents. > - Managed deployments prepare empty applications before an owner claims them. > - Those applications start database pollers even though no company can have work. > - Health probes also query SQL, so idle databases cannot remain suspended. > - This pull request adds an explicit standby marker for empty, unclaimed Cloud apps. > - The existing signed, durable claim resumes normal processing without restarting the app. ## Linked Issues or Issue Description **What happened?** An empty, unclaimed warm application runs recurring chat, email, plugin, heartbeat, cleanup, and reconciliation queries. Its health route also opens the database. This prevents idle database compute from suspending. **Expected behavior** An explicitly marked unclaimed application should keep its HTTP process and sandbox provider plugins ready while leaving the database idle. A successful signed claim should resume normal API behavior and background processing. Claimed and self-hosted instances should keep their current behavior. **Steps to reproduce** Start an empty Cloud-managed application and leave it unclaimed. Observe database activity while repeatedly requesting `/api/health`. Before this change, periodic queries continue without company data. Related: #15153 reduces allocation during chat polling. This change suppresses polling only for explicitly marked, empty, unclaimed Cloud apps. ## What Changed - Add `PAPERCLIP_CLOUD_WARM_STANDBY=1`. Check company emptiness once after restoring the persisted Cloud runtime identity. Missing Cloud configuration, existing data, or a persisted claim leaves normal processing active. - Gate recurring database pollers with an in-memory predicate. Keep startup preparation and sandbox provider plugin loading intact. - Serve unclaimed health probes without session or database reads and report `warmStandby: true`. Serve standby pages/assets directly from the UI router, bypassing session, bearer, tenant, and dynamic handlers. Refuse API requests and all WebSocket upgrades before authentication can query SQL or seed company data. - Exit standby after the existing signed identity assertion commits. Normal timers resume at their next tick; a restart restores the claim even with stale provider variables. - Document the marker, readiness semantics, rollout checks, and rollback. ## Verification - `pnpm -r typecheck` passed. A final server typecheck also passed after adding tests. - `pnpm build` passed. - Focused standby, signed claim, restart, health, static/Vite routing, hostname, HMR, and live-events suites: 72 passed after the review fixes. Includes real HTTP upgrade admission before/after claim. - `pnpm test:run` was attempted locally; both superseded runs were stopped after encountering checkout/platform failures. A clean-checkout rerun eliminated ancestor skill-directory lookup failures. The company-skills/runtime-cache families encounter macOS read-only-directory rename failures (`EACCES`); all three company-skills failures reproduce on unmodified base `bf14f803d5`. The initial full run also reported one native runner API test failure; an isolated comparison on both revisions was blocked by local embedded PostgreSQL startup failures. The full [Linux CI run](https://github.com/paperclipai/paperclip/actions/runs/37423263993) passed on final commit `5016c415ea`, including all server and workspace test shards, browser suites, typecheck, build, and release canary. This is not a claim that the full local suite passed. - Isolated full server with local PostgreSQL: after startup and connection expiry, 70 health probes, 70 page requests carrying valid synthetic tenant credentials, and 70 rejected WebSocket upgrades over 70 seconds observed zero app database connections. The signed claim completed in 62 ms and normal polling resumed (475 database transactions over 12 seconds). Restart with stale provider variables restored the durable claim. The latency is local-only, not a provider wake measurement. - Apex review: **5/5** on `5016c415ea`, both earlier threads resolved, no open recommendations. - No live-provider test or production deployment was performed. An actual database suspension/resume canary remains required before enabling the control-plane switch. ## Risks - Standby health reports HTTP readiness rather than current database connectivity. The signed claim still requires a durable database write; claimed health checks retain the SQL probe and 503 failure behavior. - Pollers resume at their usual intervals. A suspended database may add claim latency. Validate the real provider before enabling the marker. - Startup preparation and sandbox plugins remain loaded. New plugins or background loops must respect the same standby contract. - The marker is off by default. Remove it or set it to `0` and restart to roll back. No schema migration or claimed-workspace inactivity policy changes. ## Model Used OpenAI Codex, based on GPT-6. The exact serving snapshot and configured context-window size are not exposed in this session. Assistance included source review, TypeScript changes, command execution, and PostgreSQL tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted tests pass; full local suite limitations are documented above, and full Linux CI is green - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1006.0-canary.17 |
||
|
|
f2715e02bb |
fix(native): surface model capacity errors and retry automatically (#15347)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs preserve provider events and results, then decide task state. > - Codex can fail a turn because the selected model is temporarily at capacity. > - Paperclip displayed a generic native-session failure and left the task in recovery. > - A known capacity failure needs a clear message and a durable delayed retry. > - This pull request uses committed terminal evidence, the status-effect ledger, and existing dispatch gates. > - The task can continue automatically without an unbounded retry loop or a model switch. ## Linked Issues or Issue Description Related: #13993 adds launcher capacity deferrals before provider startup. This change handles a committed native Codex `serverOverloaded` terminal after provider work has started. **What happened?** A native Codex turn ended with `codexErrorInfo: serverOverloaded` and “Selected model is at capacity. Please try a different model.” Paperclip saved the result, showed a generic failure, and required recovery instead of waiting and retrying. **Expected behavior** Show the capacity error directly. Schedule bounded retries after a short delay. Preserve saved work and honor current execution gates. **Steps to reproduce** 1. Run a native Codex task. 2. Commit a `turn.failed` event with `serverOverloaded`, then accept a failed result for the same turn. 3. Inspect the task error and its recovery state. The regression test reproduces this without a paid provider call. **Paperclip version or commit** Reproduced against master `16b7db35ffa0f9a95913c8cbdeea3d595435691f`. ## What Changed - Classify capacity failures from committed runner events and the pinned execution identity. Preserve accepted results and display the specific capacity error. - Atomically persist one scheduled successor with the status decision. Retry after one minute, then two minutes. Share the existing failure budget and stop after two automatic retries. - Reuse dispatch gates for ownership, task holds, dependencies, budget, and locks. Wait for predecessor execution, finalization, and cleanup before claiming a retry. - Preserve pending reviewer authority and suppress retries after reassignment or a successor claim. - Consume the failed run's resume receipt and delivered wake input. Rebuild ordinary continuation from the failed run so explicitly resumed tasks can retry without borrowing one-run authorization. - Label scheduled retries “Model at capacity.” Add a Storybook example and avoid duplicate punctuation in the existing retry card. - Add regression coverage and document the runtime contract. ## Verification - Targeted recovery and UI suites: 125 tests passed. Additional final cleanup and lock regressions passed. - Latest-head continuation and authorization regressions: 57 tests passed, including all four capacity integration tests against an isolated PostgreSQL database. - Rechecked all four capacity integration tests with the full runner's isolated `PAPERCLIP_HOME`, config, temporary directory, and serial fork settings: passed. - `pnpm -r typecheck`: passed. Final server typecheck passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Browser: checked the real Storybook card. It shows the capacity message, automatic retry time, and existing Retry now action. - `pnpm test:run`: attempted and restarted after an interruption. The resumed run started before the final review correction and was stopped after recorded workspace/native test failures and the five-minute Git streaming timeout already documented on master. Exit 130; no complete local full-suite pass is claimed. Current-head isolated capacity tests and the complete CI suite pass. - A combined local recovery-suite attempt also hit PostgreSQL initialization failures at this macOS host's global shared-memory limit (32 slots). The focused recovery suites passed separately; final-head capacity tests also pass with the full runner's isolated environment settings. - Latest-head GitHub checks: all 56 checks green, including general and serialized test shards, browser E2E, typecheck, build, Runner verification, and canary dry run. No merge conflicts. - Greptile: 5/5 on `813a470b8c18c05aeb7e31e63e573c3a3a2f5cac`, with zero unresolved findings after fixing the resumed-task continuation issue. ## Risks - Capacity retries can repeat a task turn after partial work. They start a fresh provider session with task history and wait for predecessor cleanup. They do not resume the failed turn. - Retries retain the configured model unless an operator changes configuration. Persistent overload consumes the existing failure budget and then needs an explicit retry or model change. - Only run-bound native Codex `serverOverloaded` failures qualify. Usage-limit exhaustion, model/account incompatibility, unbound text, and unknown provider failures retain their current recovery behavior. - No schema migration is required. ## Model Used - OpenAI GPT-6 through Codex. The exact model ID and context-window size are not exposed to this session. Used reasoning, repository tools, code execution, and browser verification. No subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1006.0-canary.16 |
||
|
|
a259543902 |
docs(connections): explain encrypted history and qualification environment checks
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ec1e5a330b |
merge: apply core provider review fixes to setup and qualification
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
fe1da1ca25 |
fix(connections): recover browser logins and isolate Hermes gateway keys
Allow a personal gateway to be saved before a new agent exists. Update the new-agent connection selector regression. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
925d203823 |
test(connections): cover bounded history and fix Storybook selector types
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e1e83b51dd |
merge: synchronize provider UI and evaluations with current master
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
28c4615258 |
fix(evals): verify execution environment and cancel on hangup
Block unavailable requested environments, verify saved and actual execution targets, handle SIGHUP during paid campaigns, and correct the gateway Storybook submit selector. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
586950c9c9 |
Merge origin/master and renumber provider-default migration to 0301
Preserve the new master schema snapshot and use a later timestamp for the idempotent Google constraint migration. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3bb36afcd9 |
fix(connections): wire provider catalog setup and no-auth execution
Keep the credential form and compatible personal-account picker with their catalog definitions. Add no-auth runtime preparation and abort-aware Claude code input. Move independent Grok and Gemini live-workflow fixes to the qualification follow-up. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
20740a710d |
test(connections): use valid protocol contracts in no-auth cases
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c00dbd8ffe |
fix(connections): secure Grok history and repair no-auth and login cancellation
Store Grok history encrypted on the authorized task checkpoint. Restore regular files into a temporary home without a shared host symlink. Permit no-auth routes without a vault credential and keep their session identity stable. Abort pending Claude browser-code input so cancellation and timeout can dispose the login process. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
16b7db35ff |
Shorten planning skills and measure task decomposition (#15296)
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
44e4979d23 |
Capture project and repository update lifecycle events (#15306)
## Thinking Path > - Existing lifecycle capture records agent transitions and project creation. > - Project and repository edits need matching project update records. > - The record must commit with the mutation so failed writes cannot lose a hook. > - Creation with repositories is one creation, and repository replacement is one update. > - Archive-only changes preserve project state without creating a hook. > - Plugin consumption and provider behavior are separate work. ## Linked Issues or Issue Description **Problem or motivation** Project edits and repository/workspace changes lack durable lifecycle records. A future resource plugin needs those changes captured alongside the existing project creation hook. Archive-only changes must produce no hook. **Proposed solution** Allow project `update` records in the existing lifecycle journal. Record project and workspace mutations in their database transaction while holding the project row lock. Suppress intermediate workspace hooks during project creation and aggregate repository replacement. **Alternatives considered** Route-only hooks miss shared service callers. Recording after commit can lose an event. Emitting a hook for each child mutation exposes intermediate repository state. **Roadmap alignment** This completes project lifecycle capture begun in #15280. Plugin delivery, VM/volume provisioning, and backfill remain separate. Searches found no duplicate project lifecycle work; related #13306 concerns decision events on the in-process plugin bus. ## What Changed - Record project edits and workspace additions, updates, and removals as project `update` events. - Commit each event atomically with its mutation under the project row lock. - Keep project creation with repositories to one creation event and repository replacement to one aggregate update. - Ignore archive-only changes and retain workspace records. - Extend the journal action constraint and document project update capture. ## Verification - `pnpm -r typecheck` passed on the narrowed scope. - 57 tests passed across six lifecycle, project, repository, and chat-project suites. - Seven managed-sandbox workspace route tests and the CLI lagging-worktree migration regression passed (65 targeted tests total). - Full GitHub CI passed on `8ba6f97f22`; all required gates are green. - Greptile scored the final project-only commit 5/5 with zero unresolved review threads. - The branch is current with `master` and has no merge conflicts. - `git diff --check` and a local secret/PII scan passed. ## Risks - Apply migration `0300_chunky_chamber.sql` before running the new server. It permits project update actions and tolerates older JavaScript worktree backups that omitted the prior CHECK constraint. - Event-write failure intentionally rolls back the project or repository mutation. - Records contain identity and action; future consumers must load current authorized project/workspace data. - Plugin consumption, provider calls, volume cleanup, and backfill are outside this PR. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code execution, and tool use. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1006.0-canary.15 |
||
|
|
88e35f27fe |
feat(connections): add advanced setup and repeatable live qualification
Reuse connector rows, access controls, and agent setup components. Add persistent subscription, API key, and custom gateway choices with matching dropdowns and provider logos. Include grouped review stories and an opt-in browser campaign for local or staging targets with private credential and evidence handling. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
fa9b3d22a1 |
refactor(connections): keep runtime and browser-login changes independently usable
Keep the provider runtime PR below the review file limit. Preserve browser sign-in compatibility and provider branding here; ship advanced setup, stories, and the qualification harness in the linked UI PR. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1477d1ecea |
test: remove the no-op sequential describe modifier (#15286)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs the server and runner test suites with Vitest. > - Vitest 5 removes the deprecated `describe.sequential` property, so the pending Vitest 5 upgrade fails the type-check and test jobs. > - `describe.sequential` only changes behaviour inside a `describe.concurrent` suite, or when `sequence.concurrent` is on. > - This repository has neither, so the modifier changed nothing at run time. > - The benefit is that the Vitest 5 upgrade can land, and the test files lose a modifier that did no work. ## Linked Issues or Issue Description Refs: #12969 ## What Changed - Replace every `describe.sequential` use with a plain `describe` call. - Drop the `{ concurrent: false }` suite option from the two runner test files. - Add a comment to `server/vitest.config.ts` that records why these suites must run one test at a time. - Leave the package manifests and the lockfile unchanged. ## Why the modifier did nothing The Vitest documentation states that `describe.sequential` is useful to run tests in sequence inside a `describe.concurrent` suite, or with the `--sequence.concurrent` option. `sequence.concurrent` defaults to `false`. This repository satisfies neither condition: - No test file uses `describe.concurrent`, `it.concurrent`, or `test.concurrent`. - `server/vitest.config.ts` sets `sequence.concurrent: false`, with `maxWorkers: 1`, `maxConcurrency: 1`, and `isolate: true`. - `packages/paperclip-runner/vitest.config.ts` sets no `sequence` block, so the `false` default applies. `packages/db` and `cli` already run the same embedded-Postgres suites with a plain `describe`, and those jobs are green. The server package was the only outlier. The modifier did carry one real piece of knowledge: these suites need their tests to run one at a time. The new comment in `server/vitest.config.ts` records that reason next to the setting that enforces it. ## Verification - `git grep` for `describe.sequential` returns nothing outside `node_modules`. - The author ran the changed server test files under the installed Vitest 4, and the results match the results without this change. - Two very large embedded-Postgres test files exceeded the author's local memory limit, so the CI test jobs cover those two. - The two changed runner test files have pre-existing local failures caused by a missing Rust toolchain and a missing global `pnpm` binary. The failures are identical with and without this change. - The author type-checked the changed files and found no new error. - CI must pass the typecheck, build, server test, and runner verify jobs. ## Risks - Low risk. Suite execution stays serial, because the Vitest config enforces it. - The change adds no dependency and changes no package manifest or lockfile. - A future change that turns `sequence.concurrent` on would break these suites. The new config comment warns against it. ## Model Used - Claude Sonnet 5 — code edits and local verification. - OpenAI Codex, GPT-5 — the earlier revision of this branch. ## Test plan - [x] Every CI check reaches a terminal green state. A pending or queued check is not a pass. - [x] The `Typecheck + Release Registry` job passes. This change must not introduce a type error. - [x] The `Build` job passes. - [x] The server test jobs and the runner verify jobs pass. - [x] Greptile re-reviews this commit set and posts a passing verdict. The dependabot waiver does not apply to this pull request. - [x] `mergeable` reads `MERGEABLE` as a terminal value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the related public issue with `Refs: #12969` - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Priya Raman <priya.raman@paperclip.ing> --------- Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com> |
||
|
|
90182b4f8b |
Report bounded Cloud portfolio failure diagnostics (#15340)
## Thinking Path > - Paperclip manages work for AI agents and their companies. > - Cloud instances proxy a trusted user's portfolio through the control plane. > - A failed request returns a generic error to the client. > - That replacement error loses the failure phase and network code in Sentry. > - Operators need bounded evidence without upstream messages or credentials. > - This PR adds safe diagnostics while preserving the existing request behavior. ## Linked Issues or Issue Description Refs #10850, which added the portfolio proxy. No open PR for this diagnostic gap was found. **What happened?** A rejected portfolio fetch becomes a generic 502 in Sentry. The event cannot distinguish a connection reset, deadline, HTTP response failure, or body failure. The route's previous warning also included the original error and a stack identifier. **Steps to reproduce** Make the portfolio proxy's fetch reject with a TypeError whose cause has `code: ECONNRESET`. The client correctly receives the generic 502, but the captured replacement error loses that code. **Expected behavior** Keep the existing client response. Attach only bounded server-side diagnostic fields to the failure event. Do not retry the request or expose the original error. ## What Changed - Add a typed portfolio error with a private frozen diagnostic record: phase, upstream HTTP status, elapsed milliseconds, and an allowlisted network code. - Read at most four error/cause objects through own data properties. Unknown codes, messages, getters, and out-of-range values do not enter the record. - Replace the route's raw-error warnings with safe fields. Send a plain error plus event-local context through the existing optional Sentry gate. Keep the route callsite and default fingerprint policy. - Preserve authentication, trusted headers, cookies, exact HTTP error bodies, cache behavior, the ten-second deadline, and one fetch per request. Public responses receive no diagnostic fields. - Document the fields and test HTTP behavior, privacy, and event isolation with the real Sentry SDK. ## Verification - `PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1 pnpm exec vitest run server/src/__tests__/cloud-portfolio-error.test.ts server/src/__tests__/cloud-routes.test.ts server/src/__tests__/sentry.test.ts server/src/__tests__/run-failure-sentry-real-sdk.test.ts`: 65 passed, with the audited optional SDK installed. - Full local `pnpm -r typecheck` and `pnpm build` passed on Node 24.21.0 and pnpm 9.15.4. - Independent review found no blockers and independently passed all 65 tests, including the real SDK checks, on this exact commit. - Full Linux CI passed on this exact commit and provides aggregate suite coverage (54 successful checks, 2 intentional skips). A duplicate full local aggregate was not run. - The first SDK contract job failed before tests when npm could not resolve an OpenTelemetry transitive package. A subsequent empty-cache install first encountered a missing tarball, then succeeded after the registry artifact became available. All 6 real-SDK tests passed against that fresh install; the single unchanged-head CI retry passed. The SDK pin, workflow, and dependency files are unchanged. - The first browser shard 8 run timed out waiting for the inbox retry reply after 45 seconds. The unchanged isolated case passed (1/1), and the test, UI handler, fixture, and recovery files match the base commit. The failed log contains no wakeup POST before the test's immediate navigation; a navigation/request timing race is suspected but unproven without a trace. The single unchanged-head shard retry passed (20 passed, 1 skipped); no timeout or source change was made. - Greptile reviewed this exact commit at 5/5 with no unresolved review threads. - No live portfolio request was replayed. Route tests use controlled local upstream responses. - The added diff passed the secret and PII scan and `git diff --check`. ## Risks This is a diagnostic change. It does not identify or repair the origin of a connection reset. Unknown transport failures remain `unknown`. Elapsed values outside 0–60,000 ms become null. Default Sentry fingerprinting remains enabled; exact historical group membership is not guaranteed. No retry, migration, deployment, or configuration change is included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and independent agent review. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab56046a8e |
refactor(connections): split UI and qualification into follow-up PR
Keep the provider and runtime implementation in this PR. Preserve the completed interface, review stories and live browser qualification on a linked follow-up branch for a bounded review. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
7cbd61349a |
Merge remote-tracking branch 'origin/master' into codex/provider-routing-connections
* origin/master: fix(tasks): surface Codex ChatGPT model rejection (#15299) fix(runner): continue restart-interrupted Codex turns (#15297) fix(agents): grant configuration access by default (#15283) Guard routine UUID lookups without narrowing valid inputs (#15313) fix(ssh): transport project repositories as their own git checkouts (#14782) feat(exe-dev): copy a source VM with exe.dev cp (#14975) Diagnose native model rejections and preserve repairable reviews (#15304) Verify native semantic input against its raw wire digest (#15301) Retry pooled execution projection reads after disconnects (#15300) Fix browser polling for unsaved agent chats (#15298) test(ui): keep initial reasoning with the saved reply (#15107) docs: add a product feature map (#15288) |
||
|
|
3290d97417 |
Scope the release smoke opening question to its card (#15339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The release lane checks onboarding before it promotes a nightly candidate. > - The first task shows an unanswered-question summary and an open question card. > - Both display the same prompt, so the smoke test's page-wide text locator fails strict matching. > - This pull request scopes that assertion to the open interaction card. > - The smoke still requires the actual prompt and verifies that onboarding does not start an agent run. ## Linked Issues or Issue Description **What happened?** [Scheduled release run 37441718130](https://github.com/paperclipai/paperclip/actions/runs/37441718130) failed `smoke_nightly / smoke` on both attempts. The page-wide `getByText("What would you like to do?")` matched two elements: the timeline summary and the open question card. This skipped `publish_nightly`. **Expected behavior** The test confirms that the seeded task's open question card displays the exact prompt. A summary row alone must not satisfy that check. **Steps to reproduce** Run the Docker onboarding smoke against published candidate `2026.1006.0-canary.11`, then run the release-smoke Playwright spec. The old assertion fails after the greeting appears. **Paperclip version or commit** Release workflow commit `f858207161ba29c01c82f4674aef83d91b74480f`; published candidate `2026.1006.0-canary.11`. The assertion is unchanged on the current master base `9b3fe260bac576d622ffdfefb923db73cc5273f7`. **Deployment mode** Docker smoke harness, authenticated/private, with its existing mock provider. Related: #13166 added the first-task chat assertion. I also reviewed open #12316, which fixes a separate bootstrap race and does not change this selector. Searches found no duplicate selector fix or public issue. ## What Changed - Scope the exact opening-prompt text to `task-chat-interaction`. - Explain why the timeline summary cannot satisfy the assertion. - Preserve the rest of the smoke flow, including the 15-second check for no agent runs. ## Verification - Local Docker harness ran the exact failed published candidate, `2026.1006.0-canary.11`, with the existing mock provider. - The old spec reproduced the same two-element strict-mode error in Chrome. - The fixed spec passed the full authenticated onboarding flow and the 15-second no-run check: 1/1, 32.6 seconds. The executed spec copy was byte-identical to the changed repository file; only artifact output paths and the local port were overridden. - Independent Chrome checks: the scoped locator passes with both summary and card present. Summary-only, wrong prompt, hidden prompt, and duplicate active prompts each fail as intended. - `pnpm typecheck` and `pnpm build` passed on Node 24.21.0. Canonical Linux CI passed all 55 checks (53 success, 2 intentional Storybook skips), including the full aggregate tests, build, typecheck, all eight E2E shards, and post-ready security scan. - Greptile scored the exact head `5a30e667396600a052a4041c819652c19b8c68ff` at 5/5 with no actionable issues. Independent review found no issues; there are no review threads. - `git diff --check` passed. The diff and PR text were scanned for secrets and private identifiers. ## Risks Low risk: one assertion in a release smoke test changes. A missing or hidden prompt still fails, and multiple matching prompts in interaction cards still fail strict mode. This does not change the product, release selection, or publishing. The scheduled release lane still needs its next normal successful run to prove nightly recovery. ## Model Used OpenAI GPT-6 through Codex, with reasoning, shell tools, and an independent Codex review. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b1caf043e |
test(connections): add repeatable live provider qualification
Exercise Apps and new-agent setup with real credentials, isolated targets, browser login handoffs, independently verified artifacts and follow-ups. Preserve safe failure diagnostics, private evidence, cleanup, provenance and incomplete coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
01e756bc61 |
fix(connections): preserve provider workspaces and follow-up context
Keep selected model and skill paths current across disposable connection homes. Preserve private Grok history and assigned OpenCode and Hermes workspaces. Fix artifact helper paths and cover the boundaries with regression checks. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9b3fe260ba |
fix(tasks): surface Codex ChatGPT model rejection (#15299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native Runner runs report provider errors and can also save a structured failure result. > - Codex rejects a model that the user's ChatGPT account cannot use. > - Master now diagnoses this rejection, but the task's compact run data omits its message. The thread can still call it a generic run failure. > - Users need the account restriction and a clear step to repair the model selection. > - This pull request shows the existing diagnosis as “Model unavailable” on the task. ## Linked Issues or Issue Description **What happened?** Codex returns HTTP 400 with `invalid_request_error` and the message `The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account.` Master now stores an actionable diagnosis for this rejection. The task thread still labels it “Run failed” and does not receive its error text in compact run data. **Expected behavior** Show the account restriction on the task and in the run error. Tell the user to choose a supported model or clear the task's model override before retrying. **Steps to reproduce** 1. Use a Native Runner Codex agent signed in with a ChatGPT account. 2. Select a model that produces the rejection above and start a task. 3. Let the runner save its generic failed result. Inspect the task's failure marker and recovery notice. **Additional context** Refs: #15304. That merged PR diagnoses the provider failure and preserves worker and review recovery rules. This PR adds its task-facing message and guidance without changing that diagnosis or those rules. Refs: #13134. That PR improves model discovery for ChatGPT accounts. This PR exposes the rejection when a configured model still fails at execution time. ## What Changed - Return a bounded model rejection message in compact issue-run data. - Show “Model unavailable” and model-change guidance in the task thread and recovery notice. - Add database and UI regression tests, including both the original rejection text and master's fixed diagnosis. Document the new failure message. ## Verification - Focused merged-branch validation: 279 tests passed across activity service, native provider failure observation and PostgreSQL integration, TaskChatThread, and ExecutionBlockerNotice. - `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` passed on the merged source tree. - Greptile reviewed conflict-resolution commit `1db39c4e636a92e2e76ba224b71c6f1e2f556c81` at 5/5 with no findings or inline comments. - All 55 check runs completed without failure on that commit. The two optional Storybook jobs were skipped. The legacy Snyk status passed. The branch has no merge conflicts. - UI regression assertion: the task's failure marker says “Model unavailable” and retains the account restriction. It no longer says that this failure happened after a final response. ## Risks - Provider recognition and recovery are owned by the existing master implementation. This PR exposes only bounded error text for failed runs with `native_provider_model_rejected`. - Historical runs with the generic `adapter_failed` code are not reclassified. No stored run is rewritten. - Existing retry and reconciliation gates remain in place. No schema migration or model configuration change is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser inspection. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
63f3aa2dbf |
fix(runner): continue restart-interrupted Codex turns (#15297)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs retain their provider conversation across server restarts. > - A dead runner can restore a Codex conversation after its active turn is lost. > - The old recovery path synthesized a failed task result from that interruption. > - The task then required operator action even though its conversation and workspace were available. > - This pull request preserves the interruption cause and uses the admitted restart attempt for one continuation in the same conversation. > - The agent can reconcile unfinished actions and complete the current request without resending the original task. ## Linked Issues or Issue Description Related: #12845 added native restart recovery. #15042 covers admission during shutdown. #14796 covers legacy shutdown recovery. This change covers a lost native Codex turn after successful conversation restoration. **What happened?** After a server restart killed the local runner, Paperclip restored the saved Codex thread. Runnerd found that the old turn was no longer active. It synthesized a failed terminal and a needs-review result from the last progress message. The transport also discarded the terminal error when it reconstructed thread history. The task failed instead of continuing. **Expected behavior** After proving that the old process stopped and admitting a bounded recovery attempt, resume the current request in the same conversation. Preserve the workspace. Inspect unfinished actions before proceeding. Keep real provider failures, accepted results, intentional stops, unknown unreconciled effects, and exhausted attempts subject to their existing rules. **Steps to reproduce** 1. Start a local native Codex run and leave its turn active. 2. Kill the isolated runner and provider processes, as can happen during a server restart. 3. Restore the same provider thread with no active turn. 4. Observe the synthetic task failure. The new real-process regression reproduces this boundary with a scripted provider. ## What Changed - Record an explicit recoverable process-loss cause without inventing a task result. - Preserve terminal errors and prior turns in reconstructed provider history. Recover the authoritative saved result when adopting an accepted continuation. - Send one continuation in the same conversation for an admitted dead-runner recovery. Require reconciliation of unfinished commands and external actions. - Persist the interrupted terminal before submission and retain the existing recovery marker across another controller loss. - Keep provider attempt limits, terminal failures, and intentional cancellation behavior. - Add red/green regressions, real process-kill coverage, restart checkpoint coverage, and retry-budget coverage. Document the behavior and run-log evidence. ## Verification - Red: the new native runtime regression rejected with `NativeProviderTerminalFailure` on the original code; the Rust restore regression found a missing recovery cause. - Red/green: if restoring the conversation fails and replacement is allowed, the replacement receives the full task and fresh-session handoff. Both prepared and legacy execution inputs are covered. - Red: a second controller crash after the provider accepted the continuation caused an extra `turn/start`. The regression now proves there are exactly two submissions total: the original and its continuation. - Green: focused runtime, backend, driver recovery, and real-process restart suites (207 tests). After the final history/result changes, driver recovery and real-process restart suites passed again (36 tests). - Green: complete Codex transport suite (186 tests), server restart classification/database integration suites (34 tests), and Rust Codex provider suite (92 passed, 2 ignored). - Full `pnpm -r typecheck` and `pnpm build` passed on `60141e649`. The subsequent replacement-prompt guard passed the Runner TypeScript check and the complete runtime plus process-restart suites (148 tests). - Local full-suite attempt: `pnpm test:run` reported two failures in the untouched chat integration suite. Both passed individually, and the complete chat suite passed on rerun (1,063 tests). After all remote test shards passed, the duplicate serial local run was stopped with SIGINT; it is not claimed as a full local-suite pass. - Latest-head CI (`b1297dcd4`): 55 successful checks and 4 intentionally skipped checks, including all test shards, typecheck, build, native Runner verification, end-to-end tests, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37405346303). - Greptile: 5/5 on the latest head, with no unresolved review threads. The PR is mergeable. - The process tests use the real runner binary and a scripted Codex provider. They do not call a live model service. ## Risks - This changes local Codex recovery after process loss. A continuation can execute more work in the retained conversation. Its prompt requires state inspection before repeating an uncertain action; the system does not replay tool calls. - Recovery shares the existing three-attempt budget and one-shot continuation marker. Real failures and older unmarked failed checkpoints are not reopened. - No database migration or API change is required. ## Model Used - OpenAI Codex, based on GPT-6. The exact model ID and context-window size are not exposed in this session. - Capabilities: reasoning, source inspection, tool use, code editing, code execution, and test analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1006.0-canary.13 |
||
|
|
a1ab55a56d |
fix(agents): grant configuration access by default (#15283)
## Thinking Path > - Paperclip manages agent work and permissions for a company. > - Agent setup can require one agent to configure another agent. > - New standard agents could create agents, but they had no direct configuration grant. > - Existing agents must keep their current permissions after an upgrade. > - The requested new-agent defaults include 13 more direct permissions for suggestions, skills, tools, audit, inbox, and task assignment. > - This pull request adds the 14-grant set at creation and approval activation, retains it through invitations, and keeps existing agents unchanged. > - The change keeps low trust and built-in agents on their narrower permissions. ## Linked Issues or Issue Description **What existing behavior does this improve?** The agent creation and approval flows, and the agent configuration authorization path. **Current behavior** A new standard agent can create an agent. It cannot make protected changes to a peer agent unless an operator adds an `agents:configure` grant. **Proposed behavior** Add direct grants for `agents:configure`, `agents:suggest-changes`, `skills:create`, `skills:suggest-changes`, `tools:manage_connections`, `tools:manage_profiles`, `tools:view_audit`, `audit:view_agent_actions`, `tools:use`, `tools:manage_runtime`, `inbox:manage`, `tasks:assign`, `tasks:assign_scope`, and `tasks:manage_active_checkouts` to new standard agents. Scope `tasks:assign_scope` to the agent’s reporting subtree. Keep existing agents and their grants unchanged. **Reason and benefit** New standard agents can complete agent setup and the requested tool, skill, inbox, and task workflows under existing route, scope, and approval checks. **Breaking changes** Existing agents keep their current permissions. There is no permission migration. Related PR: #12212 adds scoped grant routes. This PR changes the default grant. ## What Changed - Add the 14 requested direct grants in agent creation and approval activation. Keep them when invitation approval replaces grants, preserving explicit scopes. Remove them when an agent is deleted. - Apply defaults only to new agents. Remove the existing-agent permission migration, its snapshot, and its journal entry. Preserve existing scoped grants. - Add permission, scope, invitation, and existing-agent regression tests. Document all default and excluded permissions. - Prevent agent keys from creating, changing, or restoring host-executed process or local adapter command settings, and from restoring workspace commands through rollback. ## Verification - Run `node_modules/.bin/vitest run server/src/__tests__/agent-default-configure-grants.test.ts server/src/__tests__/invite-join-grants.test.ts packages/db/src/migration-snapshot-drift.test.ts`. All 15 tests pass. These cover new-agent defaults, unchanged existing grants, pending approval, invitation grants, and migration history. - Run `node_modules/.bin/tsc -p server/tsconfig.json --noEmit`. It passes. - Run `packages/db/node_modules/.bin/tsx packages/db/src/check-migration-numbering.ts` and `packages/db/node_modules/.bin/tsx packages/db/src/check-migration-safety.ts`. Both pass. - Confirm that this PR has no files under `packages/db/src/migrations/` in its final diff. - GitHub CI at `e8173b2c41` passes build, typecheck, server suites, browser shards, and canary verification. One unchanged OpenCode transport test reached its five-second timeout in the first Runner shard run. That test passes locally in 2.55 seconds. The shard passed on one retry. All 54 checks pass, with two expected skips. - Greptile gives this exact head 5/5. Security review passes. The PR has no unresolved review threads or merge conflicts. ## Risks - The 14 default permissions apply only to new standard agents. Existing permissions stay unchanged. Low trust and managed built-in agents are excluded. - New grants include connection, runtime, and active-checkout management. Existing company, responsible-user, scope, and approval checks remain in force. The default inbox grant carries a responsible-user-only scope so it cannot override another user's inbox settings. - Protected changes still need the responsible user's authority when that check applies. Company boundaries and approval gates still apply. - Agent-authenticated requests cannot configure process adapters or host command settings on local adapters. Known provider credential references remain allowed, as do narrow plain authentication overrides on new peers when the caller uses an AI connection pool. Arbitrary environment settings remain blocked. Board operators retain the host configuration paths. > I checked `ROADMAP.md`. This is a narrow fix to existing agent configuration behavior. ## Model Used - OpenAI Codex, GPT-6 series. This runtime did not expose its exact hosted model ID or context window. This revision uses OpenAI GPT-6 through Codex with tool use, code execution, and tests. The runtime does not expose the exact hosted model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1006.0-canary.12 |