mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
codex/plugin-task-execution
21
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
892b0b3606 |
fix: scope quota reports to authorized subscription accounts (#14994)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2ca0d26a99 |
fix(connections): recover missing personal AI credentials in chat (#15376)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e38d6d16b6 |
feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect accounts and choose an agent harness and model. > - The runtime change in #14970 supports custom providers on those connections. > - Normal setup must stay simple while advanced users can choose a compatible gateway. > - Shared connector rows and access controls keep these choices consistent. > - This pull request refines the agent setup UI and adds review stories and repeatable browser qualification. > - The qualification checks real tools and downloaded outputs, not only a successful run status. ## Linked Issues or Issue Description Refs #14970, #37, #13083, #14104, #14565, #12692. The core implementation in #14970 is merged. This branch incorporates its squash commit and targets `master`. Both PRs contain our implementation. #14016 is a reference only and is not a dependency. This PR has 96 changed files. ## What Changed - Complete model-provider connector presentation beside other connectors. Each row uses the existing Connect action and connection list. Tags are stored without category UI. The base PR includes the provider forms and routes. - Show persistent Subscription, API Key, and Advanced choices. Label Advanced as Custom Gateway. Reuse provider logos, connection lists, and permissions controls. Default access to the organization and all agents when permitted; keep narrowing controls under Advanced. - Keep Configure reachable before subscription sign-in, so users can select a supported environment when the default cannot sign in. Testing and saving still require a connection. Show the execution environment in Configure. Preserve the confirmed Connect choice. Editing a method, credential, saved account, or advanced choice requires that current choice to connect before testing or saving. Use matching model and thinking-effort dropdowns and retain connection icons in selected values. - Preserve the new harness model default when switching an existing OpenCode agent to Codex or Claude, and resolve user-selected model names with the effective harness. - Load popular OpenRouter models through the shared connection-model discovery path. Keep explicit model lists and manual model entry available. - Group onboarding, connection setup, agent runtime, management, recovery, and production-component stories under AI Connections / Provider routing. - Add an explicit-only provider-connections browser suite for managed local or existing local/staging targets. Use private browser profiles and credential handoffs. Support human-assisted subscription sign-in without sharing passwords or tokens in reports. - Verify persisted connection identity, runtime probes, tool execution, exact artifact bytes, completion, and context-dependent follow-up. Retain source/model provenance, cost bounds, closed error diagnostics, original failures, and cleanup evidence. - Add Gemini startup-model and skill-root fixes, Grok private-history detection, ACP filesystem regression fixtures, selected-workspace handling for local Hermes, and artifact-helper workspace fallback. - Keep managed Grok runtime homes disposable. Remove host-side transcript retention/restoration because private file modes do not isolate same-user agent processes. Ignore earlier development archives and use a fresh task handoff when history is unavailable. Verify the absence of restored transcripts with a separate same-user process. - Capture stopped-run diagnostics before deleting an attached-company fixture agent. Track creation and owned sign-in receipts; revoke only this attempt's accounts and never adopt a concurrent campaign's newly created account. Preserve failure signals and final status through cleanup. - Require the requested environment in the saved agent and every run, including follow-ups. Reject a forced incompatible target. Keep one cancellation state through startup, every cell, reporting, and teardown for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after interruption. Document qualification limits. ## Verification - Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes master `d9f600043`. The security fix in `a758fde31` passes full workspace typecheck, production build, and 119 connection/Grok regressions. The unchanged UI passes all 126 configuration/model-discovery tests and token gates. The final published-guide correction passes Grok adapter typecheck. Earlier head `eebd8225c` passed the complete deterministic runner suite (1,404 Vitest tests and 128 Node tests) and all CI jobs. Current-head CI run `37520147514` passed all 47 jobs, including the full sharded Vitest and browser matrix, production build, and canary dry run. All 55 checks completed: 53 successes and two expected skips. The current-head security scan passed, Greptile is 5/5, and no review threads remain open. - A separate same-user process reproduced reading a restored Grok transcript before the security fix. The regression now finds no transcript. Existing fresh-session fallback and ordinary session metadata behavior pass. - The final account-choice and cleanup fixes pass 85 setup tests and 26 qualification-harness tests. Regressions verify that editing a connection invalidates confirmation, Configure remains reachable before sign-in, diagnostics are captured before fixture deletion, and concurrent campaigns cannot adopt or revoke each other's accounts. UI and E2E typechecks pass. - The Storybook build and actual Chromium production-component stories passed during this change. Review the neighboring AI Connections / Provider routing stories, regular connector rows, three connection modes, model discovery, and the single execution-environment control in Configure. - Cancellation smoke verified authenticated cleanup before browser close for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during startup and reporting, missing-file ACP resource errors, and preserved permission denials. Both ACP runtime versions and 54 ACPX/Grok regressions passed. The deterministic connection-intent browser suite passed two tests. - Historical local qualification retained 43 passing API/gateway cells out of 46, with downloaded outputs and follow-up receipts. These attempts span earlier builds; they do not qualify this exact commit or staging. Subscription combinations, Gemini overloads, and the unresolved follow-up failure remain recorded rather than counted as passing. - Use `pnpm test:e2e:runner -- --list --suite provider-connections` to inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md` for credentials, target URL, sign-in assistance, budget, evidence, and cleanup. Paid live tests remain opt-in. ## Risks - The core implementation in #14970 is merged. This PR adds no database migration of its own. - Subscription login needs an interactive provider session. Dedicated accounts and staging qualification remain follow-up work; this PR does not certify every login combination for production. - Managed Grok transcript resume is deferred until provider history has an OS isolation or authorized broker solution. Follow-ups start fresh with Paperclip task context; earlier live Grok results do not qualify this behavior. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Live overloads and one unresolved follow-up timeout remain recorded. The stock CLI is unchanged, and those cases are not marked as passing. - Real-provider tests spend credits and use private credential/evidence directories. The launcher requires explicit selection and checks target ownership. It must not attach to a developer's database by accident. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local remain outside custom provider setup. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9f7057e122 |
feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use a harness, a model, and a credential to run tasks. > - Connections already store credentials and control who can use them. > - Custom providers also need an endpoint and a supported API format. > - A per-agent endpoint would duplicate credentials and access rules. > - This pull request stores routing on the connection and projects it into the harness. > - Isolated credentials and protocol checks keep the selected connection authoritative. ## Linked Issues or Issue Description Refs #37, #13083, #14104, #14565, #12692. #14016 is a reference only. This PR has its own schema, vault persistence, routing validation, runtime projection, and tests. None of the five commits in #14016 is an ancestor of this branch. We do not depend on or plan to merge it. #14967 addresses task-pinned account pools. #14422 addresses another provider integration. This is the first of two linked PRs. Merge this connection change before #15341, which refines agent setup and adds the qualification harness. The split keeps each review below 100 changed files. Provider catalog entries have usable setup forms in this PR. Local browser subscription sign-in is included. ## What Changed - Store non-secret routing metadata on AI connections. Vault provider API keys, including Bedrock bearer API keys. Reject general AWS access keys. - Enforce company, owner, human audience, agent access, connection status, and protocol checks before resolving credentials. Keep reconnect destinations immutable and retain connection identity during key rotation. - Project OpenRouter and compatible custom endpoints into Codex, Claude, OpenCode, and local Hermes. Carry these settings through both legacy and native runner transports. Clear conflicting host credentials and redact keys from diagnostics. - Preserve older OpenRouter accounts and native personal defaults. Add Google API-key accounts and migration `0306` for the two provider-default constraints. - Run local Claude and Codex subscription sign-in behind the existing browser sign-in card. Use private attempt homes and owner-bound completion instead of a copied terminal command. - Seed isolated Gemini authentication and preserve OpenCode workspace permissions. Keep the selected connection authoritative. The independent Gemini and Grok workflow fixes are in #15341. - Keep native OpenCode custom gateway keys in a runner-owned selected-model proxy; the harness config contains only a session-scoped capability. Honor runtime outgoing proxy and certificate settings. Preserve streamed responses and revoke the proxy on close or startup failure. - Allow ordinary members to connect native personal accounts before an agent exists. - Repair routed accounts from task cards using the saved provider destination, protocol, model aliases, and connection identity. - Add provider catalog definitions, model discovery, pinned logos, and complete native and routed setup forms. Allow a personal routed connection before a new agent exists. Keep endpoint authentication keys out of Hermes terminal children. - Recover cancelled or restarted browser sign-in with a clear restart action. Support no-auth endpoints without a vault credential. Add isolation and recovery regressions and runtime documentation. ## Verification - Updated with `origin/master` at `22a3ea341`. Migration `0306` follows the new master migration and passes migration and snapshot checks. - The integrated connection regressions passed 152 tests and 50 native OpenCode driver tests, including key-free child-shell configuration reads, authenticated/no-auth forwarding, streaming, model/path restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass. Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34 executable also completed a turn through the proxy against a local synthetic provider; the reusable key was absent from its config. A second real-executable smoke passed with an HTTPS CONNECT proxy and runtime-specific synthetic certificate trust. Certificate-file and certificate-directory regressions pass. - Task-card repair passed 48 tests, including OpenRouter, Bedrock, and custom gateway reconnect cases. UI typecheck and token gates passed. - The prior core regression set passed 133 tests across new-agent setup, provider forms, browser sign-in, routing projection, and connection authorization. Token gates and UI typecheck passed. - Full workspace typecheck and production build passed again after the latest integration and credential-proxy fix. The merged deterministic runner E2E suite passed 1,400 Vitest tests and 128 Node tests. - Full workspace typecheck passed on the prior linked combined implementation. Production build, Storybook build, 1,316 browser-harness Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the latest master integration and regenerated migration. This exact head passed 54 remote checks with four expected skips and Greptile 5/5; no review threads remain open. An unchanged server fixture had a random six-character issue-prefix collision on its first attempt. All 245 tests passed locally and the single CI retry passed. - A provider-free terminal check used the cited supported Hermes source and dummy keys. Gateway and OpenRouter terminal children could not read the selected key. - The broad local Vitest attempt passed 15,442 tests but was not green. It had an embedded-Postgres startup failure, an HTTP logger timeout, an origin socket error, and a browser cancellation wait timeout. The cancellation wait was corrected. The relevant connection tests and the full origin test file passed separately. Latest-head CI must pass before merge. - Prior credential-backed acceptance exercised task creation, tool use, artifact delivery, completion, and context-dependent follow-up. Claude legacy and native runners passed Bedrock with `us-east-1` and `us.anthropic.claude-sonnet-4-6`. - Historical local qualification retained 43 passing API/gateway cells out of 46. Those attempts span earlier builds. They do not qualify this exact commit or staging. All subscription combinations and staging remain unqualified. - Verify native subscription and API-key setup. Connect a regular provider catalog row. Verify an incompatible harness and a changed reconnect URL are rejected. Use #15341 for the complete browser campaign. ## Risks - Migration `0306` changes two check constraints. It preserves rows and is safe to reapply. It takes normal constraint-change locks. - Credential projection touches several harnesses. CLI upgrades can change provider configuration and session behavior. - The native OpenCode proxy adds a loopback hop, pins requests to the selected model, limits request bodies to 16 MiB, rejects redirects, and expires at session close. It prevents reusable keys in the child configuration; it is not an OS isolation boundary against a process debugger running as the same user. - Custom endpoints must be reachable from the agent environment. Saving a connection does not prove connectivity. Bedrock keys require rotation before expiry. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Provider overloads and an unresolved follow-up timeout also affect live Gemini qualification. We have not patched the installed CLI or marked those cases as passing. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS identity, arbitrary auth headers, and custom routing for other harnesses are excluded. - These PRs do not establish production or staging qualification for every provider and login method. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7d59de6113 |
feat(connections): probe provider usage limits on demand (#14936)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections store the AI accounts used by legacy and native runners. > - Subscription accounts can reach session, weekly, model, or paid usage limits. > - Operators need to read these limits for a specific stored account before making a routing decision. > - This pull request adds an on-demand usage probe to the connection service and account detail. > - The result preserves provider limits, reset times, paid usage, and unknown values for later consumers. ## Linked Issues or Issue Description **Subsystem affected** Shared contracts, the connection service and API, and the account detail UI. **Problem or motivation** Managed AI accounts lack a common operation to read their current usage limits. A local harness probe can read a different login from the account selected for an agent. **Proposed solution** Add `aiConnectionService.probeUsage()` and a board-only connection usage endpoint. Probe the selected credential grant on request. Support Codex, Claude, and Grok subscriptions, plus OpenRouter API key limits. **Alternatives considered** Harness-specific automatic polling would couple the read to execution and can read ambient credentials. This change uses the managed connection credential and leaves scheduling and admission decisions to later work. **Roadmap alignment** This extends the existing Personal & Shared AI Accounts capability. It adds no routing or quota enforcement. Related: Refs #14459 for managed OpenAI quota reads; Refs #14781 and Refs #13379 for downstream pacing and budget work. This operation reads one requested account across all three subscription providers. ## What Changed - Add typed usage snapshots and a probe capability flag to managed AI connections. - Normalize Codex, Claude, Grok, and OpenRouter responses. Keep model scopes, provider admission, reset periods, and paid allowances separate. Preserve unknown values. - Enforce company membership, credential audience, grant identity, and connection lifecycle before reading the stored secret. - Add a board-only `GET /api/companies/:companyId/ai-connections/:connectionId/usage` endpoint with `no-store` responses. - Add manual **Check usage** and **Refresh** actions to account details. Show compact usage bars, resets, admission and overage status; remove repeated descriptions and account-default copy. Clear previous results during a new request or error. - Add Storybook previews using the production account components for all four providers, initial checks, loading, and permission errors. - Add provider, authorization, runner selection, API, and UI coverage. Document provider sources and live qualification. ## Verification - Initial provider, authorization, selection, API, and UI validation passed (96 focused tests): `pnpm exec vitest run server/src/services/ai-connection-usage.test.ts server/src/__tests__/ai-connections.test.ts ui/src/components/ai-connections/AiConnectionUsagePanel.test.tsx server/src/__tests__/openapi-routes.test.ts`. - `pnpm -r typecheck` passes for the initial implementation. After simplifying the UI, 9 usage-panel and date-helper tests, UI typecheck, token gates, and Storybook build pass. The initial feature module boundary check also passed. - Real Codex, Claude, and Grok credentials were saved to encrypted disposable connections. The actual usage HTTP route returned 200 with `status: ok`. Legacy and native runner selection checks passed. The tests started no model turn and exchanged no refresh token. The disposable databases and vaults were removed. - Live Claude responses added structured scoped limits. Live Grok responses omitted included-plan usage. Tests now cover both shapes and preserve the Grok omission as unknown. - The full workspace build passes. A full local test run hit a heartbeat feedback timeout. That case passes in isolation. The duplicate local run was stopped after all remote checks passed. The Slack ordering and OpenCode transport CI flakes also pass in isolation and on the CI rerun. - Current head: `ff3d479029a1c4248190323e221b2803cfb0d79d`. All 54 active checks pass. Two Storybook checks are intentionally skipped by the workflow. Greptile is 5/5 with no unresolved review findings; the branch is mergeable. ## Risks - Subscription usage endpoints can change. Credentials can lack usage-read permission. The probe returns explicit errors without fresh limits in these cases. - A successful probe can contain partial data. Missing utilization or admission remains unknown. An enabled paid-usage switch does not prove a funded balance. - This change adds no migration. It does not change runner admission or automatic provider selection. Provider requests use fixed endpoints, disabled redirects, bounded response sizes, and a 15-second deadline. ## Model Used OpenAI Codex, GPT-6, with reasoning, file editing, shell execution, and HTTP tools. The session does not expose the exact runtime model variant or context window size. Real provider credentials were used only for the authorized live checks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e00d10d5d5 |
fix(connections): repair stale AI defaults from agent settings (#14916)
## Thinking Path > - Paperclip manages AI agents and controls the credentials used for their work. > - Managed AI connections resolve each responsible user's provider default. > - Agent settings created another account but kept the old default selected. > - A rejected provider test left the old account marked as connected. > - Claude ACP reported a typed login failure as a generic terminal-access error. > - This pull request repairs the selected account or selects the new login explicitly. > - Agents can save and run with the repaired credential, and failed logins request sign-in. ## Linked Issues or Issue Description - Fixes #14831. - Refs #13867. Environment failures remain separate from credential-health failures. ## What Changed - Add an agent-settings action to reconnect an unavailable personal default in place. Keep its connection, grant, default, and agent access. - State that a new account becomes the user's provider default. Select its returned grant before changing the agent binding. Keep the actual sign-in method. - Show default-update errors and allow retry without another provider login. - Show the agent-access choice. Connection managers start with company-wide access for their own tasks. Other members start with access for the current agent. - Use the server's connection-manager permission in the shared list response. This includes members with a custom management grant. - Mark credentials as needing attention after an explicit login rejection in Test or Save. This includes API-key 401 and 403 responses. Network, quota, and server failures keep the credential health unchanged. - Reuse the credential-generation check so an old failure cannot invalidate a newer reconnect. - Route Claude's typed provider `access` failure to the existing login-recovery flow. Replace its generic terminal-access fallback with a sign-in message. - Add regression tests and update the AI Connections documentation. ## Verification - Red: the UI tests failed on the missing reconnect action, unused returned grant, missing access choice, and lost default-update error. The server tests failed because rejected credentials stayed connected. The real ACP fixture returned `acpx_turn_failed` for typed login failures. - Green: 156 tests passed across the AI connection, hiring, agent field, and New Agent suites. All 37 environment-route tests passed. The Claude ACP authentication fixtures also passed. - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - The full local `pnpm test:run` passed 707 files and 14,503 tests, then exited with an agent-conversation timeout and embedded PostgreSQL startup failures in unchanged suites. The isolated conversation and migration tests passed on rerun. Later local test groups did not run after this failure. - [All CI gates passed](https://github.com/paperclipai/paperclip/actions/runs/37012669356) on commit `38513dfe2`. This includes the full test matrix, browser tests, typecheck, build, Runner checks, and canary dry run. - Greptile reviewed commit `38513dfe2` and returned 5/5 with no open findings. - The regression tests use a real embedded database and a real ACP fixture process. Live provider sign-in requires a valid account and was not run. ## Risks - Connecting a new account from agent settings changes the user's provider default. The dialog states this before sign-in. - The displayed access choice can allow all company agents to use the account for its owner's tasks. Reconnect keeps the existing access. Server permissions still control installs. - Claude's typed `access` category maps to the provider's `auth_required` signal. Tool and workspace request failures retain their existing classification. - No database migration or provider credential format changes are required. ## Model Used - OpenAI GPT-6 through Codex. The exact served model identifier and context window are not exposed in this session. Capabilities used: reasoning, repository tools, code editing, and command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2f6fa3b6dc |
fix: recover provider authentication inside tasks (#14629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a working model provider connection to run a task. > - A provider can reject a stored credential after the task starts. > - The failed run must ask the responsible user to repair that connection. > - This pull request adds that request directly to the task and reuses Connections sign-in. > - The user can choose an API key or subscription, then continue the task with a fresh session. ## Linked Issues or Issue Description **What happened?** A run that ended with `acpx_auth_required` or another known provider authentication error did not immediately offer an inline way to connect the provider. A repair form could also lock the user to the failed account's sign-in method. **Expected behavior** Show a provider connection card in the task as soon as the authentication failure is saved. Allow the responsible user to connect or repair the provider with any supported sign-in method. Keep the connection name automatic and resume the task after successful setup. **Steps to reproduce** 1. Run a task with a supported provider and an expired or invalid credential. 2. Let the run fail with a provider authentication error. 3. Open the task and attempt to repair the connection. Related work: Refs #13724 and #13726. This change adds the inline task repair flow and method choice. ## What Changed - Classify provider authentication failures and create one connection request for the current task. A persisted blocked classification suppresses automatic retries only after the repair card is created; unsupported providers retain their existing recovery path. - Mark only the attributed, unchanged credential as needing sign-in. Preserve credentials that were refreshed after the failed run started. - Reuse the provider sign-in controls inside the task. Allow API key and subscription choices for Claude, Codex, and Grok. Keep names hidden and generate a default from the user, provider, and method. - Keep the existing account when reconnecting with the same method. Create and select another account when the method changes. Validate updates to explicit agent bindings through the normal agent save path. - Require explicit adoption for legacy agent authentication. Validate in the agent environment, then commit the binding, connection install, audit, and card completion in one transaction. Keep failed setup and account selection visible and retryable. - Add regression tests and update the specification and Connections documentation. ## Verification - Fresh local verification: 199 tests passed across the inline form, provider method selector, default naming, authentication and recovery classifiers, run liveness, OpenAPI routes, database adoption/rollback, and Cursor execution suites. The adoption database suite also passed against disposable Docker PostgreSQL. - Full repository `pnpm build` and `pnpm -r typecheck` passed on the latest commit. Token gates are clean. - Embedded browser: opened real task cards from seeded authentication failures; switched Claude from API key to subscription and back; switched Codex from subscription to API key; confirmed the name field stays hidden. Provider sign-in was not completed with real credentials. - The broad local `pnpm test:run` started before review fixes and was interrupted after the working tree changed; it is not counted as a passing full run. Fresh focused tests passed. CI supplies the full test and browser suite results for the current commit. - CI is green on commit `4b97a4e447045ff3d7516525a187a5d1d21e0d4c`: 54 checks passed and two Storybook checks were skipped by their path rules. The workspace preview job passed on one rerun after a local-server startup timeout; its rerun passed 835 tests. - Greptile is 5/5 on the same commit with no actionable findings and no unresolved review threads. ## Risks - Incorrect authentication classification could prompt for a connection unnecessarily. Tests exclude tool authorization, quota, and unrelated runtime failures. - A method change selects the new personal provider default, which also applies to other agents that use that user's default. Explicit account bindings use the existing permission and runtime validation path. - Credential invalidation must not race with refresh or reconnect. The code compares the saved credential generation and grant update time under locks. - No database migration or new credential storage format is required. ## Model Used OpenAI GPT-6 through Codex. The exact model ID and context window size were not exposed in this session. Capabilities used: reasoning, repository editing, shell commands, database tests, and embedded-browser interaction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused tests listed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f8750627f |
fix(codex): preserve typed ACP quota classification and reset time (#13945)
## Thinking Path > - Paperclip runs agents and recovers failed tasks. > - Codex ACP can report usage exhaustion as a typed terminal failure. > - The shared engine removes provider text before it stores the result. > - Codex did not use the existing terminal classifier hook, so a quota failure became a generic turn failure. > - This change classifies explicit usage exhaustion before that text is removed. > - Recovery can then use quota backoff and the supported reset clock. ## Linked Issues or Issue Description **What happened?** A Codex ACP `limit` failure with explicit usage-exhaustion text produced `acpx_turn_failed`. Recovery could schedule the ordinary short retries because the quota classification and reset time were lost. **What did you expect to happen?** Keep the failure visible and use the existing provider-quota wait. Preserve a supported reset timestamp without storing provider text. **Steps to reproduce** Return a terminal ACP session failure with category `limit` and title `You've hit your usage limit for GPT-5. Switch to another model now, or try again at 4:30 PM (America/Chicago).` The real child-process regression tests exercise both pinned ACPX versions in persistent and oneshot modes. **Paperclip version** Base commit: `32573876d4`. Related work: #13651 and #13831 provide the shared hook and Claude classification. #11854 includes quota handling as part of optional credential rotation, but reads the already-sanitized result; this patch handles the typed terminal boundary without adding rotation. #13549 reads recovery text after this boundary and cannot recover discarded provider text. #9011 concerns the Codex CLI backoff. This change leaves those other mechanisms in place. ## What Changed - Register a Codex terminal-failure classifier with the existing ACP engine hook. - Recognize explicit usage exhaustion only in a typed `limit` failure. - Reuse the Codex reset-time parser and existing provider-quota recovery fields. - Leave context, turn, rate, budget, storage-capacity, and unknown failures on their existing paths. - Test real ACP children, both dependency patches, privacy, and recovery classification. Document the boundary. ## Verification - 88 focused tests passed across Codex ACP, parsing, and server recovery classification. - An initial cross-adapter run passed 53 tests, including the existing Claude quota suite. - Removing only the classifier registration makes five new integration tests fail; restoring it passes all 19 new tests. - `pnpm -r typecheck` and `pnpm build` passed. - Full Linux CI passed on `b10d60e002`, including all unit/integration shards, browser shards, typecheck/build, runner checks, and release/package gates. - The duplicate local `pnpm test:run` reported three skill-cache failures and two runner-suite failures. It is not claimed as a full local pass. - The three skill-cache failures reproduce in a clean worktree at the unchanged base commit (3 failed, 107 passed, 3 skipped across the skills and runner files). The cache rename reports `EACCES` on macOS. - An isolated runner-suite run reports two embedded PostgreSQL startup failures before its assertions (37 tests pass). No runner source is changed; the Linux CI runner and server suites passed. - Greptile reviewed the current head at 5/5 with no actionable findings or unresolved threads. ## Risks - Explicit usage-exhaustion failures now wait for quota recovery instead of short generic retries. - Unknown wording retains the existing behavior. The generic historical terminal-limit message alone cannot establish quota exhaustion. - Reset parsing keeps the existing Codex clock formats. A missing or unsupported reset uses the existing quota backoff. - No schema, credential, UI, dependency, or deployment changes. Provider text stays in memory and is absent from results and logs. ## Model Used OpenAI GPT-6 via Codex. Exact runtime model variant and context-window size were not exposed. Used reasoning, repository inspection, editing, and local tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9f30eb10dd |
fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path > - Paperclip manages AI agents and keeps their work attached to tasks. > - Chat connectors carry user messages and agent replies between a provider and those tasks. > - Each extra startup and context reset delays a reply. > - Managed account metadata was lost during adapter decoding, so compatible follow-ups started fresh. > - This branch fixes the reset and measures the remaining preparation, execution, and delivery costs. > - The changes must preserve account isolation, authorization, durable output, and recovery ownership. ## Linked Issues or Issue Description Refs #13699. The related service lifecycle work in #13410 and #13408 is separate; this branch focuses on task-bound chat response latency. **What happened?** Managed AI follow-ups started new provider sessions even after their configuration fingerprint stayed stable. The Codex codec removes unknown fields. The resume check then read the removed credential identity and treated it as a credential change. **Expected behavior** Compatible follow-ups resume the correct provider session. Changes to credentials, responsible users, permissions, or task configuration retain their reset behavior. **Steps to reproduce** 1. Use a Slack connector with a managed AI connection. 2. Send a message, then send a same-thread follow-up. 3. Inspect the configuration reset reason and the provider session identity. **Paperclip version or commit** Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`. **Deployment mode** Cloud staging with a native Codex runner. ## What Changed - Read saved credential identity before adapter decoding discards it. - Remove the internal credential identity from adapter-facing session params. - Test the real Codex codec and missing, changed, or unmanaged identity cases. - Preserve configured warm Codex runners and flush refreshed credentials after every turn. - Fence detached or closing session handles from successor credential ownership. - Stage current Codex launch credentials after restoring durable session history, uploading launch assets only once. - Reuse a runner binary already in the retained sandbox only when its SHA-256 matches the controller-owned artifact; still verify required capabilities before launch. - Lock the task before the run when saving results, preventing deadlocks with task updates. - Scope reusable projectless sandboxes to the company, environment, task, agent, and runtime configuration; verify Daytona sentinels for that scope. - Admit a new authorized chat message after a fully committed failed run and verified process cleanup. - Send compact deltas for verified plain-text Slack continuations. Match the actual prior run and current comment identity/body; exclude edited historical comments and prior agent output, preserve genuine brief edits and the full bootstrap fallback. - Keep attachments, omitted input, questions, approvals, recovery, and other providers on their existing framing. - Document managed session compatibility, credential lifecycle, and compact continuation boundaries. ## Verification - Workspace/session coverage: 156 tests passed. - Native session and credential ownership coverage: 390 tests passed, including exact artifact reuse, mismatches, failed probes, timeouts, and explicit artifact overrides. - Explicit continuation and durable chat authorization coverage: 172 tests passed. - Session resume and launch preparation coverage: 416 tests passed. - Result persistence coverage: 15 tests passed. The new concurrency test reproduced a PostgreSQL deadlock before the lock-order fix. - Environment lifecycle coverage: 92 tests passed, including projectless reuse and task/agent isolation at both selection and atomic handoff. - Daytona plugin coverage: 237 tests passed; 6 gated tests skipped. Standalone plugin build passed. - Compact Slack continuation and native resume coverage: 69 tests passed, including full-bootstrap retention, matching message authors/bodies, current-delivery selection, rejection of duplicate identities and historical comments, brief edits, and attachment/recovery fallbacks. - Final frozen-head `pnpm test:run` on repository-supported Node 26: 668 suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped. The sole failure was a local `socket hang up` in `issue-recovery-actions.test.ts`, not an authorization assertion mismatch. All 57 tests in that suite passed three fresh reruns, and the suite passed latest-head CI. The full local invocation is therefore not claimed green. - An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup failure; that 11-test catalog suite passes on Node 26 and in CI. No test behavior or timeout was relaxed. - Full local typecheck and build passed. Latest-head CI is green; Greptile is 5/5 with no unresolved review threads. - Two real Slack baseline replies took 25.1 and 24.6 seconds (24.9-second mean). Three same-thread signed probes on this head took 23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample and a modest wall-clock improvement, not a large or statistically established speedup. - In that same thread, uncached provider input fell from 8,514 tokens before compact input to 694–765 tokens afterward. The current delivery uses a 362-character delta; the full 19–21k-character bootstrap remains available for failed resume. Verified runner artifact preparation fell from about 1.2 seconds to 0.6 seconds. - A fresh thread created a separate task, sandbox, and provider session with full bootstrap (24.4 seconds). Its follow-up reused its own sandbox/session and compact input (28.3 seconds, including 16 seconds of model execution). Model variability and process startup remain substantial. - A signed duplicate webhook produced exactly one user comment, one successful run, and one final Slack reply. Slack's API independently confirmed the actual replies and a public task URL without an internal or pool hostname. - Earlier signed probes verified recovery after a failed run and reuse across a server deployment. The final idle test observed Daytona report the sandbox as stopped, then delivered a new reply in 19.9 seconds using the same sandbox/provider-session identity and compact input. Slack’s API confirmed that reply. - Live probes use signed synthetic inbound webhooks and real outbound Slack delivery, read back through Slack’s API. The final browser recheck found the Mac locked and the Slack tab blocked by another extension, so this is not claimed as full UI E2E proof. - This is a review branch. Do not merge until the maintainer reviews it. ## Risks - Incorrect session reuse could mix account or task context. Missing or changed identities continue to reset, and existing authorization checks remain in place. - Warm mode remains opt-in. Remote warm mode requires a reusable sandbox lease. Retained processes keep credentials until they close, so idle expiry and ownership fences are required. - A fresh user message may continue after a committed provider failure. Approval, current authorization, process termination, and prior-result checks remain required. - Projectless sandbox reuse is task- and agent-scoped. Missing or mismatched ownership cannot replace an existing lease; existing workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox per task/agent, so provider auto-stop and deletion policies still determine idle compute and storage costs. Fleet defaults are unchanged. - Compact prompts apply only after proven resume and a matching prior-run delta. Missing or specialized context falls back to full input; fresh sessions always receive the full bootstrap. - No schema or migration changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, tool use, and test execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9335b7db10 |
fix(runner): validate inherited environments and replace stale sandbox binaries (#13677)
Resolve the effective environment for account adoption and adapter tests. Preserve saved-agent overrides when the request omits environmentId, and treat explicit null as inheritance from the instance. Reject sandbox runners that lack unlimited-runtime and connection-lease-renewal capabilities. Stage the bundled runner before launch when the image binary is stale. Add regression coverage for environment precedence, fail-closed validation, adapter switches, API-key reverification, and runner artifact fallback. Document the operational workaround for older controllers. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9ed55f6931 |
fix: allow concurrent agent runs on one OpenAI or xAI subscription connection (#13452)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts local command-line sessions and stores provider credentials through managed connections > - One OpenAI Codex or xAI Grok subscription connection held a credential lease for the full agent run > - A second run then waited for the first run, and concurrent runs could overwrite a newer credential > - The write-back must compare fresh credentials while it holds the row lock > - This pull request removes the run lease, keeps the revocation guard, and bounds Codex timestamps against the host clock > - The benefit is safe concurrent use of one subscription connection with newest-credential selection ## Linked Issues or Issue Description **What happened?** A managed OpenAI Codex or xAI Grok subscription connection held a credential lease for the full agent run. A second run waited for the first run to finish. The write-back gate also rejected any row change before it compared credential freshness. **Expected behavior** Concurrent runs should start on one subscription connection. The server should keep the newest valid credential and reject a credential write after a person revokes the connection. **Steps to reproduce** 1. Start two runs that use one OpenAI or xAI subscription connection. 2. Let both provider tools refresh the credential. 3. Finish the runs in either order. 4. Confirm that the newest valid credential remains in the connection. **Paperclip version or commit** `ec25bf1e4a81d1729a6d7276e6a486587cff4a0b` **Deployment mode** Local dev (`pnpm dev`), built from source. **Agent adapter(s) involved** Codex. The server path also covers xAI Grok subscription connections. **Database mode** Embedded PGlite for local development, and external Postgres for deployments. **Access context** Both board and agent runs can use managed connections. **Additional context** The provider command-line tool refreshes credentials inside the sandbox. The server copies the result back after the run. Two long runs can still refresh one token hours apart, so the provider can reject the second refresh. The server cannot observe that provider call. ## What Changed - Remove the full-run credential lease for OpenAI Codex and xAI Grok subscription connections. - Lock and re-read the connection row before credential write-back. - Accept only a strictly newer credential, while keeping the connection revocation guard. - Reject Codex freshness timestamps more than five minutes ahead of the host clock. - Add tests for both completion orders, xAI cleanup, revocation, the Codex time bound, and agent hiring. - Update the connection and run-log documentation. ## Verification - `server/src/__tests__/ai-connections.test.ts` passes with 42 tests. - `packages/adapters/codex-local/src/server/codex-auth-merge-decision.test.ts` covers the five-minute boundary and the one-millisecond overflow. - `packages/adapters/codex-local/src/server/codex-auth-merge.test.ts` passes. - `server/src/__tests__/agent-hire-ai-connections.test.ts` covers OpenAI and Anthropic. - The project type check reports no new error in changed files. - GitHub Actions must pass on this pull request. ## Risks The write-back now permits concurrent runs, so the provider may reject a later refresh when both long runs use one token. The server keeps the revocation guard and rejects future-dated Codex timestamps. No schema change occurs. ## Model Used OpenAI Codex, GPT-5. Context window and exact deployment build are not exposed in this run. The model used tool calls, code inspection, and test verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ee81cee76d |
fix: allow concurrent agent runs on one Anthropic subscription connection (#13445)
## Thinking Path > - Paperclip manages AI agents and their work. > - Agent runs use provider connections and subscription credentials. > - File-backed subscriptions need a lease while a run writes new credentials. > - Anthropic subscriptions pass a token and do not write a credential file. > - The current lease blocks concurrent Anthropic runs without protecting shared state. > - This pull request limits the lease to file-backed subscriptions and tests both paths. > - The result lets Anthropic runs share one connection safely while file-backed credentials remain serialized. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change improves concurrent use of one subscription connection. Anthropic runs no longer wait for a lease that protects no shared file. **Subsystem affected** server/ — REST API and orchestration services, with related server tests and connection documentation. **Current behavior** The server takes a connection lease for every subscription invocation. The lease blocks a second Anthropic run while the first runtime remains open. **Proposed behavior** The server takes the lease only when the subscription writes a provider authentication file back to the grant. Anthropic runs can proceed at the same time. **Reason and benefit** Anthropic passes its token through an environment variable and writes no file. Removing this unnecessary wait improves concurrency and preserves credential safety for OpenAI and xAI. **Breaking changes** None. OpenAI and xAI keep the existing lease and retry behavior. ## What Changed - Limit the AI credential lease to subscriptions that write a provider authentication file. - Add coverage for concurrent Anthropic runs and hired-agent runs. - Keep the existing OpenAI contention assertions as a sensitivity control. - Update the AI connection documentation to describe the lease boundary. ## Verification - Run `server/src/__tests__/ai-connections.test.ts`. - Run `server/src/__tests__/agent-hire-ai-connections.test.ts`. - Run `server/src/__tests__/heartbeat-ai-subscription-contention.test.ts`. - Run the server type check. - Review the pull request checks after GitHub starts continuous integration. ## Risks The main risk is an incorrect provider classification. The write-back condition remains unchanged, and OpenAI contention tests keep the file-backed lease behavior under test. No schema or API contract changes occur. ## Model Used OpenAI GPT-5. Exact runtime model ID and context window are not exposed in this execution. The model used tool calls, shell commands, and GitHub operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a2e7ffdc34 |
fix(runtime): validate sandbox paths and preserve live controller leases (#13432)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can run in remote sandboxes. > - Connection checks must use the selected execution target. > - Recovery must respect the controller that owns an active run. > - A host path or PID does not describe a remote sandbox. > - This pull request checks sandbox paths on the target and preserves live controller leases. ## Linked Issues or Issue Description **What happened?** Selecting an AI account for a sandbox agent could fail because the Claude ACP environment check tried to create the sandbox directory on the Paperclip host. The recovery sweep could also interrupt a sandbox run while its controller lease was still valid. It treated a PID absent from the local host as proof that the run had stopped. **Expected behavior** ACP checks directories on the selected execution target. Recovery leaves a run with a live controller lease alone. Its final database write rejects a stale snapshot after renewal, a claim, a controller change, or a runtime change. **Steps to reproduce** 1. Test a Claude ACP sandbox agent with a directory that cannot be created on the host. The check fails before this fix. 2. Give a running sandbox task a valid controller lease and a PID absent from the host. Run the stale-lock sweep without an in-memory handle. The sweep interrupts the run before this fix. 3. Renew or replace the controller between the sweep's read and write. The old snapshot must not end that controller's run. **Paperclip version or commit** Rebased onto `origin/master` at `0e9b24c8216171c26c8358ba387d77858e02c7a9`. All seven regression cases still fail against this base. Refs #13438, which supplies the managed hiring and task-connection behavior, and #13433, which preserves non-assignee subscription comment wakes. This PR preserves both upstream changes and addresses the two remaining sandbox failures. ## What Changed - Resolve and create Claude ACP test directories through the execution-target helpers. - Preserve active legacy controller leases during stale-lock recovery, including finalization after a task becomes terminal. - Recheck the controller, lease, runtime mode, and native ownership in the terminal database write. - Add two sandbox-directory cases and five database-backed controller-lease cases. - Document the target used for ACP directory checks. ## Verification - Red: all seven new cases fail against `0e9b24c82` without these two implementation changes. - Before the final upstream sync, 192 focused tests passed. Full `pnpm test:run` coverage completed using the repository's group/shard runner: all general server and workspace groups passed, and all 147 serialized server suites passed across the initial run and isolated continuations. Five cold-import timeout suites passed with `--experimental.fsModuleCache`; their assertions and deadlines were unchanged. - The hiring routes, default-selection service, and upstream hiring tests match `origin/master` exactly. The two remaining fixes are unchanged by the final rebase. - On final head `bd28d5cefbdf7084acc3759199cd9661907e7a26`, all 372 focused tests pass across 16 suites covering both upstream changes and these fixes. Two timeouts in the combined run (database setup and an existing ACP case) pass in isolated reruns with fresh test homes and temporary directories. `pnpm -r typecheck` and `pnpm build` also pass on this head. - All 32 active checks pass on final head `bd28d5cef`, including the full test matrix and browser shards, in [CI run 34904204849](https://github.com/paperclipai/paperclip/actions/runs/34904204849). Two Storybook checks are skipped by path filters. - Greptile reviewed final head `bd28d5cef` at 5/5 with no findings or unresolved review threads. ## Risks - A failed remote directory check still blocks connection adoption. - A live controller retains finalization authority after its task becomes terminal. Cleanup waits for ownership to expire and must pass the final ownership check. - No schema or credential-storage changes. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository inspection, code execution, and API tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e9b24c821 |
fix: resume subscription comment wakes and preserve retry status (#13433)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Concurrent task runs can share a managed AI subscription with one credential lease. > - #13438 added durable retries for busy subscriptions and task-lock checks. > - A comment wake can run for an agent who is not the task assignee. Such a run never owns the task lock, so the new checks suppress its retry. > - This pull request preserves those comment wakes while keeping the lock checks for assignee runs. > - It also shows the subscription wait in task status and preserves the failure count through the database projection. ## Linked Issues or Issue Description Refs #13438. **What happened?** When a subscription is busy, a non-assignee comment wake is cancelled without a successor. Repeated subscription waits also display an inflated attempt count because the execution query omits the preserved failure count. **Expected behavior** An eligible comment wake waits and resumes without claiming the assignee's task lock. The task shows “Waiting for AI subscription”. Waiting does not consume provider-failure retries. Assignee retries still stop when the task lock is cleared or transferred. **Steps to reproduce** 1. Configure an agent with a managed subscription and concurrent runs. 2. Hold its credential lease in one run. 3. Mention the agent on a task assigned to another actor. 4. Release the lease and inspect whether the comment wake has a scheduled retry. The test uses a real embedded Postgres database and a real credential lease. Provider execution uses a fixture. The reassignment test applies a database mutation after the real checkout. No live provider account is needed. **Paperclip version or commit** Based on `f912ecaac` from #13438. Before the reconciliation fix, the rebased regression at `2063cfd13` failed because the comment wake had no scheduled retry. The original configuration failure was reproduced on `5282cabde` before #13438 merged. **Deployment mode** Built from source with embedded Postgres for local verification. ## What Changed - Record non-assignee comment-wake authority at admission while holding the task and run locks. A later reassignment cannot grant this exception. Preserve it through repeated subscription waits. - Keep master’s lock checks for assignee retries, including the check inside the scheduling transaction. - Show the subscription wait and preserve its failure count in the execution projection query. - Limit pre-provider wait receipts to fresh executions. A persisted native execution input must retain its recovery path. - Add lease, retry, projection, cancellation, reassignment, pause, revocation, service-recreation, and ownership regression coverage. - Use master’s 60–120 second retry interval and document the resulting behavior. ## Verification - Red: after rebase, the comment-wake test failed with no scheduled retry; the other 50 contention and retry-scheduling tests passed. - Red/green: reassigning the task immediately after real checkout reproduced an unwanted successor for both assignment and comment wakes. Both tests pass after recording authority at admission. - Green: all 285 targeted tests pass across 12 suites, including #13438’s four cancellation-race phases, its hiring/connection suite, run-dispatch integration tests, and both reassignment regressions. - `pnpm -r typecheck` passes after rebase. - `pnpm build` passes on the final revision. - The full general and serialized test suites pass in [CI run 34899514681](https://github.com/paperclipai/paperclip/actions/runs/34899514681) on `99f1e4c1d`. Full-suite verification ran in CI; local verification used the 285 targeted tests. - All 32 active PR checks pass, including build, typecheck, native runner verification, canary dry run, and all browser shards. Two Storybook checks are skipped by path filters. - Greptile reviewed `99f1e4c1d` at 5/5 with no outstanding findings. All review threads are resolved. - `git diff --check` and `node scripts/check-module-boundaries.mjs` pass. ## Risks - The non-assignee exception requires authority recorded under admission locks or its server-created subscription retry. Tests cover caller-supplied flags, assignment and comment reassignment races, repeated comment waits, and assignee lock protections. - A wait uses master’s 60–120 second interval. A long-held lease can produce multiple wait records. - Existing native sessions can have provider effects. They must not receive a fresh-execution receipt that permits replay. - No schema migration or credential permission change is required. ## Model Used OpenAI Codex, model `gpt-6-astra`, with repository search, code editing, shell execution, and test tools. The session does not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f912ecaacf |
fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents hire other agents and assign tasks to them. > - Managed AI connections must follow those hires across legacy and native runners. > - Missing accounts should pause task execution and let the user connect from the task. > - Subscription contention must wait without asking for new credentials. > - This pull request fixes these paths and the native tool and Daytona staging failures found during live tests. > - The result is a working hire, subtask, and connection setup flow on local and remote runners. ## Linked Issues or Issue Description **What happened?** A managed Claude or Codex agent could hire a teammate without a usable AI binding. Cross-provider hiring could fail before the user had a chance to connect the new provider. First-time task setup did not show the existing AI credential form inline. A busy subscription could request a new connection. Native API replies could stop the parent after a hire had already committed. Fresh Daytona sandboxes could fail to extract read-only skill directories created on macOS. **Expected behavior** Compatible hires inherit the managed connection choice. A hire for another provider uses the responsible user's default. If that account is missing, the hire succeeds and the task asks for a connection. Completing setup in the task resumes work automatically. Explicit child auth settings and existing unmanaged login paths keep precedence. Shared-account access checks remain in force. **Steps to reproduce** 1. Connect a Claude or Codex parent with a managed AI account. 2. Ask it to hire one agent of each provider and create a self-assigned subtask. 3. Assign work to both hires without connecting the second provider first. 4. Connect the missing provider from its task card. 5. Check that all tasks finish and same-provider work uses the original account. 6. Repeat with native runners and fresh Daytona sandboxes. The opt-in browser suite in `tests/hiring-ai-connections/README.md` performs these steps. **Paperclip version or commit** The live failures were reproduced from `f2c5e54dc`. The branch is rebased onto `5282cabde`. **Deployment mode** Isolated local development instance. Legacy CLI and native runners. Local execution and ephemeral Daytona sandboxes. Related work: Refs #13247 for managed AI connections. Refs #13268 for legacy credential-reference inheritance, which this branch preserves. Refs #13432 for a concurrent managed-inheritance fix. This PR also covers cross-provider task setup, subscription waits, native API replies, and Daytona extraction. It permits missing responsible-user defaults at hire time; restricted shared selections still fail. ## What Changed - Apply managed connection defaults to both agent creation routes. Preserve explicit auth choices and legacy credential-reference inheritance. - Allow hires before their responsible user connects the provider. Keep approval gates, company boundaries, and shared-account access checks. - Reuse the production AI credential form inside the pending task card. Resume the task after setup. - Retry subscription lease contention without consuming the provider-failure allowance or creating a connection request. - Require task execution-lock ownership when scheduling, promoting, and dispatching subscription retries. Recheck ownership under the issue row lock. - Rename the HTTP operation identity at the native tool boundary so it cannot override the runner's operation identity. - Delay directory permission restoration during Daytona extraction. Preserve the final read-only modes. - Add database-backed regressions, real browser acceptance tests, and Storybook states. Document setup and run-log behavior. ## Verification - Six real browser scenarios passed: both parent providers on legacy local and legacy Daytona; native Codex locally; native Claude on Daytona. Each scenario hires both providers, completes a self-subtask and assigned work, and connects the missing provider inline with automatic continuation. - Successful runs verify the account, responsible user, runner mode, and Daytona lease. All 18 test sandboxes were deleted. - Live authentication used API keys. Subscription inheritance, lease contention, and retry have integration coverage. Fresh subscription OAuth sign-in was not automated. - Red/green tests reproduced missing bindings, missing inline forms, subscription contention, native API reply failure, and GNU tar permission failure. - Seven Storybook browser checks passed. They cover both providers, method selection, narrow layout, completion, cancellation, and invalid credentials. - Full local suite coverage completed before rebase. Initial timing and fixture startup failures passed unchanged on isolated reruns. The first full command did not exit cleanly; the remaining workspace and serialized groups were completed separately. - After rebase, 107 hiring/auth/retry tests and 59 native API, task-card, and Daytona tests passed. The full workspace typecheck, production build, and token gates passed again. Storybook build passed before rebase. - Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and four explicit cancellation-race cases passed. Eight cross-provider cases cover stale auth keys on both creation routes and both runner types. Server typecheck and build passed. - Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks passed. Two optional Storybook jobs were skipped. The full server, workspace, browser, native runner, build, typecheck, and release checks passed. - Three unchanged tests initially failed on a busy port, a chat row-lock race, and preview-server readiness. Each affected job passed after one CI rerun. Isolated local checks also passed: 41 credential tests, the chat-concurrency case, and 25 preview-runtime tests. - Greptile reviewed the final commit at 5/5. Both review threads are resolved. GitHub reports no merge conflicts. ## Risks - A missing personal account now defers authentication to the first task. Explicit incompatible bindings and restricted shared accounts still fail at hire time. - An inherited personal default uses the responsible user's existing authorization to install access for the new agent. It never copies credentials or another user's identity. - Subscription contention retries after a delay and rechecks task eligibility. It does not consume the provider-failure budget. - Native hiring uses the existing managed API-tools opt-in. Remote native runners require a matching Linux binary and provider pack, as documented in the acceptance README. - No schema changes. Live tests make paid provider calls and remain opt-in. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository inspection, code execution, browser automation, and API tools. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
827ba8a434 |
fix: cancel stalled sandbox startup without waiting for setup (#13352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox runs prepare credentials and files before an agent starts. > - Stop must work during that preparation. > - ACPX registered cancellation, but did not handle it while a setup command was waiting. > - Daytona cleanup waited for that same command before stopping the sandbox. > - This change stops the run's sandbox first and requires proof before abandoning setup. ## Linked Issues or Issue Description **What happened?** Stop left a sandbox run active when remote credential setup stalled. The run kept its connection lease until the sandbox was stopped separately. **Expected behavior** Stop terminates the selected run's sandbox, prevents later setup from launching the agent, and lets run cleanup finish. It must not report success without proof from the provider. **Steps to reproduce** 1. Start an ACPX agent in Daytona. 2. Hold a command during remote credential or file setup. 3. Select Stop before the agent starts. 4. Before this fix, cleanup waits for the held command and never reaches sandbox stop. Related: #13351 exposed this during connection acceptance testing. #12150 addresses scheduler load and session initialization limits, a separate startup problem. ## What Changed - Handle cancellation during ACPX sandbox preparation with a host-owned stop callback. - Pass an explicit active-work cancellation flag through environment cleanup. - Stop Daytona before draining setup commands. Keep normal graceful cleanup. - Require an exact run and lease termination receipt. Keep ownership of outstanding requests when stop cannot be verified. - Reject late setup work and defer sandbox resume until old requests settle. - Add regression tests and document the cancellation boundary. ## Verification - Red: both the stalled ACPX setup test and the Daytona cancellation test failed before the fix because Stop never reached the provider. - Green: adapter and Daytona suites passed, along with cancellation boundary and database-backed receipt/isolation tests. - Two real Daytona probes ran a five-minute setup command. Cancellation returned matching stopped receipts in 6.01 and 6.533 seconds. No later setup command ran. Both test sandboxes were deleted. - Repository typecheck and build passed locally. The full required CI suite passed, including all server/workspace tests, browser shards, Paperclip Runner verification, and canary packaging. The duplicate local full-suite run was interrupted after CI passed; it is not claimed as a completed local pass. - No UI changes. The live probe uses the actual adapter cancellation boundary and Daytona plugin; it is not a browser acceptance test. ## Risks - Provider stop failures remain unacknowledged. The adapter keeps ownership while its original requests remain active. - A retry can receive a settling-work error until old provider requests finish. - Cancelled sandboxes are stopped and retained under their existing provider expiry policy. - Local execution and cancellation after the agent turn starts keep their current behavior. - Other providers must return an exact termination receipt to permit early setup cancellation. No database migration. ## Model Used OpenAI GPT-6 through Codex, with code execution and tool use. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d1f0c20af |
fix: let responsible users choose either AI subscription or API key (#13351)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - AI Connections select the account used for each run. > - A responsible-user binding must follow the person whose work the agent performs. > - The saved sign-in method currently blocks users with another method for the same provider. > - This pull request resolves a personal default by company, user, and provider. > - Each user can use a subscription or API key with the same bot and model. ## Linked Issues or Issue Description Refs #13247, #13248, #13346, #13347. **What happened?** A bot configured with a Claude subscription rejects another responsible user’s Claude API key. Inline repair also limits that person to the original sign-in method. **Expected behavior** The same bot uses each responsible user’s default Claude account, whether it is a subscription or API key. Explicit shared account selections remain fixed. **Steps to reproduce** 1. User A connects a Claude subscription and creates a bot using the responsible user’s connection. 2. User C connects a personal Claude API key. 3. User C runs the same bot. Before this fix, credential resolution fails. ## What Changed - Add a personal provider-default table. Preserve legacy per-method preferences and backfill the most recently updated preference, including unavailable defaults. Repeated migration does not replace a selection. A database trigger propagates old-server default updates without treating new accounts as replacement defaults. - Resolve responsible-user bindings by provider. Retain the method as a wire compatibility hint for old servers. Explicit selections still require the exact method and grant. - Use the selected account’s method for credential isolation, refresh locking, and run attribution. - Update onboarding, agent setup, the picker, and inline task repair. Keep existing authentication components and harness/model settings. - Add mixed-method runtime, migration, repair, and Storybook coverage. Include upstream’s duplicate Anthropic option fix through the base branch. ## Verification - Focused resolver, migration, connection-intent, onboarding, agent setup, model, and connector UI suites: 331 tests passed. - Onboarding and new-agent regression suites passed during the initial focused run. - UI typecheck, token gates, and Storybook build passed. - Live browser checks passed for Claude and Codex API-default execution, switching both back to subscriptions, and both existing shared-account bots. Bot configuration remained unchanged. - One Daytona startup command stalled before Claude launched. The test run was cancelled, its sandbox stopped, and the same account/task passed on retry. Startup cancellation remains a separate environment finding; this PR does not change that command transport. - Browser review: all eight assertions passed in the new mixed-method story, including shared selection, return to responsible-user selection, and unchanged harness/model. - Repository build and typecheck passed after refreshing upstream dependencies. Final resolver and historical rollback verification: 38 tests passed. All latest-head CI gates passed, including server/workspace/serialized suites, browser E2E, build, typecheck, and runner verification. The extra serial local full-suite run was stopped after equivalent CI passed; focused local checks completed. ## Risks - Users with both historical method defaults get their most recently updated preference as the initial provider default. They can change it explicitly in Connections. - A revoked or unavailable default blocks. Connecting an additional account does not silently replace it. - Existing legacy authentication is unchanged. Managed responsible-user bindings intentionally stop pinning a method. - Live staging: the same Claude and Codex bots completed real API-key runs after changing only the personal default, then completed subscription runs after restoring the original defaults. Read-only database verification confirms unchanged bot configuration and actual method attribution. Distinct-user concurrency is covered by automated real-database tests with synthetic credentials, not two live human logins. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, code execution, and browser interaction. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e704e1c9ae |
fix: keep AI subscription locks safe through transaction pooling (#13347)
## Thinking Path > - Paperclip manages AI agents and their credentials. > - AI Connections must preserve account ownership from login through execution. > - A subscription environment check can pass and leave a session advisory lock on a pooled database backend. > - The next execution then reports that the subscription is busy. > - Onboarding can also lose its completion callback when the connection list refreshes before the login poll. > - This change keeps each credential lease in one transaction and keeps the login controller mounted until completion. ## Linked Issues or Issue Description Refs: #13247, #13248, #13344. **What happened?** A real hosted subscription login saved successfully and passed its environment check. The first task then failed with an AI connection busy error. A connection-list update could also leave onboarding on Connecting. **Expected behavior** A completed environment check releases its credential lease. A saved login completes onboarding without another sign-in. **Steps to reproduce** 1. Run Paperclip with a transaction-pooling PostgreSQL proxy. 2. Connect a Claude subscription through onboarding and reuse it for an agent. 3. Start a task after its environment check succeeds. 4. For the UI race, refresh the managed connection list before the login completion poll. ## What Changed - Hold the grant lock inside one transaction. Roll back on cleanup or failed acquisition. - Disable the lease transaction's idle timeout so long provider executions retain the lock. - Keep the onboarding login controller mounted while its authorization URL is active; restore saved-account retry if later agent creation fails. - Add lock-lifetime and Claude/Codex completion-order regressions. - Document the pooling requirement. ## Verification - 104 focused connection and onboarding tests passed. - Workspace typecheck, production build, Storybook build, and token gates passed. - All 34 latest-head checks are terminal: 32 passed and two Storybook workflows skipped by their normal filters. CI covers all server/workspace/serialized test shards, browser tests, typecheck, build, runner verification, and canary packaging. The additional local full-suite run is still in progress. - Real browser sign-ins saved Claude and OpenAI subscriptions on both new and existing staging stacks. On the patched new stack, both providers completed tasks and immediate repeat executions with the same saved accounts. Read-only database inspection confirmed live transaction-scoped locks and no remaining locks after completion. On the final deployed head, both shared accounts on the existing stack also completed actual tasks. Both personal accounts passed a fresh environment check followed immediately by execution after a deployment restart. ## Risks - Each active subscription reserves one database connection and holds an otherwise idle transaction until cleanup. This is required to pin the backend through a transaction pool. - Old session locks from earlier versions may need operator cleanup after active runs drain. This change does not bypass existing locks. - No database migration, credential routing, or legacy authentication change. ## Model Used OpenAI GPT-6 in Codex, with code execution and browser tools. The exact deployment model ID and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b2acc674be |
fix: enable isolated subscription login on authenticated self-hosted instances (#13344)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - AI Connections store reusable provider credentials with company and owner boundaries. > - Self-hosted instances can require board authentication while running agents on the local server. > - The login UI treated those instances as unsupported and displayed instructions without a command. > - Removing that restriction must not expose the server operator's existing CLI account. > - This pull request enables owner-scoped login attempts and reuses the existing login UI. > - Users can connect Codex and Claude subscriptions on an authenticated self-hosted instance. ## Linked Issues or Issue Description **What happened?** On an authenticated self-hosted instance, OpenAI subscription setup displayed terminal instructions with no command and a disabled Connect button. Claude also could not complete the local connection flow. **Expected behavior** An authorized board user can prepare an isolated sign-in attempt, sign in on the server, and save the verified account as a Connection. This must not import another user's or the operator's ambient credentials. **Steps to reproduce** Run Paperclip in authenticated mode with a local environment. Open Connections, choose OpenAI or Anthropic, and select Subscription. The previous UI never enabled local login preparation. Related: #13247, #13248, and #10751. This fix preserves the restriction on remote access to the operator's ambient Claude login. ## What Changed - Allow company-authorized users to create, check, cancel, and complete their own isolated local login attempts. - Keep ambient Claude credential import restricted to the local operator. - Support isolated Claude credential files without falling back to the host account or mutating process-wide environment variables. - Use Codex device authorization so sign-in does not depend on a browser callback to the remote server's localhost. - Gate server-host login on authenticated public deployments unless a trusted runtime host is configured. Publish the capability through health so setup shows supported alternatives. - Read isolated Claude credential files through bounded, descriptor-bound opens with ownership, permission, and symlink checks. Try the alternate filename after malformed JSON. - Reuse shared login instructions and lifecycle hooks in onboarding, agent setup, and Connections. Show health-query failures explicitly. - Document authenticated self-hosted behavior and add authorization, isolation, lifecycle, and UI regression tests. ## Verification - Passed 71 focused tests across connection routes, credential isolation, legacy compatibility, the shared login hook, and agent setup. - Passed 89 onboarding regression tests. - Passed `pnpm -r typecheck`, `pnpm build`, Storybook build, and `pnpm check:token-gates`. - Completed real Codex device authorization and Claude browser authorization on an authenticated Linux self-hosted instance. Both accounts were detected automatically and saved as Connected. Both completed attempt directories were removed. - These live checks cover login, credential validation, and connection creation. They do not establish a new model execution or long-running refresh result. - Review follow-up: 68 focused checks passed after rerunning one route socket error; the full route/health rerun passed all 49 tests. The 70-test onboarding suite also passed. Final workspace typecheck, production build, and Storybook build passed again. - The broad local run exposed an instance-name assumption in two new assertions. The fixture now uses an explicit non-default instance, and all 32 connection tests passed with a different inherited instance name. The superseded broad run was stopped; this is not a claim that the full local suite completed. Full CI results will be recorded before merge. ## Risks - The server must have the provider CLI installed. Users still run the displayed command on the server that hosts Paperclip. - Authorization checks must keep login attempts scoped to the company, owner, provider, and reconnect target. Regression tests cover cross-user and cross-company access. - Existing local-trusted Claude behavior stays available. Authenticated remote users cannot use its ambient import path. - No database migration, dependency change, agent binding change, or provider routing change is included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and browser tools. The runtime does not expose a more specific model version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d6232e7b0 |
feat: reuse provider sign-in across AI connection workflows (#13248)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect provider accounts during onboarding and agent setup. > - They should reuse and manage those accounts through the existing Connectors interface. > - A second login wizard would diverge from the established provider workflows. > - This pull request composes the existing sign-in components into Connections and agent configuration. > - Users can select accounts without changing their agent's harness or model. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add compact AI-account management to the existing Connectors pages. - Reuse AgentProviderConnection, AdapterLoginPanel, AdapterLoginChrome, and authentication controllers. - Add the shared connection picker to agent setup/settings and task requests. - Preserve onboarding's sequence and reuse existing accounts. - Add local-login recovery, retry, cancellation, and React StrictMode handling. - Add interactive Storybook scenarios, design-guide examples, and app acceptance checks. This is part 2 of the AI Connections change. The runtime foundation in #13247 is merged. This PR now targets master. ## Verification - Updated against master `47ded8bf9`, including the landed runtime foundation and upstream task-search changes. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All CI test, browser, build, packaging, and runner jobs passed on final head `dd17d3211931dd70aaa6ea619d83a7f9966dd18e`. The fresh Greptile review is 5/5, the security scan passed, and there are no unresolved review threads. The final CI aggregate gates passed. - Local general-server coverage passed 11,804 tests; three port-collision failures passed in an isolated 25-test rerun. All 6,111 UI tests passed. CLI coverage passed 484 tests; its remaining doctor test requires port 3199, which is occupied by an unrelated report server on this Mac. The complete CLI suite passed in CI. - Live browser testing verified Codex API-key reconnect inside a task card on desktop and phone. Real provider runs resumed and completed with unchanged connection/grant identity and agent routing. - Tested opening, cancelling, reopening, and completing connection creation. A regression confirms Connect another account cannot submit the new-agent form or copy provider keys into agent settings. - Added shared inline repair, automatic local sign-in checks, and responsive connection dialogs. Standalone Daytona installation ignores workspace configuration and suppresses dependency scripts. Its standalone build also passed with CI's exact pnpm 9.15.4. - Destructive live tests are excluded by default. Explicit opt-in, local deployment checks, and matching disposable fixture identities are required before any mutation. ## Risks - Local Codex/Grok creation requires the connection-specific terminal login command. - Browser sign-in uses the existing supported-environment controllers. - This update verifies live local Claude detection and Codex API-key task repair. New subscription authorization/refresh and independent-human/native-runner isolation were not reverified in this update. - No agent automatically adopts managed Connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
47ded8bf97 |
feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need credentials for a specific provider and sign-in method. > - Connections already owns accounts, grants, and access permissions. > - AI authentication should use those same boundaries. > - This pull request adds the storage, API, adoption, and runtime foundation. > - Legacy agents keep their authentication until they explicitly adopt a managed connection. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add AI-purpose/runtime-auth contracts and an additive, idempotent migration. - Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and catalog entries. - Store credentials on grants. Resolve responsible-user defaults or explicit permitted grants. - Isolate managed credentials and provider sessions across accounts. Block missing credentials without ambient fallback. - Keep imported legacy secrets unchanged during reconnect. Use independent local Codex/Grok sign-in attempts for rotating credentials. - Add authorization, migration, concurrent refresh, retry, cancellation, and legacy-compatibility tests. This is part 1 of a two-PR stack. The app UI follows in #13248. Merge the foundation first. ## Verification - Updated against master `04e364236`, preserving upstream provider login and connector workflows. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All current-head CI checks passed on `2a996560a`, including all server/workspace tests, browser shards, runner verification, typecheck, build, and canary dry run. Greptile reviewed that commit at 5/5 with no unresolved threads. Earlier local full-suite attempts hit the Mac PostgreSQL shared-memory limit; the complete suites passed in CI. - Renumbered the additive AI migration to `0276` after upstream migrations and regenerated its snapshot. Existing legacy agents retain their configuration. - Added local login status checks, owner-scoped retry, managed OpenCode remote homes, credential-aware model discovery, and task connection-repair delivery. ## Risks - Managed credential failures intentionally block execution. They do not restore legacy fallback. - Preview-era copied Codex/Grok subscriptions require independent reconnect. - The integrated branch has live provider acceptance coverage. This update verifies local Claude detection and Codex API-key task repair; it does not add a new subscription authorization/refresh or Daytona stress pass. - Runtime-auth connections must stay excluded from tool and channel handling. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |