Files
PaperClipAI/doc/connections/AI-CONNECTIONS.md
DottaandPaperclip 7d59de6113 feat(connections): probe provider usage limits on demand (#14936)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections store the AI accounts used by legacy and native runners.
> - Subscription accounts can reach session, weekly, model, or paid
usage limits.
> - Operators need to read these limits for a specific stored account
before making a routing decision.
> - This pull request adds an on-demand usage probe to the connection
service and account detail.
> - The result preserves provider limits, reset times, paid usage, and
unknown values for later consumers.

## Linked Issues or Issue Description

**Subsystem affected**

Shared contracts, the connection service and API, and the account detail
UI.

**Problem or motivation**

Managed AI accounts lack a common operation to read their current usage
limits. A local harness probe can read a different login from the
account selected for an agent.

**Proposed solution**

Add `aiConnectionService.probeUsage()` and a board-only connection usage
endpoint. Probe the selected credential grant on request. Support Codex,
Claude, and Grok subscriptions, plus OpenRouter API key limits.

**Alternatives considered**

Harness-specific automatic polling would couple the read to execution
and can read ambient credentials. This change uses the managed
connection credential and leaves scheduling and admission decisions to
later work.

**Roadmap alignment**

This extends the existing Personal & Shared AI Accounts capability. It
adds no routing or quota enforcement. Related: Refs #14459 for managed
OpenAI quota reads; Refs #14781 and Refs #13379 for downstream pacing
and budget work. This operation reads one requested account across all
three subscription providers.

## What Changed

- Add typed usage snapshots and a probe capability flag to managed AI
connections.
- Normalize Codex, Claude, Grok, and OpenRouter responses. Keep model
scopes, provider admission, reset periods, and paid allowances separate.
Preserve unknown values.
- Enforce company membership, credential audience, grant identity, and
connection lifecycle before reading the stored secret.
- Add a board-only `GET
/api/companies/:companyId/ai-connections/:connectionId/usage` endpoint
with `no-store` responses.
- Add manual **Check usage** and **Refresh** actions to account details.
Show compact usage bars, resets, admission and overage status; remove
repeated descriptions and account-default copy. Clear previous results
during a new request or error.
- Add Storybook previews using the production account components for all
four providers, initial checks, loading, and permission errors.
- Add provider, authorization, runner selection, API, and UI coverage.
Document provider sources and live qualification.

## Verification

- Initial provider, authorization, selection, API, and UI validation
passed (96 focused tests): `pnpm exec vitest run
server/src/services/ai-connection-usage.test.ts
server/src/__tests__/ai-connections.test.ts
ui/src/components/ai-connections/AiConnectionUsagePanel.test.tsx
server/src/__tests__/openapi-routes.test.ts`.
- `pnpm -r typecheck` passes for the initial implementation. After
simplifying the UI, 9 usage-panel and date-helper tests, UI typecheck,
token gates, and Storybook build pass. The initial feature module
boundary check also passed.
- Real Codex, Claude, and Grok credentials were saved to encrypted
disposable connections. The actual usage HTTP route returned 200 with
`status: ok`. Legacy and native runner selection checks passed. The
tests started no model turn and exchanged no refresh token. The
disposable databases and vaults were removed.
- Live Claude responses added structured scoped limits. Live Grok
responses omitted included-plan usage. Tests now cover both shapes and
preserve the Grok omission as unknown.
- The full workspace build passes. A full local test run hit a heartbeat
feedback timeout. That case passes in isolation. The duplicate local run
was stopped after all remote checks passed. The Slack ordering and
OpenCode transport CI flakes also pass in isolation and on the CI rerun.


- Current head: `ff3d479029a1c4248190323e221b2803cfb0d79d`. All 54
active checks pass. Two Storybook checks are intentionally skipped by
the workflow. Greptile is 5/5 with no unresolved review findings; the
branch is mergeable.

## Risks

- Subscription usage endpoints can change. Credentials can lack
usage-read permission. The probe returns explicit errors without fresh
limits in these cases.
- A successful probe can contain partial data. Missing utilization or
admission remains unknown. An enabled paid-usage switch does not prove a
funded balance.
- This change adds no migration. It does not change runner admission or
automatic provider selection. Provider requests use fixed endpoints,
disabled redirects, bounded response sizes, and a 15-second deadline.

## Model Used

OpenAI Codex, GPT-6, with reasoning, file editing, shell execution, and
HTTP tools. The session does not expose the exact runtime model variant
or context window size. Real provider credentials were used only for the
authorized live checks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 11:47:14 -05:00

34 KiB
Raw Permalink Blame History

AI Connections

AI accounts use the existing Apps/Connections substrate. Manage them at /:company/apps; select their use beside an agent's harness/model settings. Onboarding, new-agent setup, account creation/reconnect, and inline task requests reuse AdapterLoginPanel, its existing login controllers, and AdapterLoginChrome. Onboarding and new-agent setup retain upstream's SavedProviderKeySelect and useSavedProviderKeys, including saved-key references and account-specific Codex homes. Managed default/shared accounts are additional choices in that same selector. Selecting “Sign in to another account” survives background refreshes; Claude authorization paste keeps upstream's immediate Connecting feedback.

Storybook's simulated controllers and page annotations do not run in the app.

Compatibility and selection

The shared AI_CONNECTION_CAPABILITIES contract defines these combinations:

Provider Sign-in method Existing harness
Claude / Anthropic Claude subscription token or Anthropic API key Claude
OpenAI ChatGPT/Codex subscription or OpenAI API key Codex
OpenRouter API key OpenCode, with an openrouter/ model
Grok / xAI Grok subscription or xAI API key Grok

Native runner supports the corresponding existing Codex, OpenCode, and Claude ACP profiles. Connections creation and reconnect mount AgentProviderConnection, the same provider tiles, method controls, API entry, and AdapterLoginPanel used by agent setup. Supported sandbox environments use onboarding's existing browser sign-in controllers. Self-hosted installations use the shared terminal sign-in instructions described below and require no sandbox. Environment selection does not change agent execution settings. API keys are validated against fixed provider endpoints; redirects and caller-supplied validation URLs are rejected.

runtimeConfig.aiConnection contains provider, mode, and method. For responsible-user selections, method is a legacy wire hint retained for rolling upgrades; the resolver uses the selected account’s actual method:

  • responsible_user: resolve the run's responsible user's personal provider default, using that account's subscription or API key. The method hint does not restrict the responsible user's account.
  • shared: use the named connectionId and grantId, with audience and agent access checks.
  • delegated: retained only to read legacy bindings. It cannot bypass human access; a personal credential remains available only for its owner's tasks. New configuration offers personal defaults or shared accounts.

“Which humans can use this credential?” is the sole permission for whose work can use the account. “Just me” means the personal owner; shared accounts allow selected company members or every company member. The separate agent-access setting determines which agents can use it. There is no additional AI agent authorization, and old delegation records do not override the human audience.

A connection choice never changes the harness, model, or provider routing. Changing those separately may make a binding incompatible; saving then requires a compatible choice. Agent configuration cannot grant access to another account.

Personal defaults are unique per company, user, and provider. A Claude bot can use one user’s subscription and another user’s API key without changing its harness or model. Explicit shared selections remain pinned to the selected account and method. The first successful personal connection sets a default only when none exists. The additive ai_provider_defaults table preserves the legacy per-method preferences. Migration selects each user’s most recently updated provider preference (including unavailable accounts), and rerunning it never overwrites a provider default. New writes maintain the legacy table for older servers. A database trigger propagates older servers’ explicit default updates to the provider default. Inserting an additional method default does not replace an existing provider default.

Revocation retains the unavailable default; connecting another account does not silently replace it. Change it explicitly on the account detail page.

Agent settings offer Reconnect account when the current personal default needs attention. This repairs the same connection and keeps its default and agent access. Connect another account states that the new account will become the user's provider default. It selects the returned grant before adopting the binding and retains the actual sign-in method. A failed default update stays visible and can be retried without another login. New account setup shows an agent-access checkbox, enabled for all company agents by default for connection managers. The owner can limit access to the current agent. This access applies only to the owner's tasks. Reconnect never expands existing access.

Storage and API

AI connections pair connectionPurpose: ai with transport: runtime_auth. Database checks and the shared discriminator enforce the pair. These entries cannot participate in tool discovery, MCP gateways, execution, or channels. Anthropic offers Claude subscription and Claude API key; the unsupported duplicate REST API option is excluded. Catalog validation also pairs AI metadata with runtime authentication and rejects unsupported sign-in methods. Provider artwork and source provenance live in ui/public/brands/apps/manifest.json; OpenRouter uses its official sign-in assets, and OpenAI/Grok reuse the repository's pinned Lobe Icons source and license.

Provider/method metadata lives in config.ai. Credentials live on the existing grant through encrypted vault secret references, with existing consumer bindings. Safe provider-reported account identity is optional; secret references, tokens, and authentication paths are never account labels.

Company-scoped /api/companies/:companyId/ai-connections operations provide list, API-key creation/reconnect, personal defaults, completed login references, and active-run attribution. The list includes canManageConnections, evaluated by the same server permission check as creation, including custom tools:manage_connections grants. Existing Connections operations handle naming, access, and revocation. Mutation authorization is enforced server-side. OpenAPI documents the new board-only operations. Agent-originated configuration and environment tests resolve the authenticated request’s responsible user; an agent ID is never a personal-account owner. A missing responsible identity blocks personal-default resolution.

Subscription login attempts retain their company, owner, method, access intent, and reconnect target in the existing durable authentication session. Duplicate completion returns the same connection/grant. Abandoned or expired attempts cannot save a healthy connection. Reconnect preserves the connection ID, bindings, customized name, and access settings. A completed connection remains even if subsequent agent creation fails or is cancelled.

Account adoption during Save and the agent runtime test use the selected agent environment, or the instance default when no override is set. An unavailable remote environment blocks validation rather than probing the server host. The runtime test accepts the form’s prospective adapter selection before it is saved. For a saved-agent test, omitting environmentId uses the agent’s saved override. Sending environmentId: null tests a change back to the instance default.

On-demand usage limits

The AI account detail has Check usage. It reads the selected account only when requested; opening the account, listing connections, and starting either runner do not invoke this probe. There is no scheduling, provider selection, quota enforcement, credit purchase, or automatic retry attached to it.

aiConnectionService(db).probeUsage(companyId, userId, connectionId, grantId?) is the common server operation for legacy and native runner connections. GET /api/companies/:companyId/ai-connections/:connectionId/usage?grantId=... exposes it to board users. The UI client calls aiConnectionsApi.probeUsage. The optional grant ID must belong to this connection; omit it only when exactly one authorized grant exists. The company membership, personal owner/shared human audience, connection lifecycle, and credential ownership checks run before any provider request. Agent API keys cannot call this board endpoint. Internal execution callers must also use the normal runner connection selection checks for agent installation and responsible-user authority before using its result.

The returned AiConnectionUsage contains a timestamp and probe status (ok, unsupported, unavailable, or error), provider/source, every reported limit window, model/feature scope, exact window duration when known, reset time, utilization/remaining percentage, absolute allowance when reported, and overage settings/balance. The list summary advertises usageProbeSupported. limitReached describes that window; allowed preserves explicit provider admission for its limit group. A model-specific exhausted window must not be treated as an account-wide denial. Codex workspace spend controls are also returned when present. Percentages remain provider percentages: 0.5 means 0.5%, 99.99 stays below exhaustion, and values above 100 remain visible.

Capability means the provider has a probe implementation, not that every token has permission to read it. Claude setup-token credentials characterized in this repo request only user:inference; the provider can require user:profile for usage access. A 403 returns permission_denied with no limits, distinct from expired authentication or exhausted capacity. Minting another inference-only token does not establish usage access. This probe cannot add scopes to a stored token.

Connection Read-only source Observations
Codex subscription GET https://chatgpt.com/backend-api/wham/usage with stored access token and account ID Primary/secondary windows, additional feature/model limits, provider admission, workspace spend control, credit balance
Claude subscription GET https://api.anthropic.com/api/oauth/usage, with stored OAuth token and anthropic-beta: oauth-2025-04-20 Legacy and structured session/weekly/scoped windows, monthly extra usage, structured spend when legacy extra usage is absent
Grok subscription GET https://cli-chat-proxy.grok.com/v1/billing?format=credits, with stored token and x-xai-token-auth: xai-grok-cli Included-plan utilization/period when reported, separately reported on-demand allowance and prepaid balance in USD cents
OpenRouter API key GET https://openrouter.ai/api/v1/key Key credit cap, reset cadence, free-model daily request cap when returned
Other API-key methods No supported single-key allowance endpoint Explicit unsupported; no provider or secret read

These subscription sources are provider-client endpoints, not a promise of a stable public API. Codex's current endpoint/shape is grounded in the official backend client and response models. Claude uses the same OAuth endpoint as the existing adapter probe; official usage documentation describes plan windows, extra usage, and usage-check rate limiting. Grok's endpoint, optional fields, and USD-cent units follow its official billing implementation. It prefers creditUsagePercent and currentPeriod, with used/monthlyLimit as the legacy included-allowance fallback. The official Grok commands describe its usage screen. OpenRouter documents its key limits endpoint.

Unknown values are null, never assumed zero. A present Grok Cent message {} means zero under its documented proto3 encoding; an absent message or omitted included-usage percentage remains unknown. Claude's is_active dashboard flag does not establish whether requests are allowed. Named Claude limit groups keep their own identities and scopes, including session and weekly_all groups. They do not replace the legacy account-wide windows. An enabled extra-usage switch or remaining spend allowance does not establish a funded usable balance; overage available stays unknown unless the provider confirms it or reports it disabled/exhausted. An unlimited OpenRouter key does not establish account balance. A Grok reset timestamp alone does not establish weekly/monthly cadence. API rate limiting of the probe is an error, not an exhausted subscription. Malformed/empty responses and network failures return no fresh limits. Responses are no-store, bounded to 256 KiB and 15 seconds, use fixed endpoints with redirects disabled, and contain neither credentials nor raw provider errors. Expired credentials report authentication_required; the probe does not exchange refresh tokens or change connection health.

Verification:

pnpm exec vitest run server/src/services/ai-connection-usage.test.ts server/src/__tests__/ai-connections.test.ts ui/src/components/ai-connections/AiConnectionUsagePanel.test.tsx server/src/__tests__/openapi-routes.test.ts
pnpm check:token-gates

Fixtures prove normalization, credential isolation and manual UI behavior. Live provider/account entitlement qualification is separate; provider-side omissions remain unknown and must not be used downstream as proof of available capacity. On 2026-10-02, a real local Codex access token was saved to an encrypted, disposable test connection. The actual connection usage HTTP route returned 200 with status: ok, a weekly window at 100% used, its reset time, denied included usage, and available overage credits. The same stored connection passed selection checks for both codex_local and paperclip_runner with the native Codex provider. This verifies the connection API/vault/provider path and both selection paths; it does not claim a live runner turn or browser acceptance test. The disposable database and vault home were removed after verification. After the operator signed in again on 2026-10-02, real Claude Keychain and Grok file credentials were saved to encrypted disposable connections. Both the provider requests and actual connection usage HTTP route returned 200 with status: ok. Selection checks passed for claude_local, native Claude, ACP Claude, grok_local, and ACP Grok through paperclip_runner. Claude returned session and weekly usage at 0%, a scoped weekly window, and enabled extra usage. Grok returned a weekly period and zero on-demand allowance and prepaid balance, but omitted included-plan usage; it remains unknown rather than being reported as 0% or available capacity. The live shapes prompted regressions for Claude structured limits/spend and Grok legacy usage, monetary units, absent percentages, and proto3 zero messages. The disposable database and encrypted vault were removed. These checks started no provider turn and exchanged no refresh tokens.

Runtime isolation

Provider authentication failures, including acpx_auth_required, adapter login requirements, and expired/invalidated refresh tokens, create an AI connection card as the failed run is finalized. The card names the provider and uses the same inline connection/reconnect controls as missing-account setup. Pending cards are deduplicated. Creating a repair card persists a blocked run classification that suppresses immediate and periodic generic retries until the responsible user repairs the connection. Unsupported providers and failures that could not create a card retain their existing recovery path. Tool permission errors and provider quota failures do not request model authentication.

An attributed managed credential is marked as needing reauthorization only if its stored generation still matches the failed run. Late failures cannot invalidate a refreshed or reconnected credential. Repair preserves the selected account and its permissions when the same sign-in method is selected. The card also lets the user switch between API key and subscription authentication for providers that support both. Switching creates a separate account, then selects it as the user's provider default or validates and updates the agent's explicit account binding. The original account is retained. Accepting the card resumes with a fresh session through the existing durable continuation delivery.

For compatible legacy agents, the card offers the responsible person's provider connection without changing authentication automatically. After connecting, Use connection and continue checks agent-update permissions and validates in the agent's execution environment before committing the agent binding, connection install, audit, and card completion in one transaction. A failed validation or completion leaves the request pending and the old agent configuration and access intact. Unsupported harness/provider routes are not guessed.

Codex ACP terminal failures with category limit and explicit usage-exhaustion wording enter provider-quota recovery. A supported reset clock uses the existing Codex parser; when none is available, recovery uses its existing quota backoff. Context, turn, rate, storage-capacity and configured-budget limits retain their existing handling. The adapter inspects bounded provider text only in memory and retains recovery labels and a parsed timestamp, without copying the text to run results or logs. A historical generic terminal-limit message alone does not establish quota exhaustion.

prepareManagedAiRuntime is shared by runs, environment tests, and adoption. Test and Save mark the tested account as needing attention when its provider hello test rejects authentication or its API-key check returns 401 or 403. Network, quota, runtime, and environment failures do not change credential health. The same generation check protects a newer reconnect from a late test result. Claude ACP's typed access failure is its provider auth_required signal and enters the existing sign-in recovery path, including when only the generic terminal-access fallback message is available. Claude ACP validates working directories on the selected execution target. A sandbox directory does not need to exist on the Paperclip server. When the agent has no configured directory, the test uses the remote target's working directory. It checks responsible identity, membership, compatibility, connection health, human audience and agent installation before reading credentials. Missing credentials produce an actionable configuration failure; responsible-user task runs use the existing connection-request interaction, marked purpose: ai. A runtime-auth request cannot satisfy, reuse, or supersede a tool request for the same provider. AI-only methods are excluded from agent tool discovery.

Each invocation receives a private authentication home and only the selected grant's credentials. Inherited credential variables are cleared. Conflicting project authentication and provider-routing overrides are rejected. Managed failure cannot reactivate host or legacy credentials.

A subscription invocation takes no lease. Two invocations of one grant, from the same or a different provider account, run at the same time. At cleanup, each invocation re-reads the credential stored at that moment under a row lock on the grant, then compares it against its own refreshed copy using the provider's own freshness field: Codex compares last_refresh and bounds it against the host clock; Grok compares expires_at. The newer credential persists; a tie or an unparseable freshness value keeps the stored credential, so a spent single-use refresh token never overwrites a good one. Refreshes are merged only into the originating active grant, with a revocation check. Temporary homes are removed on normal completion or failure.

A fresh task execution cannot enter subscription contention. The freshest-write rule above resolves the conflict instead. A run that already entered this wait keeps a durable scheduled retry, checked every 60–120 seconds. The task shows “Waiting for AI subscription”. It does not request a reconnect, and it does not consume its provider-failure retry allowance. Each attempt rechecks task eligibility, ownership, budget, and current credential access. Revocation and other configuration failures still require user action. Authorized comment wakes that started as non-assignee runs can resume without claiming the assignee’s execution lock. Admission records this authority while holding the task and run locks. A reassignment during preflight cannot grant it. Assignee retries must still own that lock. Already-started native sessions retain their existing same-run recovery path; they must not be replaced by a fresh execution with a pre-provider receipt.

Session reuse includes grant identity, responsible user, and credential generation. A changed identity starts a fresh provider session. Native Codex (paperclip_runner) honors the configured warm lifecycle. It copies refreshed credentials back to the current invocation before deleting that invocation's private home. The session-owned credential stays private until idle timeout or explicit closure. Each follow-up rechecks current authorization and account identity before reusing the session; changing identity retires the previous owner. Other managed harnesses retain per-turn cleanup. A suspended native execution whose credential identity changed must restart as a new execution.

After a verified provider resume, plain-text Slack follow-ups send the new authorized message delta instead of repeating the full task framing. The saved run and current message identities and bodies must match. Actual brief edits still arrive; historical Slack task titles are not repeated as new directions. Attachments, omitted input, questions, approvals, and recovery retain their full framing. A fresh provider session always receives the complete bootstrap. Retained sandbox runner binaries are reused only after an exact SHA-256 match with the controller artifact and the normal capability checks. Run-scoped credential changes still require provider process rotation.

Warm sandbox execution requires both reuseLease: true and runnerLifecycleMode: "warm" on the environment. runnerIdleTimeoutMs bounds idle process retention. Chat tasks without a project reuse a sandbox only within the same company, environment, task, agent, and runtime configuration. They do not need an artificial project workspace. Other tasks, other agents, and ad-hoc connection tests cannot claim that retained sandbox. Daytona verifies a matching workspace sentinel before accepting either a workspace-scoped or task-scoped lease.

Revocation blocks new invocations and refresh persistence. A running provider process may already hold credentials. The revoke confirmation lists attributed active runs and exposes the existing Stop action; it does not promise immediate provider-side revocation.

Stop during sandbox preparation

ACPX startup registers cancellation while it materializes the remote auth home and stages files. Stop requests termination of that run's sandbox. Daytona closes admission and stops the sandbox before waiting for outstanding setup commands. The host requires a receipt for the exact company, run, and provider lease before abandoning the blocked setup RPC. Normal completion still drains work gracefully.

Late setup responses cannot launch the agent. A cancelled sandbox cannot resume while its old provider requests are still settling; a retry receives an explicit error instead. If termination cannot be verified, the adapter keeps ownership until the outstanding operation settles, and Stop is not acknowledged as complete. Local execution and cancellation of an already-running agent turn are unchanged.

Legacy adoption

Migration 0273 indexes only explicitly owned personal secrets with a recognized provider/method and matching agent configuration. It keeps original secret references and leaves every agent's legacy authentication unchanged. Reconnecting an indexed account creates a private grant credential instead of rotating the legacy secret. Subsequent reconnects rotate that private credential. Unknown ownership and filesystem-only subscriptions remain unresolved. The migration is repeatable and does not classify unknown credentials as company-shared.

Imported accounts initially need validation. Agent settings show “Existing authentication — not managed by Connections” until adoption. The adoption confirmation names the binding and affected agent. Saving runs a provider hello test in that agent's environment before replacing authentication. After adoption, the server preserves the managed binding and will not restore legacy fallback.

Local subscription sign-in

Local installations do not need a sandbox to connect a subscription. Connections, onboarding, and agent setup share LocalProviderLoginInstructions and useLocalAiLogin. In local-trusted mode, Claude checks the operator’s existing Claude Code login. Authenticated self-hosted users instead get a separate CLAUDE_CONFIG_DIR for claude auth login; checking and saving only read that attempt’s credential files, never the server operator’s account or Keychain.

Codex and Grok start a separate terminal sign-in for each connection or reconnect. The shared component shows a server-generated command with a fresh CODEX_HOME or GROK_HOME. Codex uses file credential storage in that home and login --device-auth, so signing in from another computer does not depend on a localhost callback. The home is never seeded with the operator's existing login: copying a rotating refresh token would allow managed runs to invalidate credentials still used by legacy agents or the operator's terminal. The user completes browser sign-in from that command, then clicks Connect. This does not require a sandbox or change the host login.

Attempts reuse adapter_auth_sessions, binding company, owner, provider, access intent, reconnect target, and a 30-minute expiry. Validation and completion are serialized; duplicate completion returns the saved connection. Restart retains the attempt. Cancellation and expiry remove the attempt home, and the startup/ periodic cleanup sweep retries expired directories. Successful completion persists credentials to the encrypted grant and removes the temporary login home. Refreshes subsequently update only that grant. Reconnect preserves IDs and access settings.

Starting an isolated attempt requires normal company-scoped AI-connection creation permission. Checks, completion, cancellation, and resumption are owner-bound. Authenticated users cannot import host credentials or use another user’s attempt. Claude Keychain reads remain limited to the explicit local-trusted default-home import. A failed verification creates no healthy connection. Preview-era Codex/Grok managed connections without the isolated-subscription marker require reconnect before another managed execution; unmanaged legacy agents retain their existing authentication paths.

Verification

server/src/__tests__/ai-connections.test.ts exercises storage, isolation, defaults, human audiences, agent access, reconnect races, refresh ownership, concurrent subscription write-backs, migration replay, and redacted API failures against a real embedded database. Existing login, adapter, tool, and channel suites cover their shared integration paths. The onboarding tests cover managed reuse and keeping a successfully connected account after failed agent creation.

The Storybook review index retains deterministic authentication states and interaction checks. Run pnpm build-storybook, then pnpm exec playwright test --config tests/ai-connections-review/playwright.config.ts. Also run token gates, repository typecheck, tests, and build before handoff. Live connect → reuse → run → reconnect verification still requires valid provider credentials and a supported login/runtime environment; fixtures do not prove it.

For an isolated running test drive, also run:

AI_CONNECTIONS_TEST_COMPANY_ID=<company-id> pnpm exec playwright test --config tests/ai-connections-app/playwright.config.ts

Set AI_CONNECTIONS_TEST_URL when the test drive uses a port other than 3100.

These browser checks exercise the production list/detail pages, rejected API-key validation, cancellation, focus restoration, and adoption without saving agent changes. They submit an explicitly invalid fixture key and do not prove successful authentication with a live account.

Local sign-in checks

Local subscription screens share the same credential check on entry and when the window regains focus. Waiting screens also poll until sign-in verifies. A successful check shows the account is signed in; only Connect creates or reconnects the grant. In local-trusted mode, Claude checks the local operator’s Claude Code login. Authenticated Claude users, plus all Codex and Grok users, check only their connection-specific login home. The health response selects credential isolation, not whether a self-hosted user may sign in.

Leaving and returning to a local sign-in screen resumes its active attempt. Navigation does not delete a directory referenced by a copied command. Start sign-in again explicitly cancels the old attempt; abandoned attempts expire after 30 minutes. Commands create their directory if necessary, and completed/expired attempts are cleaned up through the existing lifecycle.

Disposable live inline-repair test

The normal app test configuration excludes *.live.spec.ts. To run the destructive inline-repair scenario, set AI_REPAIR_TEST_ALLOW_DESTRUCTIVE=1 and use a separate loopback local_trusted instance. Set AI_REPAIR_TEST_DISPOSABLE_MARKER to a fresh 32-character lowercase hexadecimal value. The company, single Codex agent, single personal OpenAI API connection, and issue must all be named AI Repair QA <marker> (the issue uses that title). Supply their IDs with AI_CONNECTIONS_TEST_COMPANY_ID, AI_REPAIR_TEST_CONNECTION_ID, and AI_REPAIR_TEST_ISSUE_ID, and the disposable provider key with AI_REPAIR_TEST_KEY. The test verifies these boundaries before revoking credentials or submitting work. Delete the disposable instance and revoke its provider key after the test; failed tests may leave a paused task for inspection.

Authenticated public deployments must configure a trusted runtime host (PAPERCLIP_TRUSTED_MCP_RUNTIME_HOST or PAPERCLIP_TOOL_RUNTIME_TRUSTED_HOST) before offering server-host subscription login, matching the local stdio runtime boundary. Health reports this capability so setup can offer a supported environment or API key instead of an unusable terminal command. Private authenticated self-hosted instances support isolated local login without that extra setting. Isolated Claude credential files must be private, owned by the server user, bounded, and free of symlinks.

Hiring and delegated work

When a managed agent creates or hires another agent without an explicit AI binding or adapter auth setting, the server inherits its compatible managed connection choice. Explicit credentials, blank overrides, credential directories, and provider routing settings for the child provider take precedence. Unrelated provider keys do not suppress the default. Unmanaged parents keep their existing authentication path. The new agent resolves the responsible user's account at execution time; it never copies the parent's credentials or identity. Same-provider hires preserve subscription/API-key choice. A different provider selects the responsible user's default for that provider. Native Codex and ACPX/Claude provider selections follow the same compatibility rules.

Hiring may succeed before that personal account exists or while it needs repair, including hires awaiting board approval. The first assigned task then shows an AI connection card. First-time setup presents the provider's subscription/API controls inside the task. The inline form omits the connection name field and names new accounts from the user's display name, provider, and selected authentication method (for example, dotta's Claude API account). Reconnecting preserves the existing account name. Connecting installs access for that agent and resumes the pending work automatically. Explicit incompatible bindings and shared-account permission denials still fail; hiring never expands a restricted shared account's audience.

Concurrent runs of one subscription do not wait for each other. No credential lease exists to hold them, so a fresh task execution cannot enter a contention wait. A run that already entered this wait keeps its scheduled retries. It does not request new credentials, and it does not consume the provider-failure retry allowance. Each retry revalidates the account, and existing run-dispatch rules still suppress cancelled, reassigned, or otherwise ineligible work. An assignee retry must still keep execution-lock ownership at scheduling, promotion, and dispatch.

server/src/__tests__/agent-hire-ai-connections.test.ts covers both creation routes, both providers and methods, approval gates, native provider mapping, shared access boundaries, and concurrent runs of one subscription for both providers. The opt-in tests/hiring-ai-connections/README.md describes real browser hiring, subtask, connection, and automatic-resume checks on local and Daytona environments, plus the production component Storybook checks.

Managed session compatibility

Resume checks compare the selected account identity with the server-owned metadata in the saved task session. Read this metadata before decoding the adapter session: adapter codecs intentionally discard unknown fields. A missing identity, a different grant or responsible user, or a changed credential generation requires a fresh session. The metadata is removed before passing session params to an adapter. Temporary authentication-home paths do not change the configuration fingerprint. These checks do not relax current connection authorization.