Files
PaperClipAI/doc/runner-api-tools.md
T
DottaandPaperclip f912ecaacf fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents hire other agents and assign tasks to them.
> - Managed AI connections must follow those hires across legacy and
native runners.
> - Missing accounts should pause task execution and let the user
connect from the task.
> - Subscription contention must wait without asking for new
credentials.
> - This pull request fixes these paths and the native tool and Daytona
staging failures found during live tests.
> - The result is a working hire, subtask, and connection setup flow on
local and remote runners.

## Linked Issues or Issue Description

**What happened?**

A managed Claude or Codex agent could hire a teammate without a usable
AI binding. Cross-provider hiring could fail before the user had a
chance to connect the new provider. First-time task setup did not show
the existing AI credential form inline. A busy subscription could
request a new connection. Native API replies could stop the parent after
a hire had already committed. Fresh Daytona sandboxes could fail to
extract read-only skill directories created on macOS.

**Expected behavior**

Compatible hires inherit the managed connection choice. A hire for
another provider uses the responsible user's default. If that account is
missing, the hire succeeds and the task asks for a connection.
Completing setup in the task resumes work automatically. Explicit child
auth settings and existing unmanaged login paths keep precedence.
Shared-account access checks remain in force.

**Steps to reproduce**

1. Connect a Claude or Codex parent with a managed AI account.
2. Ask it to hire one agent of each provider and create a self-assigned
subtask.
3. Assign work to both hires without connecting the second provider
first.
4. Connect the missing provider from its task card.
5. Check that all tasks finish and same-provider work uses the original
account.
6. Repeat with native runners and fresh Daytona sandboxes. The opt-in
browser suite in `tests/hiring-ai-connections/README.md` performs these
steps.

**Paperclip version or commit**

The live failures were reproduced from `f2c5e54dc`. The branch is
rebased onto `5282cabde`.

**Deployment mode**

Isolated local development instance. Legacy CLI and native runners.
Local execution and ephemeral Daytona sandboxes.

Related work: Refs #13247 for managed AI connections. Refs #13268 for
legacy credential-reference inheritance, which this branch preserves.
Refs #13432 for a concurrent managed-inheritance fix. This PR also
covers cross-provider task setup, subscription waits, native API
replies, and Daytona extraction. It permits missing responsible-user
defaults at hire time; restricted shared selections still fail.

## What Changed

- Apply managed connection defaults to both agent creation routes.
Preserve explicit auth choices and legacy credential-reference
inheritance.
- Allow hires before their responsible user connects the provider. Keep
approval gates, company boundaries, and shared-account access checks.
- Reuse the production AI credential form inside the pending task card.
Resume the task after setup.
- Retry subscription lease contention without consuming the
provider-failure allowance or creating a connection request.
- Require task execution-lock ownership when scheduling, promoting, and
dispatching subscription retries. Recheck ownership under the issue row
lock.
- Rename the HTTP operation identity at the native tool boundary so it
cannot override the runner's operation identity.
- Delay directory permission restoration during Daytona extraction.
Preserve the final read-only modes.
- Add database-backed regressions, real browser acceptance tests, and
Storybook states. Document setup and run-log behavior.

## Verification

- Six real browser scenarios passed: both parent providers on legacy
local and legacy Daytona; native Codex locally; native Claude on
Daytona. Each scenario hires both providers, completes a self-subtask
and assigned work, and connects the missing provider inline with
automatic continuation.
- Successful runs verify the account, responsible user, runner mode, and
Daytona lease. All 18 test sandboxes were deleted.
- Live authentication used API keys. Subscription inheritance, lease
contention, and retry have integration coverage. Fresh subscription
OAuth sign-in was not automated.
- Red/green tests reproduced missing bindings, missing inline forms,
subscription contention, native API reply failure, and GNU tar
permission failure.
- Seven Storybook browser checks passed. They cover both providers,
method selection, narrow layout, completion, cancellation, and invalid
credentials.
- Full local suite coverage completed before rebase. Initial timing and
fixture startup failures passed unchanged on isolated reruns. The first
full command did not exit cleanly; the remaining workspace and
serialized groups were completed separately.
- After rebase, 107 hiring/auth/retry tests and 59 native API,
task-card, and Daytona tests passed. The full workspace typecheck,
production build, and token gates passed again. Storybook build passed
before rebase.

- Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and
four explicit cancellation-race cases passed. Eight cross-provider cases
cover stale auth keys on both creation routes and both runner types.
Server typecheck and build passed.

- Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks
passed. Two optional Storybook jobs were skipped. The full server,
workspace, browser, native runner, build, typecheck, and release checks
passed.
- Three unchanged tests initially failed on a busy port, a chat row-lock
race, and preview-server readiness. Each affected job passed after one
CI rerun. Isolated local checks also passed: 41 credential tests, the
chat-concurrency case, and 25 preview-runtime tests.
- Greptile reviewed the final commit at 5/5. Both review threads are
resolved. GitHub reports no merge conflicts.

## Risks

- A missing personal account now defers authentication to the first
task. Explicit incompatible bindings and restricted shared accounts
still fail at hire time.
- An inherited personal default uses the responsible user's existing
authorization to install access for the new agent. It never copies
credentials or another user's identity.
- Subscription contention retries after a delay and rechecks task
eligibility. It does not consume the provider-failure budget.
- Native hiring uses the existing managed API-tools opt-in. Remote
native runners require a matching Linux binary and provider pack, as
documented in the acceptance README.
- No schema changes. Live tests make paid provider calls and remain
opt-in.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository
inspection, code execution, browser automation, and API tools. The
context-window size is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 16:15:48 -05:00

8.9 KiB

Runner API escape hatch

search_api and call_api extend the native runner when an available dedicated operation cannot express the requested work. Existing tools remain preferred; agents do not have to search before using them. Only two tool definitions are advertised. The API catalog is returned on demand, never injected into the initial prompt.

Controlled rollout

The escape hatch is disabled by default. Set PAPERCLIP_RUNNER_API_TOOLS_ENABLED=true on the server to enable it. For an initial company rollout, also set PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS to a comma-separated list of company UUIDs. An unset list allows every company; an explicitly empty list allows none. IDs must match exactly.

The server always requires the explicit true flag, including for server-owned bindings. A binding can disable these tools for a baseline eval but cannot enable them without operator opt-in. Setting the flag to false disables them. The server checks this switch when advertising tools, when accepting a call, and immediately before HTTP dispatch after preparing any files. Existing dedicated tools remain available. Operators must update the environment of each server process and restart it for deployment-level changes; this environment switch is not a live settings API.

Evaluate selected companies first. Compare success, unnecessary fallback calls, cost, and latency against the dedicated-tool baseline before widening. Keep the switch disabled if authorization, replay, or cost accounting fails.

Discovery and requests

{"query":"create project","limit":5}

Search is deterministic lexical ranking over OpenAPI paths, summaries and the old skill reference. It supports task/issue and other terminology, exact METHOD /api/path/{parameter} lookup, and opaque query/catalog-bound pagination. Results include resolved request schemas, response descriptions, authorization metadata, work modes, examples where available, and relevant dedicated tools with their supported parameters. limit defaults to five and is capped at ten.

{"operationId":"PATCH /api/projects/{id}","pathParams":{"id":"PROJECT_UUID"},"body":{"description":"Updated project description"}}

The catalog determines method and path. companyId is filled from the active binding. Scalars and arrays are accepted in query. body defaults to JSON; contentType supports text and raw uploads. files accepts entries containing exactly one authorized artifactId or task-workspace path, and an optional multipart field. No arbitrary URL, headers, authentication, or remote file URL can be supplied. Routes still validate payloads and enforce permissions.

Requests have a 30-second HTTP timeout, 16 KiB URL limit and 10 MiB payload/response transfer limit. Responses above 24 KiB and binary responses become company-owned assets with retrievable references; text previews are limited to 2,000 bytes. Tool responses identify the HTTP route with apiOperationId. The native protocol reserves operationId and callId for semantic tool-call identity; API metadata must not masquerade as that envelope. Saved mutation receipts are normalized at the tool boundary as well, without repeating their HTTP request. All redirects are refused. Oversized or interrupted mutation responses have an unknown outcome, requiring inspection before another mutation. Mutation responses with HTTP 5xx, HTTP 408, redirects, or malformed JSON also retain an unknown outcome. A server may have committed the write before it failed to return a valid response.

Authority and replay

The server revalidates the active native run, assigned task and actor, then creates a server-held agent JWT bound to that company and run. Requests go through the actual HTTP router with its authorization, validation and domain audit behavior. An additional runner.api_called receipt attributes mutations to the run even where older route audit events omit that field. The run and work mode are checked again after asynchronous file preparation, so a stopped run cannot dispatch an upload prepared under its earlier binding.

Ask and pre-acceptance Plan permit reads through the escape hatch. Existing dedicated-tool exceptions are unchanged. Runner-owned checkout, completion, status/assignment transitions, approval decisions and execution-control actions cannot be bypassed through generic calls. Routine creation, schedule/trigger changes and manual/public routine execution require the existing scheduling clients. Direct workspace runtime commands, runtime-slot stop/restart, case automation retries and skill test-run controls also require their existing execution clients. Gateway session credentials cannot enter generic results. Routine metadata remains readable; annotation threads, comments and thread resolution remain available through the fallback. API-only ordinary fields, such as a task's billingCode, remain accessible even when a dedicated tool covers other fields on that endpoint.

Mutation call IDs reserve a durable receipt in the run's existing resultJson before dispatch. Replays return the recorded result. Reusing an ID with different arguments is rejected. A crash after reservation leaves an unknown outcome and never automatically resends the mutation. The limit is 512 mutation receipts per run. No database migration is needed.

Workspace uploads use the existing workspace resource containment checks, no-symlink file opens covering every path component, and bounded descriptor reads. Local uploads require Linux or macOS; authorized artifacts work on other hosts. Lifecycle-sensitive endpoints require an inline JSON object, so a raw uploaded JSON file cannot hide protected fields from policy checks. Artifacts must belong to the bound company. Secret-value access, credential management, secret proposals and company exports require their existing secure clients. Search describes these operations as restricted. call_api rejects them before creating a replay receipt or making an HTTP request. Safe secret metadata listing remains available. Agent credentials are never returned to the model. Streaming, WebSocket, MCP and authentication handshakes are documented as protocol operations requiring their existing clients.

Catalog maintenance

runner-api-catalog.ts builds from the server OpenAPI registry. Experimental pipeline, Cases and smoke-lab routes now share their validators with discovery. Seven Cases/pipeline route shapes are multiplexed by resource identity: the Cases router intentionally forwards unknown resources to the pipeline router. Their separate catalog entries explain which resource identifier is required. Registry authorization descriptions are documentation; actual route checks are authoritative.

Regenerate old-skill enrichment after editing its API reference:

node scripts/generate-runner-api-reference.mjs
node scripts/generate-runner-api-reference.mjs --check
node scripts/generate-runner-experimental-api-metadata.mjs
node scripts/generate-runner-experimental-api-metadata.mjs --check

Mounted-route coverage tests include experimental routes. Three WebSocket mounts are explicitly classified in the catalog. Shared protocol-action catalogs, provider projections and generated compatibility checks include both tools.

Verification and paid evals

The companion paperclip-evals worktree contains evals/runner-api-tools. Its README documents explicit case/model selectors, the cumulative budget ledger, fixture reset, progressive batches, and Evalbook generation. No command defaults to running the entire paid suite. Capability, forced operation contracts and paired common-operation regressions are reported separately.

Provider-free integration tests exercise real runnerd → PRP → authority → HTTP, route validation and audit, stale bindings, Ask/Plan restrictions, identity spoofing, file containment, uncertain mutation receipts and fixture isolation.

The Evalbook viewer uses the existing shared viewer and stylesheet on master. The report retains actual persisted-state summaries for private local inspection; public replay continues to withhold company-state details.

The ACPX sidecar includes the upstream terminal-usage accounting correction from origin/codex/evalbook-default-chat-sept6. Its qualified Claude executable requires Linux x64. The first macOS stage records a zero-cost ACPX admission failure. A later user-authorized OpenCode/OpenRouter Sonnet profile reached a real HTTP read, but the attempt failed on a missing harness completion contract and incomplete terminal accounting. The harness contract is corrected. The missing fourth request was subsequently recovered from the matching OpenRouter session and generation billing record; the original failed attempt remains immutable. New attempts retain an append-only, flushed event journal and bounded provider trace outside disposable runtime files. Provider-free startup succeeds for OpenRouter Sonnet and DeepSeek. See doc/plans/2026-09-07-runner-api-production-readiness.md for remaining release gates.