mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents hire other agents and assign tasks to them. > - Managed AI connections must follow those hires across legacy and native runners. > - Missing accounts should pause task execution and let the user connect from the task. > - Subscription contention must wait without asking for new credentials. > - This pull request fixes these paths and the native tool and Daytona staging failures found during live tests. > - The result is a working hire, subtask, and connection setup flow on local and remote runners. ## Linked Issues or Issue Description **What happened?** A managed Claude or Codex agent could hire a teammate without a usable AI binding. Cross-provider hiring could fail before the user had a chance to connect the new provider. First-time task setup did not show the existing AI credential form inline. A busy subscription could request a new connection. Native API replies could stop the parent after a hire had already committed. Fresh Daytona sandboxes could fail to extract read-only skill directories created on macOS. **Expected behavior** Compatible hires inherit the managed connection choice. A hire for another provider uses the responsible user's default. If that account is missing, the hire succeeds and the task asks for a connection. Completing setup in the task resumes work automatically. Explicit child auth settings and existing unmanaged login paths keep precedence. Shared-account access checks remain in force. **Steps to reproduce** 1. Connect a Claude or Codex parent with a managed AI account. 2. Ask it to hire one agent of each provider and create a self-assigned subtask. 3. Assign work to both hires without connecting the second provider first. 4. Connect the missing provider from its task card. 5. Check that all tasks finish and same-provider work uses the original account. 6. Repeat with native runners and fresh Daytona sandboxes. The opt-in browser suite in `tests/hiring-ai-connections/README.md` performs these steps. **Paperclip version or commit** The live failures were reproduced from `f2c5e54dc`. The branch is rebased onto `5282cabde`. **Deployment mode** Isolated local development instance. Legacy CLI and native runners. Local execution and ephemeral Daytona sandboxes. Related work: Refs #13247 for managed AI connections. Refs #13268 for legacy credential-reference inheritance, which this branch preserves. Refs #13432 for a concurrent managed-inheritance fix. This PR also covers cross-provider task setup, subscription waits, native API replies, and Daytona extraction. It permits missing responsible-user defaults at hire time; restricted shared selections still fail. ## What Changed - Apply managed connection defaults to both agent creation routes. Preserve explicit auth choices and legacy credential-reference inheritance. - Allow hires before their responsible user connects the provider. Keep approval gates, company boundaries, and shared-account access checks. - Reuse the production AI credential form inside the pending task card. Resume the task after setup. - Retry subscription lease contention without consuming the provider-failure allowance or creating a connection request. - Require task execution-lock ownership when scheduling, promoting, and dispatching subscription retries. Recheck ownership under the issue row lock. - Rename the HTTP operation identity at the native tool boundary so it cannot override the runner's operation identity. - Delay directory permission restoration during Daytona extraction. Preserve the final read-only modes. - Add database-backed regressions, real browser acceptance tests, and Storybook states. Document setup and run-log behavior. ## Verification - Six real browser scenarios passed: both parent providers on legacy local and legacy Daytona; native Codex locally; native Claude on Daytona. Each scenario hires both providers, completes a self-subtask and assigned work, and connects the missing provider inline with automatic continuation. - Successful runs verify the account, responsible user, runner mode, and Daytona lease. All 18 test sandboxes were deleted. - Live authentication used API keys. Subscription inheritance, lease contention, and retry have integration coverage. Fresh subscription OAuth sign-in was not automated. - Red/green tests reproduced missing bindings, missing inline forms, subscription contention, native API reply failure, and GNU tar permission failure. - Seven Storybook browser checks passed. They cover both providers, method selection, narrow layout, completion, cancellation, and invalid credentials. - Full local suite coverage completed before rebase. Initial timing and fixture startup failures passed unchanged on isolated reruns. The first full command did not exit cleanly; the remaining workspace and serialized groups were completed separately. - After rebase, 107 hiring/auth/retry tests and 59 native API, task-card, and Daytona tests passed. The full workspace typecheck, production build, and token gates passed again. Storybook build passed before rebase. - Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and four explicit cancellation-race cases passed. Eight cross-provider cases cover stale auth keys on both creation routes and both runner types. Server typecheck and build passed. - Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks passed. Two optional Storybook jobs were skipped. The full server, workspace, browser, native runner, build, typecheck, and release checks passed. - Three unchanged tests initially failed on a busy port, a chat row-lock race, and preview-server readiness. Each affected job passed after one CI rerun. Isolated local checks also passed: 41 credential tests, the chat-concurrency case, and 25 preview-runtime tests. - Greptile reviewed the final commit at 5/5. Both review threads are resolved. GitHub reports no merge conflicts. ## Risks - A missing personal account now defers authentication to the first task. Explicit incompatible bindings and restricted shared accounts still fail at hire time. - An inherited personal default uses the responsible user's existing authorization to install access for the new agent. It never copies credentials or another user's identity. - Subscription contention retries after a delay and rechecks task eligibility. It does not consume the provider-failure budget. - Native hiring uses the existing managed API-tools opt-in. Remote native runners require a matching Linux binary and provider pack, as documented in the acceptance README. - No schema changes. Live tests make paid provider calls and remain opt-in. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository inspection, code execution, browser automation, and API tools. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
156 lines
8.9 KiB
Markdown
156 lines
8.9 KiB
Markdown
# Runner API escape hatch
|
|
|
|
`search_api` and `call_api` extend the native runner when an available dedicated
|
|
operation cannot express the requested work. Existing tools remain preferred;
|
|
agents do not have to search before using them. Only two tool definitions are
|
|
advertised. The API catalog is returned on demand, never injected into the
|
|
initial prompt.
|
|
|
|
## Controlled rollout
|
|
|
|
The escape hatch is disabled by default. Set
|
|
`PAPERCLIP_RUNNER_API_TOOLS_ENABLED=true` on the server to enable it. For an
|
|
initial company rollout, also set `PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS` to a
|
|
comma-separated list of company UUIDs. An unset list allows every company;
|
|
an explicitly empty list allows none. IDs must match exactly.
|
|
|
|
The server always requires the explicit `true` flag, including for server-owned
|
|
bindings. A binding can disable these tools for a baseline eval but cannot enable
|
|
them without operator opt-in. Setting the flag to `false` disables them. The server checks this switch when advertising tools,
|
|
when accepting a call, and immediately before HTTP dispatch after preparing any
|
|
files. Existing dedicated tools remain available. Operators must update the
|
|
environment of each server process and restart it for deployment-level changes;
|
|
this environment switch is not a live settings API.
|
|
|
|
Evaluate selected companies first. Compare success, unnecessary fallback calls,
|
|
cost, and latency against the dedicated-tool baseline before widening. Keep the
|
|
switch disabled if authorization, replay, or cost accounting fails.
|
|
|
|
## Discovery and requests
|
|
|
|
```json
|
|
{"query":"create project","limit":5}
|
|
```
|
|
|
|
Search is deterministic lexical ranking over OpenAPI paths, summaries and the
|
|
old skill reference. It supports task/issue and other terminology, exact
|
|
`METHOD /api/path/{parameter}` lookup, and opaque query/catalog-bound pagination.
|
|
Results include resolved request schemas, response descriptions, authorization
|
|
metadata, work modes, examples where available, and relevant dedicated tools
|
|
with their supported parameters. `limit` defaults to five and is capped at ten.
|
|
|
|
```json
|
|
{"operationId":"PATCH /api/projects/{id}","pathParams":{"id":"PROJECT_UUID"},"body":{"description":"Updated project description"}}
|
|
```
|
|
|
|
The catalog determines method and path. `companyId` is filled from the active
|
|
binding. Scalars and arrays are accepted in `query`. `body` defaults to JSON;
|
|
`contentType` supports text and raw uploads. `files` accepts entries containing
|
|
exactly one authorized `artifactId` or task-workspace `path`, and an optional
|
|
multipart `field`. No arbitrary URL, headers, authentication, or remote file URL
|
|
can be supplied. Routes still validate payloads and enforce permissions.
|
|
|
|
Requests have a 30-second HTTP timeout, 16 KiB URL limit and 10 MiB payload/response
|
|
transfer limit. Responses above 24 KiB and binary responses become company-owned
|
|
assets with retrievable references; text previews are limited to 2,000 bytes.
|
|
Tool responses identify the HTTP route with `apiOperationId`. The native protocol
|
|
reserves `operationId` and `callId` for semantic tool-call identity; API metadata
|
|
must not masquerade as that envelope. Saved mutation receipts are normalized at
|
|
the tool boundary as well, without repeating their HTTP request.
|
|
All redirects are refused. Oversized or interrupted mutation responses have an
|
|
unknown outcome, requiring inspection before another mutation.
|
|
Mutation responses with HTTP 5xx, HTTP 408, redirects, or malformed JSON also
|
|
retain an unknown outcome. A server may have committed the write before it
|
|
failed to return a valid response.
|
|
|
|
## Authority and replay
|
|
|
|
The server revalidates the active native run, assigned task and actor, then
|
|
creates a server-held agent JWT bound to that company and run. Requests go
|
|
through the actual HTTP router with its authorization, validation and domain
|
|
audit behavior. An additional `runner.api_called` receipt attributes mutations
|
|
to the run even where older route audit events omit that field.
|
|
The run and work mode are checked again after asynchronous file preparation, so
|
|
a stopped run cannot dispatch an upload prepared under its earlier binding.
|
|
|
|
Ask and pre-acceptance Plan permit reads through the escape hatch. Existing
|
|
dedicated-tool exceptions are unchanged. Runner-owned checkout, completion,
|
|
status/assignment transitions, approval decisions and execution-control actions
|
|
cannot be bypassed through generic calls. Routine creation, schedule/trigger
|
|
changes and manual/public routine execution require the existing scheduling
|
|
clients. Direct workspace runtime commands, runtime-slot stop/restart, case
|
|
automation retries and skill test-run controls also require their existing
|
|
execution clients. Gateway session credentials cannot enter generic results.
|
|
Routine metadata remains readable; annotation threads, comments and thread
|
|
resolution remain available through the fallback. API-only ordinary fields, such as a
|
|
task's `billingCode`, remain accessible even when a dedicated tool covers other
|
|
fields on that endpoint.
|
|
|
|
Mutation call IDs reserve a durable receipt in the run's existing `resultJson`
|
|
before dispatch. Replays return the recorded result. Reusing an ID with different
|
|
arguments is rejected. A crash after reservation leaves an unknown outcome and
|
|
never automatically resends the mutation. The limit is 512 mutation receipts per
|
|
run. No database migration is needed.
|
|
|
|
Workspace uploads use the existing workspace resource containment checks,
|
|
no-symlink file opens covering every path component, and bounded descriptor reads.
|
|
Local uploads require Linux or macOS; authorized artifacts work on other hosts.
|
|
Lifecycle-sensitive endpoints require an inline JSON object, so a raw uploaded
|
|
JSON file cannot hide protected fields from policy checks. Artifacts must belong
|
|
to the bound company. Secret-value access, credential management, secret proposals
|
|
and company exports require their existing secure clients. Search describes these
|
|
operations as restricted. `call_api` rejects them before creating a replay receipt
|
|
or making an HTTP request. Safe secret metadata listing remains available.
|
|
Agent credentials are never returned to the model. Streaming, WebSocket, MCP and authentication
|
|
handshakes are documented as protocol operations requiring their existing clients.
|
|
|
|
## Catalog maintenance
|
|
|
|
`runner-api-catalog.ts` builds from the server OpenAPI registry. Experimental
|
|
pipeline, Cases and smoke-lab routes now share their validators with discovery.
|
|
Seven Cases/pipeline route shapes are multiplexed by resource identity: the Cases
|
|
router intentionally forwards unknown resources to the pipeline router. Their
|
|
separate catalog entries explain which resource identifier is required. Registry
|
|
authorization descriptions are documentation; actual route checks are authoritative.
|
|
|
|
Regenerate old-skill enrichment after editing its API reference:
|
|
|
|
```sh
|
|
node scripts/generate-runner-api-reference.mjs
|
|
node scripts/generate-runner-api-reference.mjs --check
|
|
node scripts/generate-runner-experimental-api-metadata.mjs
|
|
node scripts/generate-runner-experimental-api-metadata.mjs --check
|
|
```
|
|
|
|
Mounted-route coverage tests include experimental routes. Three WebSocket mounts
|
|
are explicitly classified in the catalog. Shared protocol-action catalogs,
|
|
provider projections and generated compatibility checks include both tools.
|
|
|
|
## Verification and paid evals
|
|
|
|
The companion `paperclip-evals` worktree contains `evals/runner-api-tools`.
|
|
Its README documents explicit case/model selectors, the cumulative budget ledger,
|
|
fixture reset, progressive batches, and Evalbook generation. No command defaults
|
|
to running the entire paid suite. Capability, forced operation contracts and
|
|
paired common-operation regressions are reported separately.
|
|
|
|
Provider-free integration tests exercise real runnerd → PRP → authority → HTTP,
|
|
route validation and audit, stale bindings, Ask/Plan restrictions, identity
|
|
spoofing, file containment, uncertain mutation receipts and fixture isolation.
|
|
|
|
The Evalbook viewer uses the existing shared viewer and stylesheet on master.
|
|
The report retains actual persisted-state summaries for private local inspection;
|
|
public replay continues to withhold company-state details.
|
|
|
|
The ACPX sidecar includes the upstream terminal-usage accounting correction from
|
|
`origin/codex/evalbook-default-chat-sept6`. Its qualified Claude executable requires
|
|
Linux x64. The first macOS stage records a zero-cost ACPX admission failure. A later user-authorized
|
|
OpenCode/OpenRouter Sonnet profile reached a real HTTP read, but the attempt failed
|
|
on a missing harness completion contract and incomplete terminal accounting. The
|
|
harness contract is corrected. The missing fourth request was subsequently
|
|
recovered from the matching OpenRouter session and generation billing record;
|
|
the original failed attempt remains immutable. New attempts retain an append-only,
|
|
flushed event journal and bounded provider trace outside disposable runtime files.
|
|
Provider-free startup succeeds for OpenRouter Sonnet and DeepSeek. See
|
|
`doc/plans/2026-09-07-runner-api-production-readiness.md` for remaining release gates.
|