mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.
## Linked Issues or Issue Description
**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.
**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.
**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.
Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.
## What Changed
- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.
## Verification
- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.
## Risks
- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.
## Model Used
OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
248 lines
15 KiB
Markdown
248 lines
15 KiB
Markdown
# Runner API escape hatch
|
||
|
||
`search_api` and `call_api` extend the native runner when an available dedicated
|
||
operation cannot express the requested work. Existing tools remain preferred;
|
||
agents do not have to search before using them. Only two tool definitions are
|
||
advertised. The API catalog is returned on demand, never injected into the
|
||
initial prompt.
|
||
|
||
## Default availability and operator controls
|
||
|
||
The escape hatch is enabled by default. No environment variable is required.
|
||
The native `hire_agent` tool shares this availability policy.
|
||
Set `PAPERCLIP_RUNNER_API_TOOLS_ENABLED=false` on the server to disable these
|
||
tools. An explicit `true` also enables them; other explicit values fail closed.
|
||
|
||
Operators can restrict availability with `PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS`,
|
||
a comma-separated list of company UUIDs. An unset list allows every company;
|
||
an explicitly empty list allows none. IDs must match exactly. This restriction
|
||
also applies when the enabled flag is unset.
|
||
|
||
A server-owned binding can disable these tools for a baseline eval but cannot
|
||
override an operator restriction. The server checks the policy when advertising
|
||
tools, when accepting a call, and immediately before HTTP dispatch after
|
||
preparing any files. Existing dedicated tools remain available. Operators must
|
||
update the environment of each server process and restart it for deployment-level
|
||
changes; this environment switch is not a live settings API.
|
||
|
||
All company, run, work-mode, credential, and lifecycle checks below still apply.
|
||
|
||
## Discovery and requests
|
||
|
||
```json
|
||
{"query":"create project","limit":5}
|
||
```
|
||
|
||
Search is deterministic lexical ranking over OpenAPI paths, summaries and the
|
||
old skill reference. It supports task/issue and other terminology, exact
|
||
`METHOD /api/path/{parameter}` lookup, and opaque query/catalog-bound pagination.
|
||
Results include resolved request schemas, response descriptions, authorization
|
||
metadata, work modes, examples where available, and relevant dedicated tools
|
||
with their supported parameters. `limit` defaults to five and is capped at ten.
|
||
|
||
```json
|
||
{"operationId":"PATCH /api/projects/{id}","pathParams":{"id":"PROJECT_UUID"},"body":{"description":"Updated project description"}}
|
||
```
|
||
|
||
The catalog determines method and path. `companyId` is filled from the active
|
||
binding. Scalars and arrays are accepted in `query`. `body` defaults to JSON;
|
||
`contentType` supports text and raw uploads. `files` accepts entries containing
|
||
exactly one authorized `artifactId` or task-workspace `path`, and an optional
|
||
multipart `field`. No arbitrary URL, headers, authentication, or remote file URL
|
||
can be supplied. Routes still validate payloads and enforce permissions.
|
||
|
||
Requests have a 16 KiB URL limit and 10 MiB request/upload limit. **New response
|
||
captures are limited to 1 GiB of decoded bytes.** This is separate from the 24 KiB
|
||
inline/page limit. The receiver rejects an oversized Content-Length before
|
||
reading and counts actual bytes before writing, including chunked or compressed
|
||
responses. Oversized responses return `api_response_too_large`; narrow the query
|
||
or use the endpoint's own pagination. A mutation may already have committed, so
|
||
inspect its state rather than retrying it to obtain a smaller response.
|
||
|
||
Connection setup and stalled response reads time out after 30 seconds. An active
|
||
capture has a 10-minute total download deadline. Responses above 24 KiB stream
|
||
into a private temporary file, then into a company-owned asset. Memory stays
|
||
bounded by the inline prefix and stream buffers. Binary responses also become
|
||
assets; text previews are limited to 2,000 bytes.
|
||
|
||
Large captures reserve 1 GiB against a **durable 4 GiB per-run capture budget**
|
||
before creating a file. A completed capture settles to its actual byte count;
|
||
failed/interrupted captures retain the full reservation to bound retry loops.
|
||
A process restart does not reset that budget. Small inline responses and reads
|
||
of existing assets need no reservation. Concurrent large captures are limited
|
||
to **two per company and four per server process**, including the storage upload
|
||
and temporary-file cleanup. Budget exhaustion returns `api_response_capture_limit`;
|
||
concurrency exhaustion returns `api_response_capture_busy` without queuing more
|
||
large transfers. A **20 GiB company-wide quota** counts all stored `runner-api` snapshots plus
|
||
unattached reservations. Admission uses a company database lock, so runs and
|
||
server processes share the same quota. Existing snapshots from before this
|
||
change count too. Attaching an asset converts its reservation to actual stored
|
||
bytes; deleting the asset frees that capacity. Small binary snapshots also need
|
||
storage admission. Ordinary inline text/JSON and existing-asset pages do not.
|
||
|
||
Operators can set `PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES` to a positive
|
||
safe integer of at least 1 GiB. Invalid values fall back to 20 GiB. No zero or
|
||
unlimited setting is accepted. This quota covers API snapshots, not all company
|
||
attachments. Storage capacity and backend limits still apply.
|
||
|
||
Handled pre-storage failures release the company reservation after temporary
|
||
file cleanup, but retain the run's charge against retry loops. After a metadata
|
||
transaction fails, a locking read must prove the asset was not committed before
|
||
the uploaded object is removed. Confirmed cleanup refunds the company quota. A crash, failed
|
||
cleanup, or ambiguous storage write keeps an unattached reservation. Operators
|
||
must reconcile possible orphan files/objects before deleting that reservation
|
||
from `runner_api_response_reservations`; no automatic expiry silently refunds
|
||
possibly occupied space. Linked rows are removed with their asset. Deleting a
|
||
run does not release its unattached company reservations.
|
||
|
||
Temporary files are removed on success or failure. Long captures revalidate the
|
||
active run at least every MiB or at the next chunk after one second, and again
|
||
before returning the snapshot. Stopping the run stops its download.
|
||
|
||
To inspect saved text without creating another artifact, call its authorized
|
||
content operation with `responseText`:
|
||
|
||
```json
|
||
{"operationId":"GET /api/assets/{assetId}/content","pathParams":{"assetId":"RETURNED_ARTIFACT_ID"},"responseText":{"offsetBytes":0,"limitBytes":8192}}
|
||
```
|
||
|
||
The result contains `data` as text (including JSON), plus `responseText` with
|
||
`offsetBytes`, `nextOffsetBytes`, and `totalBytes`. Continue at `nextOffsetBytes`
|
||
until it is null. Windows end at UTF-8 boundaries. The limit defaults to 24 KiB
|
||
and accepts 4–24,576 bytes. The server rejects binary content, invalid UTF-8,
|
||
and offsets inside a code point or beyond the response. This option only works
|
||
with GET; never repeat a mutation to retrieve another part of its response.
|
||
Read the saved artifact for a stable snapshot instead of paging a changing live
|
||
response. A text-window call against a live response above 24 KiB also returns
|
||
that snapshot's artifact reference; continue on its content operation. A saved
|
||
asset page never creates another asset. An unpaged asset read also fetches only
|
||
a bounded preview and returns the existing reference instead of copying the file. Offsets and total sizes use safe integer
|
||
byte counts, including values above 2 GiB. Existing assets larger than the 1 GiB
|
||
capture limit remain readable because each request transfers only a bounded range.
|
||
Request/upload limits and all route
|
||
authorization remain.
|
||
Saved asset pages use authenticated HTTP byte ranges. The storage provider reads
|
||
only the requested window, with at most two extra bytes for UTF-8/EOF handling.
|
||
The client validates `Content-Range`, the total size, and the received byte count;
|
||
it rejects unsupported or inconsistent ranges instead of downloading the whole
|
||
asset for every page. Other GET routes are fetched once in full to create the
|
||
snapshot, so use the returned asset for subsequent pages. S3 uses streamed
|
||
multipart uploads for large snapshots. Asset sizes are stored as PostgreSQL
|
||
`bigint`, preserving the existing numeric API shape.
|
||
|
||
### Media files and upload limits
|
||
|
||
The 1 GiB limit covers new snapshots returned by `call_api`, such as large JSON
|
||
exports or binary API downloads. It does not raise attachment upload limits or
|
||
limit files an agent creates and edits inside its workspace. Saved asset downloads
|
||
stream from storage and support byte ranges, including video seeking.
|
||
|
||
`PAPERCLIP_ATTACHMENT_MAX_BYTES` separately defaults to 10 MiB for uploads and
|
||
native file handoffs. `call_api` uploads also have their own 10 MiB limit. Several
|
||
upload and handoff paths buffer complete files in memory; raising those defaults
|
||
to GiB sizes requires streaming ingestion and corresponding admission/budget
|
||
controls first. For a future video attachment workflow, 2 GiB per streamed file
|
||
is a reasonable default, with an operator override and storage quotas. Do not
|
||
claim that this response-paging change enables GiB attachment uploads.
|
||
|
||
Tool responses identify the HTTP route with `apiOperationId`. The native protocol
|
||
reserves `operationId` and `callId` for semantic tool-call identity; API metadata
|
||
must not masquerade as that envelope. Saved mutation receipts are normalized at
|
||
the tool boundary as well, without repeating their HTTP request.
|
||
All redirects are refused. Interrupted mutation responses have an
|
||
unknown outcome, requiring inspection before another mutation.
|
||
Mutation responses with HTTP 5xx, HTTP 408, redirects, or malformed JSON also
|
||
retain an unknown outcome. A server may have committed the write before it
|
||
failed to return a valid response.
|
||
|
||
## Authority and replay
|
||
|
||
The server revalidates the active native run, assigned task and actor, then
|
||
creates a server-held agent JWT bound to that company and run. Requests go
|
||
through the actual HTTP router with its authorization, validation and domain
|
||
audit behavior. An additional `runner.api_called` receipt attributes mutations
|
||
to the run even where older route audit events omit that field.
|
||
The run and work mode are checked again after asynchronous file preparation, so
|
||
a stopped run cannot dispatch an upload prepared under its earlier binding.
|
||
|
||
Ask and pre-acceptance Plan permit reads through the escape hatch. Existing
|
||
dedicated-tool exceptions are unchanged. Runner-owned checkout, completion,
|
||
status/assignment transitions, approval decisions and execution-control actions
|
||
cannot be bypassed through generic calls. Routine creation, schedule/trigger
|
||
changes and manual/public routine execution require the existing scheduling
|
||
clients. Direct workspace runtime commands, runtime-slot stop/restart, case
|
||
automation retries and skill test-run controls also require their existing
|
||
execution clients. Gateway session credentials cannot enter generic results.
|
||
Routine metadata remains readable; annotation threads, comments and thread
|
||
resolution remain available through the fallback. API-only ordinary fields, such as a
|
||
task's `billingCode`, remain accessible even when a dedicated tool covers other
|
||
fields on that endpoint.
|
||
|
||
Mutation call IDs reserve a durable receipt in the run's existing `resultJson`
|
||
before dispatch. Replays return the recorded result. Reusing an ID with different
|
||
arguments is rejected. A crash after reservation leaves an unknown outcome and
|
||
never automatically resends the mutation. The limit is 512 mutation receipts per
|
||
run. No database migration is needed.
|
||
|
||
Workspace uploads use the existing workspace resource containment checks,
|
||
no-symlink file opens covering every path component, and bounded descriptor reads.
|
||
Local uploads require Linux or macOS; authorized artifacts work on other hosts.
|
||
Lifecycle-sensitive endpoints require an inline JSON object, so a raw uploaded
|
||
JSON file cannot hide protected fields from policy checks. Artifacts must belong
|
||
to the bound company. Secret-value access, credential management, secret proposals
|
||
and company exports require their existing secure clients. Search describes these
|
||
operations as restricted. `call_api` rejects them before creating a replay receipt
|
||
or making an HTTP request. Safe secret metadata listing remains available.
|
||
Agent credentials are never returned to the model. Streaming, WebSocket, MCP and authentication
|
||
handshakes are documented as protocol operations requiring their existing clients.
|
||
|
||
## Catalog maintenance
|
||
|
||
`runner-api-catalog.ts` builds from the server OpenAPI registry. Experimental
|
||
pipeline, Cases and smoke-lab routes now share their validators with discovery.
|
||
Seven Cases/pipeline route shapes are multiplexed by resource identity: the Cases
|
||
router intentionally forwards unknown resources to the pipeline router. Their
|
||
separate catalog entries explain which resource identifier is required. Registry
|
||
authorization descriptions are documentation; actual route checks are authoritative.
|
||
|
||
Regenerate old-skill enrichment after editing its API reference:
|
||
|
||
```sh
|
||
node scripts/generate-runner-api-reference.mjs
|
||
node scripts/generate-runner-api-reference.mjs --check
|
||
node scripts/generate-runner-experimental-api-metadata.mjs
|
||
node scripts/generate-runner-experimental-api-metadata.mjs --check
|
||
```
|
||
|
||
Mounted-route coverage tests include experimental routes. Three WebSocket mounts
|
||
are explicitly classified in the catalog. Shared protocol-action catalogs,
|
||
provider projections and generated compatibility checks include both tools.
|
||
|
||
## Verification and paid evals
|
||
|
||
The companion `paperclip-evals` worktree contains `evals/runner-api-tools`.
|
||
Its README documents explicit case/model selectors, the cumulative budget ledger,
|
||
fixture reset, progressive batches, and Evalbook generation. No command defaults
|
||
to running the entire paid suite. Capability, forced operation contracts and
|
||
paired common-operation regressions are reported separately.
|
||
|
||
Provider-free integration tests exercise real runnerd → PRP → authority → HTTP,
|
||
route validation and audit, stale bindings, Ask/Plan restrictions, identity
|
||
spoofing, file containment, uncertain mutation receipts and fixture isolation.
|
||
|
||
The Evalbook viewer uses the existing shared viewer and stylesheet on master.
|
||
The report retains actual persisted-state summaries for private local inspection;
|
||
public replay continues to withhold company-state details.
|
||
|
||
The ACPX sidecar includes the upstream terminal-usage accounting correction from
|
||
`origin/codex/evalbook-default-chat-sept6`. Its qualified Claude executable requires
|
||
Linux x64. The first macOS stage records a zero-cost ACPX admission failure. A later user-authorized
|
||
OpenCode/OpenRouter Sonnet profile reached a real HTTP read, but the attempt failed
|
||
on a missing harness completion contract and incomplete terminal accounting. The
|
||
harness contract is corrected. The missing fourth request was subsequently
|
||
recovered from the matching OpenRouter session and generation billing record;
|
||
the original failed attempt remains immutable. New attempts retain an append-only,
|
||
flushed event journal and bounded provider trace outside disposable runtime files.
|
||
Provider-free startup succeeds for OpenRouter Sonnet and DeepSeek. See
|
||
`doc/plans/2026-09-07-runner-api-production-readiness.md` for remaining release gates.
|