Files
PaperClipAI/tests/runner-e2e/api-response-reading.ts
DottaandPaperclip 3447609d22 fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.

## Linked Issues or Issue Description

**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.

**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.

**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.

Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.

## What Changed

- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.

## Risks

- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 10:36:32 -05:00

135 lines
8.6 KiB
TypeScript

import { createHash } from "node:crypto";
import type { RunnerTaskFixture } from "./types.js";
import type { RunnerApi } from "./api.js";
import type { IssueAttachment } from "../../packages/shared/src/types/issue.js";
/** Verify delivered bytes through the same public download path the user receives. */
export async function readResponseProof(api: RunnerApi, issueId: string, runId: string) {
const attachments = await api.get<IssueAttachment[]>(`/api/issues/${issueId}/attachments`);
const proofs = attachments.filter(attachment =>
attachment.issueId === issueId && attachment.originatingRunId === runId &&
attachment.originalFilename === "api-response-proof.txt");
if (proofs.length !== 1) throw new Error(`Expected one proof attachment from the tested run; observed ${proofs.length}`);
const proof = proofs[0]!;
const response = await api.request.get(`/api/attachments/${encodeURIComponent(proof.id)}/content?download=1`);
if (!response.ok()) throw new Error(`Proof download returned ${response.status()}`);
const bytes = await response.body();
if (bytes.length !== proof.byteSize || createHash("sha256").update(bytes).digest("hex") !== proof.sha256) {
throw new Error("Downloaded proof disagrees with stored attachment bytes");
}
return { attachmentId: proof.id, content: bytes.toString("utf8"), sha256: proof.sha256 };
}
export function responseEvidenceCode(nonce: string) {
return createHash("sha256").update(`bounded-response-evidence:${nonce}`).digest("hex");
}
export function responseEvidenceDescription(nonce: string) {
return `${"Synthetic diagnostic padding.\n".repeat(1400)}\nEvidence code: ${responseEvidenceCode(nonce)}\n`;
}
/** Correlate trusted PRP call inputs, results, and completion receipts. */
export function gradeApiResponsePaging(events: readonly { eventType?: string; payload?: unknown }[], sourceIssueId: string) {
const object = (value: unknown): Record<string, unknown> =>
value !== null && typeof value === "object" && !Array.isArray(value) ? value as Record<string, unknown> : {};
const inputs = new Map<string, Record<string, unknown>>();
const results = new Map<string, Record<string, unknown>>();
const completed = new Set<string>();
for (const event of events) {
const payload = object(object(object(event.payload).prpEvent).payload);
const item = object(payload.item);
if (event.eventType === "item.started" && item.type === "tool_use" && item.name === "call_api" && typeof item.id === "string") {
inputs.set(item.id, object(item.input));
}
if (event.eventType === "item.completed" && item.type === "tool_result" && typeof item.id === "string" &&
(item.tool_use_id === undefined || item.tool_use_id === item.id)) results.set(item.id, object(item.result));
if (event.eventType === "tool.execution.completed" && payload.name === "call_api" && payload.status === "completed" && typeof payload.executionId === "string") {
completed.add(payload.executionId);
}
}
const calls = [...inputs].flatMap(([id, input]) => {
const result = results.get(id);
return completed.has(id) && result?.ok === true && [200, 206].includes(Number(result.status)) ? [{ input, result }] : [];
});
const sourceOperation = "GET /api/issues/{id}";
const assetOperation = "GET /api/assets/{assetId}/content";
const sources = calls.filter(call => call.input.operationId === sourceOperation && call.result.apiOperationId === sourceOperation &&
object(call.input.pathParams).id === sourceIssueId);
let failure = "No successful call_api source read with a saved response artifact";
for (const source of sources) {
const artifact = object(source.result.artifact);
const assetId = artifact.artifactId;
const total = artifact.byteSize;
if (typeof assetId !== "string" || typeof total !== "number" || !Number.isSafeInteger(total) || total <= 24 * 1024 || total > 1024 * 1024 ||
typeof artifact.sha256 !== "string" || !/^[a-f0-9]{64}$/.test(artifact.sha256) || artifact.url !== `/api/assets/${assetId}/content`) continue;
const pages = new Map<number, { bytes: Buffer; next: number | null }>();
let malformed = false;
for (const call of calls) {
if (call.input.operationId !== assetOperation || object(call.input.pathParams).assetId !== assetId) continue;
const requested = object(call.input.responseText);
const returned = object(call.result.responseText);
const offset = requested.offsetBytes;
const limit = requested.limitBytes;
const next = returned.nextOffsetBytes;
const data = call.result.data;
if (call.result.apiOperationId !== assetOperation || typeof offset !== "number" || !Number.isSafeInteger(offset) || offset < 0 || offset >= total ||
typeof limit !== "number" || !Number.isSafeInteger(limit) || limit < 1 || limit > 8192 || returned.offsetBytes !== offset || returned.totalBytes !== total ||
typeof data !== "string") { malformed = true; break; }
const bytes = Buffer.from(data, "utf8");
const end = offset + bytes.length;
if (bytes.length === 0 || bytes.length > limit || end > total || (end === total ? next !== null : next !== end)) {
malformed = true; break;
}
const prior = pages.get(offset);
if (prior && (!prior.bytes.equals(bytes) || prior.next !== next)) { malformed = true; break; }
pages.set(offset, { bytes, next: next as number | null });
}
if (malformed) { failure = "Artifact paging request, returned offsets, bytes, or EOF disagree"; continue; }
const chunks: Buffer[] = [];
const offsets: number[] = [];
let offset = 0;
let eofReached = false;
while (pages.has(offset)) {
const page = pages.get(offset)!;
chunks.push(page.bytes); offsets.push(offset);
if (page.next === null) { eofReached = true; break; }
offset = page.next;
}
if (!eofReached || offsets.length < 2) { failure = "Saved artifact pages do not cover offset zero through EOF without gaps"; continue; }
const bytes = Buffer.concat(chunks);
const digest = createHash("sha256").update(bytes).digest("hex");
if (bytes.length !== total || digest !== artifact.sha256) { failure = "Paged bytes disagree with the source response artifact digest"; continue; }
try {
if (object(JSON.parse(bytes.toString("utf8"))).id !== sourceIssueId) { failure = "Paged artifact belongs to another source issue"; continue; }
} catch { failure = "Paged source response is not complete JSON"; continue; }
return { passed: true, sourceIssueId, artifactId: assetId, totalBytes: total, pageCount: offsets.length, offsets, eofReached, sha256: digest };
}
return { passed: false, sourceIssueId, failure };
}
export const apiResponseReadingTask: RunnerTaskFixture = {
id: "saved-text-pages", label: "Read saved large API evidence", groups: [],
workMode: "standard", flow: "single_turn", expectedRunCount: 1,
attemptTimeoutMs: { local: 10 * 60_000, daytona: 10 * 60_000 },
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: nonce => `Read large API evidence ${nonce}`,
// Delivered files render as an attachment card, which can replace the finish summary.
buildVisibleMarker: () => "api-response-proof.txt",
buildPrompt: nonce => [
"Inspect diagnostic evidence task {{API_RESPONSE_SOURCE_ID}} using the Paperclip API tools.",
"Discover GET /api/issues/{id}, call it for that task, and retain the returned response artifact.",
"Read the saved artifact through GET /api/assets/{assetId}/content with responseText: {offsetBytes:0,limitBytes:8192}.",
"Continue using responseText.nextOffsetBytes until null. Extract the Evidence code at the end of its description.",
"The large response must be read with bounded responseText pages. Do not use other tools or API routes to obtain the evidence; if bounded reading fails, report the failure instead of substituting a different reader.",
"Write only that code followed by a newline to api-response-proof.txt in the current workspace and deliver that file as an attachment. Do not edit the evidence task or create child tasks.",
`Verify the proof file and finish this task with paperclip_finish, reportedWorkDisposition done, and summary API_RESPONSE_READ_${nonce}.`,
].join("\n"),
buildMatchers: (nonce, execution) => [
{ kind: "file_exact", path: "api-response-proof.txt", expected: `${responseEvidenceCode(nonce)}\n` },
{ kind: "message_contains", expected: "api-response-proof.txt" },
{ kind: "issue_status", expected: "done" },
{ kind: "run_status", expected: "succeeded" },
{ kind: "runtime_mode", expected: execution.profile.expectedRuntimeMode },
{ kind: "environment", expected: execution.environment.id },
],
};