Files
PaperClipAI/tests/runner-e2e/api-response-reading.test.ts
T
DottaandPaperclip 3447609d22 fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.

## Linked Issues or Issue Description

**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.

**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.

**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.

Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.

## What Changed

- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.

## Risks

- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 10:36:32 -05:00

103 lines
6.8 KiB
TypeScript

import { createHash } from "node:crypto";
import { describe, expect, it, vi } from "vitest";
import type { RunnerApi } from "./api.js";
import { apiResponseReadingTask, gradeApiResponsePaging, readResponseProof, responseEvidenceCode, responseEvidenceDescription } from "./api-response-reading.js";
describe("large API response evidence fixture", () => {
it("keeps evidence beyond both preview and inline limits and out of the assignment", () => {
const source = responseEvidenceDescription("fixture");
const code = responseEvidenceCode("fixture");
expect(Buffer.byteLength(source)).toBeGreaterThan(24 * 1024);
expect(source.indexOf(code)).toBeGreaterThan(24 * 1024);
expect(apiResponseReadingTask.buildPrompt("fixture")).not.toContain(code);
expect(apiResponseReadingTask.buildPrompt("fixture")).toContain("responseText.nextOffsetBytes");
});
it("downloads exact run-attributed proof bytes rather than trusting the agent's report", async () => {
const bytes = Buffer.from("independent evidence\n");
const proof = { id: "proof", issueId: "issue", originatingRunId: "run", originalFilename: "api-response-proof.txt", byteSize: bytes.length, sha256: createHash("sha256").update(bytes).digest("hex") };
const download = vi.fn().mockResolvedValue({ ok: () => true, body: async () => bytes });
const get = vi.fn().mockResolvedValue([proof]);
const api = { get, request: { get: download } } as unknown as RunnerApi;
expect((await readResponseProof(api, "issue", "run")).content).toBe(bytes.toString());
expect(download).toHaveBeenCalledWith("/api/attachments/proof/content?download=1");
get.mockResolvedValue([{ ...proof, originatingRunId: "other" }]);
await expect(readResponseProof(api, "issue", "run")).rejects.toThrow("observed 0");
get.mockResolvedValue([proof, proof]);
await expect(readResponseProof(api, "issue", "run")).rejects.toThrow("observed 2");
get.mockResolvedValue([{ ...proof, sha256: "mismatch" }]);
await expect(readResponseProof(api, "issue", "run")).rejects.toThrow("disagrees");
});
});
describe("saved API response pagination oracle", () => {
const sourceIssueId = "source-issue";
const assetId = "source-asset";
const bytes = Buffer.from(JSON.stringify({ id: sourceIssueId, description: "synthetic padding ".repeat(2000) }));
const operation = "GET /api/assets/{assetId}/content";
const event = (eventType: string, payload: unknown) => ({ eventType, payload: { prpEvent: { payload } } });
const call = (id: string, input: unknown, result: unknown, status = "completed") => [
event("item.started", { item: { id, type: "tool_use", name: "call_api", input } }),
event("item.completed", { item: { id, type: "tool_result", tool_use_id: id, result } }),
event("tool.execution.completed", { name: "call_api", executionId: id, status }),
];
const source = () => call("source-read", { operationId: "GET /api/issues/{id}", pathParams: { id: sourceIssueId } }, {
ok: true, status: 200, apiOperationId: "GET /api/issues/{id}", artifact: {
artifactId: assetId, byteSize: bytes.length, sha256: createHash("sha256").update(bytes).digest("hex"),
url: `/api/assets/${assetId}/content`,
},
});
const page = (offset: number, id = `page-${offset}`, artifactId = assetId) => {
const end = Math.min(offset + 8192, bytes.length);
return call(id, { operationId: operation, pathParams: { assetId: artifactId }, responseText: { offsetBytes: offset, limitBytes: 8192 } }, {
ok: true, status: 200, apiOperationId: operation, data: bytes.subarray(offset, end).toString(),
responseText: { offsetBytes: offset, totalBytes: bytes.length, nextOffsetBytes: end === bytes.length ? null : end },
});
};
const pages = () => Array.from({ length: Math.ceil(bytes.length / 8192) }, (_, i) => page(i * 8192)).flat();
it("accepts correlated source artifact pages through EOF, including a harmless identical retry", () => {
expect(gradeApiResponsePaging([...source(), ...page(0, "retry"), ...pages()], sourceIssueId).passed).toBe(true);
});
it("rejects unrelated completed calls", () => {
const unrelated = [1, 2, 3].flatMap(i => call(`unrelated-${i}`, { operationId: "GET /api/agents/me" }, { ok: true, status: 200 }));
expect(gradeApiResponsePaging(unrelated, sourceIssueId).passed).toBe(false);
});
it("rejects repeating the first page without reaching EOF", () => {
expect(gradeApiResponsePaging([...source(), ...page(0), ...page(0, "again"), ...page(0, "again-2")], sourceIssueId).passed).toBe(false);
});
it("rejects a gap even when the last page claims EOF", () => {
expect(gradeApiResponsePaging([...source(), ...pages().filter((_e, i) => i < 3 || i >= 6)], sourceIssueId).passed).toBe(false);
});
it("rejects reading a different artifact", () => {
const other = Array.from({ length: Math.ceil(bytes.length / 8192) }, (_, i) => page(i * 8192, `other-${i}`, "other-asset")).flat();
expect(gradeApiResponsePaging([...source(), ...other], sourceIssueId).passed).toBe(false);
});
it("requires the exact seeded source issue", () => {
expect(gradeApiResponsePaging([...source(), ...pages()], "other-issue").passed).toBe(false);
});
it("requires a matched successful completion for every page", () => {
const incomplete = pages().filter(e => e.eventType !== "tool.execution.completed");
expect(gradeApiResponsePaging([...source(), ...incomplete], sourceIssueId).passed).toBe(false);
});
it.each([
["requested offset", { offsetBytes: 1, totalBytes: bytes.length, nextOffsetBytes: 8192 }],
["next offset", { offsetBytes: 0, totalBytes: bytes.length, nextOffsetBytes: 8193 }],
["total size", { offsetBytes: 0, totalBytes: bytes.length + 1, nextOffsetBytes: 8192 }],
["premature EOF", { offsetBytes: 0, totalBytes: bytes.length, nextOffsetBytes: null }],
])("rejects mismatched %s metadata", (_label, responseText) => {
const malformed = call("bad-page", {
operationId: operation, pathParams: { assetId }, responseText: { offsetBytes: 0, limitBytes: 8192 },
}, { ok: true, status: 206, apiOperationId: operation, data: bytes.subarray(0, 8192).toString(), responseText });
expect(gradeApiResponsePaging([...source(), ...pages(), ...malformed], sourceIssueId).passed).toBe(false);
});
it("rejects contradictory duplicate bytes and a wrong source artifact digest", () => {
const corrupt = call("corrupt-page", {
operationId: operation, pathParams: { assetId }, responseText: { offsetBytes: 0, limitBytes: 8192 },
}, {
ok: true, status: 206, apiOperationId: operation, data: "x".repeat(8192),
responseText: { offsetBytes: 0, totalBytes: bytes.length, nextOffsetBytes: 8192 },
});
expect(gradeApiResponsePaging([...source(), ...pages(), ...corrupt], sourceIssueId).passed).toBe(false);
expect(gradeApiResponsePaging([...source(), ...corrupt, ...pages().slice(3)], sourceIssueId).passed).toBe(false);
});
});