mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 16:35:27 +02:00
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.
## Linked Issues or Issue Description
**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.
**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.
**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.
Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.
## What Changed
- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.
## Verification
- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.
## Risks
- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.
## Model Used
OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
632 lines
24 KiB
TypeScript
632 lines
24 KiB
TypeScript
import { describe, expect, it } from "vitest";
|
|
import { normalizePrpResultSignals } from "../../packages/paperclip-runner/src/protocol/result-normalization.js";
|
|
import {
|
|
connectionReviewSuite,
|
|
runnerEnvironments,
|
|
runnerMatrix,
|
|
openRouterBreadthExcludedExecutionIds,
|
|
openRouterBreadthExcludedModelIds,
|
|
openRouterBreadthProfiles,
|
|
openRouterBreadthTasks,
|
|
localIntegrityTasks,
|
|
runnerProfiles,
|
|
runnerSuites,
|
|
runnerTasks,
|
|
daytonaWarmContinuityTask,
|
|
daytonaWarmEnvironment,
|
|
isImmutableDaytonaImage,
|
|
suiteDefinitionHash,
|
|
validateRunnerCatalog,
|
|
} from "./catalog.js";
|
|
import {
|
|
buildMatrixJobs,
|
|
parseRunnerSelectors,
|
|
RunnerSelectorError,
|
|
selectRunnerExecutions,
|
|
} from "./selectors.js";
|
|
|
|
describe("runner E2E catalog", () => {
|
|
it("supplies an actionable human review in native warm completion examples", () => {
|
|
const prompts = [daytonaWarmContinuityTask.buildPrompt("nonce"), ...daytonaWarmContinuityTask.buildFollowupMessages!("nonce")];
|
|
for (const [index, prompt] of prompts.entries()) {
|
|
const match = prompt.match(/attentionRequests:(\[.*?\]),evidence:/);
|
|
expect(match).not.toBeNull();
|
|
const signals = normalizePrpResultSignals({ attentionRequests: JSON.parse(match![1]) });
|
|
expect(signals.ignoredAttentionRequests).toEqual([]);
|
|
expect(signals.actionableAttentionRequests).toHaveLength(index === 2 ? 0 : 1);
|
|
if (index < 2) expect(signals.actionableAttentionRequests[0]).toMatchObject({kind:"review", ownerClass:"human"});
|
|
expect(prompt).not.toContain("call request_human_input");
|
|
}
|
|
});
|
|
|
|
it("defines sixteen local connection-review journeys without expanding the default matrix", () => {
|
|
expect(connectionReviewSuite.expectedMatrixSize).toBe(16);
|
|
expect(new Set(connectionReviewSuite.profiles.map(profile => profile.id))).toEqual(new Set(["runner-codex", "runner-acpx-claude", "legacy-codex", "legacy-claude"]));
|
|
expect(connectionReviewSuite.environments.map(environment => environment.id)).toEqual(["local"]);
|
|
expect(connectionReviewSuite.tasks.map(task => task.toolReviewDecision)).toEqual(["approve", "decline", "always", "restart"]);
|
|
expect(connectionReviewSuite.tasks.every(task => task.flow === "governed_tool_review")).toBe(true);
|
|
});
|
|
|
|
it("tests native chat plans, tasks, and reassignment with production permission defaults", () => {
|
|
const suite = runnerSuites.find(suite => suite.id === "agent-chat")!;
|
|
for (const id of ["runner-codex", "runner-acpx-claude"]) {
|
|
const profile = suite.profiles.find(profile => profile.id === id)!;
|
|
const payload = profile.buildAgent({
|
|
executionId: "default-permissions", workspacePath: "/workspace", environmentId: "env-1", environmentFixtureId: "local",
|
|
secretRefs: {
|
|
[profile.credential]: { type: "secret_ref", secretId: "22222222-2222-4222-8222-222222222222", version: "latest" },
|
|
},
|
|
});
|
|
expect(payload.adapterConfig).not.toHaveProperty("acpxPermissionMode");
|
|
expect(payload.adapterConfig).not.toHaveProperty("codexPermissionMode");
|
|
expect(suite.tasks.map(task => task.id)).toEqual(expect.arrayContaining(["plan-handoff", "reassign-task", "create-backlog"]));
|
|
}
|
|
});
|
|
|
|
it("validates the core, local-integrity, breadth, and warm suites", () => {
|
|
expect(runnerProfiles).toHaveLength(7);
|
|
expect(openRouterBreadthProfiles).toHaveLength(4);
|
|
expect(runnerEnvironments).toHaveLength(2);
|
|
expect(runnerTasks).toHaveLength(3);
|
|
expect(localIntegrityTasks).toHaveLength(2);
|
|
expect(openRouterBreadthTasks).toHaveLength(3);
|
|
expect(runnerSuites.map((suite) => suite.expectedMatrixSize)).toEqual([
|
|
2, 8, 46, 23, 47, 52, 28, 18, 6, 6, 4, 42, 14, 10, 2,
|
|
]);
|
|
expect(validateRunnerCatalog()).toHaveLength(308);
|
|
expect(new Set(runnerMatrix.map((entry) => entry.id)).size).toBe(308);
|
|
expect(
|
|
runnerMatrix.filter((entry) => entry.suite.id === "core-compatibility"),
|
|
).toHaveLength(42);
|
|
expect(
|
|
runnerMatrix.filter(
|
|
(entry) => entry.suite.id === "local-session-integrity",
|
|
),
|
|
).toHaveLength(14);
|
|
expect(
|
|
runnerMatrix.filter(
|
|
(entry) => entry.suite.id === "openrouter-model-breadth",
|
|
),
|
|
).toHaveLength(10);
|
|
expect(
|
|
runnerMatrix.filter(
|
|
(entry) => entry.suite.id === "daytona-warm-continuity",
|
|
),
|
|
).toHaveLength(2);
|
|
expect(
|
|
runnerMatrix.filter(entry => !entry.suite.manualOnly).reduce(
|
|
(total, execution) => total + execution.task.expectedRunCount,
|
|
0,
|
|
),
|
|
).toBe(371);
|
|
expect(
|
|
runnerTasks.find((task) => task.id === "plan-revise-accept")
|
|
?.attemptTimeoutMs,
|
|
).toEqual({ local: 8 * 60_000, daytona: 12 * 60_000 });
|
|
});
|
|
|
|
it("defines the warm Daytona continuity fixture as exactly two Codex cells", () => {
|
|
expect(daytonaWarmEnvironment).toMatchObject({
|
|
id: "daytona",
|
|
configurationKey: "warm-reuse-v1",
|
|
groups: ["daytona", "warm"],
|
|
});
|
|
expect(
|
|
daytonaWarmEnvironment.buildEnvironment({
|
|
secretRefs: {
|
|
DAYTONA_API_KEY: {
|
|
type: "secret_ref",
|
|
secretId: "22222222-2222-4222-8222-222222222222",
|
|
version: "latest",
|
|
},
|
|
},
|
|
daytonaImage: `runner@sha256:${"a".repeat(64)}`,
|
|
executionId: "warm",
|
|
}),
|
|
).toMatchObject({
|
|
config: {
|
|
reuseLease: true,
|
|
runnerLifecycleMode: "warm",
|
|
autoStopInterval: 5,
|
|
autoArchiveInterval: 15,
|
|
autoDeleteInterval: 60,
|
|
},
|
|
});
|
|
expect(daytonaWarmContinuityTask).toMatchObject({
|
|
flow: "warm_three_turn",
|
|
expectedRunCount: 3,
|
|
turnTimeoutMs: 600_000,
|
|
});
|
|
expect(
|
|
daytonaWarmContinuityTask.buildFollowupMessages?.("nonce"),
|
|
).toHaveLength(2);
|
|
const initialPrompt = daytonaWarmContinuityTask.buildPrompt("nonce");
|
|
const followups =
|
|
daytonaWarmContinuityTask.buildFollowupMessages?.("nonce") ?? [];
|
|
expect(initialPrompt).toContain('"kind":"request_confirmation"');
|
|
expect(initialPrompt).toContain(
|
|
'"reviewInteractionId":"<returned interaction id>"',
|
|
);
|
|
expect(initialPrompt).toContain('"continuationPolicy":"wake_assignee"');
|
|
expect(initialPrompt).toContain(
|
|
'"prompt":"Is this warm continuity task ready to complete after turn 1?"',
|
|
);
|
|
expect(initialPrompt).not.toContain("Continue to warm continuity turn 2?");
|
|
expect(followups[0]).toContain('"kind":"request_confirmation"');
|
|
expect(followups[0]).toContain(
|
|
'"prompt":"Is this warm continuity task ready to complete after turn 2?"',
|
|
);
|
|
expect(followups[0]).toContain(
|
|
'"reviewInteractionId":"<returned interaction id>"',
|
|
);
|
|
expect(followups[1]).toContain(
|
|
'{"status":"done","comment":"PAPERCLIP_E2E_WARM_T3_nonce"}',
|
|
);
|
|
expect(followups[1]).not.toContain('"kind":"request_confirmation"');
|
|
const cells = runnerMatrix.filter(
|
|
(entry) => entry.suite.id === "daytona-warm-continuity",
|
|
);
|
|
expect(cells.map((entry) => entry.profile.id)).toEqual([
|
|
"legacy-codex",
|
|
"runner-codex",
|
|
]);
|
|
expect(cells.every((entry) => entry.environment.id === "daytona")).toBe(
|
|
true,
|
|
);
|
|
const suite = runnerSuites.find(
|
|
(candidate) => candidate.id === "daytona-warm-continuity",
|
|
)!;
|
|
expect(
|
|
suiteDefinitionHash({
|
|
...suite,
|
|
environments: [
|
|
{ ...daytonaWarmEnvironment, configurationKey: "changed" },
|
|
],
|
|
}),
|
|
).not.toBe(suiteDefinitionHash(suite));
|
|
});
|
|
|
|
it("derives the qualified local native OpenCode profiles from the ranked snapshot", () => {
|
|
expect(openRouterBreadthExcludedModelIds).toEqual(["xiaomi/mimo-v2.5"]);
|
|
expect(openRouterBreadthExcludedExecutionIds).toEqual([
|
|
"openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.plan-approve-complete",
|
|
"openrouter-model-breadth.openrouter-tencent-hy3.local.plan-approve-complete",
|
|
]);
|
|
expect(
|
|
openRouterBreadthExcludedExecutionIds.every(
|
|
(excludedExecutionId) =>
|
|
!runnerMatrix.some(
|
|
(execution) => execution.id === excludedExecutionId,
|
|
),
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
runnerMatrix
|
|
.filter(
|
|
(execution) => execution.profile.id === "openrouter-tencent-hy3",
|
|
)
|
|
.map((execution) => execution.task.id),
|
|
).toEqual(["hello-complete", "question-resume-complete"]);
|
|
expect(
|
|
openRouterBreadthProfiles.map((profile) => profile.ranking?.rank),
|
|
).toEqual([1, 3, 4, 5]);
|
|
expect(
|
|
openRouterBreadthProfiles.every(
|
|
(profile) =>
|
|
profile.adapterType === "paperclip_runner" &&
|
|
profile.provider === "opencode" &&
|
|
profile.model.startsWith("openrouter/") &&
|
|
profile.supportedEnvironments.join(",") === "local" &&
|
|
profile.modelQualification.source === "openrouter_rankings_snapshot",
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("defines deterministic two-run question and plan state machines", () => {
|
|
const localQuestion = localIntegrityTasks.find(
|
|
(task) => task.id === "structured-question-resume",
|
|
);
|
|
const restartQuestion = localIntegrityTasks.find(
|
|
(task) => task.id === "structured-question-restart-resume",
|
|
);
|
|
const question = openRouterBreadthTasks.find(
|
|
(task) => task.id === "question-resume-complete",
|
|
);
|
|
const plan = openRouterBreadthTasks.find(
|
|
(task) => task.id === "plan-approve-complete",
|
|
);
|
|
expect(question).toMatchObject({
|
|
flow: "question_resume_completion",
|
|
expectedRunCount: 2,
|
|
});
|
|
expect(question?.buildQuestionAnswer?.("nonce")).toMatchObject({
|
|
optionLabel: "Cobalt",
|
|
});
|
|
expect(localQuestion).toMatchObject({
|
|
flow: "question_resume_completion",
|
|
expectedRunCount: 2,
|
|
});
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain("ask_user_questions");
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"do not spell, quote, repeat, announce, or include PAPERCLIP_E2E_QUESTION_DONE_nonce",
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"refer to it only as “the terminal marker.”",
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
'API_ORIGIN="${PAPERCLIP_API_URL%/}"; API_ORIGIN="${API_ORIGIN%/api}"',
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
'"idempotencyKey":"question-nonce"',
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
'PATCH $API_ORIGIN/api/issues/$PAPERCLIP_TASK_ID with exactly {"status":"in_review"}',
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"Do not include `reviewInteractionId`",
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"retry only that PATCH and never POST the interaction again",
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"make exactly one completion write",
|
|
);
|
|
const legacyQuestionExitInstruction =
|
|
"In a legacy runner, after those two writes succeed, end the current response and heartbeat immediately. Do not wait, sleep, poll, or fetch the interaction; `wake_assignee` will start a new heartbeat after the user answers.";
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
legacyQuestionExitInstruction,
|
|
);
|
|
expect(restartQuestion).toMatchObject({
|
|
flow: "question_resume_completion",
|
|
expectedRunCount: 2,
|
|
restartServerBeforeQuestionAnswer: true,
|
|
});
|
|
expect(restartQuestion?.buildPrompt("nonce")).toContain(
|
|
legacyQuestionExitInstruction,
|
|
);
|
|
expect(plan).toMatchObject({
|
|
flow: "plan_approval_completion",
|
|
expectedRunCount: 2,
|
|
});
|
|
expect(plan?.buildPrompt("nonce")).toContain("exactly two numbered steps");
|
|
});
|
|
|
|
it("emits native terminal text after the terminal tool succeeds", () => {
|
|
const message = runnerTasks.find((task) => task.id === "message-marker");
|
|
const ask = runnerTasks.find((task) => task.id === "ask-question");
|
|
const plan = runnerTasks.find((task) => task.id === "plan-revise-accept");
|
|
const question = localIntegrityTasks.find(
|
|
(task) => task.id === "structured-question-resume",
|
|
);
|
|
const breadthTasks = openRouterBreadthTasks.map((task) =>
|
|
task.buildPrompt("nonce"),
|
|
);
|
|
|
|
for (const prompt of [
|
|
message?.buildPrompt("nonce"),
|
|
ask?.buildPrompt("nonce"),
|
|
plan?.buildPrompt("nonce"),
|
|
question?.buildPrompt("nonce"),
|
|
...breadthTasks,
|
|
]) {
|
|
const terminalTextInstruction = prompt?.match(/then emit (?:exactly|only)/)?.[0];
|
|
expect(terminalTextInstruction).toBeDefined();
|
|
expect(prompt!.indexOf("paperclip_finish exactly once")).toBeLessThan(
|
|
prompt!.indexOf(terminalTextInstruction!),
|
|
);
|
|
expect(prompt).toContain("Wait for that tool call to succeed");
|
|
}
|
|
|
|
for (const taskId of [
|
|
"question-resume-complete",
|
|
"plan-approve-complete",
|
|
]) {
|
|
const prompt = openRouterBreadthTasks
|
|
.find((task) => task.id === taskId)
|
|
?.buildPrompt("nonce");
|
|
expect(prompt).toContain(
|
|
"do not spell, quote, repeat, announce, or include",
|
|
);
|
|
expect(prompt).toContain("refer to it only as “the terminal marker.”");
|
|
}
|
|
|
|
const breadthHello = openRouterBreadthTasks
|
|
.find((task) => task.id === "hello-complete")
|
|
?.buildPrompt("nonce");
|
|
expect(breadthHello).toContain(
|
|
"Your first response action must be the paperclip_finish tool call",
|
|
);
|
|
expect(breadthHello).toContain(
|
|
"Do not emit any assistant text, acknowledgement, or preamble before calling it",
|
|
);
|
|
|
|
const nativeAsk = ask?.buildPrompt("nonce");
|
|
expect(nativeAsk).toContain("paperclip_finish must be your only tool call");
|
|
expect(nativeAsk).toContain(
|
|
"never call report_progress or any other tool before or after it",
|
|
);
|
|
});
|
|
|
|
it("uses only declared secret references in generated payloads", () => {
|
|
expect(
|
|
runnerMatrix.every((entry) =>
|
|
entry.requiredCredentials.includes(entry.profile.credential),
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
runnerMatrix
|
|
.filter((entry) => entry.environment.id === "daytona")
|
|
.every((entry) =>
|
|
entry.requiredCredentials.includes("DAYTONA_API_KEY"),
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("pins legacy Codex and Claude to their classic CLI engines", () => {
|
|
for (const profileId of ["legacy-codex", "legacy-claude"]) {
|
|
const execution = runnerMatrix.find(
|
|
(candidate) =>
|
|
candidate.profile.id === profileId &&
|
|
candidate.environment.id === "local",
|
|
);
|
|
expect(execution).toBeDefined();
|
|
expect(
|
|
execution!.profile.buildAgent({
|
|
environmentId: "11111111-1111-4111-8111-111111111111",
|
|
environmentFixtureId: "local",
|
|
workspacePath: "/tmp/runner-e2e-workspace",
|
|
secretRefs: {
|
|
[execution!.profile.credential]: {
|
|
type: "secret_ref",
|
|
secretId: "22222222-2222-4222-8222-222222222222",
|
|
version: "latest",
|
|
},
|
|
},
|
|
executionId: execution!.id,
|
|
}),
|
|
).toMatchObject({
|
|
adapterConfig: {
|
|
engine: "cli",
|
|
...(profileId === "legacy-codex"
|
|
? { extraArgs: ["-c", "features.shell_snapshot=false"] }
|
|
: {}),
|
|
},
|
|
});
|
|
}
|
|
});
|
|
|
|
it("binds native Codex automation auth to the encrypted OpenAI secret", () => {
|
|
const execution = runnerMatrix.find(
|
|
(candidate) =>
|
|
candidate.id === "core-compatibility.runner-codex.local.message-marker",
|
|
);
|
|
expect(execution).toBeDefined();
|
|
const secretRef = {
|
|
type: "secret_ref" as const,
|
|
secretId: "22222222-2222-4222-8222-222222222222",
|
|
version: "latest" as const,
|
|
};
|
|
const agent = execution!.profile.buildAgent({
|
|
environmentId: "11111111-1111-4111-8111-111111111111",
|
|
environmentFixtureId: "local",
|
|
workspacePath: "/tmp/runner-e2e-workspace",
|
|
secretRefs: { OPENAI_API_KEY: secretRef },
|
|
executionId: execution!.id,
|
|
});
|
|
expect(agent.adapterConfig).toMatchObject({
|
|
env: {
|
|
OPENAI_API_KEY: secretRef,
|
|
CODEX_API_KEY: secretRef,
|
|
},
|
|
});
|
|
});
|
|
|
|
it("gives legacy planning agents a direct bounded API recipe", () => {
|
|
const task = runnerTasks.find(
|
|
(candidate) => candidate.id === "plan-revise-accept",
|
|
);
|
|
const execution = runnerMatrix.find(
|
|
(candidate) =>
|
|
candidate.profile.id === "legacy-claude" &&
|
|
candidate.environment.id === "local" &&
|
|
candidate.task.id === "plan-revise-accept",
|
|
);
|
|
expect(task).toBeDefined();
|
|
expect(execution).toBeDefined();
|
|
const agent = execution!.profile.buildAgent({
|
|
environmentId: "11111111-1111-4111-8111-111111111111",
|
|
environmentFixtureId: "local",
|
|
workspacePath: "/tmp/runner-e2e-workspace",
|
|
secretRefs: {
|
|
ANTHROPIC_API_KEY: {
|
|
type: "secret_ref",
|
|
secretId: "22222222-2222-4222-8222-222222222222",
|
|
version: "latest",
|
|
},
|
|
},
|
|
executionId: execution!.id,
|
|
});
|
|
expect(agent.adapterConfig).toMatchObject({ maxTurnsPerRun: 24 });
|
|
expect(agent.instructionsBundle).toMatchObject({
|
|
files: { "AGENTS.md": expect.stringContaining("/interactions") },
|
|
});
|
|
expect(task!.buildPrompt("nonce")).toContain("request_confirmation");
|
|
expect(task!.buildPrompt("nonce")).toContain("baseRevisionId");
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"do not spell, quote, repeat, announce, or include PAPERCLIP_E2E_PLAN_DONE_nonce",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
'summary:"PAPERCLIP_E2E_PLAN_DONE_nonce"',
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain("first call get_task_context");
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"identifies the exact revised Plan revision used as the confirmation target as accepted",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"After that verification succeeds, your immediate next action must be the paperclip_finish tool call",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).not.toContain(
|
|
"trust that inline acceptance",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"those two tool calls form one indivisible response sequence",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"Do not emit assistant text, end the response or heartbeat, or stop after write_document alone",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"one atomic issue PATCH with status `done` and that exact comment",
|
|
);
|
|
const revisionRequest = task!.buildRevisionRequest?.("nonce");
|
|
expect(revisionRequest).toContain("baseRevisionId");
|
|
expect(revisionRequest).toContain(
|
|
"request_human_input must be your immediate next action",
|
|
);
|
|
});
|
|
|
|
it("requires one atomic legacy Ask completion write", () => {
|
|
const task = runnerTasks.find(
|
|
(candidate) => candidate.id === "ask-question",
|
|
);
|
|
expect(task).toBeDefined();
|
|
const prompt = task!.buildPrompt("nonce");
|
|
expect(prompt).toContain(
|
|
"make exactly one public-API write containing the marker",
|
|
);
|
|
expect(prompt).toContain(
|
|
'PATCH /api/issues/$PAPERCLIP_TASK_ID with {"status":"done","comment":"E2E_ASK_12_nonce"}',
|
|
);
|
|
expect(prompt).toContain("Do not POST to /comments");
|
|
expect(prompt).toContain("do not PATCH the status separately");
|
|
});
|
|
|
|
it("accepts only complete immutable Daytona digests", () => {
|
|
expect(
|
|
isImmutableDaytonaImage(
|
|
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:${"a".repeat(64)}`,
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
isImmutableDaytonaImage(
|
|
"ghcr.io/paperclipai/paperclip-daytona-runner@sha256:REPLACE_ME",
|
|
),
|
|
).toBe(false);
|
|
expect(
|
|
isImmutableDaytonaImage(
|
|
"ghcr.io/paperclipai/paperclip-daytona-runner:e2e-latest",
|
|
),
|
|
).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe("runner E2E selectors", () => {
|
|
it("requires an explicit billable selector", () => {
|
|
expect(() => parseRunnerSelectors([])).toThrow(RunnerSelectorError);
|
|
});
|
|
|
|
it("selects dimensions with OR within a dimension and AND across dimensions", () => {
|
|
const options = parseRunnerSelectors([
|
|
"--profile",
|
|
"legacy-codex",
|
|
"--profile",
|
|
"runner-codex",
|
|
"--environment",
|
|
"local",
|
|
]);
|
|
expect(selectRunnerExecutions(options).map((entry) => entry.id)).toEqual([
|
|
...runnerMatrix.filter(entry => entry.suite.id === "continuation" && ["legacy-codex", "runner-codex"].includes(entry.profile.id)).map(entry => entry.id),
|
|
...runnerMatrix.filter(entry => entry.suite.id === "first-task" && ["legacy-codex", "runner-codex"].includes(entry.profile.id)).map(entry => entry.id),
|
|
...runnerMatrix.filter(entry => entry.suite.id === "agent-chat" && ["legacy-codex", "runner-codex"].includes(entry.profile.id)).map(entry => entry.id),
|
|
"core-compatibility.legacy-codex.local.message-marker",
|
|
"core-compatibility.legacy-codex.local.plan-revise-accept",
|
|
"core-compatibility.legacy-codex.local.ask-question",
|
|
"core-compatibility.runner-codex.local.message-marker",
|
|
"core-compatibility.runner-codex.local.plan-revise-accept",
|
|
"core-compatibility.runner-codex.local.ask-question",
|
|
"local-session-integrity.legacy-codex.local.structured-question-resume",
|
|
"local-session-integrity.legacy-codex.local.structured-question-restart-resume",
|
|
"local-session-integrity.runner-codex.local.structured-question-resume",
|
|
"local-session-integrity.runner-codex.local.structured-question-restart-resume",
|
|
]);
|
|
});
|
|
|
|
it("selects a suite without exploding its environment matrix", () => {
|
|
const selected = selectRunnerExecutions(
|
|
parseRunnerSelectors(["--suite", "openrouter-model-breadth"]),
|
|
);
|
|
expect(selected).toHaveLength(10);
|
|
expect(
|
|
selected.every(
|
|
(entry) =>
|
|
entry.suite.id === "openrouter-model-breadth" &&
|
|
entry.environment.id === "local",
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("combines repeated groups with AND semantics", () => {
|
|
const options = parseRunnerSelectors([
|
|
"--group",
|
|
"native",
|
|
"--group",
|
|
"daytona",
|
|
]);
|
|
const selected = selectRunnerExecutions(options);
|
|
expect(selected).toHaveLength(13);
|
|
expect(
|
|
selected.every(
|
|
(entry) =>
|
|
entry.profile.generation === "native" &&
|
|
entry.environment.id === "daytona",
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("rejects unknown groups", () => {
|
|
const options = parseRunnerSelectors(["--group", "codex"]);
|
|
expect(() => selectRunnerExecutions(options)).toThrow("Unknown group");
|
|
});
|
|
|
|
it("emits one independently schedulable job per scenario", () => {
|
|
const jobs = buildMatrixJobs(
|
|
selectRunnerExecutions(parseRunnerSelectors(["--all"])),
|
|
);
|
|
expect(jobs).toHaveLength(171);
|
|
expect(jobs.filter((job) => job.needsDaytona)).toHaveLength(23);
|
|
expect(jobs.filter((job) => !job.needsDaytona)).toHaveLength(148);
|
|
expect(new Set(jobs.map((job) => job.executionId)).size).toBe(171);
|
|
expect(
|
|
jobs.find(
|
|
(job) =>
|
|
job.executionId ===
|
|
"core-compatibility.runner-acpx-claude.local.plan-revise-accept",
|
|
)?.timeoutMinutes,
|
|
).toBe(25);
|
|
expect(
|
|
jobs.find(
|
|
(job) =>
|
|
job.executionId ===
|
|
"local-session-integrity.runner-acpx-codex.local.structured-question-restart-resume",
|
|
)?.timeoutMinutes,
|
|
).toBe(32);
|
|
expect(
|
|
jobs.every((job) =>
|
|
runnerMatrix.some(
|
|
(execution) =>
|
|
execution.id === job.executionId &&
|
|
execution.profile.credential === job.credentialName,
|
|
),
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("validates bounded local parallelism", () => {
|
|
expect(
|
|
parseRunnerSelectors(["--all", "--max-parallel", "8"]).maxParallel,
|
|
).toBe(8);
|
|
expect(() =>
|
|
parseRunnerSelectors(["--all", "--max-parallel", "0"]),
|
|
).toThrow("positive integer");
|
|
});
|
|
});
|