mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 03:08:10 +02:00
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent chat uses native runner sessions to plan, delegate, and track that work. > - A user can press Stop while the native session is still starting. > - The server can acknowledge that Stop without dispatching it, then let the session submit a turn. > - This leaves chat recovery waiting for an execution that the user expected to stop. > - This PR waits for the startup handle, dispatches cancellation, and prevents a late startup from submitting a turn. > - New full-stack evals check the resulting records and outputs across Claude and Codex. > - Those evals also exposed missing ACPX readiness fields, unbounded polling, and an old-run identity check that rejected valid warm handoffs. ## Linked Issues or Issue Description **What happened?** Stop during native startup could record an acknowledged cancellation with `dispatched: false`. The provider could then begin work. A subsequent `/new` stayed queued. A remote Claude follow-up also exhausted the command journal while probing warm-session readiness: ACPX never returned the readiness fields required by the shared transport. Once readiness worked, attachment incorrectly compared the next run descriptor against the old run ID. The 25 ms polling loop could issue 4,800 commands during its two-minute wait, beyond the 500-command bound. The existing chat eval treated lifecycle logs as proof of an active provider turn, so it did not distinguish startup cancellation from active-turn cancellation. **Expected behavior** A Stop during startup must reach the pending session. A late session must not submit a prompt after Stop. Recovery must retain control when startup exceeds the bounded wait. Chat evals must check saved task state, document contents, worker identity, account binding, and duplicate effects. **Steps to reproduce** 1. Start a native Claude or Codex chat turn. 2. Press Stop after process startup is requested but before the provider turn starts. 3. Send `/new`, then send a fresh message. 4. On the affected base, cancellation can be acknowledged without dispatch and the reset stays queued. **Paperclip version or commit** The live Claude baseline reproduced this on `29d6b3509`. The branch also includes master commit `0f5fafe16`. Related work: #13678, #13686, #13693, #13291, #13738. A separate runner reliability branch also contains a startup-wait fix. Its overlap must be reconciled before merging; this branch additionally prevents prompt submission after a late startup. ## What Changed - Wait for a pending native startup before acknowledging a run-scoped Stop. Preserve the existing recovery error when that wait expires. - Keep a Stop guard on startup. Cancel a late handle before it can submit a provider turn. - Add regression tests for normal handle publication and publication after the Stop deadline. - Back off blocked warm-attachment probes. Keep the fast two-snapshot barrier, fail closed, and record changed blockers. - Add red/green tests for delayed readiness, persistent blockers, alternating readiness, and readiness near the deadline. - Publish ACPX readiness and blockers. Preserve the old authority’s event acknowledgement barrier; only settled sessions can proceed to attachment. - Bind warm ACPX descriptors to the validated next authority while retaining old-run event correlation until activation. Preserve session identity and provider profile checks. - Exercise two consecutive run rotations through a qualified fake sidecar, verifying checkpointing, provider identity, pre-activation rejection, and new-run work admission. - Separate startup and active-turn cancellation checkpoints in the browser eval. - Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells across Claude and Codex. - Cover hiring and reuse through managed AI accounts, source-based review, current blocked-task status, request replay after a lost HTTP acknowledgement, server restart continuity, and Stop/reset continuity. - Use ordinary production agent instructions. Enable API tools only for the two coordination cases that need them. - Calibrate the matchers with invalid records and outputs. Require remembered context after restart and a structured status snapshot that distinguishes the current blocker from history and task status from active execution. Compare the public issue mutation contract and relationships during read-only reporting. Preserve before/after source records in failed eval evidence. - Fix the lost-ack browser harness and verify it against a real HTTP server. Check the chat composer after restart instead of waiting for an unrelated document lifecycle event. - Document the scope and limits of each case. ## Verification - The startup regression failed on the unfixed executor and passed after the fix. - `pnpm test:e2e:runner:typecheck` passed. - `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files. - `pnpm exec vitest run server/src/services/native-runtime/native-session-executor.test.ts` passed: 385 tests. - [Baseline live campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868): Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed. Codex delegation was blocked by provider capacity. - [Eval-only startup campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786): both providers failed as expected. Both persisted `dispatched: false` and left `/new` queued. - [First fixed campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706) on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both providers. Failed cases exposed eval harness defects and remote continuity failures. All attempts remain available. - [Original workflows and stronger memory checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649) on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and startup Stop passed for both providers; Codex remote restart passed. Claude remote restart exposed the missing readiness contract. Two Codex planning cells hit provider capacity. - [Unchanged-model retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548): Codex planning and backlog creation both passed. - [18-cell campaign with ACPX readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963) on `6a98ef743`: 16/18 passed, including all local/remote Stop and committed-send cases. Claude remote continuity exposed the next-authority check, now fixed. Codex hiring produced its checklist, but the runner redacted the requested marker after it appeared as “Tracking token: …”. That content-redaction policy is unchanged and remains an explicit limitation. - [Structured status grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725) on `50448c228`: both providers passed on their first attempt, including cleanup. - [Complete read-only state grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011) on `551e13892`: both providers passed. - [Final ACPX handoff and hiring retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456) on `cbd637587`: all three Claude Daytona cases passed (restart continuity, active Stop/reset, and lost-ack replay). Codex hiring reproduced the content-redaction failure: the saved checklist contained `Tracking token: [REDACTED]` instead of the required business marker. All four cases completed cleanup successfully. [Published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/). The only subsequent commit adds the qualified-sidecar integration test; production code is identical to this live proof. - `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without paid models. - Runner TypeScript typecheck passed. All 5 warm-readiness tests pass; two failed with the prior fixed-rate loop, and the late-readiness test failed before the pacing correction. - ACPX readiness and warm-identity regressions each failed before their fixes. All 292 runner-core Rust library tests passed. The qualified-sidecar integration test passes. Rust formatting is checked. - Status-grader regressions for misleading historical mentions and previously unchecked mutations each failed before tightening the oracle and pass now. - [Latest-head CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307) passed on `a4093c8f1`: full build, type checks, test partitions, browser E2E, and native runner checks. Two unrelated tests initially failed (Sentry fixture release attribution and local-service fixture readiness); both passed locally together (35 passed, 5 optional SDK tests skipped) and on the failed-job retry. No changes were made to those tests. - Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed and all review threads are resolved. - The paid live suite is not fully green: the reproducible content-redaction case remains red. This is separate from the passing PR merge checks. No production content-redaction, prompt, model, or completion-policy change is included. - Managed-account hiring and review cases explicitly enable API tools; these do not qualify default new-user onboarding. ## Risks - Stop can wait up to 30 seconds for startup, then use the existing pending-recovery path. This does not prove that remote cleanup has finished. - Blocked warm readiness adds up to 750 ms between later probes with the two-minute remote budget, or about 32 ms with the default five-second budget. Ready sessions retain the short second barrier. - Paid evals can fail because of provider capacity or agent decisions. Each failure needs evidence-based classification. - The HTTP request replay case checks comment idempotency and duplicate effects. It does not prove replay safety for an ambiguous provider tool call. - The new suite is opt-in. It does not increase the default paid campaign. - No production prompts or model selection change. Review-handoff behavior and content-redaction policy remain separate product decisions. The latter can remove harmless business content that looks like credential syntax; the failing attempt is retained. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
632 lines
24 KiB
TypeScript
632 lines
24 KiB
TypeScript
import { describe, expect, it } from "vitest";
|
|
import { normalizePrpResultSignals } from "../../packages/paperclip-runner/src/protocol/result-normalization.js";
|
|
import {
|
|
connectionReviewSuite,
|
|
runnerEnvironments,
|
|
runnerMatrix,
|
|
openRouterBreadthExcludedExecutionIds,
|
|
openRouterBreadthExcludedModelIds,
|
|
openRouterBreadthProfiles,
|
|
openRouterBreadthTasks,
|
|
localIntegrityTasks,
|
|
runnerProfiles,
|
|
runnerSuites,
|
|
runnerTasks,
|
|
daytonaWarmContinuityTask,
|
|
daytonaWarmEnvironment,
|
|
isImmutableDaytonaImage,
|
|
suiteDefinitionHash,
|
|
validateRunnerCatalog,
|
|
} from "./catalog.js";
|
|
import {
|
|
buildMatrixJobs,
|
|
parseRunnerSelectors,
|
|
RunnerSelectorError,
|
|
selectRunnerExecutions,
|
|
} from "./selectors.js";
|
|
|
|
describe("runner E2E catalog", () => {
|
|
it("supplies an actionable human review in native warm completion examples", () => {
|
|
const prompts = [daytonaWarmContinuityTask.buildPrompt("nonce"), ...daytonaWarmContinuityTask.buildFollowupMessages!("nonce")];
|
|
for (const [index, prompt] of prompts.entries()) {
|
|
const match = prompt.match(/attentionRequests:(\[.*?\]),evidence:/);
|
|
expect(match).not.toBeNull();
|
|
const signals = normalizePrpResultSignals({ attentionRequests: JSON.parse(match![1]) });
|
|
expect(signals.ignoredAttentionRequests).toEqual([]);
|
|
expect(signals.actionableAttentionRequests).toHaveLength(index === 2 ? 0 : 1);
|
|
if (index < 2) expect(signals.actionableAttentionRequests[0]).toMatchObject({kind:"review", ownerClass:"human"});
|
|
expect(prompt).not.toContain("call request_human_input");
|
|
}
|
|
});
|
|
|
|
it("defines sixteen local connection-review journeys without expanding the default matrix", () => {
|
|
expect(connectionReviewSuite.expectedMatrixSize).toBe(16);
|
|
expect(new Set(connectionReviewSuite.profiles.map(profile => profile.id))).toEqual(new Set(["runner-codex", "runner-acpx-claude", "legacy-codex", "legacy-claude"]));
|
|
expect(connectionReviewSuite.environments.map(environment => environment.id)).toEqual(["local"]);
|
|
expect(connectionReviewSuite.tasks.map(task => task.toolReviewDecision)).toEqual(["approve", "decline", "always", "restart"]);
|
|
expect(connectionReviewSuite.tasks.every(task => task.flow === "governed_tool_review")).toBe(true);
|
|
});
|
|
|
|
it("tests native chat plans, tasks, and reassignment with production permission defaults", () => {
|
|
const suite = runnerSuites.find(suite => suite.id === "agent-chat")!;
|
|
for (const id of ["runner-codex", "runner-acpx-claude"]) {
|
|
const profile = suite.profiles.find(profile => profile.id === id)!;
|
|
const payload = profile.buildAgent({
|
|
executionId: "default-permissions", workspacePath: "/workspace", environmentId: "env-1", environmentFixtureId: "local",
|
|
secretRefs: {
|
|
[profile.credential]: { type: "secret_ref", secretId: "22222222-2222-4222-8222-222222222222", version: "latest" },
|
|
},
|
|
});
|
|
expect(payload.adapterConfig).not.toHaveProperty("acpxPermissionMode");
|
|
expect(payload.adapterConfig).not.toHaveProperty("codexPermissionMode");
|
|
expect(suite.tasks.map(task => task.id)).toEqual(expect.arrayContaining(["plan-handoff", "reassign-task", "create-backlog"]));
|
|
}
|
|
});
|
|
|
|
it("validates the core, local-integrity, breadth, and warm suites", () => {
|
|
expect(runnerProfiles).toHaveLength(7);
|
|
expect(openRouterBreadthProfiles).toHaveLength(4);
|
|
expect(runnerEnvironments).toHaveLength(2);
|
|
expect(runnerTasks).toHaveLength(3);
|
|
expect(localIntegrityTasks).toHaveLength(2);
|
|
expect(openRouterBreadthTasks).toHaveLength(3);
|
|
expect(runnerSuites.map((suite) => suite.expectedMatrixSize)).toEqual([
|
|
23, 38, 52, 28, 18, 42, 14, 10, 2,
|
|
]);
|
|
expect(validateRunnerCatalog()).toHaveLength(227);
|
|
expect(new Set(runnerMatrix.map((entry) => entry.id)).size).toBe(227);
|
|
expect(
|
|
runnerMatrix.filter((entry) => entry.suite.id === "core-compatibility"),
|
|
).toHaveLength(42);
|
|
expect(
|
|
runnerMatrix.filter(
|
|
(entry) => entry.suite.id === "local-session-integrity",
|
|
),
|
|
).toHaveLength(14);
|
|
expect(
|
|
runnerMatrix.filter(
|
|
(entry) => entry.suite.id === "openrouter-model-breadth",
|
|
),
|
|
).toHaveLength(10);
|
|
expect(
|
|
runnerMatrix.filter(
|
|
(entry) => entry.suite.id === "daytona-warm-continuity",
|
|
),
|
|
).toHaveLength(2);
|
|
expect(
|
|
runnerMatrix.filter(entry => !entry.suite.manualOnly).reduce(
|
|
(total, execution) => total + execution.task.expectedRunCount,
|
|
0,
|
|
),
|
|
).toBe(371);
|
|
expect(
|
|
runnerTasks.find((task) => task.id === "plan-revise-accept")
|
|
?.attemptTimeoutMs,
|
|
).toEqual({ local: 8 * 60_000, daytona: 12 * 60_000 });
|
|
});
|
|
|
|
it("defines the warm Daytona continuity fixture as exactly two Codex cells", () => {
|
|
expect(daytonaWarmEnvironment).toMatchObject({
|
|
id: "daytona",
|
|
configurationKey: "warm-reuse-v1",
|
|
groups: ["daytona", "warm"],
|
|
});
|
|
expect(
|
|
daytonaWarmEnvironment.buildEnvironment({
|
|
secretRefs: {
|
|
DAYTONA_API_KEY: {
|
|
type: "secret_ref",
|
|
secretId: "22222222-2222-4222-8222-222222222222",
|
|
version: "latest",
|
|
},
|
|
},
|
|
daytonaImage: `runner@sha256:${"a".repeat(64)}`,
|
|
executionId: "warm",
|
|
}),
|
|
).toMatchObject({
|
|
config: {
|
|
reuseLease: true,
|
|
runnerLifecycleMode: "warm",
|
|
autoStopInterval: 5,
|
|
autoArchiveInterval: 15,
|
|
autoDeleteInterval: 60,
|
|
},
|
|
});
|
|
expect(daytonaWarmContinuityTask).toMatchObject({
|
|
flow: "warm_three_turn",
|
|
expectedRunCount: 3,
|
|
turnTimeoutMs: 600_000,
|
|
});
|
|
expect(
|
|
daytonaWarmContinuityTask.buildFollowupMessages?.("nonce"),
|
|
).toHaveLength(2);
|
|
const initialPrompt = daytonaWarmContinuityTask.buildPrompt("nonce");
|
|
const followups =
|
|
daytonaWarmContinuityTask.buildFollowupMessages?.("nonce") ?? [];
|
|
expect(initialPrompt).toContain('"kind":"request_confirmation"');
|
|
expect(initialPrompt).toContain(
|
|
'"reviewInteractionId":"<returned interaction id>"',
|
|
);
|
|
expect(initialPrompt).toContain('"continuationPolicy":"wake_assignee"');
|
|
expect(initialPrompt).toContain(
|
|
'"prompt":"Is this warm continuity task ready to complete after turn 1?"',
|
|
);
|
|
expect(initialPrompt).not.toContain("Continue to warm continuity turn 2?");
|
|
expect(followups[0]).toContain('"kind":"request_confirmation"');
|
|
expect(followups[0]).toContain(
|
|
'"prompt":"Is this warm continuity task ready to complete after turn 2?"',
|
|
);
|
|
expect(followups[0]).toContain(
|
|
'"reviewInteractionId":"<returned interaction id>"',
|
|
);
|
|
expect(followups[1]).toContain(
|
|
'{"status":"done","comment":"PAPERCLIP_E2E_WARM_T3_nonce"}',
|
|
);
|
|
expect(followups[1]).not.toContain('"kind":"request_confirmation"');
|
|
const cells = runnerMatrix.filter(
|
|
(entry) => entry.suite.id === "daytona-warm-continuity",
|
|
);
|
|
expect(cells.map((entry) => entry.profile.id)).toEqual([
|
|
"legacy-codex",
|
|
"runner-codex",
|
|
]);
|
|
expect(cells.every((entry) => entry.environment.id === "daytona")).toBe(
|
|
true,
|
|
);
|
|
const suite = runnerSuites.find(
|
|
(candidate) => candidate.id === "daytona-warm-continuity",
|
|
)!;
|
|
expect(
|
|
suiteDefinitionHash({
|
|
...suite,
|
|
environments: [
|
|
{ ...daytonaWarmEnvironment, configurationKey: "changed" },
|
|
],
|
|
}),
|
|
).not.toBe(suiteDefinitionHash(suite));
|
|
});
|
|
|
|
it("derives the qualified local native OpenCode profiles from the ranked snapshot", () => {
|
|
expect(openRouterBreadthExcludedModelIds).toEqual(["xiaomi/mimo-v2.5"]);
|
|
expect(openRouterBreadthExcludedExecutionIds).toEqual([
|
|
"openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.plan-approve-complete",
|
|
"openrouter-model-breadth.openrouter-tencent-hy3.local.plan-approve-complete",
|
|
]);
|
|
expect(
|
|
openRouterBreadthExcludedExecutionIds.every(
|
|
(excludedExecutionId) =>
|
|
!runnerMatrix.some(
|
|
(execution) => execution.id === excludedExecutionId,
|
|
),
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
runnerMatrix
|
|
.filter(
|
|
(execution) => execution.profile.id === "openrouter-tencent-hy3",
|
|
)
|
|
.map((execution) => execution.task.id),
|
|
).toEqual(["hello-complete", "question-resume-complete"]);
|
|
expect(
|
|
openRouterBreadthProfiles.map((profile) => profile.ranking?.rank),
|
|
).toEqual([1, 3, 4, 5]);
|
|
expect(
|
|
openRouterBreadthProfiles.every(
|
|
(profile) =>
|
|
profile.adapterType === "paperclip_runner" &&
|
|
profile.provider === "opencode" &&
|
|
profile.model.startsWith("openrouter/") &&
|
|
profile.supportedEnvironments.join(",") === "local" &&
|
|
profile.modelQualification.source === "openrouter_rankings_snapshot",
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("defines deterministic two-run question and plan state machines", () => {
|
|
const localQuestion = localIntegrityTasks.find(
|
|
(task) => task.id === "structured-question-resume",
|
|
);
|
|
const restartQuestion = localIntegrityTasks.find(
|
|
(task) => task.id === "structured-question-restart-resume",
|
|
);
|
|
const question = openRouterBreadthTasks.find(
|
|
(task) => task.id === "question-resume-complete",
|
|
);
|
|
const plan = openRouterBreadthTasks.find(
|
|
(task) => task.id === "plan-approve-complete",
|
|
);
|
|
expect(question).toMatchObject({
|
|
flow: "question_resume_completion",
|
|
expectedRunCount: 2,
|
|
});
|
|
expect(question?.buildQuestionAnswer?.("nonce")).toMatchObject({
|
|
optionLabel: "Cobalt",
|
|
});
|
|
expect(localQuestion).toMatchObject({
|
|
flow: "question_resume_completion",
|
|
expectedRunCount: 2,
|
|
});
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain("ask_user_questions");
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"do not spell, quote, repeat, announce, or include PAPERCLIP_E2E_QUESTION_DONE_nonce",
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"refer to it only as “the terminal marker.”",
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
'API_ORIGIN="${PAPERCLIP_API_URL%/}"; API_ORIGIN="${API_ORIGIN%/api}"',
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
'"idempotencyKey":"question-nonce"',
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
'PATCH $API_ORIGIN/api/issues/$PAPERCLIP_TASK_ID with exactly {"status":"in_review"}',
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"Do not include `reviewInteractionId`",
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"retry only that PATCH and never POST the interaction again",
|
|
);
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
"make exactly one completion write",
|
|
);
|
|
const legacyQuestionExitInstruction =
|
|
"In a legacy runner, after those two writes succeed, end the current response and heartbeat immediately. Do not wait, sleep, poll, or fetch the interaction; `wake_assignee` will start a new heartbeat after the user answers.";
|
|
expect(localQuestion?.buildPrompt("nonce")).toContain(
|
|
legacyQuestionExitInstruction,
|
|
);
|
|
expect(restartQuestion).toMatchObject({
|
|
flow: "question_resume_completion",
|
|
expectedRunCount: 2,
|
|
restartServerBeforeQuestionAnswer: true,
|
|
});
|
|
expect(restartQuestion?.buildPrompt("nonce")).toContain(
|
|
legacyQuestionExitInstruction,
|
|
);
|
|
expect(plan).toMatchObject({
|
|
flow: "plan_approval_completion",
|
|
expectedRunCount: 2,
|
|
});
|
|
expect(plan?.buildPrompt("nonce")).toContain("exactly two numbered steps");
|
|
});
|
|
|
|
it("emits native terminal text after the terminal tool succeeds", () => {
|
|
const message = runnerTasks.find((task) => task.id === "message-marker");
|
|
const ask = runnerTasks.find((task) => task.id === "ask-question");
|
|
const plan = runnerTasks.find((task) => task.id === "plan-revise-accept");
|
|
const question = localIntegrityTasks.find(
|
|
(task) => task.id === "structured-question-resume",
|
|
);
|
|
const breadthTasks = openRouterBreadthTasks.map((task) =>
|
|
task.buildPrompt("nonce"),
|
|
);
|
|
|
|
for (const prompt of [
|
|
message?.buildPrompt("nonce"),
|
|
ask?.buildPrompt("nonce"),
|
|
plan?.buildPrompt("nonce"),
|
|
question?.buildPrompt("nonce"),
|
|
...breadthTasks,
|
|
]) {
|
|
const terminalTextInstruction = prompt?.match(/then emit (?:exactly|only)/)?.[0];
|
|
expect(terminalTextInstruction).toBeDefined();
|
|
expect(prompt!.indexOf("paperclip_finish exactly once")).toBeLessThan(
|
|
prompt!.indexOf(terminalTextInstruction!),
|
|
);
|
|
expect(prompt).toContain("Wait for that tool call to succeed");
|
|
}
|
|
|
|
for (const taskId of [
|
|
"question-resume-complete",
|
|
"plan-approve-complete",
|
|
]) {
|
|
const prompt = openRouterBreadthTasks
|
|
.find((task) => task.id === taskId)
|
|
?.buildPrompt("nonce");
|
|
expect(prompt).toContain(
|
|
"do not spell, quote, repeat, announce, or include",
|
|
);
|
|
expect(prompt).toContain("refer to it only as “the terminal marker.”");
|
|
}
|
|
|
|
const breadthHello = openRouterBreadthTasks
|
|
.find((task) => task.id === "hello-complete")
|
|
?.buildPrompt("nonce");
|
|
expect(breadthHello).toContain(
|
|
"Your first response action must be the paperclip_finish tool call",
|
|
);
|
|
expect(breadthHello).toContain(
|
|
"Do not emit any assistant text, acknowledgement, or preamble before calling it",
|
|
);
|
|
|
|
const nativeAsk = ask?.buildPrompt("nonce");
|
|
expect(nativeAsk).toContain("paperclip_finish must be your only tool call");
|
|
expect(nativeAsk).toContain(
|
|
"never call report_progress or any other tool before or after it",
|
|
);
|
|
});
|
|
|
|
it("uses only declared secret references in generated payloads", () => {
|
|
expect(
|
|
runnerMatrix.every((entry) =>
|
|
entry.requiredCredentials.includes(entry.profile.credential),
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
runnerMatrix
|
|
.filter((entry) => entry.environment.id === "daytona")
|
|
.every((entry) =>
|
|
entry.requiredCredentials.includes("DAYTONA_API_KEY"),
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("pins legacy Codex and Claude to their classic CLI engines", () => {
|
|
for (const profileId of ["legacy-codex", "legacy-claude"]) {
|
|
const execution = runnerMatrix.find(
|
|
(candidate) =>
|
|
candidate.profile.id === profileId &&
|
|
candidate.environment.id === "local",
|
|
);
|
|
expect(execution).toBeDefined();
|
|
expect(
|
|
execution!.profile.buildAgent({
|
|
environmentId: "11111111-1111-4111-8111-111111111111",
|
|
environmentFixtureId: "local",
|
|
workspacePath: "/tmp/runner-e2e-workspace",
|
|
secretRefs: {
|
|
[execution!.profile.credential]: {
|
|
type: "secret_ref",
|
|
secretId: "22222222-2222-4222-8222-222222222222",
|
|
version: "latest",
|
|
},
|
|
},
|
|
executionId: execution!.id,
|
|
}),
|
|
).toMatchObject({
|
|
adapterConfig: {
|
|
engine: "cli",
|
|
...(profileId === "legacy-codex"
|
|
? { extraArgs: ["-c", "features.shell_snapshot=false"] }
|
|
: {}),
|
|
},
|
|
});
|
|
}
|
|
});
|
|
|
|
it("binds native Codex automation auth to the encrypted OpenAI secret", () => {
|
|
const execution = runnerMatrix.find(
|
|
(candidate) =>
|
|
candidate.id === "core-compatibility.runner-codex.local.message-marker",
|
|
);
|
|
expect(execution).toBeDefined();
|
|
const secretRef = {
|
|
type: "secret_ref" as const,
|
|
secretId: "22222222-2222-4222-8222-222222222222",
|
|
version: "latest" as const,
|
|
};
|
|
const agent = execution!.profile.buildAgent({
|
|
environmentId: "11111111-1111-4111-8111-111111111111",
|
|
environmentFixtureId: "local",
|
|
workspacePath: "/tmp/runner-e2e-workspace",
|
|
secretRefs: { OPENAI_API_KEY: secretRef },
|
|
executionId: execution!.id,
|
|
});
|
|
expect(agent.adapterConfig).toMatchObject({
|
|
env: {
|
|
OPENAI_API_KEY: secretRef,
|
|
CODEX_API_KEY: secretRef,
|
|
},
|
|
});
|
|
});
|
|
|
|
it("gives legacy planning agents a direct bounded API recipe", () => {
|
|
const task = runnerTasks.find(
|
|
(candidate) => candidate.id === "plan-revise-accept",
|
|
);
|
|
const execution = runnerMatrix.find(
|
|
(candidate) =>
|
|
candidate.profile.id === "legacy-claude" &&
|
|
candidate.environment.id === "local" &&
|
|
candidate.task.id === "plan-revise-accept",
|
|
);
|
|
expect(task).toBeDefined();
|
|
expect(execution).toBeDefined();
|
|
const agent = execution!.profile.buildAgent({
|
|
environmentId: "11111111-1111-4111-8111-111111111111",
|
|
environmentFixtureId: "local",
|
|
workspacePath: "/tmp/runner-e2e-workspace",
|
|
secretRefs: {
|
|
ANTHROPIC_API_KEY: {
|
|
type: "secret_ref",
|
|
secretId: "22222222-2222-4222-8222-222222222222",
|
|
version: "latest",
|
|
},
|
|
},
|
|
executionId: execution!.id,
|
|
});
|
|
expect(agent.adapterConfig).toMatchObject({ maxTurnsPerRun: 24 });
|
|
expect(agent.instructionsBundle).toMatchObject({
|
|
files: { "AGENTS.md": expect.stringContaining("/interactions") },
|
|
});
|
|
expect(task!.buildPrompt("nonce")).toContain("request_confirmation");
|
|
expect(task!.buildPrompt("nonce")).toContain("baseRevisionId");
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"do not spell, quote, repeat, announce, or include PAPERCLIP_E2E_PLAN_DONE_nonce",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
'summary:"PAPERCLIP_E2E_PLAN_DONE_nonce"',
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain("first call get_task_context");
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"identifies the exact revised Plan revision used as the confirmation target as accepted",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"After that verification succeeds, your immediate next action must be the paperclip_finish tool call",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).not.toContain(
|
|
"trust that inline acceptance",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"those two tool calls form one indivisible response sequence",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"Do not emit assistant text, end the response or heartbeat, or stop after write_document alone",
|
|
);
|
|
expect(task!.buildPrompt("nonce")).toContain(
|
|
"one atomic issue PATCH with status `done` and that exact comment",
|
|
);
|
|
const revisionRequest = task!.buildRevisionRequest?.("nonce");
|
|
expect(revisionRequest).toContain("baseRevisionId");
|
|
expect(revisionRequest).toContain(
|
|
"request_human_input must be your immediate next action",
|
|
);
|
|
});
|
|
|
|
it("requires one atomic legacy Ask completion write", () => {
|
|
const task = runnerTasks.find(
|
|
(candidate) => candidate.id === "ask-question",
|
|
);
|
|
expect(task).toBeDefined();
|
|
const prompt = task!.buildPrompt("nonce");
|
|
expect(prompt).toContain(
|
|
"make exactly one public-API write containing the marker",
|
|
);
|
|
expect(prompt).toContain(
|
|
'PATCH /api/issues/$PAPERCLIP_TASK_ID with {"status":"done","comment":"E2E_ASK_12_nonce"}',
|
|
);
|
|
expect(prompt).toContain("Do not POST to /comments");
|
|
expect(prompt).toContain("do not PATCH the status separately");
|
|
});
|
|
|
|
it("accepts only complete immutable Daytona digests", () => {
|
|
expect(
|
|
isImmutableDaytonaImage(
|
|
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:${"a".repeat(64)}`,
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
isImmutableDaytonaImage(
|
|
"ghcr.io/paperclipai/paperclip-daytona-runner@sha256:REPLACE_ME",
|
|
),
|
|
).toBe(false);
|
|
expect(
|
|
isImmutableDaytonaImage(
|
|
"ghcr.io/paperclipai/paperclip-daytona-runner:e2e-latest",
|
|
),
|
|
).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe("runner E2E selectors", () => {
|
|
it("requires an explicit billable selector", () => {
|
|
expect(() => parseRunnerSelectors([])).toThrow(RunnerSelectorError);
|
|
});
|
|
|
|
it("selects dimensions with OR within a dimension and AND across dimensions", () => {
|
|
const options = parseRunnerSelectors([
|
|
"--profile",
|
|
"legacy-codex",
|
|
"--profile",
|
|
"runner-codex",
|
|
"--environment",
|
|
"local",
|
|
]);
|
|
expect(selectRunnerExecutions(options).map((entry) => entry.id)).toEqual([
|
|
...runnerMatrix.filter(entry => entry.suite.id === "continuation" && ["legacy-codex", "runner-codex"].includes(entry.profile.id)).map(entry => entry.id),
|
|
...runnerMatrix.filter(entry => entry.suite.id === "first-task" && ["legacy-codex", "runner-codex"].includes(entry.profile.id)).map(entry => entry.id),
|
|
...runnerMatrix.filter(entry => entry.suite.id === "agent-chat" && ["legacy-codex", "runner-codex"].includes(entry.profile.id)).map(entry => entry.id),
|
|
"core-compatibility.legacy-codex.local.message-marker",
|
|
"core-compatibility.legacy-codex.local.plan-revise-accept",
|
|
"core-compatibility.legacy-codex.local.ask-question",
|
|
"core-compatibility.runner-codex.local.message-marker",
|
|
"core-compatibility.runner-codex.local.plan-revise-accept",
|
|
"core-compatibility.runner-codex.local.ask-question",
|
|
"local-session-integrity.legacy-codex.local.structured-question-resume",
|
|
"local-session-integrity.legacy-codex.local.structured-question-restart-resume",
|
|
"local-session-integrity.runner-codex.local.structured-question-resume",
|
|
"local-session-integrity.runner-codex.local.structured-question-restart-resume",
|
|
]);
|
|
});
|
|
|
|
it("selects a suite without exploding its environment matrix", () => {
|
|
const selected = selectRunnerExecutions(
|
|
parseRunnerSelectors(["--suite", "openrouter-model-breadth"]),
|
|
);
|
|
expect(selected).toHaveLength(10);
|
|
expect(
|
|
selected.every(
|
|
(entry) =>
|
|
entry.suite.id === "openrouter-model-breadth" &&
|
|
entry.environment.id === "local",
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("combines repeated groups with AND semantics", () => {
|
|
const options = parseRunnerSelectors([
|
|
"--group",
|
|
"native",
|
|
"--group",
|
|
"daytona",
|
|
]);
|
|
const selected = selectRunnerExecutions(options);
|
|
expect(selected).toHaveLength(13);
|
|
expect(
|
|
selected.every(
|
|
(entry) =>
|
|
entry.profile.generation === "native" &&
|
|
entry.environment.id === "daytona",
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("rejects unknown groups", () => {
|
|
const options = parseRunnerSelectors(["--group", "codex"]);
|
|
expect(() => selectRunnerExecutions(options)).toThrow("Unknown group");
|
|
});
|
|
|
|
it("emits one independently schedulable job per scenario", () => {
|
|
const jobs = buildMatrixJobs(
|
|
selectRunnerExecutions(parseRunnerSelectors(["--all"])),
|
|
);
|
|
expect(jobs).toHaveLength(171);
|
|
expect(jobs.filter((job) => job.needsDaytona)).toHaveLength(23);
|
|
expect(jobs.filter((job) => !job.needsDaytona)).toHaveLength(148);
|
|
expect(new Set(jobs.map((job) => job.executionId)).size).toBe(171);
|
|
expect(
|
|
jobs.find(
|
|
(job) =>
|
|
job.executionId ===
|
|
"core-compatibility.runner-acpx-claude.local.plan-revise-accept",
|
|
)?.timeoutMinutes,
|
|
).toBe(25);
|
|
expect(
|
|
jobs.find(
|
|
(job) =>
|
|
job.executionId ===
|
|
"local-session-integrity.runner-acpx-codex.local.structured-question-restart-resume",
|
|
)?.timeoutMinutes,
|
|
).toBe(32);
|
|
expect(
|
|
jobs.every((job) =>
|
|
runnerMatrix.some(
|
|
(execution) =>
|
|
execution.id === job.executionId &&
|
|
execution.profile.credential === job.credentialName,
|
|
),
|
|
),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("validates bounded local parallelism", () => {
|
|
expect(
|
|
parseRunnerSelectors(["--all", "--max-parallel", "8"]).maxParallel,
|
|
).toBe(8);
|
|
expect(() =>
|
|
parseRunnerSelectors(["--all", "--max-parallel", "0"]),
|
|
).toThrow("positive integer");
|
|
});
|
|
});
|