Files
PaperClipAI/tests/runner-e2e/catalog.test.ts
T
DottaandPaperclip 43acbcc398 fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task state to provider sessions.
> - Follow-up turns must retain provider memory and carry new user
direction.
> - Lost session IDs caused repeated context and extra input tokens.
> - Native question answers and approval races could leave valid work
blocked.
> - This pull request repairs those paths and adds regression coverage.
> - Agents can continue accepted work without repeating the conversation
or losing the user's answer.

## Linked Issues or Issue Description

Refs #13574. That merged PR shortened continuation prompts and moved
question instructions into tool documentation. This change preserves
sessions and fixes failures exposed by broader testing. Related runtime
work: #13408 and #13410.

**What happened?**

Native follow-up turns could lose the provider session ID. Completion
guidance could replace the original task with its latest comment. Claude
native questions could remain pending after the user answered. Approval
during a running tool call could suspend the run before the tool
response arrived. Onboarding and chat handoff instructions also caused
repeated planning or missing plan documents.

**Expected behavior**

Reuse a valid provider session. Send only new events when that session
already has the history. Preserve the task requirements and apply later
user direction. Store the question answer and deliver it to the waiting
run. Finish governed tool responses before suspending. Execute the
accepted plan without asking for the same approval again.

**Steps to reproduce**

Run the continuation, local-session-integrity, first-task, and
agent-chat suites with native Codex and Claude. Include
provider-question-bridge, accept-while-running, and plan-handoff.

**Paperclip version or commit**

This branch is based on master d54b75011. The active full catalog run
tests 4e75881db. Later review fixes have separate regression coverage.

**Deployment mode**

Isolated local instances and Daytona sandboxes in the existing Runner
full-stack E2E harness.

## What Changed

- Retain provider session identity across turns and late usage
snapshots. Send new continuation events on session reuse, with full
context available for a fresh session.
- Preserve task requirements and later direction in completion guidance.
Return the current contract revision after a stale completion
submission.
- Bridge native Claude questions to saved Paperclip cards. Submit
answers through the saved card and resume the same run.
- Delay governed suspension until tool results settle. Add a
deterministic test barrier for approval during an active run.
- Clarify free-text question examples, explicit onboarding plans, and
execution of accepted chat plans.
- Fix continuation readiness, verified output evidence, and declared
screenshot collection.
- Qualify the legacy Claude test CLI at 2.1.277. The old 2.1.19 CLI did
not discover mounted skills. Update the existing workflow pin and
isolated launcher together.
- Refresh the Daytona image lockfile integrity pin after reviewing
master patch updates.
- Carry continuation mode as runtime metadata instead of inferring it
from user-visible text. Install the test Claude CLI without lifecycle
scripts.

## Verification

- Targeted paid verification: 20/20 cases passed across
local-environment campaigns before the rebase.
[Report](https://pages.paperclip.ing/runner-e2e-seven-fixes-35397904249/).
- Harness checks: 379 unit tests passed; harness typecheck passed.
- Latest-head PR checks: 55 passed, two intentionally skipped. Greptile
is 5/5; the security scan passes.
- Review regressions: 350 executor tests and 204 session/driver tests
passed. A script-free Claude install was verified with the actual CLI.
- Full catalog, including the explicit-only everyday suite: [run
35417932353](https://github.com/paperclipai/paperclip/actions/runs/35417932353).
Completed: **164/205 passed; 41 failed**. [Full dashboard and failure
investigation](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/).
Includes 204 case artifacts and one pre-case GitHub authorization
timeout; missing evidence is not scored as a pass. The full run tested
`4e75881db`; Final-head metadata/CLI smoke cases both passed. In the
separate [six infrastructure
retries](https://github.com/paperclipai/paperclip/actions/runs/35419769343),
the GitHub timeout case passed and all five Docker preflight failures
repeated. [Follow-up
dashboard](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/follow-up/).
- Full local typecheck and build passed on the rebased branch. The full
local unit run completed with 657 passing files, two test timeouts and
one suite setup timeout. All three affected files passed when rerun in
isolation (84 tests). The first full local run was not clean.
- Focused regression coverage includes the live question bridge,
same-run response delivery, UI routing, stale revisions, approval
overlap, and session reuse.

## Full-catalog follow-ups

- Test infrastructure: 14 Claude everyday cells probe an absent host
CLI; six cells failed pre-task GitHub/Docker qualification (GitHub
passes on retry; all five Docker cases repeat; the workflow preflight
allowlist omits their case IDs); five ACPX Codex cells cannot create
sandbox namespaces.
- Runtime: four OpenCode completion-criteria mismatches masked by
shutdown errors, one service-approval suspension failure; three Daytona
recovery failures encounter existing skill files; one duplicate
completion wake.
- Confirmed test defects: question pagination and a noncanonical plan
document key.
- Product/behavior: mismatched visible/required question sets, an
attachment instead of the requested task document, one lone-option
onboarding question, early completion instead of review, and a Codex
Mini completion-schema failure.
- The report job itself fails on trusted master’s stale patch/lock
configuration. The linked report is rebuilt with the shared renderer
from original cell results and public fixture screenshots; it excludes
private snapshots, logs and traces.

These are investigated follow-ups, not silently regraded passes.
First-task passed 51/52. The PR checks are green independently of the
broader catalog’s behavioral/infrastructure failures.

## Risks

- Session reuse depends on a valid provider identity and context
coverage. Fresh-session fallback and reset tests cover this boundary.
- Native question delivery spans saved interaction state and a live
provider run. Tests cover duplicate events, closed runs, and same-run
answers.
- Provider behavior varies. The full paid catalog may expose failures
beyond these targeted fixes; those results will be reported without
relaxing valid approval or output checks.
- The legacy Claude version update is limited to test infrastructure. No
database migration is included.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code execution,
and browser/E2E tools. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks and all
three timeout-file reruns pass; full-run timeout caveat above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 07:42:57 -05:00

602 lines
22 KiB
TypeScript

import { describe, expect, it } from "vitest";
import {
connectionReviewSuite,
runnerEnvironments,
runnerMatrix,
openRouterBreadthExcludedExecutionIds,
openRouterBreadthExcludedModelIds,
openRouterBreadthProfiles,
openRouterBreadthTasks,
localIntegrityTasks,
runnerProfiles,
runnerSuites,
runnerTasks,
daytonaWarmContinuityTask,
daytonaWarmEnvironment,
isImmutableDaytonaImage,
suiteDefinitionHash,
validateRunnerCatalog,
} from "./catalog.js";
import {
buildMatrixJobs,
parseRunnerSelectors,
RunnerSelectorError,
selectRunnerExecutions,
} from "./selectors.js";
describe("runner E2E catalog", () => {
it("defines sixteen local connection-review journeys without expanding the default matrix", () => {
expect(connectionReviewSuite.expectedMatrixSize).toBe(16);
expect(new Set(connectionReviewSuite.profiles.map(profile => profile.id))).toEqual(new Set(["runner-codex", "runner-acpx-claude", "legacy-codex", "legacy-claude"]));
expect(connectionReviewSuite.environments.map(environment => environment.id)).toEqual(["local"]);
expect(connectionReviewSuite.tasks.map(task => task.toolReviewDecision)).toEqual(["approve", "decline", "always", "restart"]);
expect(connectionReviewSuite.tasks.every(task => task.flow === "governed_tool_review")).toBe(true);
});
it("validates the core, local-integrity, breadth, and warm suites", () => {
expect(runnerProfiles).toHaveLength(7);
expect(openRouterBreadthProfiles).toHaveLength(4);
expect(runnerEnvironments).toHaveLength(2);
expect(runnerTasks).toHaveLength(3);
expect(localIntegrityTasks).toHaveLength(2);
expect(openRouterBreadthTasks).toHaveLength(3);
expect(runnerSuites.map((suite) => suite.expectedMatrixSize)).toEqual([
23, 38, 52, 24, 42, 14, 10, 2,
]);
expect(validateRunnerCatalog()).toHaveLength(205);
expect(new Set(runnerMatrix.map((entry) => entry.id)).size).toBe(205);
expect(
runnerMatrix.filter((entry) => entry.suite.id === "core-compatibility"),
).toHaveLength(42);
expect(
runnerMatrix.filter(
(entry) => entry.suite.id === "local-session-integrity",
),
).toHaveLength(14);
expect(
runnerMatrix.filter(
(entry) => entry.suite.id === "openrouter-model-breadth",
),
).toHaveLength(10);
expect(
runnerMatrix.filter(
(entry) => entry.suite.id === "daytona-warm-continuity",
),
).toHaveLength(2);
expect(
runnerMatrix.filter(entry => !entry.suite.manualOnly).reduce(
(total, execution) => total + execution.task.expectedRunCount,
0,
),
).toBe(363);
expect(
runnerTasks.find((task) => task.id === "plan-revise-accept")
?.attemptTimeoutMs,
).toEqual({ local: 8 * 60_000, daytona: 12 * 60_000 });
});
it("defines the warm Daytona continuity fixture as exactly two Codex cells", () => {
expect(daytonaWarmEnvironment).toMatchObject({
id: "daytona",
configurationKey: "warm-reuse-v1",
groups: ["daytona", "warm"],
});
expect(
daytonaWarmEnvironment.buildEnvironment({
secretRefs: {
DAYTONA_API_KEY: {
type: "secret_ref",
secretId: "22222222-2222-4222-8222-222222222222",
version: "latest",
},
},
daytonaImage: `runner@sha256:${"a".repeat(64)}`,
executionId: "warm",
}),
).toMatchObject({
config: {
reuseLease: true,
runnerLifecycleMode: "warm",
autoStopInterval: 5,
autoArchiveInterval: 15,
autoDeleteInterval: 60,
},
});
expect(daytonaWarmContinuityTask).toMatchObject({
flow: "warm_three_turn",
expectedRunCount: 3,
turnTimeoutMs: 600_000,
});
expect(
daytonaWarmContinuityTask.buildFollowupMessages?.("nonce"),
).toHaveLength(2);
const initialPrompt = daytonaWarmContinuityTask.buildPrompt("nonce");
const followups =
daytonaWarmContinuityTask.buildFollowupMessages?.("nonce") ?? [];
expect(initialPrompt).toContain('"kind":"request_confirmation"');
expect(initialPrompt).toContain(
'"reviewInteractionId":"<returned interaction id>"',
);
expect(initialPrompt).toContain('"continuationPolicy":"wake_assignee"');
expect(initialPrompt).toContain(
'"prompt":"Is this warm continuity task ready to complete after turn 1?"',
);
expect(initialPrompt).not.toContain("Continue to warm continuity turn 2?");
expect(followups[0]).toContain('"kind":"request_confirmation"');
expect(followups[0]).toContain(
'"prompt":"Is this warm continuity task ready to complete after turn 2?"',
);
expect(followups[0]).toContain(
'"reviewInteractionId":"<returned interaction id>"',
);
expect(followups[1]).toContain(
'{"status":"done","comment":"PAPERCLIP_E2E_WARM_T3_nonce"}',
);
expect(followups[1]).not.toContain('"kind":"request_confirmation"');
const cells = runnerMatrix.filter(
(entry) => entry.suite.id === "daytona-warm-continuity",
);
expect(cells.map((entry) => entry.profile.id)).toEqual([
"legacy-codex",
"runner-codex",
]);
expect(cells.every((entry) => entry.environment.id === "daytona")).toBe(
true,
);
const suite = runnerSuites.find(
(candidate) => candidate.id === "daytona-warm-continuity",
)!;
expect(
suiteDefinitionHash({
...suite,
environments: [
{ ...daytonaWarmEnvironment, configurationKey: "changed" },
],
}),
).not.toBe(suiteDefinitionHash(suite));
});
it("derives the qualified local native OpenCode profiles from the ranked snapshot", () => {
expect(openRouterBreadthExcludedModelIds).toEqual(["xiaomi/mimo-v2.5"]);
expect(openRouterBreadthExcludedExecutionIds).toEqual([
"openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.plan-approve-complete",
"openrouter-model-breadth.openrouter-tencent-hy3.local.plan-approve-complete",
]);
expect(
openRouterBreadthExcludedExecutionIds.every(
(excludedExecutionId) =>
!runnerMatrix.some(
(execution) => execution.id === excludedExecutionId,
),
),
).toBe(true);
expect(
runnerMatrix
.filter(
(execution) => execution.profile.id === "openrouter-tencent-hy3",
)
.map((execution) => execution.task.id),
).toEqual(["hello-complete", "question-resume-complete"]);
expect(
openRouterBreadthProfiles.map((profile) => profile.ranking?.rank),
).toEqual([1, 3, 4, 5]);
expect(
openRouterBreadthProfiles.every(
(profile) =>
profile.adapterType === "paperclip_runner" &&
profile.provider === "opencode" &&
profile.model.startsWith("openrouter/") &&
profile.supportedEnvironments.join(",") === "local" &&
profile.modelQualification.source === "openrouter_rankings_snapshot",
),
).toBe(true);
});
it("defines deterministic two-run question and plan state machines", () => {
const localQuestion = localIntegrityTasks.find(
(task) => task.id === "structured-question-resume",
);
const restartQuestion = localIntegrityTasks.find(
(task) => task.id === "structured-question-restart-resume",
);
const question = openRouterBreadthTasks.find(
(task) => task.id === "question-resume-complete",
);
const plan = openRouterBreadthTasks.find(
(task) => task.id === "plan-approve-complete",
);
expect(question).toMatchObject({
flow: "question_resume_completion",
expectedRunCount: 2,
});
expect(question?.buildQuestionAnswer?.("nonce")).toMatchObject({
optionLabel: "Cobalt",
});
expect(localQuestion).toMatchObject({
flow: "question_resume_completion",
expectedRunCount: 2,
});
expect(localQuestion?.buildPrompt("nonce")).toContain("ask_user_questions");
expect(localQuestion?.buildPrompt("nonce")).toContain(
"do not spell, quote, repeat, announce, or include PAPERCLIP_E2E_QUESTION_DONE_nonce",
);
expect(localQuestion?.buildPrompt("nonce")).toContain(
"refer to it only as “the terminal marker.”",
);
expect(localQuestion?.buildPrompt("nonce")).toContain(
'API_ORIGIN="${PAPERCLIP_API_URL%/}"; API_ORIGIN="${API_ORIGIN%/api}"',
);
expect(localQuestion?.buildPrompt("nonce")).toContain(
'"idempotencyKey":"question-nonce"',
);
expect(localQuestion?.buildPrompt("nonce")).toContain(
'PATCH $API_ORIGIN/api/issues/$PAPERCLIP_TASK_ID with exactly {"status":"in_review"}',
);
expect(localQuestion?.buildPrompt("nonce")).toContain(
"Do not include `reviewInteractionId`",
);
expect(localQuestion?.buildPrompt("nonce")).toContain(
"retry only that PATCH and never POST the interaction again",
);
expect(localQuestion?.buildPrompt("nonce")).toContain(
"make exactly one completion write",
);
const legacyQuestionExitInstruction =
"In a legacy runner, after those two writes succeed, end the current response and heartbeat immediately. Do not wait, sleep, poll, or fetch the interaction; `wake_assignee` will start a new heartbeat after the user answers.";
expect(localQuestion?.buildPrompt("nonce")).toContain(
legacyQuestionExitInstruction,
);
expect(restartQuestion).toMatchObject({
flow: "question_resume_completion",
expectedRunCount: 2,
restartServerBeforeQuestionAnswer: true,
});
expect(restartQuestion?.buildPrompt("nonce")).toContain(
legacyQuestionExitInstruction,
);
expect(plan).toMatchObject({
flow: "plan_approval_completion",
expectedRunCount: 2,
});
expect(plan?.buildPrompt("nonce")).toContain("exactly two numbered steps");
});
it("emits native terminal text after the terminal tool succeeds", () => {
const message = runnerTasks.find((task) => task.id === "message-marker");
const ask = runnerTasks.find((task) => task.id === "ask-question");
const plan = runnerTasks.find((task) => task.id === "plan-revise-accept");
const question = localIntegrityTasks.find(
(task) => task.id === "structured-question-resume",
);
const breadthTasks = openRouterBreadthTasks.map((task) =>
task.buildPrompt("nonce"),
);
for (const prompt of [
message?.buildPrompt("nonce"),
ask?.buildPrompt("nonce"),
plan?.buildPrompt("nonce"),
question?.buildPrompt("nonce"),
...breadthTasks,
]) {
const terminalTextInstruction = prompt?.match(/then emit (?:exactly|only)/)?.[0];
expect(terminalTextInstruction).toBeDefined();
expect(prompt!.indexOf("paperclip_finish exactly once")).toBeLessThan(
prompt!.indexOf(terminalTextInstruction!),
);
expect(prompt).toContain("Wait for that tool call to succeed");
}
for (const taskId of [
"question-resume-complete",
"plan-approve-complete",
]) {
const prompt = openRouterBreadthTasks
.find((task) => task.id === taskId)
?.buildPrompt("nonce");
expect(prompt).toContain(
"do not spell, quote, repeat, announce, or include",
);
expect(prompt).toContain("refer to it only as “the terminal marker.”");
}
const breadthHello = openRouterBreadthTasks
.find((task) => task.id === "hello-complete")
?.buildPrompt("nonce");
expect(breadthHello).toContain(
"Your first response action must be the paperclip_finish tool call",
);
expect(breadthHello).toContain(
"Do not emit any assistant text, acknowledgement, or preamble before calling it",
);
const nativeAsk = ask?.buildPrompt("nonce");
expect(nativeAsk).toContain("paperclip_finish must be your only tool call");
expect(nativeAsk).toContain(
"never call report_progress or any other tool before or after it",
);
});
it("uses only declared secret references in generated payloads", () => {
expect(
runnerMatrix.every((entry) =>
entry.requiredCredentials.includes(entry.profile.credential),
),
).toBe(true);
expect(
runnerMatrix
.filter((entry) => entry.environment.id === "daytona")
.every((entry) =>
entry.requiredCredentials.includes("DAYTONA_API_KEY"),
),
).toBe(true);
});
it("pins legacy Codex and Claude to their classic CLI engines", () => {
for (const profileId of ["legacy-codex", "legacy-claude"]) {
const execution = runnerMatrix.find(
(candidate) =>
candidate.profile.id === profileId &&
candidate.environment.id === "local",
);
expect(execution).toBeDefined();
expect(
execution!.profile.buildAgent({
environmentId: "11111111-1111-4111-8111-111111111111",
environmentFixtureId: "local",
workspacePath: "/tmp/runner-e2e-workspace",
secretRefs: {
[execution!.profile.credential]: {
type: "secret_ref",
secretId: "22222222-2222-4222-8222-222222222222",
version: "latest",
},
},
executionId: execution!.id,
}),
).toMatchObject({
adapterConfig: {
engine: "cli",
...(profileId === "legacy-codex"
? { extraArgs: ["-c", "features.shell_snapshot=false"] }
: {}),
},
});
}
});
it("binds native Codex automation auth to the encrypted OpenAI secret", () => {
const execution = runnerMatrix.find(
(candidate) =>
candidate.id === "core-compatibility.runner-codex.local.message-marker",
);
expect(execution).toBeDefined();
const secretRef = {
type: "secret_ref" as const,
secretId: "22222222-2222-4222-8222-222222222222",
version: "latest" as const,
};
const agent = execution!.profile.buildAgent({
environmentId: "11111111-1111-4111-8111-111111111111",
environmentFixtureId: "local",
workspacePath: "/tmp/runner-e2e-workspace",
secretRefs: { OPENAI_API_KEY: secretRef },
executionId: execution!.id,
});
expect(agent.adapterConfig).toMatchObject({
env: {
OPENAI_API_KEY: secretRef,
CODEX_API_KEY: secretRef,
},
});
});
it("gives legacy planning agents a direct bounded API recipe", () => {
const task = runnerTasks.find(
(candidate) => candidate.id === "plan-revise-accept",
);
const execution = runnerMatrix.find(
(candidate) =>
candidate.profile.id === "legacy-claude" &&
candidate.environment.id === "local" &&
candidate.task.id === "plan-revise-accept",
);
expect(task).toBeDefined();
expect(execution).toBeDefined();
const agent = execution!.profile.buildAgent({
environmentId: "11111111-1111-4111-8111-111111111111",
environmentFixtureId: "local",
workspacePath: "/tmp/runner-e2e-workspace",
secretRefs: {
ANTHROPIC_API_KEY: {
type: "secret_ref",
secretId: "22222222-2222-4222-8222-222222222222",
version: "latest",
},
},
executionId: execution!.id,
});
expect(agent.adapterConfig).toMatchObject({ maxTurnsPerRun: 24 });
expect(agent.instructionsBundle).toMatchObject({
files: { "AGENTS.md": expect.stringContaining("/interactions") },
});
expect(task!.buildPrompt("nonce")).toContain("request_confirmation");
expect(task!.buildPrompt("nonce")).toContain("baseRevisionId");
expect(task!.buildPrompt("nonce")).toContain(
"do not spell, quote, repeat, announce, or include PAPERCLIP_E2E_PLAN_DONE_nonce",
);
expect(task!.buildPrompt("nonce")).toContain(
'summary:"PAPERCLIP_E2E_PLAN_DONE_nonce"',
);
expect(task!.buildPrompt("nonce")).toContain("first call get_task_context");
expect(task!.buildPrompt("nonce")).toContain(
"identifies the exact revised Plan revision used as the confirmation target as accepted",
);
expect(task!.buildPrompt("nonce")).toContain(
"After that verification succeeds, your immediate next action must be the paperclip_finish tool call",
);
expect(task!.buildPrompt("nonce")).not.toContain(
"trust that inline acceptance",
);
expect(task!.buildPrompt("nonce")).toContain(
"those two tool calls form one indivisible response sequence",
);
expect(task!.buildPrompt("nonce")).toContain(
"Do not emit assistant text, end the response or heartbeat, or stop after write_document alone",
);
expect(task!.buildPrompt("nonce")).toContain(
"one atomic issue PATCH with status `done` and that exact comment",
);
const revisionRequest = task!.buildRevisionRequest?.("nonce");
expect(revisionRequest).toContain("baseRevisionId");
expect(revisionRequest).toContain(
"request_human_input must be your immediate next action",
);
});
it("requires one atomic legacy Ask completion write", () => {
const task = runnerTasks.find(
(candidate) => candidate.id === "ask-question",
);
expect(task).toBeDefined();
const prompt = task!.buildPrompt("nonce");
expect(prompt).toContain(
"make exactly one public-API write containing the marker",
);
expect(prompt).toContain(
'PATCH /api/issues/$PAPERCLIP_TASK_ID with {"status":"done","comment":"E2E_ASK_12_nonce"}',
);
expect(prompt).toContain("Do not POST to /comments");
expect(prompt).toContain("do not PATCH the status separately");
});
it("accepts only complete immutable Daytona digests", () => {
expect(
isImmutableDaytonaImage(
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:${"a".repeat(64)}`,
),
).toBe(true);
expect(
isImmutableDaytonaImage(
"ghcr.io/paperclipai/paperclip-daytona-runner@sha256:REPLACE_ME",
),
).toBe(false);
expect(
isImmutableDaytonaImage(
"ghcr.io/paperclipai/paperclip-daytona-runner:e2e-latest",
),
).toBe(false);
});
});
describe("runner E2E selectors", () => {
it("requires an explicit billable selector", () => {
expect(() => parseRunnerSelectors([])).toThrow(RunnerSelectorError);
});
it("selects dimensions with OR within a dimension and AND across dimensions", () => {
const options = parseRunnerSelectors([
"--profile",
"legacy-codex",
"--profile",
"runner-codex",
"--environment",
"local",
]);
expect(selectRunnerExecutions(options).map((entry) => entry.id)).toEqual([
...runnerMatrix.filter(entry => entry.suite.id === "continuation" && ["legacy-codex", "runner-codex"].includes(entry.profile.id)).map(entry => entry.id),
...runnerMatrix.filter(entry => entry.suite.id === "first-task" && ["legacy-codex", "runner-codex"].includes(entry.profile.id)).map(entry => entry.id),
...runnerMatrix.filter(entry => entry.suite.id === "agent-chat" && ["legacy-codex", "runner-codex"].includes(entry.profile.id)).map(entry => entry.id),
"core-compatibility.legacy-codex.local.message-marker",
"core-compatibility.legacy-codex.local.plan-revise-accept",
"core-compatibility.legacy-codex.local.ask-question",
"core-compatibility.runner-codex.local.message-marker",
"core-compatibility.runner-codex.local.plan-revise-accept",
"core-compatibility.runner-codex.local.ask-question",
"local-session-integrity.legacy-codex.local.structured-question-resume",
"local-session-integrity.legacy-codex.local.structured-question-restart-resume",
"local-session-integrity.runner-codex.local.structured-question-resume",
"local-session-integrity.runner-codex.local.structured-question-restart-resume",
]);
});
it("selects a suite without exploding its environment matrix", () => {
const selected = selectRunnerExecutions(
parseRunnerSelectors(["--suite", "openrouter-model-breadth"]),
);
expect(selected).toHaveLength(10);
expect(
selected.every(
(entry) =>
entry.suite.id === "openrouter-model-breadth" &&
entry.environment.id === "local",
),
).toBe(true);
});
it("combines repeated groups with AND semantics", () => {
const options = parseRunnerSelectors([
"--group",
"native",
"--group",
"daytona",
]);
const selected = selectRunnerExecutions(options);
expect(selected).toHaveLength(13);
expect(
selected.every(
(entry) =>
entry.profile.generation === "native" &&
entry.environment.id === "daytona",
),
).toBe(true);
});
it("rejects unknown groups", () => {
const options = parseRunnerSelectors(["--group", "codex"]);
expect(() => selectRunnerExecutions(options)).toThrow("Unknown group");
});
it("emits one independently schedulable job per scenario", () => {
const jobs = buildMatrixJobs(
selectRunnerExecutions(parseRunnerSelectors(["--all"])),
);
expect(jobs).toHaveLength(167);
expect(jobs.filter((job) => job.needsDaytona)).toHaveLength(23);
expect(jobs.filter((job) => !job.needsDaytona)).toHaveLength(144);
expect(new Set(jobs.map((job) => job.executionId)).size).toBe(167);
expect(
jobs.find(
(job) =>
job.executionId ===
"core-compatibility.runner-acpx-claude.local.plan-revise-accept",
)?.timeoutMinutes,
).toBe(25);
expect(
jobs.find(
(job) =>
job.executionId ===
"local-session-integrity.runner-acpx-codex.local.structured-question-restart-resume",
)?.timeoutMinutes,
).toBe(32);
expect(
jobs.every((job) =>
runnerMatrix.some(
(execution) =>
execution.id === job.executionId &&
execution.profile.credential === job.credentialName,
),
),
).toBe(true);
});
it("validates bounded local parallelism", () => {
expect(
parseRunnerSelectors(["--all", "--max-parallel", "8"]).maxParallel,
).toBe(8);
expect(() =>
parseRunnerSelectors(["--all", "--max-parallel", "0"]),
).toThrow("positive integer");
});
});