Files
PaperClipAI/tests/runner-e2e/continuation-flow.ts
T
DottaandPaperclip 24a15c5afa test(evals): verify durable waiting across continuation checkpoints (#15564)
## Thinking Path

> - Paperclip manages agents and durable tasks across provider runs.
> - Human questions must preserve task ownership and stop work until a
real answer arrives.
> - The continuation eval checks the final work, but its lifecycle
oracle misses several broken waiting states.
> - A correct final answer can hide a stale execution lock or a lost
intermediate answer receipt.
> - This pull request checks each wait and retains every question and
run identity.
> - A source audit records which waiting operations belong to native
runtimes and which still require legacy API calls.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The existing Product E2E continuation oracle for human question and
approval journeys.

**Current behavior**

A pending interaction can pass the waiting check even when its task has
the wrong status, a stale lock, or a retry. Only the first pending
question and first run receipts are checked at the final checkpoint.

**Proposed behavior**

Require a healthy wait on the same task and assignee. Preserve all
intermediate run receipts and every pending question's answered
identity. Allow a paused native provider question only when its pending
runtime request identifies the running native run and execution lock.

Related: #15548 and #15554 cover earlier bookkeeping slices. #15544
changes production continuation summaries; this PR changes the eval
oracle and does not overlap that fix.

## What Changed

- Check waiting state, execution ownership, and all answer/run
identities in the existing lifecycle oracle.
- Add negative calibrations for broken records and positive coverage for
both semantic waits and paused provider questions.
- Retain task activity at each checkpoint for inspection of successful
persisted mutations.
- Verify 1,000-cent company and agent budget limits before continuation
work. Disable automatic cell rerolls.
- Document runtime ownership, existing coverage, and the remaining
instruction decision. Keep production instructions unchanged.

## Verification

- `pnpm test:e2e:runner:unit`: 1,867 Vitest tests pass, one skips; 128
Node tests pass.
- `pnpm test:e2e:runner:typecheck`: passes.
- `pnpm -r typecheck`: passes.
- `pnpm build`: passes.
- Negative calibration: 27 added cases fail against the prior oracle and
pass with these checks.
- Four existing local continuation cells are selected for a separate
bounded live canary. Live results are pending; no behavioral pass is
claimed here.
- The full local `pnpm test:run` suite was not repeated because embedded
PostgreSQL was unavailable in the preceding workspace verification.
Required Linux PR CI must pass before readiness.

## Risks

The stronger oracle can expose existing product or fixture defects. A
paused provider question and a terminal semantic wait have distinct
valid states. The four-cell canary does not qualify approval/review,
dependency unblock, crash races, remote execution, or general task
quality. Task activity records successful persisted writes, not failed
API attempts; repeated progress comments are not automatically defects.
Historical eval grades remain unchanged. No production scheduling,
prompt, tool, schema or migration changes.

## Model Used

OpenAI Codex, GPT-6-based, with tool use and code execution. The exact
serving model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 08:45:36 -05:00

326 lines
13 KiB
TypeScript

import { gradeLifecycleBaseline, type LifecycleCheckpoint } from "./lifecycle-baseline.js";
import { lifecycleLiveCase, lifecycleLiveContinuation, gradeLifecycleNarrative } from "./lifecycle-live-cases.js";
import { prepareLegacyContinuationSkill, prepareContinuationBudget } from "./continuation-fixtures.js";
import { captureFirstTaskAttachments } from "./first-task-attachments.js";
import { answerableRuntimeRunIds, isSingleClaudeQuestion } from "./runtime-question-readiness.js";
import { expect, type Page } from "@playwright/test";
import path from "node:path";
import { continuationAnswerCommitted, continuationInitialReady, continuationCheckpointReady } from "./continuation-readiness.js";
import { captureLoadedContinuation } from "./continuation-screenshot.js";
import { seedContinuationContext } from "./continuation-workspace.js";
import { pollUntil, type RunnerApi } from "./api.js";
import {
chatQuestionPresentation,
sendChatMessage,
collectChatRunEvidence,
type ChatRun,
} from "./chat-flow.js";
import {
continuationScenario,
continuationScreenshotFile,
} from "./continuation-cases.js";
import {
gradeContinuation,
isContinuationPlan,
type ContinuationCheckpoint,
} from "./continuation-scoring.js";
import { createTaskThroughUi } from "./user-actions.js";
import type { LiveFixtureValues } from "./live-fixtures.js";
import type { MatrixExecution } from "./types.js";
type Row = Record<string, any>;
export async function runContinuationFlow(input: {
page: Page;
api: RunnerApi;
fixtures: LiveFixtureValues;
execution: MatrixExecution;
nonce: string;
secrets: readonly string[];
workspacePath: string;
deadlineAt: number;
restart(): Promise<void>;
observe(
issue: any,
runs: any[],
checks: ReturnType<typeof gradeContinuation>,
): void;
capture(id: string, label: string, file: string): Promise<void>;
evidence(name: string, value: unknown): Promise<void>;
}) {
const { page, api, fixtures, execution } = input;
const lifecycleProbe = lifecycleLiveCase(execution.task.id);
const scenario = lifecycleProbe
? lifecycleLiveContinuation(execution.task.id, input.nonce)
: continuationScenario(execution.task.id, input.nonce);
const checkpoints: LifecycleCheckpoint[] = [];
let budgetGuard: Awaited<ReturnType<typeof prepareContinuationBudget>> | undefined;
let issue: Row | undefined;
let runs: Row[] = [];
let checks: ReturnType<typeof gradeContinuation> = [];
const tasksPath = `/api/companies/${fixtures.company.id}/issues?limit=100`;
async function refresh() {
issue = await api.get<Row>(`/api/issues/${issue!.id}`);
const listed = await api.get<Row[]>(
`/api/companies/${fixtures.company.id}/heartbeat-runs?limit=100`,
);
runs = await Promise.all(
listed.map((r) => api.get<Row>(`/api/heartbeat-runs/${r.id}`)),
);
input.observe(issue, runs, checks);
return { issue, runs };
}
let pausedRuntimeRunIds = new Set<string>();
async function settle(prior: Set<string>, requireQuestion = false, answeredInteractionId?: string) {
let stable = "";
const previousPaused = pausedRuntimeRunIds;
await pollUntil({
label: `continuation ${scenario.id} settled`,
deadlineAt: input.deadlineAt,
intervalMs: 1000,
load: async () => ({
...await refresh(),
interactions: await api.get<Row[]>(`/api/issues/${issue!.id}/interactions`),
}),
accept: (state) => {
const paused = answerableRuntimeRunIds(state.interactions);
const idle =
continuationAnswerCommitted(state.interactions, answeredInteractionId) &&
continuationCheckpointReady({
issue: { status: state.issue.status, executionRunId: state.issue.executionRunId },
runs: state.runs.map(r => ({ id: r.id, status: r.status, runtimeMode: r.runtimeMode })),
interactions: state.interactions,
}) &&
state.runs.some((r) => !prior.has(r.id) || previousPaused.has(r.id)) &&
state.runs.every((r) => ["succeeded", "failed", "timed_out", "cancelled"].includes(r.status) ||
(r.status === "running" && paused.has(r.id))) &&
!state.issue.scheduledRetry &&
!state.issue.activeRecoveryAction &&
(!requireQuestion || continuationInitialReady(state.interactions));
const key = idle
? JSON.stringify([state.issue.status, state.issue.executionRunId,
state.runs.map(r => [r.id, r.status]), state.interactions.map(i => [i.id, i.status])])
: "";
const ready = !!key && key === stable;
stable = key;
if (ready) pausedRuntimeRunIds = paused;
return ready;
},
reject: (state) =>
state.runs.length > 12
? "Bounded continuation run count exceeded"
: state.runs.some((r) =>
["failed", "timed_out", "cancelled"].includes(r.status),
)
? `Provider execution failed: ${state.runs
.filter((r) => r.status !== "succeeded")
.map(
(r) => `${r.id} ${r.errorCode ?? r.status}: ${r.error ?? ""}`,
)
.join("; ")}`
: undefined,
});
}
async function open() {
await page.goto(
`/${fixtures.company.issuePrefix}/issues/${issue!.identifier ?? issue!.id}`,
{ waitUntil: "domcontentloaded" },
);
}
async function snapshot(phase: ContinuationCheckpoint["phase"]) {
const [tasks, summaries, comments, interactions, attachments, activity] =
await Promise.all([
api.get<Row[]>(tasksPath),
api.get<Row[]>(`/api/issues/${issue!.id}/documents`),
api.get<Row[]>(`/api/issues/${issue!.id}/comments?order=asc`),
api.get<Row[]>(`/api/issues/${issue!.id}/interactions`),
captureFirstTaskAttachments(api, [{ id: issue!.id }], input.secrets),
api.get<Row[]>(`/api/issues/${issue!.id}/activity`),
]);
const documents = await Promise.all(
summaries.map((d) =>
api.get<Row>(
`/api/issues/${issue!.id}/documents/${encodeURIComponent(d.key)}`,
),
),
);
checkpoints.push({
phase,
lifecycle: {
executionRunId: issue!.executionRunId ?? null,
scheduledRetry: issue!.scheduledRetry ?? null,
activeRecoveryAction: issue!.activeRecoveryAction ?? null,
monitorNextCheckAt: issue!.monitorNextCheckAt ?? null,
},
issue: issue as ContinuationCheckpoint["issue"],
children: tasks.filter(
(t) => t.parentId === issue!.id,
) as ContinuationCheckpoint["children"],
documents: documents as ContinuationCheckpoint["documents"],
comments,
activity,
interactions,
attachments,
runs: [...runs] as ContinuationCheckpoint["runs"],
});
await input.evidence("continuation.json", {
...scenario,
budgetGuard,
checkpoints,
checks,
});
await input.evidence("api-state.json", checkpoints.at(-1));
await open();
await captureLoadedContinuation(page, String(issue!.title), () => input.capture(
phase,
`Continuation: ${phase}`,
continuationScreenshotFile(phase),
));
}
async function answer(choice?: string) {
const interactions = await api.get<Row[]>(
`/api/issues/${issue!.id}/interactions`,
);
const questions = interactions.filter(
(i) => i.kind === "ask_user_questions" && i.status === "pending",
);
expect(questions, "one real question must be shown").toHaveLength(1);
const set = chatQuestionPresentation(questions[0].payload);
if (scenario.id === "provider-question-bridge") {
expect(isSingleClaudeQuestion(set.questions), "one choice question with only the optional provider Other field").toBe(true);
} else expect(set.questions, "ask only the requested next question").toHaveLength(1);
const before = new Set(runs.map((r) => r.id));
if (choice) {
expect(set.questions[0].answerMode, "choices must use radio controls").toBe("single_select");
const options = questions[0].payload.questionSet?.questions[0]?.options
?? questions[0].payload.questions[0]?.options ?? [];
expect(new Set(options.map((o: Row) => String(o.label).trim().toLowerCase())).size).toBeGreaterThanOrEqual(2);
await page.getByRole("radio", { name: new RegExp(`^${choice}\\b`, "i") }).last().click();
} else {
expect(set.questions[0].answerMode, "open answers must render a text field, not a choice question").toBe("text");
await page.getByTestId("question-text-answer-composer").last()
.locator('[contenteditable="true"],textarea').first().fill(scenario.answer);
}
// Claude may add a separate optional Other field after its choice page.
// Navigate every rendered page before submitting; do not invent an answer.
for (let index = 1; index < set.questions.length; index += 1) {
await page.getByRole("button", { name: "Next", exact: true }).last().click();
}
await page
.getByRole("button", {
name: set.submitLabel ?? "Submit answers",
exact: true,
})
.last()
.click();
await settle(before, false, questions[0].id);
}
async function reply(body: string) {
const before = new Set(runs.map((r) => r.id));
await sendChatMessage(page, body);
await settle(before);
}
function assertWaiting() {
const c = checkpoints.at(-1)!;
expect(
c.documents.filter((d) => !isContinuationPlan(d, c)),
"no deliverable before authorization",
).toHaveLength(0);
expect(c.attachments, "no attachment before authorization").toHaveLength(0);
expect(c.issue.status, "waiting is not complete").not.toBe("done");
}
try {
budgetGuard = await prepareContinuationBudget(api, fixtures.company.id, fixtures.agent.id);
if (execution.profile.generation === "legacy") await prepareLegacyContinuationSkill(api, fixtures.company.id, fixtures.agent.id);
await api.patch("/api/instance/settings/experimental", {
enableClassicTaskInterface: false,
});
const createdTask = await createTaskThroughUi({
page,
issuePrefix: fixtures.company.issuePrefix!,
agentName: fixtures.agent.name,
title: execution.task.buildTitle(input.nonce),
prompt: scenario.prompt,
workMode: "standard",
});
issue = await pollUntil({
label: "continuation task created",
deadlineAt: input.deadlineAt,
load: async () =>
(await api.get<Row[]>(tasksPath)).find(
(t) => t.id === createdTask.issueId,
),
accept: Boolean,
});
if (!issue) throw new Error("Missing continuation task");
await settle(new Set(), scenario.id !== "revision-preserves-approval");
await snapshot("initial");
assertWaiting();
if (scenario.id === "untrusted-evidence") {
const parentRun = runs.find((r) => r.contextSnapshot?.issueId === issue!.id);
await seedContinuationContext({
isolatedRoot: path.dirname(input.workspacePath),
recordedCwd: parentRun?.contextSnapshot?.paperclipWorkspace?.cwd,
body: scenario.context,
});
}
if (scenario.id === "completed-action-resume") {
await input.restart();
await open();
}
if (scenario.id === "question-tool-documentation") {
await answer("Afternoon");
await snapshot("answered");
assertWaiting();
await answer();
} else if (scenario.id === "provider-question-bridge") await answer(scenario.marker);
else if (scenario.id === "revision-preserves-approval")
await reply(scenario.revision);
else await answer();
if (scenario.gate) {
await snapshot(
scenario.id === "revision-preserves-approval" ? "revised" : "answered",
);
assertWaiting();
await reply(scenario.approval);
}
await snapshot("final");
} finally {
checks = gradeContinuation({
...scenario,
checkpoints,
runtimeMode: execution.profile.expectedRuntimeMode,
});
checks.push(...gradeLifecycleBaseline(checkpoints));
if (lifecycleProbe) checks.push(gradeLifecycleNarrative({
narrative: lifecycleProbe.narrative,
agentId: fixtures.agent.id,
initial: checkpoints.find(c => c.phase === "initial"),
}));
if (issue) input.observe(issue, runs, checks);
await input.evidence("continuation.json", {
...scenario,
budgetGuard,
checkpoints,
checks,
});
await input.evidence(
"continuation-run-evidence.json",
await Promise.all(
runs.map(async (r) => {
try {
return await collectChatRunEvidence(api, r as ChatRun);
} catch (error) {
return { runId: r.id, evidenceCaptureError: String(error) };
}
}),
),
);
}
const failures = checks.filter((c) => !c.passed);
if (failures.length)
throw new Error(
`Continuation matcher failures: ${failures.map((c) => `${c.id}: ${c.detail}`).join("; ")}`,
);
return { issue: issue!, runs, checks };
}