mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip manages tasks across persistent agent sessions. > - The full Runner E2E catalog exposed failures in session restoration, tool validation, and test controls. > - These failures prevented valid work from resuming or made a valid interaction fail the test. > - Invalid completion reports also reached finalization before the provider received useful feedback. > - This pull request repairs those boundaries without changing production prompts or approval policy. > - Focused regressions and fresh paid cases verify each fix. ## Linked Issues or Issue Description Follow-up to #13655. Stacked on the trusted worker prerequisite fix in #13674. **What happened?** Read-only skill uploads failed in resumed Daytona sandboxes. Invalid criterion IDs escaped tool validation. A progress event could park a run before its tool response settled. Partial question forms hid required answers. Two test assumptions rejected valid plan keys or failed to navigate an optional question page. **What did you expect to happen?** Resume identical skill bundles, give repairable feedback for malformed completion calls, preserve in-flight tool responses, show all required questions, and test the rendered workflow accurately. **Steps to reproduce** Inspect the failed cases in https://github.com/paperclipai/paperclip/actions/runs/35417932353. Fresh campaigns: https://github.com/paperclipai/paperclip/actions/runs/35444497313 and https://github.com/paperclipai/paperclip/actions/runs/35445327618. The later backup cleanup is tested in https://github.com/paperclipai/paperclip/actions/runs/35446477285. Combined report: https://pages.paperclip.ing/runner-e2e-operational-35444497313/investigation.html. ## What Changed - Compare immutable archives before reusing read-only Daytona bundles. Reject corrupted content and preserve unrelated files. - Validate exact criterion IDs before accepting completion. OpenCode returns a tool error instead of emitting a result that terminates runnerd. - Complete the activity item for rejected OpenCode calls. - Remove retired read-only harness backups without altering live files or following symlinks. A fresh paid rerun exposed this later checkpoint-cleanup failure. - Exclude progress messages from the governed-wait completion boundary. - Reject newly created question forms that omit questions or contradict their stored answer semantics. Keep historical rows readable. - Navigate all rendered question pages and recognize revision-bound descriptive plan keys in the continuation suite. ## Verification - Harness unit suite: 383 tests pass. Harness typecheck passes. - Native session executor and status corpus: 381 tests pass. - Shared question and interaction-service tests: 42 pass; native question bridge and executor: 360 pass. Daytona sync: 21 pass, including foreign-owner archives and corrupted immutable content. - OpenCode driver: 29 tests pass, including wrong, missing, and duplicate criterion IDs followed by a valid retry. - Repository typecheck and build pass. The later OpenCode activity fix also passes its package build. - The latest commit passes all 52 PR checks and Greptile 5/5. The backup-cleanup fix also passes 351 related local tests and server typecheck. Local full-suite coverage completed across runs. adapter-auth-signal-routes and pipelines-routes encountered transient socket resets; both pass on retry, and all remaining 24 serialized files pass. Paid reruns are complete: 27 of 29 unique cases pass using the latest recording per case. Both Daytona controller-restart cases still fail with runner_state_identity_mismatch; the report describes this remaining runtime issue. Eight affected cells need #13674 on master before their rerun. ## Risks Creation rejects inconsistent dual question representations but does not change historical records. Immutable bundle comparison must verify bytes before skipping extraction. Completion feedback must use the contract bound to the current run. Durable suspension and approval checks remain enforced. Production prompts are unchanged. ## Model Used OpenAI GPT-6 via Codex, with repository inspection, code editing, and test execution. The exact API model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
297 lines
11 KiB
TypeScript
297 lines
11 KiB
TypeScript
import { prepareLegacyContinuationSkill } from "./continuation-fixtures.js";
|
|
import { captureFirstTaskAttachments } from "./first-task-attachments.js";
|
|
import { answerableRuntimeRunIds, isSingleClaudeQuestion } from "./runtime-question-readiness.js";
|
|
import { expect, type Page } from "@playwright/test";
|
|
import path from "node:path";
|
|
import { continuationAnswerCommitted, continuationInitialReady } from "./continuation-readiness.js";
|
|
import { captureLoadedContinuation } from "./continuation-screenshot.js";
|
|
import { seedContinuationContext } from "./continuation-workspace.js";
|
|
import { pollUntil, type RunnerApi } from "./api.js";
|
|
import {
|
|
chatQuestionPresentation,
|
|
sendChatMessage,
|
|
collectChatRunEvidence,
|
|
type ChatRun,
|
|
} from "./chat-flow.js";
|
|
import {
|
|
continuationScenario,
|
|
continuationScreenshotFile,
|
|
} from "./continuation-cases.js";
|
|
import {
|
|
gradeContinuation,
|
|
isContinuationPlan,
|
|
type ContinuationCheckpoint,
|
|
} from "./continuation-scoring.js";
|
|
import { createTaskThroughUi } from "./user-actions.js";
|
|
import type { LiveFixtureValues } from "./live-fixtures.js";
|
|
import type { MatrixExecution } from "./types.js";
|
|
type Row = Record<string, any>;
|
|
|
|
export async function runContinuationFlow(input: {
|
|
page: Page;
|
|
api: RunnerApi;
|
|
fixtures: LiveFixtureValues;
|
|
execution: MatrixExecution;
|
|
nonce: string;
|
|
secrets: readonly string[];
|
|
workspacePath: string;
|
|
deadlineAt: number;
|
|
restart(): Promise<void>;
|
|
observe(
|
|
issue: any,
|
|
runs: any[],
|
|
checks: ReturnType<typeof gradeContinuation>,
|
|
): void;
|
|
capture(id: string, label: string, file: string): Promise<void>;
|
|
evidence(name: string, value: unknown): Promise<void>;
|
|
}) {
|
|
const { page, api, fixtures, execution } = input;
|
|
const scenario = continuationScenario(execution.task.id, input.nonce);
|
|
const checkpoints: ContinuationCheckpoint[] = [];
|
|
let issue: Row | undefined;
|
|
let runs: Row[] = [];
|
|
let checks: ReturnType<typeof gradeContinuation> = [];
|
|
const tasksPath = `/api/companies/${fixtures.company.id}/issues?limit=100`;
|
|
async function refresh() {
|
|
issue = await api.get<Row>(`/api/issues/${issue!.id}`);
|
|
const listed = await api.get<Row[]>(
|
|
`/api/companies/${fixtures.company.id}/heartbeat-runs?limit=100`,
|
|
);
|
|
runs = await Promise.all(
|
|
listed.map((r) => api.get<Row>(`/api/heartbeat-runs/${r.id}`)),
|
|
);
|
|
input.observe(issue, runs, checks);
|
|
return { issue, runs };
|
|
}
|
|
let pausedRuntimeRunIds = new Set<string>();
|
|
async function settle(prior: Set<string>, requireQuestion = false, answeredInteractionId?: string) {
|
|
let stable = "";
|
|
const previousPaused = pausedRuntimeRunIds;
|
|
await pollUntil({
|
|
label: `continuation ${scenario.id} settled`,
|
|
deadlineAt: input.deadlineAt,
|
|
intervalMs: 1000,
|
|
load: async () => ({
|
|
...await refresh(),
|
|
interactions: await api.get<Row[]>(`/api/issues/${issue!.id}/interactions`),
|
|
}),
|
|
accept: (state) => {
|
|
const paused = answerableRuntimeRunIds(state.interactions);
|
|
const idle =
|
|
continuationAnswerCommitted(state.interactions, answeredInteractionId) &&
|
|
state.runs.some((r) => !prior.has(r.id) || previousPaused.has(r.id)) &&
|
|
state.runs.every((r) => ["succeeded", "failed", "timed_out", "cancelled"].includes(r.status) ||
|
|
(r.status === "running" && paused.has(r.id))) &&
|
|
!state.issue.scheduledRetry &&
|
|
!state.issue.activeRecoveryAction &&
|
|
(!requireQuestion || continuationInitialReady(state.interactions));
|
|
const key = idle
|
|
? state.runs.map((r) => `${r.id}:${r.status}`).join()
|
|
: "";
|
|
const ready = !!key && key === stable;
|
|
stable = key;
|
|
if (ready) pausedRuntimeRunIds = paused;
|
|
return ready;
|
|
},
|
|
reject: (state) =>
|
|
state.runs.length > 12
|
|
? "Bounded continuation run count exceeded"
|
|
: state.runs.some((r) =>
|
|
["failed", "timed_out", "cancelled"].includes(r.status),
|
|
)
|
|
? `Provider execution failed: ${state.runs
|
|
.filter((r) => r.status !== "succeeded")
|
|
.map(
|
|
(r) => `${r.id} ${r.errorCode ?? r.status}: ${r.error ?? ""}`,
|
|
)
|
|
.join("; ")}`
|
|
: undefined,
|
|
});
|
|
}
|
|
async function open() {
|
|
await page.goto(
|
|
`/${fixtures.company.issuePrefix}/issues/${issue!.identifier ?? issue!.id}`,
|
|
{ waitUntil: "domcontentloaded" },
|
|
);
|
|
}
|
|
async function snapshot(phase: ContinuationCheckpoint["phase"]) {
|
|
const [tasks, summaries, comments, interactions, attachments] =
|
|
await Promise.all([
|
|
api.get<Row[]>(tasksPath),
|
|
api.get<Row[]>(`/api/issues/${issue!.id}/documents`),
|
|
api.get<Row[]>(`/api/issues/${issue!.id}/comments?order=asc`),
|
|
api.get<Row[]>(`/api/issues/${issue!.id}/interactions`),
|
|
captureFirstTaskAttachments(api, [{ id: issue!.id }], input.secrets),
|
|
]);
|
|
const documents = await Promise.all(
|
|
summaries.map((d) =>
|
|
api.get<Row>(
|
|
`/api/issues/${issue!.id}/documents/${encodeURIComponent(d.key)}`,
|
|
),
|
|
),
|
|
);
|
|
checkpoints.push({
|
|
phase,
|
|
issue: issue as ContinuationCheckpoint["issue"],
|
|
children: tasks.filter(
|
|
(t) => t.parentId === issue!.id,
|
|
) as ContinuationCheckpoint["children"],
|
|
documents: documents as ContinuationCheckpoint["documents"],
|
|
comments,
|
|
interactions,
|
|
attachments,
|
|
runs: [...runs] as ContinuationCheckpoint["runs"],
|
|
});
|
|
await input.evidence("continuation.json", {
|
|
...scenario,
|
|
checkpoints,
|
|
checks,
|
|
});
|
|
await input.evidence("api-state.json", checkpoints.at(-1));
|
|
await open();
|
|
await captureLoadedContinuation(page, String(issue!.title), () => input.capture(
|
|
phase,
|
|
`Continuation: ${phase}`,
|
|
continuationScreenshotFile(phase),
|
|
));
|
|
}
|
|
async function answer(choice?: string) {
|
|
const interactions = await api.get<Row[]>(
|
|
`/api/issues/${issue!.id}/interactions`,
|
|
);
|
|
const questions = interactions.filter(
|
|
(i) => i.kind === "ask_user_questions" && i.status === "pending",
|
|
);
|
|
expect(questions, "one real question must be shown").toHaveLength(1);
|
|
const set = chatQuestionPresentation(questions[0].payload);
|
|
if (scenario.id === "provider-question-bridge") {
|
|
expect(isSingleClaudeQuestion(set.questions), "one choice question with only the optional provider Other field").toBe(true);
|
|
} else expect(set.questions, "ask only the requested next question").toHaveLength(1);
|
|
const before = new Set(runs.map((r) => r.id));
|
|
if (choice) {
|
|
expect(set.questions[0].answerMode, "choices must use radio controls").toBe("single_select");
|
|
const options = questions[0].payload.questionSet?.questions[0]?.options
|
|
?? questions[0].payload.questions[0]?.options ?? [];
|
|
expect(new Set(options.map((o: Row) => String(o.label).trim().toLowerCase())).size).toBeGreaterThanOrEqual(2);
|
|
await page.getByRole("radio", { name: new RegExp(`^${choice}\\b`, "i") }).last().click();
|
|
} else {
|
|
expect(set.questions[0].answerMode, "open answers must render a text field, not a choice question").toBe("text");
|
|
await page.getByTestId("question-text-answer-composer").last()
|
|
.locator('[contenteditable="true"],textarea').first().fill(scenario.answer);
|
|
}
|
|
// Claude may add a separate optional Other field after its choice page.
|
|
// Navigate every rendered page before submitting; do not invent an answer.
|
|
for (let index = 1; index < set.questions.length; index += 1) {
|
|
await page.getByRole("button", { name: "Next", exact: true }).last().click();
|
|
}
|
|
await page
|
|
.getByRole("button", {
|
|
name: set.submitLabel ?? "Submit answers",
|
|
exact: true,
|
|
})
|
|
.last()
|
|
.click();
|
|
await settle(before, false, questions[0].id);
|
|
}
|
|
async function reply(body: string) {
|
|
const before = new Set(runs.map((r) => r.id));
|
|
await sendChatMessage(page, body);
|
|
await settle(before);
|
|
}
|
|
function assertWaiting() {
|
|
const c = checkpoints.at(-1)!;
|
|
expect(
|
|
c.documents.filter((d) => !isContinuationPlan(d, c)),
|
|
"no deliverable before authorization",
|
|
).toHaveLength(0);
|
|
expect(c.attachments, "no attachment before authorization").toHaveLength(0);
|
|
expect(c.issue.status, "waiting is not complete").not.toBe("done");
|
|
}
|
|
try {
|
|
if (execution.profile.generation === "legacy") await prepareLegacyContinuationSkill(api, fixtures.company.id, fixtures.agent.id);
|
|
await api.patch("/api/instance/settings/experimental", {
|
|
enableClassicTaskInterface: false,
|
|
});
|
|
await createTaskThroughUi({
|
|
page,
|
|
issuePrefix: fixtures.company.issuePrefix!,
|
|
agentName: fixtures.agent.name,
|
|
title: execution.task.buildTitle(input.nonce),
|
|
prompt: scenario.prompt,
|
|
workMode: "standard",
|
|
});
|
|
issue = await pollUntil({
|
|
label: "continuation task created",
|
|
deadlineAt: input.deadlineAt,
|
|
load: async () =>
|
|
(await api.get<Row[]>(tasksPath)).find(
|
|
(t) => t.title === execution.task.buildTitle(input.nonce),
|
|
),
|
|
accept: Boolean,
|
|
});
|
|
if (!issue) throw new Error("Missing continuation task");
|
|
await settle(new Set(), scenario.id !== "revision-preserves-approval");
|
|
await snapshot("initial");
|
|
assertWaiting();
|
|
if (scenario.id === "untrusted-evidence") {
|
|
const parentRun = runs.find((r) => r.contextSnapshot?.issueId === issue!.id);
|
|
await seedContinuationContext({
|
|
isolatedRoot: path.dirname(input.workspacePath),
|
|
recordedCwd: parentRun?.contextSnapshot?.paperclipWorkspace?.cwd,
|
|
body: scenario.context,
|
|
});
|
|
}
|
|
if (scenario.id === "completed-action-resume") {
|
|
await input.restart();
|
|
await open();
|
|
}
|
|
if (scenario.id === "question-tool-documentation") {
|
|
await answer("Afternoon");
|
|
await snapshot("answered");
|
|
assertWaiting();
|
|
await answer();
|
|
} else if (scenario.id === "provider-question-bridge") await answer(scenario.marker);
|
|
else if (scenario.id === "revision-preserves-approval")
|
|
await reply(scenario.revision);
|
|
else await answer();
|
|
if (scenario.gate) {
|
|
await snapshot(
|
|
scenario.id === "revision-preserves-approval" ? "revised" : "answered",
|
|
);
|
|
assertWaiting();
|
|
await reply(scenario.approval);
|
|
}
|
|
await snapshot("final");
|
|
} finally {
|
|
checks = gradeContinuation({
|
|
...scenario,
|
|
checkpoints,
|
|
runtimeMode: execution.profile.expectedRuntimeMode,
|
|
});
|
|
if (issue) input.observe(issue, runs, checks);
|
|
await input.evidence("continuation.json", {
|
|
...scenario,
|
|
checkpoints,
|
|
checks,
|
|
});
|
|
await input.evidence(
|
|
"continuation-run-evidence.json",
|
|
await Promise.all(
|
|
runs.map(async (r) => {
|
|
try {
|
|
return await collectChatRunEvidence(api, r as ChatRun);
|
|
} catch (error) {
|
|
return { runId: r.id, evidenceCaptureError: String(error) };
|
|
}
|
|
}),
|
|
),
|
|
);
|
|
}
|
|
const failures = checks.filter((c) => !c.passed);
|
|
if (failures.length)
|
|
throw new Error(
|
|
`Continuation matcher failures: ${failures.map((c) => `${c.id}: ${c.detail}`).join("; ")}`,
|
|
);
|
|
return { issue: issue!, runs, checks };
|
|
}
|