Files
PaperClipAI/tests/runner-e2e/continuation.test.ts
DottaandPaperclip c1f6c3310a fix(runner): repair catalog runtime and grading boundaries (#13676)
## Thinking Path

> - Paperclip manages tasks across persistent agent sessions.
> - The full Runner E2E catalog exposed failures in session restoration,
tool validation, and test controls.
> - These failures prevented valid work from resuming or made a valid
interaction fail the test.
> - Invalid completion reports also reached finalization before the
provider received useful feedback.
> - This pull request repairs those boundaries without changing
production prompts or approval policy.
> - Focused regressions and fresh paid cases verify each fix.

## Linked Issues or Issue Description

Follow-up to #13655. Stacked on the trusted worker prerequisite fix in
#13674.

**What happened?**

Read-only skill uploads failed in resumed Daytona sandboxes. Invalid
criterion IDs escaped tool validation. A progress event could park a run
before its tool response settled. Partial question forms hid required
answers. Two test assumptions rejected valid plan keys or failed to
navigate an optional question page.

**What did you expect to happen?**

Resume identical skill bundles, give repairable feedback for malformed
completion calls, preserve in-flight tool responses, show all required
questions, and test the rendered workflow accurately.

**Steps to reproduce**

Inspect the failed cases in
https://github.com/paperclipai/paperclip/actions/runs/35417932353. Fresh
campaigns:
https://github.com/paperclipai/paperclip/actions/runs/35444497313 and
https://github.com/paperclipai/paperclip/actions/runs/35445327618. The
later backup cleanup is tested in
https://github.com/paperclipai/paperclip/actions/runs/35446477285.
Combined report:
https://pages.paperclip.ing/runner-e2e-operational-35444497313/investigation.html.

## What Changed

- Compare immutable archives before reusing read-only Daytona bundles.
Reject corrupted content and preserve unrelated files.
- Validate exact criterion IDs before accepting completion. OpenCode
returns a tool error instead of emitting a result that terminates
runnerd.
- Complete the activity item for rejected OpenCode calls.
- Remove retired read-only harness backups without altering live files
or following symlinks. A fresh paid rerun exposed this later
checkpoint-cleanup failure.
- Exclude progress messages from the governed-wait completion boundary.
- Reject newly created question forms that omit questions or contradict
their stored answer semantics. Keep historical rows readable.
- Navigate all rendered question pages and recognize revision-bound
descriptive plan keys in the continuation suite.

## Verification

- Harness unit suite: 383 tests pass. Harness typecheck passes.
- Native session executor and status corpus: 381 tests pass.
- Shared question and interaction-service tests: 42 pass; native
question bridge and executor: 360 pass. Daytona sync: 21 pass, including
foreign-owner archives and corrupted immutable content.
- OpenCode driver: 29 tests pass, including wrong, missing, and
duplicate criterion IDs followed by a valid retry.
- Repository typecheck and build pass. The later OpenCode activity fix
also passes its package build.
- The latest commit passes all 52 PR checks and Greptile 5/5. The
backup-cleanup fix also passes 351 related local tests and server
typecheck. Local full-suite coverage completed across runs.
adapter-auth-signal-routes and pipelines-routes encountered transient
socket resets; both pass on retry, and all remaining 24 serialized files
pass. Paid reruns are complete: 27 of 29 unique cases pass using the
latest recording per case. Both Daytona controller-restart cases still
fail with runner_state_identity_mismatch; the report describes this
remaining runtime issue. Eight affected cells need #13674 on master
before their rerun.

## Risks

Creation rejects inconsistent dual question representations but does not
change historical records. Immutable bundle comparison must verify bytes
before skipping extraction. Completion feedback must use the contract
bound to the current run. Durable suspension and approval checks remain
enforced. Production prompts are unchanged.

## Model Used

OpenAI GPT-6 via Codex, with repository inspection, code editing, and
test execution. The exact API model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 09:27:48 -05:00

241 lines
10 KiB
TypeScript

import { createHash } from "node:crypto";
import { mkdtemp, mkdir, writeFile, readFile, realpath, symlink, rm } from "node:fs/promises";
import os from "node:os";
import path from "node:path";
import { seedContinuationContext } from "./continuation-workspace.js";
import { packageEvidence } from "./evidence.js";
import { continuationScreenshotFile } from "./continuation-cases.js";
import { describe, expect, it } from "vitest";
import {
CONTINUATION_CASES,
continuationScenario,
} from "./continuation-cases.js";
import {
gradeContinuation,
type ContinuationCheckpoint,
} from "./continuation-scoring.js";
import { runnerMatrix } from "./catalog.js";
function recording(id = "clarification-not-approval") {
const scenario = continuationScenario(id, "nonce");
const initial: ContinuationCheckpoint = {
phase: "initial",
issue: { id: "parent", status: "in_review" },
children:
id === "completed-action-resume"
? [
{
id: "child",
title: scenario.childTitle,
status: "done",
assigneeAgentId: "agent",
},
]
: [],
documents: [],
attachments: [],
comments: [],
interactions: [],
runs: [{ id: "first", status: "succeeded", runtimeMode: "native" }],
};
const answered = { ...structuredClone(initial), phase: "answered" as const };
const final: ContinuationCheckpoint = {
...structuredClone(initial),
phase: "final",
issue: { id: "parent", status: "done" },
documents: [
{
key: "output",
body: `Welcome ${scenario.marker} at ${scenario.fact}.`,
latestRevisionId: "revision",
},
],
runs: [
...initial.runs,
{ id: "second", status: "succeeded", runtimeMode: "native" },
],
};
return {
...scenario,
runtimeMode: "native",
checkpoints: scenario.gate ? [initial, answered, final] : [initial, final],
};
}
const failures = (r: ReturnType<typeof recording>) =>
gradeContinuation(r)
.filter((c) => !c.passed)
.map((c) => c.id);
describe("continuation behavioral evaluation", () => {
it("registers all five cases for both runtime generations and providers", () => {
const matrix = runnerMatrix.filter((c) => c.suite.id === "continuation");
expect(matrix).toHaveLength(23);
expect(new Set(matrix.map((c) => c.profile.id))).toEqual(
new Set([
"legacy-codex",
"legacy-claude",
"runner-codex",
"runner-acpx-claude",
]),
);
expect(matrix.every((c) => !c.suite.manualOnly)).toBe(true);
});
it.each(CONTINUATION_CASES.filter(id => !["question-tool-documentation", "provider-question-bridge"].includes(id)))("accepts a complete %s recording", (id) =>
expect(failures(recording(id))).toEqual([]),
);
it("accepts a revision-bound descriptive plan key without counting it as final output", () => {
const r = recording("revision-preserves-approval");
for (const c of r.checkpoints) {
c.documents.push({ key: "welcome-note-plan", body: "After approval, write the note.", latestRevisionId: "plan-v1" });
c.interactions.push({ kind: "request_confirmation", payload: { target: { type: "issue_document", issueId: "parent", key: "welcome-note-plan", revisionId: "plan-v1" } } });
}
expect(failures(r)).toEqual([]);
r.checkpoints[0].documents.at(-1)!.latestRevisionId = "unapproved-v2";
expect(failures(r)).toContain("initial.no-premature-output");
});
it("does not treat an arbitrary deliverable targeted for confirmation as a plan", () => {
const r = recording();
r.checkpoints[0].documents.push({ key: "welcome-note", body: r.marker, latestRevisionId: "v1" });
r.checkpoints[0].interactions.push({ kind: "request_confirmation", payload: { target: { type: "issue_document", issueId: "parent", key: "welcome-note", revisionId: "v1" } } });
expect(failures(r)).toContain("initial.no-premature-output");
});
it("fails premature output even when the final result is correct", () => {
const r = recording();
r.checkpoints[0].documents.push({ key: "output", body: r.marker });
expect(failures(r)).toContain("initial.no-premature-output");
});
it("fails mistaken approval from clarification", () => {
const r = recording();
r.checkpoints[1].issue.status = "done";
expect(failures(r)).toContain("answered.no-premature-output");
});
it("fails scope revision that silently drops approval", () => {
const r = recording("revision-preserves-approval");
r.checkpoints.splice(1, 1);
expect(failures(r)).toContain("approval-boundary-recorded");
});
it.each(["old", "injected"] as const)(
"fails output containing the %s scope",
(key) => {
const r = recording("untrusted-evidence");
r.checkpoints.at(-1)!.documents[0].body += r[key];
expect(failures(r)).toContain("updated-output");
},
);
it("fails duplicate child creation and replacement with an identical title", () => {
const r = recording("completed-action-resume");
r.checkpoints
.at(-1)!
.children.push({ ...r.checkpoints[0].children[0], id: "duplicate" });
expect(failures(r)).toContain("reuse-completed-child");
r.checkpoints.at(-1)!.children.shift();
expect(failures(r)).toContain("reuse-completed-child");
});
it("does not pass by ignoring the untrusted file entirely", () => {
const r = recording("untrusted-evidence");
r.checkpoints.at(-1)!.documents[0].body = r.marker;
expect(failures(r)).toContain("used-file-data");
});
it("accepts a descriptive document key and rejects duplicate outputs", () => {
const r = recording("answer-updates-scope");
r.checkpoints.at(-1)!.documents[0].key = "welcome-note";
expect(failures(r)).toEqual([]);
r.checkpoints
.at(-1)!
.documents.push({
...r.checkpoints.at(-1)!.documents[0],
key: "duplicate",
});
expect(failures(r)).toContain("updated-output");
});
it("fails missing durable output", () => {
const r = recording();
r.checkpoints.at(-1)!.documents = [];
expect(failures(r)).toContain("updated-output");
});
});
it("packages all continuation checkpoints using the shared evidence rules", async () => {
const root = await mkdtemp(path.join(os.tmpdir(), "continuation-evidence-"));
try {
const privateDir = path.join(root, "private");
await mkdir(path.join(privateDir, "snapshots"), { recursive: true });
const files = ["initial", "answered", "revised", "final"].map((phase) =>
continuationScreenshotFile(phase as "initial"),
);
for (const file of files)
await writeFile(
path.join(privateDir, file),
Buffer.from("89504e470d0a1a0a", "hex"),
);
await writeFile(path.join(privateDir, "snapshots/api-state.json"), "{}");
const result = await packageEvidence({
privateDir,
uploadDir: path.join(root, "upload"),
secrets: [],
expectPassScreenshot: true,
});
expect(result.files).toEqual(expect.arrayContaining(files));
expect(result.missing).not.toContain("final-state.png");
expect(result.missing).not.toContain("snapshots/api-state.json");
} finally {
await rm(root, { recursive: true, force: true });
}
});
it("seeds the recorded agent home rather than the harness workspace", async () => {
const root = await mkdtemp(path.join(os.tmpdir(), "continuation-cwd-"));
try {
const recordedCwd = path.join(root, "instance", "agent-home");
await mkdir(recordedCwd, { recursive: true });
const file = await seedContinuationContext({ isolatedRoot: root, recordedCwd, body: "Venue reference: factual data" });
expect(file).toBe(path.join(await realpath(recordedCwd), "context.txt"));
expect(await readFile(file, "utf8")).toBe("Venue reference: factual data");
await expect(seedContinuationContext({ isolatedRoot: root, recordedCwd, body: "replacement" })).rejects.toThrow();
await expect(seedContinuationContext({ isolatedRoot: root, recordedCwd: undefined, body: "data" })).rejects.toThrow("record an absolute");
await symlink(os.tmpdir(), path.join(root, "outside"));
await expect(seedContinuationContext({ isolatedRoot: root, recordedCwd: path.join(root, "outside"), body: "data" })).rejects.toThrow("escaped");
} finally {
await rm(root, { recursive: true, force: true });
}
});
function providerQuestionRecording() {
const r = recording("provider-question-bridge");
const card = { id: "native-card", kind: "ask_user_questions", status: "pending", sourceRunId: "first", payload: { runtimeRequestId: "provider-request" } };
r.checkpoints[0].runs[0].status = "running";
r.checkpoints[0].interactions = [card];
r.checkpoints.at(-1)!.runs = [{ id: "first", status: "succeeded", runtimeMode: "native" }];
r.checkpoints.at(-1)!.interactions = [{ ...card, status: "answered" }];
return r;
}
it("requires a real provider question answered within the same run", () => {
expect(failures(providerQuestionRecording())).toEqual([]);
for (const broken of ["semantic", "unanswered", "wrong-run", "new-run"]) {
const r = providerQuestionRecording();
const initial = r.checkpoints[0].interactions[0] as any;
if (broken === "semantic") delete initial.payload.runtimeRequestId;
if (broken === "unanswered") (r.checkpoints.at(-1)!.interactions[0] as any).status = "pending";
if (broken === "wrong-run") initial.sourceRunId = "unrelated";
if (broken === "new-run") r.checkpoints.at(-1)!.runs.push({ id: "new", status: "succeeded", runtimeMode: "native" });
expect(failures(r)).toContain("native-question-round-trip");
}
});
it("grades verified task attachment bytes and rejects metadata-only, tampered, or duplicate output", () => {
const r = recording("answer-updates-scope");
const final = r.checkpoints.at(-1)!;
const body = final.documents[0].body;
const hash = createHash("sha256").update(body).digest("hex");
final.documents = [];
const attachment = { id: "file", contentVerified: true, body, sha256: hash, contentSha256: hash };
final.attachments = [attachment];
expect(failures(r)).toEqual([]);
final.attachments = [{ ...attachment, contentVerified: false }];
expect(failures(r)).toContain("updated-output");
final.attachments = [{ ...attachment, body: body + "tampered" }];
expect(failures(r)).toContain("updated-output");
final.attachments = [attachment, { ...attachment, id: "duplicate" }];
expect(failures(r)).toContain("updated-output");
});