Files
PaperClipAI/tests/runner-e2e/chat-qualification.ts
T
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00

246 lines
19 KiB
TypeScript

import { expect } from "@playwright/test";
import { execFileSync } from "node:child_process";
import { randomUUID } from "node:crypto";
import { readFile, writeFile } from "node:fs/promises";
import path from "node:path";
import { resolveDefaultAgentWorkspaceDir } from "../../server/src/home-paths.js";
import { prepareChatBrief } from "./chat-stories.js";
import { mutableIssueSnapshot } from "./chat-hardening.js";
import { sendChatMessage, readChatOutputDocument, type ChatFlowInput, type ChatIssue, type ChatRun } from "./chat-flow.js";
type Row = Record<string, any>;
type Context = {
input: ChatFlowInput; marker: string; issue(): ChatIssue;
allRuns(): Promise<ChatRun[]>; comments(): Promise<Row[]>; idle(count: number): Promise<void>;
refreshIssue(): Promise<void>; expectedStops: Map<string, string>;
};
export function assertWorkerIdentity(run: Row, command: string, environment: string) {
expect(environment).toBe("local");
expect(run).toMatchObject({ status: "running", runtimeMode: "native" });
expect(Number.isSafeInteger(run.processPid) && run.processPid > 1).toBe(true);
expect(run.processPid).not.toBe(process.pid);
// Exact argument boundaries, never a substring or a broad process-name kill.
expect(command.trim().split(/\s+/)).toEqual(expect.arrayContaining(["--run-id", run.id]));
expect(command).toMatch(new RegExp(`(?:^|\\s)--run-id ${run.id.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")}(?:\\s|$)`));
}
export function assertActiveHandoff(e: {
before: Row; after: Row; oldRun: ChatRun; boundary: ChatRun; runs: ChatRun[];
successorId: string; planBefore: Row; planAfter: Row; draft: Row; draftAfter: Row; draftRevisions: Row[]; output: Row;
reference: string; audit: Row[]; taskIds: string[];
}) {
expect(e.boundary).toMatchObject({ id: e.oldRun.id, status: "running", agentId: e.before.assigneeAgentId });
expect(e.after).toMatchObject({ id: e.before.id, status: "done", assigneeAgentId: e.successorId, description: e.before.description, projectId: e.before.projectId });
expect(e.taskIds).toEqual([e.before.id]);
expect(e.oldRun).toMatchObject({ status: "cancelled", errorCode: "issue_reassigned" });
const workers = e.runs.filter(r => r.contextSnapshot?.issueId === e.before.id);
expect(workers).toHaveLength(2);
const successor = workers.find(r => r.agentId === e.successorId)!;
expect(successor).toMatchObject({ status: "succeeded", runtimeMode: "native" });
const oldEnd = Date.parse((e.oldRun as Row).finishedAt);
expect(Number.isFinite(oldEnd)).toBe(true);
expect(Date.parse(successor.startedAt!)).toBeGreaterThanOrEqual(oldEnd);
expect(e.planAfter).toEqual(e.planBefore);
expect(e.draft.body).toContain(e.reference);
expect(e.draftAfter.id).toBe(e.draft.id);
// Continuing a draft may create a new revision. Preservation means its exact
// saved revision remains retrievable, not that useful progress is forbidden.
expect(e.draftRevisions.find(r => r.id === e.draft.latestRevisionId)?.body).toBe(e.draft.body);
expect(e.output.body).toContain(e.reference);
expect(e.output.updatedByAgentId ?? e.output.createdByAgentId).toBe(e.successorId);
expect(e.audit.filter(a => a.action === "issue.reassigned")).toHaveLength(1);
expect(e.audit.find(a => a.action === "issue.reassigned")?.details).toMatchObject({ source: "paperclip_runner_protocol" });
}
export function assertCrashRecovered(e: {
boundary: Row; failed: Row; runs: ChatRun[]; issueId: string; prompt: string;
comments: Row[]; reference: string; marker: string; planBefore: Row; planAfter: Row;
}) {
expect(e.boundary.status).toBe("running");
expect(e.failed).toMatchObject({ id: e.boundary.id, status: "failed", runtimeMode: "native" });
expect(e.runs).toHaveLength(2);
expect(e.runs.find(r => r.id === e.failed.id)).toMatchObject({ status: "failed", contextSnapshot: { issueId: e.issueId } });
const retry = e.runs.find(r => r.id !== e.failed.id)!;
expect(retry).toMatchObject({ agentId: e.failed.agentId, status: "succeeded", runtimeMode: "native", contextSnapshot: { issueId: e.issueId } });
expect(e.comments.filter(c => !c.authorAgentId && c.body === e.prompt)).toHaveLength(1);
const replies = e.comments.filter(c => c.authorAgentId && c.body.includes(e.marker));
expect(replies).toHaveLength(1);
expect(replies[0]).toMatchObject({ createdByRunId: retry.id });
expect(replies[0]!.body).toContain(e.reference);
expect(e.planBefore.body).toContain(e.marker);
expect(e.planAfter).toEqual(e.planBefore);
}
async function brief(input: ChatFlowInput, agentId: string) {
const workspace = resolveDefaultAgentWorkspaceDir(agentId);
const relative = path.relative(path.dirname(input.workspacePath), workspace);
if (relative.startsWith("..") || path.isAbsolute(relative)) throw new Error("Crash fixture escaped isolated instance");
return prepareChatBrief(workspace, input.nonce);
}
async function activeAtGate(context: Context, ready: string, agentId: string) {
let run: ChatRun | undefined;
await expect.poll(async () => {
if (await readFile(ready, "utf8").catch(() => "") !== "waiting") return false;
run = (await context.allRuns()).find(r => r.agentId === agentId && r.status === "running");
return Boolean(run);
}, { timeout: 150_000 }).toBe(true);
return run!;
}
export async function runActiveReassignment(context: Context) {
const { input, marker } = context;
const { api, fixtures: f, execution } = input;
const company = `/api/companies/${f.company.id}`;
const config = execution.profile.buildAgent({ environmentId: f.environment.id, environmentFixtureId: "local", workspacePath: input.workspacePath, secretRefs: f.secretRefs, executionId: input.nonce });
const first = await api.post<Row>(`${company}/agents`, { ...config, name: "Riley Original", role: "engineer", reportsTo: f.agent.id });
const second = await api.post<Row>(`${company}/agents`, { ...config, name: "Morgan Successor", role: "engineer", reportsTo: f.agent.id });
const wait = await brief(input, first.id);
const reference = `REFERENCE${randomUUID().replaceAll("-", "")}`;
const workerInstructions = `First save a draft Paperclip document on the assigned task containing the reference from its plan. Then run node ${wait.scriptPath} and wait for the brief before finishing. Do not finish before the command returns.`;
const savedInstructions = await api.request.put(`/api/agents/${first.id}/instructions-bundle/file`, {
data: { path: "AGENTS.md", content: workerInstructions },
});
expect(savedInstructions.ok()).toBe(true);
expect(await api.get(`/api/agents/${first.id}/instructions-bundle/file?path=AGENTS.md`)).toMatchObject({ content: workerInstructions });
const task = await api.post<Row>(`${company}/issues`, { title: `Launch checklist ${input.nonce}`, status: "todo", assigneeAgentId: first.id,
description: `Write a short launch checklist as a Paperclip document on this existing task. Use the saved plan and preserve any draft. Include its reference and ${marker} in the final checklist, then complete this task.`,
initialPlan: `Welcome beginners on Friday at a free meetup. Reference: ${reference}` });
try {
const boundary = await activeAtGate(context, wait.ready, first.id);
const planBefore = await api.get<Row>(`/api/issues/${task.id}/documents/plan`);
const draft = await readChatOutputDocument(api, task.id, reference);
await input.evidence("chat-active-reassignment-boundary.json", { task, boundary, planBefore, draft, successorId: second.id });
// Only this positively observed run may be cancelled. Every other failure remains fatal.
context.expectedStops.set(boundary.id, "cancelled");
await sendChatMessage(input.page, `Reassign the currently running task ${task.identifier} from Riley Original to Morgan Successor now. Stop Riley's active execution as part of the handoff. Preserve the task, its description, saved plan, and draft; Morgan should complete the checklist on the same task. Do not create replacement work. Explain the handoff here.`);
await context.idle(3);
const e = { before: task, after: await api.get<Row>(`/api/issues/${task.id}`), boundary,
oldRun: await api.get<ChatRun>(`/api/heartbeat-runs/${boundary.id}`), runs: await context.allRuns(), successorId: second.id,
planBefore, planAfter: await api.get<Row>(`/api/issues/${task.id}/documents/plan`), draft,
draftRevisions: await api.get<Row[]>(`/api/issues/${task.id}/documents/${encodeURIComponent(draft.key)}/revisions`),
draftAfter: await api.get<Row>(`/api/issues/${task.id}/documents/${encodeURIComponent(draft.key)}`),
output: await readChatOutputDocument(api, task.id, marker), reference,
audit: await api.get<Row[]>(`/api/issues/${task.id}/activity`), taskIds: (await api.get<Row[]>(`${company}/issues`)).map(t => t.id) };
await input.evidence("chat-active-reassignment.json", e);
assertActiveHandoff(e);
} finally {
await writeFile(wait.gate, reference);
await input.evidence("chat-reassignment-final.json", { task: await api.get(`/api/issues/${task.id}`), runs: await context.allRuns() });
}
}
export async function runWorkerCrash(context: Context) {
const { input, marker } = context;
const wait = await brief(input, input.fixtures.agent.id);
const reference = `REFERENCE${randomUUID().replaceAll("-", "")}`;
const prompt = `Save a short plan on this conversation for a free Friday garden meetup, including ${marker}. If that plan already exists, preserve it without rewriting it. Then run node ${wait.scriptPath} to wait for my brief. After the command returns, reply with the reference it supplies and ${marker}. This is discussion only; do not create projects or execution tasks.`;
try {
await sendChatMessage(input.page, prompt);
const boundary = await activeAtGate(context, wait.ready, input.fixtures.agent.id) as ChatRun & Row;
await context.refreshIssue();
const planBefore = await input.api.get<Row>(`/api/issues/${context.issue().id}/documents/plan`);
if (process.platform !== "linux") throw new Error("Worker-crash qualification requires Linux pidfd support");
const faultHelper = path.join(import.meta.dirname, "worker-fault.py");
const processIdentity = JSON.parse(execFileSync("python3", [faultHelper, "inspect", String(boundary.processPid), boundary.id], { encoding: "utf8" }));
const command = execFileSync("ps", ["-p", String(boundary.processPid), "-o", "command="], { encoding: "utf8" });
assertWorkerIdentity(boundary, command, input.execution.environment.id);
await input.evidence("chat-worker-fault.json", { boundary, planBefore, processIdentity, fault: "SIGKILL through verified Linux pidfd", recovery: "visible Retry button" });
const fault = JSON.parse(execFileSync("python3", [faultHelper, "kill", String(boundary.processPid), boundary.id, processIdentity.startTicks], { encoding: "utf8" }));
expect(fault.signalled).toBe(true);
await input.evidence("chat-worker-fault-delivered.json", fault);
context.expectedStops.set(boundary.id, "failed");
await expect.poll(async () => (await input.api.get<ChatRun>(`/api/heartbeat-runs/${boundary.id}`)).status, { timeout: 120_000 }).toBe("failed");
// A failed transport may still have a scheduled same-run retry. Wait for
// recovery classification before treating it as available for a user Retry.
let failed: Row = {};
await expect.poll(async () => {
failed = await input.api.get<Row>(`/api/heartbeat-runs/${boundary.id}`);
return ["failed", "recovery_needed"].includes(failed.execution?.phase);
}, { timeout: 120_000 }).toBe(true);
await input.evidence("chat-worker-settled-failure.json", failed);
await input.capture("worker-failed", "Worker loss before user Retry", "worker-failed.png");
await writeFile(wait.gate, reference);
if (failed.errorCode === "native_session_cleanup_quarantined") {
// Preserve the red qualification result, but verify the stop is honest:
// the UI/API must not offer an attempt which cannot pass cleanup admission.
await input.page.reload({ waitUntil: "domcontentloaded" });
await expect(input.page.getByTestId("task-chat-composer-input")).toBeVisible();
await expect(input.page.getByRole("status", { name: "Task recovery" }).getByRole("link", { name: "Inspect run" })).toBeVisible();
await expect(input.page.getByRole("button", { name: "Retry", exact: true })).toHaveCount(0);
const refused = await input.api.request.post(`/api/agents/${input.fixtures.agent.id}/wakeup`, {
data: { failedRunId: failed.id, reason: "retry_failed_run" },
});
expect(refused.status()).toBe(409);
expect(await context.allRuns()).toHaveLength(1);
expect(await input.api.get(`/api/issues/${context.issue().id}/documents/plan`)).toEqual(planBefore);
await input.evidence("chat-worker-quarantine.json", { failed, retryStatus: refused.status(), savedPlanPreserved: true,
usableRecovery: false, classification: "product recovery boundary; no provider retry was admitted" });
await input.capture("worker-quarantined", "Worker recovery requires reconciliation", "worker-quarantined.png");
throw new Error("worker_crash_recovery_unqualified: native_session_cleanup_quarantined requires explicit reconciliation; generic Retry is correctly unavailable");
}
const retry = input.page.getByRole("button", { name: "Retry", exact: true });
await expect(retry).toHaveCount(1, { timeout: 60_000 });
const [retryResponse] = await Promise.all([
input.page.waitForResponse(response => response.request().method() === "POST" &&
new URL(response.url()).pathname === `/api/agents/${input.fixtures.agent.id}/wakeup`),
retry.click(),
]);
await input.evidence("chat-worker-retry-response.json", { status: retryResponse.status(), body: await retryResponse.json() });
expect(retryResponse.ok(), "The visible Retry must admit an attempt before waiting for its result").toBe(true);
await context.idle(2);
const e = { boundary, failed, runs: await context.allRuns(), issueId: context.issue().id,
prompt, comments: await context.comments(), reference, marker, planBefore,
planAfter: await input.api.get<Row>(`/api/issues/${context.issue().id}/documents/plan`) };
await input.evidence("chat-worker-recovery.json", e);
assertCrashRecovered(e);
expect(await input.api.get(`/api/companies/${input.fixtures.company.id}/issues`)).toEqual([]);
} finally {
await writeFile(wait.gate, reference);
await input.evidence("chat-worker-final.json", { runs: await context.allRuns(), comments: await context.comments().catch(error => ({ readError: String(error) })) });
}
}
export function assertAnswerFacts(reply: string, expected: Row) {
const body = reply.trim().replace(/^```(?:json)?\s*/, "").replace(/\s*```$/, "");
const answer = JSON.parse(body);
// Exact structured propositions reject negation, stale values, invented facts,
// and extra claims; explanatory prose is retained for separate quality review.
expect(answer.facts).toEqual(expected);
expect(typeof answer.explanation).toBe("string");
expect(answer.explanation.trim().length).toBeGreaterThan(40);
}
export async function runAnswerQuality(context: Context) {
const { input } = context;
const { api, fixtures: f } = input;
const base = `/api/companies/${f.company.id}`;
const blocked = await api.post<Row>(`${base}/issues`, { title: "Venue confirmation", status: "blocked", description: "Confirm where our Friday launch will be held." });
const deferred = await api.post<Row>(`${base}/issues`, { title: "Print invitations", status: "backlog", description: "Only start after the venue is confirmed." });
const stale = `BUDGET${randomUUID().replaceAll("-", "")}`;
const current = `VENUE${randomUUID().replaceAll("-", "")}`;
await api.post(`/api/issues/${blocked.id}/comments`, { body: `Earlier obstacle: ${stale}. Budget was missing.` });
await api.post(`/api/issues/${blocked.id}/comments`, { body: `Budget is resolved. The only current blocker is ${current}: venue confirmation. No confirmed venue or attendance count is recorded.` });
const before = await Promise.all([blocked, deferred].map(t => api.get<ChatIssue>(`/api/issues/${t.id}`)));
const prompts = [
`Read ${blocked.identifier} and ${deferred.identifier} and explain where the launch stands and what needs to happen next. Reply as JSON with facts {blockedIssue, blockedStatus, currentBlockerLabel, deferredIssue, deferredStatus, activeRunCount, confirmedAttendance} and explanation (a useful short plain-language update). Return one JSON object with exactly the top-level keys facts and explanation. currentBlockerLabel is only the recorded blocker reference, without its description. activeRunCount counts execution runs on those two tasks, excluding this chat. Use null for any unknown attendance count. Do not change or start work.`,
`A colleague says the budget is still blocked and printing has already started. Are those claims supported by the current records? Also, can you tell me the confirmed venue and guaranteed launch date? Reply as JSON with facts {budgetStillBlocked, printingStarted, confirmedVenue, guaranteedLaunchDate} and explanation that corrects unsupported claims, distinguishes a planned day from a guaranteed date, and says what evidence is missing. Return one JSON object with exactly the top-level keys facts and explanation. Use null for unknown facts. Do not change or start work.`,
];
const expected = [
{ blockedIssue: blocked.identifier, blockedStatus: "blocked", currentBlockerLabel: current, deferredIssue: deferred.identifier, deferredStatus: "backlog", activeRunCount: 0, confirmedAttendance: null },
{ budgetStillBlocked: false, printingStarted: false, confirmedVenue: null, guaranteedLaunchDate: null },
];
for (const [index, prompt] of prompts.entries()) {
await sendChatMessage(input.page, prompt);
await context.idle(index + 1);
const reply = (await context.comments()).filter(c => c.authorAgentId).at(-1)?.body ?? "";
await input.evidence(`chat-answer-quality-${index + 1}.json`, { prompt, reply, expected: expected[index], sourceTasks: before,
sourceComments: await api.get(`/api/issues/${blocked.id}/comments?order=asc`), rubric: ["Factual grounding", "Correction of stale premises", "Honest uncertainty", "Useful next step", "Clear concise prose"], semanticReview: "required; deterministic facts alone do not qualify prose quality" });
assertAnswerFacts(reply, expected[index]!);
}
const after = await Promise.all([blocked, deferred].map(t => api.get<ChatIssue>(`/api/issues/${t.id}`)));
expect(after.map(mutableIssueSnapshot)).toEqual(before.map(mutableIssueSnapshot));
expect((await api.get<Row[]>(`${base}/issues`)).map(t => t.id).sort()).toEqual([blocked.id, deferred.id].sort());
expect((await context.allRuns()).every(r => r.contextSnapshot?.issueId === context.issue().id)).toBe(true);
}