Files
PaperClipAI/tests/runner-e2e/chat-qualification.test.ts
T
DottaandPaperclip 74a9730acb fix: continue native agent chats after worker loss (#13813)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat uses native workers to run Claude and Codex
conversations.
> - A worker crash leaves a cleanup hold because its provider did not
acknowledge suspension.
> - A new user message must not reuse that unverified session or repeat
old tool calls.
> - The existing continuation path can preserve history and start a
fresh session, but local cleanup ownership remained held.
> - This pull request verifies the stopped local owners and releases
only their cleanup hold for a new user turn.

## Linked Issues or Issue Description

Refs #13775.

**What happened?**

After a native worker crashed, both providers retained cleanup
quarantine. A saved plan survived, but the conversation could not
produce another answer.

**Expected behavior**

Once the old worker and provider process groups have stopped, a new user
message can continue in a fresh session with the saved work and prior
action history.

**Steps to reproduce**

Run the opt-in `agent-chat-qualification` suite with case
`worker-crash-retry` on native Codex and native Claude. The fixture
saves a plan, kills the exact worker through a Linux pidfd, releases a
local read-only brief, and sends a new message.

## What Changed

- Verify the exact local worker stop receipt, provider identity
receipts, released leases, and retained state before retiring a native
cleanup hold.
- Recheck process liveness and state before admission. Keep the old run
and durable session files intact.
- Use the existing explicit conversation continuation path. Generic
Retry remains blocked for cleanup quarantine, including on the old
failed-run marker after a successful continuation.
- Extend the live oracle to require a successful fresh session, correct
predecessor context, unchanged plan, one original message, and one
answer containing a reference introduced after the crash.
- Add physical-proof and database-backed admission tests. Document the
precise qualification scope.

## Verification

- Live Product E2E: **2/2 passed**, **2/2 cleanup passed**, with real
native `gpt-5.6-sol` and `claude-sonnet-5`, Chromium, server, database,
and public APIs.
- Core recovery proof source:
`3592b04c2bc76e23795fcdf964720e38a409dc4d`. Suite definition version 8:
`9867367994d81a0c726956d91f2c7fddab6417a12f41b5cef3f7e62f3be417da`.
- [Core recovery
campaign](https://github.com/paperclipai/paperclip/actions/runs/35741746990)
· [Public evidence
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35741746990-1/).
- Both cells verify the real crash boundary, blocked generic Retry,
unchanged saved plan, one original prompt, one fresh successor with
predecessor context, and one run-attributed answer containing the
post-crash reference. Billing coverage is partial because the crashed
runs did not report complete usage; missing cost is not zero cost.
- [Master
baseline](https://github.com/paperclipai/paperclip/actions/runs/35737364443):
both providers stopped at quarantine, with cleanup passing.
- [Follow-up baseline without the
fix](https://github.com/paperclipai/paperclip/actions/runs/35738638866):
Codex produced a complete red result. Claude reached the same error, but
its artifact upload was canceled.
- [Complete Claude baseline with the same version-8
definition](https://github.com/paperclipai/paperclip/actions/runs/35740555177):
red at cleanup quarantine, cleanup passed, source
`6479a90c5754044356b39a9278b9a3e92ce8e55e`.
- Earlier candidate attempts remain retained:
[first](https://github.com/paperclipai/paperclip/actions/runs/35738449214)
passed Codex and found a Claude fixture wait race;
[second](https://github.com/paperclipai/paperclip/actions/runs/35740408065)
exposed the normalized session-open receipt mismatch. Both corrections
are in the final source.
- Targeted server suites: 570 passed before the final two additional
receipt regression cases. The physical-proof suite, including those
cases, passed 46/46. Eval oracle and catalog: 40 passed. Server and
Product E2E typechecks passed.
- **All 54 PR checks passed on final head `db6f775df`**, including
repository typecheck, build, tests, browser shards, and canary dry run.
The canary job required one retry after its runner received a shutdown
signal. On the earlier core proof head, two timing-sensitive tests
passed in isolation and on a single CI retry.
- Final UI regression checks: 5 passed; UI typecheck and token gates
passed. Updated eval oracle/catalog: 40 passed; eval typecheck passed.
- Final version-9 two-provider campaign: **2/2 passed, 2/2 cleanup
passed**, including the browser assertion that the quarantined
historical run never regains Try again. [Final
campaign](https://github.com/paperclipai/paperclip/actions/runs/35747416013),
source `db6f775dfff405e1514ec02fedb0450d42c7dad2`, definition hash
`bf5abf1cc45cb6dad4e082fbf818b8fa0f4d8c282776a7e98762b22919eadab7`.
[Final public evidence
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35747416013-1/).
Final screenshots and retained state inspected for both providers; cost
coverage remains partial.

## Risks

- Missing or conflicting stop evidence keeps the conversation blocked.
This change does not kill an unverified process.
- This qualifies new user input after local worker loss. It does not
enable automatic replay, exact-session recovery, remote crash recovery,
or native onboarding defaults.
- Old action outcomes remain part of the continuation. A process exit is
not proof that an action did not happen.
- No schema migration or production prompt change.

## Model Used

OpenAI Codex, GPT-6. The runtime does not expose a more specific model
identifier or context-window size. Used reasoning, repository
inspection, code editing, shell tools, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 13:01:45 -05:00

89 lines
7.1 KiB
TypeScript

import { execFileSync } from "node:child_process";
import path from "node:path";
import { describe, expect, it } from "vitest";
import { assertActiveHandoff, assertAnswerFacts, assertCrashRecovered, assertWorkerIdentity } from "./chat-qualification.js";
import { runnerMatrix } from "./catalog.js";
import { buildRunnerE2EProcessEnvironment } from "./harness-env.js";
const run = { id: "old", nativeSessionId: "old-session", agentId: "original", companyId: "company", runtimeMode: "native", status: "cancelled", errorCode: "issue_reassigned", finishedAt: "2026-09-21T00:01:00Z", contextSnapshot: { issueId: "task" } };
const successor = { ...run, id: "next", agentId: "successor", status: "succeeded", startedAt: "2026-09-21T00:01:01Z" };
const before = { id: "task", assigneeAgentId: "original", description: "Preserve scope", projectId: null };
const plan = { body: "Friday REFERENCE", latestRevisionId: "revision" };
const handoff = { before, after: { ...before, status: "done", assigneeAgentId: "successor" }, oldRun: run, boundary: { ...run, status: "running" }, runs: [run, successor], successorId: "successor", planBefore: plan, planAfter: plan, draft: { id: "draft", latestRevisionId: "v1", body: "REFERENCE" }, draftAfter: { id: "draft", body: "REFERENCE" }, draftRevisions: [{ id: "v1", body: "REFERENCE" }], output: { body: "REFERENCE", createdByAgentId: "successor" }, reference: "REFERENCE", audit: [{ action: "issue.reassigned", details: { source: "paperclip_runner_protocol" } }], taskIds: ["task"] };
const recovery = { boundary: { ...run, status: "running" }, failed: { ...run, status: "failed" }, runs: [{ ...run, status: "failed" }, { ...successor, agentId: "original", nativeSessionId: "fresh-session", contextSnapshot: { issueId: "task", previousRunId: "old", forceFreshSession: true } }], issueId: "task", prompt: "Read my brief", comments: [{ body: "Read my brief" }, { body: "REFERENCE MARKER", authorAgentId: "original", createdByRunId: "next" }], reference: "REFERENCE", marker: "MARKER", planBefore: { body: "plan MARKER", latestRevisionId: "v1" }, planAfter: { body: "plan MARKER", latestRevisionId: "v1" } };
describe("remaining native chat qualification", () => {
it("calibrates the pidfd helper against reuse and wrong-identity faults", () => {
execFileSync("python3", [path.join(import.meta.dirname, "worker-fault.test.py")], { stdio: "pipe" });
});
it("requires a stopped original worker before exactly one successor executes", () => {
expect(() => assertActiveHandoff(handoff)).not.toThrow();
expect(() => assertActiveHandoff({ ...handoff, draftAfter: { id: "draft", body: "Expanded REFERENCE" },
output: { body: "Expanded REFERENCE", createdByAgentId: "original", updatedByAgentId: "successor" } })).not.toThrow();
for (const change of [
{ boundary: { ...handoff.boundary, status: "succeeded" } },
{ oldRun: { ...run, status: "succeeded" } },
{ oldRun: { ...run, errorCode: "unrelated_cancellation" } },
{ runs: [run, { ...successor, startedAt: "2026-09-21T00:00:59Z" }] },
{ runs: [run, successor, { ...successor, id: "duplicate" }] },
{ runs: [run] },
{ taskIds: ["replacement"] },
{ planAfter: { ...plan, body: "rewritten" } },
{ draft: { body: "no saved progress" } },
{ draftAfter: { id: "replaced", body: "REFERENCE" } },
{ draftRevisions: [] },
{ draftRevisions: [{ id: "v1", body: "overwritten original" }] },
{ output: { body: "REFERENCE", createdByAgentId: "original" } },
{ audit: [] },
]) expect(() => assertActiveHandoff({ ...handoff, ...change })).toThrow();
});
it("requires successful run-attributed recovery preserving input and saved work", () => {
expect(() => assertCrashRecovered(recovery)).not.toThrow();
for (const change of [
{ boundary: { ...recovery.boundary, status: "succeeded" } },
{ failed: { ...recovery.failed, id: "unrelated" } },
{ runs: [recovery.runs[0]!, { ...successor, status: "failed" }] },
{ runs: [recovery.runs[0]!, { ...successor, contextSnapshot: { issueId: "new-chat" } }] },
{ runs: [recovery.runs[0]!, { ...recovery.runs[1]!, nativeSessionId: "old-session" }] },
{ runs: [recovery.runs[0]!, { ...recovery.runs[1]!, contextSnapshot: { issueId: "task", previousRunId: "unrelated", forceFreshSession: true } }] },
{ comments: [...recovery.comments, recovery.comments[0]!] },
{ comments: [recovery.comments[0]!, { ...recovery.comments[1], createdByRunId: "old" }] },
{ comments: [recovery.comments[0]!, { ...recovery.comments[1], body: "MARKER" }] },
{ planAfter: { ...recovery.planAfter, latestRevisionId: "rewritten" } },
]) expect(() => assertCrashRecovered({ ...recovery, ...change })).toThrow();
});
it("refuses unknown PIDs, nonnative or nonlocal processes and partial run identities", () => {
const running = { id: "run-123", status: "running", runtimeMode: "native", processPid: 123456 };
expect(() => assertWorkerIdentity(running, "node runner --run-id run-123 --worker", "local")).not.toThrow();
for (const command of ["node server", "node runner --run-id run-1234", "node runner run-123 --run-id other"])
expect(() => assertWorkerIdentity(running, command, "local")).toThrow();
for (const processPid of [undefined, 0, 1, -20, process.pid, 1.2])
expect(() => assertWorkerIdentity({ ...running, processPid }, "node runner --run-id run-123", "local")).toThrow();
expect(() => assertWorkerIdentity(running, "node runner --run-id run-123", "daytona")).toThrow();
expect(() => assertWorkerIdentity({ ...running, runtimeMode: "legacy" }, "node runner --run-id run-123", "local")).toThrow();
});
it("grades factual propositions rather than matching words in misleading prose", () => {
const facts = { currentBlocker: "VENUE", confirmedAttendance: null, printingStarted: false };
const explanation = "The venue is still unconfirmed, so printing remains deferred. No attendance count is recorded.";
expect(() => assertAnswerFacts(JSON.stringify({ facts, explanation }), facts)).not.toThrow();
for (const wrong of [
{ ...facts, currentBlocker: "BUDGET" }, { ...facts, confirmedAttendance: 40 },
{ ...facts, printingStarted: true }, { ...facts, madeUpMetric: 50 }, {},
]) expect(() => assertAnswerFacts(JSON.stringify({ facts: wrong, explanation }), facts)).toThrow();
expect(() => assertAnswerFacts(JSON.stringify({ facts, explanation: "" }), facts)).toThrow();
expect(() => assertAnswerFacts("VENUE confirmedAttendance null printingStarted false", facts)).toThrow();
});
it("exposes exactly six explicit local native cells with bounded run counts", () => {
const cells = runnerMatrix.filter(c => c.suite.id === "agent-chat-qualification");
expect(cells).toHaveLength(6);
for (const cell of cells) {
expect(cell.suite.manualOnly).toBe(true);
expect(cell.environment.id).toBe("local");
expect(cell.profile.generation).toBe("native");
expect(cell.task.expectedRunCount).toBe(cell.task.id === "active-reassignment" ? 3 : 2);
expect(buildRunnerE2EProcessEnvironment({}, [cell]).PAPERCLIP_RUNNER_API_TOOLS_ENABLED)
.toBe(cell.task.id === "grounded-answer-quality" ? "true" : undefined);
}
});
});