mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 16:35:27 +02:00
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat needs reliable native execution before native runners become the onboarding default. > - Existing stories covered idle reassignment and controller restart, but not an executing worker handoff or worker process loss. > - Status answer tests also need to reject stale claims and invented facts. > - This pull request adds six opt-in full-stack cells with independent state assertions and retained evidence. > - The probes exposed a misleading Retry across server projection and recovery-banner paths; the fix reports the blocked recovery honestly. > - The tests preserve failures without changing recovery policy, production prompts, or onboarding defaults. ## Linked Issues or Issue Description Refs: #13762. Related: #13765 (Retry targets the latest failed attempt), #13753 (task context ownership), #13746 (native recovery work). ## What Changed - Add active reassignment with saved draft and plan preservation, old-worker cancellation, and successor completion checks. - Preserve recovery-needed projection when native cleanup fails before its coordinator exists, refuse a generic retry that would immediately fail again, and replace the recovery banner's misleading Retry with Inspect run. - Add verified local worker process loss with a required successful continuation; retain a failing qualification result when recovery is unavailable, while independently verifying the UI/API refuse doomed retries. - Add two-turn factual answer checks for current blockers, stale claims, inactive backlog work, and unknown facts. Retain prose for separate semantic review. - Add positive and negative oracle calibration and document fault isolation, cleanup, billing, and qualification limits. ## Verification - Eval TypeScript check passes. - All 442 eval support tests pass locally. The 89 focused server tests and server typecheck pass. Six recovery-banner UI tests and token gates pass. - Initial new-cell campaign: https://github.com/paperclipai/paperclip/actions/runs/35657128077. All six results are retained; four failed on fixture-contract issues and two exposed real worker cleanup quarantine. - All 26 existing native onboarding cells: https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26 passed on master846336e5a, all cleanup passed). - Intermediate handoff/fault campaign: https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four retained failures: two overly strict draft oracles, two real crash quarantines). - Final active handoff: https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2 passed on cf6d4ae3a; both cleanup passed). - Clarified answer-quality fixtures: https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2 passed on 4a26f10be; both cleanup passed; all four answers semantically reviewed). - Quarantine guard regression campaign: https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both API requests correctly refused with 409/no second run, but exposed a separate misleading Retry in the recovery banner and a fixture wait on a non-admitted run). - Final quarantine guard verification: https://github.com/paperclipai/paperclip/actions/runs/35661067305 (147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one retained run, unchanged saved plan, and successful disposable cleanup. Both evals intentionally remain red with `worker_crash_recovery_unqualified`; no successful continuation exists). The preceding campaign 35658772755 never ran provider cases because GitHub artifact finalization returned HTTP 403. - Full repository CI passes on147e42f7e: typecheck, tests, build, and browser gates. One unchanged local-service-supervisor readiness test failed initially; its six-test file passed in isolation and the failed shard passed on its single rerun. Latest-head rollup: 54 successful, 2 intentionally skipped, no failed or pending checks. Greptile is 5/5 with zero unresolved findings. - See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts and semantic review. Published reports: [onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/), [handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/), [grounded answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/), [crash guards and unqualified recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/). ## Risks - Paid cells are explicit-only and local-only. The fault fixture signals only the exact native run PID after checking its identity. - Live worker-loss probes currently fail on cleanup quarantine for both providers. The eval must remain red until there is a usable recovery, even when preservation and refusal checks pass. Verified cleanup with a fresh attempt versus exact-session resume remains a product decision. - Structured facts alone do not qualify prose quality; semantic review remains separate. - Onboarding uses the existing runtime switch after the real wizard and before provider execution. Native UI selection and public defaults remain unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact served snapshot and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
87 lines
6.7 KiB
TypeScript
87 lines
6.7 KiB
TypeScript
import { execFileSync } from "node:child_process";
|
|
import path from "node:path";
|
|
import { describe, expect, it } from "vitest";
|
|
import { assertActiveHandoff, assertAnswerFacts, assertCrashRecovered, assertWorkerIdentity } from "./chat-qualification.js";
|
|
import { runnerMatrix } from "./catalog.js";
|
|
import { buildRunnerE2EProcessEnvironment } from "./harness-env.js";
|
|
|
|
const run = { id: "old", agentId: "original", companyId: "company", runtimeMode: "native", status: "cancelled", errorCode: "issue_reassigned", finishedAt: "2026-09-21T00:01:00Z", contextSnapshot: { issueId: "task" } };
|
|
const successor = { ...run, id: "next", agentId: "successor", status: "succeeded", startedAt: "2026-09-21T00:01:01Z" };
|
|
const before = { id: "task", assigneeAgentId: "original", description: "Preserve scope", projectId: null };
|
|
const plan = { body: "Friday REFERENCE", latestRevisionId: "revision" };
|
|
const handoff = { before, after: { ...before, status: "done", assigneeAgentId: "successor" }, oldRun: run, boundary: { ...run, status: "running" }, runs: [run, successor], successorId: "successor", planBefore: plan, planAfter: plan, draft: { id: "draft", latestRevisionId: "v1", body: "REFERENCE" }, draftAfter: { id: "draft", body: "REFERENCE" }, draftRevisions: [{ id: "v1", body: "REFERENCE" }], output: { body: "REFERENCE", createdByAgentId: "successor" }, reference: "REFERENCE", audit: [{ action: "issue.reassigned", details: { source: "paperclip_runner_protocol" } }], taskIds: ["task"] };
|
|
const recovery = { boundary: { ...run, status: "running" }, failed: { ...run, status: "failed" }, runs: [{ ...run, status: "failed" }, { ...successor, agentId: "original" }], issueId: "task", prompt: "Read my brief", comments: [{ body: "Read my brief" }, { body: "REFERENCE MARKER", authorAgentId: "original", createdByRunId: "next" }], reference: "REFERENCE", marker: "MARKER", planBefore: { body: "plan MARKER", latestRevisionId: "v1" }, planAfter: { body: "plan MARKER", latestRevisionId: "v1" } };
|
|
|
|
describe("remaining native chat qualification", () => {
|
|
it("calibrates the pidfd helper against reuse and wrong-identity faults", () => {
|
|
execFileSync("python3", [path.join(import.meta.dirname, "worker-fault.test.py")], { stdio: "pipe" });
|
|
});
|
|
it("requires a stopped original worker before exactly one successor executes", () => {
|
|
expect(() => assertActiveHandoff(handoff)).not.toThrow();
|
|
expect(() => assertActiveHandoff({ ...handoff, draftAfter: { id: "draft", body: "Expanded REFERENCE" },
|
|
output: { body: "Expanded REFERENCE", createdByAgentId: "original", updatedByAgentId: "successor" } })).not.toThrow();
|
|
for (const change of [
|
|
{ boundary: { ...handoff.boundary, status: "succeeded" } },
|
|
{ oldRun: { ...run, status: "succeeded" } },
|
|
{ oldRun: { ...run, errorCode: "unrelated_cancellation" } },
|
|
{ runs: [run, { ...successor, startedAt: "2026-09-21T00:00:59Z" }] },
|
|
{ runs: [run, successor, { ...successor, id: "duplicate" }] },
|
|
{ runs: [run] },
|
|
{ taskIds: ["replacement"] },
|
|
{ planAfter: { ...plan, body: "rewritten" } },
|
|
{ draft: { body: "no saved progress" } },
|
|
{ draftAfter: { id: "replaced", body: "REFERENCE" } },
|
|
{ draftRevisions: [] },
|
|
{ draftRevisions: [{ id: "v1", body: "overwritten original" }] },
|
|
{ output: { body: "REFERENCE", createdByAgentId: "original" } },
|
|
{ audit: [] },
|
|
]) expect(() => assertActiveHandoff({ ...handoff, ...change })).toThrow();
|
|
});
|
|
it("requires successful run-attributed recovery preserving input and saved work", () => {
|
|
expect(() => assertCrashRecovered(recovery)).not.toThrow();
|
|
for (const change of [
|
|
{ boundary: { ...recovery.boundary, status: "succeeded" } },
|
|
{ failed: { ...recovery.failed, id: "unrelated" } },
|
|
{ runs: [recovery.runs[0]!, { ...successor, status: "failed" }] },
|
|
{ runs: [recovery.runs[0]!, { ...successor, contextSnapshot: { issueId: "new-chat" } }] },
|
|
{ comments: [...recovery.comments, recovery.comments[0]!] },
|
|
{ comments: [recovery.comments[0]!, { ...recovery.comments[1], createdByRunId: "old" }] },
|
|
{ comments: [recovery.comments[0]!, { ...recovery.comments[1], body: "MARKER" }] },
|
|
{ planAfter: { ...recovery.planAfter, latestRevisionId: "rewritten" } },
|
|
]) expect(() => assertCrashRecovered({ ...recovery, ...change })).toThrow();
|
|
});
|
|
it("refuses unknown PIDs, nonnative or nonlocal processes and partial run identities", () => {
|
|
const running = { id: "run-123", status: "running", runtimeMode: "native", processPid: 123456 };
|
|
expect(() => assertWorkerIdentity(running, "node runner --run-id run-123 --worker", "local")).not.toThrow();
|
|
for (const command of ["node server", "node runner --run-id run-1234", "node runner run-123 --run-id other"])
|
|
expect(() => assertWorkerIdentity(running, command, "local")).toThrow();
|
|
for (const processPid of [undefined, 0, 1, -20, process.pid, 1.2])
|
|
expect(() => assertWorkerIdentity({ ...running, processPid }, "node runner --run-id run-123", "local")).toThrow();
|
|
expect(() => assertWorkerIdentity(running, "node runner --run-id run-123", "daytona")).toThrow();
|
|
expect(() => assertWorkerIdentity({ ...running, runtimeMode: "legacy" }, "node runner --run-id run-123", "local")).toThrow();
|
|
});
|
|
it("grades factual propositions rather than matching words in misleading prose", () => {
|
|
const facts = { currentBlocker: "VENUE", confirmedAttendance: null, printingStarted: false };
|
|
const explanation = "The venue is still unconfirmed, so printing remains deferred. No attendance count is recorded.";
|
|
expect(() => assertAnswerFacts(JSON.stringify({ facts, explanation }), facts)).not.toThrow();
|
|
for (const wrong of [
|
|
{ ...facts, currentBlocker: "BUDGET" }, { ...facts, confirmedAttendance: 40 },
|
|
{ ...facts, printingStarted: true }, { ...facts, madeUpMetric: 50 }, {},
|
|
]) expect(() => assertAnswerFacts(JSON.stringify({ facts: wrong, explanation }), facts)).toThrow();
|
|
expect(() => assertAnswerFacts(JSON.stringify({ facts, explanation: "" }), facts)).toThrow();
|
|
expect(() => assertAnswerFacts("VENUE confirmedAttendance null printingStarted false", facts)).toThrow();
|
|
});
|
|
it("exposes exactly six explicit local native cells with bounded run counts", () => {
|
|
const cells = runnerMatrix.filter(c => c.suite.id === "agent-chat-qualification");
|
|
expect(cells).toHaveLength(6);
|
|
for (const cell of cells) {
|
|
expect(cell.suite.manualOnly).toBe(true);
|
|
expect(cell.environment.id).toBe("local");
|
|
expect(cell.profile.generation).toBe("native");
|
|
expect(cell.task.expectedRunCount).toBe(cell.task.id === "active-reassignment" ? 3 : 2);
|
|
expect(buildRunnerE2EProcessEnvironment({}, [cell]).PAPERCLIP_RUNNER_API_TOOLS_ENABLED)
|
|
.toBe(cell.task.id === "grounded-answer-quality" ? "true" : undefined);
|
|
}
|
|
});
|
|
});
|