Files
PaperClipAI/tests/runner-e2e/chat-cases.ts
T
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00

69 lines
3.6 KiB
TypeScript

import type { RunnerTaskFixture } from "./types.js";
// Markers cross the rich-text composer and Markdown persistence boundary.
// Alphanumeric text has identical visible and stored representations.
export function chatMarker(
prefix: "CHAT" | "DRAFT" | "OLDCONTEXT",
nonce: string,
) {
return `${prefix}${nonce.replace(/[^a-zA-Z0-9]/g, "")}`;
}
export const CHAT_CASES = [
["create-backlog", "Save a planned task without starting work", 2],
["reassign-task", "Reassign existing work and preserve queued context", 2],
["continuity-restart", "Conversation continuity across restart", 3],
["new-session", "Fresh context within preserved history", 2],
["stop-new-resume", "Stop, reset, and resume", 3],
["plan-handoff", "Draft, revise, approve, and hand off a plan", 4],
["clarify-reuse", "Clarify and reuse an existing project", 3],
["multi-repository", "Create a project with multiple repository URLs", 2],
] as const;
export type ChatCase = (typeof CHAT_CASES)[number][0];
const HARDENING_CASES = [
["stop-startup-new-resume", "Stop during startup, reset, and resume", 3],
["hire-delegate-reuse", "Hire through chat, delegate, and reuse the same teammate", 5],
["blocked-status-review", "Read the actual blocker and hand source material to a reviewer", 4],
["committed-send-retry", "Recover a lost send acknowledgement without repeating committed work", 2],
...CHAT_CASES.filter(([id]) => ["stop-new-resume", "continuity-restart"].includes(id)),
] as const;
function buildChatTasks(definitions: readonly (readonly [string, string, number])[]): readonly RunnerTaskFixture[] {
return definitions.map(
([id, label, expectedRunCount]) => ({
id,
label,
groups: ["chat"],
flow: "agent_chat",
workMode: "standard",
expectedRunCount, // Provider turns, including cancelled turns and handed-off work; reset runs are separate.
attemptTimeoutMs: { local: 15 * 60_000, daytona: 15 * 60_000 },
expectedTerminalState: { issue: "in_review", run: "succeeded" },
buildTitle: (nonce) => `Chat acceptance ${id} ${nonce}`,
buildPrompt: (nonce) => `Let's discuss ${nonce}.`,
buildVisibleMarker: (nonce) => chatMarker("CHAT", nonce),
buildMatchers: () => [{ kind: "issue_status", expected: "in_review" }],
}),
);
}
export const chatTasks = buildChatTasks(CHAT_CASES);
export const chatHardeningTasks = buildChatTasks(HARDENING_CASES);
export const chatStoryTasks = buildChatTasks([
["enable-disable-resume", "Enable Agent Chat, pause access, and resume preserved history", 2],
["followup-while-running", "Deliver a follow-up while a provider turn is running", 2],
["revise-while-running", "Change instructions during active work and save the updated plan", 2],
]).map(task => ({ ...task, ...(task.id === "enable-disable-resume" ? {} : { minimumExpectedRunCount: 1 }) }));
export function chatNeedsApiTools(suiteId: string, caseId: string): boolean {
return (suiteId === "agent-chat-qualification" && caseId === "grounded-answer-quality") || suiteId === "agent-chat-hardening" && ["hire-delegate-reuse", "blocked-status-review"].includes(caseId);
}
export function isManagedHiringCase(suiteId: string, caseId: string): boolean {
return (suiteId === "everyday-workflows" && caseId === "hire-reuse") ||
(suiteId === "agent-chat-hardening" && caseId === "hire-delegate-reuse");
}
export const chatQualificationTasks = buildChatTasks([
["active-reassignment", "Reassign an executing task and preserve its saved work", 3],
["worker-crash-retry", "Recover from worker process loss through visible Retry", 2],
["grounded-answer-quality", "Ground status, correct stale claims, and acknowledge uncertainty", 2],
]);