Files
PaperClipAI/tests/runner-e2e/chat-stories.test.ts
DottaandPaperclip 846336e5a0 test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path

> - Paperclip lets people manage agents through ongoing conversations.
> - Chat users can change instructions while a provider is already
working.
> - Existing chat evals wait for each turn to settle before the next
message.
> - They cannot prove delivery during active work or the saved effect of
a correction.
> - Existing fixtures also enable Agent Chat through the API rather than
the settings UI.
> - This PR adds bounded browser workflows and checks their persisted
outcomes.

## Linked Issues or Issue Description

Refs #13741, #13752, #13750.

**What happened?**

The chat suites cover planning, delegation, status, and recovery. They
lack active-turn follow-ups and the experimental settings lifecycle. A
sequential conversation can pass even if messages sent during work are
lost.

**Expected behavior**

A follow-up submitted during a provider turn survives and affects the
final reply. A changed launch day appears in the saved plan. Disabling
Agent Chat rejects new messages while preserving history; re-enabling
resumes the same conversation.

**Steps to reproduce**

Run the explicit `agent-chat-stories` suite. It selects three local
cases for each native Claude and Codex profile. An ordinary provider
command waits for a fixture brief file so the browser can send the
follow-up at an observed active-run boundary.

## What Changed

- Add six opt-in Product E2E cells for settings, active follow-ups, and
plan corrections.
- Drive experimental settings through the UI and verify disabled sends
are rejected by the public API.
- Use a bounded file wait in the actual isolated agent workspace, with
provider-written readiness and an undisclosed brief reference.
- Grade persisted user messages, final replies, native run outcomes, and
exact saved plan fields.
- Accept active-turn steering or one queued successor; reject lost
input, duplicate input, and stale outputs.
- Allow one steered run or two sequential runs throughout the shared
harness, while preserving exact counts for other cases.
- Require a single marker-bearing response attributed to the final
provider run.
- Unload the development browser client before restarting the server,
avoiding reconnect/navigation races without weakening the post-restart
memory check.
- Add browser regressions for restart isolation and asynchronously saved
settings switches.
- Document prepared-agent setup, native onboarding limits, and the
separate API-tool rollout gate.

## Verification

- Eval TypeScript check passed.
- Eval support suite: 436 tests passed in 39 files.
- New oracle calibration: six tests passed, including plausible invalid
outcomes.
- Browser support regressions: seven tests passed; the restart
regression was observed failing before the fix.
- Catalog discovery selects exactly six local native cases and leaves
default paid selection unchanged.
- [Consolidated existing native chat
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/):
master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a
browser navigation timeout across restart; the page request returned 200
and the chat rendered.
- [Nine targeted restart/replay
cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/)
passed on `1fe2fe275`, including the original failure, across native
Claude/Codex and local/Daytona; all cleanup passed.
- [Initial six-story
campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817)
retained all six failures: asynchronous switch assertions, unavailable
fixture paths, and rich-text escaping in raw command comparisons. The
corrected fixtures preserve the same behavioral assertions.
- [Six-story campaign
v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/)
on `8232773a0`: 4/6 passed (both settings cases and both Claude
interruptions). Codex could not see the host-temp fixture outside its
workspace; this failed before follow-up delivery was exercised.
- [Four affected interruption
cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/)
all passed, including cleanup, on definition v3 / `ad6ac0545`. Files
live inside the actual agent workspace and the observed run workspace is
verified. Both providers saved Friday in the real plan with the
undisclosed brief reference; follow-ups persisted while the original run
was active. Together with both unchanged settings cases from v2, all six
new scenario variants have passing live evidence.
- Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful
checks, two intentional skips, zero pending/failing checks; mergeable
and clean. Fresh Greptile 5/5, zero unresolved findings.
- Full typecheck, tests, build, and browser CI passed remotely. One
earlier head encountered a signoff-policy browser timing failure; the
final head passed that shard.
- Local pnpm wrapper could not fetch its version/signature metadata in
the restricted environment; local eval checks used the installed Node
executables. Repo-wide validation was completed by GitHub Actions.

## Risks

These are eval-only changes. The file wait is a timing fixture in the
isolated agent workspace, not a production runner hook. Native Codex
host-filesystem isolation stays unchanged. It has a two-minute limit and
is released in `finally`. The prepared-agent settings case is not full
native onboarding: the wizard currently offers legacy adapters. The
disabled-entry assertion uses full document navigation, which clears the
prior React Query cache; preserved history is checked through the public
API and re-enabled chat. No production prompt, rollout default, adapter
behavior, or credential policy changes. Active-task reassignment and
worker-crash recovery remain outside these new cases.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 16:18:06 -05:00

74 lines
4.9 KiB
TypeScript

import { describe, expect, it } from "vitest";
import { assertInterruptedChat, prepareChatBrief } from "./chat-stories.js";
import { mkdtemp, readFile, rm, writeFile } from "node:fs/promises";
import { tmpdir } from "node:os";
import path from "node:path";
import { execFile } from "node:child_process";
import { promisify } from "node:util";
import { runnerMatrix } from "./catalog.js";
import { buildRunnerE2EProcessEnvironment } from "./harness-env.js";
const run = { id: "run", companyId: "company", agentId: "agent", status: "succeeded", runtimeMode: "native", contextSnapshot: { issueId: "chat" } };
const valid = {
first: "Read the brief", followup: "Change the launch day", reference: "BRIEF123", marker: "UPDATED123", issueId: "chat",
boundaryRun: { ...run, status: "running" }, activeAtFollowup: { ...run, status: "running" },
comments: [{ id: "first", body: "Read the brief" }, { id: "followup", body: "Change the launch day" },
{ id: "answer", authorAgentId: "agent", createdByRunId: "run", body: "BRIEF123 UPDATED123" }],
runs: [run], revisedPlan: JSON.stringify({ launchDay: "Friday", reference: "BRIEF123", revision: "UPDATED123" }),
};
describe("active chat follow-up oracle", () => {
it("creates the missing fixture directory and waits for the host-supplied brief", async () => {
const root = await mkdtemp(path.join(tmpdir(), "chat-brief-test-"));
try {
const { gate, ready, scriptPath } = await prepareChatBrief(path.join(root, "new workspace"), "fixture");
const command = promisify(execFile)(process.execPath, [scriptPath], { timeout: 5_000 });
try {
await expect.poll(() => readFile(ready, "utf8").catch(() => "")).toBe("waiting");
} finally {
await writeFile(gate, "BRIEF fixture result");
expect((await command).stdout.trim()).toBe("BRIEF fixture result");
}
} finally {
await rm(root, { recursive: true, force: true });
}
});
it("accepts either steering or one queued successor, with the saved correction", () => {
expect(() => assertInterruptedChat(valid)).not.toThrow();
expect(() => assertInterruptedChat({ ...valid, runs: [run, { ...run, id: "successor" }], comments: [...valid.comments.slice(0, 2), { ...valid.comments[2]!, createdByRunId: "successor" }] })).not.toThrow();
});
it("rejects follow-ups sent after the active boundary", () => {
expect(() => assertInterruptedChat({ ...valid, boundaryRun: run })).toThrow();
expect(() => assertInterruptedChat({ ...valid, activeAtFollowup: run })).toThrow();
expect(() => assertInterruptedChat({ ...valid, activeAtFollowup: { ...valid.activeAtFollowup, id: "other" } })).toThrow();
});
it("rejects lost or duplicated input, missing file evidence, and stale plans", () => {
expect(() => assertInterruptedChat({ ...valid, comments: [...valid.comments, { ...valid.comments[2]!, id: "duplicate-reply" }] })).toThrow();
expect(() => assertInterruptedChat({ ...valid, comments: [...valid.comments.slice(0, 2), { ...valid.comments[2]!, createdByRunId: "unrelated" }] })).toThrow();
expect(() => assertInterruptedChat({ ...valid, comments: valid.comments.filter(c => c.id !== "followup") })).toThrow();
expect(() => assertInterruptedChat({ ...valid, comments: [...valid.comments, { id: "duplicate", body: valid.followup }] })).toThrow();
expect(() => assertInterruptedChat({ ...valid, comments: [...valid.comments.slice(0, 2), { id: "answer", authorAgentId: "agent", body: "UPDATED123" }] })).toThrow();
expect(() => assertInterruptedChat({ ...valid, revisedPlan: valid.revisedPlan.replace("Friday", "Monday") })).toThrow();
expect(() => assertInterruptedChat({ ...valid, revisedPlan: "The updated plan is saved." })).toThrow();
});
it("rejects failed, unfinished, duplicated, or unrelated execution", () => {
for (const status of ["running", "queued", "failed", "cancelled"])
expect(() => assertInterruptedChat({ ...valid, runs: [{ ...run, status }] })).toThrow();
expect(() => assertInterruptedChat({ ...valid, runs: [] })).toThrow();
expect(() => assertInterruptedChat({ ...valid, runs: [run, run, run] })).toThrow();
expect(() => assertInterruptedChat({ ...valid, runs: [run, run] })).toThrow();
expect(() => assertInterruptedChat({ ...valid, runs: [{ ...run, id: "other" }] })).toThrow();
expect(() => assertInterruptedChat({ ...valid, runs: [{ ...run, contextSnapshot: { issueId: "other" } }] })).toThrow();
});
it("keeps setup and interruption stories native, local, opt-in, and outside API-tool overrides", () => {
const cells = runnerMatrix.filter(cell => cell.suite.id === "agent-chat-stories");
expect(cells).toHaveLength(6);
for (const cell of cells) {
expect(cell.suite.manualOnly).toBe(true);
expect(cell.profile.generation).toBe("native");
expect(cell.environment.id).toBe("local");
expect(buildRunnerE2EProcessEnvironment({}, [cell]).PAPERCLIP_RUNNER_API_TOOLS_ENABLED).toBeUndefined();
}
});
});