mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip lets people manage agents through ongoing conversations. > - Chat users can change instructions while a provider is already working. > - Existing chat evals wait for each turn to settle before the next message. > - They cannot prove delivery during active work or the saved effect of a correction. > - Existing fixtures also enable Agent Chat through the API rather than the settings UI. > - This PR adds bounded browser workflows and checks their persisted outcomes. ## Linked Issues or Issue Description Refs #13741, #13752, #13750. **What happened?** The chat suites cover planning, delegation, status, and recovery. They lack active-turn follow-ups and the experimental settings lifecycle. A sequential conversation can pass even if messages sent during work are lost. **Expected behavior** A follow-up submitted during a provider turn survives and affects the final reply. A changed launch day appears in the saved plan. Disabling Agent Chat rejects new messages while preserving history; re-enabling resumes the same conversation. **Steps to reproduce** Run the explicit `agent-chat-stories` suite. It selects three local cases for each native Claude and Codex profile. An ordinary provider command waits for a fixture brief file so the browser can send the follow-up at an observed active-run boundary. ## What Changed - Add six opt-in Product E2E cells for settings, active follow-ups, and plan corrections. - Drive experimental settings through the UI and verify disabled sends are rejected by the public API. - Use a bounded file wait in the actual isolated agent workspace, with provider-written readiness and an undisclosed brief reference. - Grade persisted user messages, final replies, native run outcomes, and exact saved plan fields. - Accept active-turn steering or one queued successor; reject lost input, duplicate input, and stale outputs. - Allow one steered run or two sequential runs throughout the shared harness, while preserving exact counts for other cases. - Require a single marker-bearing response attributed to the final provider run. - Unload the development browser client before restarting the server, avoiding reconnect/navigation races without weakening the post-restart memory check. - Add browser regressions for restart isolation and asynchronously saved settings switches. - Document prepared-agent setup, native onboarding limits, and the separate API-tool rollout gate. ## Verification - Eval TypeScript check passed. - Eval support suite: 436 tests passed in 39 files. - New oracle calibration: six tests passed, including plausible invalid outcomes. - Browser support regressions: seven tests passed; the restart regression was observed failing before the fix. - Catalog discovery selects exactly six local native cases and leaves default paid selection unchanged. - [Consolidated existing native chat report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/): master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a browser navigation timeout across restart; the page request returned 200 and the chat rendered. - [Nine targeted restart/replay cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/) passed on `1fe2fe275`, including the original failure, across native Claude/Codex and local/Daytona; all cleanup passed. - [Initial six-story campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817) retained all six failures: asynchronous switch assertions, unavailable fixture paths, and rich-text escaping in raw command comparisons. The corrected fixtures preserve the same behavioral assertions. - [Six-story campaign v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/) on `8232773a0`: 4/6 passed (both settings cases and both Claude interruptions). Codex could not see the host-temp fixture outside its workspace; this failed before follow-up delivery was exercised. - [Four affected interruption cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/) all passed, including cleanup, on definition v3 / `ad6ac0545`. Files live inside the actual agent workspace and the observed run workspace is verified. Both providers saved Friday in the real plan with the undisclosed brief reference; follow-ups persisted while the original run was active. Together with both unchanged settings cases from v2, all six new scenario variants have passing live evidence. - Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful checks, two intentional skips, zero pending/failing checks; mergeable and clean. Fresh Greptile 5/5, zero unresolved findings. - Full typecheck, tests, build, and browser CI passed remotely. One earlier head encountered a signoff-policy browser timing failure; the final head passed that shard. - Local pnpm wrapper could not fetch its version/signature metadata in the restricted environment; local eval checks used the installed Node executables. Repo-wide validation was completed by GitHub Actions. ## Risks These are eval-only changes. The file wait is a timing fixture in the isolated agent workspace, not a production runner hook. Native Codex host-filesystem isolation stays unchanged. It has a two-minute limit and is released in `finally`. The prepared-agent settings case is not full native onboarding: the wizard currently offers legacy adapters. The disabled-entry assertion uses full document navigation, which clears the prior React Query cache; preserved history is checked through the public API and re-enabled chat. No production prompt, rollout default, adapter behavior, or credential policy changes. Active-task reassignment and worker-crash recovery remain outside these new cases. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
161 lines
11 KiB
TypeScript
161 lines
11 KiB
TypeScript
import { expect } from "@playwright/test";
|
|
import { mkdir, readFile, writeFile } from "node:fs/promises";
|
|
import path from "node:path";
|
|
import { randomUUID } from "node:crypto";
|
|
import { sendChatMessage, type ChatFlowInput, type ChatIssue, type ChatRun } from "./chat-flow.js";
|
|
import { setSettingsToggle } from "./settings-toggle.js";
|
|
import { resolveDefaultAgentWorkspaceDir } from "../../server/src/home-paths.js";
|
|
|
|
type Comment = { id: string; body: string; authorAgentId?: string; createdByRunId?: string };
|
|
type Context = {
|
|
input: ChatFlowInput; marker: string; issue(): ChatIssue;
|
|
idle(count: number): Promise<void>; allRuns(): Promise<ChatRun[]>; comments(): Promise<Comment[]>;
|
|
};
|
|
|
|
/** Seed only the disabled setting; the browser performs the user's opt-in. */
|
|
export async function enableChatThroughSettings(input: ChatFlowInput) {
|
|
await input.api.patch("/api/instance/settings/experimental", { enableAgentChat: false, enableClassicTaskInterface: false });
|
|
await setChatEnabled(input, true);
|
|
}
|
|
|
|
async function setChatEnabled(input: ChatFlowInput, enabled: boolean) {
|
|
await input.page.goto(`/${input.fixtures.company.issuePrefix}/company/settings/instance/experimental`, { waitUntil: "domcontentloaded" });
|
|
const toggle = input.page.getByRole("switch", { name: "Toggle agent chat experimental setting" });
|
|
await setSettingsToggle(toggle, enabled);
|
|
await expect.poll(async () => (await input.api.get<{ enableAgentChat: boolean }>("/api/instance/settings/experimental")).enableAgentChat).toBe(enabled);
|
|
}
|
|
|
|
export function assertInterruptedChat(input: {
|
|
first: string; followup: string; reference: string; marker: string; issueId: string;
|
|
boundaryRun: ChatRun; activeAtFollowup: ChatRun; comments: Comment[]; runs: ChatRun[];
|
|
revisedPlan?: string;
|
|
}) {
|
|
expect(input.boundaryRun.status).toBe("running");
|
|
expect(input.activeAtFollowup).toMatchObject({ id: input.boundaryRun.id, status: "running" });
|
|
for (const message of [input.first, input.followup]) {
|
|
expect(input.comments.filter(comment => !comment.authorAgentId && comment.body === message)).toHaveLength(1);
|
|
}
|
|
const replies = input.comments.filter(comment => comment.authorAgentId);
|
|
const final = replies.at(-1)!;
|
|
expect(final).toBeTruthy();
|
|
expect(final.body).toContain(input.reference);
|
|
expect(final.body).toContain(input.marker);
|
|
expect(replies.filter(reply => reply.body.includes(input.marker))).toHaveLength(1);
|
|
expect(input.runs.length).toBeGreaterThanOrEqual(1);
|
|
expect(input.runs.length).toBeLessThanOrEqual(2);
|
|
expect(input.runs.every(run => run.status === "succeeded" && run.runtimeMode === "native")).toBe(true);
|
|
expect(new Set(input.runs.map(run => run.id)).size).toBe(input.runs.length);
|
|
expect(input.runs.every(run => run.agentId === input.boundaryRun.agentId && run.contextSnapshot?.issueId === input.issueId)).toBe(true);
|
|
expect(final.authorAgentId).toBe(input.boundaryRun.agentId);
|
|
const responseRun = input.runs.length === 1
|
|
? input.boundaryRun
|
|
: input.runs.find(run => run.id !== input.boundaryRun.id)!;
|
|
expect(final.createdByRunId).toBe(responseRun.id);
|
|
expect(input.runs.some(run => run.id === input.boundaryRun.id)).toBe(true);
|
|
if (input.revisedPlan !== undefined) {
|
|
const body = input.revisedPlan.trim().replace(/^```(?:json)?\s*/, "").replace(/\s*```$/, "");
|
|
expect(JSON.parse(body)).toEqual({ launchDay: "Friday", reference: input.reference, revision: input.marker });
|
|
}
|
|
}
|
|
|
|
export async function prepareChatBrief(workspacePath: string, nonce: string) {
|
|
await mkdir(workspacePath, { recursive: true });
|
|
const gate = path.join(workspacePath, `chat-brief-${nonce}.txt`);
|
|
const ready = `${gate}.waiting`;
|
|
const scriptPath = path.join(workspacePath, `wait-for-brief-${nonce}.cjs`);
|
|
const script = `const fs=require("node:fs");fs.writeFileSync(${JSON.stringify(ready)},"waiting");const end=Date.now()+120000;const timer=setInterval(()=>{if(fs.existsSync(${JSON.stringify(gate)})){console.log(fs.readFileSync(${JSON.stringify(gate)},"utf8"));clearInterval(timer)}else if(Date.now()>end){clearInterval(timer);process.exitCode=1}},100);`;
|
|
await writeFile(scriptPath, script, "utf8");
|
|
return { gate, ready, scriptPath };
|
|
}
|
|
|
|
export async function runChatInterruption(context: Context & { refreshIssue(): Promise<void> }) {
|
|
const { input, marker, idle, allRuns, comments } = context;
|
|
// Unprojected host /tmp files are intentionally hidden from native Codex.
|
|
// These projectless chats use the normal agent-home workspace, not the
|
|
// harness's separate project fixture directory. Keep the asset inside it.
|
|
const agentWorkspace = resolveDefaultAgentWorkspaceDir(input.fixtures.agent.id);
|
|
const relative = path.relative(path.dirname(input.workspacePath), agentWorkspace);
|
|
if (relative.startsWith("..") || path.isAbsolute(relative)) throw new Error("Chat fixture workspace escaped the isolated instance");
|
|
const { gate, ready, scriptPath } = await prepareChatBrief(agentWorkspace, input.nonce);
|
|
const reference = `BRIEF${randomUUID().replaceAll("-", "")}`;
|
|
// Real provider tool execution waits on an ordinary fixture file. No runner
|
|
// hooks, provider results, database records, or task outcomes are fabricated.
|
|
// Keep JavaScript in a seeded fixture file: rich-text input serializes raw
|
|
// operators as Markdown escapes, which should not alter a timing fixture.
|
|
const command = `node ${scriptPath}`;
|
|
const first = `Our launch is planned for Monday. First run this bounded command to wait for the brief reference file I am supplying: ${command}\nAfter it returns, acknowledge the reference from the file here. This is discussion only; do not create projects or tasks.`;
|
|
const revise = input.execution.task.id === "revise-while-running";
|
|
const followup = revise
|
|
? `Change the launch day to Friday. Save the plan on this conversation as JSON with exactly launchDay, reference (from the brief file), and revision ("${marker}"). Then reply here with the brief reference and ${marker}. Do not create projects or execution tasks.`
|
|
: `After reading the brief, reply here with its reference and ${marker}. Keep this in the current conversation; do not create projects or tasks.`;
|
|
let boundaryRun: ChatRun | undefined;
|
|
let activeAtFollowup: ChatRun | undefined;
|
|
try {
|
|
await sendChatMessage(input.page, first);
|
|
await expect.poll(async () => {
|
|
if (await readFile(ready, "utf8").catch(() => "") !== "waiting") return false;
|
|
boundaryRun = (await allRuns()).find(run => run.status === "running");
|
|
return Boolean(boundaryRun);
|
|
}, { timeout: 120_000 }).toBe(true);
|
|
expect(boundaryRun!.contextSnapshot?.paperclipWorkspace).toMatchObject({ cwd: agentWorkspace });
|
|
await context.refreshIssue();
|
|
await sendChatMessage(input.page, followup);
|
|
await expect.poll(async () => (await comments()).filter(comment => !comment.authorAgentId && comment.body === followup).length,
|
|
{ timeout: 30_000 }).toBe(1);
|
|
activeAtFollowup = await input.api.get<ChatRun>(`/api/heartbeat-runs/${boundaryRun!.id}`);
|
|
await input.evidence("chat-interruption-boundary.json", { first, followup, boundaryRun, activeAtFollowup, comments: await comments() });
|
|
expect(activeAtFollowup.status, "Follow-up must persist while the first provider run is active").toBe("running");
|
|
await writeFile(gate, reference, "utf8");
|
|
await expect.poll(async () => (await comments()).filter(comment => comment.authorAgentId).at(-1)?.body ?? "", { timeout: 240_000 }).toContain(marker);
|
|
await idle(1); // Providers may queue a new run or steer the current run.
|
|
const plan = revise ? await input.api.get<{ body: string }>(`/api/issues/${context.issue().id}/documents/plan`) : undefined;
|
|
const evidence = { first, followup, reference, marker, issueId: context.issue().id, boundaryRun: boundaryRun!, activeAtFollowup,
|
|
comments: await comments(), runs: await allRuns(), ...(plan ? { revisedPlan: plan.body } : {}) };
|
|
await input.evidence("chat-interruption.json", evidence);
|
|
assertInterruptedChat(evidence);
|
|
expect(await input.api.get(`/api/companies/${input.fixtures.company.id}/issues`)).toEqual([]);
|
|
expect(await input.api.get(`/api/companies/${input.fixtures.company.id}/projects`)).toEqual([]);
|
|
} finally {
|
|
// Always release a provider that is still waiting, including failed UI runs.
|
|
await writeFile(gate, reference, "utf8");
|
|
const chatPath = `/api/companies/${input.fixtures.company.id}/chats/${input.fixtures.agent.id}`;
|
|
const current = await input.api.get<ChatIssue | null>(chatPath).catch(() => null);
|
|
await input.evidence("chat-interruption-state.json", {
|
|
issue: current,
|
|
runs: await allRuns().catch(error => ({ readError: String(error) })),
|
|
comments: current ? await input.api.get(`/api/issues/${current.id}/comments?order=asc`).catch(error => ({ readError: String(error) })) : [],
|
|
plan: current && revise ? await input.api.get(`/api/issues/${current.id}/documents/plan`).catch(error => ({ readError: String(error) })) : null,
|
|
});
|
|
}
|
|
}
|
|
|
|
export async function runChatSettingsLifecycle(context: Context) {
|
|
const { input, idle, comments, allRuns, marker } = context;
|
|
const remembered = `REMEMBER${marker}`;
|
|
const route = `/${input.fixtures.company.issuePrefix}/chats/${input.fixtures.agent.id}`;
|
|
await sendChatMessage(input.page, `Remember ${remembered} in this conversation. Acknowledge only; no tasks or projects.`);
|
|
await idle(1);
|
|
const before = { issue: context.issue(), comments: await comments(), runs: await allRuns() };
|
|
await setChatEnabled(input, false);
|
|
// Full document navigation clears the client query cache. Unlike SPA
|
|
// navigation, this verifies the disabled entry page without cached history.
|
|
await input.page.goto(route, { waitUntil: "domcontentloaded" });
|
|
await expect(input.page.getByText(/Agent Chat is disabled/)).toBeVisible();
|
|
const blocked = await input.api.request.post(`/api/issues/${before.issue.id}/comments`, { data: { body: `DISABLED${marker}` } });
|
|
expect(blocked.status()).toBe(404);
|
|
expect(await comments()).toEqual(before.comments);
|
|
expect((await allRuns()).map(run => ({ id: run.id, status: run.status })))
|
|
.toEqual(before.runs.map(run => ({ id: run.id, status: run.status })));
|
|
await setChatEnabled(input, true);
|
|
await input.page.goto(route, { waitUntil: "domcontentloaded" });
|
|
await sendChatMessage(input.page, `What reference did I ask you to remember? Reply with it and ${marker}. No new work.`);
|
|
await idle(2);
|
|
expect(context.issue().id).toBe(before.issue.id);
|
|
expect(context.issue().conversationSessionGeneration).toBe(before.issue.conversationSessionGeneration);
|
|
const reply = (await comments()).filter(comment => comment.authorAgentId).at(-1)?.body ?? "";
|
|
expect(reply).toContain(remembered);
|
|
expect(reply).toContain(marker);
|
|
expect(await input.api.get(`/api/companies/${input.fixtures.company.id}/issues`)).toEqual([]);
|
|
await input.evidence("chat-settings-lifecycle.json", { before, after: { issue: context.issue(), comments: await comments(), runs: await allRuns() } });
|
|
}
|