Files
PaperClipAI/tests/runner-e2e/chat-flow.test.ts
DottaandPaperclip 7bc03e0acd feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat uses native runners to save plans and coordinate tasks.
> - Provider defaults differed across harnesses and could stop
unattended work at a second permission gate.
> - Agents also lacked a dedicated tool to move existing work to another
agent safely.
> - This change defaults native providers to full automatic permission
for provider tools and connected tools.
> - A guarded reassignment tool preserves task identity, stops the
previous run, and schedules the new owner once.
> - Codex and Claude chat acceptance tests now use production permission
defaults.

## Linked Issues or Issue Description

**Subsystem affected**

Native runner, ACPX Claude permission policy, task authority, and Agent
Chat acceptance tests.

**Problem or motivation**

A user can authorize an agent to save a plan or create a task, but
Claude's default provider gate can still stop that action. Reassignment
needs a dedicated operation that preserves context and avoids concurrent
owners or unintended recovery runs.

**Proposed solution**

Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to
`never`. Apply the defaults at configuration, execution, fresh-session,
resume, driver, and proxy boundaries. Keep explicit permission settings
and server-side company, claim, task-mode, and approval checks. Add
`reassign_task` with version checks, durable idempotency, audited
cancellation, and guarded successor scheduling.

**Alternatives considered**

A Paperclip-only allowlist still blocks provider tools and other
connections during unattended work. Full automatic permission is the
requested product default. Recreating a task discards its identity and
history. Updating assignment without stopping the previous run can leave
two agents working on the same task.

**Roadmap alignment**

This extends the existing planning, delegated work, governed tool
access, and recovery features. It adds no new service or schema
migration. Recent related tasks and open PRs were checked for duplicate
work.

**Additional context**

Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner
startup). The stacked legacy-adapter companion is #13693. This also
fixes the deployed-server artifact fallback needed to stage the current
runner binary.

## What Changed

- Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex
to `never`, including missing settings at direct driver and proxy entry
points. These defaults cover provider tools and connected tools.
Preserve explicitly configured restrictive modes.
- Include assigned approval reads using canonical side-effect
classifications, so verifying a recorded approval does not trigger
another provider gate. Paperclip approval decisions still enforce
controller authority.
- Carry the new permission mode through server configuration, execution
contracts, recovery identity, TypeScript, and Rust. Keep
`approve-paperclip` as an optional restricted mode, with exact SDK rules
and closed unknown requests. It is not a default.
- Add `reassign_task` to the semantic catalog, controller, mock
authority, and generated contracts.
- Guard reassignment with company authorization, expected owner and
version, protected-state checks, and durable retry receipts.
- Honor explicit backlog task creation atomically with the initial plan,
without scheduling a wake. Preserve backlog holds regardless of
dependency readiness.
- Stop active work before changing ownership. Restore the prior owner
through a guarded, idempotent wake if final handoff validation fails.
Keep intentional reassignment stops out of failure recovery. Preserve
backlog and blocked states without waking them early.
- Add authorization, concurrency, replay, stop, and permission boundary
regressions. Add Codex and Claude chat reassignment cases and run native
chat cases with production defaults.
- Clarify shared runner guidance: save plans and Paperclip documents
directly with `write_document`; create and register a local file only
when a downloadable file is requested.
- Document provider defaults and the operator choices for existing
agents.

## Verification

- Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks
passed**, with two intentional skips. [PR
checks](https://github.com/paperclipai/paperclip/pull/13686/checks).
- Greptile reviewed that exact head at **5/5**. The security reviewer
acknowledged the intended full-auto default, and the acknowledged
discussions are resolved.
- Full workspace `pnpm -r typecheck` and `pnpm build` passed locally
after rebasing onto current master. Targeted adapter/server, runner,
API, default/resume, and heartbeat configuration tests passed.
- **All six real-provider acceptance cases passed on their first
attempt, with cleanup passing:** plan handoff, task reassignment, and
backlog creation/status, each on native Claude and Codex. Evidence
records Claude's effective `approve-all` mode. [Campaign and
downloadable
evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548).
- The live campaign tested combined revision
`a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only
a heartbeat test expectation correction; application code is unchanged
from that live-tested revision.
- The campaign's result-enforcement job passed. Its separate report
publisher failed because the trusted workflow's `patchedDependencies`
configuration differs from its frozen lockfile. All six results and
screenshots remain available as GitHub artifacts. The overall manual
workflow is red for this publishing failure.
- Full-suite coverage is supplied by the passing CI partitions. The
separate unsharded local run was stopped after the corresponding CI
partitions passed; it is not counted as a completed local run.
- Reassignment tests cover stale state, cross-company access, denied
authority, cancellation failure, compensating wake, and idempotent
retries. Backlog tests verify the original creation audit, saved plan,
exact task count, and absence of task-bound runs.

## Risks

- Agents with no explicit permission mode now receive full provider tool
permission, including connected tools. This is a deliberate broad
default. Existing explicit restrictive modes still apply. Controller
authorization, company isolation, workspace boundaries, and Paperclip
governance remain in force.
- Reassignment crosses run cancellation and task ownership transactions.
Durable stop intent, revalidation, audit receipts, and guarded queue
dispatch cover interruptions and retries.
- The new permission enum requires a current runner artifact. The remote
artifact fallback uses the same resolved controller binary for upload
and execution.
- Live provider behavior remains subject to the selected model. Targeted
live results do not qualify the full catalog.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact deployment model ID and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:49:18 -05:00

417 lines
17 KiB
TypeScript

import { describe, it, expect, vi } from "vitest";
import type { AskUserQuestionsPayload } from "../../packages/shared/src/types/issue.js";
import {
assertChatBacklogCreation,
assertChatHandoff,
assertChatReassignment,
assertChatTaskHandoff,
chatQuestionPresentation,
chatRunFailure,
chatTaskCompletionFailure,
createChatIdleFailureDetector,
collectChatRunEvidence,
readRunningChatLog,
readChatOutputDocument,
isChatClarificationReply,
assertChatExecutionOutput,
isResetRun,
type ChatIssue,
type ChatRun,
} from "./chat-flow.js";
import type { RunnerApi } from "./api.js";
import { chatMarker } from "./chat-cases.js";
import { runnerMatrix } from "./catalog.js";
import { isPublicRunnerScreenshotRoute } from "./screenshot-policy.js";
import { classifyFailure, shouldRetryFailure } from "./failure-classifier.js";
const source: ChatIssue = {
id: "chat",
companyId: "co",
title: "Chat",
status: "in_review",
assigneeAgentId: "agent",
};
const task: ChatIssue = {
...source,
id: "work",
parentId: null,
projectId: "project",
};
const plan = {
body: "# Relevant plan",
latestRevisionId: "revision",
updatedAt: "2026-09-11T10:00:00Z",
};
const run: ChatRun = {
id: "run",
companyId: "co",
agentId: "agent",
status: "succeeded",
startedAt: "2026-09-11T10:00:01Z",
};
describe("chat acceptance contracts", () => {
it("accepts concrete information requests without requiring question punctuation", () => {
expect(isChatClarificationReply("What is the club name?")).toBe(true);
expect(
isChatClarificationReply(
"Before assigning the welcome-note work, please share:\n\n1. Club name and intended readers.\n2. Format, length, and tone.\n3. Required details, sender, and deadline.",
),
).toBe(true);
expect(
isChatClarificationReply("Tell me the intended audience and format."),
).toBe(true);
expect(
isChatClarificationReply(
"Thanks — before assigning the welcome-note drafting work, I need a compact brief covering:\n\n- Club and audience: club name and intended readers.\n- Purpose: welcome or next steps.\n- Required content: dates, links, and contacts.\n- Voice: tone and sender.\n- Delivery constraints: format, length, and deadline.\n- Examples or policies: existing notes and approval requirements.",
),
).toBe(true);
expect(isChatClarificationReply("I'll need your details about the audience and format.")).toBe(true);
expect(isChatClarificationReply("We need some information about the club and intended readers.")).toBe(true);
expect(isChatClarificationReply("I needed a compact brief before I assigned the work.")).toBe(false);
expect(isChatClarificationReply("I need a compact brief:")).toBe(false);
expect(isChatClarificationReply("I need information.")).toBe(false);
expect(isChatClarificationReply("I need to create the task and write the note.")).toBe(false);
expect(isChatClarificationReply("Please share:")).toBe(false);
expect(isChatClarificationReply("Please share.")).toBe(false);
expect(
isChatClarificationReply("Asked the user clarifying questions about their club."),
).toBe(false);
expect(
isChatClarificationReply("I created the task and started writing the welcome note."),
).toBe(false);
});
it("rejects superseded plan requirements in executed output, independently of plan history", () => {
expect(() => assertChatExecutionOutput("Welcome CHAT123.", "CHAT123", "DRAFT123")).not.toThrow();
expect(() => assertChatExecutionOutput("Welcome DRAFT123 and CHAT123.", "CHAT123", "DRAFT123")).toThrow();
expect(() => assertChatExecutionOutput("Welcome DRAFT123.", "CHAT123", "DRAFT123")).toThrow();
});
it("keeps chat markers literal across rich-text and Markdown boundaries", () => {
for (const prefix of ["CHAT", "DRAFT", "OLDCONTEXT"] as const) {
expect(chatMarker(prefix, "abc123-1")).toBe(`${prefix}abc1231`);
expect(chatMarker(prefix, "abc_123-1")).toMatch(/^[a-zA-Z0-9]+$/);
}
expect(chatMarker("OLDCONTEXT", "one-1")).not.toBe(
chatMarker("CHAT", "one-1"),
);
expect(chatMarker("CHAT", "one-1")).not.toBe(chatMarker("CHAT", "two-1"));
});
it("covers existing workflows and native reassignment on the chosen local profiles", () => {
const matrix = runnerMatrix.filter(
(cell) => cell.suite.id === "agent-chat",
);
expect(matrix).toHaveLength(28);
expect(new Set(matrix.map((cell) => cell.profile.id))).toEqual(
new Set([
"legacy-codex",
"legacy-claude",
"runner-codex",
"runner-acpx-claude",
]),
);
expect(new Set(matrix.map((cell) => cell.task.id)).size).toBe(8);
expect(
matrix.every(
(cell) =>
cell.environment.id === "local" &&
cell.task.expectedTerminalState.issue === "in_review",
),
).toBe(true);
});
it("rejects missing plan, chat children, wrong assignments, and execution before the plan", () => {
expect(() => assertChatHandoff(task, plan, [run], source)).not.toThrow();
for (const invalid of [
{ ...task, parentId: "chat" },
{ ...task, projectId: null },
{ ...task, assigneeAgentId: "other" },
])
expect(() => assertChatHandoff(invalid, plan, [run], source)).toThrow();
expect(() =>
assertChatHandoff(task, { ...plan, body: "" }, [run], source),
).toThrow();
expect(() =>
assertChatHandoff(
task,
{ ...plan, updatedAt: "2026-09-11T10:00:02Z" },
[run],
source,
),
).toThrow();
expect(() => assertChatHandoff(task, plan, [], source)).toThrow();
});
it("requires a plan for plan handoff, while direct requests need only normal task assignment", () => {
expect(() => assertChatTaskHandoff(task, [run], source)).not.toThrow();
expect(() =>
assertChatHandoff(task, { ...plan, body: "" }, [run], source),
).toThrow();
expect(() =>
assertChatTaskHandoff({ ...task, projectId: null }, [run], source),
).toThrow();
});
it.each(["project-description", "welcome-note", "output"])("finds committed %s output without accepting a copied plan or a claim", async (key) => {
const output = {
...plan,
id: "description-doc",
issueId: "work",
key,
body: "A completed description with CHAT123.",
createdByAgentId: "agent",
};
const get = vi.fn(async (path: string) => {
if (path === "/api/issues/work/documents")
return [{ key: "plan" }, { key }];
if (path === `/api/issues/work/documents/${key}`)
return output;
throw new Error(`Unexpected document read: ${path}`);
});
const api = { get } as Pick<RunnerApi, "get">;
await expect(readChatOutputDocument(api, "work", "CHAT123")).resolves.toBe(
output,
);
await expect(readChatOutputDocument(api, "work", "WRONG123")).rejects.toThrow(
"no non-plan output document",
);
get.mockImplementation(async () => [{ key: "plan" }]);
await expect(readChatOutputDocument(api, "work", "CHAT123")).rejects.toThrow(
"document keys: plan",
);
});
it("uses durable free-text labels, multi-selection, and the supplied submit label", () => {
const payload: AskUserQuestionsPayload = {
version: 1,
submitLabel: "Send brief",
questions: [
{
id: "audience",
prompt: "Who is it for?",
selectionMode: "multi",
required: true,
options: [
{ id: "members", label: "New members" },
{
id: "custom",
label: "Another audience or occasion",
freeText: true,
},
],
},
],
};
const presentation = chatQuestionPresentation(payload);
expect(presentation.submitLabel).toBe("Send brief");
expect(presentation.questions[0]).toMatchObject({
answerMode: "multi_select",
customAnswer: { enabled: true, label: "Another audience or occasion" },
});
const nativePayload: AskUserQuestionsPayload = {
...payload,
questionSet: {
schema: "paperclip.question_set.v1",
submitLabel: "Continue",
questions: [
{
id: "audience",
prompt: "Who is it for?",
required: true,
answerMode: "text",
},
],
},
};
expect(chatQuestionPresentation(nativePayload)).toBe(
nativePayload.questionSet,
);
});
it("retains reset events without requesting a provider log, and does not hide missing real logs", async () => {
const get = vi.fn().mockResolvedValue([{ type: "session_reset" }]);
const reset = { ...run, resultJson: { conversationReset: true } };
await expect(collectChatRunEvidence({ get }, reset)).resolves.toEqual({
runId: run.id,
log: null,
events: [{ type: "session_reset" }],
});
expect(get.mock.calls).toEqual([
[`/api/heartbeat-runs/${run.id}/events?limit=1000`],
]);
get.mockRejectedValue(new Error("Run log not found"));
await expect(collectChatRunEvidence({ get }, run)).rejects.toThrow(
"Run log not found",
);
});
it("retains events for an unstarted dependency-blocked wake without asking for a nonexistent log", async () => {
const get = vi.fn().mockResolvedValue([]);
const suppressed = { ...run, status: "cancelled", errorCode: "issue_dependencies_blocked", startedAt: null };
expect((await collectChatRunEvidence({ get }, suppressed)).log).toBeNull();
expect(get).toHaveBeenCalledTimes(1);
get.mockRejectedValue(new Error("Run log not found"));
await expect(collectChatRunEvidence({ get }, { ...suppressed, startedAt: "2026-09-18T00:00:00Z" })).rejects.toThrow("Run log not found");
await expect(collectChatRunEvidence({ get }, { ...suppressed, errorCode: "provider_transport_failed" })).rejects.toThrow("Run log not found");
});
it("waits for a newly running provider's log file without swallowing server failures", async () => {
const get = vi.fn().mockResolvedValue({ status: () => 404 });
const api = { request: { get } } as unknown as Pick<RunnerApi, "request">;
await expect(readRunningChatLog(api, "starting")).resolves.toBeUndefined();
get.mockResolvedValue({
status: () => 200,
ok: () => true,
json: async () => ({ content: "streamed reply" }),
});
await expect(readRunningChatLog(api, "running")).resolves.toBe(
"streamed reply",
);
get.mockResolvedValue({ status: () => 500, ok: () => false });
await expect(readRunningChatLog(api, "broken")).rejects.toThrow(
"log returned 500",
);
});
it("fails terminal execution errors without preempting active retries or recovery", () => {
const failed = {
...run,
status: "failed",
error: "provider rejected request",
};
expect(chatTaskCompletionFailure(task, [failed])).toContain(
"provider rejected request",
);
expect(
chatTaskCompletionFailure(task, [failed, { ...run, status: "queued" }]),
).toBeUndefined();
expect(
chatTaskCompletionFailure({ ...task, scheduledRetry: { id: "retry" } }, [
failed,
]),
).toBeUndefined();
expect(
chatTaskCompletionFailure(
{ ...task, activeRecoveryAction: { id: "recovery" } },
[failed],
),
).toBeUndefined();
});
it("fails stable contradictory idle states promptly without paid retries or transient false alarms", () => {
const detect = createChatIdleFailureDetector(3);
const settled = {
resolved: true,
status: "blocked",
conversationState: "waiting",
providerRunCount: 3,
activeRuns: [] as string[],
};
expect(detect(settled)).toBeUndefined();
expect(detect({ ...settled, activeRuns: ["running"] })).toBeUndefined();
expect(detect(settled)).toBeUndefined();
const failure = detect(settled);
expect(failure).toContain("chat_idle_state_invariant");
expect(classifyFailure(failure)).toBe("candidate_failure");
expect(shouldRetryFailure(classifyFailure(failure))).toBe(false);
expect(detect({ ...settled, status: "in_review" })).toBeUndefined();
expect(detect(settled)).toBeUndefined();
expect(detect({ ...settled, providerRunCount: 2 })).toBeUndefined();
expect(detect({ ...settled, status: "in_progress" })).toBeUndefined();
expect(detect(settled)).toBeUndefined();
});
it("fails promptly on terminal provider failures while permitting only expected cancellations", () => {
expect(chatRunFailure([run])).toBeUndefined();
expect(chatRunFailure([{ ...run, status: "running" }])).toBeUndefined();
expect(
chatRunFailure([
{
...run,
status: "failed",
errorCode: "permission_denied",
error: "sandbox unavailable",
},
]),
).toContain("run run failed (permission_denied): sandbox unavailable");
expect(chatRunFailure([{ ...run, status: "cancelled" }])).toContain(
"cancelled",
);
expect(
chatRunFailure([{ ...run, status: "cancelled" }], true),
).toBeUndefined();
});
it("separates reset control runs from provider runs without treating failures as resets", () => {
expect(isResetRun(run)).toBe(false);
expect(isResetRun({ ...run, status: "failed" })).toBe(false);
expect(
isResetRun({ ...run, contextSnapshot: { conversationReset: true } }),
).toBe(true);
});
it("only publishes screenshots of the exact disposable chat", () => {
const target = {
issuePrefix: "E2E",
issueId: "chat",
issueIdentifier: null,
chatAgentId: "fixture-agent",
};
expect(
isPublicRunnerScreenshotRoute(
"http://127.0.0.1:3199/E2E/chats/fixture-agent",
target,
),
).toBe(true);
expect(
isPublicRunnerScreenshotRoute(
"http://127.0.0.1:3199/E2E/chats/another-agent",
target,
),
).toBe(false);
expect(
isPublicRunnerScreenshotRoute(
"https://example.com/E2E/chats/fixture-agent",
target,
),
).toBe(false);
});
});
describe("reassignment outcome oracle", () => {
const evidence = () => ({ readyId: "ready", queuedId: "queued", teammateId: "riley",
tasks: [{ id: "ready", companyId: "co", title: "Ready", status: "done", assigneeAgentId: "riley" }, { id: "queued", companyId: "co", title: "Later", status: "backlog", assigneeAgentId: "riley" }],
runs: [{ id: "successor", companyId: "co", agentId: "riley", status: "succeeded", runtimeMode: "native", contextSnapshot: { issueId: "ready" } }],
audit: [{ action: "issue.reassigned", details: { source: "paperclip_runner_protocol" } }], outputBody: "Launch CHECK123", marker: "CHECK123",
});
it("accepts persisted ownership and exactly one successful successor", () => {
expect(() => assertChatReassignment(evidence())).not.toThrow();
});
it.each(["owner", "duplicate", "missing-run", "missing-audit", "backlog-started", "missing-output"])("rejects %s evidence", defect => {
const data = evidence();
if (defect === "owner") data.tasks[0]!.assigneeAgentId = "old";
if (defect === "duplicate") data.tasks.push({ ...data.tasks[0]!, id: "replacement" });
if (defect === "missing-run") data.runs = [];
if (defect === "missing-audit") data.audit = [];
if (defect === "backlog-started") data.runs.push({ ...data.runs[0]!, id: "early", contextSnapshot: { issueId: "queued" } });
if (defect === "missing-output") data.outputBody = "I reassigned it";
expect(() => assertChatReassignment(data)).toThrow();
});
});
describe("backlog creation outcome oracle", () => {
const evidence = () => ({
tasks: [{ id: "held", companyId: "co", title: "Later", status: "backlog", parentId: null, assigneeAgentId: "planner" }],
runs: [] as ChatRun[], ownerId: "planner", marker: "PLAN123",
plan: { body: "Three steps PLAN123", latestRevisionId: "revision-1", updatedAt: "2026-09-19T00:00:00Z" },
activity: [{ action: "issue.created", details: { status: "backlog", source: "paperclip_runner_protocol" } }],
});
it("accepts one planned backlog task with no execution", () => {
expect(() => assertChatBacklogCreation(evidence())).not.toThrow();
});
it("does not mistake the creating conversation for task execution", () => {
const data = evidence();
data.runs.push({ id: "creator", companyId: "co", agentId: "planner", status: "succeeded", contextSnapshot: { issueId: "conversation" } });
expect(() => assertChatBacklogCreation(data)).not.toThrow();
});
it.each(["duplicate", "status", "started-then-stopped", "corrected-after-creation", "owner", "plan"])("rejects %s", defect => {
const data = evidence();
if (defect === "duplicate") data.tasks.push({ ...data.tasks[0]!, id: "duplicate" });
if (defect === "status") data.tasks[0]!.status = "todo";
if (defect === "started-then-stopped") data.runs.push({ id: "early", companyId: "co", agentId: "planner", status: "cancelled", contextSnapshot: { issueId: "held" } });
if (defect === "corrected-after-creation") data.activity[0]!.details.status = "todo";
if (defect === "owner") data.tasks[0]!.assigneeAgentId = "other";
if (defect === "plan") data.plan.body = "I saved a plan";
expect(() => assertChatBacklogCreation(data)).toThrow();
});
});