Files
PaperClipAI/tests/runner-e2e/blocker-flow.ts
T
DottaandPaperclip d30b03bd8c test: add persistent E2E coverage for human blocker decisions (#14707)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Agents use the coordination skill when work needs human authority or
a scope decision.
> - PR #14188 replaced automatic manager escalation with direct blocker
handling.
> - This behavior needs real browser, server, database, and provider
tests.
> - The test must verify saved human input, task ownership, and resumed
work.
> - This pull request adds six reusable Product E2E cases and improves
the skill examples that they exercise.

## Linked Issues or Issue Description

Refs #14188.

The merged change needs repeatable behavior coverage. The new suite
tests missing administrator access, missing hiring permission, and
requester scope questions. Searches found no duplicate blocker-guidance
suite. This extends the existing eval system described in ROADMAP.md.

## What Changed

- Add the explicit-only `blocker-guidance` Product E2E suite. It has
three local scenarios on legacy Codex and legacy Claude.
- Use the production UI and public APIs to create work, save a
human-only question or confirmation, answer it after reload, and resume
the same task.
- Check requester identity, ownership history, manager activity, hiring,
saved answers, and completion. Keep direct text input as a separate UX
result.
- Save pending and final screenshots, API checkpoints, skill hashes,
provider evidence, and billing data through the existing report
pipeline.
- Isolate the Claude fixture home. Verify the served skill bytes before
dispatch so an old installed skill cannot silently replace the evaluated
skill.
- Improve the coordination and hiring skill examples. Include the
human-only policy, requester address, wake behavior, and handling of
authorized scope changes.
- Grader v5 requires the exact approved public welcome note as a new
worker comment. Browser input checks reject unwritable scope cards
before clicking, and confirmation direction must be saved in the
resolution before the worker wakes.
- Add grader calibration and browser-input tests. Update the fixture
guide and generated capability inventories.

## Verification

- `pnpm build`: passed after rebasing onto current master.
- `pnpm -r typecheck`: passed.
- `pnpm test:e2e:runner:typecheck`: passed.
- `pnpm test:e2e:runner:unit`: 742 tests passed.
- `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests
passed.
- `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells
found.
- Capability inventory and generated-contract checks: passed.
- `pnpm exec vitest run
server/src/__tests__/hiring-operational-examples.test.ts`: four tests
passed after synchronizing the generated API reference and section
anchor.
- Full general and serialized test suites: passed in CI on
`6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm
test:run` was interrupted after complete CI coverage passed; it is not
claimed as a completed local full-suite run.
- Final GitHub checks: 54 passed, two optional Storybook checks skipped.
The runtime-exposure startup test hit a 10-second readiness timeout
once, passed a targeted local reproduction, and its CI shard passed the
single retry without code changes.
- Current-head Greptile: 5/5, clean check, zero unresolved threads.
- Historical live measurement on September 29 at
`4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell
runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex
`gpt-5.6-sol` passed 8/9. These runs predate this rebase.
- Version 5 changes the scope answer to an exact approved publication
draft. The historical runs do not qualify that new requirement; the
two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec`
passed Codex and failed Claude. Claude posted the correct salary-free
sentence but omitted its required reference line from that comment,
placing the reference in a separate completion message. The
`public-welcome-note` check correctly failed. An earlier Claude
database-startup failure was retained separately; its fresh-instance
retry reached the model. This pilot is not a six-cell qualification.
- The failed Codex scope case omitted `addresseeUserId`. The strict
routing check remains. All 18 attempts had clean evidence manifests and
passed cleanup.
- To repeat with provider credentials: `pnpm test:e2e:runner -- --suite
blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is
excluded from `--all`.

## Risks

- The live suite measures variable model behavior. The retained 17/18
historical result and the current 1/2 scope pilot are not all-pass
qualifications. These paid cases are opt-in; their observed model
failures remain visible independently of deterministic CI checks.
- A separate generic task-replacement diagnostic still exposed a Claude
refusal. The ordinary cases use specific business decisions. The
diagnostic is not a standalone catalog case in this change.
- Earlier measurements included an old installed Claude skill and test
defects. Their grades remain retained and are not combined with the
three final repetitions.
- Skill examples can affect when agents ask for human input. Downstream
permission checks still apply.
- Native runners, Daytona, agent-requester routing, and real external
connection authorization are outside this suite.
- Raw provider traces and credentials remain private. No screenshots,
raw reports, secrets, workflow changes, or lockfile changes are
committed.

## Model Used

OpenAI GPT-6 through Codex assisted with this change. The exact deployed
variant and context window size are not exposed in this session. The
assistant used reasoning, repository edits, tool use, and shell
execution. The evaluated models were `gpt-5.6-sol` and
`claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:56:54 -05:00

138 lines
8.9 KiB
TypeScript

import { expect, type Page } from "@playwright/test";
import { createHash } from "node:crypto";
import { readFile } from "node:fs/promises";
import path from "node:path";
import { pollUntil, type RunnerApi } from "./api.js";
import { blockerScenario } from "./blocker-cases.js";
import { setupBlockerFixtures } from "./blocker-fixtures.js";
import { BLOCKER_GRADER_VERSION, gradeBlocker, gradeBlockerInputUx, pendingBlockerInput, type BlockerCheckpoint } from "./blocker-scoring.js";
import { answerBlockerThroughUi } from "./blocker-input.js";
import { collectChatRunEvidence } from "./chat-flow.js";
import type { LiveFixtureValues } from "./live-fixtures.js";
import type { MatrixExecution } from "./types.js";
import { createTaskThroughUi } from "./user-actions.js";
type Row = Record<string, any>;
export async function runBlockerFlow(input: {
page: Page; api: RunnerApi; fixtures: LiveFixtureValues; execution: MatrixExecution;
nonce: string; workspacePath: string; deadlineAt: number;
observe(issue: any, runs: any[], checks: ReturnType<typeof gradeBlocker>): void;
capture(id: string, label: string, file: string): Promise<void>;
evidence(name: string, value: unknown): Promise<void>;
}) {
const { page, api, fixtures, execution } = input;
const scenario = blockerScenario(execution.task.id, input.nonce);
const checkpoints: BlockerCheckpoint[] = [];
let issue: Row | undefined;
let runs: Row[] = [];
let managerId = "";
let checks: ReturnType<typeof gradeBlocker> = [];
const company = `/api/companies/${fixtures.company.id}`;
const hashes: Record<string, string> = {};
const grade = (requireFinal: boolean) => gradeBlocker({ caseId: scenario.id, assigneeId: fixtures.agent.id,
managerId, marker: scenario.marker, checkpoints, requireFinal });
async function state(phase: BlockerCheckpoint["phase"]): Promise<BlockerCheckpoint> {
issue = await api.get<Row>(`/api/issues/${issue!.id}`);
const listed = await api.get<Row[]>(`${company}/heartbeat-runs?limit=100`);
runs = await Promise.all(listed.map(r => api.get<Row>(`/api/heartbeat-runs/${r.id}`)));
const [issues, agents, interactions, comments, activity, approvals] = await Promise.all([
api.get<Row[]>(`${company}/issues`), api.get<Row[]>(`${company}/agents`),
api.get<Row[]>(`/api/issues/${issue.id}/interactions`), api.get<Row[]>(`/api/issues/${issue.id}/comments?order=asc`),
api.get<Row[]>(`/api/issues/${issue.id}/activity`), api.get<Row[]>(`${company}/approvals`),
]);
input.observe(issue, runs, checks);
return { phase, issue, issues, runs, agents, interactions, comments, activity, approvals };
}
async function settle(phase: BlockerCheckpoint["phase"]) {
let lastKey = "";
return pollUntil({ label: `blocker ${phase}`, deadlineAt: input.deadlineAt, intervalMs: 1000,
load: () => state(phase), accept: s => {
const idle = s.runs.length > 0 && s.runs.every(r => ["succeeded", "failed", "timed_out", "cancelled"].includes(r.status)) &&
!s.issue.scheduledRetry && !s.issue.activeRecoveryAction;
const waitingRuns = checkpoints.find(c => c.phase === "waiting")?.runs ?? [];
const ready = idle && (phase === "waiting" || s.runs.some(r => !waitingRuns.some(old => old.id === r.id)));
const key = ready ? JSON.stringify([s.issue.status, s.runs.map(r => [r.id, r.status]), s.interactions.map(i => [i.id, i.status])]) : "";
const stable = !!key && key === lastKey;
lastKey = key;
return stable;
}, reject: s => s.runs.length > 4 ? "bounded run count exceeded" :
s.runs.some(r => ["failed", "timed_out", "cancelled"].includes(r.status)) ? "provider run failed; see retained run records" : undefined });
}
async function open() {
await page.goto(`/${fixtures.company.issuePrefix}/issues/${issue!.identifier ?? issue!.id}`, { waitUntil: "domcontentloaded" });
await expect(page.getByRole("heading", { name: String(issue!.title), exact: true })).toBeVisible();
await expect(page.getByTestId("issue-chat-skeleton")).toHaveCount(0);
}
function assertChecks() {
const failed = checks.filter(c => !c.passed);
input.observe(issue, runs, checks);
if (failed.length) throw new Error(`Blocker outcome checks failed: ${failed.map(c => c.id).join(", ")}`);
}
try {
for (const file of ["skills/paperclip/SKILL.md", "skills/paperclip/references/api-reference.md", "skills/paperclip-create-agent/SKILL.md",
...["blocker-cases.ts", "blocker-flow.ts", "blocker-input.ts", "blocker-fixtures.ts", "blocker-scoring.ts"].map(f => `tests/runner-e2e/${f}`)]) {
hashes[file] = createHash("sha256").update(await readFile(path.resolve(import.meta.dirname, "../..", file))).digest("hex");
}
const setup = await setupBlockerFixtures(input);
managerId = setup.manager.id;
const skill = (await api.get<Row[]>(`${company}/skills`)).find(s => s.key === "paperclipai/paperclip/paperclip");
if (!skill) throw new Error("Missing assigned operational skill");
const servedHashes: Record<string, string> = {};
for (const file of ["SKILL.md", "references/api-reference.md"]) {
const served = await api.get<{ content: string }>(`${company}/skills/${skill.id}/files?path=${encodeURIComponent(file)}`);
servedHashes[`skills/paperclip/${file}`] = createHash("sha256").update(served.content).digest("hex");
}
await input.evidence("blocker-skill-source.json", { skillId: skill.id, servedHashes,
agentSkills: await api.get(`/api/agents/${fixtures.agent.id}/skills`) });
for (const [file, hash] of Object.entries(servedHashes)) {
if (hash !== hashes[file]) throw new Error(`Bundled operational skill differs from evaluated source: ${file}`);
}
await api.patch("/api/instance/settings/experimental", { enableClassicTaskInterface: false });
await createTaskThroughUi({ page, issuePrefix: fixtures.company.issuePrefix!, agentName: fixtures.agent.name,
title: execution.task.buildTitle(input.nonce), prompt: scenario.prompt, workMode: "standard" });
issue = await pollUntil({ label: "browser-created blocker task", deadlineAt: input.deadlineAt,
load: async () => (await api.get<Row[]>(`${company}/issues`)).find(i => i.title === execution.task.buildTitle(input.nonce)), accept: Boolean });
if (!issue) throw new Error("No browser-created task");
checkpoints.push(await settle("waiting"));
await open();
await input.capture("blocker-waiting", "Saved blocker decision", "decision-pending.png");
checks = grade(false);
assertChecks();
// Reload proves the waiting interaction is durable before a real UI answer.
await page.reload();
const question = pendingBlockerInput(checkpoints[0])!;
await answerBlockerThroughUi(page, question, scenario.answer);
await pollUntil({ label: "saved browser decision", deadlineAt: input.deadlineAt,
load: () => api.get<Row[]>(`/api/issues/${issue!.id}/interactions`),
accept: rows => rows.some(i => i.id === question.id && i.status === (question.kind === "ask_user_questions" ? "answered" : "rejected")) });
checkpoints.push(await settle("final"));
checks = grade(true);
await open();
assertChecks();
const reply = checkpoints.at(-1)!.comments.find(c => c.authorAgentId === fixtures.agent.id && String(c.body).includes(scenario.marker));
expect(reply, "the worker's acknowledgement must be persisted").toBeTruthy();
// A later worker follow-up may refer to its earlier acknowledgement.
// Durable authorship/completion are graded above; capture the latest reply.
const bubble = page.getByTestId("task-chat-agent-bubble").filter({ hasText: scenario.marker }).last();
await expect(bubble).toBeVisible();
await bubble.scrollIntoViewIfNeeded();
await input.capture("final-state", "Original worker completed after human answer", "final-state.png");
assertChecks();
return { issue: issue!, runs, checks };
} catch (error) {
checks.push({ id: "workflow-completed", passed: false, detail: error instanceof Error ? error.message : String(error) });
if (issue) input.observe(issue, runs, checks);
throw error;
} finally {
let lastObservation: unknown;
if (issue) lastObservation = await state("final").catch(error => ({ evidenceError: String(error) }));
await input.evidence("blocker-runs.json", await Promise.all(runs.map(run =>
collectChatRunEvidence(api, run as Parameters<typeof collectChatRunEvidence>[1])
.catch(error => ({ runId: run.id, evidenceError: String(error) })))));
await input.evidence("api-state.json", { capturePhase: "blocker-final", issue, runs, checks, lastObservation });
await input.evidence("blocker-guidance.json", { schema: BLOCKER_GRADER_VERSION, graderVersion: BLOCKER_GRADER_VERSION,
inputUx: gradeBlockerInputUx(checkpoints.find(c => c.phase === "waiting")), caseId: scenario.id,
prompt: scenario.prompt, answer: scenario.answer, hashes, managerId, assigneeId: fixtures.agent.id, checks, checkpoints, lastObservation });
}
}