mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip manages work for AI agents. > - Agents use the coordination skill when work needs human authority or a scope decision. > - PR #14188 replaced automatic manager escalation with direct blocker handling. > - This behavior needs real browser, server, database, and provider tests. > - The test must verify saved human input, task ownership, and resumed work. > - This pull request adds six reusable Product E2E cases and improves the skill examples that they exercise. ## Linked Issues or Issue Description Refs #14188. The merged change needs repeatable behavior coverage. The new suite tests missing administrator access, missing hiring permission, and requester scope questions. Searches found no duplicate blocker-guidance suite. This extends the existing eval system described in ROADMAP.md. ## What Changed - Add the explicit-only `blocker-guidance` Product E2E suite. It has three local scenarios on legacy Codex and legacy Claude. - Use the production UI and public APIs to create work, save a human-only question or confirmation, answer it after reload, and resume the same task. - Check requester identity, ownership history, manager activity, hiring, saved answers, and completion. Keep direct text input as a separate UX result. - Save pending and final screenshots, API checkpoints, skill hashes, provider evidence, and billing data through the existing report pipeline. - Isolate the Claude fixture home. Verify the served skill bytes before dispatch so an old installed skill cannot silently replace the evaluated skill. - Improve the coordination and hiring skill examples. Include the human-only policy, requester address, wake behavior, and handling of authorized scope changes. - Grader v5 requires the exact approved public welcome note as a new worker comment. Browser input checks reject unwritable scope cards before clicking, and confirmation direction must be saved in the resolution before the worker wakes. - Add grader calibration and browser-input tests. Update the fixture guide and generated capability inventories. ## Verification - `pnpm build`: passed after rebasing onto current master. - `pnpm -r typecheck`: passed. - `pnpm test:e2e:runner:typecheck`: passed. - `pnpm test:e2e:runner:unit`: 742 tests passed. - `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests passed. - `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells found. - Capability inventory and generated-contract checks: passed. - `pnpm exec vitest run server/src/__tests__/hiring-operational-examples.test.ts`: four tests passed after synchronizing the generated API reference and section anchor. - Full general and serialized test suites: passed in CI on `6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm test:run` was interrupted after complete CI coverage passed; it is not claimed as a completed local full-suite run. - Final GitHub checks: 54 passed, two optional Storybook checks skipped. The runtime-exposure startup test hit a 10-second readiness timeout once, passed a targeted local reproduction, and its CI shard passed the single retry without code changes. - Current-head Greptile: 5/5, clean check, zero unresolved threads. - Historical live measurement on September 29 at `4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex `gpt-5.6-sol` passed 8/9. These runs predate this rebase. - Version 5 changes the scope answer to an exact approved publication draft. The historical runs do not qualify that new requirement; the two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec` passed Codex and failed Claude. Claude posted the correct salary-free sentence but omitted its required reference line from that comment, placing the reference in a separate completion message. The `public-welcome-note` check correctly failed. An earlier Claude database-startup failure was retained separately; its fresh-instance retry reached the model. This pilot is not a six-cell qualification. - The failed Codex scope case omitted `addresseeUserId`. The strict routing check remains. All 18 attempts had clean evidence manifests and passed cleanup. - To repeat with provider credentials: `pnpm test:e2e:runner -- --suite blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is excluded from `--all`. ## Risks - The live suite measures variable model behavior. The retained 17/18 historical result and the current 1/2 scope pilot are not all-pass qualifications. These paid cases are opt-in; their observed model failures remain visible independently of deterministic CI checks. - A separate generic task-replacement diagnostic still exposed a Claude refusal. The ordinary cases use specific business decisions. The diagnostic is not a standalone catalog case in this change. - Earlier measurements included an old installed Claude skill and test defects. Their grades remain retained and are not combined with the three final repetitions. - Skill examples can affect when agents ask for human input. Downstream permission checks still apply. - Native runners, Daytona, agent-requester routing, and real external connection authorization are outside this suite. - Raw provider traces and credentials remain private. No screenshots, raw reports, secrets, workflow changes, or lockfile changes are committed. ## Model Used OpenAI GPT-6 through Codex assisted with this change. The exact deployed variant and context window size are not exposed in this session. The assistant used reasoning, repository edits, tool use, and shell execution. The evaluated models were `gpt-5.6-sol` and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
138 lines
8.9 KiB
TypeScript
138 lines
8.9 KiB
TypeScript
import { expect, type Page } from "@playwright/test";
|
|
import { createHash } from "node:crypto";
|
|
import { readFile } from "node:fs/promises";
|
|
import path from "node:path";
|
|
import { pollUntil, type RunnerApi } from "./api.js";
|
|
import { blockerScenario } from "./blocker-cases.js";
|
|
import { setupBlockerFixtures } from "./blocker-fixtures.js";
|
|
import { BLOCKER_GRADER_VERSION, gradeBlocker, gradeBlockerInputUx, pendingBlockerInput, type BlockerCheckpoint } from "./blocker-scoring.js";
|
|
import { answerBlockerThroughUi } from "./blocker-input.js";
|
|
import { collectChatRunEvidence } from "./chat-flow.js";
|
|
import type { LiveFixtureValues } from "./live-fixtures.js";
|
|
import type { MatrixExecution } from "./types.js";
|
|
import { createTaskThroughUi } from "./user-actions.js";
|
|
type Row = Record<string, any>;
|
|
|
|
export async function runBlockerFlow(input: {
|
|
page: Page; api: RunnerApi; fixtures: LiveFixtureValues; execution: MatrixExecution;
|
|
nonce: string; workspacePath: string; deadlineAt: number;
|
|
observe(issue: any, runs: any[], checks: ReturnType<typeof gradeBlocker>): void;
|
|
capture(id: string, label: string, file: string): Promise<void>;
|
|
evidence(name: string, value: unknown): Promise<void>;
|
|
}) {
|
|
const { page, api, fixtures, execution } = input;
|
|
const scenario = blockerScenario(execution.task.id, input.nonce);
|
|
const checkpoints: BlockerCheckpoint[] = [];
|
|
let issue: Row | undefined;
|
|
let runs: Row[] = [];
|
|
let managerId = "";
|
|
let checks: ReturnType<typeof gradeBlocker> = [];
|
|
const company = `/api/companies/${fixtures.company.id}`;
|
|
const hashes: Record<string, string> = {};
|
|
const grade = (requireFinal: boolean) => gradeBlocker({ caseId: scenario.id, assigneeId: fixtures.agent.id,
|
|
managerId, marker: scenario.marker, checkpoints, requireFinal });
|
|
async function state(phase: BlockerCheckpoint["phase"]): Promise<BlockerCheckpoint> {
|
|
issue = await api.get<Row>(`/api/issues/${issue!.id}`);
|
|
const listed = await api.get<Row[]>(`${company}/heartbeat-runs?limit=100`);
|
|
runs = await Promise.all(listed.map(r => api.get<Row>(`/api/heartbeat-runs/${r.id}`)));
|
|
const [issues, agents, interactions, comments, activity, approvals] = await Promise.all([
|
|
api.get<Row[]>(`${company}/issues`), api.get<Row[]>(`${company}/agents`),
|
|
api.get<Row[]>(`/api/issues/${issue.id}/interactions`), api.get<Row[]>(`/api/issues/${issue.id}/comments?order=asc`),
|
|
api.get<Row[]>(`/api/issues/${issue.id}/activity`), api.get<Row[]>(`${company}/approvals`),
|
|
]);
|
|
input.observe(issue, runs, checks);
|
|
return { phase, issue, issues, runs, agents, interactions, comments, activity, approvals };
|
|
}
|
|
async function settle(phase: BlockerCheckpoint["phase"]) {
|
|
let lastKey = "";
|
|
return pollUntil({ label: `blocker ${phase}`, deadlineAt: input.deadlineAt, intervalMs: 1000,
|
|
load: () => state(phase), accept: s => {
|
|
const idle = s.runs.length > 0 && s.runs.every(r => ["succeeded", "failed", "timed_out", "cancelled"].includes(r.status)) &&
|
|
!s.issue.scheduledRetry && !s.issue.activeRecoveryAction;
|
|
const waitingRuns = checkpoints.find(c => c.phase === "waiting")?.runs ?? [];
|
|
const ready = idle && (phase === "waiting" || s.runs.some(r => !waitingRuns.some(old => old.id === r.id)));
|
|
const key = ready ? JSON.stringify([s.issue.status, s.runs.map(r => [r.id, r.status]), s.interactions.map(i => [i.id, i.status])]) : "";
|
|
const stable = !!key && key === lastKey;
|
|
lastKey = key;
|
|
return stable;
|
|
}, reject: s => s.runs.length > 4 ? "bounded run count exceeded" :
|
|
s.runs.some(r => ["failed", "timed_out", "cancelled"].includes(r.status)) ? "provider run failed; see retained run records" : undefined });
|
|
}
|
|
async function open() {
|
|
await page.goto(`/${fixtures.company.issuePrefix}/issues/${issue!.identifier ?? issue!.id}`, { waitUntil: "domcontentloaded" });
|
|
await expect(page.getByRole("heading", { name: String(issue!.title), exact: true })).toBeVisible();
|
|
await expect(page.getByTestId("issue-chat-skeleton")).toHaveCount(0);
|
|
}
|
|
function assertChecks() {
|
|
const failed = checks.filter(c => !c.passed);
|
|
input.observe(issue, runs, checks);
|
|
if (failed.length) throw new Error(`Blocker outcome checks failed: ${failed.map(c => c.id).join(", ")}`);
|
|
}
|
|
try {
|
|
for (const file of ["skills/paperclip/SKILL.md", "skills/paperclip/references/api-reference.md", "skills/paperclip-create-agent/SKILL.md",
|
|
...["blocker-cases.ts", "blocker-flow.ts", "blocker-input.ts", "blocker-fixtures.ts", "blocker-scoring.ts"].map(f => `tests/runner-e2e/${f}`)]) {
|
|
hashes[file] = createHash("sha256").update(await readFile(path.resolve(import.meta.dirname, "../..", file))).digest("hex");
|
|
}
|
|
const setup = await setupBlockerFixtures(input);
|
|
managerId = setup.manager.id;
|
|
const skill = (await api.get<Row[]>(`${company}/skills`)).find(s => s.key === "paperclipai/paperclip/paperclip");
|
|
if (!skill) throw new Error("Missing assigned operational skill");
|
|
const servedHashes: Record<string, string> = {};
|
|
for (const file of ["SKILL.md", "references/api-reference.md"]) {
|
|
const served = await api.get<{ content: string }>(`${company}/skills/${skill.id}/files?path=${encodeURIComponent(file)}`);
|
|
servedHashes[`skills/paperclip/${file}`] = createHash("sha256").update(served.content).digest("hex");
|
|
}
|
|
await input.evidence("blocker-skill-source.json", { skillId: skill.id, servedHashes,
|
|
agentSkills: await api.get(`/api/agents/${fixtures.agent.id}/skills`) });
|
|
for (const [file, hash] of Object.entries(servedHashes)) {
|
|
if (hash !== hashes[file]) throw new Error(`Bundled operational skill differs from evaluated source: ${file}`);
|
|
}
|
|
await api.patch("/api/instance/settings/experimental", { enableClassicTaskInterface: false });
|
|
await createTaskThroughUi({ page, issuePrefix: fixtures.company.issuePrefix!, agentName: fixtures.agent.name,
|
|
title: execution.task.buildTitle(input.nonce), prompt: scenario.prompt, workMode: "standard" });
|
|
issue = await pollUntil({ label: "browser-created blocker task", deadlineAt: input.deadlineAt,
|
|
load: async () => (await api.get<Row[]>(`${company}/issues`)).find(i => i.title === execution.task.buildTitle(input.nonce)), accept: Boolean });
|
|
if (!issue) throw new Error("No browser-created task");
|
|
checkpoints.push(await settle("waiting"));
|
|
await open();
|
|
await input.capture("blocker-waiting", "Saved blocker decision", "decision-pending.png");
|
|
checks = grade(false);
|
|
assertChecks();
|
|
// Reload proves the waiting interaction is durable before a real UI answer.
|
|
await page.reload();
|
|
const question = pendingBlockerInput(checkpoints[0])!;
|
|
await answerBlockerThroughUi(page, question, scenario.answer);
|
|
await pollUntil({ label: "saved browser decision", deadlineAt: input.deadlineAt,
|
|
load: () => api.get<Row[]>(`/api/issues/${issue!.id}/interactions`),
|
|
accept: rows => rows.some(i => i.id === question.id && i.status === (question.kind === "ask_user_questions" ? "answered" : "rejected")) });
|
|
checkpoints.push(await settle("final"));
|
|
checks = grade(true);
|
|
await open();
|
|
assertChecks();
|
|
const reply = checkpoints.at(-1)!.comments.find(c => c.authorAgentId === fixtures.agent.id && String(c.body).includes(scenario.marker));
|
|
expect(reply, "the worker's acknowledgement must be persisted").toBeTruthy();
|
|
// A later worker follow-up may refer to its earlier acknowledgement.
|
|
// Durable authorship/completion are graded above; capture the latest reply.
|
|
const bubble = page.getByTestId("task-chat-agent-bubble").filter({ hasText: scenario.marker }).last();
|
|
await expect(bubble).toBeVisible();
|
|
await bubble.scrollIntoViewIfNeeded();
|
|
await input.capture("final-state", "Original worker completed after human answer", "final-state.png");
|
|
assertChecks();
|
|
return { issue: issue!, runs, checks };
|
|
} catch (error) {
|
|
checks.push({ id: "workflow-completed", passed: false, detail: error instanceof Error ? error.message : String(error) });
|
|
if (issue) input.observe(issue, runs, checks);
|
|
throw error;
|
|
} finally {
|
|
let lastObservation: unknown;
|
|
if (issue) lastObservation = await state("final").catch(error => ({ evidenceError: String(error) }));
|
|
await input.evidence("blocker-runs.json", await Promise.all(runs.map(run =>
|
|
collectChatRunEvidence(api, run as Parameters<typeof collectChatRunEvidence>[1])
|
|
.catch(error => ({ runId: run.id, evidenceError: String(error) })))));
|
|
await input.evidence("api-state.json", { capturePhase: "blocker-final", issue, runs, checks, lastObservation });
|
|
await input.evidence("blocker-guidance.json", { schema: BLOCKER_GRADER_VERSION, graderVersion: BLOCKER_GRADER_VERSION,
|
|
inputUx: gradeBlockerInputUx(checkpoints.find(c => c.phase === "waiting")), caseId: scenario.id,
|
|
prompt: scenario.prompt, answer: scenario.answer, hashes, managerId, assigneeId: fixtures.agent.id, checks, checkpoints, lastObservation });
|
|
}
|
|
}
|