Files
PaperClipAI/tests/runner-e2e/plan-task-flow.ts
T
DottaandPaperclip 16b7db35ff Shorten planning skills and measure task decomposition (#15296)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Planning guidance helps agents choose owners and dependencies.
> - The runtime skill favors few tasks, but the catalog skill requires a
child-task breakdown.
> - Both add repeated process instructions that can distract from the
requested outcome.
> - This change keeps the ownership and dependency rules and removes the
required matrix and repeated checklist.
> - A bounded Product E2E comparison measures saved outcomes and task
handoffs before qualification.

## Linked Issues or Issue Description

Refs #11057. Related measurement work: #15218.

**What existing behavior does this improve?**
Planning and delegation through the runtime plan-to-tasks and bundled
task-planning skills.

**Current behavior**
The two skills contain about 1,900 words and conflicting guidance on
whether plans require child tasks.

**Proposed behavior**
Keep cohesive work with one owner. Split only for a real owner, parallel
output, dependency, independent review, or follow-up lifecycle. Preserve
existing authorization and planning mechanics.

## What Changed

- Shorten both skills to about 400 words combined. Preserve their keys
and installed-version behavior.
- Remove the duplicate operational-skill pointer and regenerate affected
source metadata.
- Add twelve explicit Product E2E cells: four scenarios with current,
short and disabled planning skills.
- Use the current task composer and actual create-response ID; calibrate
public skill APIs and browser creation without providers.
- Eliminate an observed collision in chat-test company prefixes with a
per-suite sequence.
- Grade saved documents, exact author/run attribution, child count,
prerequisite execution order, review boundaries and completion handoffs.
- Retain current skill bytes and report source, selections, run
accounting and failures.

## Verification

- `pnpm test:e2e:runner:typecheck`: pass.
- `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks
pass.
- `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve
local Codex cells.
- Archived current skills match master
`72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly.
- Corrected fixture: three real public-API/database calibrations pass
with zero provider runs; all 35 evaluator checks and Product E2E
typecheck pass.
- Setup campaign
[37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253)
was canceled after source review found unsupported bundled edits and
automatic core reinstallation. Its paid-cell step was skipped: zero
provider runs, no behavioral grade.
- The next setup
[37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799)
failed before task creation on the old title-field selector: zero actual
runs, original FAIL retained, cleanup passed. A real browser/API
calibration of the new helper passes with paused non-provider agents and
zero runs.
- Full local typecheck/build pass. Full local tests retain one unchanged
five-minute Git streaming timeout (also fails isolated), 9,591 passes
and 5,796 skips. CI's chat failure was a proven random fixture-prefix
collision; five affected cases pass after the test-only repair.
- Paid behavior comparison and new-head CI/review remain pending. This
PR remains a draft.

## Risks

- The shorter text may change delegation decisions. Live outcomes are
not yet qualified.
- The initial comparison uses one profile and one attempt per cell. It
cannot establish cross-model reliability or cost trends.
- Disabled means unassigned company-owned copies; the company library
remains discoverable. This does not qualify global removal, automatic
accepted-plan wiring changes, or installed-copy migration.
- Skill availability does not prove a model read or cognitively used it.
- No provider/tool protocol, permission, timeout or runtime lifecycle
behavior changes in production.

## Model Used

OpenAI Codex (GPT-6), with repository inspection, code editing and tool
use. The exact backend model ID and context-window size are not exposed
in this session. The declared eval model is native Codex `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:36:01 -05:00

140 lines
9.8 KiB
TypeScript

import { expect, type Page } from "@playwright/test";
import { mkdir } from "node:fs/promises";
import path from "node:path";
import { pollUntil, type RunnerApi } from "./api.js";
import { collectChatRunEvidence } from "./chat-flow.js";
import { FixtureRegistry } from "./fixture-registry.js";
import type { LiveFixtureValues } from "./live-fixtures.js";
import type { MatrixExecution } from "./types.js";
import { createPlanTaskThroughUi } from "./plan-task-ui.js";
import { PLAN_BASE_SHA, PLAN_BUDGET_CENTS, PLAN_MAX_RUNS, parsePlanCase, planDefinitionDigest, planScenario } from "./plan-task-cases.js";
import { gradePlanTask, type PlanCheck, type PlanDocument, type PlanObservation, type PlanRow } from "./plan-task-scoring.js";
import { planExposure, planInvocationPrompt, planInvocations, type PlanInvocation } from "./plan-task-exposure.js";
import { preparePlanSkills, selectPlanSkills, verifyPlanSelection, verifyPlanSkills } from "./plan-task-skills.js";
type Input = {
page: Page; api: RunnerApi; fixtures: LiveFixtureValues; execution: MatrixExecution;
nonce: string; workspacePath: string; deadlineAt: number;
observe(issue: PlanRow, runs: PlanRow[], checks: PlanCheck[]): void;
capture(id: string, label: string, file: string): Promise<void>;
evidence(name: string, value: unknown): Promise<void>;
};
export async function runPlanTaskFlow(input: Input) {
const { api, page, fixtures: f, execution } = input;
const { variant, caseId } = parsePlanCase(execution.task.id);
const scenario = planScenario(caseId, input.nonce);
const company = `/api/companies/${f.company.id}`;
const source: PlanRow = { origin: api.baseURL, variant, base: PLAN_BASE_SHA, definitionDigest: planDefinitionDigest, skills: [], agents: [] };
let parent: PlanRow | undefined, alexId = "", rileyId = "", failure: string | undefined;
let observed: PlanObservation = { issues: [], runs: [], documents: [], comments: [], activity: [], interactions: [], wakes: [] };
let checks: PlanCheck[] = [];
let invocations: PlanInvocation[] = [];
let runEvidence: PlanRow[] = [];
const exposure = () => planExposure({ runs: observed.runs, evidence: runEvidence, leadId: f.agent.id, parentId: parent?.id ?? "", invocations });
async function captureRuns() {
runEvidence = await Promise.all(observed.runs.map(run => collectChatRunEvidence(api, run as Parameters<typeof collectChatRunEvidence>[1])
.catch(error => ({ runId: run.id, evidenceError: String(error) }))));
}
const grade = () => gradePlanTask({ caseId, marker: scenario.marker, parentId: parent?.id ?? "", leadId: f.agent.id,
alexId, rileyId, observation: observed, maxRuns: PLAN_MAX_RUNS, origin: api.baseURL });
async function observe() {
const [issues, list] = await Promise.all([api.get<PlanRow[]>(`${company}/issues`), api.get<PlanRow[]>(`${company}/heartbeat-runs?limit=100`)]);
if (list.length >= 100) throw new Error("Company run list is truncated; cannot qualify accounting");
const runs = await Promise.all(list.map(run => api.get<PlanRow>(`/api/heartbeat-runs/${run.id}`)));
const details = await Promise.all(issues.map(async issue => {
const base = `/api/issues/${issue.id}`;
const [full, docs, comments, activity, interactions, wakes] = await Promise.all([
api.get<PlanRow>(base), api.get<PlanRow[]>(`${base}/documents`), api.get<PlanRow[]>(`${base}/comments?order=asc`),
api.get<PlanRow[]>(`${base}/activity`), api.get<PlanRow[]>(`${base}/interactions`), api.get<PlanRow>(`${base}/diagnostics/wakes`),
]);
const documents = await Promise.all(docs.map(async doc => ({ ...doc, issueId: issue.id,
revisions: await api.get<PlanRow[]>(`${base}/documents/${encodeURIComponent(doc.key)}/revisions`) } as PlanDocument)));
return { full, documents, comments, activity, interactions, wakes: { ...wakes, issueId: issue.id } };
}));
observed = { issues: details.map(d => d.full), runs, documents: details.flatMap(d => d.documents),
comments: details.flatMap(d => d.comments), activity: details.flatMap(d => d.activity),
interactions: details.flatMap(d => d.interactions), wakes: details.map(d => d.wakes) };
if (parent) { parent = observed.issues.find(i => i.id === parent!.id) ?? parent; input.observe(parent, runs, checks); }
return observed;
}
try {
// Public fixture APIs and normal skill versioning only. No provider begins
// until the served bytes and all three agents' selections are verified.
await api.patch(`${company}/budgets`, { budgetMonthlyCents: PLAN_BUDGET_CENTS });
source.skills = await preparePlanSkills(api, f.company.id, variant);
const registry = new FixtureRegistry();
for (const [id, name, role, capabilities] of [
["alex", "Alex Metrics", "engineer", "Owns arithmetic verification, operational metrics, and signed order summaries."],
["riley", "Riley Copy", "qa", "Owns publication copy, release verification and independent arithmetic audits."],
] as const) {
registry.register<PlanRow>({ id, async setup() {
const workspacePath = path.join(input.workspacePath, id);
await mkdir(workspacePath, { recursive: true });
const payload = execution.profile.buildAgent({ environmentId: f.environment.id, environmentFixtureId: "local",
workspacePath, secretRefs: f.secretRefs, executionId: input.nonce });
return api.post(`${company}/agents`, { ...payload, name, role, title: name, capabilities, reportsTo: f.agent.id,
instructionsBundle: { entryFile: "AGENTS.md", files: { "AGENTS.md": `You are ${name}. ${capabilities} Deliver accurate work that meets your assignment.` } } });
} });
}
const team = await registry.setupAll();
alexId = (team.values.get("alex") as PlanRow).id;
rileyId = (team.values.get("riley") as PlanRow).id;
for (const id of [f.agent.id, alexId, rileyId]) {
await api.patch(`/api/agents/${id}/permissions`, { canCreateAgents: false, canAssignTasks: id === f.agent.id });
await api.patch(`/api/agents/${id}/budgets`, { budgetMonthlyCents: PLAN_BUDGET_CENTS });
const skills = await selectPlanSkills(api, f.company.id, id, source.skills);
source.agents.push({ id, skills });
if (id === f.agent.id) invocations = planInvocations(source.skills, skills);
}
source.invocations = invocations;
source.requestPrompt = planInvocationPrompt(scenario.prompt, invocations);
await input.evidence("plan-task-source.json", source);
await api.patch("/api/instance/settings/experimental", { enableClassicTaskInterface: false });
const created = await createPlanTaskThroughUi({ page, companyId: f.company.id, issuePrefix: f.company.issuePrefix!,
agentName: f.agent.name, prompt: source.requestPrompt });
parent = await pollUntil({ label: "browser-created planning-guidance task", deadlineAt: input.deadlineAt,
load: () => api.get<PlanRow>(`/api/issues/${created.id}`), accept: i => i.id === created.id });
if (!parent) throw new Error("Browser task not found");
let previous = "", same = 0;
await pollUntil({ label: "planning work and handoffs settle", deadlineAt: input.deadlineAt, intervalMs: 1_500, load: observe,
accept: o => {
const idle = o.runs.length > 0 && o.runs.every(r => ["succeeded", "failed", "timed_out", "cancelled"].includes(r.status));
const done = o.issues.length > 0 && o.issues.every(i => i.status === "done" && !i.scheduledRetry && !i.activeRecoveryAction);
const signature = idle && done ? JSON.stringify(o) : "";
same = signature && signature === previous ? same + 1 : 0; previous = signature;
return same >= 2;
}, reject: o => o.runs.length > PLAN_MAX_RUNS ? "bounded run count exceeded" :
o.runs.some(r => ["failed", "timed_out", "cancelled"].includes(r.status)) ? "actual run failed; retain original failure" : undefined });
await verifyPlanSkills(api, f.company.id, source.skills);
for (const id of [f.agent.id, alexId, rileyId]) await verifyPlanSelection(api, f.company.id, id, source.skills);
await captureRuns();
checks = [...grade().checks, exposure()];
input.observe(parent, observed.runs, checks);
await page.goto(`/${f.company.issuePrefix}/issues/${parent.identifier ?? parent.id}`, { waitUntil: "domcontentloaded" });
await expect(page.getByRole("heading", { name: parent.title, exact: true })).toBeVisible();
await expect(page.getByTestId("issue-chat-skeleton")).toHaveCount(0);
await input.capture("final-state", "Saved planning outcome and task graph", "final-state.png");
if (checks.some(c => !c.passed)) throw new Error(`Planning outcome failed: ${checks.filter(c => !c.passed).map(c => c.id).join(", ")}`);
return { issue: parent, runs: observed.runs, checks };
} catch (error) {
failure = error instanceof Error ? error.message : String(error);
checks = [...grade().checks, exposure(), { id: "workflow-completed", passed: false, detail: failure }];
if (parent) input.observe(parent, observed.runs, checks);
throw error;
} finally {
let captureError: string | undefined;
try { await observe(); } catch (error) { captureError = String(error); }
const finalSkills = await Promise.all([f.agent.id, alexId, rileyId].filter(Boolean).map(async id => ({ id,
skills: await api.get(`/api/agents/${id}/skills?companyId=${f.company.id}`).catch(error => ({ error: String(error) })) })));
await captureRuns();
await input.evidence("plan-task-runs.json", runEvidence);
await input.evidence("plan-task-guidance.json", { variant, caseId, scenario, source, finalSkills, parentId: parent?.id,
leadId: f.agent.id, alexId, rileyId, budgetCents: PLAN_BUDGET_CENTS, maxRuns: PLAN_MAX_RUNS,
observation: observed, result: grade(), exposure: exposure(), failure, captureError });
await input.evidence("api-state.json", { capturePhase: "plan-task-final", issue: parent, runs: observed.runs, checks, failure, captureError });
}
}