Files
PaperClipAI/tests/runner-e2e/plan-task-scoring.ts
T
DottaandPaperclip 16b7db35ff Shorten planning skills and measure task decomposition (#15296)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Planning guidance helps agents choose owners and dependencies.
> - The runtime skill favors few tasks, but the catalog skill requires a
child-task breakdown.
> - Both add repeated process instructions that can distract from the
requested outcome.
> - This change keeps the ownership and dependency rules and removes the
required matrix and repeated checklist.
> - A bounded Product E2E comparison measures saved outcomes and task
handoffs before qualification.

## Linked Issues or Issue Description

Refs #11057. Related measurement work: #15218.

**What existing behavior does this improve?**
Planning and delegation through the runtime plan-to-tasks and bundled
task-planning skills.

**Current behavior**
The two skills contain about 1,900 words and conflicting guidance on
whether plans require child tasks.

**Proposed behavior**
Keep cohesive work with one owner. Split only for a real owner, parallel
output, dependency, independent review, or follow-up lifecycle. Preserve
existing authorization and planning mechanics.

## What Changed

- Shorten both skills to about 400 words combined. Preserve their keys
and installed-version behavior.
- Remove the duplicate operational-skill pointer and regenerate affected
source metadata.
- Add twelve explicit Product E2E cells: four scenarios with current,
short and disabled planning skills.
- Use the current task composer and actual create-response ID; calibrate
public skill APIs and browser creation without providers.
- Eliminate an observed collision in chat-test company prefixes with a
per-suite sequence.
- Grade saved documents, exact author/run attribution, child count,
prerequisite execution order, review boundaries and completion handoffs.
- Retain current skill bytes and report source, selections, run
accounting and failures.

## Verification

- `pnpm test:e2e:runner:typecheck`: pass.
- `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks
pass.
- `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve
local Codex cells.
- Archived current skills match master
`72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly.
- Corrected fixture: three real public-API/database calibrations pass
with zero provider runs; all 35 evaluator checks and Product E2E
typecheck pass.
- Setup campaign
[37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253)
was canceled after source review found unsupported bundled edits and
automatic core reinstallation. Its paid-cell step was skipped: zero
provider runs, no behavioral grade.
- The next setup
[37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799)
failed before task creation on the old title-field selector: zero actual
runs, original FAIL retained, cleanup passed. A real browser/API
calibration of the new helper passes with paused non-provider agents and
zero runs.
- Full local typecheck/build pass. Full local tests retain one unchanged
five-minute Git streaming timeout (also fails isolated), 9,591 passes
and 5,796 skips. CI's chat failure was a proven random fixture-prefix
collision; five affected cases pass after the test-only repair.
- Paid behavior comparison and new-head CI/review remain pending. This
PR remains a draft.

## Risks

- The shorter text may change delegation decisions. Live outcomes are
not yet qualified.
- The initial comparison uses one profile and one attempt per cell. It
cannot establish cross-model reliability or cost trends.
- Disabled means unassigned company-owned copies; the company library
remains discoverable. This does not qualify global removal, automatic
accepted-plan wiring changes, or installed-copy migration.
- Skill availability does not prove a model read or cognitively used it.
- No provider/tool protocol, permission, timeout or runtime lifecycle
behavior changes in production.

## Model Used

OpenAI Codex (GPT-6), with repository inspection, code editing and tool
use. The exact backend model ID and context-window size are not exposed
in this session. The declared eval model is native Codex `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:36:01 -05:00

107 lines
7.7 KiB
TypeScript

import { isDeepStrictEqual } from "node:util";
import type { PlanCase } from "./plan-task-cases.js";
export type PlanRow = Record<string, any>;
export interface PlanDocument { id: string; latestRevisionNumber: number; issueId: string; key: string; body: string; latestRevisionId: string; revisions: PlanRow[] }
export interface PlanObservation {
issues: PlanRow[]; runs: PlanRow[]; documents: PlanDocument[];
comments: PlanRow[]; activity: PlanRow[]; interactions: PlanRow[]; wakes: PlanRow[];
}
export interface PlanCheck { id: string; passed: boolean; detail: string }
export const PLAN_GRADER = "paperclip.plan-task-guidance.v1";
export function planDocumentJson(body: unknown): unknown {
if (typeof body !== "string") return undefined;
const text = body.trim();
const fenced = text.match(/^```(?:json)?\s*\n([\s\S]*?)\n```$/i);
try { return JSON.parse(fenced?.[1] ?? text); } catch { return undefined; }
}
const time = (value: unknown) => typeof value === "string" ? Date.parse(value) : NaN;
export function gradePlanTask(input: {
caseId: PlanCase; marker: string; parentId: string; leadId: string; alexId: string; rileyId: string;
observation: PlanObservation; maxRuns: number; origin: string;
}) {
const { caseId, marker, parentId, leadId, alexId, rileyId, observation: o } = input;
const checks: PlanCheck[] = [];
const check = (id: string, passed: boolean, detail: string) => checks.push({ id, passed, detail });
const parent = o.issues.find(i => i.id === parentId);
const children = o.issues.filter(i => i.id !== parentId);
const alex = children.filter(i => i.assigneeAgentId === alexId);
const riley = children.filter(i => i.assigneeAgentId === rileyId);
const expectedChildren = caseId === "cohesive" ? 0 : caseId === "review" ? 1 : 2;
const pending = o.interactions.some(i => ["pending", "open"].includes(i.status)) || o.wakes.some(w => w.truncated ||
w.events?.some((event: PlanRow) => event.kind === "wake_request" && !["completed", "skipped", "failed", "cancelled"].includes(event.status) &&
!(event.status === "coalesced" && o.runs.some(r => r.id === event.runId && r.status === "succeeded"))));
check("tasks-completed", Boolean(parent) && o.issues.every(i => i.status === "done" && !i.scheduledRetry && !i.activeRecoveryAction) && !pending &&
o.issues.every(i => o.wakes.filter(w => w.issueId === i.id && Array.isArray(w.events) && !w.truncated).length === 1),
"Every requested task must be done with no waiting decision, retry, recovery or pending wake; an adverse review is a completed deliverable.");
check("bounded-native-runs", o.runs.length > 0 && o.runs.length <= input.maxRuns && o.runs.every(r => r.status === "succeeded" && r.runtimeMode === "native" &&
r.runnerInstanceId && r.nativeSessionId && !r.retryOfRunId && time(r.finishedAt) >= time(r.startedAt)),
`Observed ${o.runs.length} total run records; all must be successful native executions with identity and finite ordered timestamps, without retries.`);
check("minimal-owned-work", children.length === expectedChildren && (caseId === "cohesive" ||
(riley.length === 1 && (caseId === "review" ? alex.length === 0 : alex.length === 1))),
`Expected ${expectedChildren} independent work items for the stated owners; found ${children.length}.`);
function document(issue: PlanRow | undefined, owner: string, expected: unknown, label: string) {
const docs = issue ? o.documents.filter(d => d.issueId === issue.id && d.key === "result") : [];
const doc = docs.length === 1 ? docs[0] : undefined;
const revision = doc?.revisions.find(r => r.id === doc.latestRevisionId);
const writes = o.activity.filter(a => ["issue.document_created", "issue.document_updated"].includes(a.action) &&
a.entityId === issue?.id && a.details?.documentId === doc?.id && a.details?.key === "result" &&
a.details?.revisionNumber === doc?.latestRevisionNumber);
const write = writes.length === 1 ? writes[0] : undefined;
const run = o.runs.find(r => r.id === write?.runId);
const runIssue = run?.nativeIssueId ?? run?.contextSnapshot?.issueId ?? run?.contextSnapshot?.taskId;
check(`${label}-saved-output`, !!doc && isDeepStrictEqual(planDocumentJson(doc.body), expected), "The saved result must match the independently calculated business output.");
check(`${label}-authorship`, !!revision && revision.createdByAgentId === owner && !revision.createdByUserId &&
write?.agentId === owner && run?.agentId === owner && runIssue === issue?.id && revision.body === doc?.body,
"The latest document revision must join its exact author and run on the owned task; assignment or a claimed signature alone is insufficient.");
return doc;
}
const order = { marker, units: 5, total: 29 };
const verdict = { marker, verdict: "reject", correctTotal: 29, difference: 1 };
let expected: unknown = order;
if (caseId === "parallel") {
document(alex[0], alexId, { marker, seats: 24 }, "alex");
document(riley[0], rileyId, { marker, welcome: "Hola, equipo" }, "riley");
expected = { marker, seats: 24, welcome: "Hola, equipo" };
}
if (caseId === "dependency") {
const source = document(alex[0], alexId, order, "alex");
expected = { marker, releasedTotal: 29, sourceRevisionId: source?.latestRevisionId };
document(riley[0], rileyId, expected, "riley");
const downstream = o.runs.filter(r => (r.nativeIssueId ?? r.contextSnapshot?.issueId ?? r.contextSnapshot?.taskId) === riley[0]?.id);
check("prerequisite-before-execution", downstream.length > 0 && downstream.every(r => time(r.startedAt) >= time(alex[0]?.completedAt)),
"No downstream provider execution may precede the prerequisite's persisted completion; late creation is also a valid way to wait.");
}
if (caseId === "review") {
document(riley[0], rileyId, verdict, "independent-review");
expected = verdict;
check("review-write-boundary", !o.comments.some(c => c.issueId === parentId && c.authorAgentId === rileyId) &&
!o.documents.filter(d => d.issueId === parentId).some(d => d.revisions.some(r => r.createdByAgentId === rileyId)),
"The independent reviewer must deliver on its own task without parent writes.");
}
document(parent, leadId, expected, "parent");
check("handoff-before-parent-completion", !!parent && children.every(c => time(parent.completedAt) >= time(c.completedAt)),
"Parent completion must follow completion of every delegated deliverable.");
const parentReplies = o.comments.filter(c => c.issueId === parentId && c.authorAgentId === leadId && typeof c.body === "string");
check("visible-result-link", parentReplies.some(c => [...String(c.body).matchAll(/\[[^\]]+\]\(([^)]+)\)/g)].some(match => {
try {
const url = new URL(match[1]!, input.origin);
const parts = url.pathname.split("/").filter(Boolean);
const issueIndex = parts.indexOf("issues");
if (issueIndex < 0) return false;
const id = parts[issueIndex + 1];
return url.origin === new URL(input.origin).origin && url.hash === "#document-result" && [parent?.id, parent?.identifier].includes(id);
} catch { return false; }
})),
"The lead must publish a clickable link to the saved result.");
const parallelOffered = caseId === "parallel" && alex.length === 1 && riley.length === 1
? Math.max(time(alex[0].createdAt), time(riley[0].createdAt)) < Math.min(time(alex[0].completedAt), time(riley[0].completedAt)) : null;
// Scheduling opportunity is reported separately from correctness. Host/provider
// serialization does not turn correct independently owned work into failure.
return { schema: PLAN_GRADER, passed: checks.every(c => c.passed), checks,
measurements: { childCount: children.length, unnecessaryChildCount: Math.max(0, children.length - expectedChildren),
runCount: o.runs.length, parallelOfferedBeforeFirstCompletion: parallelOffered } };
}