mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
107 lines
7.7 KiB
TypeScript
107 lines
7.7 KiB
TypeScript
import { isDeepStrictEqual } from "node:util";
|
|
import type { PlanCase } from "./plan-task-cases.js";
|
|
|
|
export type PlanRow = Record<string, any>;
|
|
export interface PlanDocument { id: string; latestRevisionNumber: number; issueId: string; key: string; body: string; latestRevisionId: string; revisions: PlanRow[] }
|
|
export interface PlanObservation {
|
|
issues: PlanRow[]; runs: PlanRow[]; documents: PlanDocument[];
|
|
comments: PlanRow[]; activity: PlanRow[]; interactions: PlanRow[]; wakes: PlanRow[];
|
|
}
|
|
export interface PlanCheck { id: string; passed: boolean; detail: string }
|
|
export const PLAN_GRADER = "paperclip.plan-task-guidance.v1";
|
|
export function planDocumentJson(body: unknown): unknown {
|
|
if (typeof body !== "string") return undefined;
|
|
const text = body.trim();
|
|
const fenced = text.match(/^```(?:json)?\s*\n([\s\S]*?)\n```$/i);
|
|
try { return JSON.parse(fenced?.[1] ?? text); } catch { return undefined; }
|
|
}
|
|
const time = (value: unknown) => typeof value === "string" ? Date.parse(value) : NaN;
|
|
|
|
export function gradePlanTask(input: {
|
|
caseId: PlanCase; marker: string; parentId: string; leadId: string; alexId: string; rileyId: string;
|
|
observation: PlanObservation; maxRuns: number; origin: string;
|
|
}) {
|
|
const { caseId, marker, parentId, leadId, alexId, rileyId, observation: o } = input;
|
|
const checks: PlanCheck[] = [];
|
|
const check = (id: string, passed: boolean, detail: string) => checks.push({ id, passed, detail });
|
|
const parent = o.issues.find(i => i.id === parentId);
|
|
const children = o.issues.filter(i => i.id !== parentId);
|
|
const alex = children.filter(i => i.assigneeAgentId === alexId);
|
|
const riley = children.filter(i => i.assigneeAgentId === rileyId);
|
|
const expectedChildren = caseId === "cohesive" ? 0 : caseId === "review" ? 1 : 2;
|
|
const pending = o.interactions.some(i => ["pending", "open"].includes(i.status)) || o.wakes.some(w => w.truncated ||
|
|
w.events?.some((event: PlanRow) => event.kind === "wake_request" && !["completed", "skipped", "failed", "cancelled"].includes(event.status) &&
|
|
!(event.status === "coalesced" && o.runs.some(r => r.id === event.runId && r.status === "succeeded"))));
|
|
check("tasks-completed", Boolean(parent) && o.issues.every(i => i.status === "done" && !i.scheduledRetry && !i.activeRecoveryAction) && !pending &&
|
|
o.issues.every(i => o.wakes.filter(w => w.issueId === i.id && Array.isArray(w.events) && !w.truncated).length === 1),
|
|
"Every requested task must be done with no waiting decision, retry, recovery or pending wake; an adverse review is a completed deliverable.");
|
|
check("bounded-native-runs", o.runs.length > 0 && o.runs.length <= input.maxRuns && o.runs.every(r => r.status === "succeeded" && r.runtimeMode === "native" &&
|
|
r.runnerInstanceId && r.nativeSessionId && !r.retryOfRunId && time(r.finishedAt) >= time(r.startedAt)),
|
|
`Observed ${o.runs.length} total run records; all must be successful native executions with identity and finite ordered timestamps, without retries.`);
|
|
check("minimal-owned-work", children.length === expectedChildren && (caseId === "cohesive" ||
|
|
(riley.length === 1 && (caseId === "review" ? alex.length === 0 : alex.length === 1))),
|
|
`Expected ${expectedChildren} independent work items for the stated owners; found ${children.length}.`);
|
|
|
|
function document(issue: PlanRow | undefined, owner: string, expected: unknown, label: string) {
|
|
const docs = issue ? o.documents.filter(d => d.issueId === issue.id && d.key === "result") : [];
|
|
const doc = docs.length === 1 ? docs[0] : undefined;
|
|
const revision = doc?.revisions.find(r => r.id === doc.latestRevisionId);
|
|
const writes = o.activity.filter(a => ["issue.document_created", "issue.document_updated"].includes(a.action) &&
|
|
a.entityId === issue?.id && a.details?.documentId === doc?.id && a.details?.key === "result" &&
|
|
a.details?.revisionNumber === doc?.latestRevisionNumber);
|
|
const write = writes.length === 1 ? writes[0] : undefined;
|
|
const run = o.runs.find(r => r.id === write?.runId);
|
|
const runIssue = run?.nativeIssueId ?? run?.contextSnapshot?.issueId ?? run?.contextSnapshot?.taskId;
|
|
check(`${label}-saved-output`, !!doc && isDeepStrictEqual(planDocumentJson(doc.body), expected), "The saved result must match the independently calculated business output.");
|
|
check(`${label}-authorship`, !!revision && revision.createdByAgentId === owner && !revision.createdByUserId &&
|
|
write?.agentId === owner && run?.agentId === owner && runIssue === issue?.id && revision.body === doc?.body,
|
|
"The latest document revision must join its exact author and run on the owned task; assignment or a claimed signature alone is insufficient.");
|
|
return doc;
|
|
}
|
|
const order = { marker, units: 5, total: 29 };
|
|
const verdict = { marker, verdict: "reject", correctTotal: 29, difference: 1 };
|
|
let expected: unknown = order;
|
|
if (caseId === "parallel") {
|
|
document(alex[0], alexId, { marker, seats: 24 }, "alex");
|
|
document(riley[0], rileyId, { marker, welcome: "Hola, equipo" }, "riley");
|
|
expected = { marker, seats: 24, welcome: "Hola, equipo" };
|
|
}
|
|
if (caseId === "dependency") {
|
|
const source = document(alex[0], alexId, order, "alex");
|
|
expected = { marker, releasedTotal: 29, sourceRevisionId: source?.latestRevisionId };
|
|
document(riley[0], rileyId, expected, "riley");
|
|
const downstream = o.runs.filter(r => (r.nativeIssueId ?? r.contextSnapshot?.issueId ?? r.contextSnapshot?.taskId) === riley[0]?.id);
|
|
check("prerequisite-before-execution", downstream.length > 0 && downstream.every(r => time(r.startedAt) >= time(alex[0]?.completedAt)),
|
|
"No downstream provider execution may precede the prerequisite's persisted completion; late creation is also a valid way to wait.");
|
|
}
|
|
if (caseId === "review") {
|
|
document(riley[0], rileyId, verdict, "independent-review");
|
|
expected = verdict;
|
|
check("review-write-boundary", !o.comments.some(c => c.issueId === parentId && c.authorAgentId === rileyId) &&
|
|
!o.documents.filter(d => d.issueId === parentId).some(d => d.revisions.some(r => r.createdByAgentId === rileyId)),
|
|
"The independent reviewer must deliver on its own task without parent writes.");
|
|
}
|
|
document(parent, leadId, expected, "parent");
|
|
check("handoff-before-parent-completion", !!parent && children.every(c => time(parent.completedAt) >= time(c.completedAt)),
|
|
"Parent completion must follow completion of every delegated deliverable.");
|
|
const parentReplies = o.comments.filter(c => c.issueId === parentId && c.authorAgentId === leadId && typeof c.body === "string");
|
|
check("visible-result-link", parentReplies.some(c => [...String(c.body).matchAll(/\[[^\]]+\]\(([^)]+)\)/g)].some(match => {
|
|
try {
|
|
const url = new URL(match[1]!, input.origin);
|
|
const parts = url.pathname.split("/").filter(Boolean);
|
|
const issueIndex = parts.indexOf("issues");
|
|
if (issueIndex < 0) return false;
|
|
const id = parts[issueIndex + 1];
|
|
return url.origin === new URL(input.origin).origin && url.hash === "#document-result" && [parent?.id, parent?.identifier].includes(id);
|
|
} catch { return false; }
|
|
})),
|
|
"The lead must publish a clickable link to the saved result.");
|
|
const parallelOffered = caseId === "parallel" && alex.length === 1 && riley.length === 1
|
|
? Math.max(time(alex[0].createdAt), time(riley[0].createdAt)) < Math.min(time(alex[0].completedAt), time(riley[0].completedAt)) : null;
|
|
// Scheduling opportunity is reported separately from correctness. Host/provider
|
|
// serialization does not turn correct independently owned work into failure.
|
|
return { schema: PLAN_GRADER, passed: checks.every(c => c.passed), checks,
|
|
measurements: { childCount: children.length, unnecessaryChildCount: Math.max(0, children.length - expectedChildren),
|
|
runCount: o.runs.length, parallelOfferedBeforeFirstCompletion: parallelOffered } };
|
|
}
|