mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
146 lines
9.0 KiB
TypeScript
146 lines
9.0 KiB
TypeScript
import { describe, expect, it } from "vitest";
|
|
import { readFileSync } from "node:fs";
|
|
import { PLAN_CASES, PLAN_MAX_RUNS, PLAN_SKILLS, PLAN_VARIANTS, parsePlanCase, planHash, planScenario } from "./plan-task-cases.js";
|
|
import { gradePlanTask, planDocumentJson, type PlanObservation } from "./plan-task-scoring.js";
|
|
import { runnerSuites, runnerMatrix } from "./catalog.js";
|
|
import { parseRunnerSelectors, selectRunnerExecutions } from "./selectors.js";
|
|
|
|
const t = (n: number) => `2026-10-05T12:00:${String(n).padStart(2, "0")}.000Z`;
|
|
function valid(caseId: typeof PLAN_CASES[number]) {
|
|
const o: PlanObservation = { issues: [], runs: [], documents: [], comments: [], activity: [], interactions: [], wakes: [] };
|
|
function add(id: string, agentId: string, value: unknown, start: number, end: number) {
|
|
const body = JSON.stringify(value), runId = `run-${id}`, documentId = `doc-${id}`, revisionId = `rev-${id}`;
|
|
o.issues.push({ id, identifier: `T-${id}`, assigneeAgentId: agentId, status: "done", createdAt: t(0), completedAt: t(end) });
|
|
o.runs.push({ id: runId, agentId, status: "succeeded", runtimeMode: "native", runnerInstanceId: "runner", nativeSessionId: "session", nativeIssueId: id, startedAt: t(start), finishedAt: t(end) });
|
|
o.documents.push({ id: documentId, issueId: id, key: "result", body, latestRevisionId: revisionId, latestRevisionNumber: 1,
|
|
revisions: [{ id: revisionId, body, createdByAgentId: agentId }] });
|
|
o.activity.push({ action: "issue.document_created", entityId: id, agentId, runId, details: { key: "result", documentId, revisionNumber: 1 } });
|
|
o.wakes.push({ issueId: id, events: [], truncated: false });
|
|
}
|
|
const marker = "PLANfixture", order = { marker, units: 5, total: 29 };
|
|
let output: unknown = order;
|
|
if (caseId === "parallel") {
|
|
add("alex", "alex-agent", { marker, seats: 24 }, 2, 10);
|
|
add("riley", "riley-agent", { marker, welcome: "Hola, equipo" }, 3, 12);
|
|
output = { marker, seats: 24, welcome: "Hola, equipo" };
|
|
} else if (caseId === "dependency") {
|
|
add("alex", "alex-agent", order, 2, 10);
|
|
output = { marker, releasedTotal: 29, sourceRevisionId: "rev-alex" };
|
|
add("riley", "riley-agent", output, 11, 15);
|
|
} else if (caseId === "review") {
|
|
output = { marker, verdict: "reject", correctTotal: 29, difference: 1 };
|
|
add("riley", "riley-agent", output, 2, 10);
|
|
}
|
|
add("parent", "lead", output, 16, 20);
|
|
o.comments.push({ issueId: "parent", authorAgentId: "lead", body: "[Result](/T/issues/T-parent#document-result)" });
|
|
return { caseId, marker, parentId: "parent", leadId: "lead", alexId: "alex-agent", rileyId: "riley-agent", observation: o, maxRuns: PLAN_MAX_RUNS, origin: "http://127.0.0.1:3100" };
|
|
}
|
|
|
|
const passed = (value: ReturnType<typeof valid>, id: string) => gradePlanTask(value).checks.find(c => c.id === id)?.passed;
|
|
|
|
describe("planning guidance outcome calibration", () => {
|
|
it.each(PLAN_CASES)("accepts independently saved %s work", id => expect(gradePlanTask(valid(id)).passed).toBe(true));
|
|
it.each(PLAN_CASES)("rejects missing %s evidence", id => {
|
|
const input = valid(id); input.observation.documents = [];
|
|
expect(gradePlanTask(input).passed).toBe(false);
|
|
});
|
|
it("does not reward copying an expected answer without the specialist's authorship", () => {
|
|
const input = valid("parallel"); input.observation.documents[0]!.revisions[0]!.createdByAgentId = "lead";
|
|
expect(passed(input, "alex-authorship")).toBe(false);
|
|
});
|
|
it.each(["missing-run", "wrong-task", "wrong-revision", "duplicate-write"])("rejects %s attribution", kind => {
|
|
const input = valid("cohesive"), o = input.observation;
|
|
if (kind === "missing-run") o.activity[0]!.runId = "unknown";
|
|
if (kind === "wrong-task") o.runs[0]!.nativeIssueId = "other";
|
|
if (kind === "wrong-revision") o.activity[0]!.details.revisionNumber = 2;
|
|
if (kind === "duplicate-write") o.activity.push(structuredClone(o.activity[0]!));
|
|
expect(passed(input, "parent-authorship")).toBe(false);
|
|
});
|
|
it("rejects a correct-looking answer with wrong arithmetic", () => {
|
|
const input = valid("cohesive"); input.observation.documents[0]!.body = JSON.stringify({ marker: input.marker, units: 5, total: 30 });
|
|
expect(passed(input, "parent-saved-output")).toBe(false);
|
|
});
|
|
it("rejects unnecessary children for cohesive work", () => {
|
|
const input = valid("cohesive"); input.observation.issues.push({ id: "unnecessary", assigneeAgentId: "lead", status: "done" });
|
|
expect(passed(input, "minimal-owned-work")).toBe(false);
|
|
});
|
|
it("rejects downstream execution before its prerequisite", () => {
|
|
const input = valid("dependency"); input.observation.runs[1]!.startedAt = t(3);
|
|
expect(passed(input, "prerequisite-before-execution")).toBe(false);
|
|
});
|
|
it("rejects an invented upstream revision", () => {
|
|
const input = valid("dependency"); const doc = input.observation.documents[1]!;
|
|
doc.body = doc.body.replace("rev-alex", "invented");
|
|
expect(passed(input, "riley-saved-output")).toBe(false);
|
|
});
|
|
it("rejects missing dependency timestamps", () => {
|
|
const input = valid("dependency"); delete input.observation.issues[0]!.completedAt;
|
|
expect(passed(input, "prerequisite-before-execution")).toBe(false);
|
|
});
|
|
it("rejects closing the parent before the delegated output is complete", () => {
|
|
const input = valid("review"); input.observation.issues.at(-1)!.completedAt = t(5);
|
|
expect(passed(input, "handoff-before-parent-completion")).toBe(false);
|
|
});
|
|
it("rejects an adverse review stranded as blocked", () => {
|
|
const input = valid("review"); input.observation.issues[0]!.status = "blocked";
|
|
expect(passed(input, "tasks-completed")).toBe(false);
|
|
});
|
|
it("rejects reviewer writes on the parent", () => {
|
|
const input = valid("review"); input.observation.comments.push({ issueId: "parent", authorAgentId: "riley-agent", body: "reject" });
|
|
expect(passed(input, "review-write-boundary")).toBe(false);
|
|
});
|
|
it.each(["failed", "cancelled", "timed_out"])("retains %s runs as failures", status => {
|
|
const input = valid("parallel"); input.observation.runs[0]!.status = status;
|
|
expect(passed(input, "bounded-native-runs")).toBe(false);
|
|
});
|
|
it("rejects missing wake evidence and pending recovery", () => {
|
|
const input = valid("cohesive"); input.observation.wakes = [];
|
|
expect(passed(input, "tasks-completed")).toBe(false);
|
|
const recovery = valid("cohesive"); recovery.observation.issues[0]!.scheduledRetry = { reason: "retry" };
|
|
expect(passed(recovery, "tasks-completed")).toBe(false);
|
|
});
|
|
it.each(["/T-parent#document-result", "/T/issues/other#document-result", "https://unrelated.example/T/issues/T-parent#document-result", "/T/issues/T-parent#document-other"])("rejects unrelated result link %s", url => {
|
|
const input = valid("cohesive"); input.observation.comments[0]!.body = `[Result](${url})`;
|
|
expect(passed(input, "visible-result-link")).toBe(false);
|
|
});
|
|
it("records serial scheduling without confusing it with wrong output", () => {
|
|
const input = valid("parallel"); input.observation.issues[1]!.createdAt = t(11);
|
|
expect(gradePlanTask(input).measurements.parallelOfferedBeforeFirstCompletion).toBe(false);
|
|
expect(gradePlanTask(input).passed).toBe(true);
|
|
});
|
|
it("parses whole JSON documents or one JSON fence without scraping arbitrary prose", () => {
|
|
expect(planDocumentJson('```json\n{"value":1}\n```')).toEqual({ value: 1 });
|
|
expect(planDocumentJson('Claim: {"value":1}')).toBeUndefined();
|
|
});
|
|
});
|
|
|
|
describe("bounded planning comparison admission", () => {
|
|
it("declares only 12 explicit local Codex cells with one attempt", () => {
|
|
const suite = runnerSuites.find(s => s.id === "plan-task-guidance")!;
|
|
expect(suite.manualOnly).toBe(true);
|
|
expect(runnerMatrix.filter(e => e.suite.id === suite.id)).toHaveLength(12);
|
|
expect(suite.tasks.every(t => t.automaticRetryPolicy === "single_attempt" && t.expectedRunCount === 8)).toBe(true);
|
|
expect(selectRunnerExecutions(parseRunnerSelectors(["--all"])).some(e => e.suite.id === suite.id)).toBe(false);
|
|
});
|
|
it("keeps task prompts identical across variants", () => {
|
|
const suite = runnerSuites.find(s => s.id === "plan-task-guidance")!;
|
|
for (const id of PLAN_CASES) expect(new Set(suite.tasks.filter(t => parsePlanCase(t.id).caseId === id).map(t => t.buildPrompt("same"))).size).toBe(1);
|
|
expect(PLAN_VARIANTS).toHaveLength(3);
|
|
});
|
|
it("retains current source and materially smaller candidate guidance", () => {
|
|
for (const skill of PLAN_SKILLS) {
|
|
const old = readFileSync(new URL(`../../${skill.current}`, import.meta.url), "utf8");
|
|
const short = readFileSync(new URL(`../../${skill.short}`, import.meta.url), "utf8");
|
|
expect(Buffer.byteLength(short)).toBeLessThan(Buffer.byteLength(old) / 3);
|
|
expect(planHash(short)).not.toBe(planHash(old));
|
|
}
|
|
});
|
|
it("does not give the arithmetic or review answers in the scenario", () => {
|
|
for (const id of PLAN_CASES) {
|
|
const prompt = planScenario(id, "same").prompt;
|
|
expect(prompt).not.toContain('"total":29');
|
|
expect(prompt).not.toContain('"verdict":"reject"');
|
|
}
|
|
});
|
|
});
|