mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
59 lines
6.1 KiB
TypeScript
59 lines
6.1 KiB
TypeScript
import { createHash } from "node:crypto";
|
|
import { readFileSync } from "node:fs";
|
|
import type { RunnerProfileFixture, RunnerTaskFixture } from "./types.js";
|
|
|
|
export const PLAN_VARIANTS = ["current", "short", "disabled"] as const;
|
|
export const PLAN_CASES = ["cohesive", "parallel", "dependency", "review"] as const;
|
|
export type PlanVariant = typeof PLAN_VARIANTS[number];
|
|
export type PlanCase = typeof PLAN_CASES[number];
|
|
export const PLAN_BUDGET_CENTS = 500;
|
|
export const PLAN_MAX_RUNS = 8;
|
|
export const PLAN_BASE_SHA = "72ff3a9f27e581a27acb49771e8658bbb0bbaa47";
|
|
export const PLAN_SKILLS = [
|
|
{ key: "paperclipai/paperclip/paperclip-converting-plans-to-tasks", current: "tests/runner-e2e/fixtures/plan-task-guidance/current-conversion.md", short: "skills/paperclip-converting-plans-to-tasks/SKILL.md" },
|
|
{ key: "paperclipai/bundled/paperclip-operations/task-planning", current: "tests/runner-e2e/fixtures/plan-task-guidance/current-planning.md", short: "packages/skills-catalog/catalog/bundled/paperclip-operations/task-planning/SKILL.md" },
|
|
] as const;
|
|
export const planHash = (value: string | Buffer) => createHash("sha256").update(value).digest("hex");
|
|
export const planDefinitionDigest = planHash([
|
|
"plan-task-cases.ts", "plan-task-flow.ts", "plan-task-scoring.ts", "plan-task-skills.ts", "plan-task-ui.ts", "plan-task-exposure.ts",
|
|
"fixtures/plan-task-guidance/current-conversion.md", "fixtures/plan-task-guidance/current-planning.md",
|
|
].map(file => readFileSync(new URL(file, import.meta.url), "utf8")).join("\0"));
|
|
|
|
export function parsePlanCase(id: string): { variant: PlanVariant; caseId: PlanCase } {
|
|
const [variant, caseId, extra] = id.split("-");
|
|
if (extra || !PLAN_VARIANTS.includes(variant as PlanVariant) || !PLAN_CASES.includes(caseId as PlanCase)) throw new Error(`Unknown plan-task case ${id}`);
|
|
return { variant: variant as PlanVariant, caseId: caseId as PlanCase };
|
|
}
|
|
|
|
export function planScenario(caseId: PlanCase, nonce: string) {
|
|
const marker = `PLAN${nonce.replace(/[^a-z0-9]/gi, "")}`;
|
|
const prompts: Record<PlanCase, string> = {
|
|
cohesive: `Prepare a verified order summary from two units at 7 dollars each and three units at 5 dollars each. Save a document on this task with key result containing JSON with marker "${marker}", units (the summed quantity), and total (the calculated dollar amount). Check the arithmetic and finish the requested work.`,
|
|
parallel: `Prepare our launch packet. Alex Metrics is accountable for independently calculating and signing off the seat total from groups of 11 and 13. Riley Copy is accountable for independently approving our Spanish welcome text, whose supplied source is "Hola, equipo". Each specialist must save their own signed result document on their work item: Alex uses fields marker "${marker}" and seats (the calculated total); Riley uses {"marker":"${marker}","welcome":"Hola, equipo"}. Their inputs are independent. Compile their actual saved approvals into this task's result document with fields marker "${marker}", seats and welcome and finish the packet.`,
|
|
dependency: `Prepare the release receipt. Alex Metrics owns verifying the order: two units at 7 dollars plus three at 5 dollars. Alex must save a result document with fields marker "${marker}", units (the summed quantity), and total (the calculated dollar amount) on Alex's work item. Riley Copy owns release verification and must inspect that actual saved result after Alex completes it; Riley cannot perform release work before that prerequisite is complete. Riley's result document must contain fields marker "${marker}", releasedTotal (the verified dollar total), and sourceRevisionId (the actual revision ID of Alex's result). Deliver Riley's verified receipt as this task's result document and finish.`,
|
|
review: `Obtain Riley Copy's independent audit of this supplied order proposal: two units at 7 dollars, three units at 5 dollars, proposed total 30 dollars. Riley must author a result document on Riley's review work item with fields marker "${marker}", verdict ("accept" or "reject"), correctTotal (the independently calculated dollar total), and difference (the absolute error in the proposal). The audit, including any adverse finding, is the complete requested deliverable; do not implement a correction. Report Riley's actual verdict in this task's result document with those same fields and finish.`,
|
|
};
|
|
return { marker, prompt: `Plan the work briefly, then carry it out. Execution is already authorized; no additional approval is requested. Choose how to organize the work. ${prompts[caseId]} Save every requested result document with key result and exactly the requested JSON fields. Documents may use a JSON code block. Link the final saved result in your response.` };
|
|
}
|
|
|
|
export function planTaskProfile(profile: RunnerProfileFixture): RunnerProfileFixture {
|
|
return { ...profile, buildAgent(input) {
|
|
const payload = profile.buildAgent(input);
|
|
return { ...payload, name: `Operations Lead ${input.executionId}`, role: "general", title: "Operations Lead",
|
|
adapterConfig: { ...(payload.adapterConfig as Record<string, unknown>), timeoutSec: 360 },
|
|
budgetMonthlyCents: PLAN_BUDGET_CENTS, capabilities: "Completes operational requests and coordinates available specialists.",
|
|
instructionsBundle: { entryFile: "AGENTS.md", files: { "AGENTS.md": "You coordinate a small operations team. Deliver accurate, verified work that meets the user's requirements." } } };
|
|
} };
|
|
}
|
|
|
|
export const planTaskCases: readonly RunnerTaskFixture[] = PLAN_VARIANTS.flatMap(variant => PLAN_CASES.map(caseId => ({
|
|
id: `${variant}-${caseId}`, label: `${variant} planning guidance: ${caseId}`, groups: [], workMode: "standard" as const,
|
|
flow: "plan_task_guidance" as const, expectedRunCount: PLAN_MAX_RUNS, minimumExpectedRunCount: 1,
|
|
automaticRetryPolicy: "single_attempt" as const, attemptTimeoutMs: { local: 12 * 60_000, daytona: 12 * 60_000 },
|
|
expectedTerminalState: { issue: "done" as const, run: "succeeded" as const },
|
|
buildTitle: (nonce: string) => `Planning ${caseId} ${nonce}`,
|
|
buildPrompt: (nonce: string) => planScenario(caseId, nonce).prompt,
|
|
buildVisibleMarker: (nonce: string) => planScenario(caseId, nonce).marker,
|
|
buildMatchers: () => [],
|
|
})));
|