## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
5.5 KiB
Planning guidance utility
plan-task-guidance is an explicit-only Product E2E suite. It runs the real UI,
public API, native Codex and saved task graph on four scenarios: one cohesive
order summary; two independent specialist approvals; a release depending on a
saved upstream result; and an independent adverse audit.
Each scenario compares current (archived from master
72ff3a9f27e581a27acb49771e8658bbb0bbaa47), short (checkout production files),
and disabled planning skills. The fixture creates editable company-owned copies through the public API,
preserving each source file byte-for-byte, then explicitly selects those copies
before any task starts. Bundled/catalog skills remain read-only. Both runtime plan conversion
and catalog task planning change together. All agents get the same treatment.
Disabled leaves those copies unassigned, with an empty explicit skill selection.
The company library and automatically restored core inventory remain available;
this is a selection ablation, not proof that a determined agent cannot discover
unassigned guidance.
The eval does not exercise automatic accepted-plan selection or installed-skill
migration, and cannot by itself qualify deleting that wiring.
The business prompts do not prescribe task counts or particular tools. Enabled variants append explicit references to the two assigned runtime skill names; the unassigned control appends none. This measures explicitly invoked guidance, not automatic discovery. A separate exposure gate requires the initial lead run to submit exactly those native skill inputs, bound to the immutable SKILL.md file digests, and records provider turn acceptance. The controller emits the submission receipt before a provider turn ID exists: this is a run-bound submission contract, not a claimed provider call-ID join or proof of cognitive consumption. Native protocol calibration separately verifies mapping into the isolated provider skill directory. Missing or wrong exposure fails qualification even when the output checks pass.
The scenario prompts do not prescribe task counts or particular tools. They do specify which specialist is accountable and what business output is required. The grader checks independent arithmetic, actual latest document revisions and author/run attribution through exact document-ID/revision-number activity joins, owned work items, prerequisite execution order, reviewer write boundaries, and completion handoffs. It reports parallel scheduling opportunity separately from correctness. No run can pass through an agent-authored self-assessment.
Twelve local cells, one attempt each, 12-minute cell deadline, 360-second provider deadline, eight total run records, and 500-cent company/per-agent budget hard stops bound the initial comparison. Coordination wakes count. Failed or missing evidence stays failed. No automatic retries. The initial paid selection is one exact pilot, then the eleven remaining cells only if fixture admission is sound. All cells use the same native Codex model and production tool contracts. This is a bounded pilot, not a cross-model reliability, speed, or cost claim.
pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
pnpm test:e2e:runner -- --list --suite plan-task-guidance
pnpm test:e2e:runner -- --id plan-task-guidance.runner-codex.local.short-cohesive --max-parallel 1
Each attempt keeps plan-task-source.json, plan-task-guidance.json,
plan-task-runs.json, API state, screenshots and the existing run/accounting
reports. Selections and served hashes establish availability, not successful
skill invocation or cognitive consumption. Word/UTF-8 byte counts are exact file
measurements; provider context, token usage and billing coverage come from actual
retained run records. Cleanup, credential scanning and public projection use the
existing Product E2E pipeline. Do not publish raw sessions or hidden reasoning.
Setup attempt 37399550253
on source 843238f43cb266f6ca0bc9d255881558ffd2eb71 was cancelled after source
review identified unsupported bundled edits and automatic core reinstallation.
The paid-cell step was skipped: zero provider runs and no behavioral grade.
The corrected fixture uses company copies and selection only.
The next setup campaign, 37401094799,
measured source 370e51d110836b942e5f90567d2bbe260bcc0f3a using trusted workflow
revision 0e0b63e5a551388ac4601ed982b3e8f1c772f123 (workflow blob
0600886144d3e22ea2e4a38329a79177882f3948). It remains an original FAIL,
classified by the harness as candidate_failure. Inspection shows a browser
fixture error before task creation: the shared helper waited for the old Task
title field. runIds, the company run ledger, and planning observation are all
empty; no model executed. Cleanup passed. The billing summary's runCount=1 is
its minimum-one placeholder (billing.ts), not evidence of a provider run;
runtime/actual charges remain unmetered. No original result is regraded.
The corrected planning-only browser helper uses the current description composer, explicitly selects the owner, and captures the public task-create response ID. A real browser/server/database calibration creates the exact prompt/assignment with paused non-provider agents, confirms zero run rows, and deletes its company. It passes. Before any model execution, prompts also explicitly name the already required result document key and exact JSON fields, avoiding an unstated oracle format assumption. The legacy helper and production UI remain unchanged.