Files
PaperClipAI/doc/plans/2026-10-05-plan-task-guidance.md
DottaandPaperclip 16b7db35ff Shorten planning skills and measure task decomposition (#15296)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Planning guidance helps agents choose owners and dependencies.
> - The runtime skill favors few tasks, but the catalog skill requires a
child-task breakdown.
> - Both add repeated process instructions that can distract from the
requested outcome.
> - This change keeps the ownership and dependency rules and removes the
required matrix and repeated checklist.
> - A bounded Product E2E comparison measures saved outcomes and task
handoffs before qualification.

## Linked Issues or Issue Description

Refs #11057. Related measurement work: #15218.

**What existing behavior does this improve?**
Planning and delegation through the runtime plan-to-tasks and bundled
task-planning skills.

**Current behavior**
The two skills contain about 1,900 words and conflicting guidance on
whether plans require child tasks.

**Proposed behavior**
Keep cohesive work with one owner. Split only for a real owner, parallel
output, dependency, independent review, or follow-up lifecycle. Preserve
existing authorization and planning mechanics.

## What Changed

- Shorten both skills to about 400 words combined. Preserve their keys
and installed-version behavior.
- Remove the duplicate operational-skill pointer and regenerate affected
source metadata.
- Add twelve explicit Product E2E cells: four scenarios with current,
short and disabled planning skills.
- Use the current task composer and actual create-response ID; calibrate
public skill APIs and browser creation without providers.
- Eliminate an observed collision in chat-test company prefixes with a
per-suite sequence.
- Grade saved documents, exact author/run attribution, child count,
prerequisite execution order, review boundaries and completion handoffs.
- Retain current skill bytes and report source, selections, run
accounting and failures.

## Verification

- `pnpm test:e2e:runner:typecheck`: pass.
- `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks
pass.
- `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve
local Codex cells.
- Archived current skills match master
`72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly.
- Corrected fixture: three real public-API/database calibrations pass
with zero provider runs; all 35 evaluator checks and Product E2E
typecheck pass.
- Setup campaign
[37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253)
was canceled after source review found unsupported bundled edits and
automatic core reinstallation. Its paid-cell step was skipped: zero
provider runs, no behavioral grade.
- The next setup
[37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799)
failed before task creation on the old title-field selector: zero actual
runs, original FAIL retained, cleanup passed. A real browser/API
calibration of the new helper passes with paused non-provider agents and
zero runs.
- Full local typecheck/build pass. Full local tests retain one unchanged
five-minute Git streaming timeout (also fails isolated), 9,591 passes
and 5,796 skips. CI's chat failure was a proven random fixture-prefix
collision; five affected cases pass after the test-only repair.
- Paid behavior comparison and new-head CI/review remain pending. This
PR remains a draft.

## Risks

- The shorter text may change delegation decisions. Live outcomes are
not yet qualified.
- The initial comparison uses one profile and one attempt per cell. It
cannot establish cross-model reliability or cost trends.
- Disabled means unassigned company-owned copies; the company library
remains discoverable. This does not qualify global removal, automatic
accepted-plan wiring changes, or installed-copy migration.
- Skill availability does not prove a model read or cognitively used it.
- No provider/tool protocol, permission, timeout or runtime lifecycle
behavior changes in production.

## Model Used

OpenAI Codex (GPT-6), with repository inspection, code editing and tool
use. The exact backend model ID and context-window size are not exposed
in this session. The declared eval model is native Codex `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:36:01 -05:00

6.2 KiB

Planning guidance reduction — 5 October 2026

Status: implementation and eval preparation; no live result yet.

Decision and scope

Shorten the runtime plan-to-tasks skill and the bundled task-planning skill together. Preserve skill identities, installed-version behavior, automatic accepted-plan selection, authorization, and native/legacy mechanics. Remove the duplicated operational-skill pointer. Do not merge without the human's authorization.

Base: 72ff3a9f27e581a27acb49771e8658bbb0bbaa47. Prior work on PR #15218 and its separate checklist remains untouched.

Bounded comparison

Four ordinary execution scenarios: cohesive work; independent specialist outputs; a real prerequisite between specialists; an independent adverse review. Compare current, short, and disabled planning skills using public company-copy creation and selection APIs in isolated companies. Archived current skill bytes come from the base above. The disabled variant is an unassigned-selection ablation (the library remains discoverable), not a production removal or an accepted-plan continuation test. Other instructions, tools, models, permissions, fixtures, graders, budgets, and deadlines stay matched.

Initial scope: native Codex default profile, local, 12 cells, one attempt each, 500-cent company and per-agent hard stops, at most eight total run records per cell, including coordination wakes. One exact pilot cell precedes the remaining selection. No automatic retries or outcome rerolls. User authorization for paid runs persists. Extra profiles/repetitions are not part of this initial comparison.

Judge independently saved results, authorship, child boundaries, dependency order, review delivery, and parent completion. Retain all run records, failures, missing evidence, source/fixture hashes, skill selections and served bytes, screenshots, usage and billing coverage. Byte/word reductions are not token or cost savings. Skill availability is not proof that a model read or cognitively used a skill.

Run credential-free source, catalog, type and grader calibration checks before providers. Preserve positive, plausible wrong, and missing-evidence calibration. Report exact scenario pairs rather than equal aggregate totals. Short guidance is eligible only if the retained comparison supports it; inconclusive evidence stays explicit. Automatic removal and existing installed copies need separate migration and continuation qualification before a deletion recommendation can be shipped.

Credential-free admission

  • Product E2E TypeScript compilation: pass.
  • Product E2E support: 1,286 Vitest tests and 128 Node checks pass.
  • Catalog discovery: twelve explicit local Codex cells; excluded from --all.
  • Archived source hashes: conversion 08cb036df0e05b1d704dc0cd547c4e37b73078597c072a85ff982a2bb9b3a370; planning 9c52a44a30ec8d306119da51bf298e9e3e6c382a9a9559ffd3054bf5e0c75f36. Both match the named base Git blobs exactly.
  • Generated capability inventory and catalog are synchronized.
  • No provider call has started. Live outcomes remain unqualified.

Source review found two fixture defects before provider execution: bundled and catalog skills cannot be edited, and deleted core skills are automatically restored. Campaign 37399550253 at 843238f43cb266f6ca0bc9d255881558ffd2eb71 was cancelled; its paid-cell step was skipped, with zero provider runs and no behavioral grade. Corrected setup creates editable, byte-identical company copies and varies their explicit selection. No bundled deletion or in-place edit occurs.

Corrected fixture admission: all three current/short/unassigned variants pass real company-skill creation, content readback and native-agent selection APIs against a disposable PostgreSQL database. No heartbeat run rows were created. The evaluator's 35 positive/wrong/missing-evidence checks pass, including rejection of a bare issue-ID URL without the issues route. Product E2E typecheck passes. Initial local DB checks were skipped until the dependency's missing library symlinks were restored with its supplied postinstall script; skipped checks were not counted as passes.

The next setup campaign, 37401094799, measured source 370e51d110836b942e5f90567d2bbe260bcc0f3a using trusted workflow revision 0e0b63e5a551388ac4601ed982b3e8f1c772f123 (workflow blob 0600886144d3e22ea2e4a38329a79177882f3948). It remains an original FAIL, classified by the harness as candidate_failure. Inspection shows a browser fixture error before task creation: the shared helper waited for the old Task title field. runIds, the company run ledger, and planning observation are all empty; no model executed. Cleanup passed. The billing summary's runCount=1 is its minimum-one placeholder (billing.ts), not evidence of a provider run; runtime/actual charges remain unmetered. No original result is regraded.

The corrected planning-only browser helper uses the current description composer, explicitly selects the owner, and captures the public task-create response ID. A real browser/server/database calibration creates the exact prompt/assignment with paused non-provider agents, confirms zero run rows, and deletes its company. It passes. Before any model execution, prompts also explicitly name the already required result document key and exact JSON fields, avoiding an unstated oracle format assumption. The legacy helper and production UI remain unchanged.

Repository checks at the previous source: full typecheck and build pass. Local full tests: 9,591 pass, 5,796 skip and one unchanged macOS real-Git streaming test hits its five-minute deadline; an isolated check also fails. Preserve this limitation. Normal CI found a different, exact cause: truncated random fixture company prefixes collided in chat integration setup. An initial per-suite counter fixed within-run collisions but review found that its values repeat against an external database on subsequent runs. The corrected prefix contains the complete company UUID, with no truncation; distinct company IDs therefore produce distinct prefixes within and across runs. All five affected native-modal variants pass locally (1,058 unrelated cases filtered). This is a test-only repair; the measured planning fixture and production guidance remain frozen at 8538cfce4c2defdedc2efba7519bcfce880c4e00.