mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 21:03:51 +02:00
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
88 lines
5.5 KiB
Markdown
88 lines
5.5 KiB
Markdown
# Planning guidance utility
|
|
|
|
`plan-task-guidance` is an explicit-only Product E2E suite. It runs the real UI,
|
|
public API, native Codex and saved task graph on four scenarios: one cohesive
|
|
order summary; two independent specialist approvals; a release depending on a
|
|
saved upstream result; and an independent adverse audit.
|
|
|
|
Each scenario compares current (archived from master
|
|
`72ff3a9f27e581a27acb49771e8658bbb0bbaa47`), short (checkout production files),
|
|
and disabled planning skills. The fixture creates editable company-owned copies through the public API,
|
|
preserving each source file byte-for-byte, then explicitly selects those copies
|
|
before any task starts. Bundled/catalog skills remain read-only. Both runtime plan conversion
|
|
and catalog task planning change together. All agents get the same treatment.
|
|
Disabled leaves those copies unassigned, with an empty explicit skill selection.
|
|
The company library and automatically restored core inventory remain available;
|
|
this is a selection ablation, not proof that a determined agent cannot discover
|
|
unassigned guidance.
|
|
The eval does not exercise automatic accepted-plan selection or installed-skill
|
|
migration, and cannot by itself qualify deleting that wiring.
|
|
|
|
The business prompts do not prescribe task counts or particular tools. Enabled
|
|
variants append explicit references to the two assigned runtime skill names; the
|
|
unassigned control appends none. This measures explicitly invoked guidance, not
|
|
automatic discovery. A separate exposure gate requires the initial lead run to
|
|
submit exactly those native skill inputs, bound to the immutable SKILL.md file
|
|
digests, and records provider turn acceptance. The controller emits the submission
|
|
receipt before a provider turn ID exists: this is a run-bound submission contract,
|
|
not a claimed provider call-ID join or proof of cognitive consumption. Native
|
|
protocol calibration separately verifies mapping into the isolated provider skill
|
|
directory. Missing or wrong exposure fails qualification even when the output
|
|
checks pass.
|
|
|
|
The scenario prompts do not prescribe task counts or particular tools. They do
|
|
specify which specialist is accountable and what business output is required.
|
|
The grader checks independent arithmetic, actual latest document revisions and
|
|
author/run attribution through exact document-ID/revision-number activity joins,
|
|
owned work items, prerequisite execution order, reviewer write boundaries, and
|
|
completion handoffs. It reports parallel scheduling opportunity separately from
|
|
correctness. No run can pass through an agent-authored self-assessment.
|
|
|
|
Twelve local cells, one attempt each, 12-minute cell deadline, 360-second provider
|
|
deadline, eight total run records, and 500-cent company/per-agent budget hard stops
|
|
bound the initial comparison. Coordination wakes count. Failed or missing evidence
|
|
stays failed. No automatic retries. The initial paid selection is one exact pilot,
|
|
then the eleven remaining cells only if fixture admission is sound. All cells use
|
|
the same native Codex model and production tool contracts. This is a bounded pilot,
|
|
not a cross-model reliability, speed, or cost claim.
|
|
|
|
```sh
|
|
pnpm test:e2e:runner:typecheck
|
|
pnpm test:e2e:runner:unit
|
|
pnpm test:e2e:runner -- --list --suite plan-task-guidance
|
|
pnpm test:e2e:runner -- --id plan-task-guidance.runner-codex.local.short-cohesive --max-parallel 1
|
|
```
|
|
|
|
Each attempt keeps `plan-task-source.json`, `plan-task-guidance.json`,
|
|
`plan-task-runs.json`, API state, screenshots and the existing run/accounting
|
|
reports. Selections and served hashes establish availability, not successful
|
|
skill invocation or cognitive consumption. Word/UTF-8 byte counts are exact file
|
|
measurements; provider context, token usage and billing coverage come from actual
|
|
retained run records. Cleanup, credential scanning and public projection use the
|
|
existing Product E2E pipeline. Do not publish raw sessions or hidden reasoning.
|
|
|
|
Setup attempt [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253)
|
|
on source `843238f43cb266f6ca0bc9d255881558ffd2eb71` was cancelled after source
|
|
review identified unsupported bundled edits and automatic core reinstallation.
|
|
The paid-cell step was skipped: zero provider runs and no behavioral grade.
|
|
The corrected fixture uses company copies and selection only.
|
|
|
|
The next setup campaign, [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799),
|
|
measured source `370e51d110836b942e5f90567d2bbe260bcc0f3a` using trusted workflow
|
|
revision `0e0b63e5a551388ac4601ed982b3e8f1c772f123` (workflow blob
|
|
`0600886144d3e22ea2e4a38329a79177882f3948`). It remains an original FAIL,
|
|
classified by the harness as `candidate_failure`. Inspection shows a browser
|
|
fixture error before task creation: the shared helper waited for the old Task
|
|
title field. `runIds`, the company run ledger, and planning observation are all
|
|
empty; no model executed. Cleanup passed. The billing summary's runCount=1 is
|
|
its minimum-one placeholder (`billing.ts`), not evidence of a provider run;
|
|
runtime/actual charges remain unmetered. No original result is regraded.
|
|
|
|
The corrected planning-only browser helper uses the current description composer,
|
|
explicitly selects the owner, and captures the public task-create response ID. A
|
|
real browser/server/database calibration creates the exact prompt/assignment with
|
|
paused non-provider agents, confirms zero run rows, and deletes its company. It
|
|
passes. Before any model execution, prompts also explicitly name the already
|
|
required result document key and exact JSON fields, avoiding an unstated oracle
|
|
format assumption. The legacy helper and production UI remain unchanged.
|