Files
PaperClipAI/tests/runner-e2e/plan-task-exposure.ts
T
DottaandPaperclip 16b7db35ff Shorten planning skills and measure task decomposition (#15296)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Planning guidance helps agents choose owners and dependencies.
> - The runtime skill favors few tasks, but the catalog skill requires a
child-task breakdown.
> - Both add repeated process instructions that can distract from the
requested outcome.
> - This change keeps the ownership and dependency rules and removes the
required matrix and repeated checklist.
> - A bounded Product E2E comparison measures saved outcomes and task
handoffs before qualification.

## Linked Issues or Issue Description

Refs #11057. Related measurement work: #15218.

**What existing behavior does this improve?**
Planning and delegation through the runtime plan-to-tasks and bundled
task-planning skills.

**Current behavior**
The two skills contain about 1,900 words and conflicting guidance on
whether plans require child tasks.

**Proposed behavior**
Keep cohesive work with one owner. Split only for a real owner, parallel
output, dependency, independent review, or follow-up lifecycle. Preserve
existing authorization and planning mechanics.

## What Changed

- Shorten both skills to about 400 words combined. Preserve their keys
and installed-version behavior.
- Remove the duplicate operational-skill pointer and regenerate affected
source metadata.
- Add twelve explicit Product E2E cells: four scenarios with current,
short and disabled planning skills.
- Use the current task composer and actual create-response ID; calibrate
public skill APIs and browser creation without providers.
- Eliminate an observed collision in chat-test company prefixes with a
per-suite sequence.
- Grade saved documents, exact author/run attribution, child count,
prerequisite execution order, review boundaries and completion handoffs.
- Retain current skill bytes and report source, selections, run
accounting and failures.

## Verification

- `pnpm test:e2e:runner:typecheck`: pass.
- `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks
pass.
- `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve
local Codex cells.
- Archived current skills match master
`72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly.
- Corrected fixture: three real public-API/database calibrations pass
with zero provider runs; all 35 evaluator checks and Product E2E
typecheck pass.
- Setup campaign
[37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253)
was canceled after source review found unsupported bundled edits and
automatic core reinstallation. Its paid-cell step was skipped: zero
provider runs, no behavioral grade.
- The next setup
[37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799)
failed before task creation on the old title-field selector: zero actual
runs, original FAIL retained, cleanup passed. A real browser/API
calibration of the new helper passes with paused non-provider agents and
zero runs.
- Full local typecheck/build pass. Full local tests retain one unchanged
five-minute Git streaming timeout (also fails isolated), 9,591 passes
and 5,796 skips. CI's chat failure was a proven random fixture-prefix
collision; five affected cases pass after the test-only repair.
- Paid behavior comparison and new-head CI/review remain pending. This
PR remains a draft.

## Risks

- The shorter text may change delegation decisions. Live outcomes are
not yet qualified.
- The initial comparison uses one profile and one attempt per cell. It
cannot establish cross-model reliability or cost trends.
- Disabled means unassigned company-owned copies; the company library
remains discoverable. This does not qualify global removal, automatic
accepted-plan wiring changes, or installed-copy migration.
- Skill availability does not prove a model read or cognitively used it.
- No provider/tool protocol, permission, timeout or runtime lifecycle
behavior changes in production.

## Model Used

OpenAI Codex (GPT-6), with repository inspection, code editing and tool
use. The exact backend model ID and context-window size are not exposed
in this session. The declared eval model is native Codex `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:36:01 -05:00

47 lines
3.5 KiB
TypeScript

import path from "node:path";
import { createHash } from "node:crypto";
import type { PlanRow } from "./plan-task-scoring.js";
import type { PlanSkillSource } from "./plan-task-skills.js";
export type PlanInvocation = { name: string; assetDigest: string };
/** Native context uses one immutable SKILL.md file for each company copy. */
export function planInvocations(sources: PlanSkillSource[], snapshot: PlanRow): PlanInvocation[] {
return sources.filter(s => s.selected).map(source => {
const entries = snapshot.entries?.filter((e: PlanRow) => e.key === source.key && e.desired);
const entry = entries?.length === 1 ? entries[0] : undefined;
if (!entry || !/^[a-zA-Z0-9_-]+$/.test(entry.runtimeName)) throw new Error("Missing unique assigned planning runtime name");
const manifest = [{ path: "SKILL.md", sha256: source.sha256, mode: 0o444, size: source.bytes }];
return { name: entry.runtimeName, assetDigest: createHash("sha256").update(JSON.stringify(manifest)).digest("hex") };
});
}
export function planInvocationPrompt(prompt: string, invocations: PlanInvocation[]) {
return invocations.length ? `${prompt}\n\nUse the assigned planning guidance: ${invocations.map(s => `/${s.name}`).join(" and ")}.` : prompt;
}
/** Exact initial lead run identity; provider turn acceptance joins its actual turn ID.
* turn.submitted is emitted before the provider assigns that ID, so its skill
* receipt establishes submission within this run, not a provider call-ID join. */
export function planExposure(input: { runs: PlanRow[]; evidence: PlanRow[]; leadId: string; parentId: string; invocations: PlanInvocation[] }) {
const initial = input.runs.filter(r => r.agentId === input.leadId && r.nativeIssueId === input.parentId)
.sort((a, b) => Date.parse(a.startedAt) - Date.parse(b.startedAt))[0];
const records = input.evidence.filter(r => r.runId === initial?.id);
const events: PlanRow[] = records.length === 1 && Array.isArray(records[0].events) ? records[0].events : [];
const submitted = events.filter(e => e.runId === initial?.id && e.eventType === "turn.submitted");
const prp = submitted.length === 1 ? submitted[0].payload?.prpEvent : undefined;
const actual = prp?.payload?.skillInputs ?? [];
const accepts = events.filter(e => e.runId === initial?.id && e.eventType === "turn.accepted");
const acceptedTurn = accepts.length === 1 ? accepts[0].payload?.prpEvent : undefined;
const accepted = typeof acceptedTurn?.turnId === "string" && acceptedTurn.turnId.length > 0 &&
acceptedTurn.payload?.turnId === acceptedTurn.turnId && events.some(e => e.runId === initial?.id && e.eventType === "turn.started" &&
e.payload?.prpEvent?.turnId === acceptedTurn.turnId);
const expected = input.invocations;
const matched = Array.isArray(actual) && actual.length === expected.length && expected.every(wanted => actual.filter((s: PlanRow) =>
s.type === "skill" && s.name === wanted.name && typeof s.path === "string" &&
path.isAbsolute(s.path) && s.path.endsWith(`/runtime-context-assets/bundles/${wanted.assetDigest}/SKILL.md`)).length === 1);
return { id: "assigned-guidance-delivered", passed: Boolean(initial && prp && accepted && matched),
detail: `Initial lead run must submit exactly ${expected.length} native skill inputs tied to the verified immutable file digests, and have a provider-accepted turn. This verifies the submission contract, not model consumption.`,
selectedCount: expected.length, observedInputCount: Array.isArray(actual) ? actual.length : null };
}