Files
PaperClipAI/doc/plans/2026-10-02-hiring-skill-drafting-evidence.json
T
DottaandPaperclip 862a5758ba fix(agents): reduce hiring templates to role descriptions (#14985)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for work.
> - New agents receive role instructions from onboarding, the hiring skill, or a team package.
> - These sources repeat harness procedures and impose generic work policies.
> - They can crowd out the task and the harness instructions.
> - This pull request reduces those sources to short role descriptions.
> - It preserves configuration, skills, authentication, reporting lines, and approval controls.
> - The benefit is less repeated instruction text with explicit coverage for the default hiring path.

## Linked Issues or Issue Description

Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection.

Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes.

## What Changed

- Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets.
- Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples.
- Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog.
- Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions.
- Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents.
- Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse.
- Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison.

Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing.

## Verification

- PASS: 99 focused server tests and eight shipped-catalog tests.
- PASS: catalog generation and validation for four shipped teams.
- PASS: hiring skill validation.
- PASS: `pnpm -r typecheck`.
- PASS: `pnpm build`.
- INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419.
- PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests).
- PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog.
- PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck.
- PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads.
- PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge.

The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence.

## Risks

- New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome.
- Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break.
- Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review.
- The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback.
- Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-02 17:56:42 -05:00

80 lines
3.3 KiB
JSON

{
"kind": "independent-drafting-simulation",
"date": "2026-10-02",
"scope": "Three synthetic hire requests; read-only. No API calls, permanent hires, provider campaign, or live harness qualification.",
"skillSources": [
{
"path": "skills/paperclip-create-agent/SKILL.md",
"sha256": "00e3063a40bbe32592171fa613956bc1b503fe0a41c4973c8dae7b8cdf8c148a"
},
{
"path": "skills/paperclip-create-agent/references/agent-instruction-templates.md",
"sha256": "9ce87fc420b41e951cab664655886d0ee50fa253cb5b8647b4a7afa4a6082d2d"
},
{
"path": "skills/paperclip-create-agent/references/baseline-role-guide.md",
"sha256": "0841c5c483526d8d64c7217df7164bca7f3f03edcf4578f2baefcca01cd156b0"
},
{
"path": "skills/paperclip-create-agent/references/draft-review-checklist.md",
"sha256": "1cbe0577a9888f31467f5c7beea01d0aa6435da2a0f577b5fdd3b5fe773984ad"
},
{
"path": "skills/paperclip-create-agent/references/agents/coder.md",
"sha256": "e04a25fbc8bb29910b2f4a3719730b8840553b0524ce1ec570617701aca20650"
},
{
"path": "skills/paperclip-create-agent/references/agents/securityengineer.md",
"sha256": "9eecf981844ab1f5896dd1bae6b76acce8471c448841608c5655497d691ba1a0"
},
{
"path": "skills/paperclip-create-agent/references/api-reference.md",
"sha256": "f43bd2f74d60a9f019af1c037584f2cb4bb2b1f66f503a60eca80278676bded1"
}
],
"observed": [
{
"request": "Backend engineer Ada; legacy Claude; repo-maintenance skill; CTO reporting line",
"instructions": "You are agent Ada, a backend software engineer at Northstar. You own implementation of assigned backend service changes.",
"desiredSkills": [
"northstar/repo-maintenance"
],
"adapterType": "claude_local",
"model": "claude-sonnet-4-6",
"timerEnabled": false,
"literalCredentials": false
},
{
"request": "Release Coordinator Riley; native Codex; no schedule",
"instructions": "You are agent Riley, a Release Coordinator at Northstar. You own coordination of release readiness and publication requests.",
"desiredSkills": [],
"adapterType": "paperclip_runner",
"provider": "codex",
"model": "gpt-5.4",
"timerEnabled": false,
"literalCredentials": false
},
{
"request": "Security Engineer Sam; legacy Claude; exact user instruction; confidential workflow installed",
"instructions": "Review only the authentication service. Keep private findings in the incident-review workflow. Do not edit code.",
"requestedInstructionsPreservedExactly": true,
"desiredSkills": [
"northstar/incident-review"
],
"adapterType": "claude_local",
"model": "claude-sonnet-4-6",
"timerEnabled": false,
"literalCredentials": false
}
],
"remainingBeforeMutation": [
"Confirm live endpoint and adapter schemas.",
"Submit through authorized hire API and reconcile returned pending approval; no hire may be claimed from this simulation."
],
"limits": [
"Outcome projection of independent subagent drafting, not a live full-stack measurement.",
"No baseline drafting simulation, timing, billing, or downstream work comparison.",
"Does not establish non-regression for removed workflow or domain policies."
]
}