Files
PaperClipAI/tests/runner-e2e/hiring-template-cases.ts
DottaandPaperclip 862a5758ba fix(agents): reduce hiring templates to role descriptions (#14985)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for work.
> - New agents receive role instructions from onboarding, the hiring skill, or a team package.
> - These sources repeat harness procedures and impose generic work policies.
> - They can crowd out the task and the harness instructions.
> - This pull request reduces those sources to short role descriptions.
> - It preserves configuration, skills, authentication, reporting lines, and approval controls.
> - The benefit is less repeated instruction text with explicit coverage for the default hiring path.

## Linked Issues or Issue Description

Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection.

Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes.

## What Changed

- Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets.
- Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples.
- Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog.
- Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions.
- Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents.
- Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse.
- Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison.

Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing.

## Verification

- PASS: 99 focused server tests and eight shipped-catalog tests.
- PASS: catalog generation and validation for four shipped teams.
- PASS: hiring skill validation.
- PASS: `pnpm -r typecheck`.
- PASS: `pnpm build`.
- INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419.
- PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests).
- PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog.
- PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck.
- PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads.
- PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge.

The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence.

## Risks

- New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome.
- Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break.
- Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review.
- The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback.
- Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-02 17:56:42 -05:00

60 lines
4.3 KiB
TypeScript

import { createHash } from "node:crypto";
import { readFileSync } from "node:fs";
import type { RunnerProfileFixture, RunnerTaskFixture } from "./types.js";
export const HIRING_TEMPLATE_GRADER_VERSION = "paperclip.hiring-templates.v1";
export const HIRING_TEMPLATE_SKILL_KEY = "paperclipai/paperclip/paperclip-create-agent";
export const HIRING_TEMPLATE_READ_FILES = [
"SKILL.md",
"references/agent-instruction-templates.md",
"references/draft-review-checklist.md",
"references/agents/coder.md",
] as const;
export const HIRING_TEMPLATE_SOURCE_FILES = [
"server/src/services/default-agent-instructions.ts",
"server/src/onboarding-assets/default/AGENTS.md",
...HIRING_TEMPLATE_READ_FILES.map(file => `skills/paperclip-create-agent/${file}`),
"skills/paperclip-create-agent/references/baseline-role-guide.md",
] as const;
export const hiringTemplateDefinitionDigest = createHash("sha256").update(
["cases", "scoring", "flow"].map(part => readFileSync(new URL(`./hiring-template-${part}.ts`, import.meta.url), "utf8")).join("\n"),
).digest("hex");
export function hiringTemplateProfile(profile: RunnerProfileFixture): RunnerProfileFixture {
return {
...profile,
buildAgent(input) {
const { instructionsBundle: _fixtureInstructions, ...payload } = profile.buildAgent(input);
return { ...payload, name: `Hiring Lead ${input.executionId}`, role: "ceo", title: "CEO" };
},
};
}
export const hiringTemplateInputs = [" Launch Queue!! ", "API__Key Rotation", "Mixed CASE / 42", "---Already--Ready---"] as const;
export function hiringTemplateScenario(nonce: string, projectName: string) {
const marker = `HIRE${nonce.replace(/[^a-zA-Z0-9]/g, "")}`;
const hireName = `Casey Coder ${nonce}`;
return {
marker, hireName, inputs: hiringTemplateInputs,
initialPrompt: [
`Use the paperclip-create-agent skill to hire exactly one coding teammate named ${hireName}, reporting to you, with your native runner, model, and available AI connection.`,
"Before hiring, read that skill's SKILL.md, references/agent-instruction-templates.md, references/draft-review-checklist.md, and references/agents/coder.md. Use the AGENTS.md example from that coder reference as the teammate's managed instructions, filling all agentName, companyName, managerTitle, and issuePrefix placeholders from this hire, company, your actual title, and the company issue prefix. Keep skills and reporting information in their configuration.",
`Have that coder prepare a JSON test fixture for a label normalizer in one assigned task in ${projectName}. Save the fixture as a Paperclip document with an object containing reference: ${JSON.stringify(marker)} and entries, an ordered array of objects with input and value. Inputs: ${JSON.stringify(hiringTemplateInputs)}.`,
"For each value, trim its input, lowercase it, replace each sequence outside ASCII a-z and 0-9 with one hyphen, then trim leading and trailing hyphens. Link the task here and let the coder complete it.",
].join("\n\n"),
reusePrompt: (identifier: string) => `Have the existing ${hireName} prepare a second saved JSON fixture in one new task in ${projectName}, assigned to that same coder. Use the inputs and normalization rule from ${identifier}, but use underscores instead of hyphens for the value separators and trim leading and trailing underscores. Use reference: ${JSON.stringify(`REUSE${marker}`)}. Include the original fixture in the handoff. Preserve the original document and task, and let the coder complete the new task.`,
statusPrompt: (first: string, second: string) => `Briefly report the recorded owner and status of ${first} and ${second}. Just report; do not create or change work.`,
};
}
export const hiringTemplateTasks: readonly RunnerTaskFixture[] = [{
id: "hire-coder-template-reuse", label: "Hire with production templates and reuse the coder",
groups: ["chat"], flow: "agent_chat", workMode: "standard", expectedRunCount: 5,
attemptTimeoutMs: { local: 15 * 60_000, daytona: 15 * 60_000 },
expectedTerminalState: { issue: "in_review", run: "succeeded" },
buildTitle: nonce => `Production hiring templates ${nonce}`,
buildPrompt: nonce => hiringTemplateScenario(nonce, "the selected project").initialPrompt,
buildVisibleMarker: nonce => hiringTemplateScenario(nonce, "").marker,
buildMatchers: () => [{ kind: "issue_status", expected: "in_review" }],
}];