mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - New agents receive role instructions from onboarding, the hiring skill, or a team package. > - These sources repeat harness procedures and impose generic work policies. > - They can crowd out the task and the harness instructions. > - This pull request reduces those sources to short role descriptions. > - It preserves configuration, skills, authentication, reporting lines, and approval controls. > - The benefit is less repeated instruction text with explicit coverage for the default hiring path. ## Linked Issues or Issue Description Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection. Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes. ## What Changed - Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets. - Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples. - Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog. - Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions. - Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents. - Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse. - Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison. Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing. ## Verification - PASS: 99 focused server tests and eight shipped-catalog tests. - PASS: catalog generation and validation for four shipped teams. - PASS: hiring skill validation. - PASS: `pnpm -r typecheck`. - PASS: `pnpm build`. - INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419. - PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests). - PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog. - PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck. - PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads. - PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge. The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence. ## Risks - New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome. - Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break. - Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review. - The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback. - Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing>
144 lines
11 KiB
TypeScript
144 lines
11 KiB
TypeScript
import { readFile } from "node:fs/promises";
|
|
import { loadDefaultAgentInstructionsBundle } from "../../server/src/services/default-agent-instructions.js";
|
|
import { readChatOutputDocument, type ChatFlowInput, type ChatIssue, type ChatRun } from "./chat-flow.js";
|
|
import { collectRunEvents } from "./run-observations.js";
|
|
import { HIRING_TEMPLATE_GRADER_VERSION, HIRING_TEMPLATE_READ_FILES, HIRING_TEMPLATE_SKILL_KEY,
|
|
HIRING_TEMPLATE_SOURCE_FILES, hiringTemplateDefinitionDigest, hiringTemplateScenario } from "./hiring-template-cases.js";
|
|
import { gradeHiringTemplate, hiringTemplateHash, hiringTemplateSize, type HiringAgent, type HiringDocument,
|
|
type HiringInstructionSnapshot, type HiringTemplateEvidence } from "./hiring-template-scoring.js";
|
|
import type { RunnerApi } from "./api.js";
|
|
|
|
export async function readHiringInstructions(api: Pick<RunnerApi, "get">, agentId: string): Promise<HiringInstructionSnapshot> {
|
|
const bundle = await api.get<{ mode: string | null; entryFile: string; files: Array<{ path: string; binary?: boolean }> }>(`/api/agents/${agentId}/instructions-bundle`);
|
|
const files: Record<string, string> = {};
|
|
// Agent-home notes may be added by a run. Retain the entry and legacy CEO
|
|
// policy files so the baseline and candidate default bundles can be compared.
|
|
for (const file of bundle.files.filter(file => !file.binary && (file.path === bundle.entryFile || ["HEARTBEAT.md", "SOUL.md", "TOOLS.md"].includes(file.path)))) {
|
|
const detail = await api.get<{ content: string }>(`/api/agents/${agentId}/instructions-bundle/file?path=${encodeURIComponent(file.path)}`);
|
|
files[file.path] = detail.content;
|
|
}
|
|
return { entryFile: bundle.entryFile, mode: bundle.mode, files };
|
|
}
|
|
|
|
async function readHiringSkillSelections(api: Pick<RunnerApi, "get">, companyId: string, agentId: string) {
|
|
const snapshot = await api.get<{ desiredSkills: string[]; desiredSkillEntries: unknown[] }>(`/api/agents/${agentId}/skills?companyId=${companyId}`);
|
|
return { desiredSkills: snapshot.desiredSkills, desiredSkillEntries: snapshot.desiredSkillEntries };
|
|
}
|
|
|
|
export async function readHiringTemplateSources(api: Pick<RunnerApi, "get">, companyId: string, leadId: string) {
|
|
const expectedCeoFiles = await loadDefaultAgentInstructionsBundle("ceo");
|
|
const expectedSourceHashes: Record<string, string> = {}, servedSourceHashes: Record<string, string> = {};
|
|
const sourceSizes: Record<string, ReturnType<typeof hiringTemplateSize>> = {};
|
|
for (const relative of [...HIRING_TEMPLATE_SOURCE_FILES, ...Object.keys(expectedCeoFiles).map(file => `server/src/onboarding-assets/ceo/${file}`)]) {
|
|
const content = await readFile(new URL(`../../${relative}`, import.meta.url), "utf8");
|
|
expectedSourceHashes[relative] = hiringTemplateHash(content);
|
|
sourceSizes[relative] = hiringTemplateSize(content);
|
|
}
|
|
const leadInstructions = await readHiringInstructions(api, leadId);
|
|
for (const [file, content] of Object.entries(leadInstructions.files)) servedSourceHashes[`server/src/onboarding-assets/ceo/${file}`] = hiringTemplateHash(content);
|
|
const skills = await api.get<Array<{ id: string; key: string }>>(`/api/companies/${companyId}/skills`);
|
|
const creator = skills.find(skill => skill.key === HIRING_TEMPLATE_SKILL_KEY);
|
|
if (!creator) throw new Error("Production hiring skill is absent from the company library");
|
|
const assigned = await api.get<{ desiredSkills: string[] }>(`/api/agents/${leadId}/skills?companyId=${companyId}`);
|
|
let coderReference = "";
|
|
for (const file of [...HIRING_TEMPLATE_READ_FILES, "references/baseline-role-guide.md"]) {
|
|
const detail = await api.get<{ content: string }>(`/api/companies/${companyId}/skills/${creator.id}/files?path=${encodeURIComponent(file)}`);
|
|
servedSourceHashes[`skills/paperclip-create-agent/${file}`] = hiringTemplateHash(detail.content);
|
|
if (file === "references/agents/coder.md") coderReference = detail.content;
|
|
}
|
|
// The loader and execution contract are source provenance, not served skill
|
|
// reads. Their bytes are retained separately from the equality checks.
|
|
delete expectedSourceHashes["server/src/services/default-agent-instructions.ts"];
|
|
delete expectedSourceHashes["server/src/onboarding-assets/default/AGENTS.md"];
|
|
return { expectedCeoFiles, leadInstructions, expectedSourceHashes, servedSourceHashes,
|
|
sourceSizes, assignedSkills: assigned.desiredSkills, coderReference, creatorSkillId: creator.id,
|
|
loaderHash: hiringTemplateHash(await readFile(new URL("../../server/src/services/default-agent-instructions.ts", import.meta.url), "utf8")),
|
|
executionContractHash: hiringTemplateHash(await readFile(new URL("../../server/src/onboarding-assets/default/AGENTS.md", import.meta.url), "utf8")) };
|
|
}
|
|
|
|
export function renderHiringCoderExample(reference: string, agentName: string, companyName: string, managerTitle: string, issuePrefix: string) {
|
|
const example = reference.match(/```md\s*\n([\s\S]*?)\n```/)?.[1];
|
|
if (!example?.trim()) throw new Error("Production coder reference has no AGENTS.md example");
|
|
return example.replaceAll("{{agentName}}", agentName).replaceAll("{{companyName}}", companyName).replaceAll("{{managerTitle}}", managerTitle).replaceAll("{{issuePrefix}}", issuePrefix);
|
|
}
|
|
|
|
export async function runHiringTemplateFlow(context: {
|
|
input: ChatFlowInput; issue(): ChatIssue; turn(message: string, count: number): Promise<void>;
|
|
tasks(): Promise<ChatIssue[]>; allRuns(): Promise<ChatRun[]>;
|
|
}) {
|
|
const { input, turn, tasks, allRuns } = context;
|
|
const { api, fixtures: f } = input;
|
|
const company = `/api/companies/${f.company.id}`;
|
|
const account = f.aiConnection;
|
|
if (!account) throw new Error("Hiring-template fixture requires a managed execution account");
|
|
const issuePrefix = f.company.issuePrefix;
|
|
if (!issuePrefix) throw new Error("Hiring-template fixture requires the actual company issue prefix");
|
|
const evidence: HiringTemplateEvidence = {
|
|
leadId: f.agent.id, chatIssueId: "", hireName: "", marker: "", projectId: "", inputs: [], expectedCeoFiles: {},
|
|
expectedSourceHashes: {}, servedSourceHashes: {}, assignedSkills: [], coderExample: "",
|
|
agents: [], tasks: [], runs: [], readRuns: [], connectionId: account.connectionId, binding: account.binding,
|
|
};
|
|
let source: Awaited<ReturnType<typeof readHiringTemplateSources>> | undefined;
|
|
let scenario: ReturnType<typeof hiringTemplateScenario> | undefined;
|
|
async function refresh() {
|
|
const observed = await Promise.all([api.get<HiringAgent[]>(`${company}/agents`), tasks(), allRuns()]);
|
|
[evidence.agents, evidence.tasks, evidence.runs] = observed;
|
|
evidence.readRuns = await Promise.all(evidence.runs.map(async run => {
|
|
const events = await collectRunEvents<{ seq?: number; eventType?: string; payload?: unknown; createdAt?: string }>(
|
|
(afterSeq, limit) => api.get(`/api/heartbeat-runs/${run.id}/events?afterSeq=${afterSeq}&limit=${limit}`),
|
|
);
|
|
return { runId: run.id, agentId: run.agentId, events };
|
|
}));
|
|
}
|
|
try {
|
|
await api.patch(`${company}/budgets`, { budgetMonthlyCents: 1_000 });
|
|
await api.patch(`/api/agents/${f.agent.id}/budgets`, { budgetMonthlyCents: 1_000 });
|
|
source = await readHiringTemplateSources(api, f.company.id, f.agent.id);
|
|
Object.assign(evidence, source);
|
|
await input.evidence("hiring-template-source.json", { ...source, definitionDigest: hiringTemplateDefinitionDigest });
|
|
if (Object.entries(source.expectedSourceHashes).some(([file, hash]) => source!.servedSourceHashes[file] !== hash)) {
|
|
throw new Error("Hiring-template served source differs from the evaluated revision; comparison is uncomparable");
|
|
}
|
|
const project = await api.post<{ id: string; name: string }>(`${company}/projects`, {
|
|
name: `Label fixtures ${input.nonce}`, description: "Repository-free JSON normalization fixtures.",
|
|
});
|
|
scenario = hiringTemplateScenario(input.nonce, project.name);
|
|
Object.assign(evidence, { hireName: scenario.hireName, marker: scenario.marker, inputs: scenario.inputs, projectId: project.id,
|
|
coderExample: renderHiringCoderExample(source.coderReference, scenario.hireName, f.company.name, String((await api.get<{ title: string }>(`/api/agents/${f.agent.id}`)).title), issuePrefix) });
|
|
await turn(scenario.initialPrompt, 2);
|
|
evidence.chatIssueId = context.issue().id;
|
|
const initial = await tasks();
|
|
if (initial.length !== 1) throw new Error("Hiring-template first turn must create exactly one coder task");
|
|
evidence.first = await readChatOutputDocument(api, initial[0]!.id, scenario.marker) as HiringDocument;
|
|
await refresh();
|
|
const hired = evidence.agents.find(agent => agent.name === scenario!.hireName);
|
|
if (!hired) throw new Error("Hiring-template coder was not created");
|
|
evidence.hiredInstructions = await readHiringInstructions(api, hired.id);
|
|
evidence.hiredSkills = await readHiringSkillSelections(api, f.company.id, hired.id);
|
|
await input.capture("hiring-template-created", "Production CEO hired a coder and delivered its first fixture", "hiring-template-created.png");
|
|
await input.evidence("hiring-template-initial.json", evidence);
|
|
await turn(scenario.reusePrompt(initial[0]!.identifier ?? initial[0]!.id), 4);
|
|
const second = (await tasks()).find(task => task.id !== evidence.first!.issueId);
|
|
if (!second) throw new Error("Hiring-template reuse task is absent");
|
|
evidence.second = await readChatOutputDocument(api, second.id, `REUSE${scenario.marker}`) as HiringDocument;
|
|
evidence.firstAfterReuse = await api.get<HiringDocument>(`/api/issues/${evidence.first.issueId}/documents/${encodeURIComponent(evidence.first.key)}`);
|
|
await turn(scenario.statusPrompt(initial[0]!.identifier ?? initial[0]!.id, second.identifier ?? second.id), 5);
|
|
evidence.hiredInstructionsAfterReuse = await readHiringInstructions(api, hired.id);
|
|
evidence.hiredSkillsAfterReuse = await readHiringSkillSelections(api, f.company.id, hired.id);
|
|
await refresh();
|
|
const result = gradeHiringTemplate(evidence);
|
|
for (const check of result.checks) input.check?.(`hiringTemplates.${check.dimension}.${check.id}`, check.passed, check.detail);
|
|
await input.evidence("hiring-template.json", { schema: HIRING_TEMPLATE_GRADER_VERSION, definitionDigest: hiringTemplateDefinitionDigest,
|
|
budgetGuard: { companyMonthlyCents: 1_000, leadMonthlyCents: 1_000 }, scenario, source, evidence, result });
|
|
const failed = result.checks.filter(check => !check.passed);
|
|
if (!result.outcomePassed) throw new Error(`Hiring-template workflow outcome failed: ${failed.filter(check => check.dimension === "outcome").map(check => check.id).join(", ")}`);
|
|
if (result.comparisonStatus === "uncomparable") throw new Error(`Hiring-template source coverage is uncomparable: ${failed.filter(check => check.dimension === "coverage").map(check => check.id).join(", ")}`);
|
|
} finally {
|
|
let observationError: string | undefined;
|
|
try { await refresh(); } catch (error) { observationError = String(error); }
|
|
await input.evidence("hiring-template.json", { schema: HIRING_TEMPLATE_GRADER_VERSION, definitionDigest: hiringTemplateDefinitionDigest,
|
|
budgetGuard: { companyMonthlyCents: 1_000, leadMonthlyCents: 1_000 },
|
|
scenario, source, evidence, result: gradeHiringTemplate(evidence), observationError });
|
|
}
|
|
}
|