mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - New agents receive role instructions from onboarding, the hiring skill, or a team package. > - These sources repeat harness procedures and impose generic work policies. > - They can crowd out the task and the harness instructions. > - This pull request reduces those sources to short role descriptions. > - It preserves configuration, skills, authentication, reporting lines, and approval controls. > - The benefit is less repeated instruction text with explicit coverage for the default hiring path. ## Linked Issues or Issue Description Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection. Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes. ## What Changed - Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets. - Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples. - Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog. - Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions. - Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents. - Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse. - Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison. Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing. ## Verification - PASS: 99 focused server tests and eight shipped-catalog tests. - PASS: catalog generation and validation for four shipped teams. - PASS: hiring skill validation. - PASS: `pnpm -r typecheck`. - PASS: `pnpm build`. - INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419. - PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests). - PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog. - PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck. - PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads. - PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge. The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence. ## Risks - New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome. - Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break. - Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review. - The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback. - Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing>
171 lines
13 KiB
TypeScript
171 lines
13 KiB
TypeScript
import { createHash } from "node:crypto";
|
|
import { HIRING_TEMPLATE_READ_FILES, HIRING_TEMPLATE_SKILL_KEY } from "./hiring-template-cases.js";
|
|
import type { ChatIssue, ChatRun } from "./chat-flow.js";
|
|
|
|
const record = (value: unknown): Record<string, any> =>
|
|
value && typeof value === "object" && !Array.isArray(value) ? value as Record<string, any> : {};
|
|
export const hiringTemplateHash = (content: string) => createHash("sha256").update(content).digest("hex");
|
|
export const hiringTemplateSize = (content: string) => ({ bytes: Buffer.byteLength(content), words: content.trim().split(/\s+/).filter(Boolean).length });
|
|
export interface HiringInstructionSnapshot { entryFile: string; mode: string | null; files: Record<string, string> }
|
|
export interface HiringReadRun { runId: string; agentId: string; events: Array<{ eventType?: string; payload?: unknown; createdAt?: string }> }
|
|
export interface HiringAgent {
|
|
id: string; name: string; role?: string; reportsTo?: string | null; adapterType?: string; createdAt?: string;
|
|
adapterConfig?: Record<string, any>; runtimeConfig?: Record<string, any>;
|
|
}
|
|
export interface HiringDocument { issueId: string; key: string; body: string; latestRevisionId: string; createdByAgentId?: string }
|
|
export interface HiringTemplateEvidence {
|
|
leadId: string; chatIssueId: string; hireName: string; marker: string; projectId: string; inputs: readonly string[];
|
|
expectedCeoFiles: Record<string, string>; leadInstructions?: HiringInstructionSnapshot;
|
|
expectedSourceHashes: Record<string, string>; servedSourceHashes: Record<string, string>;
|
|
assignedSkills: string[]; coderExample: string; hiredInstructions?: HiringInstructionSnapshot;
|
|
hiredInstructionsAfterReuse?: HiringInstructionSnapshot; hiredSkills?: unknown; hiredSkillsAfterReuse?: unknown; agents: HiringAgent[];
|
|
connectionId: string; binding: unknown; tasks: ChatIssue[]; runs: ChatRun[];
|
|
first?: HiringDocument; firstAfterReuse?: HiringDocument; second?: HiringDocument;
|
|
readRuns: HiringReadRun[];
|
|
}
|
|
|
|
function expectedFixture(inputs: readonly string[], marker: string, separator: "-" | "_") {
|
|
return { reference: marker, entries: inputs.map(input => ({
|
|
input, value: input.trim().toLowerCase().replace(/[^a-z0-9]+/g, separator).replace(/^[-_]+|[-_]+$/g, ""),
|
|
})) };
|
|
}
|
|
function sameJson(left: unknown, right: unknown): boolean {
|
|
if (Array.isArray(left) && Array.isArray(right)) return left.length === right.length && left.every((v, i) => sameJson(v, right[i]));
|
|
if (left && right && typeof left === "object" && typeof right === "object") {
|
|
const keys = Object.keys(left).sort();
|
|
return sameJson(keys, Object.keys(right).sort()) && keys.every(k => sameJson(record(left)[k], record(right)[k]));
|
|
}
|
|
return left === right;
|
|
}
|
|
function documentMatches(document: HiringDocument | undefined, expected: unknown) {
|
|
try { return sameJson(JSON.parse((document?.body ?? "").trim().replace(/^```(?:json)?\s*/, "").replace(/\s*```$/, "")), expected); }
|
|
catch { return false; }
|
|
}
|
|
|
|
function payload(event: { payload?: unknown }) {
|
|
const outer = record(event.payload);
|
|
return record(outer.prpEvent).payload ? record(record(outer.prpEvent).payload) : outer;
|
|
}
|
|
function matchingReadPath(value: unknown, file: string) {
|
|
return typeof value === "string" && value.replace(/\\/g, "/").endsWith(`/paperclip-create-agent/${file}`);
|
|
}
|
|
function commandReads(command: unknown, file: string) {
|
|
if (typeof command !== "string" || /[|;&><`$\\\r\n]/.test(command)) return false;
|
|
// Admit only direct reads with explicit arguments. Unknown options, scripts,
|
|
// expansions and help/version output leave source coverage uncomparable.
|
|
const matches = [...command.matchAll(/"[^"\n]*"|'[^'\n]*'|[^\s"']+/g)];
|
|
if (!matches.length || command.slice(0, matches[0]!.index).trim()) return false;
|
|
for (let i = 1; i < matches.length; i++) {
|
|
if (!/^\s+$/.test(command.slice(matches[i - 1]!.index! + matches[i - 1]![0].length, matches[i]!.index))) return false;
|
|
}
|
|
const last = matches.at(-1)!;
|
|
if (command.slice(last.index! + last[0].length).trim()) return false;
|
|
const tokens = matches.map(match => match[0].replace(/^['"]|['"]$/g, ""));
|
|
const program = tokens.shift();
|
|
if (program === "head" || program === "tail") {
|
|
if (tokens[0] === "-n" || tokens[0] === "-c") {
|
|
tokens.shift();
|
|
const count = tokens.shift();
|
|
if (!count || !/^[1-9]\d*$/.test(count)) return false;
|
|
} else if (/^-[nc][1-9]\d*$/.test(tokens[0] ?? "")) tokens.shift();
|
|
} else if (program === "sed") {
|
|
if (tokens.shift() !== "-n") return false;
|
|
if (tokens[0] === "-e") tokens.shift();
|
|
const script = tokens.shift() ?? "";
|
|
// Only printing from the beginning is supported, without editing or shell
|
|
// execution. Other valid sed programs still require different evidence.
|
|
if (!/^(?:p|1p|1,[1-9]\d*p)$/.test(script)) return false;
|
|
} else if (program !== "cat") return false;
|
|
if (tokens[0] === "--") tokens.shift();
|
|
return tokens.length > 0 && tokens.every(token => token.length > 0 && !token.startsWith("-") && !/["'*?\[\]]/.test(token))
|
|
&& tokens.some(token => matchingReadPath(token, file));
|
|
}
|
|
|
|
/** Only completed provider read operations count; prose and discovery listings do not. */
|
|
export function hiringTemplateReadReceipts(readRuns: HiringReadRun[], leadId: string) {
|
|
const receipts: Array<{ file: string; runId: string; itemId: string; createdAt?: string }> = [];
|
|
for (const run of readRuns.filter(run => run.agentId === leadId)) {
|
|
const starts = new Map<string, Record<string, any>>();
|
|
for (const event of run.events) {
|
|
const p = payload(event), item = record(p.item);
|
|
if (event.eventType === "item.started" && item.id) starts.set(item.id, item);
|
|
if (event.eventType !== "item.completed" && event.eventType !== "tool.execution.completed") continue;
|
|
const start = starts.get(item.tool_use_id ?? item.id) ?? item;
|
|
const name = String(start.name ?? "").toLowerCase();
|
|
const input = record(start.input);
|
|
const result = record(item.result);
|
|
const toolRead = item.type === "tool_result" && /^(read|read_file|file_read)$/.test(name)
|
|
&& item.isError !== true && item.is_error !== true && result.isError !== true
|
|
&& result.is_error !== true && !result.error
|
|
&& Object.keys(result).length > 0;
|
|
const commandRead = item.type === "commandExecution" && item.exitCode === 0 && item.status === "completed"
|
|
&& typeof item.aggregatedOutput === "string" && item.aggregatedOutput.length > 0;
|
|
const executionRead = event.eventType === "tool.execution.completed" && p.status === "completed";
|
|
const canonicalFileRead = executionRead && p.transport === "builtin" && p.operation === "read" && p.readOnly === true
|
|
&& typeof p.outputBytes === "number" && p.outputBytes > 0;
|
|
for (const file of HIRING_TEMPLATE_READ_FILES) {
|
|
if ((toolRead && [input.file_path, input.path, input.filePath].some(path => matchingReadPath(path, file)))
|
|
|| (commandRead && commandReads(item.command, file))
|
|
|| (executionRead && p.exitCode === 0 && p.outputBytes > 0 && commandReads(p.name, file))
|
|
|| (canonicalFileRead && matchingReadPath(p.target, file))) {
|
|
receipts.push({ file, runId: run.runId, itemId: String(item.id ?? p.executionId ?? ""), createdAt: event.createdAt });
|
|
}
|
|
}
|
|
}
|
|
}
|
|
return receipts;
|
|
}
|
|
|
|
export function gradeHiringTemplate(e: HiringTemplateEvidence) {
|
|
const checks: Array<{ id: string; dimension: "outcome" | "coverage"; passed: boolean; detail: string }> = [];
|
|
const check = (id: string, dimension: "outcome" | "coverage", passed: boolean, detail: string) => checks.push({ id, dimension, passed, detail });
|
|
const hires = e.agents.filter(a => a.id !== e.leadId), hired = hires[0];
|
|
const lead = e.agents.find(a => a.id === e.leadId);
|
|
check("one-coder-hire", "outcome", hires.length === 1 && hired?.name === e.hireName && hired.role === "engineer"
|
|
&& hired.reportsTo === e.leadId && hired.adapterType === "paperclip_runner", "Exactly one permanent coding teammate reports to the lead.");
|
|
check("execution-account", "outcome", Boolean(e.connectionId) && Boolean(hired && lead)
|
|
&& hired?.adapterConfig?.model === lead?.adapterConfig?.model
|
|
&& sameJson(hired?.runtimeConfig?.aiConnection, e.binding), "Hire keeps the lead model and managed account binding.");
|
|
const ids = [e.first?.issueId, e.second?.issueId];
|
|
check("two-worker-tasks", "outcome", ids.every(Boolean) && new Set(ids).size === 2 && e.tasks.length === 2
|
|
&& ids.every(id => e.tasks.some(t => t.id === id && t.status === "done" && t.assigneeAgentId === hired?.id
|
|
&& t.projectId === e.projectId && !t.parentId)), "Both independent tasks are completed by the same coder in the chosen project.");
|
|
check("five-successful-turns", "outcome", e.runs.length === 5 && e.runs.every(r => r.status === "succeeded" && r.runtimeMode === "native")
|
|
&& e.runs.filter(r => r.agentId === e.leadId && r.contextSnapshot?.issueId === e.chatIssueId).length === 3
|
|
&& ids.every(id => {
|
|
const runs = e.runs.filter(r => r.contextSnapshot?.issueId === id);
|
|
return runs.length === 1 && runs[0]?.agentId === hired?.id
|
|
&& record(runs[0]?.contextSnapshot?.aiConnection).connectionId === e.connectionId;
|
|
}), "Five turns include one actual coder execution for each task, attributed to the managed account.");
|
|
check("initial-json-artifact", "outcome", documentMatches(e.first, expectedFixture(e.inputs, e.marker, "-"))
|
|
&& e.first?.createdByAgentId === hired?.id, "Independent computation checks each original input/value and worker authorship.");
|
|
check("reused-json-artifact", "outcome", documentMatches(e.second, expectedFixture(e.inputs, `REUSE${e.marker}`, "_"))
|
|
&& e.second?.createdByAgentId === hired?.id, "The reused coder applies the changed separator to every input.");
|
|
check("original-preserved", "outcome", Boolean(e.first) && sameJson(e.first, e.firstAfterReuse), "Reuse preserves the original document, revision and task identity.");
|
|
const requiredSources = [...HIRING_TEMPLATE_READ_FILES.map(file => `skills/paperclip-create-agent/${file}`),
|
|
"skills/paperclip-create-agent/references/baseline-role-guide.md",
|
|
...Object.keys(e.expectedCeoFiles).map(file => `server/src/onboarding-assets/ceo/${file}`)];
|
|
check("source-fingerprints", "coverage", requiredSources.every(file => /^[a-f0-9]{64}$/.test(e.expectedSourceHashes[file] ?? ""))
|
|
&& Object.entries(e.expectedSourceHashes).every(([file, hash]) => e.servedSourceHashes[file] === hash), "Served instruction and hiring source bytes match the evaluated revision.");
|
|
check("production-ceo-bundle", "coverage", e.leadInstructions?.mode === "managed" && e.leadInstructions.entryFile === "AGENTS.md"
|
|
&& Object.keys(e.expectedCeoFiles).length > 0 && sameJson(e.leadInstructions.files, e.expectedCeoFiles), "The API-created lead receives this revision's default CEO bundle, without a fixture override.");
|
|
check("assigned-hiring-skill", "coverage", e.assignedSkills.includes(HIRING_TEMPLATE_SKILL_KEY), "Production CEO defaults assign the hiring skill.");
|
|
const receipts = hiringTemplateReadReceipts(e.readRuns, e.leadId);
|
|
check("production-source-reads", "coverage", HIRING_TEMPLATE_READ_FILES.every(file => receipts.some(receipt => receipt.file === file && receipt.itemId
|
|
&& Date.parse(receipt.createdAt ?? "") <= Date.parse(hired?.createdAt ?? ""))), "Completed lead read receipts before the hire prove the explicitly requested skill, guide, checklist and coder example paths. Unrecognized or missing reads leave coverage uncomparable.");
|
|
const instruction = e.hiredInstructions?.files["AGENTS.md"] ?? "";
|
|
// The same fixture can compare a long historical role and a short candidate:
|
|
// it requires the supplied role example, never a fixed size or new wording.
|
|
const expectedCoder = e.coderExample.trim();
|
|
check("supplied-coder-instructions", "coverage", expectedCoder.length > 0 && e.hiredInstructions?.mode === "managed"
|
|
&& e.hiredInstructions.entryFile === "AGENTS.md" && instruction.trim() === expectedCoder, "Saved hired instructions use the source revision's coder example with its company/name placeholders filled.");
|
|
check("hired-instructions-durable", "coverage", Boolean(e.hiredInstructions) && sameJson(e.hiredInstructions, e.hiredInstructionsAfterReuse), "The same saved instruction bundle survives the reused worker execution.");
|
|
check("hired-skills-durable", "coverage", Boolean(e.hiredSkills) && sameJson(e.hiredSkills, e.hiredSkillsAfterReuse), "The saved skill selections survive the reused worker execution.");
|
|
return {
|
|
checks, readReceipts: receipts,
|
|
outcomePassed: checks.filter(c => c.dimension === "outcome").every(c => c.passed),
|
|
comparisonStatus: checks.filter(c => c.dimension === "coverage").every(c => c.passed) ? "comparable" : "uncomparable",
|
|
instructionSizes: { ceo: Object.fromEntries(Object.entries(e.leadInstructions?.files ?? {}).map(([file, content]) => [file, hiringTemplateSize(content)])), coder: hiringTemplateSize(instruction) },
|
|
};
|
|
}
|