mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - New agents receive role instructions from onboarding, the hiring skill, or a team package. > - These sources repeat harness procedures and impose generic work policies. > - They can crowd out the task and the harness instructions. > - This pull request reduces those sources to short role descriptions. > - It preserves configuration, skills, authentication, reporting lines, and approval controls. > - The benefit is less repeated instruction text with explicit coverage for the default hiring path. ## Linked Issues or Issue Description Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection. Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes. ## What Changed - Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets. - Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples. - Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog. - Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions. - Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents. - Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse. - Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison. Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing. ## Verification - PASS: 99 focused server tests and eight shipped-catalog tests. - PASS: catalog generation and validation for four shipped teams. - PASS: hiring skill validation. - PASS: `pnpm -r typecheck`. - PASS: `pnpm build`. - INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419. - PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests). - PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog. - PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck. - PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads. - PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge. The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence. ## Risks - New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome. - Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break. - Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review. - The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback. - Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing>
78 lines
4.4 KiB
TypeScript
78 lines
4.4 KiB
TypeScript
import type { RunnerTaskFixture } from "./types.js";
|
|
|
|
// Markers cross the rich-text composer and Markdown persistence boundary.
|
|
// Alphanumeric text has identical visible and stored representations.
|
|
export function chatMarker(
|
|
prefix: "CHAT" | "DRAFT" | "OLDCONTEXT",
|
|
nonce: string,
|
|
) {
|
|
return `${prefix}${nonce.replace(/[^a-zA-Z0-9]/g, "")}`;
|
|
}
|
|
|
|
export const CHAT_CASES = [
|
|
["create-backlog", "Save a planned task without starting work", 2],
|
|
["reassign-task", "Reassign existing work and preserve queued context", 2],
|
|
["continuity-restart", "Conversation continuity across restart", 3],
|
|
["new-session", "Fresh context within preserved history", 2],
|
|
["stop-new-resume", "Stop, reset, and resume", 3],
|
|
["plan-handoff", "Draft, revise, approve, and hand off a plan", 4],
|
|
["clarify-reuse", "Clarify and reuse an existing project", 3],
|
|
["multi-repository", "Create a project with multiple repository URLs", 2],
|
|
] as const;
|
|
export type ChatCase = (typeof CHAT_CASES)[number][0];
|
|
const HARDENING_CASES = [
|
|
["stop-startup-new-resume", "Stop during startup, reset, and resume", 3],
|
|
["hire-delegate-reuse", "Hire through chat, delegate, and reuse the same teammate", 5],
|
|
["blocked-status-review", "Read the actual blocker and hand source material to a reviewer", 4],
|
|
["committed-send-retry", "Recover a lost send acknowledgement without repeating committed work", 2],
|
|
...CHAT_CASES.filter(([id]) => ["stop-new-resume", "continuity-restart"].includes(id)),
|
|
] as const;
|
|
function buildChatTasks(definitions: readonly (readonly [string, string, number])[]): readonly RunnerTaskFixture[] {
|
|
return definitions.map(
|
|
([id, label, expectedRunCount]) => ({
|
|
id,
|
|
label,
|
|
groups: ["chat"],
|
|
flow: "agent_chat",
|
|
workMode: "standard",
|
|
expectedRunCount, // Provider turns, including cancelled turns and handed-off work; reset runs are separate.
|
|
attemptTimeoutMs: { local: 15 * 60_000, daytona: 15 * 60_000 },
|
|
expectedTerminalState: { issue: "in_review", run: "succeeded" },
|
|
buildTitle: (nonce) => `Chat acceptance ${id} ${nonce}`,
|
|
buildPrompt: (nonce) => `Let's discuss ${nonce}.`,
|
|
buildVisibleMarker: (nonce) => chatMarker("CHAT", nonce),
|
|
buildMatchers: () => [{ kind: "issue_status", expected: "in_review" }],
|
|
}),
|
|
);
|
|
}
|
|
export const chatTasks = buildChatTasks(CHAT_CASES);
|
|
export const chatHardeningTasks = buildChatTasks(HARDENING_CASES);
|
|
export const chatStoryTasks = buildChatTasks([
|
|
["enable-disable-resume", "Enable Agent Chat, pause access, and resume preserved history", 2],
|
|
["followup-while-running", "Deliver a follow-up while a provider turn is running", 2],
|
|
["revise-while-running", "Change instructions during active work and save the updated plan", 2],
|
|
]).map(task => ({ ...task, ...(task.id === "enable-disable-resume" ? {} : { minimumExpectedRunCount: 1 }) }));
|
|
|
|
export function chatNeedsApiTools(suiteId: string, caseId: string): boolean {
|
|
return suiteId === "hiring-templates" || (suiteId === "agent-chat-qualification" && caseId === "grounded-answer-quality") || suiteId === "agent-chat-hardening" && ["hire-delegate-reuse", "blocked-status-review"].includes(caseId);
|
|
}
|
|
export function isManagedHiringCase(suiteId: string, caseId: string): boolean {
|
|
return (suiteId === "hiring-templates" && caseId === "hire-coder-template-reuse") || (suiteId === "everyday-workflows" && caseId === "hire-reuse") ||
|
|
(suiteId === "agent-chat-hardening" && caseId === "hire-delegate-reuse");
|
|
}
|
|
|
|
export const chatQualificationTasks = buildChatTasks([
|
|
["active-reassignment", "Reassign an executing task and preserve its saved work", 3],
|
|
["worker-crash-retry", "Recover from worker process loss through visible Retry", 2],
|
|
["grounded-answer-quality", "Ground status, correct stale claims, and acknowledge uncertainty", 2],
|
|
]);
|
|
|
|
export const chatCompletionTasks = buildChatTasks([
|
|
["handoff-completion-idle", "Report a delegated result after the chat goes idle", 4],
|
|
["handoff-completion-busy", "Queue a delegated result behind an active chat reply", 5],
|
|
["handoff-completion-multiple", "Report multiple delegated results as they finish", 7],
|
|
["handoff-completion-restart", "Recover pending completion delivery across a server restart", 5],
|
|
]).map(task => ({ ...task, minimumExpectedRunCount: 2 }));
|
|
|
|
export const chatConfirmationTasks = buildChatTasks([["confirmation-ambiguous", "Clarify ambiguous approval, then resolve only the chosen cards", 4], ["unanswered-question-return", "Move on, reopen a historical question, and deliver the late answer", 3]]);
|