Files
DottaandPaperclip 2f0c485dec fix(skills): ship the completion helper with the installed skill (#15554)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Legacy agents use the Paperclip skill to save task status and
comments.
> - The skill names a script relative to the task workspace.
> - That script exists only in the Paperclip source repository.
> - Agents in other workspaces can hit a missing command or search for
it.
> - This PR ships the helper inside the skill and uses the installed
skill path.
> - The repository command remains available through a forwarding
wrapper.

## Linked Issues or Issue Description

Fixes #9527. Refs #15548 for the preceding runtime checkout guidance.
Related: #6052 addresses LF line endings for the repository helper; this
change addresses helper delivery and path resolution.

## What Changed

- Bundle the existing issue update helper with the Paperclip skill.
Preserve its HTTP checks, echoed-status check and two-attempt limit.
- Resolve the command from the installed skill directory. Use a verified
PATCH when that path is unavailable, without searching the filesystem.
- Keep the repository command as a wrapper that works from any
directory.
- Test shell execution and exact status/comment payloads through both
provider skill-home layouts, including paths with spaces.
- Add helper sources and existing verification tests to stock-harness
admission. Record an absent historical helper explicitly. Add the
missing declaration for the admission fingerprint export.

## Verification

- Complete directly affected source suites: 30 tests pass. They cover
skill delivery, preserved multiline comments and links, authentication
headers, empty responses, mismatched status, transient retries and
definitive rejections.
- Product E2E typecheck passes. Support suites: 1,835 Vitest tests pass,
one is skipped; 128 Node tests pass.
- Full local build and workspace typecheck pass.
- Full local repository tests are not claimed as passed. Embedded
PostgreSQL was unavailable in this worktree during the preceding task;
Linux CI will run the repository gates.
- The authorized matched Codex/Claude comparison is pending. It uses the
existing assigned-skill case and original oracle, one initial attempt
per profile and variant.
- CI and a fresh Greptile review are pending. Keep this PR in draft
until readiness gates complete.

## Risks

- Correct path resolution depends on the harness supplying the installed
skill path. The instructions use verified PATCH when that path is
unavailable.
- The helper still requires Bash, curl and jq. Its existing retry and
response-verification behavior is unchanged.
- Tests use the shared skill-directory symlink mechanism and an HTTP
fixture. Real provider completion behavior still requires the bounded
live comparison.
- This fix does not redesign native completion, legacy recovery or
ambiguous transport handling.

## Model Used

OpenAI Codex, GPT-6 family. The exact model build and context window are
not exposed in this session. Used code editing, shell tools and test
execution.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 07:21:30 -05:00

187 lines
11 KiB
TypeScript

import { createHash } from "node:crypto";
import { readFileSync } from "node:fs";
import type { RunnerProfileFixture } from "./types.js";
import { readStockInstructionVariant } from "./stock-harness-instruction-variant.mjs";
// These are an independent product contract, not imports of the implementation
// constants. Otherwise a larger shipped manual could silently update the oracle.
export const STOCK_HIRE_IDENTITY = "You are an agent in a Paperclip company.\n";
export const STOCK_TEMPLATE_IDENTITY = "You are agent";
const SKILL_SOURCES = ["../../skills/paperclip/SKILL.md", "../../skills/paperclip/references/issue-documents.md",
"../../scripts/paperclip-issue-update.sh", "../../skills/paperclip/scripts/paperclip-issue-update.sh"];
export function stockHarnessSkillSources() {
return SKILL_SOURCES.map(source => {
try {
return { path: source.replace(/^\.\.\/\.\.\//, ""), present: true,
sha256: createHash("sha256").update(readFileSync(new URL(source, import.meta.url))).digest("hex") };
} catch (error) {
if ((source.endsWith("/references/issue-documents.md") || source === "../../skills/paperclip/scripts/paperclip-issue-update.sh")
&& (error as NodeJS.ErrnoException).code === "ENOENT")
return { path: source.replace(/^\.\.\/\.\.\//, ""), present: false, sha256: null };
throw error;
}
});
}
const REMOVED_PROCEDURES = [
"Execution contract:",
"Start actionable work in this heartbeat",
"clear final disposition",
"After 2 consecutive failures of the same control-plane write",
"Use child issues for parallel or long delegated work",
"a successful process exit or final response is not sufficient",
];
export function productionDefaultHireProfile(profile: RunnerProfileFixture): RunnerProfileFixture {
return {
...profile,
buildAgent(input) {
const { instructionsBundle: _fixtureManual, ...agent } = profile.buildAgent(input);
return agent;
},
};
}
export interface StockHarnessEvidence {
schema: "paperclip.stock-harness.v1";
generation: "legacy" | "native";
agentId: string;
budgets: { companyMonthlyCents: unknown; agentMonthlyCents: unknown };
bundle: { entryFile?: string; files: Array<{ path: string; content: string }> };
invocations: Array<{ runId: string; prompt: unknown; promptMetrics?: Record<string, unknown>; conversationMode?: boolean }>;
runIds: string[];
}
export function gradeStockHire(evidence: Pick<StockHarnessEvidence, "bundle" | "budgets">,
instructionVariant = readStockInstructionVariant()) {
const checks: Array<{ id: string; passed: boolean; detail: string }> = [];
const check = (id: string, passed: boolean, detail: string) => checks.push({ id, passed, detail });
check("budget-hard-stops", evidence.budgets.companyMonthlyCents === 1_000 && evidence.budgets.agentMonthlyCents === 1_000,
"Public company and agent records must both retain the 1,000-cent monthly hard stops.");
check("default-hire-bundle", evidence.bundle.entryFile === "AGENTS.md" &&
evidence.bundle.files.length === 1 && evidence.bundle.files[0]?.path === "AGENTS.md" &&
evidence.bundle.files[0]?.content === instructionVariant.content,
instructionVariant.variant === "reduced"
? "The public managed bundle must contain only the eight-word shipped identity, without an injected QA manual."
: "The comparison bundle must exactly match the independently pinned historical manual; this is structural evidence, not a task outcome.");
return checks;
}
export function gradeStockHarness(evidence: StockHarnessEvidence, instructionVariant = readStockInstructionVariant()) {
const checks = gradeStockHire(evidence, instructionVariant);
const check = (id: string, passed: boolean, detail: string) => checks.push({ id, passed, detail });
check("provider-runs-present", evidence.runIds.length > 0 && evidence.runIds.every(Boolean),
"The lifecycle oracle must reach actual provider runs; missing runs cannot pass instruction delivery.");
if (evidence.generation === "legacy") {
const prompts = evidence.invocations.filter(row => typeof row.prompt === "string" && row.prompt.length > 0);
check("invocation-evidence-complete", evidence.runIds.length > 0 && evidence.runIds.every(runId =>
prompts.some(row => row.runId === runId)),
"Every legacy run must have a public adapter.invoke event with its actual nonempty prompt.");
if (instructionVariant.variant === "reduced") {
check("generic-procedures-absent", prompts.length > 0 && prompts.every(row =>
REMOVED_PROCEDURES.every(procedure => !String(row.prompt).includes(procedure))),
"Neither startup nor continuation may reintroduce the removed generic manual.");
} else {
const classified = prompts.every(row => typeof row.conversationMode === "boolean" &&
typeof row.promptMetrics?.heartbeatPromptChars === "number" &&
Number.isFinite(row.promptMetrics.heartbeatPromptChars) && row.promptMetrics.heartbeatPromptChars >= 0);
const fresh = prompts.filter(row => Number(row.promptMetrics?.heartbeatPromptChars) > 0);
const resumed = prompts.filter(row => row.promptMetrics?.heartbeatPromptChars === 0);
check("historical-invocation-classification", prompts.length > 0 && classified,
"Every invocation needs public run-scoped conversation mode and explicit template-delivery metrics.");
check("historical-startup-contract", classified && fresh.length > 0 && fresh.every(row => row.conversationMode
? String(row.prompt).includes("Continue your Paperclip conversation using the supplied chat mode directive.") &&
String(row.prompt).includes("After 2 consecutive failures of the same control-plane write")
: String(row.prompt).includes("Execution contract:") && String(row.prompt).includes("Final disposition checklist:")),
`Check each of ${fresh.length} fresh invocations against its historical task or conversation template.`);
check("historical-continuation-contract", classified && resumed.every(row =>
String(row.prompt).includes("## Paperclip Resume Delta") && (row.conversationMode
? !String(row.prompt).includes("Execution contract:")
: String(row.prompt).includes("Execution contract: take concrete action") &&
String(row.prompt).includes("a successful process exit or final response is not sufficient"))),
`Check all ${resumed.length} observed continuation invocations separately; zero observations do not claim live continuation coverage.`);
}
const fresh = prompts.filter(row => Number(row.promptMetrics?.heartbeatPromptChars) > 0);
check("fresh-default-delivered", fresh.length > 0 && fresh.every(row =>
String(row.prompt).includes(STOCK_TEMPLATE_IDENTITY) && String(row.prompt).includes("Connection tools:") &&
String(row.prompt).includes("connections_search")),
"At least one fresh invocation must carry default identity and connection guidance.");
}
return checks;
}
/** Public API receipts only; no provider-memory claim or private runner hook. */
export async function captureStockHarness(input: {
api: { get<T>(path: string): Promise<T> };
agentId: string;
companyId: string;
generation: "legacy" | "native";
runIds: string[];
}): Promise<StockHarnessEvidence> {
const bundle = await input.api.get<{ entryFile: string; files: Array<{ path: string }> }>(
`/api/agents/${input.agentId}/instructions-bundle`,
);
const [company, agent] = await Promise.all([
input.api.get<{ budgetMonthlyCents: unknown }>(`/api/companies/${input.companyId}`),
input.api.get<{ budgetMonthlyCents: unknown }>(`/api/agents/${input.agentId}`),
]);
const files = await Promise.all(bundle.files.map(async file => {
const detail = await input.api.get<{ content: string }>(
`/api/agents/${input.agentId}/instructions-bundle/file?path=${encodeURIComponent(file.path)}`,
);
return { path: file.path, content: detail.content };
}));
const events = await Promise.all(input.runIds.map(async runId => ({
runId,
run: await input.api.get<{ id: string; companyId: string; agentId: string; contextSnapshot?: { conversationMode?: boolean } }>(
`/api/heartbeat-runs/${runId}`,
),
rows: await input.api.get<Array<{ eventType?: string; payload?: Record<string, unknown> }>>(
`/api/heartbeat-runs/${runId}/events?limit=1000`,
),
})));
if (events.some(({ runId, run }) => run.id !== runId || run.companyId !== input.companyId || run.agentId !== input.agentId ||
!run.contextSnapshot || typeof run.contextSnapshot !== "object" || Array.isArray(run.contextSnapshot) ||
(run.contextSnapshot.conversationMode !== undefined && typeof run.contextSnapshot.conversationMode !== "boolean"))) {
throw new Error("Stock invocation mode receipt is missing or bound to a different run/company/agent.");
}
return {
schema: "paperclip.stock-harness.v1", generation: input.generation,
budgets: { companyMonthlyCents: company.budgetMonthlyCents, agentMonthlyCents: agent.budgetMonthlyCents },
agentId: input.agentId, bundle: { entryFile: bundle.entryFile, files }, runIds: input.runIds,
invocations: events.flatMap(({ runId, rows, run }) => rows.filter(row => row.eventType === "adapter.invoke")
.map(row => ({ runId, prompt: row.payload?.prompt, promptMetrics: row.payload?.promptMetrics as Record<string, unknown> | undefined,
conversationMode: run.contextSnapshot?.conversationMode === true }))),
};
}
export function stockHarnessSourceDigest() {
const hash = createHash("sha256");
for (const source of [
"stock-harness.ts", "checkout-activity.ts", "context-integrity-cases.ts", "context-integrity-scoring.ts",
"context-integrity-flow.ts", "chat-cases.ts", "chat-flow.ts", "live-fixtures.ts", "runner.spec.ts",
"stock-harness-checks.mjs", "stock-harness-admission.ts", "stock-harness-manifest.ts",
"stock-harness-instruction-variant.mjs", "stock-harness-instruction-variant.d.mts", "stock-harness-instruction-variant.test.mjs",
"fixtures/stock-harness/historical-default-agents.md", "automatic-retry.ts", "automatic-retry.test.ts", "types.ts", "catalog.ts", "launch.ts",
"../../packages/adapter-utils/src/acpx-engine/execute.ts",
"../../packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.ts",
"../../packages/adapter-utils/src/acpx-engine/ephemeral-session-environment.test.ts",
"../../packages/paperclip-runner/scripts/generate-capability-contract.mjs",
"../../packages/paperclip-runner/scripts/check-capability-inventory.mjs",
"../../packages/paperclip-runner/scripts/lib/capability-inventory.mjs",
"../../packages/paperclip-runner/spec/capability/source-contract.json",
"../../packages/paperclip-runner/spec/capability/capabilities.yaml",
"../../packages/paperclip-runner/spec/capability/eval-traceability.yaml",
"../../packages/paperclip-runner/spec/capability/mcp-tool-map.yaml",
"../../packages/paperclip-runner/spec/capability/inventory.schema.json",
"../../packages/paperclip-runner/src/generated/capability-contract.ts",
"../../packages/paperclip-runner/docs/capability-contract.md",
...["capabilities.yaml", "mcp-tool-map.yaml", "eval-traceability.yaml", "capability-contract.md", "downstream-handoff.md"]
.map(file => `../../packages/paperclip-runner/generated/capability/${file}`),
"../../server/src/onboarding-assets/default/AGENTS.md",
"../../packages/adapter-utils/src/server-utils.ts",
"../../packages/shared/src/connection-intent-guidance.ts",
]) hash.update(source).update(readFileSync(new URL(source, import.meta.url)));
hash.update(JSON.stringify(stockHarnessSkillSources()));
return hash.digest("hex");
}