mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip manages AI agents and their work. > - Native Runner agents report completion through finish and block tools. > - The providers receive different descriptions for those tools. > - Completion guidance belongs with the tools that enforce the result. > - This pull request shares the descriptions and refreshes retained catalogs. > - A separate native suite checks completion and blocking on production defaults. > - Legacy agents retain their separate skill and API paths. ## Linked Issues or Issue Description Refs: #14920, #14948, #14985. **Current behavior** Native Codex and MCP bridges describe finish and block differently. Retained provider sessions can keep old descriptions. **Proposed behavior** Native providers receive the same finish and block descriptions. The descriptions cover report selection, validation feedback, returned outcomes, approval gates and final-answer timing. Retained native sessions refresh from v13 to v14. **Reason and benefit** Put the completion procedure next to its native tool. Preserve stock base instructions, schemas, permissions and terminal semantics. This PR now stands alone on master. It contains no reduced manual, shared prompt or operational-skill changes from #14948. ## What Changed - Add canonical native finish and block descriptions. Use them in direct Codex and both native MCP bridges. - Advance the native tool contract to v14. Cover old-v13 refresh without replacing task identity or prior history. - Check authenticated tool catalogs, provider start/resume frames and serialized daemon catalogs. - Add an independent, explicit-only native completion suite. Preserve the original assigned-skill durable-document journey. Pair it with a concrete whole-task blocker across Codex, ACPX Claude and OpenCode. - Verify the actual public production default bundle and budgets before execution. Require independent durable disposition, native result/terminal receipts and observable provider-final ordering. - Correct the blocker browser oracle to accept the requested explanation. Keep exact owner/action/scope checks. Calibrate positive, missing and contradictory replies. - Preserve only actual `tool_call` terminal names (`paperclip_finish` / `paperclip_block`) in the native compatibility run-log projection. Require the same named call ID through its finishing result; retain all other redaction boundaries. - Admit verified hosted shallow checkout/build hydration and bind the selected runnerd to exact source/archive/binary provenance. Hosted cells truthfully reuse the existing trusted build; local admission executes Rust calibration. Forward only public source/run identifiers through both launcher preflight subprocess paths. - Enforce single attempts in the launcher for opted-in fixtures. Keep ordinary retry policy unchanged. Run exact-source, credential-free admission before credential loading. ## Verification - Frozen candidate: `d6e59e4712a3158ab4cd7d58deff1389b4578c21`, based on master `59c07ede72dc08b8aba149a01cc11e0b7a204621`; historical descriptions: `e74ed61a69fbdd8b3a8f15dd6456bc3140246e33`. Exactly the five original native production files and six unit tests differ. Both carry identical corrected fixtures, strict named finishing-call grader, closed compatibility carrier and admission. Defaults, profiles/models/auth/permissions and manifest bytes match. - Actual launcher `prepareNativeCompletionPreflight` → `verifyNativeCompletionPreflight` admission passes on both exact refs with zero providers: candidate 132 / historical 127 selected TypeScript assertions, 128 Node calibrations and one Rust normalization calibration each; E2E typecheck, manifest checks, selected binary provenance and six-cell discovery pass. Each has 257 explicitly skipped unrelated assertions, not coverage. The credential-free environment calibration exercises both real prepare/verify subprocess options with public hosted identifiers and rejects credential/ambient overrides. Complete actual launcher prepare→verify also passes on both frozen refs with explicitly synthetic hosted metadata/verified archives, separately labeled as calibration rather than a trusted GitHub run. Exact framed provenance parsing and mock source identity are calibrated without relaxing the real verifier. - [Complete matched qualification report](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-master-qualification.md), [immutable manifest](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-manifest.json) and [closed retained audit/hashes](https://github.com/paperclipai/paperclip/blob/532066620b88e8731a5211fbe1cbc48ce8c7dd1a/doc/plans/2026-10-02-native-completion-calibrated-results/comparison.json) are inspectable. All six candidate cells pass; historical descriptions pass five. Paired outcomes: **zero new failures, one new pass (Codex blocker), five unchanged passes, zero pending pairs**. [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37098728980) and [historical campaign](https://github.com/paperclipai/paperclip/actions/runs/37098815696) each execute six original attempt-1 native runs, with no campaign retry and successful cleanup. Their trusted workflow revision is `215586d127e97c9301d86e769a39a15c13298ca2`, separate from measured source. [Candidate public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098728980-1/index.html) and [historical public HTML](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37098815696-1/index.html) retain declared screenshots. - Independent candidate evidence agrees with all original grades: 51 strict native checks, 12 served-default/budget checks and 21 original skill/document checks pass. The historical Codex blocker saves the correct whole-task blocker but omits the required marker from its actual provider final and identical saved reply. This is not semantic-summary fallback. Its original browser/matcher failure stays retained; the additional native snapshot/grade and workspace before/after digest were never written and are not fabricated by the separate API/PRP audit. Historical Codex completion has one failed finish followed by success within the same native run; the public receipt records no failure reason. All twelve runs and their usage remain counted. Reported model-cost subtotals are $0.00421482 historical/$0.00437391 candidate; Codex/Claude zero entries have unknown billing type, actual invoices are unverified and hosted execution cost is unmetered. One matched trial supports no extra failure within these six cases, not broad statistical or coding-quality equivalence. - Initial hosted `e18c2cf9` / `459455ac` and subsequent `0a9c5a7` / `00a761b` cohorts each stopped before providers in all twelve cells. The latter failed a mocked-receipt unit test under ambient hosted metadata; all source/build proofs passed. [All twelve later setup receipts](https://github.com/paperclipai/paperclip/blob/402ee94c52273ad58de355ae9a7d562dd22f8101/doc/plans/2026-10-02-native-completion-qualified-hosted-setup.json) are retained. [Exact failed setup receipts](https://github.com/paperclipai/paperclip/blob/27653eb1a8f8ce839776d760f4563f672e5a706c/doc/plans/2026-10-02-native-completion-master-hosted-setup.json) and the original manifest remain intact. Local sandbox-denied loopback and stale anchor-expectation attempts are retained separately; unchanged appropriate assertions were corrected/admitted before paid dispatch. Old anonymous OpenCode streams are not assigned inferred tool names or retroactively passed. - Full provider-free E2E support previously passed 927 tests in 67 files. Exact-head d6 normal CI run `37098409915`, attempt 1 passes full repository typecheck/build/tests, Runner Rust/static checks, all browser shards/aggregate and canary: 52 check-runs pass, four intentional skips, Snyk passes. Fresh Greptile check `111132956342` is 5/5 with zero unresolved threads. Source-specific deterministic tests do not substitute for the bounded live comparison. - Earlier native source `9138f570c341c251a5727c32d6615ce238bc8e03` is archived. Its [complete reduced-manual-context report](https://github.com/paperclipai/paperclip/blob/9138f570c341c251a5727c32d6615ce238bc8e03/doc/plans/2026-10-02-native-completion-live-comparison.md) remains intact, including original failures, grader limits and provider-free replay. It is not current-master-context qualification. ## Risks Changed tool text can change model behavior. The completed six-pair qualification shows no extra failing outcomes in this bounded trial; other tasks and repeated-run variance remain unmeasured. Observable final ordering does not prove provider feedback consumption. Public evidence can fail closed if a provider does not expose the required result sequence. This slice does not remove native fixed prompts or measure general coding quality. No database, schema, permission or legacy completion changes occur. ## Model Used OpenAI Codex, GPT-6 family, with code inspection, execution and tool use. The exact deployment ID and context-window size are not exposed in this session. They are unavailable rather than inferred from the model menu. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
298 lines
23 KiB
JavaScript
298 lines
23 KiB
JavaScript
// Independent, credential-free admission for the native completion comparison.
|
|
import { spawnSync } from "node:child_process";
|
|
import { createHash } from "node:crypto";
|
|
import { existsSync, mkdirSync, readdirSync, readFileSync, writeFileSync } from "node:fs";
|
|
import { join, resolve } from "node:path";
|
|
import {
|
|
NATIVE_COMPLETION_SOURCE_CONTRACT, NATIVE_COMPLETION_SOURCE_FILES,
|
|
nativeCompletionSourceFingerprint, nativeSourceSha256,
|
|
} from "./native-completion-source-contract.mjs";
|
|
|
|
import { inspectNativeCompletionSourceMetadata, inspectNativeCompletionRunnerd, NATIVE_COMPLETION_RUNNERD_PATH } from "./native-completion-git-source.mjs";
|
|
|
|
const root = resolve(import.meta.dirname, "../..");
|
|
export const NATIVE_COMPLETION_PREFLIGHT_SCHEMA = "paperclip.native-completion-preflight.v1";
|
|
export const NATIVE_COMPLETION_CELL_IDS = ["runner-codex", "runner-acpx-claude", "runner-opencode"].flatMap(profile =>
|
|
["assigned-skill-explicit-invocation", "native-blocked-report"].map(task => `native-completion.${profile}.local.${task}`));
|
|
export function nativeCompletionGates(variant) {
|
|
if (!["candidate", "historical"].includes(variant)) throw new Error("Unknown native completion source variant");
|
|
return [
|
|
{ id: "NC-schemas", name: "Variant-native completion schemas", cwd: "packages/paperclip-runner",
|
|
files: ["src/contracts/completion-result.test.ts"],
|
|
required: ["distinguishes user-facing answer content", "propagates answer and wait descriptions", "requires a reason code for verification that was not run"] },
|
|
{ id: "NC-tools", name: "Native tool delivery and retained provider catalogs", cwd: "packages/paperclip-runner", files: [
|
|
"src/drivers/codex/codex-app-server-driver.lifecycle.test.ts", "src/drivers/runner-tool-bridge.test.ts",
|
|
"src/drivers/opencode/mcp-bridge.test.ts", "src/live/runnerd-codex-transport.test.ts"],
|
|
testPattern: "binds to loopback|preserves stock Codex instructions on|includes ACPX terminal tools|preserves answer and internal wait descriptions",
|
|
required: ["runner semantic MCP bridge binds to loopback", "OpenCode MCP bridge binds to loopback",
|
|
"preserves stock Codex instructions on task recovery", "preserves stock Codex instructions on prepared recovery",
|
|
"preserves stock Codex instructions on direct recovery", "includes ACPX terminal tools in the authenticated bridge catalog",
|
|
...(variant === "candidate" ? ["codex", "opencode", "claude_managed", "aws_agentcore", "acpx"].map(provider =>
|
|
`preserves answer and internal wait descriptions in the serialized native ${provider} tool catalog`) : [])] },
|
|
{ id: "NC-resume", name: "Variant-native fingerprint and checkpoint refresh", cwd: ".",
|
|
files: ["server/src/services/native-runtime/native-session-resume.test.ts"], testPattern: "refreshes retained",
|
|
required: ["refreshes retained 'completion descriptions'", "refreshes retained 'task-bound human-input description'",
|
|
...(variant === "candidate" ? ["native completion tool guidance"] : [])] },
|
|
{ id: "NC-eval", name: "Independent native oracle, defaults, admission and attempt policy", cwd: ".",
|
|
config: "tests/runner-e2e/vitest.config.ts", files: [
|
|
"tests/runner-e2e/native-completion-scoring.test.ts", "tests/runner-e2e/native-completion-defaults.test.ts",
|
|
"tests/runner-e2e/native-completion-admission.test.ts", "tests/runner-e2e/automatic-retry.test.ts",
|
|
"tests/runner-e2e/context-integrity.test.ts", "tests/runner-e2e/context-integrity-evidence.test.ts",
|
|
"tests/runner-e2e/context-integrity-flow-logs.test.ts"],
|
|
required: ["allows toolchain paths and excludes every present or future credential",
|
|
"retains credential-free prerequisites inside the exact campaign root",
|
|
"retains the first transient_infrastructure failure without a second attempt",
|
|
"retains the first provider_variance failure without a second attempt"] },
|
|
];
|
|
}
|
|
export const NATIVE_COMPLETION_COMMAND_GATE_IDS = ["NC-node", "NC-typecheck", "NC-manifest", "NC-discovery", "NC-rust-carrier"];
|
|
export function nativeCompletionCommandGateIds(mode) {
|
|
return NATIVE_COMPLETION_COMMAND_GATE_IDS.map(id => mode === "trusted_hosted_archive" && id === "NC-rust-carrier" ? "NC-hosted-runnerd" : id);
|
|
}
|
|
export function gradeNativeCompletionGate(gate, report, exitCode) {
|
|
const assertions = (report?.testResults ?? []).flatMap(file => file.assertionResults ?? []);
|
|
const requirements = gate.required.map(name => ({ name, passed: assertions.some(assertion =>
|
|
assertion.fullName?.includes(name) && assertion.status === "passed") }));
|
|
const files = gate.files.map(file => ({ file, passed: (report?.testResults ?? []).some(result =>
|
|
result.name?.endsWith(file) && result.status === "passed" && result.assertionResults?.some(assertion => assertion.status === "passed") &&
|
|
result.assertionResults.every(assertion => assertion.status === "passed" || (gate.testPattern && assertion.status === "skipped" &&
|
|
!gate.required.some(name => assertion.fullName?.includes(name))))) }));
|
|
return { id: gate.id, name: gate.name, passed: exitCode === 0 && requirements.every(row => row.passed) && files.every(row => row.passed),
|
|
exitCode, files, requirements, total: report?.numTotalTests ?? 0, passedTests: report?.numPassedTests ?? 0,
|
|
failedTests: report?.numFailedTests ?? 0, pendingTests: report?.numPendingTests ?? 0 };
|
|
}
|
|
|
|
export function gradeNativeCompletionDiscovery(output, exitCode) {
|
|
const lines = output.trim().split(/\r?\n/).filter(Boolean);
|
|
const rows = lines.slice(1).map(line => line.split("\t"));
|
|
const ids = rows.map(row => row[0]);
|
|
return { passed: exitCode === 0 && lines[0] === "ID\tSUITE\tGENERATION\tPROVIDER\tMODEL\tCREDENTIALS" &&
|
|
rows.length === 6 && new Set(ids).size === 6 && NATIVE_COMPLETION_CELL_IDS.every(id => ids.includes(id)) &&
|
|
rows.every(row => row.length === 6 && row[1] === "native-completion" && row[2] === "native" && row[4]?.trim() && row[5]?.trim()),
|
|
executionIds: ids, expectedCells: 6, expectedTurns: 6, maximumAttemptsPerCell: 1 };
|
|
}
|
|
/** capture() stores compact provenance JSON on its first line, then diagnostics. */
|
|
export function gradeNativeCompletionHostedRunnerdEvidence(text, expected) {
|
|
try { return JSON.stringify(JSON.parse(text.split(/\r?\n/, 1)[0])) === JSON.stringify(expected); }
|
|
catch { return false; }
|
|
}
|
|
export function gradeNativeCompletionRustCarrier(output, exitCode) {
|
|
return exitCode === 0 && /^test provider_events::tests::preserves_closed_compatibility_terminal_tool_identity \.\.\. ok$/m.test(output) &&
|
|
/test result: ok\. 1 passed; 0 failed;/.test(output);
|
|
}
|
|
export function nativeCompletionPrerequisiteEnvironment(source) {
|
|
return Object.fromEntries(["PATH", "HOME", "TMPDIR", "TMP", "TEMP", "SYSTEMROOT", "LANG", "LC_ALL",
|
|
"CARGO_HOME", "RUSTUP_HOME", "CI", "GITHUB_ACTIONS", "GITHUB_RUN_ID", "GITHUB_RUN_ATTEMPT", "PAPERCLIP_RUNNER_E2E_SOURCE_SHA"].flatMap(name => source[name] === undefined ? [] : [[name, source[name]]]));
|
|
}
|
|
function currentSource() {
|
|
const source = nativeCompletionSourceFingerprint();
|
|
return { ...source, ...inspectNativeCompletionSourceMetadata({ repositoryRoot: root,
|
|
sourceFiles: NATIVE_COMPLETION_SOURCE_FILES, baseSha: source.baseSha, variant: source.variant,
|
|
shallowParentAnchors: NATIVE_COMPLETION_SOURCE_CONTRACT.shallowParentAnchors,
|
|
}) };
|
|
}
|
|
function buildOutputFingerprint() {
|
|
const hash = createHash("sha256");
|
|
function visit(relative) {
|
|
for (const entry of readdirSync(join(root, relative), { withFileTypes: true }).sort((a, b) => a.name.localeCompare(b.name))) {
|
|
const file = join(relative, entry.name);
|
|
if (entry.isDirectory()) visit(file);
|
|
else if (entry.isFile()) { const bytes = readFileSync(join(root, file)); hash.update(JSON.stringify([file, bytes.length])).update(bytes); }
|
|
else throw new Error(`Unexpected build output: ${file}`);
|
|
}
|
|
}
|
|
for (const directory of ["packages/shared/dist", "packages/plugins/sdk/dist", "packages/paperclip-runner/dist"]) {
|
|
if (!existsSync(join(root, directory, "index.js"))) throw new Error(`Missing prerequisite build output: ${directory}`);
|
|
visit(directory);
|
|
}
|
|
const daemon = runnerdBinary();
|
|
const bytes = readFileSync(daemon);
|
|
hash.update(JSON.stringify(["debug/paperclip-runnerd", bytes.length])).update(bytes);
|
|
return hash.digest("hex");
|
|
}
|
|
function runnerdBinary() {
|
|
return join(root, `${NATIVE_COMPLETION_RUNNERD_PATH}${process.platform === "win32" ? ".exe" : ""}`);
|
|
}
|
|
function runnerdProvenance(source, env) {
|
|
return inspectNativeCompletionRunnerd({ repositoryRoot: root, sourceSha: source.sha,
|
|
sourceFingerprint: source.fingerprint, environment: env });
|
|
}
|
|
export function assertNativeCompletionPreflightReceipt(report, current) {
|
|
const expected = [...nativeCompletionGates(current.variant).map(gate => gate.id), ...nativeCompletionCommandGateIds(current.runnerdProvenance?.mode)];
|
|
const hostedGate = report?.gates?.find(gate => gate.id === "NC-hosted-runnerd");
|
|
const truthfulHosted = current.runnerdProvenance?.mode !== "trusted_hosted_archive" ||
|
|
hostedGate?.executed === false && hostedGate?.calibration === "not_executed" && hostedGate?.total === 0 &&
|
|
hostedGate?.passedTests === 0 && hostedGate?.reuse === "trusted_same_run_build";
|
|
if (!truthfulHosted || report?.schema !== NATIVE_COMPLETION_PREFLIGHT_SCHEMA || report.passed !== true || report.providerCalls !== 0 ||
|
|
report.live !== "not_run" || report.sourceSha !== current.sha || !/^[a-f0-9]{40}$/.test(current.sha ?? "") ||
|
|
![current.fingerprint, current.fixtureFingerprint, current.manifestFingerprint, current.buildOutputFingerprint]
|
|
.every(value => typeof value === "string" && /^[a-f0-9]{64}$/.test(value)) ||
|
|
report.sourceFingerprint !== current.fingerprint || report.fixtureFingerprint !== current.fixtureFingerprint ||
|
|
report.manifestFingerprint !== current.manifestFingerprint || report.variant !== current.variant ||
|
|
report.baseSha !== NATIVE_COMPLETION_SOURCE_CONTRACT.baseSha || report.archiveSha !== NATIVE_COMPLETION_SOURCE_CONTRACT.archiveSha ||
|
|
report.layering !== true || current.layering !== true || report.immutable !== true || current.immutable !== true ||
|
|
!/^[a-f0-9]{64}$/.test(current.sourceMetadataFingerprint ?? "") ||
|
|
report.sourceMetadataFingerprint !== current.sourceMetadataFingerprint ||
|
|
!Array.isArray(report.sourceMetadataErrors) || report.sourceMetadataErrors.length !== 0 || current.sourceMetadataErrors?.length !== 0 ||
|
|
!Array.isArray(report.sourceErrors) || report.sourceErrors.length !== 0 || current.sourceErrors?.length !== 0 ||
|
|
report.setup?.passed !== true || report.setup.sdkExitCode !== 0 || report.setup.runnerTypeScriptExitCode !== 0 ||
|
|
report.setup.runnerdExitCode !== (current.runnerdProvenance?.mode === "trusted_hosted_archive" ? null : 0) ||
|
|
current.runnerdProvenance?.passed !== true || !["trusted_hosted_archive", "fresh_local_build"].includes(current.runnerdProvenance?.mode) ||
|
|
JSON.stringify(report.setup.runnerdProvenance) !== JSON.stringify(current.runnerdProvenance) ||
|
|
report.setup.runnerdSelectedPath !== current.runnerdProvenance.selectedPath ||
|
|
report.setup.runnerdSha256 !== current.runnerdProvenance.binarySha256 ||
|
|
!/^[a-f0-9]{64}$/.test(current.runnerdProvenance.binarySha256 ?? "") ||
|
|
report.setup.runnerdSourceSha !== current.sha || report.setup.runnerdSourceFingerprint !== current.fingerprint ||
|
|
report.setup.buildOutputFingerprint !== current.buildOutputFingerprint ||
|
|
!Array.isArray(report.setup.evidence) || report.setup.evidence.length !== 3 ||
|
|
["setup-sdk.txt", "setup-runner-typescript.txt", "setup-runnerd.txt"].some(file => report.setup.evidence.filter(row =>
|
|
row.file === file && /^[a-f0-9]{64}$/.test(row.sha256 ?? "")).length !== 1) ||
|
|
!Array.isArray(report.gates) || report.gates.length !== expected.length ||
|
|
expected.some(id => report.gates.filter(gate => gate.id === id && gate.passed === true && gate.exitCode === 0).length !== 1) ||
|
|
report.expectedCells !== 6 || report.expectedTurns !== 6 || report.maximumAttemptsPerCell !== 1) {
|
|
throw new Error("Native completion requires passing prerequisites for this exact immutable source SHA, variant and fingerprint.");
|
|
}
|
|
return report;
|
|
}
|
|
function runCommand(command, args, env, timeout = 10 * 60_000, cwd = root) {
|
|
return spawnSync(command, args, { cwd, env, encoding: "utf8", timeout });
|
|
}
|
|
function capture(output, name, run) {
|
|
const file = `${name}.txt`, bytes = `${run?.stdout ?? ""}\n${run?.stderr ?? ""}`;
|
|
writeFileSync(join(output, file), bytes);
|
|
return { file, sha256: nativeSourceSha256(bytes) };
|
|
}
|
|
export function assertRetainedNativeCompletionEvidence(output, evidence) {
|
|
if (!evidence || typeof evidence.file !== "string" || !/^[A-Za-z0-9_.-]+$/.test(evidence.file) ||
|
|
nativeSourceSha256(readFileSync(join(output, evidence.file))) !== evidence.sha256)
|
|
throw new Error("Native completion retained prerequisite evidence changed or is unavailable.");
|
|
return readFileSync(join(output, evidence.file), "utf8");
|
|
}
|
|
function manifestCommands(env) {
|
|
return ["generate-capability-contract.mjs", "check-capability-inventory.mjs"].map(file => runCommand(process.execPath,
|
|
[join(root, "packages/paperclip-runner/scripts", file), ...(file.startsWith("generate-") ? ["--check"] : [])], env));
|
|
}
|
|
|
|
export function main(args = process.argv.slice(2)) {
|
|
if (args.includes("--list")) {
|
|
console.log(JSON.stringify({ variants: ["candidate", "historical"], gates: nativeCompletionGates("candidate"),
|
|
commands: { local: nativeCompletionCommandGateIds("fresh_local_build"), hosted: nativeCompletionCommandGateIds("trusted_hosted_archive") },
|
|
cells: NATIVE_COMPLETION_CELL_IDS, providerCalls: 0 }, null, 2)); return;
|
|
}
|
|
if (args.some(arg => !arg.startsWith("--output-dir=") && !arg.startsWith("--verify=")))
|
|
throw new Error("Use --list, --output-dir=<path> or --verify=<receipt>; never paid providers.");
|
|
const source = currentSource(), env = nativeCompletionPrerequisiteEnvironment(process.env);
|
|
const verify = args.find(arg => arg.startsWith("--verify="))?.slice("--verify=".length);
|
|
if (verify) {
|
|
const report = assertNativeCompletionPreflightReceipt(JSON.parse(readFileSync(verify, "utf8")), {
|
|
...source, buildOutputFingerprint: buildOutputFingerprint(), runnerdProvenance: runnerdProvenance(source, env),
|
|
});
|
|
const output = resolve(verify, "..");
|
|
for (const evidence of report.setup.evidence) assertRetainedNativeCompletionEvidence(output, evidence);
|
|
for (const gate of nativeCompletionGates(source.variant)) {
|
|
const retained = report.gates.find(row => row.id === gate.id);
|
|
const grade = gradeNativeCompletionGate(gate, JSON.parse(assertRetainedNativeCompletionEvidence(output, retained.evidence)), 0);
|
|
if (!grade.passed) throw new Error(`Missing or failed retained native prerequisite assertions: ${gate.id}`);
|
|
}
|
|
for (const id of nativeCompletionCommandGateIds(report.setup.runnerdProvenance.mode)) {
|
|
const retained = report.gates.find(row => row.id === id), text = assertRetainedNativeCompletionEvidence(output, retained.evidence);
|
|
if (id === "NC-node" && (!/# fail 0\b/.test(text) || !/# tests [1-9]\d*\b/.test(text)))
|
|
throw new Error("Missing passing native admission calibrations.");
|
|
if (id === "NC-discovery" && !gradeNativeCompletionDiscovery(text.split("\n\n")[0], 0).passed)
|
|
throw new Error("Native completion six-cell discovery changed.");
|
|
if (id === "NC-rust-carrier" && !gradeNativeCompletionRustCarrier(text, 0))
|
|
throw new Error("Missing passing retained Rust terminal-tool carrier calibration.");
|
|
if (id === "NC-hosted-runnerd" && (retained.executed !== false || retained.calibration !== "not_executed" ||
|
|
retained.total !== 0 || retained.passedTests !== 0 ||
|
|
!gradeNativeCompletionHostedRunnerdEvidence(text, report.setup.runnerdProvenance)))
|
|
throw new Error("Hosted runnerd reuse must retain exact binary proof and report Rust calibration not executed.");
|
|
if (id === "NC-manifest" && manifestCommands(env).some(run => run.status !== 0))
|
|
throw new Error("Generated native capability manifests are stale.");
|
|
}
|
|
console.log(JSON.stringify(report)); return report;
|
|
}
|
|
const output = resolve(args.find(arg => arg.startsWith("--output-dir="))?.slice("--output-dir=".length) ??
|
|
join(root, "tests/runner-e2e/results", `native-completion-preflight-${new Date().toISOString().replaceAll(":", "-")}`));
|
|
mkdirSync(output, { recursive: true });
|
|
const report = { schema: NATIVE_COMPLETION_PREFLIGHT_SCHEMA, sourceSha: source.sha, sourceFingerprint: source.fingerprint,
|
|
fixtureFingerprint: source.fixtureFingerprint, manifestFingerprint: source.manifestFingerprint,
|
|
variant: source.variant, baseSha: source.baseSha, archiveSha: source.archiveSha,
|
|
sourceErrors: source.sourceErrors, immutable: source.immutable, layering: source.layering,
|
|
sourceMetadata: source.sourceMetadata, sourceMetadataErrors: source.sourceMetadataErrors,
|
|
sourceMetadataFingerprint: source.sourceMetadataFingerprint,
|
|
measuredAt: new Date().toISOString(), providerCalls: 0, live: "not_run", expectedCells: 6, expectedTurns: 6, maximumAttemptsPerCell: 1,
|
|
setup: { passed: false, sdkExitCode: null, runnerTypeScriptExitCode: null, runnerdExitCode: null, evidence: [] }, gates: [], passed: false };
|
|
if (!source.sha || source.sourceErrors.length || !source.layering || !source.immutable) {
|
|
writeFileSync(join(output, "preflight.json"), JSON.stringify(report, null, 2) + "\n");
|
|
console.error(`Native completion source admission failed before providers: ${[...source.sourceErrors, ...source.sourceMetadataErrors].join("; ")}`);
|
|
process.exitCode = 1; return report;
|
|
}
|
|
const sdk = runCommand(process.execPath, [join(root, "scripts/ensure-plugin-build-deps.mjs")], env, 5 * 60_000);
|
|
report.setup.sdkExitCode = sdk.status; report.setup.evidence.push(capture(output, "setup-sdk", sdk));
|
|
const runner = sdk.status === 0 ? runCommand("pnpm", ["--filter", "@paperclipai/paperclip-runner", "build:typescript"], env) : null;
|
|
report.setup.runnerTypeScriptExitCode = runner?.status ?? null;
|
|
report.setup.evidence.push(capture(output, "setup-runner-typescript", runner));
|
|
const hosted = env.GITHUB_ACTIONS === "true";
|
|
const runnerd = sdk.status === 0 && runner?.status === 0 && !hosted ? runCommand("cargo",
|
|
["build", "--locked", "--offline", "--workspace", "--bins"], env,
|
|
10 * 60_000, join(root, "packages/paperclip-runner/runner")) : null;
|
|
report.setup.runnerdExitCode = runnerd?.status ?? null;
|
|
report.setup.runnerdProvenance = runnerdProvenance(source, env);
|
|
report.setup.evidence.push(capture(output, "setup-runnerd", hosted
|
|
? { stdout: JSON.stringify(report.setup.runnerdProvenance), stderr: "Rust build and unit calibration not executed in this hosted cell; reusing trusted same-run build." }
|
|
: runnerd));
|
|
report.setup.passed = sdk.status === 0 && runner?.status === 0 && (hosted || runnerd?.status === 0) && report.setup.runnerdProvenance.passed;
|
|
if (report.setup.passed) {
|
|
report.setup.buildOutputFingerprint = buildOutputFingerprint();
|
|
report.setup.runnerdSha256 = report.setup.runnerdProvenance.binarySha256;
|
|
report.setup.runnerdSelectedPath = report.setup.runnerdProvenance.selectedPath;
|
|
report.setup.runnerdSourceSha = source.sha;
|
|
report.setup.runnerdSourceFingerprint = source.fingerprint;
|
|
for (const gate of nativeCompletionGates(source.variant)) {
|
|
console.log(`Checking ${gate.id}: ${gate.name}`);
|
|
const file = `${gate.id}.json`, run = runCommand(process.execPath, [join(root, "node_modules/vitest/vitest.mjs"), "run", ...gate.files,
|
|
...(gate.config ? ["--config", gate.config] : []), ...(gate.testPattern ? ["--testNamePattern", gate.testPattern] : []),
|
|
"--reporter=default", "--reporter=json", `--outputFile.json=${join(output, file)}`], env, 10 * 60_000, resolve(root, gate.cwd));
|
|
capture(output, gate.id, run);
|
|
let result; try { result = JSON.parse(readFileSync(join(output, file), "utf8")); } catch { /* Missing discovery fails closed. */ }
|
|
report.gates.push({ ...gradeNativeCompletionGate(gate, result, run.status),
|
|
evidence: existsSync(join(output, file)) ? { file, sha256: nativeSourceSha256(readFileSync(join(output, file))) } : null });
|
|
}
|
|
const commands = [
|
|
{ id: "NC-node", command: process.execPath, args: ["--test", "--test-reporter=tap", "tests/runner-e2e/native-completion-checks.test.mjs", "tests/runner-e2e/native-completion-source-contract.test.mjs", "tests/runner-e2e/native-completion-git-source.test.mjs"] },
|
|
{ id: "NC-typecheck", command: process.execPath, args: ["node_modules/typescript/bin/tsc", "-p", "tests/runner-e2e/tsconfig.json"] },
|
|
...(!hosted ? [{ id: "NC-rust-carrier", command: "cargo", args: ["test", "--locked", "--offline",
|
|
"-p", "paperclip-runner-core", "--lib", "provider_events::tests::preserves_closed_compatibility_terminal_tool_identity", "--", "--exact"],
|
|
cwd: join(root, "packages/paperclip-runner/runner") }] : []),
|
|
{ id: "NC-discovery", command: process.execPath, args: ["--import", "./server/node_modules/tsx/dist/loader.mjs", "tests/runner-e2e/launch.ts", "--list", "--suite", "native-completion"] },
|
|
];
|
|
for (const command of commands) {
|
|
console.log(`Checking ${command.id}`);
|
|
const run = runCommand(command.command, command.args, env, 10 * 60_000, command.cwd ?? root);
|
|
report.gates.push({ id: command.id, passed: run.status === 0 &&
|
|
(command.id !== "NC-discovery" || gradeNativeCompletionDiscovery(run.stdout, run.status).passed) &&
|
|
(command.id !== "NC-rust-carrier" || gradeNativeCompletionRustCarrier(`${run.stdout ?? ""}\n${run.stderr ?? ""}`, run.status)) &&
|
|
(command.id !== "NC-node" || /# fail 0\b/.test(run.stdout) && /# tests [1-9]\d*\b/.test(run.stdout)),
|
|
exitCode: run.status, evidence: capture(output, command.id, run) });
|
|
}
|
|
if (hosted) report.gates.push({ id: "NC-hosted-runnerd", name: "Trusted hosted runnerd provenance; Rust calibration not executed",
|
|
passed: report.setup.runnerdProvenance.passed, exitCode: 0, executed: false, calibration: "not_executed",
|
|
reuse: "trusted_same_run_build", total: 0, passedTests: 0, evidence: capture(output, "NC-hosted-runnerd", {
|
|
stdout: JSON.stringify(report.setup.runnerdProvenance), stderr: "Rust unit calibration must be qualified by the separate exact-source local receipt and normal CI." }) });
|
|
const manifest = manifestCommands(env), combined = { stdout: manifest.map(run => run.stdout ?? "").join("\n"), stderr: manifest.map(run => run.stderr ?? "").join("\n") };
|
|
report.gates.push({ id: "NC-manifest", passed: manifest.every(run => run.status === 0),
|
|
exitCode: manifest.find(run => run.status !== 0)?.status ?? 0, evidence: capture(output, "NC-manifest", combined) });
|
|
const after = currentSource();
|
|
report.passed = report.gates.every(gate => gate.passed) && after.immutable && after.layering &&
|
|
after.sha === source.sha && after.fingerprint === source.fingerprint && after.sourceErrors.length === 0 &&
|
|
after.sourceMetadataFingerprint === source.sourceMetadataFingerprint && after.sourceMetadataErrors.length === 0 &&
|
|
buildOutputFingerprint() === report.setup.buildOutputFingerprint &&
|
|
JSON.stringify(runnerdProvenance(after, env)) === JSON.stringify(report.setup.runnerdProvenance);
|
|
}
|
|
writeFileSync(join(output, "preflight.json"), JSON.stringify(report, null, 2) + "\n");
|
|
console.log(`Native completion prerequisite evidence: ${join(output, "preflight.json")}`);
|
|
if (!report.passed) process.exitCode = 1;
|
|
return report;
|
|
}
|
|
if (process.argv[1] && resolve(process.argv[1]) === resolve(import.meta.filename)) main();
|