mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip manages AI agents and their provider connections. > - Product E2E checks real tasks through the browser, server, and runner. > - Grok qualification needs separate API-key and subscription evidence. > - Product subscription tests and direct Grok protocol evals need explicit credential delivery. > - This change supplies each credential only to its selected profile and prepares the pinned binary. > - Maintainer authorization and protected-environment gates remain required. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13882. The Grok feature branch has a manual subscription qualification profile. The trusted master workflow must admit its selected credential and prepare the same verified binary and artifact verifier as the API profile. Direct protocol evals also need the selected xAI key and pinned Grok binary. These prerequisites do not register or schedule the new profiles on master. ## What Changed - Deliver `GROK_AUTH_JSON` from the protected paid environment only when the selected profile requests that credential. - Install the checksum-verified Grok binary for the local subscription profile. - Prepare the pinned artifact verifier for the manual subscription suite. - Extend workflow security assertions to cover the new credential and profile. - Add the ACPX Grok credential mapping to the trusted-master catalog, then deliver only the selected `XAI_API_KEY` to direct protocol cells and install the target’s checksum-verified Grok binary before packaging. - Allow a direct-protocol concurrency override from two cases up to the existing configured ceiling; it can only lower concurrency. - Document the Grok protocol workflow and its API-only credential boundary. - Render missing LLM usage and cost as Unavailable, and label partial observations with coverage. Preserve raw records, grades, and actual zero costs. - Preserve measured campaign source metadata during report regeneration instead of inheriting the renderer checkout or CI event; skip empty legacy source records when recovering older provenance. ## Verification - Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54 reported checks successful, two intentional skips, Greptile 5/5, and zero unresolved review threads. [CI run](https://github.com/paperclipai/paperclip/actions/runs/35890978288). - After merging current master, all 17 workflow security/image tests and 23 catalog/workflow policy tests passed. The trusted catalog also generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per shard. The new policy tests execute the concurrency guard against valid, out-of-range, and malformed values. - The Grok branch separately passed 450 Product harness unit tests, including private company credential staging, cleanup, and token-fragment redaction. - The fresh-login native subscription smoke passed three repetitions of tool execution, session resume, restrictive permissions, and cleanup. These are setup evidence; full subscription Product qualification remains pending. - All 72 focused report/billing/history/catalog tests and the Product harness typecheck passed for the report-display change. The initial sandbox run could not open the tsx IPC socket; the permitted rerun passed. A zero-provider-call replay of the actual 16-cell Grok campaign preserved all result records, grades, timing, and source provenance while correcting missing usage labels. - Review the thirteen-file diff. Provider credentials still enter only the selected paid-test step; default-branch, numeric-actor, and environment restrictions are unchanged. ## Risks This admits a refreshable subscription credential to explicitly selected trusted tests. Store it only in `runner-e2e-paid`, use a test login, and remove it after qualification. Unselected profiles receive an empty value. Pull requests cannot trigger the paid workflow. This PR changes no fleet admission or actor allowlist. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
458 lines
20 KiB
JavaScript
458 lines
20 KiB
JavaScript
import assert from "node:assert/strict";
|
|
import { spawnSync } from "node:child_process";
|
|
import { existsSync, readdirSync, readFileSync } from "node:fs";
|
|
import path from "node:path";
|
|
import { fileURLToPath } from "node:url";
|
|
import test from "node:test";
|
|
|
|
const repoRoot = path.resolve(
|
|
path.dirname(fileURLToPath(import.meta.url)),
|
|
"..",
|
|
"..",
|
|
);
|
|
|
|
function readWorkflow(name) {
|
|
return readFileSync(path.join(repoRoot, ".github/workflows", name), "utf8");
|
|
}
|
|
|
|
test("chaos verification isolates callers that verify the same source commit", () => {
|
|
const chaosWorkflow = readWorkflow("runner-chaos-evals.yml");
|
|
const group = chaosWorkflow.match(/^ group: (.+)$/m)?.[1];
|
|
assert.ok(group, "chaos verification must define its concurrency group");
|
|
|
|
// GitHub supplies the top-level caller's workflow name to reusable calls.
|
|
const resolveGroup = (caller, ref) => group
|
|
.replaceAll("${{ github.workflow }}", readWorkflow(caller).match(/^name: (.+)$/m)[1])
|
|
.replaceAll("${{ inputs.ref || github.ref }}", ref)
|
|
.toLowerCase();
|
|
const sha = "a".repeat(40);
|
|
const callers = ["cloud-readiness.yml", "release.yml", "runner-chaos-evals.yml"];
|
|
const groups = callers.map((caller) => resolveGroup(caller, sha));
|
|
assert.equal(new Set(groups).size, callers.length,
|
|
"Cloud readiness, Release, and standalone evals must not cancel each other");
|
|
assert.ok(groups.every((value) => !value.includes("${{")), "resolve every group input");
|
|
assert.notEqual(resolveGroup("cloud-readiness.yml", sha),
|
|
resolveGroup("cloud-readiness.yml", "b".repeat(40)), "different sources remain independent");
|
|
assert.match(chaosWorkflow, /cancel-in-progress: true/);
|
|
});
|
|
|
|
test("canary reuses exact-source proof while stable keeps full verification", () => {
|
|
const releaseWorkflow = readWorkflow("release.yml");
|
|
const canary = releaseWorkflow.split(" verify_canary:\n")[1].split("\n publish_canary:")[0];
|
|
assert.match(canary, /github\.repository == 'paperclipai\/paperclip' && github\.event_name == 'push' && github\.ref == 'refs\/heads\/master'/);
|
|
assert.match(canary, /actions: read/);
|
|
assert.match(canary, /ref: \$\{\{ github\.sha \}\}/);
|
|
assert.match(canary, /SOURCE_SHA: \$\{\{ github\.sha \}\}/);
|
|
assert.match(canary, /run: node scripts\/cloud-source-verification\.mjs "\$SOURCE_SHA"/);
|
|
assert.doesNotMatch(canary, /release-verify\.yml|continue-on-error|always\(\)/);
|
|
assert.match(releaseWorkflow, /publish_canary:\n\s+if: github\.event_name == 'push'\n\s+needs: verify_canary/);
|
|
// The stable lane is gated on the stable channel since the nightly lane
|
|
// was added; a `needs:` line (for example a preflight job) may sit between
|
|
// the gate and the delegation.
|
|
// The stable preflight resolves source_ref to an immutable SHA exactly
|
|
// once; verification must consume that pin, not re-resolve the ref.
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/verify_stable:\n\s+if: github\.event_name == 'workflow_dispatch' && inputs\.channel == 'stable'\n(?:\s+needs: [^\n]+\n)?\s+uses: \.\/\.github\/workflows\/release-verify\.yml\n\s+with:\n\s+ref: \$\{\{ needs\.preflight_stable\.outputs\.sha \}\}/,
|
|
);
|
|
assert.doesNotMatch(
|
|
releaseWorkflow,
|
|
/verify_(?:canary|stable):[\s\S]*?pnpm test:run(?:\n|$)/,
|
|
);
|
|
});
|
|
|
|
test("source proof requires every source check and does not wait on image publication", () => {
|
|
const readiness = readWorkflow("cloud-readiness.yml");
|
|
const proof = readiness.split(" source_verified:\n")[1].split("\n ready:")[0];
|
|
assert.match(proof, /name: Cloud source verified v1/);
|
|
assert.match(proof, /needs: \[verify\]/);
|
|
assert.match(proof, /node --test scripts\/cloud-source-verification.test.mjs/);
|
|
assert.match(proof, /SOURCE_SHA: \$\{\{ github\.sha \}\}/);
|
|
assert.doesNotMatch(proof, /always\(\)|continue-on-error|needs:.*(?:image|artifacts)/);
|
|
assert.match(readiness.split(" ready:\n")[1], /needs: \[verify, image, artifacts\]/);
|
|
});
|
|
|
|
test("onboard smoke container binds beyond loopback so the mapped port is reachable", () => {
|
|
const dockerfile = readFileSync(
|
|
path.join(repoRoot, "docker/Dockerfile.onboard-smoke"),
|
|
"utf8",
|
|
);
|
|
|
|
// `onboard --yes` without an explicit --bind prefers trusted-local
|
|
// defaults and writes a loopback bind, which Docker port mapping cannot
|
|
// reach. The smoke container must pin a non-loopback preset.
|
|
assert.match(dockerfile, /onboard --yes --bind lan/);
|
|
});
|
|
|
|
test("promotion selection guards against sources that predate their channel tooling", () => {
|
|
const releaseWorkflow = readWorkflow("release.yml");
|
|
|
|
// Promotions run the source commit's release.sh, so selection must reject
|
|
// sources whose tooling does not know the target channel yet.
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/git show "\$\{sha\}:scripts\/release\.sh" \| grep -qF 'canary\|nightly'/,
|
|
);
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/git show "\$\{sha\}:scripts\/release\.sh" \| grep -qF 'canary\|nightly\|beta\|stable\)'/,
|
|
);
|
|
});
|
|
|
|
test("candidate-branch betas are validated and fully verified before publish", () => {
|
|
const releaseWorkflow = readWorkflow("release.yml");
|
|
|
|
// Candidate heads are new commits: selection must pin the naming
|
|
// convention and publication must be gated on full verification.
|
|
assert.match(releaseWorkflow, /candidate\/beta-\*\)/);
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/verify_beta_candidate:\n\s+needs: select_beta\n\s+if: needs\.select_beta\.outputs\.mode == 'candidate'\n\s+uses: \.\/\.github\/workflows\/release-verify\.yml/,
|
|
);
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/needs\.verify_beta_candidate\.result == 'success'/,
|
|
);
|
|
});
|
|
|
|
test("post-publish beta smoke survives the skipped candidate-verification ancestor", () => {
|
|
const releaseWorkflow = readWorkflow("release.yml");
|
|
|
|
// publish_beta's needs chain contains verify_beta_candidate, which is
|
|
// skipped on promote-mode betas. An `if:` without a status-check function
|
|
// gets an implicit success() that evaluates that chain transitively and
|
|
// silently skips the smoke. The condition must stay explicit.
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/smoke_beta:\n\s+needs: publish_beta\n\s+if: \$\{\{ !cancelled\(\) && needs\.publish_beta\.result == 'success' && !inputs\.dry_run \}\}/,
|
|
);
|
|
});
|
|
|
|
test("published canaries are gated by the exact-version onboarding browser smoke", () => {
|
|
const releaseWorkflow = readWorkflow("release.yml");
|
|
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/publish_canary:[\s\S]*?outputs:\n\s+canary_version: \$\{\{ steps\.canary_tag\.outputs\.version \}\}/,
|
|
);
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/smoke_canary_onboarding:\n\s+needs: publish_canary\n\s+if: needs\.publish_canary\.result == 'success'/,
|
|
);
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/PAPERCLIPAI_VERSION: \$\{\{ needs\.publish_canary\.outputs\.canary_version \}\}/,
|
|
);
|
|
assert.match(releaseWorkflow, /test:canary-onboarding-smoke/);
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/smoke_canary_onboarding:[\s\S]*?uses: actions\/checkout@[0-9a-f]{40} # v7[\s\S]*?uses: pnpm\/action-setup@[0-9a-f]{40} # v6[\s\S]*?uses: actions\/setup-node@[0-9a-f]{40} # v7/,
|
|
);
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/smoke_canary_onboarding:[\s\S]*?Install test dependencies\n\s+run: pnpm install --frozen-lockfile/,
|
|
);
|
|
assert.doesNotMatch(
|
|
releaseWorkflow.match(
|
|
/smoke_canary_onboarding:[\s\S]*?(?=\n # ----- Nightly lane)/,
|
|
)?.[0] ?? "",
|
|
/cache: pnpm/,
|
|
);
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/name: Smoke exact published canary through onboarding\n\s+env:\n\s+PAPERCLIP_CANARY_SMOKE_SERVER_LOG: \$\{\{ runner\.temp \}\}\/canary-onboarding-server\.log/,
|
|
);
|
|
assert.match(
|
|
releaseWorkflow,
|
|
/smoke_canary_onboarding:[\s\S]*?uses: actions\/upload-artifact@[0-9a-f]{40} # v7/,
|
|
);
|
|
assert.match(releaseWorkflow, /canary-onboarding-server\.log/);
|
|
assert.match(releaseWorkflow, /tests\/canary-onboarding\/playwright-report/);
|
|
});
|
|
|
|
test("every lane's tag push degrades to recovery instructions when rejected", () => {
|
|
const releaseWorkflow = readWorkflow("release.yml");
|
|
|
|
// GITHUB_TOKEN may not create refs pointing at workflow-modifying commits
|
|
// from dispatch or scheduled runs; a rejected tag push after a successful
|
|
// npm publish must surface runbook recovery commands, not a bare error.
|
|
const occurrences = releaseWorkflow.match(/## Tag push rejected/g) ?? [];
|
|
assert.equal(
|
|
occurrences.length,
|
|
3,
|
|
"nightly, beta, and stable each carry the recovery summary",
|
|
);
|
|
});
|
|
|
|
test("release smoke workflow extends the container readiness budget for CI", () => {
|
|
const smokeWorkflow = readWorkflow("release-smoke.yml");
|
|
const harness = readFileSync(
|
|
path.join(repoRoot, "scripts/docker-onboard-smoke.sh"),
|
|
"utf8",
|
|
);
|
|
|
|
// CI containers cold-install paperclipai and embedded postgres, so the
|
|
// workflow must extend the harness's local-default readiness budget.
|
|
assert.match(smokeWorkflow, /SMOKE_READY_TIMEOUT_SECONDS=\d+/);
|
|
const ciBudget = Number(
|
|
smokeWorkflow.match(/SMOKE_READY_TIMEOUT_SECONDS=(\d+)/)[1],
|
|
);
|
|
assert.ok(
|
|
ciBudget >= 300,
|
|
`CI readiness budget ${ciBudget}s should be at least 300s`,
|
|
);
|
|
|
|
assert.match(
|
|
harness,
|
|
/SMOKE_READY_TIMEOUT_SECONDS="\$\{SMOKE_READY_TIMEOUT_SECONDS:-\d+\}"/,
|
|
);
|
|
assert.match(
|
|
harness,
|
|
/wait_for_http "\$PAPERCLIP_PUBLIC_URL\/api\/health" "\$SMOKE_READY_TIMEOUT_SECONDS" 1/,
|
|
);
|
|
});
|
|
|
|
test("release verify workflow covers the same split test surface as stable PR verification", () => {
|
|
const verifyWorkflow = readWorkflow("release-verify.yml");
|
|
|
|
assert.match(verifyWorkflow, /workflow_call:/);
|
|
assert.match(
|
|
verifyWorkflow,
|
|
/node \.\/scripts\/release-package-map\.mjs check/,
|
|
);
|
|
assert.match(verifyWorkflow, /pnpm -r typecheck/);
|
|
assert.match(verifyWorkflow, /pnpm build/);
|
|
const runnerScripts = JSON.parse(readFileSync(path.join(repoRoot, "packages/paperclip-runner/package.json"), "utf8")).scripts;
|
|
const runnerChecks = [...verifyWorkflow.matchAll(/^ checks: (.+)$/gm)]
|
|
.flatMap(([, checks]) => checks.split(" "));
|
|
assert.deepEqual(runnerChecks, runnerScripts["check:all"].split(" && ")
|
|
.map((command) => command.replace(/^pnpm run /, "")));
|
|
assert.match(verifyWorkflow, /pnpm --filter @paperclipai\/paperclip-runner "\$check"/);
|
|
assert.match(verifyWorkflow, /runner_workflow_evals:/);
|
|
assert.match(verifyWorkflow, /runner_chaos_evals:/);
|
|
assert.match(
|
|
verifyWorkflow,
|
|
/uses: \.\/\.github\/workflows\/runner-chaos-evals\.yml/,
|
|
);
|
|
assert.match(
|
|
verifyWorkflow,
|
|
/runner_workflow_evals:[\s\S]*?Install dependencies\n\s+run: pnpm install --no-frozen-lockfile[\s\S]*?Run deterministic Runner workflow scorer tests/,
|
|
);
|
|
assert.match(verifyWorkflow, /pnpm test:runner-workflow-evals/);
|
|
|
|
const buildJob = verifyWorkflow.match(/ build:\n[\s\S]*?(?=\n [A-Za-z0-9_-]+:|$)/)?.[0] ?? "";
|
|
assert.match(buildJob, /persist-credentials: false/);
|
|
assert.doesNotMatch(buildJob, /cache: pnpm/);
|
|
|
|
for (const group of ["general-server-without-chat", "general-chat", "general-workspaces-a", "general-workspaces-b"]) {
|
|
assert.match(verifyWorkflow, new RegExp(`group: ${group}`));
|
|
}
|
|
for (const [group, count] of [["general-server-without-chat", 10], ["general-chat", 3]]) {
|
|
const rows = [...verifyWorkflow.matchAll(new RegExp(`group: ${group}\\n\\s+group_label: [^\\n]+\\n\\s+shard_index: (\\d+)\\n\\s+shard_count: (\\d+)`, "g"))];
|
|
assert.deepEqual(rows.map((row) => [Number(row[1]), Number(row[2])]),
|
|
Array.from({ length: count }, (_, index) => [index, count]));
|
|
}
|
|
for (const shardIndex of [0, 1, 2, 3, 4]) {
|
|
assert.match(verifyWorkflow, new RegExp(`shard_index: ${shardIndex}[\\s\\S]*?shard_count: 5`));
|
|
}
|
|
|
|
// workspaces-a splits with Vitest native --shard in pr.yml; release
|
|
// verification must keep the same two-shard coverage.
|
|
for (const shardIndex of [0, 1]) {
|
|
assert.match(
|
|
verifyWorkflow,
|
|
new RegExp(
|
|
`group: general-workspaces-a[\\s\\S]*?shard_index: ${shardIndex}\\n\\s+shard_count: 2`,
|
|
),
|
|
);
|
|
}
|
|
|
|
assert.match(verifyWorkflow, /pnpm test:run:general -- --group/);
|
|
assert.match(verifyWorkflow, /pnpm test:run:serialized -- --shard-index/);
|
|
});
|
|
|
|
test("Runner eval workflows pin actions and gate paid live execution", () => {
|
|
const actionPinWorkflows = [
|
|
readWorkflow("release-verify.yml"),
|
|
readWorkflow("runner-live-evals.yml"),
|
|
readWorkflow("runner-chaos-evals.yml"),
|
|
readWorkflow("runner-full-stack-e2e.yml"),
|
|
readWorkflow("e2e.yml"),
|
|
readWorkflow("runner-protocol-live-evals.yml"),
|
|
];
|
|
|
|
for (const workflow of actionPinWorkflows) {
|
|
const remoteUses = workflow
|
|
.split("\n")
|
|
.filter(
|
|
(line) =>
|
|
/^\s*(?:-\s*)?uses: /.test(line) && !line.includes("uses: ./"),
|
|
);
|
|
assert.ok(
|
|
remoteUses.length > 0,
|
|
"expected at least one remote action reference",
|
|
);
|
|
for (const line of remoteUses) {
|
|
assert.match(line, /uses: [^@\s]+@[0-9a-f]{40}(?:\s+# .+)?$/);
|
|
}
|
|
}
|
|
|
|
const liveWorkflow = actionPinWorkflows[1];
|
|
assert.match(liveWorkflow, /RUNNER_LIVE_EVALS_NIGHTLY_ENABLED == 'true'/);
|
|
assert.match(liveWorkflow, /REF: \$\{\{ github\.ref \}\}/);
|
|
assert.match(liveWorkflow, /refs\/heads\/\$DEFAULT_BRANCH/);
|
|
assert.match(
|
|
liveWorkflow,
|
|
/DEFAULT_BRANCH: \$\{\{ github\.event\.repository\.default_branch \}\}/,
|
|
);
|
|
assert.match(liveWorkflow, /RUNNER_E2E_ALLOWED_ACTOR_IDS/);
|
|
assert.match(liveWorkflow, /needs: authorize/);
|
|
assert.match(liveWorkflow, /environment:\n\s+name: runner-e2e-paid/);
|
|
assert.match(
|
|
liveWorkflow,
|
|
/OPENAI_API_KEY: \$\{\{ secrets\.OPENAI_API_KEY \}\}/,
|
|
);
|
|
|
|
const paidWorkflowNames = [
|
|
"e2e.yml",
|
|
"runner-full-stack-e2e.yml",
|
|
"runner-live-evals.yml",
|
|
"runner-protocol-live-evals.yml",
|
|
];
|
|
const paidWorkflowNameSet = new Set(paidWorkflowNames);
|
|
const providerSecretReference =
|
|
/secrets(?:\.(?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|GROK_AUTH_JSON|DAYTONA_API_KEY)\b|\[['"](?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|GROK_AUTH_JSON|DAYTONA_API_KEY)['"]\])/g;
|
|
for (const name of readdirSync(path.join(repoRoot, ".github/workflows"))) {
|
|
if (!/\.ya?ml$/.test(name)) continue;
|
|
const workflow = readWorkflow(name);
|
|
if ([...workflow.matchAll(providerSecretReference)].length > 0) {
|
|
assert.ok(
|
|
paidWorkflowNameSet.has(name),
|
|
`${name} must not receive provider credentials`,
|
|
);
|
|
}
|
|
}
|
|
|
|
for (const name of paidWorkflowNames) {
|
|
const workflow = readWorkflow(name);
|
|
const triggerHeader = workflow.slice(0, workflow.indexOf("\njobs:\n"));
|
|
assert.doesNotMatch(
|
|
triggerHeader,
|
|
/^\s{2}(?:pull_request|pull_request_target|push|workflow_call|workflow_run):/m,
|
|
);
|
|
assert.match(triggerHeader, /^\s{2}workflow_dispatch:/m);
|
|
assert.match(workflow, /^ authorize:/m);
|
|
|
|
const jobBlocks = workflow
|
|
.slice(workflow.indexOf("\njobs:\n") + "\njobs:\n".length)
|
|
.split(/\n(?= [A-Za-z0-9_-]+:\n)/);
|
|
const providerJobs = jobBlocks.filter(
|
|
(block) => [...block.matchAll(providerSecretReference)].length > 0,
|
|
);
|
|
assert.ok(providerJobs.length > 0, `${name} needs a provider-secret job`);
|
|
for (const block of providerJobs) {
|
|
assert.match(block, /\n environment:\n name: runner-e2e-paid\n/);
|
|
assert.match(
|
|
block,
|
|
/\n steps:(?: &[A-Za-z0-9_-]+)?\n(?:\s*\n)* - name: Reauthorize[^\n]*\n/,
|
|
`${name} must reauthorize as the first provider-job step`,
|
|
);
|
|
const reauthorize = block.indexOf(" - name: Reauthorize");
|
|
assert.ok(reauthorize > 0);
|
|
assert.ok(block.indexOf("actions/checkout@") > reauthorize);
|
|
assert.ok(block.search(providerSecretReference) > reauthorize);
|
|
assert.match(block, /github\.actor_id/);
|
|
assert.match(block, /github\.triggering_actor/);
|
|
assert.match(block, /RUNNER_E2E_ALLOWED_ACTOR_IDS/);
|
|
assert.match(block, /refs\/heads\/\$DEFAULT_BRANCH/);
|
|
assert.doesNotMatch(block, /^\s+cache: pnpm$/m);
|
|
}
|
|
}
|
|
|
|
const fullStackWorkflow = readWorkflow("runner-full-stack-e2e.yml");
|
|
for (const [secret, condition] of Object.entries({
|
|
OPENAI_API_KEY: "matrix.credentialName == 'OPENAI_API_KEY'",
|
|
ANTHROPIC_API_KEY: "matrix.credentialName == 'ANTHROPIC_API_KEY'",
|
|
OPENROUTER_API_KEY: "matrix.credentialName == 'OPENROUTER_API_KEY'",
|
|
DAYTONA_API_KEY: "matrix.environmentId == 'daytona'",
|
|
})) {
|
|
assert.ok(
|
|
fullStackWorkflow.includes(
|
|
`${secret}: \${{ ${condition} && secrets.${secret} || '' }}`,
|
|
),
|
|
`${secret} must be scoped to only the matrix cells that require it`,
|
|
);
|
|
}
|
|
const historyPublisher = fullStackWorkflow.slice(
|
|
fullStackWorkflow.indexOf(" publish_history:"),
|
|
fullStackWorkflow.indexOf(" pages:"),
|
|
);
|
|
assert.doesNotMatch(historyPublisher, /^\s+cache: pnpm$/m);
|
|
|
|
for (const name of [
|
|
"runner-full-stack-e2e.yml",
|
|
"runner-live-evals.yml",
|
|
"runner-protocol-live-evals.yml",
|
|
]) {
|
|
const workflow = readWorkflow(name);
|
|
const crons = [...workflow.matchAll(/cron:\s*"([^"]+)"/g)].map(
|
|
(match) => match[1],
|
|
);
|
|
assert.equal(crons.length, 1, `${name} must have one schedule`);
|
|
assert.match(crons[0], /^\d{1,2} \d{1,2} \* \* 0$/);
|
|
}
|
|
|
|
const chaosWorkflow = actionPinWorkflows[2];
|
|
const runnerBlock = chaosWorkflow.match(
|
|
/- name: Run Runner fault and replay suites[\s\S]*?run: \|([\s\S]*?)(?=\n\s+- name: Build server test dependencies)/,
|
|
)?.[1];
|
|
const serverBlock = chaosWorkflow.match(
|
|
/- name: Run server finalization and recovery suites[\s\S]*?run: \|([\s\S]*?)(?=\n\s+- name: Upload chaos eval bundle)/,
|
|
)?.[1];
|
|
assert.ok(runnerBlock, "expected Runner chaos test command");
|
|
assert.ok(serverBlock, "expected server chaos test command");
|
|
for (const [base, block] of [
|
|
[path.join(repoRoot, "packages/paperclip-runner"), runnerBlock],
|
|
[path.join(repoRoot, "server"), serverBlock],
|
|
]) {
|
|
const listedTestPaths =
|
|
block.match(/src\/[A-Za-z0-9_./-]+\.test\.ts/g) ?? [];
|
|
assert.ok(listedTestPaths.length > 0, "expected chaos workflow test paths");
|
|
assert.equal(
|
|
new Set(listedTestPaths).size,
|
|
listedTestPaths.length,
|
|
"chaos workflow test paths must be unique",
|
|
);
|
|
for (const testPath of listedTestPaths) {
|
|
assert.ok(
|
|
existsSync(path.join(base, testPath)),
|
|
`chaos workflow test path does not exist: ${testPath}`,
|
|
);
|
|
}
|
|
}
|
|
});
|
|
|
|
|
|
test("direct Grok qualification installs the pinned binary and scopes its API key", () => {
|
|
const workflow = readWorkflow("runner-protocol-live-evals.yml");
|
|
assert.ok(workflow.includes("XAI_API_KEY: ${{ matrix.credentialName == 'XAI_API_KEY' && secrets.XAI_API_KEY || '' }}"));
|
|
assert.ok(workflow.includes("if [ -f packages/grok-acp/install.mjs ]; then"));
|
|
assert.ok(workflow.indexOf("node packages/grok-acp/install.mjs") < workflow.indexOf("pnpm --filter @paperclipai/paperclip-runner deploy --prod"));
|
|
assert.ok(!workflow.includes("secrets.GROK_AUTH_JSON"));
|
|
});
|
|
|
|
test("direct protocol concurrency override only lowers the configured ceiling", () => {
|
|
const workflow = readWorkflow("runner-protocol-live-evals.yml");
|
|
const start = workflow.indexOf(' if [ -n "${REQUESTED_MAX_PARALLEL:-}" ]; then');
|
|
const end = workflow.indexOf(" node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs catalog", start);
|
|
assert.ok(start > 0 && end > start);
|
|
const script = workflow.slice(start, end) + '\nprintf "%s" "$MAX_PARALLEL"\n';
|
|
for (const [requested, expected] of [["", "8"], ["2", "2"], ["8", "8"], ["1", null], ["9", null], ["0", null], ["-1", null], ["2.5", null], ["garbage", null], ["9999999999999999999999", null]]) {
|
|
const result = spawnSync("bash", ["-eu", "-c", script], {
|
|
env: { ...process.env, MAX_PARALLEL: "8", REQUESTED_MAX_PARALLEL: requested }, encoding: "utf8",
|
|
});
|
|
assert.equal(result.status, expected === null ? 1 : 0, requested);
|
|
if (expected !== null) assert.equal(result.stdout, expected);
|
|
}
|
|
});
|