Files
PaperClipAI/tests/runner-e2e/workflow-security.test.ts
T
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00

828 lines
35 KiB
TypeScript

import { readdir, readFile } from "node:fs/promises";
import path from "node:path";
import { describe, expect, it } from "vitest";
const repositoryRoot = path.resolve(import.meta.dirname, "../..");
// PR #13470 uses the code-owner-reviewed default branch for this first-party workflow.
const ordinaryPrTrustedWorkflowRevision = "master";
const fullStackTestNeeds =
/needs:\s*\[\s*authorize,\s*target_lock,\s*catalog,\s*daytona_image,\s*build_runner_artifacts,\s*build_remote_provider_pack,?\s*\]/u;
const buildRunnerNeeds =
/needs:\s*\[\s*authorize,\s*target_lock,\s*catalog,?\s*\]/u;
const buildRemoteProviderPackNeeds =
/needs:\s*\[\s*authorize,\s*target_lock,\s*catalog,\s*daytona_image,\s*build_runner_artifacts,?\s*\]/u;
const everydayOracleImage =
"python@sha256:9d2e5553305c7c7b0097999bb17187c69b921ccd6bc9d40e4bb5ebe652c00285";
describe("public repository paid workflow security", () => {
it("keeps the manual EC2 image build credential-free and pins the authorized target", async () => {
const workflow = await readFile(path.join(repositoryRoot, ".github/workflows/docker-runner-check.yml"), "utf8");
const manual = workflow.slice(workflow.indexOf(" authorize_manual:"));
expect(manual.match(/AWS_CI_TRUSTED_USER_IDS/gu)).toHaveLength(2);
expect(manual.match(/test "\$REPOSITORY_ID" = 1170821064/gu)).toHaveLength(2);
expect(manual).toContain('runs-on: runs-on/fleet=paperclip-public-pr-x64/env=public-ci');
expect(manual).toContain('repos/$REPOSITORY/git/ref/heads/$TARGET_BRANCH');
expect(manual).toContain('ref: ${{ needs.authorize_manual.outputs.target_sha }}');
expect(manual).toContain('SOURCE_SHA: ${{ needs.authorize_manual.outputs.target_sha }}');
expect([...manual.matchAll(/secrets\.([A-Z_]+)/gu)].map(match => match[1])).toEqual(["GITHUB_TOKEN"]);
expect(manual).toContain('pnpm install --resolution-only --ignore-scripts --no-frozen-lockfile');
expect(manual).toContain('PAPERCLIP_RUNNER_LOCK_SHA256=$lock_sha');
expect(manual).toContain('docker logout ghcr.io');
});
it("uses the reviewed master branch for the first-party trusted PR workflow", async () => {
const ordinaryPrWorkflow = await readFile(
path.join(repositoryRoot, ".github/workflows/pr.yml"),
"utf8",
);
const trustedWorkflowCalls = [
...ordinaryPrWorkflow.matchAll(
/^\s+uses:\s+(paperclipai\/paperclip\/\.github\/workflows\/pr-trusted\.yml)@([^\s#]+)$/gmu,
),
];
expect(trustedWorkflowCalls).toHaveLength(1);
expect(trustedWorkflowCalls[0]?.[1]).toBe(
"paperclipai/paperclip/.github/workflows/pr-trusted.yml",
);
expect(trustedWorkflowCalls[0]?.[2]).toBe(
ordinaryPrTrustedWorkflowRevision,
);
});
it("keeps pnpm bootstrap registry telemetry out of trusted workflow setup", async () => {
for (const workflowName of [
"runner-full-stack-e2e.yml",
"pr-trusted.yml",
]) {
const workflow = await readFile(
path.join(repositoryRoot, ".github/workflows", workflowName),
"utf8",
);
const pnpmSetupSteps = workflow
.split(/\n(?= {6}- )/u)
.filter((step) => step.includes("uses: pnpm/action-setup@"));
expect(pnpmSetupSteps, workflowName).toHaveLength(
workflowName === "pr-trusted.yml" ? 8 : 7,
);
for (const step of pnpmSetupSteps) {
expect(step, workflowName).toContain('NPM_CONFIG_AUDIT: "false"');
expect(step, workflowName).toContain('NPM_CONFIG_FUND: "false"');
expect(step, workflowName).toContain(
'NPM_CONFIG_UPDATE_NOTIFIER: "false"',
);
}
for (const variable of [
"NPM_CONFIG_AUDIT",
"NPM_CONFIG_FUND",
"NPM_CONFIG_UPDATE_NOTIFIER",
]) {
expect(
workflow.match(new RegExp(`${variable}:`, "gu")),
workflowName,
).toHaveLength(pnpmSetupSteps.length);
}
}
});
it("installs a modern Node runtime before every trusted pnpm bootstrap", async () => {
const workflows = [
{
name: "runner-full-stack-e2e.yml",
expectedCachedSetupNodeSteps: 4,
},
{
name: "pr-trusted.yml",
// PR #13300 restores shared stores directly without setup-node cache writes.
expectedCachedSetupNodeSteps: 0,
},
];
for (const { name, expectedCachedSetupNodeSteps } of workflows) {
const workflow = await readFile(
path.join(repositoryRoot, ".github/workflows", name),
"utf8",
);
const steps = workflow.split(/\n(?= {6}- )/u);
const pnpmSetupStepIndexes = steps.flatMap((step, index) =>
step.includes("uses: pnpm/action-setup@") ? [index] : [],
);
expect(pnpmSetupStepIndexes, name).toHaveLength(
name === "pr-trusted.yml" ? 8 : 7,
);
for (const pnpmSetupStepIndex of pnpmSetupStepIndexes) {
const pnpmSetupStep = steps[pnpmSetupStepIndex]!;
const nodeBootstrapStep = steps[pnpmSetupStepIndex - 1]!;
expect(nodeBootstrapStep, name).toContain("uses: actions/setup-node@");
expect(nodeBootstrapStep, name).not.toContain("cache: pnpm");
const nodeVersionMatch = nodeBootstrapStep.match(
/^\s*node-version:\s*["']?(\d+)(?:\.(\d+))?/mu,
);
expect(nodeVersionMatch, name).not.toBeNull();
const nodeMajor = Number(nodeVersionMatch![1]);
const nodeMinor = Number(nodeVersionMatch![2] ?? 0);
expect(
nodeMajor > 22 || (nodeMajor === 22 && nodeMinor >= 13),
`${name} must install Node >=22.13 before pnpm/action-setup`,
).toBe(true);
const conditionPattern = /^ {6}- if:\s*(.+)$/mu;
expect(nodeBootstrapStep.match(conditionPattern)?.[1] ?? null).toBe(
pnpmSetupStep.match(conditionPattern)?.[1] ?? null,
);
}
expect(workflow.match(/^\s+cache: pnpm$/gmu) ?? [], name).toHaveLength(
expectedCachedSetupNodeSteps,
);
}
});
it("gates every provider-secret job with stable actor IDs", async () => {
const workflows = await Promise.all(
[
"runner-full-stack-e2e.yml",
"runner-live-evals.yml",
"runner-protocol-live-evals.yml",
"e2e.yml",
].map(async (name) => ({
name,
contents: await readFile(
path.join(repositoryRoot, ".github/workflows", name),
"utf8",
),
})),
);
for (const { name, contents } of workflows) {
const authorize = contents.indexOf(" authorize:");
const reauthorize = contents.indexOf("Reauthorize");
const paidCheckout = contents.indexOf("actions/checkout@", reauthorize);
const providerAccess = contents.search(
/(?:OPENAI|ANTHROPIC|OPENROUTER|DAYTONA)_API_KEY:\s*\$\{\{\s*[^}]*secrets\./,
);
expect(
authorize,
`${name} must have an authorization job`,
).toBeGreaterThan(0);
expect(
reauthorize,
`${name} must reauthorize partial job reruns`,
).toBeGreaterThan(authorize);
expect(
paidCheckout,
`${name} must authorize before checkout`,
).toBeGreaterThan(reauthorize);
expect(
providerAccess,
`${name} must authorize before provider access`,
).toBeGreaterThan(reauthorize);
expect(contents).toContain("RUNNER_E2E_ALLOWED_ACTOR_IDS");
expect(contents).toContain("github.actor_id");
expect(contents).toContain("github.triggering_actor");
expect(contents).toContain("refs/heads/$DEFAULT_BRANCH");
expect(contents).toContain("needs: authorize");
expect(contents).toContain("name: runner-e2e-paid");
expect(contents).not.toMatch(
/^\s*(?:pull_request|pull_request_target|push|workflow_call|workflow_run):/m,
);
const actionReferences = [
...contents.matchAll(/^\s*(?:-\s*)?uses:\s*([^\s#]+)/gm),
].map((match) => match[1]!);
expect(actionReferences.length).toBeGreaterThan(0);
for (const reference of actionReferences) {
expect(reference).toMatch(/^[^@]+@[0-9a-f]{40}$/);
}
}
const fullStack = workflows[0]!.contents;
const paidJob = fullStack.slice(
fullStack.indexOf(" test:"),
fullStack.indexOf(" report:"),
);
const authorizeJob = fullStack.slice(
fullStack.indexOf(" authorize:"),
fullStack.indexOf(" target_lock:"),
);
const targetLockJob = fullStack.slice(
fullStack.indexOf(" target_lock:"),
fullStack.indexOf(" catalog:"),
);
const daytonaImageJob = fullStack.slice(
fullStack.indexOf(" daytona_image:"),
fullStack.indexOf(" build_runner_artifacts:"),
);
expect(daytonaImageJob).toMatch(buildRunnerNeeds);
expect(daytonaImageJob).toContain(
"runs-on: ${{ needs.authorize.outputs.test_runner }}",
);
expect(daytonaImageJob).not.toContain("name: runner-e2e-paid");
expect(daytonaImageJob).not.toMatch(
/(?:(?:OPENAI|ANTHROPIC|OPENROUTER|DAYTONA|XAI)_API_KEY|GROK_AUTH_JSON)/,
);
expect(authorizeJob).toContain(
"aws_runner='runs-on/fleet=paperclip-public-pr-x64/env=public-ci'",
);
expect(authorizeJob).toContain("github_runner='ubuntu-latest'");
expect(authorizeJob).toContain(
"AWS_PAID_RUNNER_ENABLED: ${{ vars.RUNNER_E2E_AWS_ENABLED }}",
);
expect(authorizeJob).toContain(
"playwright_channel: ${{ steps.runner.outputs.playwright_channel }}",
);
expect(authorizeJob).toContain('echo "playwright_channel=chrome"');
expect(authorizeJob).toContain('echo "playwright_channel="');
expect(authorizeJob).toContain(
"Resolve requested repository branch to an immutable commit",
);
expect(authorizeJob).toContain(
"repos/$REPOSITORY/branches/$encoded_branch",
);
expect(authorizeJob).toContain('echo "sha=$target_sha"');
expect(authorizeJob).toContain(
"target_ref: ${{ steps.target.outputs.ref }}",
);
expect(authorizeJob).toContain('echo "ref=refs/heads/$TARGET_BRANCH"');
expect(authorizeJob).not.toContain("actions/checkout@");
expect(authorizeJob).not.toContain("pnpm install");
expect(targetLockJob).toContain("name: Resolve target pnpm lockfile");
expect(targetLockJob).toContain("needs: authorize");
expect(targetLockJob).toContain(
"ref: ${{ needs.authorize.outputs.target_sha }}",
);
expect(targetLockJob).toContain("persist-credentials: false");
expect(targetLockJob).toContain(
"pnpm install --ignore-scripts --no-frozen-lockfile --lockfile-only",
);
expect(targetLockJob).toContain(
"artifact_id: ${{ steps.upload.outputs.artifact-id }}",
);
expect(targetLockJob).toContain("lock_sha256:");
expect(targetLockJob).toContain(
"runner-e2e-target-pnpm-lock-${{ github.run_id }}-${{ github.run_attempt }}",
);
expect(targetLockJob).not.toContain("name: runner-e2e-paid");
expect(targetLockJob).not.toMatch(
/(?:OPENAI|ANTHROPIC|OPENROUTER|DAYTONA)_API_KEY/,
);
expect(paidJob).toContain(
"runs-on: ${{ needs.authorize.outputs.test_runner }}",
);
expect(paidJob).toMatch(fullStackTestNeeds);
expect(paidJob).toContain("name: runner-e2e-paid");
expect(paidJob).toMatch(
/Reauthorize paid execution before provider access[\s\S]*actions\/checkout@[0-9a-f]{40}[\s\S]*persist-credentials: false[\s\S]*Download resolved target lockfile/,
);
expect(paidJob).toContain(
"pnpm exec playwright install --with-deps --only-shell chromium",
);
expect(paidJob).toContain("name: Qualify preinstalled Chrome");
expect(paidJob).toContain(
"if: needs.authorize.outputs.playwright_channel == 'chrome'",
);
expect(paidJob).toContain("google-chrome --version");
expect(paidJob).toContain(
"if: needs.authorize.outputs.playwright_channel != 'chrome'",
);
expect(paidJob).toContain(
"PAPERCLIP_PLAYWRIGHT_CHANNEL: ${{ needs.authorize.outputs.playwright_channel }}",
);
expect(paidJob).not.toContain(
"pnpm exec playwright install --with-deps chromium",
);
const paidInstall = paidJob.indexOf(
"pnpm install --frozen-lockfile --ignore-scripts",
);
const daytonaPluginPreparation = paidJob.indexOf(
"Prepare bundled Daytona plugin without dependency lifecycle scripts",
);
const awsFfmpegInstall = paidJob.indexOf(
"- name: Install Playwright FFmpeg on AWS runner",
);
const hostedChromiumInstall = paidJob.indexOf(
"- name: Install Chromium headless shell on GitHub-hosted fallback",
);
const everydayOraclePreparation = paidJob.indexOf(
"- name: Prepare pinned Python artifact oracle image",
);
const paidExecution = paidJob.indexOf("- name: Run paid cell");
expect(paidInstall).toBeGreaterThan(0);
expect(daytonaPluginPreparation).toBeGreaterThan(paidInstall);
expect(awsFfmpegInstall).toBeGreaterThan(daytonaPluginPreparation);
expect(hostedChromiumInstall).toBeGreaterThan(awsFfmpegInstall);
expect(everydayOraclePreparation).toBeGreaterThan(hostedChromiumInstall);
expect(paidExecution).toBeGreaterThan(awsFfmpegInstall);
expect(paidExecution).toBeGreaterThan(daytonaPluginPreparation);
expect(paidExecution).toBeGreaterThan(everydayOraclePreparation);
const grokPreparation = paidJob.indexOf("- name: Install checksum-verified Grok executable");
expect(grokPreparation).toBeGreaterThan(paidInstall);
expect(paidExecution).toBeGreaterThan(grokPreparation);
expect(paidJob).toContain("if: matrix.environmentId == 'local' && (matrix.profileId == 'runner-acpx-grok' || matrix.profileId == 'runner-acpx-grok-subscription')");
expect(paidJob).toContain("run: node packages/grok-acp/install.mjs");
const everydayOracleStep = paidJob.slice(
everydayOraclePreparation,
paidExecution,
);
expect(everydayOracleStep).toContain(
"if: (matrix.suiteId == 'everyday-workflows' || matrix.suiteId == 'grok-qualification' || matrix.suiteId == 'grok-subscription-qualification') && (matrix.caseId == 'build-revise' || matrix.caseId == 'delegate-feedback' || matrix.caseId == 'agent-review-handoff' || matrix.caseId == 'hire-reuse' || matrix.caseId == 'recover-controller' || matrix.caseId == 'stop-redirect')",
);
expect(everydayOracleStep).toContain(
`oracle_image='${everydayOracleImage}'`,
);
expect(everydayOracleStep).toContain(
"timeout 30s docker version --format '{{.Server.Version}}'",
);
expect(everydayOracleStep).toContain(
'timeout 120s docker pull "$oracle_image"',
);
expect(everydayOracleStep).toContain(
"timeout 30s docker image inspect \"$oracle_image\" --format '{{.Id}}'",
);
const artifactSource = await readFile(
path.join(repositoryRoot, "tests/runner-e2e/everyday-artifact.py"),
"utf8",
);
const artifactImage = artifactSource.match(
/^SANDBOX_IMAGE = '([^']+)'$/mu,
)?.[1];
expect(artifactImage).toBe(everydayOracleImage);
expect(everydayOracleStep).not.toMatch(/secrets\./u);
const awsFfmpegStep = paidJob.slice(
awsFfmpegInstall,
hostedChromiumInstall,
);
expect(awsFfmpegStep).toContain(
"if: needs.authorize.outputs.playwright_channel == 'chrome'",
);
expect(awsFfmpegStep).toContain("pnpm exec playwright install ffmpeg");
expect(awsFfmpegStep).not.toMatch(
/(?:OPENAI|ANTHROPIC|OPENROUTER|DAYTONA)_API_KEY/,
);
const preparedBeforeProviderAccess = paidJob.slice(
daytonaPluginPreparation,
paidExecution,
);
expect(preparedBeforeProviderAccess).toContain(
"if: matrix.environmentId == 'daytona'",
);
expect(preparedBeforeProviderAccess).toContain(
'test -f "$daytona_root/pnpm-lock.yaml"',
);
expect(preparedBeforeProviderAccess).toContain(
"pnpm install --ignore-workspace --frozen-lockfile --ignore-scripts",
);
expect(preparedBeforeProviderAccess).not.toContain("--no-lockfile");
expect(preparedBeforeProviderAccess).toContain(
"node scripts/link-plugin-dev-sdk.mjs",
);
expect(preparedBeforeProviderAccess).toContain(
'"@paperclipai/plugin-daytona"',
);
expect(preparedBeforeProviderAccess).toContain('"@paperclipai/plugin-sdk"');
expect(preparedBeforeProviderAccess).toContain(
'realpath "$daytona_root/node_modules/@paperclipai/plugin-sdk"',
);
expect(preparedBeforeProviderAccess).toContain(
'pnpm --dir "$daytona_root" build',
);
expect(preparedBeforeProviderAccess).not.toContain("secrets.");
expect(preparedBeforeProviderAccess).not.toContain("pnpm rebuild");
expect(paidJob.slice(0, paidExecution)).not.toMatch(
/secrets\.(?:OPENAI|ANTHROPIC|OPENROUTER|DAYTONA)_API_KEY/,
);
expect(authorizeJob).toContain('echo "max_parallel_limit=100"');
expect(fullStack).toContain('[ "$MAX_PARALLEL_LIMIT" -gt 100 ]');
expect(fullStack).toContain("REQUESTED_MAX_PARALLEL: ${{ inputs.max_parallel }}");
expect(fullStack).toContain('[ "$REQUESTED_MAX_PARALLEL" -gt "$MAX_PARALLEL" ]');
expect(fullStack).toContain('[[ "$REQUESTED_MAX_PARALLEL" =~ ^[1-9][0-9]{0,2}$ ]]');
expect(fullStack).toContain(
'[ "$MAX_PARALLEL" -gt "$MAX_PARALLEL_LIMIT" ]',
);
expect(fullStack).toContain(
"group: runner-full-stack-e2e-${{ github.event_name == 'workflow_dispatch' && inputs.target_branch != '' && inputs.target_branch != github.event.repository.default_branch && format('development-{0}', inputs.target_branch) || format('protected-{0}', github.run_id) }}",
);
expect(fullStack).toContain(
"cancel-in-progress: ${{ github.event_name == 'workflow_dispatch' && inputs.target_branch != '' && inputs.target_branch != github.event.repository.default_branch }}",
);
expect(daytonaImageJob).toMatch(
/- if: needs\.catalog\.outputs\.needs_daytona == 'true'\n\s+uses: actions\/checkout@[0-9a-f]{40}/u,
);
for (const stepName of [
"Download resolved target lockfile",
"Restore resolved target lockfile",
]) {
expect(daytonaImageJob).toMatch(
new RegExp(
`- name: ${stepName}\\n\\s+if: needs\\.catalog\\.outputs\\.needs_daytona == 'true'`,
"u",
),
);
}
expect(daytonaImageJob).toMatch(
/- name: No Daytona image needed\n\s+id: local_only\n\s+if: needs\.catalog\.outputs\.needs_daytona != 'true'/u,
);
expect(daytonaImageJob).toContain('echo "source_revision="');
expect(daytonaImageJob).toContain('echo "content_id="');
expect(daytonaImageJob).toContain(
"IMAGE_CACHE: ghcr.io/paperclipai/paperclip-daytona-runner:e2e-buildcache-amd64",
);
expect(daytonaImageJob).toContain(
"TARGET_REF: ${{ needs.authorize.outputs.target_ref }}",
);
expect(daytonaImageJob).toContain(
"DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}",
);
const cacheRead = daytonaImageJob.indexOf(
'--cache-from "type=registry,ref=${IMAGE_CACHE}"',
);
const trustedTargetCheck = daytonaImageJob.indexOf(
'if [ "$TARGET_REF" = "refs/heads/$DEFAULT_BRANCH" ]; then',
);
const cacheWrite = daytonaImageJob.indexOf(
'--cache-to "type=registry,ref=${IMAGE_CACHE},mode=max"',
);
expect(cacheRead).toBeGreaterThan(0);
expect(trustedTargetCheck).toBeGreaterThan(cacheRead);
expect(cacheWrite).toBeGreaterThan(trustedTargetCheck);
expect(daytonaImageJob.slice(trustedTargetCheck, cacheWrite)).not.toContain(
"secrets.",
);
const targetCodeJobs = [
fullStack.slice(
fullStack.indexOf(" catalog:"),
fullStack.indexOf(" daytona_image:"),
),
fullStack.slice(
fullStack.indexOf(" daytona_image:"),
fullStack.indexOf(" build_runner_artifacts:"),
),
fullStack.slice(
fullStack.indexOf(" build_runner_artifacts:"),
fullStack.indexOf(" build_remote_provider_pack:"),
),
fullStack.slice(
fullStack.indexOf(" build_remote_provider_pack:"),
fullStack.indexOf(" test:"),
),
paidJob,
];
for (const targetCodeJob of targetCodeJobs) {
const checkout = targetCodeJob.indexOf("actions/checkout@");
const downloadLock = targetCodeJob.indexOf(
"Download resolved target lockfile",
);
const restoreLock = targetCodeJob.indexOf(
"Restore resolved target lockfile",
);
const setupNode = targetCodeJob.indexOf("actions/setup-node@");
const install = targetCodeJob.indexOf("pnpm install --frozen-lockfile");
expect(checkout).toBeGreaterThan(0);
expect(downloadLock).toBeGreaterThan(checkout);
expect(restoreLock).toBeGreaterThan(downloadLock);
if (setupNode >= 0) {
expect(setupNode).toBeGreaterThan(restoreLock);
}
if (install >= 0) {
expect(install).toBeGreaterThan(restoreLock);
}
expect(targetCodeJob).toContain(
"artifact-ids: ${{ needs.target_lock.outputs.artifact_id }}",
);
expect(targetCodeJob).toContain(
"EXPECTED_LOCK_SHA256: ${{ needs.target_lock.outputs.lock_sha256 }}",
);
}
expect(fullStack.match(/Download resolved target lockfile/g)).toHaveLength(
5,
);
expect(fullStack.match(/Restore resolved target lockfile/g)).toHaveLength(
5,
);
expect(
fullStack.match(
/ref: \$\{\{ needs\.authorize\.outputs\.target_sha \}\}/g,
),
).toHaveLength(6);
expect(fullStack.match(/ref: \$\{\{ github\.sha \}\}/g)).toHaveLength(2);
expect(fullStack.match(/persist-credentials: false/g)).toHaveLength(8);
expect(fullStack).not.toContain("ref: ${{ inputs.target_branch }}");
expect(fullStack).toContain(
"PAPERCLIP_RUNNER_SOURCE_REVISION=${TARGET_SHA}",
);
const reportJob = fullStack.slice(
fullStack.indexOf(" report:"),
fullStack.indexOf(" publish_history:"),
);
const historyJob = fullStack.slice(fullStack.indexOf(" publish_history:"));
expect(reportJob).toContain("ref: ${{ github.sha }}");
expect(reportJob).not.toContain(
"ref: ${{ needs.authorize.outputs.target_sha }}",
);
expect(reportJob).not.toContain("Download resolved target lockfile");
expect(historyJob).toContain("ref: ${{ github.sha }}");
expect(historyJob).not.toContain(
"ref: ${{ needs.authorize.outputs.target_sha }}",
);
expect(historyJob).not.toContain("Download resolved target lockfile");
expect(fullStack).toContain(
"if: always() && !cancelled() && needs.catalog.result == 'success'",
);
for (const targetProvenanceJob of [paidJob, reportJob]) {
expect(targetProvenanceJob).toContain(
"PAPERCLIP_RUNNER_E2E_SOURCE_SHA: ${{ needs.authorize.outputs.target_sha }}",
);
expect(targetProvenanceJob).toContain(
"PAPERCLIP_RUNNER_E2E_SOURCE_REF: ${{ needs.authorize.outputs.target_ref }}",
);
}
for (const [secret, condition] of Object.entries({
OPENAI_API_KEY: "matrix.credentialName == 'OPENAI_API_KEY'",
ANTHROPIC_API_KEY: "matrix.credentialName == 'ANTHROPIC_API_KEY'",
OPENROUTER_API_KEY: "matrix.credentialName == 'OPENROUTER_API_KEY'",
XAI_API_KEY: "matrix.credentialName == 'XAI_API_KEY'",
GROK_AUTH_JSON: "matrix.credentialName == 'GROK_AUTH_JSON'",
DAYTONA_API_KEY: "matrix.environmentId == 'daytona'",
})) {
expect(fullStack).toContain(
`${secret}: \${{ ${condition} && secrets.${secret} || '' }}`,
);
}
});
it("keeps provider credentials inside explicitly gated paid workflows", async () => {
const workflowDirectory = path.join(repositoryRoot, ".github/workflows");
const allowedProviderWorkflows = new Set([
"e2e.yml",
"runner-full-stack-e2e.yml",
"runner-live-evals.yml",
"runner-protocol-live-evals.yml",
]);
const names = (await readdir(workflowDirectory)).filter((name) =>
/\.ya?ml$/.test(name),
);
for (const name of names) {
const contents = await readFile(
path.join(workflowDirectory, name),
"utf8",
);
const providerSecretReferences = [
...contents.matchAll(
/secrets(?:\.(?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|GROK_AUTH_JSON|DAYTONA_API_KEY)\b|\[['"](?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|GROK_AUTH_JSON|DAYTONA_API_KEY)['"]\])/g,
),
];
if (providerSecretReferences.length > 0) {
expect(
allowedProviderWorkflows.has(name),
`${name} must not receive provider credentials`,
).toBe(true);
}
}
});
it("runs paid scheduled campaigns only on Sundays", async () => {
const workflows = await Promise.all(
[
"runner-full-stack-e2e.yml",
"runner-live-evals.yml",
"runner-protocol-live-evals.yml",
].map((name) =>
readFile(path.join(repositoryRoot, ".github/workflows", name), "utf8"),
),
);
for (const workflow of workflows) {
const crons = [...workflow.matchAll(/cron:\s*"([^"]+)"/g)].map(
(match) => match[1]!,
);
expect(crons).toHaveLength(1);
expect(crons[0]).toMatch(/^\d{1,2} \d{1,2} \* \* 0$/);
expect(workflow).toContain("workflow_dispatch:");
}
});
it("builds runner outputs once without provider credentials and verifies them in every paid cell", async () => {
const workflow = await readFile(
path.join(repositoryRoot, ".github/workflows/runner-full-stack-e2e.yml"),
"utf8",
);
const buildJobStart = workflow.indexOf(" build_runner_artifacts:");
const testJobStart = workflow.indexOf(" test:", buildJobStart);
const reportJobStart = workflow.indexOf(" report:", testJobStart);
const buildJob = workflow.slice(buildJobStart, testJobStart);
const testJob = workflow.slice(testJobStart, reportJobStart);
expect(buildJobStart).toBeGreaterThan(0);
expect(testJobStart).toBeGreaterThan(buildJobStart);
expect(buildJob).toMatch(buildRunnerNeeds);
expect(buildJob).toMatch(buildRemoteProviderPackNeeds);
expect(buildJob).not.toContain("environment:");
expect(buildJob).not.toContain("secrets.");
expect(
buildJob.match(/pnpm install --frozen-lockfile --ignore-scripts/g),
).toHaveLength(2);
expect(buildJob).toContain(
"pnpm --filter @paperclipai/paperclip-runner build:typescript",
);
expect(buildJob).toContain(
"pnpm --filter @paperclipai/paperclip-runner build:runner-binaries",
);
expect(buildJob).toContain(
"node packages/paperclip-runner/scripts/build-provider-pack.mjs",
);
expect(buildJob).toContain(
"node packages/paperclip-runner/scripts/materialize-opencode-binary.mjs",
);
expect(buildJob).toContain("runner-e2e-build-bundle.tar.gz.sha256");
expect(buildJob).toContain("runner-e2e-provider-pack.tar.gz.sha256");
expect(buildJob).toContain(
"build_artifact_name: ${{ steps.build_artifact_name.outputs.name }}",
);
expect(buildJob).toContain(
"needs.build_runner_artifacts.outputs.build_artifact_name",
);
expect(buildJob).toContain(
"provider_pack_artifact_name: ${{ steps.provider_pack_artifact_name.outputs.name }}",
);
expect(buildJob).toContain(
"runner-e2e-build-${TARGET_SHA}-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}",
);
expect(buildJob).toContain(
"runner-e2e-provider-pack-${TARGET_SHA}-${GITHUB_RUN_ID}-${GITHUB_RUN_ATTEMPT}",
);
expect(buildJob).not.toContain("runner-e2e-build-${GITHUB_SHA}");
expect(buildJob).not.toContain("runner-e2e-provider-pack-${GITHUB_SHA}");
expect(workflow).toContain("needs_runner_typescript=");
expect(workflow).toContain("needs_native_binaries=");
expect(workflow).toContain("needs_remote_provider_pack=");
expect(testJob).toMatch(fullStackTestNeeds);
expect(testJob).toContain("Download immutable campaign outputs");
expect(
testJob.match(
/if: startsWith\(matrix\.profileId, 'runner-'\) \|\| matrix\.suiteId == 'openrouter-model-breadth'/gu,
),
).toHaveLength(2);
expect(testJob).toContain("Download immutable remote provider pack");
expect(testJob).toContain(
"needs.build_runner_artifacts.outputs.build_artifact_name",
);
expect(testJob).toContain(
"needs.build_remote_provider_pack.outputs.provider_pack_artifact_name",
);
expect(testJob).toContain("sha256sum --check");
expect(testJob.indexOf("sha256sum --check")).toBeLessThan(
testJob.indexOf("tar --extract"),
);
expect(testJob).toContain(
"test -x packages/paperclip-runner/runner/target/debug/paperclip-runnerd",
);
expect(testJob).toContain(".payload.runnerSourceRevision == $revision");
expect(workflow).toContain("Qualify local provider Node interpreter");
expect(testJob).toContain(
"Materialize verified pinned OpenCode executable",
);
expect(testJob).toContain(
"matrix.profileId == 'legacy-opencode' || matrix.profileId == 'runner-opencode' || matrix.suiteId == 'openrouter-model-breadth'",
);
expect(testJob).toContain(
"node packages/paperclip-runner/scripts/materialize-opencode-binary.mjs",
);
expect(testJob).not.toContain("postinstall.mjs");
expect(testJob).not.toContain("pnpm rebuild");
expect(testJob).not.toContain("build:typescript");
expect(testJob).not.toContain("build:runner-binaries");
expect(testJob).not.toContain("build-provider-pack.mjs");
});
it("uses the reviewed AWS Chrome channel without weakening the local executable override", async () => {
const config = await readFile(
path.join(repositoryRoot, "tests/runner-e2e/playwright.config.ts"),
"utf8",
);
expect(config).toContain(
"process.env.PAPERCLIP_PLAYWRIGHT_CHANNEL?.trim()",
);
expect(config).toContain("{ channel: playwrightChannel }");
expect(config).toContain(
"process.env.PAPERCLIP_RUNNER_E2E_CHROMIUM_EXECUTABLE?.trim()",
);
expect(config).toContain(
"PAPERCLIP_PLAYWRIGHT_CHANNEL and PAPERCLIP_RUNNER_E2E_CHROMIUM_EXECUTABLE are mutually exclusive",
);
});
it("resolves each trusted reporting lockfile before frozen install and AWS credentials", async () => {
const workflow = await readFile(
path.join(repositoryRoot, ".github/workflows/runner-full-stack-e2e.yml"), "utf8",
);
for (const jobName of ["report", "publish_history"]) {
const job = workflow.split(`\n ${jobName}:`)[1]!.split(/\n [a-z_]+:/u)[0]!;
const resolve = job.indexOf("pnpm install --lockfile-only --ignore-scripts --no-frozen-lockfile");
const install = job.indexOf("pnpm install --frozen-lockfile");
expect(job).toContain("ref: ${{ github.sha }}");
expect(job).not.toContain("resolved-target-lockfile");
expect(resolve, jobName).toBeGreaterThan(-1);
expect(install, jobName).toBeGreaterThan(resolve);
const credentials = job.indexOf("aws-actions/configure-aws-credentials@");
if (credentials >= 0) expect(install).toBeLessThan(credentials);
}
});
it("binds rerun evidence and Pages artifacts to the exact workflow attempt", async () => {
const workflow = await readFile(
path.join(repositoryRoot, ".github/workflows/runner-full-stack-e2e.yml"),
"utf8",
);
const reportStart = workflow.indexOf(" report:");
const publisherStart = workflow.indexOf(" publish_history:", reportStart);
const report = workflow.slice(reportStart, publisherStart);
const publisher = workflow.slice(publisherStart);
expect(report).toContain("actions: read");
expect(report).toContain(
'"repos/$REPOSITORY/actions/runs/$RUN_ID/jobs?filter=all&per_page=100"',
);
expect(report).toContain("gh api --paginate --slurp");
expect(report).toContain(
'"repos/$REPOSITORY/actions/runs/$RUN_ID/attempts/$attempt"',
);
expect(report).toContain("--jq '{run_attempt, run_started_at}'");
expect(report).toContain("pattern: runner-e2e-${{ github.run_id }}-*-*");
expect(report).not.toContain(
"pattern: runner-e2e-${{ github.run_id }}-${{ github.run_attempt }}-*",
);
expect(report.match(/merge-multiple: false/g)).toHaveLength(2);
expect(report).toContain("Select latest workflow attempt per cell");
expect(report).toContain("tests/runner-e2e/select-rerun-artifacts.ts");
expect(report).toContain(
"PAPERCLIP_RUNNER_E2E_REPORT_ROOT: ${{ github.workspace }}/selected-runner-e2e",
);
expect(report).toContain(
"PAPERCLIP_RUNNER_E2E_HISTORY_PUBLIC_BASE_URL: ${{ vars.RUNNER_E2E_HISTORY_PUBLIC_BASE_URL }}",
);
expect(report).toContain(
"PAPERCLIP_RUNNER_E2E_HISTORY_PREFIX: ${{ vars.RUNNER_E2E_HISTORY_PREFIX || 'runner-e2e' }}",
);
expect(
report.indexOf("Select latest workflow attempt per cell"),
).toBeLessThan(report.indexOf("Collect blob reports"));
expect(publisher).toContain(
'echo "name=github-pages-${{ github.run_id }}-${{ github.run_attempt }}"',
);
expect(publisher).toContain(
"pages_artifact_name: ${{ steps.pages_artifact_name.outputs.name }}",
);
expect(publisher).toContain(
"name: ${{ steps.pages_artifact_name.outputs.name }}",
);
expect(publisher).toContain(
"artifact_name: ${{ needs.publish_history.outputs.pages_artifact_name }}",
);
});
it("uses environment-scoped OIDC for a no-delete history publisher", async () => {
const workflow = await readFile(
path.join(repositoryRoot, ".github/workflows/runner-full-stack-e2e.yml"),
"utf8",
);
const publisher = workflow.slice(workflow.indexOf(" publish_history:"));
expect(publisher).toContain("id-token: write");
expect(publisher).toContain("name: runner-e2e-history");
expect(publisher).toContain("aws-actions/configure-aws-credentials@");
expect(publisher).toContain("RUNNER_E2E_HISTORY_AWS_ROLE_ARN");
expect(publisher).not.toContain("cache: pnpm");
expect(publisher).not.toMatch(/AWS_(?:ACCESS|SECRET)_KEY/);
expect(publisher).not.toMatch(/aws s3 (?:rm|sync .*--delete)/);
expect(workflow).toContain("history_source_ready");
expect(workflow).toContain("Verify normalized history source report");
expect(workflow).not.toContain("sanitized_screenshot=");
expect(workflow).toContain(
"Publish S3 history and Pages bundle with declared screenshots",
);
expect(workflow).toContain(
"Publish trusted summary and declared screenshots to public bundles",
);
expect(publisher).toContain(
"pnpm exec playwright install --with-deps --only-shell chromium",
);
expect(workflow).toContain(
"Package pruned dashboard with declared screenshots for GitHub Pages",
);
expect(workflow).toContain("path: runner-e2e-merged-report/pages");
expect(workflow).toContain(
"Publish latest dashboard with declared screenshots",
);
expect(workflow).not.toContain("dashboard_ready");
expect(workflow).not.toContain("Publish latest screenshot dashboard");
expect(
workflow.indexOf("pnpm test:e2e:runner:history:publish"),
).toBeLessThan(workflow.indexOf("actions/upload-pages-artifact@"));
});
});