fix(evals): support explicit Grok qualification workflows (#13878)

## Thinking Path

> - Paperclip manages AI agents and their provider connections.
> - Product E2E checks real tasks through the browser, server, and
runner.
> - Grok qualification needs separate API-key and subscription evidence.
> - Product subscription tests and direct Grok protocol evals need
explicit credential delivery.
> - This change supplies each credential only to its selected profile
and prepares the pinned binary.
> - Maintainer authorization and protected-environment gates remain
required.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13882.

The Grok feature branch has a manual subscription qualification profile.
The trusted master workflow must admit its selected credential and
prepare the same verified binary and artifact verifier as the API
profile. Direct protocol evals also need the selected xAI key and pinned
Grok binary. These prerequisites do not register or schedule the new
profiles on master.

## What Changed

- Deliver `GROK_AUTH_JSON` from the protected paid environment only when
the selected profile requests that credential.
- Install the checksum-verified Grok binary for the local subscription
profile.
- Prepare the pinned artifact verifier for the manual subscription
suite.
- Extend workflow security assertions to cover the new credential and
profile.
- Add the ACPX Grok credential mapping to the trusted-master catalog,
then deliver only the selected `XAI_API_KEY` to direct protocol cells
and install the target’s checksum-verified Grok binary before packaging.
- Allow a direct-protocol concurrency override from two cases up to the
existing configured ceiling; it can only lower concurrency.
- Document the Grok protocol workflow and its API-only credential
boundary.
- Render missing LLM usage and cost as Unavailable, and label partial
observations with coverage. Preserve raw records, grades, and actual
zero costs.
- Preserve measured campaign source metadata during report regeneration
instead of inheriting the renderer checkout or CI event; skip empty
legacy source records when recovering older provenance.

## Verification

- Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54
reported checks successful, two intentional skips, Greptile 5/5, and
zero unresolved review threads. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35890978288).

- After merging current master, all 17 workflow security/image tests and
23 catalog/workflow policy tests passed. The trusted catalog also
generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per
shard. The new policy tests execute the concurrency guard against valid,
out-of-range, and malformed values.
- The Grok branch separately passed 450 Product harness unit tests,
including private company credential staging, cleanup, and
token-fragment redaction.
- The fresh-login native subscription smoke passed three repetitions of
tool execution, session resume, restrictive permissions, and cleanup.
These are setup evidence; full subscription Product qualification
remains pending.
- All 72 focused report/billing/history/catalog tests and the Product
harness typecheck passed for the report-display change. The initial
sandbox run could not open the tsx IPC socket; the permitted rerun
passed. A zero-provider-call replay of the actual 16-cell Grok campaign
preserved all result records, grades, timing, and source provenance
while correcting missing usage labels.
- Review the thirteen-file diff. Provider credentials still enter only
the selected paid-test step; default-branch, numeric-actor, and
environment restrictions are unchanged.

## Risks

This admits a refreshable subscription credential to explicitly selected
trusted tests. Store it only in `runner-e2e-paid`, use a test login, and
remove it after qualification. Unselected profiles receive an empty
value. Pull requests cannot trigger the paid workflow. This PR changes
no fleet admission or actor allowlist.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
DottaandPaperclip authored and GitHub committed 2026-09-23 17:05:52 -05:00
1 parent 4b8ec588f3
commit b648d8cdda
13 files changed
+201 -23

No files matched your search

+3 -2
View File
@@ -891,7 +891,7 @@ jobs:
run: node packages/paperclip-runner/scripts/materialize-opencode-binary.mjs run: node packages/paperclip-runner/scripts/materialize-opencode-binary.mjs
- name: Install checksum-verified Grok executable - name: Install checksum-verified Grok executable
if: matrix.environmentId == 'local' && matrix.profileId == 'runner-acpx-grok' if: matrix.environmentId == 'local' && (matrix.profileId == 'runner-acpx-grok' || matrix.profileId == 'runner-acpx-grok-subscription')
run: node packages/grok-acp/install.mjs run: node packages/grok-acp/install.mjs
- name: Download immutable campaign outputs - name: Download immutable campaign outputs
@@ -1044,7 +1044,7 @@ jobs:
NODE NODE
- name: Prepare pinned Python artifact oracle image - name: Prepare pinned Python artifact oracle image
if: (matrix.suiteId == 'everyday-workflows' || matrix.suiteId == 'grok-qualification') && (matrix.caseId == 'build-revise' || matrix.caseId == 'delegate-feedback' || matrix.caseId == 'agent-review-handoff' || matrix.caseId == 'hire-reuse' || matrix.caseId == 'recover-controller' || matrix.caseId == 'stop-redirect') if: (matrix.suiteId == 'everyday-workflows' || matrix.suiteId == 'grok-qualification' || matrix.suiteId == 'grok-subscription-qualification') && (matrix.caseId == 'build-revise' || matrix.caseId == 'delegate-feedback' || matrix.caseId == 'agent-review-handoff' || matrix.caseId == 'hire-reuse' || matrix.caseId == 'recover-controller' || matrix.caseId == 'stop-redirect')
run: | run: |
set -euo pipefail set -euo pipefail
oracle_image='python@sha256:9d2e5553305c7c7b0097999bb17187c69b921ccd6bc9d40e4bb5ebe652c00285' oracle_image='python@sha256:9d2e5553305c7c7b0097999bb17187c69b921ccd6bc9d40e4bb5ebe652c00285'
@@ -1058,6 +1058,7 @@ jobs:
ANTHROPIC_API_KEY: ${{ matrix.credentialName == 'ANTHROPIC_API_KEY' && secrets.ANTHROPIC_API_KEY || '' }} ANTHROPIC_API_KEY: ${{ matrix.credentialName == 'ANTHROPIC_API_KEY' && secrets.ANTHROPIC_API_KEY || '' }}
OPENROUTER_API_KEY: ${{ matrix.credentialName == 'OPENROUTER_API_KEY' && secrets.OPENROUTER_API_KEY || '' }} OPENROUTER_API_KEY: ${{ matrix.credentialName == 'OPENROUTER_API_KEY' && secrets.OPENROUTER_API_KEY || '' }}
XAI_API_KEY: ${{ matrix.credentialName == 'XAI_API_KEY' && secrets.XAI_API_KEY || '' }} XAI_API_KEY: ${{ matrix.credentialName == 'XAI_API_KEY' && secrets.XAI_API_KEY || '' }}
GROK_AUTH_JSON: ${{ matrix.credentialName == 'GROK_AUTH_JSON' && secrets.GROK_AUTH_JSON || '' }}
DAYTONA_API_KEY: ${{ matrix.environmentId == 'daytona' && secrets.DAYTONA_API_KEY || '' }} DAYTONA_API_KEY: ${{ matrix.environmentId == 'daytona' && secrets.DAYTONA_API_KEY || '' }}
PAPERCLIP_E2E_DAYTONA_IMAGE: ${{ needs.daytona_image.outputs.image }} PAPERCLIP_E2E_DAYTONA_IMAGE: ${{ needs.daytona_image.outputs.image }}
PAPERCLIP_RUNNER_REMOTE_PROVIDER_PACK_PATH: ${{ github.workspace }}/packages/paperclip-runner/provider-pack PAPERCLIP_RUNNER_REMOTE_PROVIDER_PACK_PATH: ${{ github.workspace }}/packages/paperclip-runner/provider-pack
@@ -18,6 +18,10 @@ on:
type: string type: string
default: "all" default: "all"
required: false required: false
max_parallel:
description: "Optional lower concurrency (at least 2, no higher than the configured campaign limit)"
type: string
required: false
max_infrastructure_retries: max_infrastructure_retries:
description: "Automatic retries only for explicitly retryable infrastructure failures (0-3)" description: "Automatic retries only for explicitly retryable infrastructure failures (0-3)"
type: number type: number
@@ -250,6 +254,7 @@ jobs:
PAPERCLIP_PROTOCOL_EVALS_SHA: ${{ needs.authorize.outputs.evals_sha }} PAPERCLIP_PROTOCOL_EVALS_SHA: ${{ needs.authorize.outputs.evals_sha }}
MAX_PARALLEL: ${{ vars.RUNNER_E2E_MAX_PARALLEL || needs.authorize.outputs.max_parallel_default }} MAX_PARALLEL: ${{ vars.RUNNER_E2E_MAX_PARALLEL || needs.authorize.outputs.max_parallel_default }}
MAX_PARALLEL_LIMIT: ${{ needs.authorize.outputs.max_parallel_limit }} MAX_PARALLEL_LIMIT: ${{ needs.authorize.outputs.max_parallel_limit }}
REQUESTED_MAX_PARALLEL: ${{ inputs.max_parallel }}
ROSTERS: ${{ inputs.rosters || 'all' }} ROSTERS: ${{ inputs.rosters || 'all' }}
run: | run: |
set -euo pipefail set -euo pipefail
@@ -257,6 +262,13 @@ jobs:
echo "RUNNER_E2E_MAX_PARALLEL must be an integer from 2 through $MAX_PARALLEL_LIMIT for the two-shard direct suite." >&2 echo "RUNNER_E2E_MAX_PARALLEL must be an integer from 2 through $MAX_PARALLEL_LIMIT for the two-shard direct suite." >&2
exit 1 exit 1
fi fi
if [ -n "${REQUESTED_MAX_PARALLEL:-}" ]; then
if ! [[ "$REQUESTED_MAX_PARALLEL" =~ ^[1-9][0-9]{0,2}$ ]] || [ "$REQUESTED_MAX_PARALLEL" -lt 2 ] || [ "$REQUESTED_MAX_PARALLEL" -gt "$MAX_PARALLEL" ]; then
echo "max_parallel must be an integer from 2 through the configured campaign limit." >&2
exit 1
fi
MAX_PARALLEL="$REQUESTED_MAX_PARALLEL"
fi
node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs catalog \ node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs catalog \
--evals-root .paperclip-evals \ --evals-root .paperclip-evals \
--rosters "$ROSTERS" \ --rosters "$ROSTERS" \
@@ -332,6 +344,12 @@ jobs:
- name: Materialize the pinned OpenCode executable before packaging - name: Materialize the pinned OpenCode executable before packaging
run: node packages/paperclip-runner/scripts/materialize-opencode-binary.mjs run: node packages/paperclip-runner/scripts/materialize-opencode-binary.mjs
- name: Materialize the pinned Grok executable when the target includes it
run: |
if [ -f packages/grok-acp/install.mjs ]; then
node packages/grok-acp/install.mjs
fi
- name: Build runner CLI, daemon, and canonical attempt viewer - name: Build runner CLI, daemon, and canonical attempt viewer
run: | run: |
set -euo pipefail set -euo pipefail
@@ -523,6 +541,7 @@ jobs:
OPENAI_API_KEY: ${{ matrix.credentialName == 'OPENAI_API_KEY' && secrets.OPENAI_API_KEY || '' }} OPENAI_API_KEY: ${{ matrix.credentialName == 'OPENAI_API_KEY' && secrets.OPENAI_API_KEY || '' }}
ANTHROPIC_API_KEY: ${{ matrix.credentialName == 'ANTHROPIC_API_KEY' && secrets.ANTHROPIC_API_KEY || '' }} ANTHROPIC_API_KEY: ${{ matrix.credentialName == 'ANTHROPIC_API_KEY' && secrets.ANTHROPIC_API_KEY || '' }}
OPENROUTER_API_KEY: ${{ matrix.credentialName == 'OPENROUTER_API_KEY' && secrets.OPENROUTER_API_KEY || '' }} OPENROUTER_API_KEY: ${{ matrix.credentialName == 'OPENROUTER_API_KEY' && secrets.OPENROUTER_API_KEY || '' }}
XAI_API_KEY: ${{ matrix.credentialName == 'XAI_API_KEY' && secrets.XAI_API_KEY || '' }}
PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID }} PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_PROFILE_ID }}
PAPERCLIP_CLAUDE_MANAGED_AGENT_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_AGENT_ID }} PAPERCLIP_CLAUDE_MANAGED_AGENT_ID: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_AGENT_ID }}
PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION }} PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION: ${{ vars.PAPERCLIP_CLAUDE_MANAGED_AGENT_VERSION }}
@@ -205,6 +205,7 @@ The paid jobs read only the credential selected for each roster:
- `OPENAI_API_KEY` for native Codex and ACPX Codex; - `OPENAI_API_KEY` for native Codex and ACPX Codex;
- `ANTHROPIC_API_KEY` for ACPX Claude and Claude Managed; - `ANTHROPIC_API_KEY` for ACPX Claude and Claude Managed;
- `OPENROUTER_API_KEY` for native OpenCode and ACPX Pi; - `OPENROUTER_API_KEY` for native OpenCode and ACPX Pi;
- `XAI_API_KEY` for an explicitly selected ACPX Grok roster;
- short-lived GitHub OIDC workload identity for AWS AgentCore. - short-lived GitHub OIDC workload identity for AWS AgentCore.
Claude Managed also requires the four nonsecret Claude Managed also requires the four nonsecret
@@ -308,3 +309,14 @@ pnpm --filter @paperclipai/paperclip-runner \
--campaign-id gha-1-1 \ --campaign-id gha-1-1 \
--output /tmp/runner-protocol-eval-catalog.json --output /tmp/runner-protocol-eval-catalog.json
``` ```
For Grok qualification, select `protocol-live-acpx-grok` and its exact eval
revision. Set `max_parallel: 2` when the API key has a low custom rate limit;
this admits one case per shard. The override can only lower the configured
campaign ceiling and cannot change the key's provider limits. Retain any
rate-limited attempts as failures. The build installs the target's pinned,
checksum-verified Grok binary before packaging the portable runtime. Targets
without the Grok package keep their existing build behavior. This protocol
workflow uses the explicitly selected API key; subscription credentials are
not delivered by it.
@@ -64,6 +64,7 @@ export function credentialForConfig(config) {
if (config.provider === "acpx") { if (config.provider === "acpx") {
if (config.acpxAgent === "pi") return "OPENROUTER_API_KEY"; if (config.acpxAgent === "pi") return "OPENROUTER_API_KEY";
if (config.acpxAgent === "claude") return "ANTHROPIC_API_KEY"; if (config.acpxAgent === "claude") return "ANTHROPIC_API_KEY";
if (config.acpxAgent === "grok") return "XAI_API_KEY";
if (config.acpxAgent === "codex") return "OPENAI_API_KEY"; if (config.acpxAgent === "codex") return "OPENAI_API_KEY";
} }
throw new Error( throw new Error(
@@ -85,6 +85,7 @@ async function fixture() {
test("maps every qualified driver to one explicit credential boundary", () => { test("maps every qualified driver to one explicit credential boundary", () => {
assert.equal(credentialForConfig({ provider: "codex" }), "OPENAI_API_KEY"); assert.equal(credentialForConfig({ provider: "codex" }), "OPENAI_API_KEY");
assert.equal(credentialForConfig({ provider: "acpx", acpxAgent: "grok" }), "XAI_API_KEY");
assert.equal( assert.equal(
credentialForConfig({ provider: "opencode" }), credentialForConfig({ provider: "opencode" }),
"OPENROUTER_API_KEY", "OPENROUTER_API_KEY",
@@ -1,4 +1,5 @@
import assert from "node:assert/strict"; import assert from "node:assert/strict";
import { spawnSync } from "node:child_process";
import { existsSync, readdirSync, readFileSync } from "node:fs"; import { existsSync, readdirSync, readFileSync } from "node:fs";
import path from "node:path"; import path from "node:path";
import { fileURLToPath } from "node:url"; import { fileURLToPath } from "node:url";
@@ -320,7 +321,7 @@ test("Runner eval workflows pin actions and gate paid live execution", () => {
]; ];
const paidWorkflowNameSet = new Set(paidWorkflowNames); const paidWorkflowNameSet = new Set(paidWorkflowNames);
const providerSecretReference = const providerSecretReference =
/secrets(?:\.(?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|DAYTONA_API_KEY)\b|\[['"](?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|DAYTONA_API_KEY)['"]\])/g; /secrets(?:\.(?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|GROK_AUTH_JSON|DAYTONA_API_KEY)\b|\[['"](?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|GROK_AUTH_JSON|DAYTONA_API_KEY)['"]\])/g;
for (const name of readdirSync(path.join(repoRoot, ".github/workflows"))) { for (const name of readdirSync(path.join(repoRoot, ".github/workflows"))) {
if (!/\.ya?ml$/.test(name)) continue; if (!/\.ya?ml$/.test(name)) continue;
const workflow = readWorkflow(name); const workflow = readWorkflow(name);
@@ -430,3 +431,27 @@ test("Runner eval workflows pin actions and gate paid live execution", () => {
} }
} }
}); });
test("direct Grok qualification installs the pinned binary and scopes its API key", () => {
const workflow = readWorkflow("runner-protocol-live-evals.yml");
assert.ok(workflow.includes("XAI_API_KEY: ${{ matrix.credentialName == 'XAI_API_KEY' && secrets.XAI_API_KEY || '' }}"));
assert.ok(workflow.includes("if [ -f packages/grok-acp/install.mjs ]; then"));
assert.ok(workflow.indexOf("node packages/grok-acp/install.mjs") < workflow.indexOf("pnpm --filter @paperclipai/paperclip-runner deploy --prod"));
assert.ok(!workflow.includes("secrets.GROK_AUTH_JSON"));
});
test("direct protocol concurrency override only lowers the configured ceiling", () => {
const workflow = readWorkflow("runner-protocol-live-evals.yml");
const start = workflow.indexOf(' if [ -n "${REQUESTED_MAX_PARALLEL:-}" ]; then');
const end = workflow.indexOf(" node packages/paperclip-runner/scripts/runner-protocol-eval-campaign.mjs catalog", start);
assert.ok(start > 0 && end > start);
const script = workflow.slice(start, end) + '\nprintf "%s" "$MAX_PARALLEL"\n';
for (const [requested, expected] of [["", "8"], ["2", "2"], ["8", "8"], ["1", null], ["9", null], ["0", null], ["-1", null], ["2.5", null], ["garbage", null], ["9999999999999999999999", null]]) {
const result = spawnSync("bash", ["-eu", "-c", script], {
env: { ...process.env, MAX_PARALLEL: "8", REQUESTED_MAX_PARALLEL: requested }, encoding: "utf8",
});
assert.equal(result.status, expected === null ? 1 : 0, requested);
if (expected !== null) assert.equal(result.stdout, expected);
}
});
+8
View File
@@ -40,6 +40,14 @@ export interface CampaignBillingSummary {
testsWithCompleteBilling: number; testsWithCompleteBilling: number;
} }
/** Render observed subtotals without presenting absent measurements as zero. */
export function billingCoverageLabel(value: string, covered: number, total: number): string {
if (!Number.isInteger(covered) || !Number.isInteger(total) || covered <= 0 || total <= 0 || covered > total) {
return "Unavailable";
}
return covered < total ? `${value} (partial: ${covered}/${total} runs)` : value;
}
function record(value: unknown): Record<string, unknown> { function record(value: unknown): Record<string, unknown> {
return value && typeof value === "object" && !Array.isArray(value) return value && typeof value === "object" && !Array.isArray(value)
? (value as Record<string, unknown>) ? (value as Record<string, unknown>)
+11
View File
@@ -21,6 +21,7 @@ interface PublishedResult extends RunnerE2EResult {
interface PublishedCampaign { interface PublishedCampaign {
schema?: string; schema?: string;
campaignId?: string; campaignId?: string;
source?: RunnerE2ECampaign["source"];
generatedAt: string; generatedAt: string;
expected: string[]; expected: string[];
results: PublishedResult[]; results: PublishedResult[];
@@ -94,6 +95,16 @@ export async function regenerateRunnerDashboard(input: {
expected, expected,
results, results,
}); });
// A rendering refresh must not inherit the renderer's checkout or CI event.
const retainedSource = results.find((result) =>
result.source?.sha || result.source?.ref || result.source?.workflowRunUrl,
)?.source;
campaign.source = normalized.source ?? {
sha: retainedSource?.sha ?? null,
ref: retainedSource?.ref ?? null,
workflowRunUrl: retainedSource?.workflowRunUrl ?? null,
eventName: null,
};
const history = const history =
input.historyFile === null input.historyFile === null
? undefined ? undefined
+20 -13
View File
@@ -5,6 +5,7 @@ import {
type ReportExecution, type ReportExecution,
} from "./report-catalog.js"; } from "./report-catalog.js";
import type { import type {
RunnerE2EBillingSummary,
RunnerE2ECampaign, RunnerE2ECampaign,
RunnerE2EHistoryIndex, RunnerE2EHistoryIndex,
RunnerE2EResult, RunnerE2EResult,
@@ -12,6 +13,7 @@ import type {
} from "./types.js"; } from "./types.js";
import { import {
aggregateCampaignBilling, aggregateCampaignBilling,
billingCoverageLabel,
summarizeExecutionBilling, summarizeExecutionBilling,
} from "./billing.js"; } from "./billing.js";
@@ -57,8 +59,13 @@ function durationLabel(durationMs: number) {
return `${Math.floor(seconds / 60)}m ${seconds % 60}s`; return `${Math.floor(seconds / 60)}m ${seconds % 60}s`;
} }
function tokenLabel(value: number) { function tokenLabel(value: number, llm?: RunnerE2EBillingSummary["llm"]) {
return new Intl.NumberFormat("en-US").format(value); const formatted = new Intl.NumberFormat("en-US").format(value);
return llm ? billingCoverageLabel(formatted, llm.runsWithTokenUsage, llm.runCount) : formatted;
}
function llmCostLabel(value: number, llm: RunnerE2EBillingSummary["llm"]) {
return billingCoverageLabel(usdLabel(value), llm.runsWithReportedCost, llm.runCount);
} }
function usdLabel(value: number | null) { function usdLabel(value: number | null) {
@@ -264,7 +271,7 @@ function renderCase(
data-gallery-runtime="${html(entry?.result.runtimeMode ?? execution.profile.expectedRuntimeMode)}" data-gallery-runtime="${html(entry?.result.runtimeMode ?? execution.profile.expectedRuntimeMode)}"
data-gallery-status="${html(label)}" data-gallery-status="${html(label)}"
data-gallery-duration="${html(entry ? durationLabel(entry.result.durationMs) : "Not run")}" data-gallery-duration="${html(entry ? durationLabel(entry.result.durationMs) : "Not run")}"
data-gallery-tokens="${html(billing ? `${tokenLabel(billing.llm.inputTokens)} in · ${tokenLabel(billing.llm.outputTokens)} out` : "Unavailable")}" data-gallery-tokens="${html(billing ? `${tokenLabel(billing.llm.inputTokens, billing.llm)} in · ${tokenLabel(billing.llm.outputTokens, billing.llm)} out` : "Unavailable")}"
data-gallery-matchers="${html(matcherResults.length > 0 ? `${passedMatchers}/${matcherResults.length} matchers passed${notReached ? ` · ${notReached} not reached` : ""}` : "No matchers recorded")}" data-gallery-matchers="${html(matcherResults.length > 0 ? `${passedMatchers}/${matcherResults.length} matchers passed${notReached ? ` · ${notReached} not reached` : ""}` : "No matchers recorded")}"
aria-label="Open ${html(item.label)} for ${html(execution.id)} in gallery" aria-label="Open ${html(item.label)} for ${html(execution.id)} in gallery"
> >
@@ -276,8 +283,8 @@ function renderCase(
: ""; : "";
const billingStrip = billing const billingStrip = billing
? `<div class="billing-strip" aria-label="Billing for ${html(execution.id)}"> ? `<div class="billing-strip" aria-label="Billing for ${html(execution.id)}">
<div><span>Tokens</span><strong>${html(tokenLabel(billing.llm.inputTokens))} in · ${html(tokenLabel(billing.llm.outputTokens))} out</strong><small>${html(tokenLabel(billing.llm.cachedInputTokens))} cached · ${billing.llm.runsWithTokenUsage}/${billing.llm.runCount} runs covered</small></div> <div><span>Tokens</span><strong>${html(tokenLabel(billing.llm.inputTokens, billing.llm))} in · ${html(tokenLabel(billing.llm.outputTokens, billing.llm))} out</strong><small>${html(tokenLabel(billing.llm.cachedInputTokens, billing.llm))} cached · ${billing.llm.runsWithTokenUsage}/${billing.llm.runCount} runs covered</small></div>
<div><span>LLM spend</span><strong>${billing.llm.runsWithReportedCost > 0 ? html(usdLabel(billing.reportedCostUsd)) : html(billing.llm.costStatus)}</strong><small>${billing.llm.runsWithReportedCost}/${billing.llm.runCount} runs provider-priced</small></div> <div><span>LLM spend</span><strong>${html(llmCostLabel(billing.reportedCostUsd, billing.llm))}</strong><small>${billing.llm.runsWithReportedCost}/${billing.llm.runCount} runs provider-priced</small></div>
<div><span>Execution</span><strong>${billing.runtime.estimatedListCostUsd === undefined ? html(billing.runtime.costStatus === "not_metered" ? "Local · not metered" : "Cost unavailable") : `${html(usdLabel(billing.runtime.estimatedListCostUsd))} est.`}</strong><small>${html(durationLabel(billing.runtime.agentRunDurationMs))} agent${billing.runtime.leaseDurationMs === null ? "" : ` · ${html(durationLabel(billing.runtime.leaseDurationMs))} lease`}</small></div> <div><span>Execution</span><strong>${billing.runtime.estimatedListCostUsd === undefined ? html(billing.runtime.costStatus === "not_metered" ? "Local · not metered" : "Cost unavailable") : `${html(usdLabel(billing.runtime.estimatedListCostUsd))} est.`}</strong><small>${html(durationLabel(billing.runtime.agentRunDurationMs))} agent${billing.runtime.leaseDurationMs === null ? "" : ` · ${html(durationLabel(billing.runtime.leaseDurationMs))} lease`}</small></div>
</div>` </div>`
: ""; : "";
@@ -516,8 +523,8 @@ function renderHistory(history: RunnerE2EHistoryIndex | undefined) {
<td><a href="${html(campaign.publicUrl)}">${html(campaign.campaignId)}</a><small>${html(new Date(campaign.generatedAt).toLocaleString("en-US", { timeZone: "UTC" }))} UTC</small></td> <td><a href="${html(campaign.publicUrl)}">${html(campaign.campaignId)}</a><small>${html(new Date(campaign.generatedAt).toLocaleString("en-US", { timeZone: "UTC" }))} UTC</small></td>
<td>${sha ? `<code>${html(sha.slice(0, 10))}</code>` : "Unknown"}<small>${html(campaign.source.ref ?? "unknown ref")}</small></td> <td>${sha ? `<code>${html(sha.slice(0, 10))}</code>` : "Unknown"}<small>${html(campaign.source.ref ?? "unknown ref")}</small></td>
<td><span class="status history-${status}">${status}</span><small>${campaign.passed}/${campaign.selected} passed${campaign.incomplete ? ` · ${campaign.incomplete} incomplete` : ""} · ${campaign.complete ? "complete" : "partial"}</small></td> <td><span class="status history-${status}">${status}</span><small>${campaign.passed}/${campaign.selected} passed${campaign.incomplete ? ` · ${campaign.incomplete} incomplete` : ""} · ${campaign.complete ? "complete" : "partial"}</small></td>
<td>${html(tokenLabel(campaign.billing.llm.inputTokens))} / ${html(tokenLabel(campaign.billing.llm.outputTokens))}<small>input / output · ${html(tokenLabel(campaign.billing.llm.cachedInputTokens))} cached</small></td> <td>${html(tokenLabel(campaign.billing.llm.inputTokens, campaign.billing.llm))} / ${html(tokenLabel(campaign.billing.llm.outputTokens, campaign.billing.llm))}<small>input / output · ${html(tokenLabel(campaign.billing.llm.cachedInputTokens, campaign.billing.llm))} cached</small></td>
<td>${html(usdLabel(campaign.billing.reportedLlmCostUsd))}<small>${html(usdLabel(campaign.billing.estimatedRuntimeCostUsd))} runtime estimate</small></td> <td>${html(llmCostLabel(campaign.billing.reportedLlmCostUsd, campaign.billing.llm))}<small>${html(usdLabel(campaign.billing.estimatedRuntimeCostUsd))} runtime estimate</small></td>
<td>${html(durationLabel(campaign.billing.agentRunDurationMs))}<small>${html(durationLabel(campaign.billing.leaseDurationMs))} lease</small></td> <td>${html(durationLabel(campaign.billing.agentRunDurationMs))}<small>${html(durationLabel(campaign.billing.leaseDurationMs))} lease</small></td>
</tr>`; </tr>`;
}) })
@@ -607,8 +614,8 @@ function renderSuiteMatrix(input: {
const summaryHtml = summary const summaryHtml = summary
? `<div class="suite-summary" aria-label="${html(suite.label)} current campaign summary"> ? `<div class="suite-summary" aria-label="${html(suite.label)} current campaign summary">
<div><span>Pass rate</span><strong>${summary.selected > 0 ? ((summary.passed / summary.selected) * 100).toFixed(1) : "0.0"}%</strong><small>${summary.passed}/${summary.selected} passed${summary.incomplete ? ` · ${summary.incomplete} incomplete` : ""}</small></div> <div><span>Pass rate</span><strong>${summary.selected > 0 ? ((summary.passed / summary.selected) * 100).toFixed(1) : "0.0"}%</strong><small>${summary.passed}/${summary.selected} passed${summary.incomplete ? ` · ${summary.incomplete} incomplete` : ""}</small></div>
<div><span>Tokens</span><strong>${html(tokenLabel(summary.billing.llm.totalTokens))}</strong><small>${html(tokenLabel(summary.billing.llm.inputTokens))} input · ${html(tokenLabel(summary.billing.llm.outputTokens))} output</small></div> <div><span>Tokens</span><strong>${html(tokenLabel(summary.billing.llm.totalTokens, summary.billing.llm))}</strong><small>${html(tokenLabel(summary.billing.llm.inputTokens, summary.billing.llm))} input · ${html(tokenLabel(summary.billing.llm.outputTokens, summary.billing.llm))} output</small></div>
<div><span>Cost</span><strong>${html(usdLabel(summary.billing.observedAndEstimatedCostUsd))}</strong><small>reported LLM + runtime${summary.billing.judge ? " + judge" : ""} estimate</small></div> <div><span>Known spend</span><strong>${html(summary.billing.llm.runsWithReportedCost > 0 || summary.billing.estimatedRuntimeCostUsd > 0 || summary.billing.judge ? `${usdLabel(summary.billing.observedAndEstimatedCostUsd)}${summary.billing.testsWithCompleteBilling < summary.billing.testCount ? " (partial)" : ""}` : "Unavailable")}</strong><small>reported LLM + runtime${summary.billing.judge ? " + judge" : ""} estimate</small></div>
<div><span>Agent time</span><strong>${html(durationLabel(summary.billing.agentRunDurationMs))}</strong><small>${html(durationLabel(summary.billing.leaseDurationMs))} lease</small></div> <div><span>Agent time</span><strong>${html(durationLabel(summary.billing.agentRunDurationMs))}</strong><small>${html(durationLabel(summary.billing.leaseDurationMs))} lease</small></div>
<div><span>Execution</span><strong>${summary.executed}/${summary.selected}</strong><small>${summary.retries} retries · cleanup ${summary.cleanupPassed ? "passed" : "failed"}</small></div> <div><span>Execution</span><strong>${summary.executed}/${summary.selected}</strong><small>${summary.retries} retries · cleanup ${summary.cleanupPassed ? "passed" : "failed"}</small></div>
</div>` </div>`
@@ -1118,10 +1125,10 @@ export function renderRunnerE2EDashboard(input: RunnerDashboardInput) {
</header> </header>
${publicSummaryImageHref ? `<figure class="public-summary"><img src="${html(publicSummaryImageHref)}" alt="Runner E2E campaign status summary"></figure>` : ""} ${publicSummaryImageHref ? `<figure class="public-summary"><img src="${html(publicSummaryImageHref)}" alt="Runner E2E campaign status summary"></figure>` : ""}
<section class="billing-overview" aria-label="Campaign billing summary"> <section class="billing-overview" aria-label="Campaign billing summary">
<div class="billing-metric"><strong>${html(tokenLabel(campaignBilling.llm.inputTokens))}</strong><span>Input tokens</span></div> <div class="billing-metric"><strong>${html(tokenLabel(campaignBilling.llm.inputTokens, campaignBilling.llm))}</strong><span>Input tokens</span></div>
<div class="billing-metric"><strong>${html(tokenLabel(campaignBilling.llm.outputTokens))}</strong><span>Output tokens</span></div> <div class="billing-metric"><strong>${html(tokenLabel(campaignBilling.llm.outputTokens, campaignBilling.llm))}</strong><span>Output tokens</span></div>
<div class="billing-metric"><strong>${html(tokenLabel(campaignBilling.llm.cachedInputTokens))}</strong><span>Cached tokens</span></div> <div class="billing-metric"><strong>${html(tokenLabel(campaignBilling.llm.cachedInputTokens, campaignBilling.llm))}</strong><span>Cached tokens</span></div>
<div class="billing-metric"><strong>${html(usdLabel(campaignBilling.reportedLlmCostUsd))}</strong><span>LLM reported subtotal</span></div> <div class="billing-metric"><strong>${html(llmCostLabel(campaignBilling.reportedLlmCostUsd, campaignBilling.llm))}</strong><span>LLM reported subtotal</span></div>
<div class="billing-metric"><strong>${html(usdLabel(campaignBilling.estimatedRuntimeCostUsd))}</strong><span>Daytona list estimate</span></div> <div class="billing-metric"><strong>${html(usdLabel(campaignBilling.estimatedRuntimeCostUsd))}</strong><span>Daytona list estimate</span></div>
${campaignBilling.judge ? `<div class="billing-metric"><strong>${html(usdLabel(campaignBilling.judge.estimatedCostUsd))}</strong><span>Judge estimate · ${campaignBilling.judge.attempts} attempts · ${campaignBilling.judge.attemptsWithUnknownUsage} unknown usage</span><small>${html(tokenLabel(campaignBilling.judge.inputTokens))} in / ${html(tokenLabel(campaignBilling.judge.outputTokens))} out · $${campaignBilling.judge.reservedCostUsd.toFixed(6)} reserved</small></div>` : ""} ${campaignBilling.judge ? `<div class="billing-metric"><strong>${html(usdLabel(campaignBilling.judge.estimatedCostUsd))}</strong><span>Judge estimate · ${campaignBilling.judge.attempts} attempts · ${campaignBilling.judge.attemptsWithUnknownUsage} unknown usage</span><small>${html(tokenLabel(campaignBilling.judge.inputTokens))} in / ${html(tokenLabel(campaignBilling.judge.outputTokens))} out · $${campaignBilling.judge.reservedCostUsd.toFixed(6)} reserved</small></div>` : ""}
<div class="billing-metric"><strong>${html(durationLabel(campaignBilling.agentRunDurationMs))}</strong><span>Agent execution time</span></div> <div class="billing-metric"><strong>${html(durationLabel(campaignBilling.agentRunDurationMs))}</strong><span>Agent execution time</span></div>
+49
View File
@@ -67,6 +67,55 @@ function result(execution: MatrixExecution, status: "passed" | "failed") {
} }
describe("runner E2E campaign history", () => { describe("runner E2E campaign history", () => {
it("does not let empty legacy provenance hide a later measured source", async () => {
const root = await mkdtemp(path.join(os.tmpdir(), "runner-legacy-source-"));
temporaryDirectories.push(root);
const source = {
sha: "measured-source-sha", ref: "refs/heads/measured-source",
workflowRunUrl: "https://example.test/actions/runs/1",
};
const results = [
{ ...result(runnerMatrix[0]!, "passed"), schema: "paperclip.runner-e2e.result/v1" },
{ ...result(runnerMatrix[1]!, "passed"), source },
];
await writeFile(path.join(root, "normalized-results.json"), JSON.stringify({
campaignId: "legacy-source", generatedAt: results[0]!.finishedAt,
expected: results.map((entry) => entry.executionId), results,
}));
vi.stubEnv("PAPERCLIP_RUNNER_E2E_SOURCE_SHA", "renderer-sha");
vi.stubEnv("GITHUB_EVENT_NAME", "push");
await regenerateRunnerDashboard({ bundle: root });
const regenerated = JSON.parse(await readFile(path.join(root, "normalized-results.json"), "utf8"));
expect(regenerated.source).toEqual({ ...source, eventName: null });
expect(regenerated.results[0].source).toEqual({ sha: null, ref: null, workflowRunUrl: null });
expect(regenerated.passed).toBe(2);
});
it.each(["workflow_dispatch", null, undefined])("preserves retained source during regeneration (event %s)", async (eventName) => {
const root = await mkdtemp(path.join(os.tmpdir(), "runner-retained-source-"));
temporaryDirectories.push(root);
const execution = runnerMatrix[0]!;
const source = {
sha: "measured-source-sha", ref: "refs/heads/measured-source",
workflowRunUrl: "https://example.test/actions/runs/1",
};
const retainedResult = { ...result(execution, "passed"), source };
const campaign = buildRunnerCampaign({
campaignId: "retained-source", generatedAt: retainedResult.finishedAt,
expected: [execution.id], results: [retainedResult],
});
const published = { ...campaign, source: eventName === undefined ? undefined : { ...source, eventName } };
await writeFile(path.join(root, "normalized-results.json"), JSON.stringify(published));
vi.stubEnv("PAPERCLIP_RUNNER_E2E_SOURCE_SHA", "renderer-sha");
vi.stubEnv("PAPERCLIP_RUNNER_E2E_SOURCE_REF", "refs/heads/renderer");
vi.stubEnv("GITHUB_EVENT_NAME", "push");
await regenerateRunnerDashboard({ bundle: root });
const regenerated = JSON.parse(await readFile(path.join(root, "normalized-results.json"), "utf8"));
expect(regenerated.source).toEqual({ ...source, eventName: eventName ?? null });
expect(regenerated.results).toEqual(campaign.results.map((entry) => ({ ...entry, evidenceValid: true, evidenceErrors: [] })));
expect(regenerated.generatedAt).toBe(campaign.generatedAt);
});
it("retains incomplete journeys in campaign, suite and history without marking them green", () => { it("retains incomplete journeys in campaign, suite and history without marking them green", () => {
const execution = runnerMatrix.find(e => e.suite.id === "first-task")!; const execution = runnerMatrix.find(e => e.suite.id === "first-task")!;
const incomplete: RunnerE2EResult = { const incomplete: RunnerE2EResult = {
+43
View File
@@ -18,6 +18,49 @@ afterEach(async () => {
}); });
describe("runner E2E report aggregation", () => { describe("runner E2E report aggregation", () => {
it.each([
{ name: "missing", usage: null, runIds: ["run-1"], tokens: "Unavailable", cost: "Unavailable", htmlTokens: "Unavailable" },
{ name: "partial", usage: { runs: [{ usage: { inputTokens: 1250, outputTokens: 75, cachedInputTokens: 500, costUsd: 0.0125 } }, { usage: null }] }, runIds: ["run-1", "run-2"], tokens: "1250 input / 75 output / 500 cached (partial: 1/2 runs)", cost: "$0.012500 (partial: 1/2 runs)", htmlTokens: "1,250 (partial: 1/2 runs)" },
{ name: "reported", usage: { inputTokens: 1250, outputTokens: 75, cachedInputTokens: 500, costUsd: 0.0125 }, runIds: ["run-1"], tokens: "1250 input / 75 output / 500 cached", cost: "$0.012500", htmlTokens: "1,250" },
{ name: "reported zero cost", usage: { inputTokens: 1, outputTokens: 0, costUsd: 0 }, runIds: ["run-1"], tokens: "1 input / 0 output / 0 cached", cost: "$0.000000", htmlTokens: "1" },
])("renders $name usage with its actual coverage", async ({ usage, runIds, tokens, cost, htmlTokens }) => {
const root = await mkdtemp(path.join(os.tmpdir(), "runner-billing-report-"));
cleanupDirectories.push(root);
const executionId = "core-compatibility.runner-codex.local.message-marker";
const directory = path.join(root, "attempt-1");
await mkdir(directory);
const result: RunnerE2EResult = {
schema: "paperclip.runner-e2e.result/v2", executionId, suiteId: "core-compatibility",
attempt: 1, status: "passed", profileId: "runner-codex", environmentId: "local",
caseId: "message-marker", provider: "codex", model: "fixture-model", runtimeMode: "native",
startedAt: "2026-09-23T00:00:00Z", finishedAt: "2026-09-23T00:00:01Z", durationMs: 1000,
cleanup: "passed", runIds, usage,
};
const raw = JSON.stringify(result);
await writeFile(path.join(directory, "result.json"), raw);
await writeFile(path.join(directory, "final-state.png"), "fake-png");
await writeFile(path.join(directory, "evidence-manifest.json"), JSON.stringify({ files: ["final-state.png"], leaks: [], missing: [] }));
const output = path.join(root, "merged");
await execFileAsync(process.execPath, [path.join(repositoryRoot, "cli/node_modules/tsx/dist/cli.mjs"), path.join(repositoryRoot, "tests/runner-e2e/report.ts")], {
cwd: repositoryRoot,
env: { ...process.env, PAPERCLIP_RUNNER_E2E_REPORT_ROOT: root, PAPERCLIP_RUNNER_E2E_REPORT_OUT: output, PAPERCLIP_RUNNER_E2E_EXPECTED_IDS: JSON.stringify([executionId]) },
});
const markdown = await readFile(path.join(output, "summary.md"), "utf8");
expect(markdown.split("\n").find((line) => line.startsWith("Tokens: "))).toBe(`Tokens: ${tokens}`);
expect(markdown.split("\n").find((line) => line.startsWith("Provider-reported LLM cost: "))).toBe(`Provider-reported LLM cost: ${cost}`);
const dashboard = await readFile(path.join(output, "index.html"), "utf8");
expect(dashboard).toContain(`<strong>${htmlTokens}</strong><span>Input tokens</span>`);
if (usage === null) {
expect(markdown).not.toContain("0/0 | $0.000000");
expect(dashboard).toContain("<strong>Unavailable</strong><span>LLM reported subtotal</span>");
expect(dashboard).toContain("<span>Known spend</span><strong>Unavailable</strong>");
}
expect(await readFile(path.join(directory, "result.json"), "utf8")).toBe(raw);
const normalized = JSON.parse(await readFile(path.join(output, "normalized-results.json"), "utf8"));
expect(normalized.passed).toBe(1);
expect(normalized.results[0].usage).toEqual(usage);
});
it("keeps interrupted journeys incomplete unless their evidence is invalid", async () => { it("keeps interrupted journeys incomplete unless their evidence is invalid", async () => {
const root = await mkdtemp(path.join(os.tmpdir(), "runner-incomplete-report-")); const root = await mkdtemp(path.join(os.tmpdir(), "runner-incomplete-report-"));
cleanupDirectories.push(root); cleanupDirectories.push(root);
+4 -4
View File
@@ -8,7 +8,7 @@ import {
writeFile, writeFile,
} from "node:fs/promises"; } from "node:fs/promises";
import { runnerMatrix } from "./catalog.js"; import { runnerMatrix } from "./catalog.js";
import { summarizeExecutionBilling } from "./billing.js"; import { billingCoverageLabel, summarizeExecutionBilling } from "./billing.js";
import { renderRunnerE2EDashboard } from "./dashboard.js"; import { renderRunnerE2EDashboard } from "./dashboard.js";
import { import {
buildRunnerCampaign, buildRunnerCampaign,
@@ -368,9 +368,9 @@ async function main() {
"", "",
] ]
: []), : []),
`Tokens: ${billing.llm.inputTokens} input / ${billing.llm.outputTokens} output / ${billing.llm.cachedInputTokens} cached`, `Tokens: ${billingCoverageLabel(`${billing.llm.inputTokens} input / ${billing.llm.outputTokens} output / ${billing.llm.cachedInputTokens} cached`, billing.llm.runsWithTokenUsage, billing.llm.runCount)}`,
"", "",
`Provider-reported LLM cost: $${billing.reportedLlmCostUsd.toFixed(6)} (${billing.llm.runsWithReportedCost}/${billing.llm.runCount} runs priced)`, `Provider-reported LLM cost: ${billingCoverageLabel(`$${billing.reportedLlmCostUsd.toFixed(6)}`, billing.llm.runsWithReportedCost, billing.llm.runCount)}`,
"", "",
`Estimated Daytona list-price runtime cost: $${billing.estimatedRuntimeCostUsd.toFixed(6)}`, `Estimated Daytona list-price runtime cost: $${billing.estimatedRuntimeCostUsd.toFixed(6)}`,
...(billing.judge ? [`Estimated judge cost: ${billing.judge.estimatedCostUsd === null ? "unknown" : `$${billing.judge.estimatedCostUsd.toFixed(6)}`}; ${billing.judge.attempts} attempts; ${billing.judge.attemptsWithUnknownUsage} with unknown usage; $${billing.judge.reservedCostUsd.toFixed(6)} reserved`] : []), ...(billing.judge ? [`Estimated judge cost: ${billing.judge.estimatedCostUsd === null ? "unknown" : `$${billing.judge.estimatedCostUsd.toFixed(6)}`}; ${billing.judge.attempts} attempts; ${billing.judge.attemptsWithUnknownUsage} with unknown usage; $${billing.judge.reservedCostUsd.toFixed(6)} reserved`] : []),
@@ -385,7 +385,7 @@ async function main() {
const cell = publicCampaignUrl const cell = publicCampaignUrl
? `[${resolved.executionId}](${publicCampaignUrl}#execution-${encodeURIComponent(resolved.executionId)})` ? `[${resolved.executionId}](${publicCampaignUrl}#execution-${encodeURIComponent(resolved.executionId)})`
: resolved.executionId; : resolved.executionId;
return `| ${cell} | ${resolved.attempt} | ${entry.valid ? "pass" : "fail"} | ${resolved.runtimeMode} | ${Math.round(resolved.durationMs / 1000)}s | ${cellBilling.llm.inputTokens}/${cellBilling.llm.outputTokens} | $${cellBilling.reportedCostUsd.toFixed(6)} (${cellBilling.llm.costStatus}) | ${runtimeCost === undefined ? cellBilling.runtime.costStatus : `$${runtimeCost.toFixed(6)} est.`} | ${detail} |`; return `| ${cell} | ${resolved.attempt} | ${entry.valid ? "pass" : "fail"} | ${resolved.runtimeMode} | ${Math.round(resolved.durationMs / 1000)}s | ${billingCoverageLabel(`${cellBilling.llm.inputTokens}/${cellBilling.llm.outputTokens}`, cellBilling.llm.runsWithTokenUsage, cellBilling.llm.runCount)} | ${billingCoverageLabel(`$${cellBilling.reportedCostUsd.toFixed(6)}`, cellBilling.llm.runsWithReportedCost, cellBilling.llm.runCount)} | ${runtimeCost === undefined ? cellBilling.runtime.costStatus : `$${runtimeCost.toFixed(6)} est.`} | ${detail} |`;
}), }),
"", "",
]; ];
+4 -3
View File
@@ -305,7 +305,7 @@ describe("public repository paid workflow security", () => {
const grokPreparation = paidJob.indexOf("- name: Install checksum-verified Grok executable"); const grokPreparation = paidJob.indexOf("- name: Install checksum-verified Grok executable");
expect(grokPreparation).toBeGreaterThan(paidInstall); expect(grokPreparation).toBeGreaterThan(paidInstall);
expect(paidExecution).toBeGreaterThan(grokPreparation); expect(paidExecution).toBeGreaterThan(grokPreparation);
expect(paidJob).toContain("if: matrix.environmentId == 'local' && matrix.profileId == 'runner-acpx-grok'"); expect(paidJob).toContain("if: matrix.environmentId == 'local' && (matrix.profileId == 'runner-acpx-grok' || matrix.profileId == 'runner-acpx-grok-subscription')");
expect(paidJob).toContain("run: node packages/grok-acp/install.mjs"); expect(paidJob).toContain("run: node packages/grok-acp/install.mjs");
const everydayOracleStep = paidJob.slice( const everydayOracleStep = paidJob.slice(
@@ -313,7 +313,7 @@ describe("public repository paid workflow security", () => {
paidExecution, paidExecution,
); );
expect(everydayOracleStep).toContain( expect(everydayOracleStep).toContain(
"if: (matrix.suiteId == 'everyday-workflows' || matrix.suiteId == 'grok-qualification') && (matrix.caseId == 'build-revise' || matrix.caseId == 'delegate-feedback' || matrix.caseId == 'agent-review-handoff' || matrix.caseId == 'hire-reuse' || matrix.caseId == 'recover-controller' || matrix.caseId == 'stop-redirect')", "if: (matrix.suiteId == 'everyday-workflows' || matrix.suiteId == 'grok-qualification' || matrix.suiteId == 'grok-subscription-qualification') && (matrix.caseId == 'build-revise' || matrix.caseId == 'delegate-feedback' || matrix.caseId == 'agent-review-handoff' || matrix.caseId == 'hire-reuse' || matrix.caseId == 'recover-controller' || matrix.caseId == 'stop-redirect')",
); );
expect(everydayOracleStep).toContain( expect(everydayOracleStep).toContain(
`oracle_image='${everydayOracleImage}'`, `oracle_image='${everydayOracleImage}'`,
@@ -530,6 +530,7 @@ describe("public repository paid workflow security", () => {
ANTHROPIC_API_KEY: "matrix.credentialName == 'ANTHROPIC_API_KEY'", ANTHROPIC_API_KEY: "matrix.credentialName == 'ANTHROPIC_API_KEY'",
OPENROUTER_API_KEY: "matrix.credentialName == 'OPENROUTER_API_KEY'", OPENROUTER_API_KEY: "matrix.credentialName == 'OPENROUTER_API_KEY'",
XAI_API_KEY: "matrix.credentialName == 'XAI_API_KEY'", XAI_API_KEY: "matrix.credentialName == 'XAI_API_KEY'",
GROK_AUTH_JSON: "matrix.credentialName == 'GROK_AUTH_JSON'",
DAYTONA_API_KEY: "matrix.environmentId == 'daytona'", DAYTONA_API_KEY: "matrix.environmentId == 'daytona'",
})) { })) {
expect(fullStack).toContain( expect(fullStack).toContain(
@@ -557,7 +558,7 @@ describe("public repository paid workflow security", () => {
); );
const providerSecretReferences = [ const providerSecretReferences = [
...contents.matchAll( ...contents.matchAll(
/secrets(?:\.(?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|DAYTONA_API_KEY)\b|\[['"](?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|DAYTONA_API_KEY)['"]\])/g, /secrets(?:\.(?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|GROK_AUTH_JSON|DAYTONA_API_KEY)\b|\[['"](?:OPENAI_API_KEY|ANTHROPIC_API_KEY|OPENROUTER_API_KEY|XAI_API_KEY|GROK_AUTH_JSON|DAYTONA_API_KEY)['"]\])/g,
), ),
]; ];
if (providerSecretReferences.length > 0) { if (providerSecretReferences.length > 0) {