mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip manages AI agents and their provider connections. > - Product E2E checks real tasks through the browser, server, and runner. > - Grok qualification needs separate API-key and subscription evidence. > - Product subscription tests and direct Grok protocol evals need explicit credential delivery. > - This change supplies each credential only to its selected profile and prepares the pinned binary. > - Maintainer authorization and protected-environment gates remain required. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13882. The Grok feature branch has a manual subscription qualification profile. The trusted master workflow must admit its selected credential and prepare the same verified binary and artifact verifier as the API profile. Direct protocol evals also need the selected xAI key and pinned Grok binary. These prerequisites do not register or schedule the new profiles on master. ## What Changed - Deliver `GROK_AUTH_JSON` from the protected paid environment only when the selected profile requests that credential. - Install the checksum-verified Grok binary for the local subscription profile. - Prepare the pinned artifact verifier for the manual subscription suite. - Extend workflow security assertions to cover the new credential and profile. - Add the ACPX Grok credential mapping to the trusted-master catalog, then deliver only the selected `XAI_API_KEY` to direct protocol cells and install the target’s checksum-verified Grok binary before packaging. - Allow a direct-protocol concurrency override from two cases up to the existing configured ceiling; it can only lower concurrency. - Document the Grok protocol workflow and its API-only credential boundary. - Render missing LLM usage and cost as Unavailable, and label partial observations with coverage. Preserve raw records, grades, and actual zero costs. - Preserve measured campaign source metadata during report regeneration instead of inheriting the renderer checkout or CI event; skip empty legacy source records when recovering older provenance. ## Verification - Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54 reported checks successful, two intentional skips, Greptile 5/5, and zero unresolved review threads. [CI run](https://github.com/paperclipai/paperclip/actions/runs/35890978288). - After merging current master, all 17 workflow security/image tests and 23 catalog/workflow policy tests passed. The trusted catalog also generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per shard. The new policy tests execute the concurrency guard against valid, out-of-range, and malformed values. - The Grok branch separately passed 450 Product harness unit tests, including private company credential staging, cleanup, and token-fragment redaction. - The fresh-login native subscription smoke passed three repetitions of tool execution, session resume, restrictive permissions, and cleanup. These are setup evidence; full subscription Product qualification remains pending. - All 72 focused report/billing/history/catalog tests and the Product harness typecheck passed for the report-display change. The initial sandbox run could not open the tsx IPC socket; the permitted rerun passed. A zero-provider-call replay of the actual 16-cell Grok campaign preserved all result records, grades, timing, and source provenance while correcting missing usage labels. - Review the thirteen-file diff. Provider credentials still enter only the selected paid-test step; default-branch, numeric-actor, and environment restrictions are unchanged. ## Risks This admits a refreshable subscription credential to explicitly selected trusted tests. Store it only in `runner-e2e-paid`, use a test login, and remove it after qualification. Unselected profiles receive an empty value. Pull requests cannot trigger the paid workflow. This PR changes no fleet admission or actor allowlist. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
413 lines
16 KiB
TypeScript
413 lines
16 KiB
TypeScript
import { discoverReportCatalog } from "./report-catalog.js";
|
|
import path from "node:path";
|
|
import {
|
|
copyFile,
|
|
mkdir,
|
|
readFile,
|
|
readdir,
|
|
writeFile,
|
|
} from "node:fs/promises";
|
|
import { runnerMatrix } from "./catalog.js";
|
|
import { billingCoverageLabel, summarizeExecutionBilling } from "./billing.js";
|
|
import { renderRunnerE2EDashboard } from "./dashboard.js";
|
|
import {
|
|
buildRunnerCampaign,
|
|
canonicalExecutionId,
|
|
upgradeRunnerResult,
|
|
} from "./history.js";
|
|
import { resolveRunnerE2ESource } from "./source.js";
|
|
import { runnerE2ESummaryLinks } from "./summary-links.js";
|
|
import type { RunnerE2EResult } from "./types.js";
|
|
|
|
const repositoryRoot = path.resolve(import.meta.dirname, "../..");
|
|
|
|
interface EvidenceManifest {
|
|
files?: string[];
|
|
leaks?: Array<{ file: string; reason: string }>;
|
|
missing?: string[];
|
|
}
|
|
|
|
interface AggregatedResult {
|
|
result: RunnerE2EResult;
|
|
evidence: EvidenceManifest | null;
|
|
directory: string;
|
|
valid: boolean;
|
|
errors: string[];
|
|
}
|
|
|
|
async function walk(root: string): Promise<string[]> {
|
|
const entries = await readdir(root, { withFileTypes: true }).catch(() => []);
|
|
const files: string[] = [];
|
|
for (const entry of entries) {
|
|
const full = path.join(root, entry.name);
|
|
if (entry.isDirectory()) files.push(...(await walk(full)));
|
|
else if (entry.isFile()) files.push(full);
|
|
}
|
|
return files;
|
|
}
|
|
|
|
function xml(value: unknown) {
|
|
return String(value ?? "")
|
|
.replaceAll("&", "&")
|
|
.replaceAll("<", "<")
|
|
.replaceAll(">", ">")
|
|
.replaceAll('"', """)
|
|
.replaceAll("'", "'");
|
|
}
|
|
|
|
function validate(result: RunnerE2EResult, evidence: EvidenceManifest | null) {
|
|
const errors: string[] = [];
|
|
if (result.cleanup !== "passed") errors.push(`cleanup=${result.cleanup}`);
|
|
if (!evidence) errors.push("evidence manifest missing");
|
|
if (evidence?.leaks?.length)
|
|
errors.push(`secret leaks=${evidence.leaks.length}`);
|
|
if (evidence?.missing?.length)
|
|
errors.push(`missing evidence=${evidence.missing.join(",")}`);
|
|
if (
|
|
result.status === "passed" &&
|
|
!evidence?.files?.includes("final-state.png")
|
|
) {
|
|
errors.push("passing final-state screenshot missing");
|
|
}
|
|
return errors;
|
|
}
|
|
|
|
function safeEvidenceRelative(relative: string) {
|
|
if (path.isAbsolute(relative)) return null;
|
|
const segments = relative.split(/[\\/]/).filter(Boolean);
|
|
if (segments.length === 0 || segments.some((segment) => segment === ".."))
|
|
return null;
|
|
return segments;
|
|
}
|
|
|
|
async function stageDashboardEvidence(
|
|
selected: readonly AggregatedResult[],
|
|
output: string,
|
|
) {
|
|
const staged = new Map<string, { baseHref: string; files: string[] }>();
|
|
for (const entry of selected) {
|
|
const baseSegments = [
|
|
"evidence",
|
|
entry.result.executionId,
|
|
`attempt-${entry.result.attempt}`,
|
|
];
|
|
const copied: string[] = [];
|
|
for (const relative of entry.evidence?.files ?? []) {
|
|
const segments = safeEvidenceRelative(relative);
|
|
if (!segments) continue;
|
|
const source = path.join(entry.directory, ...segments);
|
|
const destination = path.join(output, ...baseSegments, ...segments);
|
|
const didCopy = await mkdir(path.dirname(destination), {
|
|
recursive: true,
|
|
})
|
|
.then(() => copyFile(source, destination))
|
|
.then(() => true)
|
|
.catch(() => false);
|
|
if (didCopy) copied.push(segments.join("/"));
|
|
}
|
|
// Playwright renames attachment files with a content hash, while runner
|
|
// results retain the stable screenshot basename used by the dashboard and
|
|
// history publisher. Materialize each declared screenshot under that
|
|
// basename when its hashed attachment is present in the evidence manifest.
|
|
// The source is still restricted to manifest-listed files, so this cannot
|
|
// expand the evidence set beyond what the test recorded.
|
|
for (const screenshot of entry.result.screenshots ?? []) {
|
|
if (copied.includes(screenshot.file)) continue;
|
|
const stem = screenshot.file.replace(/\.png$/i, "");
|
|
const escapedStem = stem.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
|
const hashedAttachment = new RegExp(
|
|
`^${escapedStem}-[0-9a-f]{8,128}\\.png$`,
|
|
"i",
|
|
);
|
|
const candidates = (entry.evidence?.files ?? []).filter((relative) => {
|
|
const basename = path.posix.basename(relative);
|
|
return (
|
|
basename === screenshot.file ||
|
|
hashedAttachment.test(basename)
|
|
);
|
|
});
|
|
if (candidates.length !== 1) continue;
|
|
const segments = safeEvidenceRelative(candidates[0]!);
|
|
if (!segments) continue;
|
|
const source = path.join(entry.directory, ...segments);
|
|
const destination = path.join(output, ...baseSegments, screenshot.file);
|
|
const didCopy = await mkdir(path.dirname(destination), {
|
|
recursive: true,
|
|
})
|
|
.then(() => copyFile(source, destination))
|
|
.then(() => true)
|
|
.catch(() => false);
|
|
if (didCopy) copied.push(screenshot.file);
|
|
}
|
|
staged.set(entry.result.executionId, {
|
|
baseHref: baseSegments.join("/"),
|
|
files: copied,
|
|
});
|
|
}
|
|
return staged;
|
|
}
|
|
|
|
async function stageDashboardBrandAssets(output: string) {
|
|
const assets = path.join(output, "assets");
|
|
await mkdir(assets, { recursive: true });
|
|
await Promise.all([
|
|
copyFile(
|
|
path.join(repositoryRoot, "ui/public/favicon-32x32.png"),
|
|
path.join(assets, "favicon-32x32.png"),
|
|
),
|
|
copyFile(
|
|
path.join(repositoryRoot, "ui/public/fonts/InterVariable.woff2"),
|
|
path.join(assets, "InterVariable.woff2"),
|
|
),
|
|
]);
|
|
}
|
|
|
|
async function main() {
|
|
const root = path.resolve(
|
|
process.env.PAPERCLIP_RUNNER_E2E_REPORT_ROOT ?? "tests/runner-e2e/results",
|
|
);
|
|
const output = path.resolve(
|
|
process.env.PAPERCLIP_RUNNER_E2E_REPORT_OUT ?? path.join(root, "merged"),
|
|
);
|
|
const expectedInput = JSON.parse(
|
|
process.env.PAPERCLIP_RUNNER_E2E_EXPECTED_IDS ?? "[]",
|
|
) as string[];
|
|
if (
|
|
!Array.isArray(expectedInput) ||
|
|
expectedInput.some((value) => typeof value !== "string")
|
|
) {
|
|
throw new Error(
|
|
"PAPERCLIP_RUNNER_E2E_EXPECTED_IDS must be a JSON string array",
|
|
);
|
|
}
|
|
const expected = expectedInput.map(canonicalExecutionId);
|
|
if (expected.length === 0 || new Set(expected).size !== expected.length) {
|
|
throw new Error(
|
|
"PAPERCLIP_RUNNER_E2E_EXPECTED_IDS must contain unique selected executions",
|
|
);
|
|
}
|
|
const resultFiles = (await walk(root)).filter(
|
|
(file) => path.basename(file) === "result.json",
|
|
);
|
|
const candidates = new Map<string, AggregatedResult[]>();
|
|
for (const resultFile of resultFiles) {
|
|
const parsed = JSON.parse(
|
|
await readFile(resultFile, "utf8"),
|
|
) as RunnerE2EResult;
|
|
if (
|
|
parsed.schema !== "paperclip.runner-e2e.result/v1" &&
|
|
parsed.schema !== "paperclip.runner-e2e.result/v2"
|
|
)
|
|
continue;
|
|
const result = upgradeRunnerResult(parsed);
|
|
const directory = path.dirname(resultFile);
|
|
const evidence = await readFile(
|
|
path.join(directory, "evidence-manifest.json"),
|
|
"utf8",
|
|
)
|
|
.then((value) => JSON.parse(value) as EvidenceManifest)
|
|
.catch(() => null);
|
|
const errors = validate(result, evidence);
|
|
const entry = {
|
|
result,
|
|
evidence,
|
|
directory,
|
|
valid: result.status === "passed" && errors.length === 0,
|
|
errors,
|
|
};
|
|
candidates.set(result.executionId, [
|
|
...(candidates.get(result.executionId) ?? []),
|
|
entry,
|
|
]);
|
|
}
|
|
|
|
const reportCatalog = discoverReportCatalog({
|
|
catalog: runnerMatrix,
|
|
expected,
|
|
results: [...candidates.values()].flatMap((entries) =>
|
|
entries.map((entry) => entry.result),
|
|
),
|
|
});
|
|
const selected: AggregatedResult[] = [];
|
|
for (const executionId of expected) {
|
|
const attempts = (candidates.get(executionId) ?? []).sort((left, right) => {
|
|
// A report root can contain artifacts from multiple local campaigns (or
|
|
// reruns downloaded from CI). Prefer retained evidence that actually
|
|
// satisfies the campaign contract, then the newest execution. Attempt
|
|
// numbers are only meaningful inside one campaign and must not let an
|
|
// older failed attempt shadow a later successful rerun.
|
|
if (left.valid !== right.valid) return left.valid ? -1 : 1;
|
|
return (
|
|
Date.parse(right.result.finishedAt) -
|
|
Date.parse(left.result.finishedAt) ||
|
|
right.result.attempt - left.result.attempt
|
|
);
|
|
});
|
|
if (attempts.length === 0) {
|
|
const now = new Date().toISOString();
|
|
const execution = reportCatalog.find(
|
|
(candidate) => candidate.id === executionId,
|
|
)!;
|
|
const missing: RunnerE2EResult = {
|
|
schema: "paperclip.runner-e2e.result/v2",
|
|
executionId,
|
|
suiteId: execution?.suite.id ?? "core-compatibility",
|
|
suiteDefinitionHash: execution?.suiteDefinitionHash,
|
|
attempt: 0,
|
|
status: "failed",
|
|
failureClass: "permanent_infrastructure",
|
|
error: "No result artifact was uploaded",
|
|
profileId: execution?.profile.id ?? "unknown",
|
|
environmentId: execution?.environment.id ?? "local",
|
|
caseId: execution?.task.id ?? "unknown",
|
|
provider: execution?.profile.provider ?? "unknown",
|
|
model: execution?.profile.model ?? "unknown",
|
|
runtimeMode: execution?.profile.expectedRuntimeMode ?? "native",
|
|
startedAt: now,
|
|
finishedAt: now,
|
|
durationMs: 0,
|
|
cleanup: "not_started",
|
|
};
|
|
selected.push({
|
|
result: missing,
|
|
evidence: null,
|
|
directory: root,
|
|
valid: false,
|
|
errors: [missing.error!],
|
|
});
|
|
} else {
|
|
selected.push(attempts[0]);
|
|
}
|
|
}
|
|
|
|
await mkdir(output, { recursive: true });
|
|
const [stagedEvidence] = await Promise.all([
|
|
stageDashboardEvidence(selected, output),
|
|
stageDashboardBrandAssets(output),
|
|
]);
|
|
const resolvedResults = selected.map((entry) => ({
|
|
...entry.result,
|
|
// Cell evidence is produced by target-controlled code. The trusted report
|
|
// stamps the immutable target selected by the authorization job instead
|
|
// of allowing retained result metadata to claim another revision.
|
|
source: resolveRunnerE2ESource(entry.result.source),
|
|
status: entry.valid ? entry.result.status : ("failed" as const),
|
|
// Evidence failures must not be mistaken for an interrupted behavior recording.
|
|
failureClass: entry.errors.length > 0
|
|
? entry.evidence?.leaks?.length ? "secret_leak" as const : "permanent_infrastructure" as const
|
|
: entry.result.failureClass,
|
|
billing: summarizeExecutionBilling(entry.result),
|
|
}));
|
|
const generatedAt = new Date().toISOString();
|
|
const campaign = buildRunnerCampaign({
|
|
campaignId:
|
|
process.env.PAPERCLIP_E2E_CAMPAIGN_ID ??
|
|
(process.env.GITHUB_RUN_ID
|
|
? `gha-${process.env.GITHUB_RUN_ID}-${process.env.GITHUB_RUN_ATTEMPT ?? "1"}`
|
|
: `report-${generatedAt.replace(/[:.]/g, "-")}`),
|
|
generatedAt,
|
|
expected,
|
|
results: resolvedResults,
|
|
});
|
|
const billing = campaign.billing;
|
|
const normalized = {
|
|
...campaign,
|
|
results: selected.map((entry, index) => ({
|
|
...resolvedResults[index]!,
|
|
evidenceValid: entry.errors.length === 0,
|
|
evidenceErrors: entry.errors,
|
|
})),
|
|
};
|
|
await writeFile(
|
|
path.join(output, "normalized-results.json"),
|
|
`${JSON.stringify(normalized, null, 2)}\n`,
|
|
);
|
|
const dashboard = renderRunnerE2EDashboard({
|
|
title: "Runner Full-Stack E2E",
|
|
generatedAt: normalized.generatedAt,
|
|
expected,
|
|
catalog: runnerMatrix,
|
|
entries: selected.map((entry, index) => ({
|
|
result: resolvedResults[index]!,
|
|
valid: entry.valid,
|
|
errors: entry.errors,
|
|
evidenceBaseHref: stagedEvidence.get(entry.result.executionId)?.baseHref,
|
|
evidenceFiles: stagedEvidence.get(entry.result.executionId)?.files,
|
|
})),
|
|
campaign,
|
|
});
|
|
await Promise.all([
|
|
writeFile(path.join(output, "dashboard.html"), dashboard, "utf8"),
|
|
writeFile(path.join(output, "index.html"), dashboard, "utf8"),
|
|
]);
|
|
|
|
const summaryLinks = runnerE2ESummaryLinks({
|
|
campaignId: normalized.campaignId,
|
|
workflowRunUrl: normalized.source.workflowRunUrl,
|
|
historyPublicBaseUrl:
|
|
process.env.PAPERCLIP_RUNNER_E2E_HISTORY_PUBLIC_BASE_URL,
|
|
historyPrefix: process.env.PAPERCLIP_RUNNER_E2E_HISTORY_PREFIX,
|
|
});
|
|
const publicCampaignUrl = summaryLinks.find(
|
|
(link) => link.kind === "campaign",
|
|
)?.url;
|
|
|
|
const summaryLines = [
|
|
"# Runner Full-Stack E2E",
|
|
"",
|
|
`Passed: ${normalized.passed}/${selected.length}`,
|
|
"",
|
|
...(summaryLinks.length > 0
|
|
? [
|
|
"## View results",
|
|
"",
|
|
...summaryLinks.map(
|
|
(link) =>
|
|
`- [${link.label}](${link.url})${link.note ? ` — ${link.note}` : ""}`,
|
|
),
|
|
"",
|
|
]
|
|
: []),
|
|
`Tokens: ${billingCoverageLabel(`${billing.llm.inputTokens} input / ${billing.llm.outputTokens} output / ${billing.llm.cachedInputTokens} cached`, billing.llm.runsWithTokenUsage, billing.llm.runCount)}`,
|
|
"",
|
|
`Provider-reported LLM cost: ${billingCoverageLabel(`$${billing.reportedLlmCostUsd.toFixed(6)}`, billing.llm.runsWithReportedCost, billing.llm.runCount)}`,
|
|
"",
|
|
`Estimated Daytona list-price runtime cost: $${billing.estimatedRuntimeCostUsd.toFixed(6)}`,
|
|
...(billing.judge ? [`Estimated judge cost: ${billing.judge.estimatedCostUsd === null ? "unknown" : `$${billing.judge.estimatedCostUsd.toFixed(6)}`}; ${billing.judge.attempts} attempts; ${billing.judge.attemptsWithUnknownUsage} with unknown usage; $${billing.judge.reservedCostUsd.toFixed(6)} reserved`] : []),
|
|
"",
|
|
"| Cell | Attempt | Result | Runtime | Duration | Tokens (in/out) | LLM reported | Runtime estimate | Detail |",
|
|
"|---|---:|---|---|---:|---:|---:|---:|---|",
|
|
...selected.map((entry, index) => {
|
|
const detail = [entry.result.error, ...entry.errors].filter(Boolean).join("; ").replaceAll("|", "\\|") || "ok";
|
|
const resolved = resolvedResults[index]!;
|
|
const cellBilling = resolved.billing!;
|
|
const runtimeCost = cellBilling.runtime.estimatedListCostUsd;
|
|
const cell = publicCampaignUrl
|
|
? `[${resolved.executionId}](${publicCampaignUrl}#execution-${encodeURIComponent(resolved.executionId)})`
|
|
: resolved.executionId;
|
|
return `| ${cell} | ${resolved.attempt} | ${entry.valid ? "pass" : "fail"} | ${resolved.runtimeMode} | ${Math.round(resolved.durationMs / 1000)}s | ${billingCoverageLabel(`${cellBilling.llm.inputTokens}/${cellBilling.llm.outputTokens}`, cellBilling.llm.runsWithTokenUsage, cellBilling.llm.runCount)} | ${billingCoverageLabel(`$${cellBilling.reportedCostUsd.toFixed(6)}`, cellBilling.llm.runsWithReportedCost, cellBilling.llm.runCount)} | ${runtimeCost === undefined ? cellBilling.runtime.costStatus : `$${runtimeCost.toFixed(6)} est.`} | ${detail} |`;
|
|
}),
|
|
"",
|
|
];
|
|
await writeFile(path.join(output, "summary.md"), summaryLines.join("\n"));
|
|
|
|
const failures = selected.filter((entry) => !entry.valid).length;
|
|
const cases = selected
|
|
.map((entry) => {
|
|
const failure = entry.valid
|
|
? ""
|
|
: `<failure message="${xml([entry.result.error, ...entry.errors].filter(Boolean).join("; "))}"/>`;
|
|
return `<testcase classname="runner-full-stack-e2e" name="${xml(entry.result.executionId)}" time="${entry.result.durationMs / 1000}">${failure}</testcase>`;
|
|
})
|
|
.join("");
|
|
const junit = `<?xml version="1.0" encoding="UTF-8"?><testsuite name="Runner Full-Stack E2E" tests="${selected.length}" failures="${failures}">${cases}</testsuite>\n`;
|
|
await writeFile(path.join(output, "junit.xml"), junit);
|
|
console.log(summaryLines.join("\n"));
|
|
if (failures > 0) process.exitCode = 1;
|
|
}
|
|
|
|
await main().catch((error) => {
|
|
console.error(error instanceof Error ? error.message : String(error));
|
|
process.exitCode = 1;
|
|
});
|