mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect accounts and choose an agent harness and model. > - The runtime change in #14970 supports custom providers on those connections. > - Normal setup must stay simple while advanced users can choose a compatible gateway. > - Shared connector rows and access controls keep these choices consistent. > - This pull request refines the agent setup UI and adds review stories and repeatable browser qualification. > - The qualification checks real tools and downloaded outputs, not only a successful run status. ## Linked Issues or Issue Description Refs #14970, #37, #13083, #14104, #14565, #12692. The core implementation in #14970 is merged. This branch incorporates its squash commit and targets `master`. Both PRs contain our implementation. #14016 is a reference only and is not a dependency. This PR has 96 changed files. ## What Changed - Complete model-provider connector presentation beside other connectors. Each row uses the existing Connect action and connection list. Tags are stored without category UI. The base PR includes the provider forms and routes. - Show persistent Subscription, API Key, and Advanced choices. Label Advanced as Custom Gateway. Reuse provider logos, connection lists, and permissions controls. Default access to the organization and all agents when permitted; keep narrowing controls under Advanced. - Keep Configure reachable before subscription sign-in, so users can select a supported environment when the default cannot sign in. Testing and saving still require a connection. Show the execution environment in Configure. Preserve the confirmed Connect choice. Editing a method, credential, saved account, or advanced choice requires that current choice to connect before testing or saving. Use matching model and thinking-effort dropdowns and retain connection icons in selected values. - Preserve the new harness model default when switching an existing OpenCode agent to Codex or Claude, and resolve user-selected model names with the effective harness. - Load popular OpenRouter models through the shared connection-model discovery path. Keep explicit model lists and manual model entry available. - Group onboarding, connection setup, agent runtime, management, recovery, and production-component stories under AI Connections / Provider routing. - Add an explicit-only provider-connections browser suite for managed local or existing local/staging targets. Use private browser profiles and credential handoffs. Support human-assisted subscription sign-in without sharing passwords or tokens in reports. - Verify persisted connection identity, runtime probes, tool execution, exact artifact bytes, completion, and context-dependent follow-up. Retain source/model provenance, cost bounds, closed error diagnostics, original failures, and cleanup evidence. - Add Gemini startup-model and skill-root fixes, Grok private-history detection, ACP filesystem regression fixtures, selected-workspace handling for local Hermes, and artifact-helper workspace fallback. - Keep managed Grok runtime homes disposable. Remove host-side transcript retention/restoration because private file modes do not isolate same-user agent processes. Ignore earlier development archives and use a fresh task handoff when history is unavailable. Verify the absence of restored transcripts with a separate same-user process. - Capture stopped-run diagnostics before deleting an attached-company fixture agent. Track creation and owned sign-in receipts; revoke only this attempt's accounts and never adopt a concurrent campaign's newly created account. Preserve failure signals and final status through cleanup. - Require the requested environment in the saved agent and every run, including follow-ups. Reject a forced incompatible target. Keep one cancellation state through startup, every cell, reporting, and teardown for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after interruption. Document qualification limits. ## Verification - Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes master `d9f600043`. The security fix in `a758fde31` passes full workspace typecheck, production build, and 119 connection/Grok regressions. The unchanged UI passes all 126 configuration/model-discovery tests and token gates. The final published-guide correction passes Grok adapter typecheck. Earlier head `eebd8225c` passed the complete deterministic runner suite (1,404 Vitest tests and 128 Node tests) and all CI jobs. Current-head CI run `37520147514` passed all 47 jobs, including the full sharded Vitest and browser matrix, production build, and canary dry run. All 55 checks completed: 53 successes and two expected skips. The current-head security scan passed, Greptile is 5/5, and no review threads remain open. - A separate same-user process reproduced reading a restored Grok transcript before the security fix. The regression now finds no transcript. Existing fresh-session fallback and ordinary session metadata behavior pass. - The final account-choice and cleanup fixes pass 85 setup tests and 26 qualification-harness tests. Regressions verify that editing a connection invalidates confirmation, Configure remains reachable before sign-in, diagnostics are captured before fixture deletion, and concurrent campaigns cannot adopt or revoke each other's accounts. UI and E2E typechecks pass. - The Storybook build and actual Chromium production-component stories passed during this change. Review the neighboring AI Connections / Provider routing stories, regular connector rows, three connection modes, model discovery, and the single execution-environment control in Configure. - Cancellation smoke verified authenticated cleanup before browser close for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during startup and reporting, missing-file ACP resource errors, and preserved permission denials. Both ACP runtime versions and 54 ACPX/Grok regressions passed. The deterministic connection-intent browser suite passed two tests. - Historical local qualification retained 43 passing API/gateway cells out of 46, with downloaded outputs and follow-up receipts. These attempts span earlier builds; they do not qualify this exact commit or staging. Subscription combinations, Gemini overloads, and the unresolved follow-up failure remain recorded rather than counted as passing. - Use `pnpm test:e2e:runner -- --list --suite provider-connections` to inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md` for credentials, target URL, sign-in assistance, budget, evidence, and cleanup. Paid live tests remain opt-in. ## Risks - The core implementation in #14970 is merged. This PR adds no database migration of its own. - Subscription login needs an interactive provider session. Dedicated accounts and staging qualification remain follow-up work; this PR does not certify every login combination for production. - Managed Grok transcript resume is deferred until provider history has an OS isolation or authorized broker solution. Follow-ups start fresh with Paperclip task context; earlier live Grok results do not qualify this behavior. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Live overloads and one unresolved follow-up timeout remain recorded. The stock CLI is unchanged, and those cases are not marked as passing. - Real-provider tests spend credits and use private credential/evidence directories. The launcher requires explicit selection and checks target ownership. It must not attach to a developer's database by accident. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local remain outside custom provider setup. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
400 lines
15 KiB
TypeScript
400 lines
15 KiB
TypeScript
import { stageDashboardBrandAssets } from "./report-assets.js";
|
|
import { discoverReportCatalog } from "./report-catalog.js";
|
|
import path from "node:path";
|
|
import {
|
|
copyFile,
|
|
mkdir,
|
|
readFile,
|
|
readdir,
|
|
writeFile,
|
|
} from "node:fs/promises";
|
|
import { runnerMatrix } from "./catalog.js";
|
|
import { billingCoverageLabel, summarizeExecutionBilling } from "./billing.js";
|
|
import { renderRunnerE2EDashboard } from "./dashboard.js";
|
|
import {
|
|
buildRunnerCampaign,
|
|
canonicalExecutionId,
|
|
upgradeRunnerResult,
|
|
} from "./history.js";
|
|
import { resolveRunnerE2ESource } from "./source.js";
|
|
import { runnerE2ESummaryLinks } from "./summary-links.js";
|
|
import type { RunnerE2EResult } from "./types.js";
|
|
|
|
const repositoryRoot = path.resolve(import.meta.dirname, "../..");
|
|
|
|
interface EvidenceManifest {
|
|
files?: string[];
|
|
leaks?: Array<{ file: string; reason: string }>;
|
|
missing?: string[];
|
|
}
|
|
|
|
interface AggregatedResult {
|
|
result: RunnerE2EResult;
|
|
evidence: EvidenceManifest | null;
|
|
directory: string;
|
|
valid: boolean;
|
|
errors: string[];
|
|
}
|
|
|
|
async function walk(root: string): Promise<string[]> {
|
|
const entries = await readdir(root, { withFileTypes: true }).catch(() => []);
|
|
const files: string[] = [];
|
|
for (const entry of entries) {
|
|
const full = path.join(root, entry.name);
|
|
if (entry.isDirectory()) files.push(...(await walk(full)));
|
|
else if (entry.isFile()) files.push(full);
|
|
}
|
|
return files;
|
|
}
|
|
|
|
function xml(value: unknown) {
|
|
return String(value ?? "")
|
|
.replaceAll("&", "&")
|
|
.replaceAll("<", "<")
|
|
.replaceAll(">", ">")
|
|
.replaceAll('"', """)
|
|
.replaceAll("'", "'");
|
|
}
|
|
|
|
function validate(result: RunnerE2EResult, evidence: EvidenceManifest | null) {
|
|
const errors: string[] = [];
|
|
if (result.cleanup !== "passed") errors.push(`cleanup=${result.cleanup}`);
|
|
if (!evidence) errors.push("evidence manifest missing");
|
|
if (evidence?.leaks?.length)
|
|
errors.push(`secret leaks=${evidence.leaks.length}`);
|
|
if (evidence?.missing?.length)
|
|
errors.push(`missing evidence=${evidence.missing.join(",")}`);
|
|
if (
|
|
result.status === "passed" &&
|
|
!evidence?.files?.includes("final-state.png")
|
|
) {
|
|
errors.push("passing final-state screenshot missing");
|
|
}
|
|
return errors;
|
|
}
|
|
|
|
function safeEvidenceRelative(relative: string) {
|
|
if (path.isAbsolute(relative)) return null;
|
|
const segments = relative.split(/[\\/]/).filter(Boolean);
|
|
if (segments.length === 0 || segments.some((segment) => segment === ".."))
|
|
return null;
|
|
return segments;
|
|
}
|
|
|
|
async function stageDashboardEvidence(
|
|
selected: readonly AggregatedResult[],
|
|
output: string,
|
|
) {
|
|
const staged = new Map<string, { baseHref: string; files: string[] }>();
|
|
for (const entry of selected) {
|
|
const baseSegments = [
|
|
"evidence",
|
|
entry.result.executionId,
|
|
`attempt-${entry.result.attempt}`,
|
|
];
|
|
const copied: string[] = [];
|
|
for (const relative of entry.evidence?.files ?? []) {
|
|
const segments = safeEvidenceRelative(relative);
|
|
if (!segments) continue;
|
|
const source = path.join(entry.directory, ...segments);
|
|
const destination = path.join(output, ...baseSegments, ...segments);
|
|
const didCopy = await mkdir(path.dirname(destination), {
|
|
recursive: true,
|
|
})
|
|
.then(() => copyFile(source, destination))
|
|
.then(() => true)
|
|
.catch(() => false);
|
|
if (didCopy) copied.push(segments.join("/"));
|
|
}
|
|
// Playwright renames attachment files with a content hash, while runner
|
|
// results retain the stable screenshot basename used by the dashboard and
|
|
// history publisher. Materialize each declared screenshot under that
|
|
// basename when its hashed attachment is present in the evidence manifest.
|
|
// The source is still restricted to manifest-listed files, so this cannot
|
|
// expand the evidence set beyond what the test recorded.
|
|
for (const screenshot of entry.result.screenshots ?? []) {
|
|
if (copied.includes(screenshot.file)) continue;
|
|
const stem = screenshot.file.replace(/\.png$/i, "");
|
|
const escapedStem = stem.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
|
const hashedAttachment = new RegExp(
|
|
`^${escapedStem}-[0-9a-f]{8,128}\\.png$`,
|
|
"i",
|
|
);
|
|
const candidates = (entry.evidence?.files ?? []).filter((relative) => {
|
|
const basename = path.posix.basename(relative);
|
|
return (
|
|
basename === screenshot.file ||
|
|
hashedAttachment.test(basename)
|
|
);
|
|
});
|
|
if (candidates.length !== 1) continue;
|
|
const segments = safeEvidenceRelative(candidates[0]!);
|
|
if (!segments) continue;
|
|
const source = path.join(entry.directory, ...segments);
|
|
const destination = path.join(output, ...baseSegments, screenshot.file);
|
|
const didCopy = await mkdir(path.dirname(destination), {
|
|
recursive: true,
|
|
})
|
|
.then(() => copyFile(source, destination))
|
|
.then(() => true)
|
|
.catch(() => false);
|
|
if (didCopy) copied.push(screenshot.file);
|
|
}
|
|
staged.set(entry.result.executionId, {
|
|
baseHref: baseSegments.join("/"),
|
|
files: copied,
|
|
});
|
|
}
|
|
return staged;
|
|
}
|
|
|
|
|
|
async function main() {
|
|
const root = path.resolve(
|
|
process.env.PAPERCLIP_RUNNER_E2E_REPORT_ROOT ?? "tests/runner-e2e/results",
|
|
);
|
|
const output = path.resolve(
|
|
process.env.PAPERCLIP_RUNNER_E2E_REPORT_OUT ?? path.join(root, "merged"),
|
|
);
|
|
const expectedInput = JSON.parse(
|
|
process.env.PAPERCLIP_RUNNER_E2E_EXPECTED_IDS ?? "[]",
|
|
) as string[];
|
|
if (
|
|
!Array.isArray(expectedInput) ||
|
|
expectedInput.some((value) => typeof value !== "string")
|
|
) {
|
|
throw new Error(
|
|
"PAPERCLIP_RUNNER_E2E_EXPECTED_IDS must be a JSON string array",
|
|
);
|
|
}
|
|
const expected = expectedInput.map(canonicalExecutionId);
|
|
if (expected.length === 0 || new Set(expected).size !== expected.length) {
|
|
throw new Error(
|
|
"PAPERCLIP_RUNNER_E2E_EXPECTED_IDS must contain unique selected executions",
|
|
);
|
|
}
|
|
const resultFiles = (await walk(root)).filter(
|
|
(file) => path.basename(file) === "result.json",
|
|
);
|
|
const candidates = new Map<string, AggregatedResult[]>();
|
|
for (const resultFile of resultFiles) {
|
|
const parsed = JSON.parse(
|
|
await readFile(resultFile, "utf8"),
|
|
) as RunnerE2EResult;
|
|
if (
|
|
parsed.schema !== "paperclip.runner-e2e.result/v1" &&
|
|
parsed.schema !== "paperclip.runner-e2e.result/v2"
|
|
)
|
|
continue;
|
|
const result = upgradeRunnerResult(parsed);
|
|
const directory = path.dirname(resultFile);
|
|
const evidence = await readFile(
|
|
path.join(directory, "evidence-manifest.json"),
|
|
"utf8",
|
|
)
|
|
.then((value) => JSON.parse(value) as EvidenceManifest)
|
|
.catch(() => null);
|
|
const errors = validate(result, evidence);
|
|
const entry = {
|
|
result,
|
|
evidence,
|
|
directory,
|
|
valid: result.status === "passed" && errors.length === 0,
|
|
errors,
|
|
};
|
|
candidates.set(result.executionId, [
|
|
...(candidates.get(result.executionId) ?? []),
|
|
entry,
|
|
]);
|
|
}
|
|
|
|
const reportCatalog = discoverReportCatalog({
|
|
catalog: runnerMatrix,
|
|
expected,
|
|
results: [...candidates.values()].flatMap((entries) =>
|
|
entries.map((entry) => entry.result),
|
|
),
|
|
});
|
|
const selected: AggregatedResult[] = [];
|
|
for (const executionId of expected) {
|
|
const attempts = (candidates.get(executionId) ?? []).sort((left, right) => {
|
|
// A report root can contain artifacts from multiple local campaigns (or
|
|
// reruns downloaded from CI). Prefer retained evidence that actually
|
|
// satisfies the campaign contract, then the newest execution. Attempt
|
|
// numbers are only meaningful inside one campaign and must not let an
|
|
// older failed attempt shadow a later successful rerun.
|
|
if (left.valid !== right.valid) return left.valid ? -1 : 1;
|
|
return (
|
|
Date.parse(right.result.finishedAt) -
|
|
Date.parse(left.result.finishedAt) ||
|
|
right.result.attempt - left.result.attempt
|
|
);
|
|
});
|
|
if (attempts.length === 0) {
|
|
const now = new Date().toISOString();
|
|
const execution = reportCatalog.find(
|
|
(candidate) => candidate.id === executionId,
|
|
)!;
|
|
const missing: RunnerE2EResult = {
|
|
schema: "paperclip.runner-e2e.result/v2",
|
|
executionId,
|
|
suiteId: execution?.suite.id ?? "core-compatibility",
|
|
suiteDefinitionHash: execution?.suiteDefinitionHash,
|
|
attempt: 0,
|
|
status: "failed",
|
|
failureClass: "permanent_infrastructure",
|
|
error: "No result artifact was uploaded",
|
|
profileId: execution?.profile.id ?? "unknown",
|
|
environmentId: execution?.environment.id ?? "local",
|
|
caseId: execution?.task.id ?? "unknown",
|
|
provider: execution?.profile.provider ?? "unknown",
|
|
model: execution?.profile.model ?? "unknown",
|
|
runtimeMode: execution?.profile.expectedRuntimeMode ?? "native",
|
|
startedAt: now,
|
|
finishedAt: now,
|
|
durationMs: 0,
|
|
cleanup: "not_started",
|
|
};
|
|
selected.push({
|
|
result: missing,
|
|
evidence: null,
|
|
directory: root,
|
|
valid: false,
|
|
errors: [missing.error!],
|
|
});
|
|
} else {
|
|
selected.push(attempts[0]);
|
|
}
|
|
}
|
|
|
|
await mkdir(output, { recursive: true });
|
|
const [stagedEvidence] = await Promise.all([
|
|
stageDashboardEvidence(selected, output),
|
|
stageDashboardBrandAssets(output),
|
|
]);
|
|
const resolvedResults = selected.map((entry) => ({
|
|
...entry.result,
|
|
// Cell evidence is produced by target-controlled code. The trusted report
|
|
// stamps the immutable target selected by the authorization job instead
|
|
// of allowing retained result metadata to claim another revision.
|
|
source: resolveRunnerE2ESource(entry.result.source),
|
|
status: entry.valid ? entry.result.status : ("failed" as const),
|
|
// Evidence failures must not be mistaken for an interrupted behavior recording.
|
|
failureClass: entry.errors.length > 0
|
|
? entry.evidence?.leaks?.length ? "secret_leak" as const : "permanent_infrastructure" as const
|
|
: entry.result.failureClass,
|
|
billing: summarizeExecutionBilling(entry.result),
|
|
}));
|
|
const generatedAt = new Date().toISOString();
|
|
const campaign = buildRunnerCampaign({
|
|
campaignId:
|
|
process.env.PAPERCLIP_E2E_CAMPAIGN_ID ??
|
|
(process.env.GITHUB_RUN_ID
|
|
? `gha-${process.env.GITHUB_RUN_ID}-${process.env.GITHUB_RUN_ATTEMPT ?? "1"}`
|
|
: `report-${generatedAt.replace(/[:.]/g, "-")}`),
|
|
generatedAt,
|
|
expected,
|
|
results: resolvedResults,
|
|
});
|
|
const billing = campaign.billing;
|
|
const normalized = {
|
|
...campaign,
|
|
results: selected.map((entry, index) => ({
|
|
...resolvedResults[index]!,
|
|
evidenceValid: entry.errors.length === 0,
|
|
evidenceErrors: entry.errors,
|
|
})),
|
|
};
|
|
await writeFile(
|
|
path.join(output, "normalized-results.json"),
|
|
`${JSON.stringify(normalized, null, 2)}\n`,
|
|
);
|
|
const dashboard = renderRunnerE2EDashboard({
|
|
title: "Runner Full-Stack E2E",
|
|
generatedAt: normalized.generatedAt,
|
|
expected,
|
|
catalog: runnerMatrix,
|
|
entries: selected.map((entry, index) => ({
|
|
result: resolvedResults[index]!,
|
|
valid: entry.valid,
|
|
errors: entry.errors,
|
|
evidenceBaseHref: stagedEvidence.get(entry.result.executionId)?.baseHref,
|
|
evidenceFiles: stagedEvidence.get(entry.result.executionId)?.files,
|
|
})),
|
|
campaign,
|
|
});
|
|
await Promise.all([
|
|
writeFile(path.join(output, "dashboard.html"), dashboard, "utf8"),
|
|
writeFile(path.join(output, "index.html"), dashboard, "utf8"),
|
|
]);
|
|
|
|
const summaryLinks = runnerE2ESummaryLinks({
|
|
campaignId: normalized.campaignId,
|
|
workflowRunUrl: normalized.source.workflowRunUrl,
|
|
historyPublicBaseUrl:
|
|
process.env.PAPERCLIP_RUNNER_E2E_HISTORY_PUBLIC_BASE_URL,
|
|
historyPrefix: process.env.PAPERCLIP_RUNNER_E2E_HISTORY_PREFIX,
|
|
});
|
|
const publicCampaignUrl = summaryLinks.find(
|
|
(link) => link.kind === "campaign",
|
|
)?.url;
|
|
|
|
const summaryLines = [
|
|
"# Runner Full-Stack E2E",
|
|
"",
|
|
`Passed: ${normalized.passed}/${selected.length}`,
|
|
"",
|
|
...(summaryLinks.length > 0
|
|
? [
|
|
"## View results",
|
|
"",
|
|
...summaryLinks.map(
|
|
(link) =>
|
|
`- [${link.label}](${link.url})${link.note ? ` — ${link.note}` : ""}`,
|
|
),
|
|
"",
|
|
]
|
|
: []),
|
|
`Tokens: ${billingCoverageLabel(`${billing.llm.inputTokens} input / ${billing.llm.outputTokens} output / ${billing.llm.cachedInputTokens} cached`, billing.llm.runsWithTokenUsage, billing.llm.runCount)}`,
|
|
"",
|
|
`Provider-reported LLM cost: ${billingCoverageLabel(`$${billing.reportedLlmCostUsd.toFixed(6)}`, billing.llm.runsWithReportedCost, billing.llm.runCount)}`,
|
|
"",
|
|
`Estimated Daytona list-price runtime cost: $${billing.estimatedRuntimeCostUsd.toFixed(6)}`,
|
|
...(billing.judge ? [`Estimated judge cost: ${billing.judge.estimatedCostUsd === null ? "unknown" : `$${billing.judge.estimatedCostUsd.toFixed(6)}`}; ${billing.judge.attempts} attempts; ${billing.judge.attemptsWithUnknownUsage} with unknown usage; $${billing.judge.reservedCostUsd.toFixed(6)} reserved`] : []),
|
|
"",
|
|
"| Cell | Attempt | Result | Runtime | Duration | Tokens (in/out) | LLM reported | Runtime estimate | Detail |",
|
|
"|---|---:|---|---|---:|---:|---:|---:|---|",
|
|
...selected.map((entry, index) => {
|
|
const detail = [entry.result.error, ...entry.errors].filter(Boolean).join("; ").replaceAll("|", "\\|") || "ok";
|
|
const resolved = resolvedResults[index]!;
|
|
const cellBilling = resolved.billing!;
|
|
const runtimeCost = cellBilling.runtime.estimatedListCostUsd;
|
|
const cell = publicCampaignUrl
|
|
? `[${resolved.executionId}](${publicCampaignUrl}#execution-${encodeURIComponent(resolved.executionId)})`
|
|
: resolved.executionId;
|
|
return `| ${cell} | ${resolved.attempt} | ${entry.valid ? "pass" : "fail"} | ${resolved.runtimeMode} | ${Math.round(resolved.durationMs / 1000)}s | ${billingCoverageLabel(`${cellBilling.llm.inputTokens}/${cellBilling.llm.outputTokens}`, cellBilling.llm.runsWithTokenUsage, cellBilling.llm.runCount)} | ${billingCoverageLabel(`$${cellBilling.reportedCostUsd.toFixed(6)}`, cellBilling.llm.runsWithReportedCost, cellBilling.llm.runCount)} | ${runtimeCost === undefined ? cellBilling.runtime.costStatus : `$${runtimeCost.toFixed(6)} est.`} | ${detail} |`;
|
|
}),
|
|
"",
|
|
];
|
|
await writeFile(path.join(output, "summary.md"), summaryLines.join("\n"));
|
|
|
|
const failures = selected.filter((entry) => !entry.valid).length;
|
|
const cases = selected
|
|
.map((entry) => {
|
|
const failure = entry.valid
|
|
? ""
|
|
: `<failure message="${xml([entry.result.error, ...entry.errors].filter(Boolean).join("; "))}"/>`;
|
|
return `<testcase classname="runner-full-stack-e2e" name="${xml(entry.result.executionId)}" time="${entry.result.durationMs / 1000}">${failure}</testcase>`;
|
|
})
|
|
.join("");
|
|
const junit = `<?xml version="1.0" encoding="UTF-8"?><testsuite name="Runner Full-Stack E2E" tests="${selected.length}" failures="${failures}">${cases}</testsuite>\n`;
|
|
await writeFile(path.join(output, "junit.xml"), junit);
|
|
console.log(summaryLines.join("\n"));
|
|
if (failures > 0) process.exitCode = 1;
|
|
}
|
|
|
|
await main().catch((error) => {
|
|
console.error(error instanceof Error ? error.message : String(error));
|
|
process.exitCode = 1;
|
|
});
|