mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 19:35:04 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip uses a paid full-stack campaign to verify runner behavior across providers and environments. > - The campaign already creates an interactive report, workflow logs, and retained evidence artifacts. > - The merge job summary shows result totals but does not link to those resources. > - Reviewers must search several workflow jobs and artifacts to find the executed cells. > - This pull request adds direct and safe links to the exact campaign, each cell, the workflow logs, and the artifacts. > - The benefit is that a reviewer can inspect a result from the Actions summary with one click. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the `Merge and enforce campaign result` summary in the `Runner Full-Stack E2E` workflow. **Subsystem affected** The runner E2E report generator and its GitHub Actions workflow are affected. **Current behavior** The summary lists each selected cell and its result. It does not link to the published campaign report, the workflow logs, or the evidence artifacts. **Proposed behavior** The summary includes a `View results` section. It links to the exact immutable campaign report, the workflow logs, and the artifacts. Each cell name links to its stable section in the campaign report. **Reason and benefit** The current summary does not show reviewers where to inspect the run. Direct links make the result evidence discoverable without manual URL construction or artifact searches. **Breaking changes** None. This change only adds links and stable HTML anchors to existing report output. **Additional context** Related: #12904. The cited successful campaign is [run 34026735033](https://github.com/paperclipai/paperclip/actions/runs/34026735033). ## What Changed - Add a safe URL builder for public campaign, workflow, and artifact links. - Add a `View results` section to the GitHub Actions campaign summary. - Link each summary table cell to its exact section in the immutable campaign report. - Add stable execution anchors to the generated dashboard. - Reject non-HTTPS, credential-bearing, malformed, and ambiguous link destinations. - Document the new links and their retention or publication timing. ## Verification - `pnpm test:e2e:runner:unit` — 116 tests passed. - `pnpm test:e2e:runner:typecheck` — passed. - `pnpm typecheck` — passed, including migration safety. - `pnpm build` — passed. - `pnpm exec prettier --check ...` for all changed files — passed. - `git diff --check origin/master...HEAD` — passed. - The full local server suite also ran. One unrelated macOS workspace-runtime file passed 157 tests and failed 4 existing path and port assumptions. Two failures compare `/var` with `/private/var`. Two failures cannot reserve a port outside a hard-coded range. This PR does not change that file or its dependencies. ## Risks - The immutable campaign link becomes available after the history publisher completes. The workflow and artifact links remain available while publication runs. - The artifact link requires GitHub access and follows the existing 30-day retention period. - Invalid configured URLs are omitted instead of being rendered into the summary. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex desktop agent with GPT-5. The runtime does not expose the context-window size. The agent used repository inspection, agentic reasoning, code execution, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/no-internal-issue-references`, `fix/sandbox-secret-resolution`, `feat/adapter-retry-backoff`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
368 lines
13 KiB
TypeScript
368 lines
13 KiB
TypeScript
import path from "node:path";
|
|
import {
|
|
copyFile,
|
|
mkdir,
|
|
readFile,
|
|
readdir,
|
|
writeFile,
|
|
} from "node:fs/promises";
|
|
import { runnerMatrix } from "./catalog.js";
|
|
import { summarizeExecutionBilling } from "./billing.js";
|
|
import { renderRunnerE2EDashboard } from "./dashboard.js";
|
|
import {
|
|
buildRunnerCampaign,
|
|
canonicalExecutionId,
|
|
upgradeRunnerResult,
|
|
} from "./history.js";
|
|
import { resolveRunnerE2ESource } from "./source.js";
|
|
import { runnerE2ESummaryLinks } from "./summary-links.js";
|
|
import type { RunnerE2EResult } from "./types.js";
|
|
|
|
const repositoryRoot = path.resolve(import.meta.dirname, "../..");
|
|
|
|
interface EvidenceManifest {
|
|
files?: string[];
|
|
leaks?: Array<{ file: string; reason: string }>;
|
|
missing?: string[];
|
|
}
|
|
|
|
interface AggregatedResult {
|
|
result: RunnerE2EResult;
|
|
evidence: EvidenceManifest | null;
|
|
directory: string;
|
|
valid: boolean;
|
|
errors: string[];
|
|
}
|
|
|
|
async function walk(root: string): Promise<string[]> {
|
|
const entries = await readdir(root, { withFileTypes: true }).catch(() => []);
|
|
const files: string[] = [];
|
|
for (const entry of entries) {
|
|
const full = path.join(root, entry.name);
|
|
if (entry.isDirectory()) files.push(...(await walk(full)));
|
|
else if (entry.isFile()) files.push(full);
|
|
}
|
|
return files;
|
|
}
|
|
|
|
function xml(value: unknown) {
|
|
return String(value ?? "")
|
|
.replaceAll("&", "&")
|
|
.replaceAll("<", "<")
|
|
.replaceAll(">", ">")
|
|
.replaceAll('"', """)
|
|
.replaceAll("'", "'");
|
|
}
|
|
|
|
function validate(result: RunnerE2EResult, evidence: EvidenceManifest | null) {
|
|
const errors: string[] = [];
|
|
if (result.status !== "passed")
|
|
errors.push(result.error ?? result.failureClass ?? "execution failed");
|
|
if (result.cleanup !== "passed") errors.push(`cleanup=${result.cleanup}`);
|
|
if (!evidence) errors.push("evidence manifest missing");
|
|
if (evidence?.leaks?.length)
|
|
errors.push(`secret leaks=${evidence.leaks.length}`);
|
|
if (evidence?.missing?.length)
|
|
errors.push(`missing evidence=${evidence.missing.join(",")}`);
|
|
if (
|
|
result.status === "passed" &&
|
|
!evidence?.files?.includes("final-state.png")
|
|
) {
|
|
errors.push("passing final-state screenshot missing");
|
|
}
|
|
return errors;
|
|
}
|
|
|
|
function safeEvidenceRelative(relative: string) {
|
|
if (path.isAbsolute(relative)) return null;
|
|
const segments = relative.split(/[\\/]/).filter(Boolean);
|
|
if (segments.length === 0 || segments.some((segment) => segment === ".."))
|
|
return null;
|
|
return segments;
|
|
}
|
|
|
|
async function stageDashboardEvidence(
|
|
selected: readonly AggregatedResult[],
|
|
output: string,
|
|
) {
|
|
const staged = new Map<string, { baseHref: string; files: string[] }>();
|
|
for (const entry of selected) {
|
|
const baseSegments = [
|
|
"evidence",
|
|
entry.result.executionId,
|
|
`attempt-${entry.result.attempt}`,
|
|
];
|
|
const copied: string[] = [];
|
|
for (const relative of entry.evidence?.files ?? []) {
|
|
const segments = safeEvidenceRelative(relative);
|
|
if (!segments) continue;
|
|
const source = path.join(entry.directory, ...segments);
|
|
const destination = path.join(output, ...baseSegments, ...segments);
|
|
const didCopy = await mkdir(path.dirname(destination), {
|
|
recursive: true,
|
|
})
|
|
.then(() => copyFile(source, destination))
|
|
.then(() => true)
|
|
.catch(() => false);
|
|
if (didCopy) copied.push(segments.join("/"));
|
|
}
|
|
staged.set(entry.result.executionId, {
|
|
baseHref: baseSegments.join("/"),
|
|
files: copied,
|
|
});
|
|
}
|
|
return staged;
|
|
}
|
|
|
|
async function stageDashboardBrandAssets(output: string) {
|
|
const assets = path.join(output, "assets");
|
|
await mkdir(assets, { recursive: true });
|
|
await Promise.all([
|
|
copyFile(
|
|
path.join(repositoryRoot, "ui/public/favicon-32x32.png"),
|
|
path.join(assets, "favicon-32x32.png"),
|
|
),
|
|
copyFile(
|
|
path.join(repositoryRoot, "ui/public/fonts/InterVariable.woff2"),
|
|
path.join(assets, "InterVariable.woff2"),
|
|
),
|
|
]);
|
|
}
|
|
|
|
async function main() {
|
|
const root = path.resolve(
|
|
process.env.PAPERCLIP_RUNNER_E2E_REPORT_ROOT ?? "tests/runner-e2e/results",
|
|
);
|
|
const output = path.resolve(
|
|
process.env.PAPERCLIP_RUNNER_E2E_REPORT_OUT ?? path.join(root, "merged"),
|
|
);
|
|
const expectedInput = JSON.parse(
|
|
process.env.PAPERCLIP_RUNNER_E2E_EXPECTED_IDS ?? "[]",
|
|
) as string[];
|
|
if (
|
|
!Array.isArray(expectedInput) ||
|
|
expectedInput.some((value) => typeof value !== "string")
|
|
) {
|
|
throw new Error(
|
|
"PAPERCLIP_RUNNER_E2E_EXPECTED_IDS must be a JSON string array",
|
|
);
|
|
}
|
|
const expected = expectedInput.map(canonicalExecutionId);
|
|
if (expected.length === 0 || new Set(expected).size !== expected.length) {
|
|
throw new Error(
|
|
"PAPERCLIP_RUNNER_E2E_EXPECTED_IDS must contain unique selected executions",
|
|
);
|
|
}
|
|
const resultFiles = (await walk(root)).filter(
|
|
(file) => path.basename(file) === "result.json",
|
|
);
|
|
const candidates = new Map<string, AggregatedResult[]>();
|
|
for (const resultFile of resultFiles) {
|
|
const parsed = JSON.parse(
|
|
await readFile(resultFile, "utf8"),
|
|
) as RunnerE2EResult;
|
|
if (
|
|
parsed.schema !== "paperclip.runner-e2e.result/v1" &&
|
|
parsed.schema !== "paperclip.runner-e2e.result/v2"
|
|
)
|
|
continue;
|
|
const result = upgradeRunnerResult(parsed);
|
|
const directory = path.dirname(resultFile);
|
|
const evidence = await readFile(
|
|
path.join(directory, "evidence-manifest.json"),
|
|
"utf8",
|
|
)
|
|
.then((value) => JSON.parse(value) as EvidenceManifest)
|
|
.catch(() => null);
|
|
const errors = validate(result, evidence);
|
|
const entry = {
|
|
result,
|
|
evidence,
|
|
directory,
|
|
valid: errors.length === 0,
|
|
errors,
|
|
};
|
|
candidates.set(result.executionId, [
|
|
...(candidates.get(result.executionId) ?? []),
|
|
entry,
|
|
]);
|
|
}
|
|
|
|
const selected: AggregatedResult[] = [];
|
|
for (const executionId of expected) {
|
|
const attempts = (candidates.get(executionId) ?? []).sort((left, right) => {
|
|
// A report root can contain artifacts from multiple local campaigns (or
|
|
// reruns downloaded from CI). Prefer retained evidence that actually
|
|
// satisfies the campaign contract, then the newest execution. Attempt
|
|
// numbers are only meaningful inside one campaign and must not let an
|
|
// older failed attempt shadow a later successful rerun.
|
|
if (left.valid !== right.valid) return left.valid ? -1 : 1;
|
|
return (
|
|
Date.parse(right.result.finishedAt) -
|
|
Date.parse(left.result.finishedAt) ||
|
|
right.result.attempt - left.result.attempt
|
|
);
|
|
});
|
|
if (attempts.length === 0) {
|
|
const now = new Date().toISOString();
|
|
const execution = runnerMatrix.find(
|
|
(candidate) => candidate.id === executionId,
|
|
);
|
|
const missing: RunnerE2EResult = {
|
|
schema: "paperclip.runner-e2e.result/v2",
|
|
executionId,
|
|
suiteId: execution?.suite.id ?? "core-compatibility",
|
|
suiteDefinitionHash: execution?.suiteDefinitionHash,
|
|
attempt: 0,
|
|
status: "failed",
|
|
failureClass: "permanent_infrastructure",
|
|
error: "No result artifact was uploaded",
|
|
profileId: execution?.profile.id ?? "unknown",
|
|
environmentId: execution?.environment.id ?? "local",
|
|
caseId: execution?.task.id ?? "unknown",
|
|
provider: execution?.profile.provider ?? "unknown",
|
|
model: execution?.profile.model ?? "unknown",
|
|
runtimeMode: execution?.profile.expectedRuntimeMode ?? "native",
|
|
startedAt: now,
|
|
finishedAt: now,
|
|
durationMs: 0,
|
|
cleanup: "not_started",
|
|
};
|
|
selected.push({
|
|
result: missing,
|
|
evidence: null,
|
|
directory: root,
|
|
valid: false,
|
|
errors: [missing.error!],
|
|
});
|
|
} else {
|
|
selected.push(attempts[0]);
|
|
}
|
|
}
|
|
|
|
await mkdir(output, { recursive: true });
|
|
const [stagedEvidence] = await Promise.all([
|
|
stageDashboardEvidence(selected, output),
|
|
stageDashboardBrandAssets(output),
|
|
]);
|
|
const resolvedResults = selected.map((entry) => ({
|
|
...entry.result,
|
|
// Cell evidence is produced by target-controlled code. The trusted report
|
|
// stamps the immutable target selected by the authorization job instead
|
|
// of allowing retained result metadata to claim another revision.
|
|
source: resolveRunnerE2ESource(entry.result.source),
|
|
status: entry.valid ? entry.result.status : ("failed" as const),
|
|
billing: summarizeExecutionBilling(entry.result),
|
|
}));
|
|
const generatedAt = new Date().toISOString();
|
|
const campaign = buildRunnerCampaign({
|
|
campaignId:
|
|
process.env.PAPERCLIP_E2E_CAMPAIGN_ID ??
|
|
(process.env.GITHUB_RUN_ID
|
|
? `gha-${process.env.GITHUB_RUN_ID}-${process.env.GITHUB_RUN_ATTEMPT ?? "1"}`
|
|
: `report-${generatedAt.replace(/[:.]/g, "-")}`),
|
|
generatedAt,
|
|
expected,
|
|
results: resolvedResults,
|
|
});
|
|
const billing = campaign.billing;
|
|
const normalized = {
|
|
...campaign,
|
|
results: selected.map((entry, index) => ({
|
|
...resolvedResults[index]!,
|
|
evidenceValid: entry.valid,
|
|
evidenceErrors: entry.errors,
|
|
})),
|
|
};
|
|
await writeFile(
|
|
path.join(output, "normalized-results.json"),
|
|
`${JSON.stringify(normalized, null, 2)}\n`,
|
|
);
|
|
const dashboard = renderRunnerE2EDashboard({
|
|
title: "Runner Full-Stack E2E",
|
|
generatedAt: normalized.generatedAt,
|
|
expected,
|
|
catalog: runnerMatrix,
|
|
entries: selected.map((entry, index) => ({
|
|
result: resolvedResults[index]!,
|
|
valid: entry.valid,
|
|
errors: entry.errors,
|
|
evidenceBaseHref: stagedEvidence.get(entry.result.executionId)?.baseHref,
|
|
evidenceFiles: stagedEvidence.get(entry.result.executionId)?.files,
|
|
})),
|
|
campaign,
|
|
});
|
|
await Promise.all([
|
|
writeFile(path.join(output, "dashboard.html"), dashboard, "utf8"),
|
|
writeFile(path.join(output, "index.html"), dashboard, "utf8"),
|
|
]);
|
|
|
|
const summaryLinks = runnerE2ESummaryLinks({
|
|
campaignId: normalized.campaignId,
|
|
workflowRunUrl: normalized.source.workflowRunUrl,
|
|
historyPublicBaseUrl:
|
|
process.env.PAPERCLIP_RUNNER_E2E_HISTORY_PUBLIC_BASE_URL,
|
|
historyPrefix: process.env.PAPERCLIP_RUNNER_E2E_HISTORY_PREFIX,
|
|
});
|
|
const publicCampaignUrl = summaryLinks.find(
|
|
(link) => link.kind === "campaign",
|
|
)?.url;
|
|
|
|
const summaryLines = [
|
|
"# Runner Full-Stack E2E",
|
|
"",
|
|
`Passed: ${normalized.passed}/${selected.length}`,
|
|
"",
|
|
...(summaryLinks.length > 0
|
|
? [
|
|
"## View results",
|
|
"",
|
|
...summaryLinks.map(
|
|
(link) =>
|
|
`- [${link.label}](${link.url})${link.note ? ` — ${link.note}` : ""}`,
|
|
),
|
|
"",
|
|
]
|
|
: []),
|
|
`Tokens: ${billing.llm.inputTokens} input / ${billing.llm.outputTokens} output / ${billing.llm.cachedInputTokens} cached`,
|
|
"",
|
|
`Provider-reported LLM cost: $${billing.reportedLlmCostUsd.toFixed(6)} (${billing.llm.runsWithReportedCost}/${billing.llm.runCount} runs priced)`,
|
|
"",
|
|
`Estimated Daytona list-price runtime cost: $${billing.estimatedRuntimeCostUsd.toFixed(6)}`,
|
|
"",
|
|
"| Cell | Attempt | Result | Runtime | Duration | Tokens (in/out) | LLM reported | Runtime estimate | Detail |",
|
|
"|---|---:|---|---|---:|---:|---:|---:|---|",
|
|
...selected.map((entry, index) => {
|
|
const detail = entry.errors.join("; ").replaceAll("|", "\\|") || "ok";
|
|
const resolved = resolvedResults[index]!;
|
|
const cellBilling = resolved.billing!;
|
|
const runtimeCost = cellBilling.runtime.estimatedListCostUsd;
|
|
const cell = publicCampaignUrl
|
|
? `[${resolved.executionId}](${publicCampaignUrl}#execution-${encodeURIComponent(resolved.executionId)})`
|
|
: resolved.executionId;
|
|
return `| ${cell} | ${resolved.attempt} | ${entry.valid ? "pass" : "fail"} | ${resolved.runtimeMode} | ${Math.round(resolved.durationMs / 1000)}s | ${cellBilling.llm.inputTokens}/${cellBilling.llm.outputTokens} | $${cellBilling.reportedCostUsd.toFixed(6)} (${cellBilling.llm.costStatus}) | ${runtimeCost === undefined ? cellBilling.runtime.costStatus : `$${runtimeCost.toFixed(6)} est.`} | ${detail} |`;
|
|
}),
|
|
"",
|
|
];
|
|
await writeFile(path.join(output, "summary.md"), summaryLines.join("\n"));
|
|
|
|
const failures = selected.filter((entry) => !entry.valid).length;
|
|
const cases = selected
|
|
.map((entry) => {
|
|
const failure = entry.valid
|
|
? ""
|
|
: `<failure message="${xml(entry.errors.join("; "))}"/>`;
|
|
return `<testcase classname="runner-full-stack-e2e" name="${xml(entry.result.executionId)}" time="${entry.result.durationMs / 1000}">${failure}</testcase>`;
|
|
})
|
|
.join("");
|
|
const junit = `<?xml version="1.0" encoding="UTF-8"?><testsuite name="Runner Full-Stack E2E" tests="${selected.length}" failures="${failures}">${cases}</testsuite>\n`;
|
|
await writeFile(path.join(output, "junit.xml"), junit);
|
|
console.log(summaryLines.join("\n"));
|
|
if (failures > 0) process.exitCode = 1;
|
|
}
|
|
|
|
await main().catch((error) => {
|
|
console.error(error instanceof Error ? error.message : String(error));
|
|
process.exitCode = 1;
|
|
});
|