Files
PaperClipAI/tests/runner-e2e/report.test.ts
T
Dotta 9ecd93a54d test(e2e): link runner campaign summaries (#12927)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip uses a paid full-stack campaign to verify runner behavior
across providers and environments.
> - The campaign already creates an interactive report, workflow logs,
and retained evidence artifacts.
> - The merge job summary shows result totals but does not link to those
resources.
> - Reviewers must search several workflow jobs and artifacts to find
the executed cells.
> - This pull request adds direct and safe links to the exact campaign,
each cell, the workflow logs, and the artifacts.
> - The benefit is that a reviewer can inspect a result from the Actions
summary with one click.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves the `Merge and enforce campaign result` summary in the
`Runner Full-Stack E2E` workflow.

**Subsystem affected**

The runner E2E report generator and its GitHub Actions workflow are
affected.

**Current behavior**

The summary lists each selected cell and its result. It does not link to
the published campaign report, the workflow logs, or the evidence
artifacts.

**Proposed behavior**

The summary includes a `View results` section. It links to the exact
immutable campaign report, the workflow logs, and the artifacts. Each
cell name links to its stable section in the campaign report.

**Reason and benefit**

The current summary does not show reviewers where to inspect the run.
Direct links make the result evidence discoverable without manual URL
construction or artifact searches.

**Breaking changes**

None. This change only adds links and stable HTML anchors to existing
report output.

**Additional context**

Related: #12904. The cited successful campaign is [run
34026735033](https://github.com/paperclipai/paperclip/actions/runs/34026735033).

## What Changed

- Add a safe URL builder for public campaign, workflow, and artifact
links.
- Add a `View results` section to the GitHub Actions campaign summary.
- Link each summary table cell to its exact section in the immutable
campaign report.
- Add stable execution anchors to the generated dashboard.
- Reject non-HTTPS, credential-bearing, malformed, and ambiguous link
destinations.
- Document the new links and their retention or publication timing.

## Verification

- `pnpm test:e2e:runner:unit` — 116 tests passed.
- `pnpm test:e2e:runner:typecheck` — passed.
- `pnpm typecheck` — passed, including migration safety.
- `pnpm build` — passed.
- `pnpm exec prettier --check ...` for all changed files — passed.
- `git diff --check origin/master...HEAD` — passed.
- The full local server suite also ran. One unrelated macOS
workspace-runtime file passed 157 tests and failed 4 existing path and
port assumptions. Two failures compare `/var` with `/private/var`. Two
failures cannot reserve a port outside a hard-coded range. This PR does
not change that file or its dependencies.

## Risks

- The immutable campaign link becomes available after the history
publisher completes. The workflow and artifact links remain available
while publication runs.
- The artifact link requires GitHub access and follows the existing
30-day retention period.
- Invalid configured URLs are omitted instead of being rendered into the
summary.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex desktop agent with GPT-5. The runtime does not expose the
context-window size. The agent used repository inspection, agentic
reasoning, code execution, and GitHub CLI tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g.
`docs/no-internal-issue-references`, `fix/sandbox-secret-resolution`,
`feat/adapter-retry-backoff`) and contains no internal Paperclip ticket
id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-06 09:09:51 -05:00

436 lines
15 KiB
TypeScript

import { execFile } from "node:child_process";
import { promisify } from "node:util";
import { mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises";
import os from "node:os";
import path from "node:path";
import { afterEach, describe, expect, it } from "vitest";
import type { RunnerE2EResult } from "./types.js";
const execFileAsync = promisify(execFile);
const repositoryRoot = path.resolve(import.meta.dirname, "../..");
const cleanupDirectories: string[] = [];
afterEach(async () => {
await Promise.all(
cleanupDirectories
.splice(0)
.map((directory) => rm(directory, { recursive: true, force: true })),
);
});
describe("runner E2E report aggregation", () => {
it("selects the latest retry and enforces cleanup and pass evidence", async () => {
const root = await mkdtemp(
path.join(os.tmpdir(), "runner-e2e-report-test-"),
);
cleanupDirectories.push(root);
const executionId = "legacy-codex.local.message-marker";
const base: RunnerE2EResult = {
schema: "paperclip.runner-e2e.result/v1",
executionId,
attempt: 2,
status: "passed",
profileId: "legacy-codex",
environmentId: "local",
caseId: "message-marker",
provider: "codex",
model: "fixture-model",
runtimeMode: "legacy",
startedAt: "2026-08-26T00:00:00.000Z",
finishedAt: "2026-08-26T00:00:01.000Z",
durationMs: 1_000,
source: {
sha: "forged-result-sha",
ref: "refs/heads/forged-result",
workflowRunUrl: "https://example.test/actions/runs/forged",
},
runIds: ["run-2"],
turnTimings: [
{
turn: 1,
submittedAt: "2026-08-26T00:00:00.000Z",
runStartedAt: "2026-08-26T00:00:00.100Z",
runFinishedAt: "2026-08-26T00:00:01.000Z",
schedulerLatencyMs: 100,
runDurationMs: 900,
responseLatencyMs: 1_000,
runId: "run-2",
leaseAcquisitionOutcome: "created",
},
],
usage: {
inputTokens: 1_250,
outputTokens: 75,
cachedInputTokens: 500,
costUsd: 0.0125,
},
cleanup: "passed",
};
for (const attempt of [1, 2]) {
const directory = path.join(root, `attempt-${attempt}`);
await mkdir(directory, { recursive: true });
const result =
attempt === 1
? {
...base,
attempt,
status: "failed" as const,
failureClass: "transient_infrastructure" as const,
}
: {
...base,
matcherResults: [
{
matcher: {
kind: "message_contains" as const,
expected: "PAPERCLIP_E2E_OK",
},
passed: true,
detail: "matched",
},
],
screenshots: [
{
id: "final-state",
label: "Final visible task state",
file: "final-state.png",
},
],
};
await writeFile(
path.join(directory, "result.json"),
JSON.stringify(result),
);
if (attempt === 2) {
await writeFile(path.join(directory, "final-state.png"), "fake-png");
}
await writeFile(
path.join(directory, "evidence-manifest.json"),
JSON.stringify({
files: attempt === 2 ? ["final-state.png"] : [],
leaks: [],
missing: [],
}),
);
}
const staleDuplicate = path.join(root, "attempt-2-stale-duplicate");
await mkdir(staleDuplicate, { recursive: true });
await writeFile(
path.join(staleDuplicate, "result.json"),
JSON.stringify({
...base,
status: "failed",
failureClass: "candidate_failure",
finishedAt: "2026-08-26T00:00:00.500Z",
}),
);
await writeFile(
path.join(staleDuplicate, "evidence-manifest.json"),
JSON.stringify({ files: [], leaks: [], missing: [] }),
);
const output = path.join(root, "merged");
await execFileAsync(
process.execPath,
[
path.join(repositoryRoot, "cli/node_modules/tsx/dist/cli.mjs"),
path.join(repositoryRoot, "tests/runner-e2e/report.ts"),
],
{
cwd: repositoryRoot,
env: {
...process.env,
PAPERCLIP_RUNNER_E2E_REPORT_ROOT: root,
PAPERCLIP_RUNNER_E2E_REPORT_OUT: output,
PAPERCLIP_RUNNER_E2E_EXPECTED_IDS: JSON.stringify([executionId]),
PAPERCLIP_RUNNER_E2E_SOURCE_SHA:
"0123456789abcdef0123456789abcdef01234567",
PAPERCLIP_RUNNER_E2E_SOURCE_REF:
"refs/heads/fix/runner-paid-source-attribution",
GITHUB_SHA: "trusted-default-workflow-sha",
GITHUB_REF: "refs/heads/master",
GITHUB_SERVER_URL: "https://github.com",
GITHUB_REPOSITORY: "paperclipai/paperclip",
GITHUB_RUN_ID: "123456",
PAPERCLIP_RUNNER_E2E_HISTORY_PUBLIC_BASE_URL:
"https://reports.example.test/",
PAPERCLIP_RUNNER_E2E_HISTORY_PREFIX: "/runner-e2e/",
},
},
);
const normalized = JSON.parse(
await readFile(path.join(output, "normalized-results.json"), "utf8"),
);
expect(normalized).toMatchObject({
schema: "paperclip.runner-e2e.campaign/v2",
selected: 1,
executed: 1,
passed: 1,
failed: 0,
retries: 1,
cleanupPassed: true,
source: {
sha: "0123456789abcdef0123456789abcdef01234567",
ref: "refs/heads/fix/runner-paid-source-attribution",
workflowRunUrl:
"https://github.com/paperclipai/paperclip/actions/runs/123456",
},
});
expect(normalized.billing).toMatchObject({
reportedLlmCostUsd: 0.0125,
llm: {
inputTokens: 1_250,
outputTokens: 75,
runsWithReportedCost: 1,
},
});
expect(normalized.results[0]).toMatchObject({
attempt: 2,
evidenceValid: true,
source: {
sha: "0123456789abcdef0123456789abcdef01234567",
ref: "refs/heads/fix/runner-paid-source-attribution",
workflowRunUrl:
"https://github.com/paperclipai/paperclip/actions/runs/123456",
},
});
const dashboard = await readFile(
path.join(output, "dashboard.html"),
"utf8",
);
expect(dashboard).toContain("Runner Full-Stack E2E");
expect(dashboard).toContain(executionId);
expect(dashboard).toContain("case-passed");
expect(dashboard).toContain(
"core-compatibility.runner-acpx-codex.daytona.message-marker",
);
expect(dashboard).toContain("case-not-selected");
expect(dashboard).toContain("<img");
expect(dashboard).toContain('class="brand-lockup"');
expect(dashboard).toContain("data-gallery-dialog");
expect(dashboard).toContain(
`id="execution-core-compatibility.${executionId}"`,
);
expect(dashboard).toContain("data-gallery-previous");
expect(dashboard).toContain("data-gallery-next");
expect(dashboard).toContain("View gallery · 1");
expect(dashboard).toContain(
"Declared PNG screenshots and sanitized structured evidence are retained with every published campaign",
);
expect(dashboard).toContain(
"Declared screenshots and sanitized structured evidence published",
);
expect(dashboard).toContain("message_contains");
expect(dashboard).toContain("Matchers and test context");
expect(dashboard).toContain("Scheduler");
expect(dashboard).toContain("Run duration");
expect(dashboard).toContain("100ms");
expect(dashboard).toContain("Campaign billing summary");
expect(dashboard).toContain("LLM reported subtotal");
expect(dashboard).toContain("Agent execution time");
expect(dashboard).toContain("Daytona lease time");
expect(dashboard).toContain("1,250 in · 75 out");
expect(dashboard).toContain("$0.0125");
expect(dashboard).toContain("unpriced or unavailable runs are excluded");
expect(dashboard).toContain('class="profile-sticky"');
expect(dashboard).toContain('class="mobile-environment-header"');
expect(dashboard).toContain("data-gallery-profile=");
expect(dashboard).toContain("data-gallery-environment=");
expect(dashboard).toContain("data-gallery-duration=");
expect(dashboard).toContain("data-gallery-tokens=");
expect(dashboard).toContain("data-gallery-matchers=");
expect(dashboard).toContain("data-report-query");
expect(dashboard).toContain("data-report-profile");
expect(dashboard).toContain("data-report-environment");
expect(dashboard).toContain("data-report-status");
expect(dashboard.indexOf('class="report-filters"')).toBeGreaterThan(
dashboard.indexOf('class="suite-nav"'),
);
expect(dashboard).not.toContain(".report-filters { position: sticky");
expect(dashboard).toContain("table-layout: fixed");
expect(dashboard).toContain('class="profile-column"');
expect(dashboard).toContain('aria-label="Previous"');
expect(dashboard).toContain('aria-label="Next"');
expect(dashboard).not.toContain("overflow: auto; max-height: calc(100vh");
expect(dashboard).toContain("@media (max-width: 1180px)");
expect(
await readFile(path.join(output, "assets", "favicon-32x32.png")),
).not.toHaveLength(0);
expect(
await readFile(path.join(output, "assets", "InterVariable.woff2")),
).not.toHaveLength(0);
expect(
await readFile(
path.join(
output,
"evidence",
`core-compatibility.${executionId}`,
"attempt-2",
"final-state.png",
),
"utf8",
),
).toBe("fake-png");
expect(await readFile(path.join(output, "index.html"), "utf8")).toBe(
dashboard,
);
const summary = await readFile(path.join(output, "summary.md"), "utf8");
expect(summary).toContain("## View results");
expect(summary).toContain(
"[Open the exact interactive campaign report](https://reports.example.test/runner-e2e/campaigns/gha-123456-1/index.html)",
);
expect(summary).toContain(
"[Open the workflow run and per-cell job logs](https://github.com/paperclipai/paperclip/actions/runs/123456)",
);
expect(summary).toContain(
"[Download the merged report and per-cell evidence](https://github.com/paperclipai/paperclip/actions/runs/123456#artifacts)",
);
expect(summary).toContain(
`[core-compatibility.${executionId}](https://reports.example.test/runner-e2e/campaigns/gha-123456-1/index.html#execution-core-compatibility.${executionId})`,
);
});
it("prefers a valid rerun over a higher attempt number from an older campaign", async () => {
const root = await mkdtemp(
path.join(os.tmpdir(), "runner-e2e-report-rerun-test-"),
);
cleanupDirectories.push(root);
const executionId = "runner-opencode.local.ask-question";
const common: RunnerE2EResult = {
schema: "paperclip.runner-e2e.result/v1",
executionId,
attempt: 2,
status: "failed",
failureClass: "candidate_failure",
profileId: "runner-opencode",
environmentId: "local",
caseId: "ask-question",
provider: "opencode",
model: "fixture-model",
runtimeMode: "native",
startedAt: "2026-08-26T00:00:00.000Z",
finishedAt: "2026-08-26T00:00:01.000Z",
durationMs: 1_000,
cleanup: "not_started",
};
const failedDirectory = path.join(root, "old-campaign", "attempt-2");
const passedDirectory = path.join(root, "new-campaign", "attempt-1");
await mkdir(failedDirectory, { recursive: true });
await mkdir(passedDirectory, { recursive: true });
await writeFile(
path.join(failedDirectory, "result.json"),
JSON.stringify(common),
);
await writeFile(
path.join(failedDirectory, "evidence-manifest.json"),
JSON.stringify({ files: [], leaks: [], missing: [] }),
);
await writeFile(
path.join(passedDirectory, "result.json"),
JSON.stringify({
...common,
attempt: 1,
status: "passed",
failureClass: undefined,
finishedAt: "2026-08-26T00:00:02.000Z",
cleanup: "passed",
}),
);
await writeFile(
path.join(passedDirectory, "evidence-manifest.json"),
JSON.stringify({
files: ["final-state.png"],
leaks: [],
missing: [],
}),
);
await writeFile(path.join(passedDirectory, "final-state.png"), "fake-png");
const output = path.join(root, "merged");
await execFileAsync(
process.execPath,
[
path.join(repositoryRoot, "cli/node_modules/tsx/dist/cli.mjs"),
path.join(repositoryRoot, "tests/runner-e2e/report.ts"),
],
{
cwd: repositoryRoot,
env: {
...process.env,
PAPERCLIP_RUNNER_E2E_REPORT_ROOT: root,
PAPERCLIP_RUNNER_E2E_REPORT_OUT: output,
PAPERCLIP_RUNNER_E2E_EXPECTED_IDS: JSON.stringify([executionId]),
},
},
);
const normalized = JSON.parse(
await readFile(path.join(output, "normalized-results.json"), "utf8"),
);
expect(normalized).toMatchObject({ passed: 1, failed: 0 });
expect(normalized.results[0]).toMatchObject({
attempt: 1,
status: "passed",
evidenceValid: true,
});
});
it("constructs the public root JUnit from fixed markup and escaped fields", async () => {
const root = await mkdtemp(
path.join(os.tmpdir(), "runner-e2e-report-junit-test-"),
);
cleanupDirectories.push(root);
const executionId = "legacy-codex.local.message-marker";
const directory = path.join(root, "attempt-1");
await mkdir(directory, { recursive: true });
await writeFile(
path.join(directory, "result.json"),
JSON.stringify({
schema: "paperclip.runner-e2e.result/v1",
executionId,
attempt: 1,
status: "failed",
failureClass: "candidate_failure",
error: `provider said \"><script>alert(1)</script>&`,
profileId: "legacy-codex",
environmentId: "local",
caseId: "message-marker",
provider: "codex",
model: "fixture-model",
runtimeMode: "legacy",
startedAt: "2026-08-26T00:00:00.000Z",
finishedAt: "2026-08-26T00:00:01.000Z",
durationMs: 1_000,
cleanup: "passed",
} satisfies RunnerE2EResult),
);
await writeFile(
path.join(directory, "evidence-manifest.json"),
JSON.stringify({ files: [], leaks: [], missing: [] }),
);
const output = path.join(root, "merged");
await expect(
execFileAsync(
process.execPath,
[
path.join(repositoryRoot, "cli/node_modules/tsx/dist/cli.mjs"),
path.join(repositoryRoot, "tests/runner-e2e/report.ts"),
],
{
cwd: repositoryRoot,
env: {
...process.env,
PAPERCLIP_RUNNER_E2E_REPORT_ROOT: root,
PAPERCLIP_RUNNER_E2E_REPORT_OUT: output,
PAPERCLIP_RUNNER_E2E_EXPECTED_IDS: JSON.stringify([executionId]),
},
},
),
).rejects.toBeDefined();
const junit = await readFile(path.join(output, "junit.xml"), "utf8");
expect(junit).toContain(
`message="provider said &quot;&gt;&lt;script&gt;alert(1)&lt;/script&gt;&amp;"`,
);
expect(junit).not.toContain("<script>");
expect(junit).not.toContain("<?xml-stylesheet");
});
});