mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task state to provider sessions. > - Follow-up turns must retain provider memory and carry new user direction. > - Lost session IDs caused repeated context and extra input tokens. > - Native question answers and approval races could leave valid work blocked. > - This pull request repairs those paths and adds regression coverage. > - Agents can continue accepted work without repeating the conversation or losing the user's answer. ## Linked Issues or Issue Description Refs #13574. That merged PR shortened continuation prompts and moved question instructions into tool documentation. This change preserves sessions and fixes failures exposed by broader testing. Related runtime work: #13408 and #13410. **What happened?** Native follow-up turns could lose the provider session ID. Completion guidance could replace the original task with its latest comment. Claude native questions could remain pending after the user answered. Approval during a running tool call could suspend the run before the tool response arrived. Onboarding and chat handoff instructions also caused repeated planning or missing plan documents. **Expected behavior** Reuse a valid provider session. Send only new events when that session already has the history. Preserve the task requirements and apply later user direction. Store the question answer and deliver it to the waiting run. Finish governed tool responses before suspending. Execute the accepted plan without asking for the same approval again. **Steps to reproduce** Run the continuation, local-session-integrity, first-task, and agent-chat suites with native Codex and Claude. Include provider-question-bridge, accept-while-running, and plan-handoff. **Paperclip version or commit** This branch is based on masterd54b75011. The active full catalog run tests4e75881db. Later review fixes have separate regression coverage. **Deployment mode** Isolated local instances and Daytona sandboxes in the existing Runner full-stack E2E harness. ## What Changed - Retain provider session identity across turns and late usage snapshots. Send new continuation events on session reuse, with full context available for a fresh session. - Preserve task requirements and later direction in completion guidance. Return the current contract revision after a stale completion submission. - Bridge native Claude questions to saved Paperclip cards. Submit answers through the saved card and resume the same run. - Delay governed suspension until tool results settle. Add a deterministic test barrier for approval during an active run. - Clarify free-text question examples, explicit onboarding plans, and execution of accepted chat plans. - Fix continuation readiness, verified output evidence, and declared screenshot collection. - Qualify the legacy Claude test CLI at 2.1.277. The old 2.1.19 CLI did not discover mounted skills. Update the existing workflow pin and isolated launcher together. - Refresh the Daytona image lockfile integrity pin after reviewing master patch updates. - Carry continuation mode as runtime metadata instead of inferring it from user-visible text. Install the test Claude CLI without lifecycle scripts. ## Verification - Targeted paid verification: 20/20 cases passed across local-environment campaigns before the rebase. [Report](https://pages.paperclip.ing/runner-e2e-seven-fixes-35397904249/). - Harness checks: 379 unit tests passed; harness typecheck passed. - Latest-head PR checks: 55 passed, two intentionally skipped. Greptile is 5/5; the security scan passes. - Review regressions: 350 executor tests and 204 session/driver tests passed. A script-free Claude install was verified with the actual CLI. - Full catalog, including the explicit-only everyday suite: [run 35417932353](https://github.com/paperclipai/paperclip/actions/runs/35417932353). Completed: **164/205 passed; 41 failed**. [Full dashboard and failure investigation](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/). Includes 204 case artifacts and one pre-case GitHub authorization timeout; missing evidence is not scored as a pass. The full run tested `4e75881db`; Final-head metadata/CLI smoke cases both passed. In the separate [six infrastructure retries](https://github.com/paperclipai/paperclip/actions/runs/35419769343), the GitHub timeout case passed and all five Docker preflight failures repeated. [Follow-up dashboard](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/follow-up/). - Full local typecheck and build passed on the rebased branch. The full local unit run completed with 657 passing files, two test timeouts and one suite setup timeout. All three affected files passed when rerun in isolation (84 tests). The first full local run was not clean. - Focused regression coverage includes the live question bridge, same-run response delivery, UI routing, stale revisions, approval overlap, and session reuse. ## Full-catalog follow-ups - Test infrastructure: 14 Claude everyday cells probe an absent host CLI; six cells failed pre-task GitHub/Docker qualification (GitHub passes on retry; all five Docker cases repeat; the workflow preflight allowlist omits their case IDs); five ACPX Codex cells cannot create sandbox namespaces. - Runtime: four OpenCode completion-criteria mismatches masked by shutdown errors, one service-approval suspension failure; three Daytona recovery failures encounter existing skill files; one duplicate completion wake. - Confirmed test defects: question pagination and a noncanonical plan document key. - Product/behavior: mismatched visible/required question sets, an attachment instead of the requested task document, one lone-option onboarding question, early completion instead of review, and a Codex Mini completion-schema failure. - The report job itself fails on trusted master’s stale patch/lock configuration. The linked report is rebuilt with the shared renderer from original cell results and public fixture screenshots; it excludes private snapshots, logs and traces. These are investigated follow-ups, not silently regraded passes. First-task passed 51/52. The PR checks are green independently of the broader catalog’s behavioral/infrastructure failures. ## Risks - Session reuse depends on a valid provider identity and context coverage. Fresh-session fallback and reset tests cover this boundary. - Native question delivery spans saved interaction state and a live provider run. Tests cover duplicate events, closed runs, and same-run answers. - Provider behavior varies. The full paid catalog may expose failures beyond these targeted fixes; those results will be reported without relaxing valid approval or output checks. - The legacy Claude version update is limited to test infrastructure. No database migration is included. ## Model Used OpenAI Codex, GPT-6 family, with repository inspection, code execution, and browser/E2E tools. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted checks and all three timeout-file reruns pass; full-run timeout caveat above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
547 lines
20 KiB
TypeScript
547 lines
20 KiB
TypeScript
import { execFile } from "node:child_process";
|
|
import { promisify } from "node:util";
|
|
import { mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises";
|
|
import os from "node:os";
|
|
import path from "node:path";
|
|
import { afterEach, describe, expect, it } from "vitest";
|
|
import type { RunnerE2EResult } from "./types.js";
|
|
|
|
const execFileAsync = promisify(execFile);
|
|
const repositoryRoot = path.resolve(import.meta.dirname, "../..");
|
|
const cleanupDirectories: string[] = [];
|
|
afterEach(async () => {
|
|
await Promise.all(
|
|
cleanupDirectories
|
|
.splice(0)
|
|
.map((directory) => rm(directory, { recursive: true, force: true })),
|
|
);
|
|
});
|
|
|
|
describe("runner E2E report aggregation", () => {
|
|
it("keeps interrupted journeys incomplete unless their evidence is invalid", async () => {
|
|
const root = await mkdtemp(path.join(os.tmpdir(), "runner-incomplete-report-"));
|
|
cleanupDirectories.push(root);
|
|
const ids: string[] = [];
|
|
for (const [caseId, validEvidence] of [["task-card-accept", true], ["task-reply-accept", false]] as const) {
|
|
const executionId = `first-task.runner-codex.local.${caseId}`;
|
|
ids.push(executionId);
|
|
const directory = path.join(root, caseId);
|
|
await mkdir(directory);
|
|
await writeFile(path.join(directory, "result.json"), JSON.stringify({
|
|
schema: "paperclip.runner-e2e.result/v2", executionId, suiteId: "first-task",
|
|
attempt: 1, status: "failed", failureClass: "candidate_failure", error: "Recording stopped before acceptance",
|
|
profileId: "runner-codex", environmentId: "local", caseId, provider: "codex", model: "fixture-model", runtimeMode: "native",
|
|
startedAt: "2026-09-15T00:00:00Z", finishedAt: "2026-09-15T00:01:00Z", durationMs: 60_000, cleanup: "passed",
|
|
firstTask: { caseId, nonce: "fixture", onboardingIssueId: "task", agentId: "agent", initialTaskIds: ["task"], instructions: [], configuredModel: null, observedModels: [], checkpoints: [],
|
|
checks: [{ id: "acceptance-recorded", passed: false, notReached: "No acceptance checkpoint", detail: "Acceptance recorded", evidence: [] }] },
|
|
} satisfies RunnerE2EResult));
|
|
if (validEvidence) await writeFile(path.join(directory, "evidence-manifest.json"), JSON.stringify({ files: [], leaks: [], missing: [] }));
|
|
}
|
|
const output = path.join(root, "merged");
|
|
await expect(execFileAsync(process.execPath, [path.join(repositoryRoot, "cli/node_modules/tsx/dist/cli.mjs"), path.join(repositoryRoot, "tests/runner-e2e/report.ts")], {
|
|
cwd: repositoryRoot, env: { ...process.env, PAPERCLIP_RUNNER_E2E_REPORT_ROOT: root, PAPERCLIP_RUNNER_E2E_REPORT_OUT: output, PAPERCLIP_RUNNER_E2E_EXPECTED_IDS: JSON.stringify(ids) },
|
|
})).rejects.toBeDefined();
|
|
const normalized = JSON.parse(await readFile(path.join(output, "normalized-results.json"), "utf8"));
|
|
expect(normalized).toMatchObject({ passed: 0, failed: 1, incomplete: 1 });
|
|
expect(normalized.results[0]).toMatchObject({ evidenceValid: true, evidenceErrors: [] });
|
|
expect(normalized.results[1]).toMatchObject({ evidenceValid: false, failureClass: "permanent_infrastructure" });
|
|
const page = await readFile(path.join(output, "index.html"), "utf8");
|
|
expect(page).toContain("Incomplete journey");
|
|
expect(page).toContain("evidence manifest missing");
|
|
});
|
|
|
|
it("selects the latest retry and enforces cleanup and pass evidence", async () => {
|
|
const root = await mkdtemp(
|
|
path.join(os.tmpdir(), "runner-e2e-report-test-"),
|
|
);
|
|
cleanupDirectories.push(root);
|
|
const executionId = "legacy-codex.local.message-marker";
|
|
const base: RunnerE2EResult = {
|
|
schema: "paperclip.runner-e2e.result/v1",
|
|
executionId,
|
|
attempt: 2,
|
|
status: "passed",
|
|
profileId: "legacy-codex",
|
|
environmentId: "local",
|
|
caseId: "message-marker",
|
|
provider: "codex",
|
|
model: "fixture-model",
|
|
runtimeMode: "legacy",
|
|
startedAt: "2026-08-26T00:00:00.000Z",
|
|
finishedAt: "2026-08-26T00:00:01.000Z",
|
|
durationMs: 1_000,
|
|
source: {
|
|
sha: "forged-result-sha",
|
|
ref: "refs/heads/forged-result",
|
|
workflowRunUrl: "https://example.test/actions/runs/forged",
|
|
},
|
|
runIds: ["run-2"],
|
|
turnTimings: [
|
|
{
|
|
turn: 1,
|
|
submittedAt: "2026-08-26T00:00:00.000Z",
|
|
runStartedAt: "2026-08-26T00:00:00.100Z",
|
|
runFinishedAt: "2026-08-26T00:00:01.000Z",
|
|
schedulerLatencyMs: 100,
|
|
runDurationMs: 900,
|
|
responseLatencyMs: 1_000,
|
|
runId: "run-2",
|
|
leaseAcquisitionOutcome: "created",
|
|
},
|
|
],
|
|
usage: {
|
|
inputTokens: 1_250,
|
|
outputTokens: 75,
|
|
cachedInputTokens: 500,
|
|
costUsd: 0.0125,
|
|
},
|
|
cleanup: "passed",
|
|
};
|
|
for (const attempt of [1, 2]) {
|
|
const directory = path.join(root, `attempt-${attempt}`);
|
|
await mkdir(directory, { recursive: true });
|
|
const result =
|
|
attempt === 1
|
|
? {
|
|
...base,
|
|
attempt,
|
|
status: "failed" as const,
|
|
failureClass: "transient_infrastructure" as const,
|
|
}
|
|
: {
|
|
...base,
|
|
matcherResults: [
|
|
{
|
|
matcher: {
|
|
kind: "message_contains" as const,
|
|
expected: "PAPERCLIP_E2E_OK",
|
|
},
|
|
passed: true,
|
|
detail: "matched",
|
|
},
|
|
],
|
|
screenshots: [
|
|
{
|
|
id: "final-state",
|
|
label: "Final visible task state",
|
|
file: "final-state.png",
|
|
},
|
|
],
|
|
};
|
|
await writeFile(
|
|
path.join(directory, "result.json"),
|
|
JSON.stringify(result),
|
|
);
|
|
if (attempt === 2) {
|
|
await writeFile(path.join(directory, "final-state.png"), "fake-png");
|
|
}
|
|
await writeFile(
|
|
path.join(directory, "evidence-manifest.json"),
|
|
JSON.stringify({
|
|
files: attempt === 2 ? ["final-state.png"] : [],
|
|
leaks: [],
|
|
missing: [],
|
|
}),
|
|
);
|
|
}
|
|
const staleDuplicate = path.join(root, "attempt-2-stale-duplicate");
|
|
await mkdir(staleDuplicate, { recursive: true });
|
|
await writeFile(
|
|
path.join(staleDuplicate, "result.json"),
|
|
JSON.stringify({
|
|
...base,
|
|
status: "failed",
|
|
failureClass: "candidate_failure",
|
|
finishedAt: "2026-08-26T00:00:00.500Z",
|
|
}),
|
|
);
|
|
await writeFile(
|
|
path.join(staleDuplicate, "evidence-manifest.json"),
|
|
JSON.stringify({ files: [], leaks: [], missing: [] }),
|
|
);
|
|
const output = path.join(root, "merged");
|
|
await execFileAsync(
|
|
process.execPath,
|
|
[
|
|
path.join(repositoryRoot, "cli/node_modules/tsx/dist/cli.mjs"),
|
|
path.join(repositoryRoot, "tests/runner-e2e/report.ts"),
|
|
],
|
|
{
|
|
cwd: repositoryRoot,
|
|
env: {
|
|
...process.env,
|
|
PAPERCLIP_RUNNER_E2E_REPORT_ROOT: root,
|
|
PAPERCLIP_RUNNER_E2E_REPORT_OUT: output,
|
|
PAPERCLIP_RUNNER_E2E_EXPECTED_IDS: JSON.stringify([executionId]),
|
|
PAPERCLIP_RUNNER_E2E_SOURCE_SHA:
|
|
"0123456789abcdef0123456789abcdef01234567",
|
|
PAPERCLIP_RUNNER_E2E_SOURCE_REF:
|
|
"refs/heads/fix/runner-paid-source-attribution",
|
|
GITHUB_SHA: "trusted-default-workflow-sha",
|
|
GITHUB_REF: "refs/heads/master",
|
|
GITHUB_SERVER_URL: "https://github.com",
|
|
GITHUB_REPOSITORY: "paperclipai/paperclip",
|
|
GITHUB_RUN_ID: "123456",
|
|
PAPERCLIP_RUNNER_E2E_HISTORY_PUBLIC_BASE_URL:
|
|
"https://reports.example.test/",
|
|
PAPERCLIP_RUNNER_E2E_HISTORY_PREFIX: "/runner-e2e/",
|
|
},
|
|
},
|
|
);
|
|
const normalized = JSON.parse(
|
|
await readFile(path.join(output, "normalized-results.json"), "utf8"),
|
|
);
|
|
expect(normalized).toMatchObject({
|
|
schema: "paperclip.runner-e2e.campaign/v2",
|
|
selected: 1,
|
|
executed: 1,
|
|
passed: 1,
|
|
failed: 0,
|
|
retries: 1,
|
|
cleanupPassed: true,
|
|
source: {
|
|
sha: "0123456789abcdef0123456789abcdef01234567",
|
|
ref: "refs/heads/fix/runner-paid-source-attribution",
|
|
workflowRunUrl:
|
|
"https://github.com/paperclipai/paperclip/actions/runs/123456",
|
|
},
|
|
});
|
|
expect(normalized.billing).toMatchObject({
|
|
reportedLlmCostUsd: 0.0125,
|
|
llm: {
|
|
inputTokens: 1_250,
|
|
outputTokens: 75,
|
|
runsWithReportedCost: 1,
|
|
},
|
|
});
|
|
expect(normalized.results[0]).toMatchObject({
|
|
attempt: 2,
|
|
evidenceValid: true,
|
|
source: {
|
|
sha: "0123456789abcdef0123456789abcdef01234567",
|
|
ref: "refs/heads/fix/runner-paid-source-attribution",
|
|
workflowRunUrl:
|
|
"https://github.com/paperclipai/paperclip/actions/runs/123456",
|
|
},
|
|
});
|
|
const dashboard = await readFile(
|
|
path.join(output, "dashboard.html"),
|
|
"utf8",
|
|
);
|
|
expect(dashboard).toContain("Runner Full-Stack E2E");
|
|
expect(dashboard).toContain(executionId);
|
|
expect(dashboard).toContain("case-passed");
|
|
expect(dashboard).toContain(
|
|
"core-compatibility.runner-acpx-codex.daytona.message-marker",
|
|
);
|
|
expect(dashboard).toContain("case-not-selected");
|
|
expect(dashboard).toContain("<img");
|
|
expect(dashboard).toContain('class="brand-lockup"');
|
|
expect(dashboard).toContain("data-gallery-dialog");
|
|
expect(dashboard).toContain(
|
|
`id="execution-core-compatibility.${executionId}"`,
|
|
);
|
|
expect(dashboard).toContain("data-gallery-previous");
|
|
expect(dashboard).toContain("data-gallery-next");
|
|
expect(dashboard).toContain("View gallery · 1");
|
|
expect(dashboard).toContain(
|
|
"Declared PNG screenshots and sanitized structured evidence are retained with every published campaign",
|
|
);
|
|
expect(dashboard).toContain(
|
|
"Declared screenshots and sanitized structured evidence published",
|
|
);
|
|
expect(dashboard).toContain("message_contains");
|
|
expect(dashboard).toContain("Matchers and test context");
|
|
expect(dashboard).toContain("Scheduler");
|
|
expect(dashboard).toContain("Run duration");
|
|
expect(dashboard).toContain("100ms");
|
|
expect(dashboard).toContain("Campaign billing summary");
|
|
expect(dashboard).toContain("LLM reported subtotal");
|
|
expect(dashboard).toContain("Agent execution time");
|
|
expect(dashboard).toContain("Daytona lease time");
|
|
expect(dashboard).toContain("1,250 in · 75 out");
|
|
expect(dashboard).toContain("$0.0125");
|
|
expect(dashboard).toContain("unpriced or unavailable runs are excluded");
|
|
expect(dashboard).toContain('class="profile-sticky"');
|
|
expect(dashboard).toContain('class="mobile-environment-header"');
|
|
expect(dashboard).toContain("data-gallery-profile=");
|
|
expect(dashboard).toContain("data-gallery-environment=");
|
|
expect(dashboard).toContain("data-gallery-duration=");
|
|
expect(dashboard).toContain("data-gallery-tokens=");
|
|
expect(dashboard).toContain("data-gallery-matchers=");
|
|
expect(dashboard).toContain("data-report-query");
|
|
expect(dashboard).toContain("data-report-profile");
|
|
expect(dashboard).toContain("data-report-environment");
|
|
expect(dashboard).toContain("data-report-status");
|
|
expect(dashboard.indexOf('class="report-filters"')).toBeGreaterThan(
|
|
dashboard.indexOf('class="suite-nav"'),
|
|
);
|
|
expect(dashboard).not.toContain(".report-filters { position: sticky");
|
|
expect(dashboard).toContain("table-layout: fixed");
|
|
expect(dashboard).toContain('class="profile-column"');
|
|
expect(dashboard).toContain('aria-label="Previous"');
|
|
expect(dashboard).toContain('aria-label="Next"');
|
|
expect(dashboard).not.toContain("overflow: auto; max-height: calc(100vh");
|
|
expect(dashboard).toContain("@media (max-width: 1180px)");
|
|
expect(
|
|
await readFile(path.join(output, "assets", "favicon-32x32.png")),
|
|
).not.toHaveLength(0);
|
|
expect(
|
|
await readFile(path.join(output, "assets", "InterVariable.woff2")),
|
|
).not.toHaveLength(0);
|
|
expect(
|
|
await readFile(
|
|
path.join(
|
|
output,
|
|
"evidence",
|
|
`core-compatibility.${executionId}`,
|
|
"attempt-2",
|
|
"final-state.png",
|
|
),
|
|
"utf8",
|
|
),
|
|
).toBe("fake-png");
|
|
expect(await readFile(path.join(output, "index.html"), "utf8")).toBe(
|
|
dashboard,
|
|
);
|
|
const summary = await readFile(path.join(output, "summary.md"), "utf8");
|
|
expect(summary).toContain("## View results");
|
|
expect(summary).toContain(
|
|
"[Open the exact interactive campaign report](https://reports.example.test/runner-e2e/campaigns/gha-123456-1/index.html)",
|
|
);
|
|
expect(summary).toContain(
|
|
"[Open the workflow run and per-cell job logs](https://github.com/paperclipai/paperclip/actions/runs/123456)",
|
|
);
|
|
expect(summary).toContain(
|
|
"[Download the merged report and per-cell evidence](https://github.com/paperclipai/paperclip/actions/runs/123456#artifacts)",
|
|
);
|
|
expect(summary).toContain(
|
|
`[core-compatibility.${executionId}](https://reports.example.test/runner-e2e/campaigns/gha-123456-1/index.html#execution-core-compatibility.${executionId})`,
|
|
);
|
|
});
|
|
|
|
it("prefers a valid rerun over a higher attempt number from an older campaign", async () => {
|
|
const root = await mkdtemp(
|
|
path.join(os.tmpdir(), "runner-e2e-report-rerun-test-"),
|
|
);
|
|
cleanupDirectories.push(root);
|
|
const executionId = "runner-opencode.local.ask-question";
|
|
const common: RunnerE2EResult = {
|
|
schema: "paperclip.runner-e2e.result/v1",
|
|
executionId,
|
|
attempt: 2,
|
|
status: "failed",
|
|
failureClass: "candidate_failure",
|
|
profileId: "runner-opencode",
|
|
environmentId: "local",
|
|
caseId: "ask-question",
|
|
provider: "opencode",
|
|
model: "fixture-model",
|
|
runtimeMode: "native",
|
|
startedAt: "2026-08-26T00:00:00.000Z",
|
|
finishedAt: "2026-08-26T00:00:01.000Z",
|
|
durationMs: 1_000,
|
|
cleanup: "not_started",
|
|
};
|
|
const failedDirectory = path.join(root, "old-campaign", "attempt-2");
|
|
const passedDirectory = path.join(root, "new-campaign", "attempt-1");
|
|
await mkdir(failedDirectory, { recursive: true });
|
|
await mkdir(passedDirectory, { recursive: true });
|
|
await writeFile(
|
|
path.join(failedDirectory, "result.json"),
|
|
JSON.stringify(common),
|
|
);
|
|
await writeFile(
|
|
path.join(failedDirectory, "evidence-manifest.json"),
|
|
JSON.stringify({ files: [], leaks: [], missing: [] }),
|
|
);
|
|
await writeFile(
|
|
path.join(passedDirectory, "result.json"),
|
|
JSON.stringify({
|
|
...common,
|
|
attempt: 1,
|
|
status: "passed",
|
|
failureClass: undefined,
|
|
finishedAt: "2026-08-26T00:00:02.000Z",
|
|
cleanup: "passed",
|
|
}),
|
|
);
|
|
await writeFile(
|
|
path.join(passedDirectory, "evidence-manifest.json"),
|
|
JSON.stringify({
|
|
files: ["final-state.png"],
|
|
leaks: [],
|
|
missing: [],
|
|
}),
|
|
);
|
|
await writeFile(path.join(passedDirectory, "final-state.png"), "fake-png");
|
|
const output = path.join(root, "merged");
|
|
await execFileAsync(
|
|
process.execPath,
|
|
[
|
|
path.join(repositoryRoot, "cli/node_modules/tsx/dist/cli.mjs"),
|
|
path.join(repositoryRoot, "tests/runner-e2e/report.ts"),
|
|
],
|
|
{
|
|
cwd: repositoryRoot,
|
|
env: {
|
|
...process.env,
|
|
PAPERCLIP_RUNNER_E2E_REPORT_ROOT: root,
|
|
PAPERCLIP_RUNNER_E2E_REPORT_OUT: output,
|
|
PAPERCLIP_RUNNER_E2E_EXPECTED_IDS: JSON.stringify([executionId]),
|
|
},
|
|
},
|
|
);
|
|
const normalized = JSON.parse(
|
|
await readFile(path.join(output, "normalized-results.json"), "utf8"),
|
|
);
|
|
expect(normalized).toMatchObject({ passed: 1, failed: 0 });
|
|
expect(normalized.results[0]).toMatchObject({
|
|
attempt: 1,
|
|
status: "passed",
|
|
evidenceValid: true,
|
|
});
|
|
});
|
|
|
|
it("materializes declared screenshots from hashed Playwright attachments", async () => {
|
|
const root = await mkdtemp(
|
|
path.join(os.tmpdir(), "runner-e2e-report-screenshot-alias-")
|
|
);
|
|
cleanupDirectories.push(root);
|
|
const executionId = "daytona-warm-continuity.legacy-codex.daytona.warm-three-turn";
|
|
const directory = path.join(root, "attempt-1");
|
|
const attachment =
|
|
"playwright-output/warm-turn/attachments/warm-turn-1-deadbeef.png";
|
|
await mkdir(path.join(directory, path.dirname(attachment)), {
|
|
recursive: true,
|
|
});
|
|
await writeFile(path.join(directory, "final-state.png"), "final-png");
|
|
await writeFile(path.join(directory, attachment), "warm-turn-png");
|
|
await writeFile(
|
|
path.join(directory, "result.json"),
|
|
JSON.stringify({
|
|
schema: "paperclip.runner-e2e.result/v1",
|
|
executionId,
|
|
attempt: 1,
|
|
status: "passed",
|
|
profileId: "legacy-codex",
|
|
environmentId: "daytona",
|
|
caseId: "warm-three-turn",
|
|
provider: "codex",
|
|
model: "fixture-model",
|
|
runtimeMode: "legacy",
|
|
startedAt: "2026-08-26T00:00:00.000Z",
|
|
finishedAt: "2026-08-26T00:00:01.000Z",
|
|
durationMs: 1_000,
|
|
cleanup: "passed",
|
|
screenshots: [
|
|
{
|
|
id: "warm-turn-1",
|
|
label: "Warm Daytona turn 1 awaiting review",
|
|
file: "warm-turn-1.png",
|
|
},
|
|
{
|
|
id: "final-state",
|
|
label: "Final visible task state",
|
|
file: "final-state.png",
|
|
},
|
|
],
|
|
} satisfies RunnerE2EResult),
|
|
);
|
|
await writeFile(
|
|
path.join(directory, "evidence-manifest.json"),
|
|
JSON.stringify({
|
|
files: ["final-state.png", attachment],
|
|
leaks: [],
|
|
missing: [],
|
|
}),
|
|
);
|
|
|
|
const output = path.join(root, "merged");
|
|
await execFileAsync(
|
|
process.execPath,
|
|
[
|
|
path.join(repositoryRoot, "cli/node_modules/tsx/dist/cli.mjs"),
|
|
path.join(repositoryRoot, "tests/runner-e2e/report.ts"),
|
|
],
|
|
{
|
|
cwd: repositoryRoot,
|
|
env: {
|
|
...process.env,
|
|
PAPERCLIP_RUNNER_E2E_REPORT_ROOT: root,
|
|
PAPERCLIP_RUNNER_E2E_REPORT_OUT: output,
|
|
PAPERCLIP_RUNNER_E2E_EXPECTED_IDS: JSON.stringify([executionId]),
|
|
},
|
|
},
|
|
);
|
|
expect(
|
|
await readFile(
|
|
path.join(output, "evidence", executionId, "attempt-1", "warm-turn-1.png"),
|
|
"utf8",
|
|
),
|
|
).toBe("warm-turn-png");
|
|
});
|
|
|
|
it("constructs the public root JUnit from fixed markup and escaped fields", async () => {
|
|
const root = await mkdtemp(
|
|
path.join(os.tmpdir(), "runner-e2e-report-junit-test-"),
|
|
);
|
|
cleanupDirectories.push(root);
|
|
const executionId = "legacy-codex.local.message-marker";
|
|
const directory = path.join(root, "attempt-1");
|
|
await mkdir(directory, { recursive: true });
|
|
await writeFile(
|
|
path.join(directory, "result.json"),
|
|
JSON.stringify({
|
|
schema: "paperclip.runner-e2e.result/v1",
|
|
executionId,
|
|
attempt: 1,
|
|
status: "failed",
|
|
failureClass: "candidate_failure",
|
|
error: `provider said \"><script>alert(1)</script>&`,
|
|
profileId: "legacy-codex",
|
|
environmentId: "local",
|
|
caseId: "message-marker",
|
|
provider: "codex",
|
|
model: "fixture-model",
|
|
runtimeMode: "legacy",
|
|
startedAt: "2026-08-26T00:00:00.000Z",
|
|
finishedAt: "2026-08-26T00:00:01.000Z",
|
|
durationMs: 1_000,
|
|
cleanup: "passed",
|
|
} satisfies RunnerE2EResult),
|
|
);
|
|
await writeFile(
|
|
path.join(directory, "evidence-manifest.json"),
|
|
JSON.stringify({ files: [], leaks: [], missing: [] }),
|
|
);
|
|
const output = path.join(root, "merged");
|
|
|
|
await expect(
|
|
execFileAsync(
|
|
process.execPath,
|
|
[
|
|
path.join(repositoryRoot, "cli/node_modules/tsx/dist/cli.mjs"),
|
|
path.join(repositoryRoot, "tests/runner-e2e/report.ts"),
|
|
],
|
|
{
|
|
cwd: repositoryRoot,
|
|
env: {
|
|
...process.env,
|
|
PAPERCLIP_RUNNER_E2E_REPORT_ROOT: root,
|
|
PAPERCLIP_RUNNER_E2E_REPORT_OUT: output,
|
|
PAPERCLIP_RUNNER_E2E_EXPECTED_IDS: JSON.stringify([executionId]),
|
|
},
|
|
},
|
|
),
|
|
).rejects.toBeDefined();
|
|
|
|
const junit = await readFile(path.join(output, "junit.xml"), "utf8");
|
|
expect(junit).toContain(
|
|
`message="provider said "><script>alert(1)</script>&"`,
|
|
);
|
|
expect(junit).not.toContain("<script>");
|
|
expect(junit).not.toContain("<?xml-stylesheet");
|
|
});
|
|
});
|