mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paid runner E2E tests verify the complete runner, control-plane, and UI path. > - A server restart could load a fresh task page while Playwright still waited on an unsettled Vite navigation lifecycle. > - The current server also ignored the isolated Vite cache path and skipped Vite's per-request HTML transform from the known-green runner snapshot. > - A one-cell paid run then exposed that download-artifact v8 removes the artifact-name directory for one pattern match. > - This pull request restores the Vite contract, proves a fresh document after restart, and accepts only the exact singleton artifact layout. > - The benefit is reliable local runner qualification without weaker UI, source, or artifact checks. ## Linked Issues or Issue Description Refs #12769 Refs #12828 Refs #12829 Refs #12833 **What happened?** The structured-question restart test could time out after the replacement server returned the task route and rendered the durable pending interaction. A focused one-cell rerun passed the paid test but failed aggregation because download-artifact v8 flattened its single artifact. **Expected behavior** The test must prove that a new document loaded after the server restart and that the same pending interaction survived. The aggregate must accept the exact documented singleton download layout while it continues to reject ambiguous or foreign artifacts. **Steps to reproduce** 1. Run the local ACPX-Codex structured-question restart-resume cell. 2. Restart the isolated server while the question waits for an answer. 3. Observe that the route and task UI can reload before Playwright settles the navigation promise. 4. Run a paid campaign with one selected cell. 5. Observe download-artifact v8 extract the sole campaign directory directly into the requested path. **Paperclip version or commit** The local campaign reproduced the navigation failure at `3586956a1b794b3cb4a9c5f57ffb7355e2b0c46d`. The one-cell aggregate reproduced the singleton layout at `f487660c0a06ba06ca140b57386f21ed39f13120`. This fix is `de4ccceff453a4b39436bf9a2eb8f03924151af7`. **Deployment mode** Local development and paid GitHub Actions. **Installation method** Built from source. **Agent adapter(s) involved** ACPX-Codex. The Vite and aggregate fixes are provider-neutral. ## What Changed - Prove a new post-restart browser document with an in-memory sentinel. - Tolerate only Playwright's navigation timeout before the exact UI and API checks run. - Honor `PAPERCLIP_VITE_CACHE_DIR` in the embedded Vite server. - Limit dependency optimization to the real UI entry. - Run `vite.transformIndexHtml` for each request while caching only the branded source template. - Accept download-artifact v8's flattened layout only for one expected cell with one unique recognized campaign. - Keep source SHA, source ref, workflow URL, execution ID, attempt, and unexpected-entry validation. - Add focused positive and negative regressions for Vite rendering and singleton artifact selection. ## Verification - Exact 45-cell local campaign https://github.com/paperclipai/paperclip/actions/runs/33888939013 passed 44/45. Its only failure was the post-restart navigation false negative fixed here. - Exact focused rerun https://github.com/paperclipai/paperclip/actions/runs/33891207957 passed the ACPX-Codex restart cell first attempt with the same session, two durable runs, the terminal marker once, and cleanup complete. - The focused Vite renderer suite passed 2/2 tests. - The focused rerun-artifact selector suite passed 12/12 tests. - Prettier and `git diff --check` passed. - An exact-head 45-cell confirmation is pending. ## Risks Low to medium risk. The Vite change restores known-green per-request transforms and isolated cache behavior. It can affect all development UI loads. The paid matrix and ordinary CI will verify that behavior. The singleton selector remains fail-closed for ambiguous layouts and validates every result source. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, and parallel focused agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
353 lines
12 KiB
TypeScript
353 lines
12 KiB
TypeScript
import { cp, mkdir, readFile, readdir } from "node:fs/promises";
|
|
import path from "node:path";
|
|
import { pathToFileURL } from "node:url";
|
|
import type { RunnerE2EResult } from "./types.js";
|
|
|
|
interface WorkflowJob {
|
|
name: string;
|
|
run_attempt: number;
|
|
started_at: string;
|
|
}
|
|
|
|
interface WorkflowJobsResponse {
|
|
jobs: WorkflowJob[];
|
|
attempts: Array<{
|
|
run_attempt: number;
|
|
run_started_at: string;
|
|
}>;
|
|
}
|
|
|
|
export interface SelectRerunArtifactsInput {
|
|
artifactRoot: string;
|
|
selectedRoot: string;
|
|
jobs: WorkflowJobsResponse;
|
|
expectedExecutionIds: readonly string[];
|
|
workflowRunId: string;
|
|
workflowRunAttempt: number;
|
|
sourceSha: string;
|
|
sourceRef: string;
|
|
workflowRunUrl: string;
|
|
}
|
|
|
|
function safeIdentifier(value: string, label: string) {
|
|
if (!value || !/^[A-Za-z0-9_.-]+$/u.test(value)) {
|
|
throw new Error(
|
|
`${label} contains unsafe characters: ${JSON.stringify(value)}`,
|
|
);
|
|
}
|
|
return value;
|
|
}
|
|
|
|
function positiveInteger(value: unknown, label: string) {
|
|
if (typeof value !== "number" || !Number.isSafeInteger(value) || value < 1) {
|
|
throw new Error(`${label} must be a positive integer`);
|
|
}
|
|
return value;
|
|
}
|
|
|
|
function timestamp(value: unknown, label: string) {
|
|
if (typeof value !== "string" || !Number.isFinite(Date.parse(value))) {
|
|
throw new Error(`${label} must be an ISO timestamp`);
|
|
}
|
|
return Date.parse(value);
|
|
}
|
|
|
|
async function walk(root: string): Promise<string[]> {
|
|
const entries = await readdir(root, { withFileTypes: true });
|
|
const files: string[] = [];
|
|
for (const entry of entries) {
|
|
const full = path.join(root, entry.name);
|
|
if (entry.isDirectory()) files.push(...(await walk(full)));
|
|
else if (entry.isFile()) files.push(full);
|
|
}
|
|
return files;
|
|
}
|
|
|
|
/**
|
|
* Selects the artifact produced by the latest workflow job for each execution.
|
|
* GitHub reruns retain earlier-attempt artifacts, so validity or result
|
|
* timestamps must never be used to choose between workflow attempts.
|
|
*/
|
|
export async function selectRerunArtifacts(input: SelectRerunArtifactsInput) {
|
|
const runId = safeIdentifier(input.workflowRunId, "workflow run ID");
|
|
const currentAttempt = positiveInteger(
|
|
input.workflowRunAttempt,
|
|
"workflow run attempt",
|
|
);
|
|
const expected = input.expectedExecutionIds.map((executionId) =>
|
|
safeIdentifier(executionId, "execution ID"),
|
|
);
|
|
if (expected.length === 0 || new Set(expected).size !== expected.length) {
|
|
throw new Error("expected execution IDs must be non-empty and unique");
|
|
}
|
|
const preexistingSelections = await readdir(input.selectedRoot).catch(
|
|
(error: NodeJS.ErrnoException) => {
|
|
if (error.code === "ENOENT") return [];
|
|
throw error;
|
|
},
|
|
);
|
|
if (preexistingSelections.length > 0) {
|
|
throw new Error("selected artifact root must start empty");
|
|
}
|
|
const expectedSet = new Set(expected);
|
|
const attemptStartedAt = new Map<number, number>();
|
|
for (const attemptMetadata of input.jobs.attempts) {
|
|
const attempt = positiveInteger(
|
|
attemptMetadata.run_attempt,
|
|
"workflow metadata attempt",
|
|
);
|
|
if (attempt > currentAttempt || attemptStartedAt.has(attempt)) {
|
|
throw new Error(`invalid workflow metadata for attempt ${attempt}`);
|
|
}
|
|
attemptStartedAt.set(
|
|
attempt,
|
|
timestamp(
|
|
attemptMetadata.run_started_at,
|
|
`workflow attempt ${attempt} start`,
|
|
),
|
|
);
|
|
}
|
|
for (let attempt = 1; attempt <= currentAttempt; attempt += 1) {
|
|
if (!attemptStartedAt.has(attempt)) {
|
|
throw new Error(`workflow metadata omitted attempt ${attempt}`);
|
|
}
|
|
}
|
|
const attemptsByExecution = new Map<string, Map<number, WorkflowJob>>();
|
|
for (const candidate of input.jobs.jobs) {
|
|
if (!expectedSet.has(candidate.name)) continue;
|
|
const attempt = positiveInteger(candidate.run_attempt, "job run attempt");
|
|
if (attempt > currentAttempt) {
|
|
throw new Error(
|
|
`job ${candidate.name} claims future workflow attempt ${attempt}`,
|
|
);
|
|
}
|
|
const jobStartedAt = timestamp(
|
|
candidate.started_at,
|
|
`job ${candidate.name} start`,
|
|
);
|
|
// GitHub's filter=all response synthesizes current-attempt rows for jobs
|
|
// retained from an earlier attempt. Their started_at remains before the
|
|
// current attempt's trusted run_started_at, so they are not real reruns.
|
|
if (jobStartedAt < attemptStartedAt.get(attempt)!) continue;
|
|
const attempts = attemptsByExecution.get(candidate.name) ?? new Map();
|
|
if (attempts.has(attempt)) {
|
|
throw new Error(
|
|
`workflow attempt ${attempt} contains duplicate job ${candidate.name}`,
|
|
);
|
|
}
|
|
attempts.set(attempt, candidate);
|
|
attemptsByExecution.set(candidate.name, attempts);
|
|
}
|
|
|
|
const latestAttemptByExecution = new Map<string, number>();
|
|
for (const executionId of expected) {
|
|
const attempts = [...(attemptsByExecution.get(executionId)?.keys() ?? [])];
|
|
if (attempts.length === 0) {
|
|
throw new Error(
|
|
`workflow job history omitted expected execution ${executionId}`,
|
|
);
|
|
}
|
|
latestAttemptByExecution.set(executionId, Math.max(...attempts));
|
|
}
|
|
|
|
const recognizedArtifactNames = new Map<
|
|
string,
|
|
{ executionId: string; workflowAttempt: number }
|
|
>();
|
|
const recognizedCampaignNames = new Map<
|
|
string,
|
|
Array<{
|
|
artifactName: string;
|
|
executionId: string;
|
|
workflowAttempt: number;
|
|
}>
|
|
>();
|
|
for (const executionId of expected) {
|
|
for (const workflowAttempt of attemptsByExecution
|
|
.get(executionId)!
|
|
.keys()) {
|
|
const artifactName = `runner-e2e-${runId}-${workflowAttempt}-${executionId}`;
|
|
recognizedArtifactNames.set(artifactName, {
|
|
executionId,
|
|
workflowAttempt,
|
|
});
|
|
const campaignName = `gha-${runId}-${workflowAttempt}-${executionId}`;
|
|
const identities = recognizedCampaignNames.get(campaignName) ?? [];
|
|
identities.push({ artifactName, executionId, workflowAttempt });
|
|
recognizedCampaignNames.set(campaignName, identities);
|
|
}
|
|
}
|
|
|
|
const artifactEntries = await readdir(input.artifactRoot, {
|
|
withFileTypes: true,
|
|
}).catch((error: NodeJS.ErrnoException) => {
|
|
if (error.code === "ENOENT") return [];
|
|
throw error;
|
|
});
|
|
const artifactDirectories = new Map<
|
|
string,
|
|
| { layout: "wrapped"; directory: string }
|
|
| { layout: "flattened"; directory: string; campaignName: string }
|
|
>();
|
|
const singletonEntry = artifactEntries[0];
|
|
const singletonCampaignIdentities = singletonEntry
|
|
? recognizedCampaignNames.get(singletonEntry.name)
|
|
: undefined;
|
|
// download-artifact v8 flattens a single pattern match into the requested
|
|
// path. Accept that shape only when the expected set and campaign identity
|
|
// make the missing artifact-name wrapper unambiguous.
|
|
if (
|
|
expected.length === 1 &&
|
|
artifactEntries.length === 1 &&
|
|
singletonEntry?.isDirectory() &&
|
|
singletonCampaignIdentities?.length === 1
|
|
) {
|
|
const identity = singletonCampaignIdentities[0]!;
|
|
artifactDirectories.set(identity.artifactName, {
|
|
layout: "flattened",
|
|
directory: path.join(input.artifactRoot, singletonEntry.name),
|
|
campaignName: singletonEntry.name,
|
|
});
|
|
} else {
|
|
for (const entry of artifactEntries) {
|
|
const identity = recognizedArtifactNames.get(entry.name);
|
|
if (!identity || !entry.isDirectory()) {
|
|
throw new Error(`downloaded unexpected runner artifact ${entry.name}`);
|
|
}
|
|
artifactDirectories.set(entry.name, {
|
|
layout: "wrapped",
|
|
directory: path.join(input.artifactRoot, entry.name),
|
|
});
|
|
}
|
|
}
|
|
|
|
const selections: Array<{
|
|
executionId: string;
|
|
workflowAttempt: number;
|
|
artifactName: string;
|
|
}> = [];
|
|
for (const executionId of expected) {
|
|
const workflowAttempt = latestAttemptByExecution.get(executionId)!;
|
|
const artifactName = `runner-e2e-${runId}-${workflowAttempt}-${executionId}`;
|
|
const artifactDirectory = artifactDirectories.get(artifactName);
|
|
// A latest job without an artifact must remain missing. Falling back to an
|
|
// older successful artifact would mask an infrastructure/upload failure.
|
|
if (!artifactDirectory) continue;
|
|
|
|
const campaignName = `gha-${runId}-${workflowAttempt}-${executionId}`;
|
|
let campaignDirectory: string;
|
|
if (artifactDirectory.layout === "flattened") {
|
|
if (artifactDirectory.campaignName !== campaignName) {
|
|
throw new Error(
|
|
`${artifactName} must contain only its exact campaign ${campaignName}`,
|
|
);
|
|
}
|
|
campaignDirectory = artifactDirectory.directory;
|
|
} else {
|
|
const topLevelEntries = await readdir(artifactDirectory.directory, {
|
|
withFileTypes: true,
|
|
});
|
|
if (
|
|
topLevelEntries.length !== 1 ||
|
|
topLevelEntries[0]?.name !== campaignName ||
|
|
!topLevelEntries[0].isDirectory()
|
|
) {
|
|
throw new Error(
|
|
`${artifactName} must contain only its exact campaign ${campaignName}`,
|
|
);
|
|
}
|
|
campaignDirectory = path.join(artifactDirectory.directory, campaignName);
|
|
}
|
|
const resultFiles = (await walk(campaignDirectory)).filter(
|
|
(file) => path.basename(file) === "result.json",
|
|
);
|
|
if (resultFiles.length === 0) {
|
|
throw new Error(`${artifactName} contains no normalized result`);
|
|
}
|
|
for (const resultFile of resultFiles) {
|
|
const result = JSON.parse(
|
|
await readFile(resultFile, "utf8"),
|
|
) as RunnerE2EResult;
|
|
if (result.executionId !== executionId) {
|
|
throw new Error(
|
|
`${artifactName} contains result for ${String(result.executionId)}`,
|
|
);
|
|
}
|
|
const source = result.source;
|
|
if (
|
|
!source ||
|
|
source.sha !== input.sourceSha ||
|
|
source.ref !== input.sourceRef ||
|
|
source.workflowRunUrl !== input.workflowRunUrl
|
|
) {
|
|
throw new Error(`${artifactName} contains result from another source`);
|
|
}
|
|
}
|
|
const destination = path.join(
|
|
input.selectedRoot,
|
|
artifactName,
|
|
campaignName,
|
|
);
|
|
await mkdir(path.dirname(destination), { recursive: true });
|
|
await cp(campaignDirectory, destination, {
|
|
recursive: true,
|
|
force: false,
|
|
errorOnExist: true,
|
|
});
|
|
selections.push({ executionId, workflowAttempt, artifactName });
|
|
}
|
|
return selections;
|
|
}
|
|
|
|
async function main() {
|
|
const required = (name: string) => {
|
|
const value = process.env[name]?.trim();
|
|
if (!value) throw new Error(`${name} is required`);
|
|
return value;
|
|
};
|
|
const expected = JSON.parse(
|
|
required("PAPERCLIP_RUNNER_E2E_EXPECTED_IDS"),
|
|
) as unknown;
|
|
if (
|
|
!Array.isArray(expected) ||
|
|
expected.some((value) => typeof value !== "string")
|
|
) {
|
|
throw new Error(
|
|
"PAPERCLIP_RUNNER_E2E_EXPECTED_IDS must be a JSON string array",
|
|
);
|
|
}
|
|
const jobs = JSON.parse(
|
|
await readFile(required("PAPERCLIP_RUNNER_E2E_JOBS_JSON"), "utf8"),
|
|
) as WorkflowJobsResponse;
|
|
if (!jobs || !Array.isArray(jobs.jobs) || !Array.isArray(jobs.attempts)) {
|
|
throw new Error("workflow jobs JSON must contain jobs and attempts arrays");
|
|
}
|
|
const runId = required("GITHUB_RUN_ID");
|
|
const serverUrl = required("GITHUB_SERVER_URL");
|
|
const repository = required("GITHUB_REPOSITORY");
|
|
const selections = await selectRerunArtifacts({
|
|
artifactRoot: required("PAPERCLIP_RUNNER_E2E_ARTIFACT_ROOT"),
|
|
selectedRoot: required("PAPERCLIP_RUNNER_E2E_SELECTED_ROOT"),
|
|
jobs,
|
|
expectedExecutionIds: expected,
|
|
workflowRunId: runId,
|
|
workflowRunAttempt: Number(required("GITHUB_RUN_ATTEMPT")),
|
|
sourceSha: required("PAPERCLIP_RUNNER_E2E_SOURCE_SHA"),
|
|
sourceRef: required("PAPERCLIP_RUNNER_E2E_SOURCE_REF"),
|
|
workflowRunUrl: `${serverUrl}/${repository}/actions/runs/${runId}`,
|
|
});
|
|
console.log(
|
|
`Selected ${selections.length}/${expected.length} latest cell artifacts`,
|
|
);
|
|
}
|
|
|
|
if (
|
|
process.argv[1] &&
|
|
import.meta.url === pathToFileURL(path.resolve(process.argv[1])).href
|
|
) {
|
|
await main().catch((error) => {
|
|
console.error(error instanceof Error ? error.message : String(error));
|
|
process.exitCode = 1;
|
|
});
|
|
}
|