Files
PaperClipAI/tests/runner-e2e/select-rerun-artifacts.ts
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00

364 lines
12 KiB
TypeScript

import { cp, mkdir, readFile, readdir } from "node:fs/promises";
import path from "node:path";
import { pathToFileURL } from "node:url";
import type { RunnerE2EResult } from "./types.js";
interface WorkflowJob {
name: string;
run_attempt: number;
started_at: string;
status?: string;
}
interface WorkflowJobsResponse {
jobs: WorkflowJob[];
attempts: Array<{
run_attempt: number;
run_started_at: string;
}>;
}
export interface SelectRerunArtifactsInput {
artifactRoot: string;
selectedRoot: string;
jobs: WorkflowJobsResponse;
expectedExecutionIds: readonly string[];
workflowRunId: string;
workflowRunAttempt: number;
sourceSha: string;
sourceRef: string;
workflowRunUrl: string;
}
function safeIdentifier(value: string, label: string) {
if (!value || !/^[A-Za-z0-9_.-]+$/u.test(value)) {
throw new Error(
`${label} contains unsafe characters: ${JSON.stringify(value)}`,
);
}
return value;
}
function positiveInteger(value: unknown, label: string) {
if (typeof value !== "number" || !Number.isSafeInteger(value) || value < 1) {
throw new Error(`${label} must be a positive integer`);
}
return value;
}
function timestamp(value: unknown, label: string) {
if (typeof value !== "string" || !Number.isFinite(Date.parse(value))) {
throw new Error(`${label} must be an ISO timestamp`);
}
return Date.parse(value);
}
async function walk(root: string): Promise<string[]> {
const entries = await readdir(root, { withFileTypes: true });
const files: string[] = [];
for (const entry of entries) {
const full = path.join(root, entry.name);
if (entry.isDirectory()) files.push(...(await walk(full)));
else if (entry.isFile()) files.push(full);
}
return files;
}
/**
* Selects the artifact produced by the latest workflow job for each execution.
* GitHub reruns retain earlier-attempt artifacts, so validity or result
* timestamps must never be used to choose between workflow attempts.
*/
export async function selectRerunArtifacts(input: SelectRerunArtifactsInput) {
const runId = safeIdentifier(input.workflowRunId, "workflow run ID");
const currentAttempt = positiveInteger(
input.workflowRunAttempt,
"workflow run attempt",
);
const expected = input.expectedExecutionIds.map((executionId) =>
safeIdentifier(executionId, "execution ID"),
);
if (expected.length === 0 || new Set(expected).size !== expected.length) {
throw new Error("expected execution IDs must be non-empty and unique");
}
const preexistingSelections = await readdir(input.selectedRoot).catch(
(error: NodeJS.ErrnoException) => {
if (error.code === "ENOENT") return [];
throw error;
},
);
if (preexistingSelections.length > 0) {
throw new Error("selected artifact root must start empty");
}
const expectedSet = new Set(expected);
const attemptStartedAt = new Map<number, number>();
for (const attemptMetadata of input.jobs.attempts) {
const attempt = positiveInteger(
attemptMetadata.run_attempt,
"workflow metadata attempt",
);
if (attempt > currentAttempt || attemptStartedAt.has(attempt)) {
throw new Error(`invalid workflow metadata for attempt ${attempt}`);
}
attemptStartedAt.set(
attempt,
timestamp(
attemptMetadata.run_started_at,
`workflow attempt ${attempt} start`,
),
);
}
for (let attempt = 1; attempt <= currentAttempt; attempt += 1) {
if (!attemptStartedAt.has(attempt)) {
throw new Error(`workflow metadata omitted attempt ${attempt}`);
}
}
const attemptsByExecution = new Map<string, Map<number, WorkflowJob>>();
for (const candidate of input.jobs.jobs) {
if (!expectedSet.has(candidate.name)) continue;
const attempt = positiveInteger(candidate.run_attempt, "job run attempt");
if (attempt > currentAttempt) {
throw new Error(
`job ${candidate.name} claims future workflow attempt ${attempt}`,
);
}
const jobStartedAt = timestamp(
candidate.started_at,
`job ${candidate.name} start`,
);
// GitHub's filter=all response synthesizes current-attempt rows for jobs
// retained from an earlier attempt. Their started_at remains before the
// current attempt's trusted run_started_at, so they are not real reruns.
if (jobStartedAt < attemptStartedAt.get(attempt)!) continue;
const attempts = attemptsByExecution.get(candidate.name) ?? new Map();
const previous = attempts.get(attempt);
if (previous) {
const hasStarted = (job: WorkflowJob) => job.status === "in_progress" || job.status === "completed";
// GitHub can retain a queued placeholder with a different job ID even
// after its replacement starts. It has no execution evidence. Keep a
// lone queued latest attempt, however, so an older pass cannot mask it.
if (previous.status === "queued" && hasStarted(candidate)) {
attempts.set(attempt, candidate);
continue;
}
if (hasStarted(previous) && candidate.status === "queued") continue;
throw new Error(
`workflow attempt ${attempt} contains duplicate job ${candidate.name}`,
);
}
attempts.set(attempt, candidate);
attemptsByExecution.set(candidate.name, attempts);
}
const latestAttemptByExecution = new Map<string, number>();
for (const executionId of expected) {
const attempts = [...(attemptsByExecution.get(executionId)?.keys() ?? [])];
if (attempts.length === 0) {
throw new Error(
`workflow job history omitted expected execution ${executionId}`,
);
}
latestAttemptByExecution.set(executionId, Math.max(...attempts));
}
const recognizedArtifactNames = new Map<
string,
{ executionId: string; workflowAttempt: number }
>();
const recognizedCampaignNames = new Map<
string,
Array<{
artifactName: string;
executionId: string;
workflowAttempt: number;
}>
>();
for (const executionId of expected) {
for (const workflowAttempt of attemptsByExecution
.get(executionId)!
.keys()) {
const artifactName = `runner-e2e-${runId}-${workflowAttempt}-${executionId}`;
recognizedArtifactNames.set(artifactName, {
executionId,
workflowAttempt,
});
const campaignName = `gha-${runId}-${workflowAttempt}-${executionId}`;
const identities = recognizedCampaignNames.get(campaignName) ?? [];
identities.push({ artifactName, executionId, workflowAttempt });
recognizedCampaignNames.set(campaignName, identities);
}
}
const artifactEntries = await readdir(input.artifactRoot, {
withFileTypes: true,
}).catch((error: NodeJS.ErrnoException) => {
if (error.code === "ENOENT") return [];
throw error;
});
const artifactDirectories = new Map<
string,
| { layout: "wrapped"; directory: string }
| { layout: "flattened"; directory: string; campaignName: string }
>();
const singletonEntry = artifactEntries[0];
const singletonCampaignIdentities = singletonEntry
? recognizedCampaignNames.get(singletonEntry.name)
: undefined;
// download-artifact v8 flattens a single pattern match into the requested
// path. Accept that shape only when the expected set and campaign identity
// make the missing artifact-name wrapper unambiguous.
if (
expected.length === 1 &&
artifactEntries.length === 1 &&
singletonEntry?.isDirectory() &&
singletonCampaignIdentities?.length === 1
) {
const identity = singletonCampaignIdentities[0]!;
artifactDirectories.set(identity.artifactName, {
layout: "flattened",
directory: path.join(input.artifactRoot, singletonEntry.name),
campaignName: singletonEntry.name,
});
} else {
for (const entry of artifactEntries) {
const identity = recognizedArtifactNames.get(entry.name);
if (!identity || !entry.isDirectory()) {
throw new Error(`downloaded unexpected runner artifact ${entry.name}`);
}
artifactDirectories.set(entry.name, {
layout: "wrapped",
directory: path.join(input.artifactRoot, entry.name),
});
}
}
const selections: Array<{
executionId: string;
workflowAttempt: number;
artifactName: string;
}> = [];
for (const executionId of expected) {
const workflowAttempt = latestAttemptByExecution.get(executionId)!;
const artifactName = `runner-e2e-${runId}-${workflowAttempt}-${executionId}`;
const artifactDirectory = artifactDirectories.get(artifactName);
// A latest job without an artifact must remain missing. Falling back to an
// older successful artifact would mask an infrastructure/upload failure.
if (!artifactDirectory) continue;
const campaignName = `gha-${runId}-${workflowAttempt}-${executionId}`;
let campaignDirectory: string;
if (artifactDirectory.layout === "flattened") {
if (artifactDirectory.campaignName !== campaignName) {
throw new Error(
`${artifactName} must contain only its exact campaign ${campaignName}`,
);
}
campaignDirectory = artifactDirectory.directory;
} else {
const topLevelEntries = await readdir(artifactDirectory.directory, {
withFileTypes: true,
});
if (
topLevelEntries.length !== 1 ||
topLevelEntries[0]?.name !== campaignName ||
!topLevelEntries[0].isDirectory()
) {
throw new Error(
`${artifactName} must contain only its exact campaign ${campaignName}`,
);
}
campaignDirectory = path.join(artifactDirectory.directory, campaignName);
}
const resultFiles = (await walk(campaignDirectory)).filter(
(file) => path.basename(file) === "result.json",
);
if (resultFiles.length === 0) {
throw new Error(`${artifactName} contains no normalized result`);
}
for (const resultFile of resultFiles) {
const result = JSON.parse(
await readFile(resultFile, "utf8"),
) as RunnerE2EResult;
if (result.executionId !== executionId) {
throw new Error(
`${artifactName} contains result for ${String(result.executionId)}`,
);
}
const source = result.source;
if (
!source ||
source.sha !== input.sourceSha ||
source.ref !== input.sourceRef ||
source.workflowRunUrl !== input.workflowRunUrl
) {
throw new Error(`${artifactName} contains result from another source`);
}
}
const destination = path.join(
input.selectedRoot,
artifactName,
campaignName,
);
await mkdir(path.dirname(destination), { recursive: true });
await cp(campaignDirectory, destination, {
recursive: true,
force: false,
errorOnExist: true,
});
selections.push({ executionId, workflowAttempt, artifactName });
}
return selections;
}
async function main() {
const required = (name: string) => {
const value = process.env[name]?.trim();
if (!value) throw new Error(`${name} is required`);
return value;
};
const expected = JSON.parse(
required("PAPERCLIP_RUNNER_E2E_EXPECTED_IDS"),
) as unknown;
if (
!Array.isArray(expected) ||
expected.some((value) => typeof value !== "string")
) {
throw new Error(
"PAPERCLIP_RUNNER_E2E_EXPECTED_IDS must be a JSON string array",
);
}
const jobs = JSON.parse(
await readFile(required("PAPERCLIP_RUNNER_E2E_JOBS_JSON"), "utf8"),
) as WorkflowJobsResponse;
if (!jobs || !Array.isArray(jobs.jobs) || !Array.isArray(jobs.attempts)) {
throw new Error("workflow jobs JSON must contain jobs and attempts arrays");
}
const runId = required("GITHUB_RUN_ID");
const serverUrl = required("GITHUB_SERVER_URL");
const repository = required("GITHUB_REPOSITORY");
const selections = await selectRerunArtifacts({
artifactRoot: required("PAPERCLIP_RUNNER_E2E_ARTIFACT_ROOT"),
selectedRoot: required("PAPERCLIP_RUNNER_E2E_SELECTED_ROOT"),
jobs,
expectedExecutionIds: expected,
workflowRunId: runId,
workflowRunAttempt: Number(required("GITHUB_RUN_ATTEMPT")),
sourceSha: required("PAPERCLIP_RUNNER_E2E_SOURCE_SHA"),
sourceRef: required("PAPERCLIP_RUNNER_E2E_SOURCE_REF"),
workflowRunUrl: `${serverUrl}/${repository}/actions/runs/${runId}`,
});
console.log(
`Selected ${selections.length}/${expected.length} latest cell artifacts`,
);
}
if (
process.argv[1] &&
import.meta.url === pathToFileURL(path.resolve(process.argv[1])).href
) {
await main().catch((error) => {
console.error(error instanceof Error ? error.message : String(error));
process.exitCode = 1;
});
}