Files
PaperClipAI/tests/runner-e2e/history.ts
T
Dotta af3023f1e3 fix(runner): repair paid provider startup paths (#12769)
## Thinking Path

> - Paperclip manages AI agents that perform work.
> - Paperclip Runner connects durable task runs to local provider
processes.
> - The full-stack paid matrix exposed failures after the runner
integrity repair.
> - Verified JavaScript entrypoints lost their relative module graph
when Linux executed them through descriptor paths.
> - Returned provider startup errors also remained pending and became
indeterminate after recovery.
> - Sparse Codex tool lifecycle events lost the `write_document`
identity before task transcript projection.
> - This pull request repairs those three boundaries and makes the
structured-question fixture deterministic.
> - The benefit is repeatable provider startup, exact failure replay,
and correct inline Plan placement.

## Linked Issues or Issue Description

Refs #12721 and #12700.

**What happened?**

The paid runner matrix failed ACPX and OpenCode startup before provider
session creation. The runner journal then replaced the original startup
error with an indeterminate recovery result. Native Codex saved a Plan
but rendered it only as a fallback card. A legacy Claude waiting reply
could also echo the reserved terminal marker before the answer arrived.

**Expected behavior**

Verified JavaScript providers must start from immutable
descriptor-backed artifacts. Returned startup failures must persist as
terminal failed command results. Native tool lifecycle updates must
preserve the `write_document` boundary. Pre-answer fixture output must
not contain the reserved terminal marker.

**Steps to reproduce**

1. Run the local provider cells in the Runner Full-Stack E2E workflow.
2. Observe ACPX and OpenCode fail during `session.open` before provider
execution.
3. Observe recovery report `execution_indeterminate` instead of the
original startup error.
4. Run the native Codex Plan cell and observe the fallback Plan card
after the tool activity row.
5. Run the legacy Claude structured-question resume cell and observe an
early marker echo in waiting prose.

**Paperclip version or commit**

`0f9452101740835ce0b1488a204bf48acd5bafc3`

**Deployment mode**

Local development with the paid GitHub Actions acceptance workflow.

## What Changed

- Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM
entrypoints before hashing and verified descriptor launch.
- Anchor ACPX dynamic provider package resolution at a
controller-derived provider-pack root and keep that root out of the
provider child environment.
- Persist executor-returned startup errors as redacted durable failed
command results while retaining indeterminate recovery for true process
death.
- Coalesce sparse native tool items by stable ID so a late
`write_document` name, input, and result reach the transcript boundary
once.
- Forbid the structured-question fixture from spelling or announcing its
reserved terminal marker before the user answers.

## Verification

- Rust and TypeScript regression tests cover durable failed replay, true
crash ambiguity, bundle closure, package-root derivation, environment
filtering, exact Codex tool lifecycle coalescing, and prompt
determinism.
- Local execution is intentionally limited to formatters and static diff
checks. GitHub Actions will run tests, type checks, builds, and security
checks.
- After ordinary CI is green, scoped paid cells will validate one ACPX
launch, one OpenCode launch, native Codex Plan projection, and legacy
Claude structured resume before a complete matrix rerun.
- Prior failing matrix:
https://github.com/paperclipai/paperclip/actions/runs/33682434315

## Risks

- Bundling changes the bytes covered by provider launch hashes.
Provider-pack generation already hashes the final built files.
- ACPX still loads qualified provider packages dynamically. The
controller supplies a normalized package root, while existing version,
digest, path, and descriptor checks remain active.
- Durable `failed` is terminal. Replays return the same redacted result
and do not execute the provider effect twice.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex based on GPT-5 with agentic reasoning, repository
inspection, code editing, Git, parallel subagents, and GitHub Actions
coordination. The exact deployed snapshot and context-window size are
not exposed to this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked related public work or described the bug in
this PR
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
ticket id
- [ ] I have run tests locally and they pass (intentionally deferred to
GitHub Actions)
- [x] I have added or updated tests where applicable
- [x] No documentation change is required for this runtime repair
- [x] I have considered and documented the risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-04 07:58:44 -05:00

239 lines
7.4 KiB
TypeScript

import { runnerMatrix, runnerSuites } from "./catalog.js";
import {
aggregateCampaignBilling,
summarizeExecutionBilling,
} from "./billing.js";
import { resolveRunnerE2ESource } from "./source.js";
import type {
RunnerE2ECampaign,
RunnerE2EHistoryCampaign,
RunnerE2EHistoryIndex,
RunnerE2EResult,
} from "./types.js";
export function canonicalExecutionId(id: string) {
if (runnerMatrix.some((execution) => execution.id === id)) return id;
const coreId = `core-compatibility.${id}`;
return runnerMatrix.some((execution) => execution.id === coreId)
? coreId
: id;
}
export function upgradeRunnerResult(result: RunnerE2EResult): RunnerE2EResult {
const executionId = canonicalExecutionId(result.executionId);
const execution = runnerMatrix.find(
(candidate) => candidate.id === executionId,
);
if (!execution) return result;
return {
...result,
executionId,
suiteId: result.suiteId ?? execution.suite.id,
suiteDefinitionHash:
result.suiteDefinitionHash ?? execution.suiteDefinitionHash,
...(result.schema === "paperclip.runner-e2e.result/v1"
? {
source: result.source ?? {
sha: null,
ref: null,
workflowRunUrl: null,
},
}
: {}),
};
}
export function buildRunnerCampaign(input: {
campaignId: string;
generatedAt: string;
expected: readonly string[];
results: readonly RunnerE2EResult[];
eventName?: string | null;
}): RunnerE2ECampaign {
const expected = input.expected.map(canonicalExecutionId);
const results = input.results.map((result) => ({
...upgradeRunnerResult(result),
billing: result.billing ?? summarizeExecutionBilling(result),
}));
const resultSource = results.find((result) => result.source)?.source;
const source = {
...resolveRunnerE2ESource(resultSource),
eventName: input.eventName ?? process.env.GITHUB_EVENT_NAME ?? null,
};
const suites = runnerSuites
.map((suite) => {
const suiteExpected = expected.filter((id) =>
id.startsWith(`${suite.id}.`),
);
if (suiteExpected.length === 0) return null;
const suiteResults = results.filter(
(result) => result.suiteId === suite.id,
);
const passed = suiteResults.filter(
(result) => result.status === "passed" && result.cleanup === "passed",
).length;
const executed = suiteResults.filter(
(result) => result.attempt > 0,
).length;
return {
suiteId: suite.id,
suiteDefinitionHash:
suiteResults[0]?.suiteDefinitionHash ??
runnerMatrix.find((execution) => execution.suite.id === suite.id)!
.suiteDefinitionHash,
expected: suite.expectedMatrixSize,
selected: suiteExpected.length,
executed,
passed,
failed: suiteExpected.length - passed,
retries: suiteResults.reduce(
(total, result) => total + Math.max(0, result.attempt - 1),
0,
),
cleanupPassed: suiteResults.every(
(result) => result.cleanup === "passed",
),
complete: suiteExpected.length === suite.expectedMatrixSize,
durationMs: suiteResults.reduce(
(total, result) => total + result.durationMs,
0,
),
billing: aggregateCampaignBilling(suiteResults),
};
})
.filter((suite): suite is NonNullable<typeof suite> => Boolean(suite));
const passed = results.filter(
(result) => result.status === "passed" && result.cleanup === "passed",
).length;
const rankingSnapshots = [
...new Map(
results.flatMap((result) =>
result.rankingSnapshot
? [
[
result.rankingSnapshot.snapshotId,
{
snapshotId: result.rankingSnapshot.snapshotId,
capturedAt: result.rankingSnapshot.capturedAt,
sourceUrl: result.rankingSnapshot.sourceUrl,
},
] as const,
]
: [],
),
).values(),
];
return {
schema: "paperclip.runner-e2e.campaign/v2",
campaignId: input.campaignId,
generatedAt: input.generatedAt,
source,
expected,
complete:
suites.length === runnerSuites.length &&
suites.every((suite) => suite.complete),
selected: expected.length,
executed: results.filter((result) => result.attempt > 0).length,
passed,
failed: expected.length - passed,
retries: results.reduce(
(total, result) => total + Math.max(0, result.attempt - 1),
0,
),
cleanupPassed: results.every((result) => result.cleanup === "passed"),
rankingSnapshots,
billing: aggregateCampaignBilling(results),
suites,
results,
};
}
export function campaignHistoryRecord(
campaign: RunnerE2ECampaign,
publicBaseUrl: string,
): RunnerE2EHistoryCampaign {
const publicUrl = `${publicBaseUrl.replace(/\/$/, "")}/campaigns/${encodeURIComponent(campaign.campaignId)}/`;
return {
campaignId: campaign.campaignId,
generatedAt: campaign.generatedAt,
source: campaign.source,
complete: campaign.complete,
selected: campaign.selected,
executed: campaign.executed,
passed: campaign.passed,
failed: campaign.failed,
retries: campaign.retries,
cleanupPassed: campaign.cleanupPassed,
publicUrl,
billing: campaign.billing,
suites: campaign.suites,
executions: campaign.results.map((result) => ({
executionId: result.executionId,
suiteId: result.suiteId ?? "core-compatibility",
profileId: result.profileId,
environmentId: result.environmentId,
caseId: result.caseId,
provider: result.provider,
model: result.model,
status: result.status,
durationMs: result.durationMs,
attempt: result.attempt,
cleanup: result.cleanup,
billing: result.billing ?? summarizeExecutionBilling(result),
})),
};
}
export function emptyRunnerHistory(): RunnerE2EHistoryIndex {
return {
schema: "paperclip.runner-e2e.history/v1",
updatedAt: new Date(0).toISOString(),
latestCampaignId: null,
latestGreenCampaignId: null,
latestBySuite: {},
latestGreenBySuite: {},
campaigns: [],
};
}
export function mergeRunnerHistory(
current: RunnerE2EHistoryIndex | null | undefined,
campaign: RunnerE2EHistoryCampaign,
): RunnerE2EHistoryIndex {
const base =
current?.schema === "paperclip.runner-e2e.history/v1"
? current
: emptyRunnerHistory();
const campaigns = [
campaign,
...base.campaigns.filter(
(candidate) => candidate.campaignId !== campaign.campaignId,
),
].sort(
(left, right) =>
Date.parse(right.generatedAt) - Date.parse(left.generatedAt),
);
const latestBySuite: Record<string, string> = {};
const latestGreenBySuite: Record<string, string> = {};
for (const candidate of campaigns) {
for (const suite of candidate.suites) {
latestBySuite[suite.suiteId] ??= candidate.campaignId;
if (suite.complete && suite.failed === 0) {
latestGreenBySuite[suite.suiteId] ??= candidate.campaignId;
}
}
}
return {
schema: "paperclip.runner-e2e.history/v1",
updatedAt: new Date().toISOString(),
latestCampaignId: campaigns[0]?.campaignId ?? null,
latestGreenCampaignId:
campaigns.find(
(candidate) => candidate.complete && candidate.failed === 0,
)?.campaignId ?? null,
latestBySuite,
latestGreenBySuite,
campaigns,
};
}