Files
PaperClipAI/tests/runner-e2e/types.ts
T
Dotta 1dceee9a4e fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path

> - Paperclip manages AI agent work and the execution state for each
task.
> - Remote agents run in sandbox environments such as Daytona.
> - Daytona keeps files while a sandbox is stopped, but deletion removes
those files.
> - Runner Codex did not copy successful remote workspace changes back
to the host workspace.
> - A warm sandbox could therefore hide data loss until Daytona replaced
or deleted the sandbox.
> - This pull request makes the host workspace durable after every
successful turn and keeps verified reusable sandboxes warm.
> - The benefit is reliable multi-turn work across warm reuse, restart,
stop, and sandbox replacement.

## Linked Issues or Issue Description

**What happened?**

A successful native Codex turn in Daytona could leave workspace changes
only in the remote sandbox. A later warm turn appeared to work because
it reused that filesystem. A replacement sandbox could start from stale
host data and lose the successful changes.

**Expected behavior**

Paperclip must merge each successful remote turn into the authoritative
host workspace before it completes the run. A verified warm lease may
reuse its remote files. A replacement lease must reconstruct the exact
durable workspace seed.

**Steps to reproduce**

1. Run Codex in a reusable Daytona environment.
2. Write a file during one successful turn.
3. Replace the Daytona sandbox before the next turn.
4. Observe that the next turn can start without the prior file on the
unpatched code.

Related remote workspace foundation: #10070.

## What Changed

- Added explicit `host_current`, `durable_seed`, and `adopt_remote`
workspace preparation modes.
- Added atomic, versioned native workspace descriptors and seed archives
under `PAPERCLIP_HOME`.
- Added real native sandbox export and three-way host merge before
terminal result completion.
- Added workspace-only recovery after a proposed result. Recovery does
not submit another provider turn or consume the provider retry budget.
- Added fail-closed handling when a sandbox with unexported changes is
gone.
- Kept healthy reusable Daytona sandboxes started for legacy Codex and
Runner Codex.
- Kept the Runner Codex process and provider session across verified
warm turns.
- Added the paid `daytona-warm-continuity` browser suite. It contains
exactly the legacy Codex and Runner Codex cells. Each cell performs
three measured turns.
- Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`.
No package script was added.
- Added no database migration. The metadata format is backward
compatible and idempotent.

## Verification

- `pnpm typecheck`
- `pnpm test:e2e:runner:unit` — 114 passed
- Native workspace, finalizer, session, and environment tests — 232
passed
- Daytona provider tests — 150 passed
- Workspace staging and merge tests — 98 passed
- Runner transport tests — 63 passed
- Legacy Codex restore tests — 5 passed
- Rust format and compile checks pass through root typecheck
- The paid Daytona suite was not run locally because the required
Daytona, OpenAI, and immutable image credentials are not present.

## Risks

- The main risk is an incorrect workspace identity or merge after a
crash. Durable descriptors bind the run, workspace, lease, provider
lease, local root, remote root, and baseline digest. Ambiguous evidence
fails closed.
- The host merge may conflict with concurrent host edits. The existing
three-way merge and exclusion rules handle this case and surface
failures.
- A deleted sandbox cannot recover unexported bytes. Paperclip now
blocks with `workspace_sync_out_unrecoverable` instead of reporting
success or rerunning the provider.
- There is no database migration. Descriptor writes and recovery are
atomic and idempotent.

## Model Used

OpenAI Codex with GPT-5. The run used agentic reasoning, repository
inspection, code execution, test execution, Git, and GitHub CLI tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-05 13:00:57 -05:00

395 lines
11 KiB
TypeScript

export const CREDENTIAL_NAMES = [
"OPENAI_API_KEY",
"ANTHROPIC_API_KEY",
"OPENROUTER_API_KEY",
"DAYTONA_API_KEY",
] as const;
export type CredentialName = (typeof CREDENTIAL_NAMES)[number];
export type RunnerGeneration = "legacy" | "native";
export type RunnerEnvironmentId = "local" | "daytona";
export type RunnerTaskWorkMode = "standard" | "planning" | "ask";
export type RunnerTaskFlow =
| "single_turn"
| "plan_revision_acceptance"
| "question_resume_completion"
| "plan_approval_completion"
| "warm_three_turn";
export interface SecretReference {
type: "secret_ref";
secretId: string;
version: "latest";
}
export type SecretReferenceMap = Partial<
Record<CredentialName, SecretReference>
>;
export interface AgentFixtureBuildInput {
environmentId: string;
environmentFixtureId: RunnerEnvironmentId;
workspacePath: string;
secretRefs: SecretReferenceMap;
executionId: string;
}
export interface EnvironmentFixtureBuildInput {
secretRefs: SecretReferenceMap;
daytonaImage?: string;
executionId: string;
}
export interface RunnerProfileFixture {
id: string;
label: string;
generation: RunnerGeneration;
groups: readonly string[];
adapterType: string;
provider: string;
model: string;
modelQualification: {
source:
| "adapter_constant"
| "qualified_runner_profile"
| "openrouter_rankings_snapshot";
qualificationId: string;
};
ranking?: {
rank: number;
canonicalModelId: string;
snapshotId: string;
capturedAt: string;
sourceUrl: string;
};
credential: Exclude<CredentialName, "DAYTONA_API_KEY">;
supportedEnvironments: readonly RunnerEnvironmentId[];
expectedRuntimeMode: RunnerGeneration;
expectedRuntimeMetadata: {
adapterType: string;
provider: string;
};
buildAgent(input: AgentFixtureBuildInput): Record<string, unknown>;
}
export interface EnvironmentFixture {
id: RunnerEnvironmentId;
/** Distinguishes materially different configurations that share a provider ID. */
configurationKey?: string;
label: string;
groups: readonly string[];
driver: "local" | "sandbox";
provider: "local" | "daytona";
credential?: "DAYTONA_API_KEY";
lifecycle: {
setup: "instance_managed" | "create_via_api";
probe: "run_context_via_api";
cleanup: "instance_shutdown" | "delete_via_api_and_destroy_leases";
};
expectedExecutionTarget: {
kind: "local" | "remote";
transport?: "sandbox";
};
buildEnvironment(
input: EnvironmentFixtureBuildInput,
): Record<string, unknown>;
}
export type Matcher =
| { kind: "message_exact"; expected: string }
| { kind: "message_contains"; expected: string }
| { kind: "message_occurrences"; expected: string; count: number }
| { kind: "message_regex"; pattern: string; flags?: string }
| { kind: "message_ordered"; expected: readonly string[] }
| { kind: "issue_status"; expected: string }
| { kind: "run_status"; expected: string }
| { kind: "runtime_mode"; expected: RunnerGeneration }
| { kind: "environment"; expected: RunnerEnvironmentId }
| { kind: "file_exists"; path: string }
| { kind: "file_exact"; path: string; expected: string }
| { kind: "file_contains"; path: string; expected: string }
| { kind: "artifact_exists"; name: string; mimeType?: string }
| { kind: "json_path"; path: string; expected: unknown }
| { kind: "json_schema"; schema: Record<string, unknown> };
export interface RunnerTaskFixture {
id: string;
label: string;
groups: readonly string[];
workMode: RunnerTaskWorkMode;
flow: RunnerTaskFlow;
expectedRunCount: number;
attemptTimeoutMs: Readonly<Record<RunnerEnvironmentId, number>>;
expectedTerminalState: {
issue: "done";
run: "succeeded";
};
buildTitle(nonce: string): string;
buildPrompt(nonce: string): string;
buildVisibleMarker(nonce: string): string;
buildRevisionRequest?(nonce: string): string;
buildFollowupMessages?(nonce: string): readonly [string, string];
turnTimeoutMs?: number;
buildQuestionAnswer?(nonce: string): {
optionLabel: string;
expectedMarker: string;
};
/** Restart the isolated Paperclip server after the waiting turn settles. */
restartServerBeforeQuestionAnswer?: boolean;
buildPlanMarkers?(nonce: string): {
draft: string;
revised: string;
};
buildMatchers(nonce: string, execution: MatrixExecution): readonly Matcher[];
}
export interface MatrixExecution {
id: string;
suite: RunnerSuiteFixture;
suiteDefinitionHash: string;
profile: RunnerProfileFixture;
environment: EnvironmentFixture;
task: RunnerTaskFixture;
groups: readonly string[];
requiredCredentials: readonly CredentialName[];
}
export interface RunnerSuiteFixture {
id: string;
label: string;
description: string;
groups: readonly string[];
profiles: readonly RunnerProfileFixture[];
environments: readonly EnvironmentFixture[];
tasks: readonly RunnerTaskFixture[];
excludedExecutionIds?: readonly string[];
expectedMatrixSize: number;
definitionMetadata?: Readonly<Record<string, unknown>>;
}
export interface MatrixJob {
executionId: string;
suiteId: string;
profileId: string;
credentialName: Exclude<CredentialName, "DAYTONA_API_KEY">;
environmentId: RunnerEnvironmentId;
caseId: string;
timeoutMinutes: number;
needsDaytona: boolean;
}
export type FailureClass =
| "candidate_failure"
| "provider_variance"
| "transient_infrastructure"
| "permanent_infrastructure"
| "secret_leak"
| "cleanup_failure";
export type RunnerE2ECostStatus =
| "reported"
| "estimated"
| "partial"
| "unpriced"
| "unavailable"
| "not_metered";
export interface RunnerE2ERuntimeUsage {
provider: RunnerEnvironmentId;
/** Sum of the selected Paperclip heartbeat-run spans. */
agentRunDurationMs: number;
/** Sum of provider lease windows when the environment exposes leases. */
leaseDurationMs: number | null;
leaseCount: number;
cpuCores?: number;
memoryGiB?: number;
diskGiB?: number;
estimatedListCostUsd?: number;
costStatus: "estimated" | "unavailable" | "not_metered";
costSource:
| "daytona_public_list_price"
| "provider_cost_unavailable"
| "local_not_metered";
pricingAsOf?: string;
pricingUrl?: string;
}
export interface RunnerE2EBillingSummary {
llm: {
runCount: number;
runsWithTokenUsage: number;
runsWithReportedCost: number;
inputTokens: number;
outputTokens: number;
cachedInputTokens: number;
totalTokens: number;
reportedCostUsd: number;
costStatus: Exclude<RunnerE2ECostStatus, "estimated" | "not_metered">;
};
runtime: RunnerE2ERuntimeUsage;
/** Provider-reported model spend only; never includes unknown/unpriced runs. */
reportedCostUsd: number;
/** Public-list-price estimate for metered execution infrastructure. */
estimatedRuntimeCostUsd: number;
/** Reported model subtotal plus the runtime list-price estimate. */
observedAndEstimatedCostUsd: number;
complete: boolean;
}
export interface RunnerE2EResult {
schema: "paperclip.runner-e2e.result/v1" | "paperclip.runner-e2e.result/v2";
executionId: string;
suiteId?: string;
suiteDefinitionHash?: string;
source?: {
sha: string | null;
ref: string | null;
workflowRunUrl: string | null;
};
rankingSnapshot?: {
snapshotId: string;
capturedAt: string;
sourceUrl: string;
rank: number;
canonicalModelId: string;
};
attempt: number;
status: "passed" | "failed";
failureClass?: FailureClass;
error?: string;
profileId: string;
environmentId: RunnerEnvironmentId;
caseId: string;
provider: string;
model: string;
runtimeMode: RunnerGeneration;
issueId?: string;
issueIdentifier?: string | null;
runIds?: string[];
turnTimings?: Array<{
turn: number;
submittedAt: string;
runStartedAt: string | null;
runFinishedAt: string | null;
schedulerLatencyMs: number | null;
runDurationMs: number | null;
responseLatencyMs: number | null;
runId: string;
leaseAcquisitionOutcome: "created" | "resumed" | "replacement" | "unknown";
}>;
startedAt: string;
finishedAt: string;
durationMs: number;
usage?: Record<string, unknown> | null;
runtimeUsage?: RunnerE2ERuntimeUsage;
billing?: RunnerE2EBillingSummary;
matcherResults?: Array<{
matcher: Matcher;
passed: boolean;
detail: string;
}>;
screenshots?: Array<{
id: string;
label: string;
file: string;
publication?: "public-runner-fixture";
}>;
cleanup: "not_started" | "passed" | "failed";
}
export interface RunnerE2ESuiteSummary {
suiteId: string;
suiteDefinitionHash: string;
expected: number;
selected: number;
executed: number;
passed: number;
failed: number;
retries: number;
cleanupPassed: boolean;
complete: boolean;
durationMs: number;
billing: RunnerE2EAggregateBillingSummary;
}
export interface RunnerE2EAggregateBillingSummary {
testCount: number;
agentRunDurationMs: number;
leaseDurationMs: number;
llm: RunnerE2EBillingSummary["llm"];
reportedLlmCostUsd: number;
estimatedRuntimeCostUsd: number;
observedAndEstimatedCostUsd: number;
testsWithCompleteBilling: number;
}
export interface RunnerE2ECampaign {
schema: "paperclip.runner-e2e.campaign/v2";
campaignId: string;
generatedAt: string;
source: {
sha: string | null;
ref: string | null;
workflowRunUrl: string | null;
eventName: string | null;
};
expected: string[];
complete: boolean;
selected: number;
executed: number;
passed: number;
failed: number;
retries: number;
cleanupPassed: boolean;
rankingSnapshots: Array<{
snapshotId: string;
capturedAt: string;
sourceUrl: string;
}>;
billing: RunnerE2EAggregateBillingSummary;
suites: RunnerE2ESuiteSummary[];
results: RunnerE2EResult[];
}
export interface RunnerE2EHistoryExecution {
executionId: string;
suiteId: string;
profileId: string;
environmentId: RunnerEnvironmentId;
caseId: string;
provider: string;
model: string;
status: "passed" | "failed";
durationMs: number;
attempt: number;
cleanup: RunnerE2EResult["cleanup"];
billing: RunnerE2EBillingSummary;
}
export interface RunnerE2EHistoryCampaign {
campaignId: string;
generatedAt: string;
source: RunnerE2ECampaign["source"];
complete: boolean;
selected: number;
executed: number;
passed: number;
failed: number;
retries: number;
cleanupPassed: boolean;
publicUrl: string;
billing: RunnerE2EAggregateBillingSummary;
suites: RunnerE2ESuiteSummary[];
executions: RunnerE2EHistoryExecution[];
}
export interface RunnerE2EHistoryIndex {
schema: "paperclip.runner-e2e.history/v1";
updatedAt: string;
latestCampaignId: string | null;
latestGreenCampaignId: string | null;
latestBySuite: Record<string, string>;
latestGreenBySuite: Record<string, string>;
campaigns: RunnerE2EHistoryCampaign[];
}