Files
PaperClipAI/tests/runner-e2e/daytona-git-streaming.ts
T
DottaandPaperclip fc36d1ccd3 test(e2e): qualify large Daytona Git workspace continuation (#14316)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A task can keep one agent process and workspace across human review
turns.
> - Large workspaces must survive each transfer between Daytona and the
host.
> - Small fixtures do not cross the previous 32 MiB filename-output
limit.
> - Correct files alone do not prove that finalization and recovery have
settled.
> - This pull request adds an explicit three-turn test with 60,000 files
and independent host checks.
> - The test detects lost files, replaced processes, and stale
finalization retries.

## Linked Issues or Issue Description

Refs: #14253, #14314, #14315, #14402, #14420.

This test covers the workspace Git streaming fix in #14253 and the
Daytona archive validation and native finalization fixes consolidated
from #14315 and #14314 into #14402. These runtime changes are merged
into master. This PR adds regression coverage.

## What Changed

- Add one explicit-only native Codex Daytona test. Broad matrix runs do
not select it.
- Create 60,000 small untracked files through ordinary provider
execution. Independently check all host file contents and the
39,828,890-byte filename manifest after each turn. Each continuation
changes all generated file contents, so stale host copies fail.
- Check spaces, newlines, leading hyphens, Unicode, and glob characters
in filenames.
- Require committed native finalization, successful workspace receipts,
no active transfer or runless cleanup, and no scheduled recovery before
each continuation and after the final turn.
- Use a fixed external instruction bundle so the same runner PID and
process identity can continue across the three browser-driven turns.
Managed agent folders intentionally stop the process for file collection
after #14420. Set the 20-minute idle window on the environment, then
verify the admitted policy on every run. Set a 25-minute Daytona
auto-stop window for this large-file test.
- Save public workspace-operation evidence when an E2E attempt fails.
- Bound this large-file fixture to 15 minutes per turn and 50 minutes
total, reserving five minutes outside the turns for setup, host
verification, and cleanup. A measured CI continuation succeeded in 11
minutes 5 seconds, exceeding the ordinary warm fixture's 10-minute
deadline. The ordinary fixture and all file, process, finalization, and
cleanup assertions remain unchanged.

## Verification

- Final PR head: `b7eed5517426507f5d7912edb8ae49163d1783b0`.
- `pnpm test:e2e:runner:unit` — 56 files and 708 tests passed locally.
The catalog retains all 407 existing cells and adds this one
explicit-only cell.
- `pnpm test:e2e:runner:typecheck` — passed locally.
- [Final-head
CI](https://github.com/paperclipai/paperclip/actions/runs/36640920233) —
all gates passed, including repository typechecks, tests, build, browser
suites, and canary dry run. The PR has 54 successful checks and two
expected skips. Fresh Greptile reviewed all seven files at 5/5; all
review threads are resolved.
- [Live single-cell Daytona
verification](https://github.com/paperclipai/paperclip/actions/runs/36637283902)
— passed on the first attempt in 1,822,299 ms (30m 22s) on `227074b7b`,
using native Codex `gpt-5.6-sol` and the verified Daytona image. All
three turns independently verified every one of the 60,000 file
contents, 39,828,890 filename bytes, and five unusual names. All runs
committed with one stable runner PID/process fingerprint,
native/provider sessions, runner instance, and sandbox. Final
browser/download assertions, all nine matchers, explicit cleanup, and
report publication passed. Results are published through the [Product
E2E history](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/) and
[eval hub](https://pages.paperclip.ing/evals/).
- The final `b7eed5517` follow-up only extends the total allowance from
45 to 50 minutes, updates its catalog assertion/version, and documents
the setup/cleanup margin. Per-turn limits and behavioral assertions are
unchanged from the successful live run; that paid run was not repeated
for this allowance-only follow-up.
- The [first integrated-head live
attempt](https://github.com/paperclipai/paperclip/actions/runs/36632127364)
is retained: turn 1 passed, then turn 2 hit the old 10-minute deadline
while finalizing. Its trace records successful completion after 11m 5s
and successful cleanup. Process diagnostics also showed the intentional
managed agent-folder stop boundary introduced in #14420. These
observations motivated the larger turn budget and fixed external
instructions.
- The broad local `pnpm test:run` began alongside the build and
encountered server setup and port-test failures. The affected setup
suites and assertions passed on focused reruns after the build, using
canonical macOS temporary paths where needed. The redundant broad local
run was stopped; full-suite success is established by final-head CI, not
by that local run.

## Risks

- The live test makes provider calls and creates a billable Daytona
sandbox. It runs only when explicitly selected and deletes its sandbox
during cleanup.
- The test creates 60,000 files and can take several minutes per
transfer. Its longer retention window applies only to this fixture.
- The test will fail on a runtime that does not include all three
required fixes.

## Model Used

OpenAI GPT-6 in Codex assisted with code, terminal tools, and test
analysis. The exact serving model suffix and context-window size were
not exposed to this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 17:49:36 -05:00

158 lines
9.1 KiB
TypeScript

import { execFile } from "node:child_process";
import { readFile, readdir } from "node:fs/promises";
import path from "node:path";
import { promisify } from "node:util";
import type { RunnerTaskFixture } from "./types.js";
import type { RunnerApi } from "./api.js";
const exec = promisify(execFile);
export const GIT_STREAMING_FILE_COUNT = 60_000;
export const GIT_STREAMING_PARENT = [...Array.from({ length: 2 }, () => "nested-".repeat(30)), "storybook-output"].join("/");
export const gitStreamingFilename = (index: number) => `${"asset-".repeat(36)}${index}.js`;
const unusualNames = [" space ", "line\nbreak", "-option", "雪-💾", "wild[?]*"];
/** Seed an ordinary empty project repository before its first public task. */
export async function setupGitStreamingWorkspace(workspacePath: string): Promise<void> {
await exec("git", ["init", "-q"], { cwd: workspacePath });
await exec("git", ["-c", "user.name=Runner E2E", "-c", "user.email=runner-e2e@example.test", "commit", "--allow-empty", "-qm", "Empty E2E workspace"], { cwd: workspacePath });
}
export function createGitStreamingTask(base: RunnerTaskFixture): RunnerTaskFixture {
const setup = [
"python3 - <<'PY'",
"from pathlib import Path",
`parent = Path(${JSON.stringify(GIT_STREAMING_PARENT)})`,
"parent.mkdir(parents=True, exist_ok=True)",
`for index in range(${GIT_STREAMING_FILE_COUNT}):`,
" (parent / ('asset-' * 36 + str(index) + '.js')).write_text('base')",
`for name in ${JSON.stringify(unusualNames)}:`,
" Path(name).write_text(name)",
`print('generated', ${GIT_STREAMING_FILE_COUNT}, 'files')`,
"PY",
].join("\n");
const verify = (previousTurn: number, nextTurn: number) => [
"Before continuing, run this command and confirm it succeeds:",
"python3 - <<'PY'",
"from pathlib import Path",
`parent = Path(${JSON.stringify(GIT_STREAMING_PARENT)})`,
`assert len(list(parent.iterdir())) == ${GIT_STREAMING_FILE_COUNT}`,
`for index in range(${GIT_STREAMING_FILE_COUNT}):`,
` assert (parent / ('asset-' * 36 + str(index) + '.js')).read_text() == ${JSON.stringify(previousTurn === 1 ? "base" : `base-turn-${previousTurn}`)}`,
`for name in ${JSON.stringify(unusualNames)}:`,
" assert Path(name).read_text() == name",
`for index in range(${GIT_STREAMING_FILE_COUNT}):`,
` (parent / ('asset-' * 36 + str(index) + '.js')).write_text('base-turn-${nextTurn}')`,
"print('all generated files preserved')",
"PY",
"Keep these generated files untracked. Do not rename, remove, add, commit, or ignore them.",
].join("\n");
return {
...base,
id: "large-path-three-turn",
label: "Large Git filename manifest across three Daytona turns",
// CI needs more than the ordinary warm fixture's ten minutes to prepare
// and copy back 60,000 files (a measured continuation took eleven minutes).
turnTimeoutMs: 15 * 60_000,
// Reserve another five minutes for setup, host checks, and cleanup outside
// the three per-turn deadlines.
attemptTimeoutMs: { ...base.attemptTimeoutMs, daytona: 50 * 60_000 },
buildTitle: nonce => `Runner E2E Git streaming ${nonce}`,
buildPrompt: nonce => [
"Execute this bounded fixture command once in the current execution workspace. It creates 60,000 small untracked files whose Git filename list exceeds 32 MiB. Keep all files untracked; do not add, commit, rename, delete, or ignore them. Do not print the filenames or file contents.",
setup,
base.buildPrompt(nonce),
].join("\n"),
buildFollowupMessages: nonce => {
const [second, third] = base.buildFollowupMessages!(nonce);
return [`${verify(1, 2)}\n${second}`, `${verify(2, 3)}\n${third}`];
},
};
}
export function gradeGitStreamingInventory(names: string[]): { generatedFiles: number; filenameBytes: number } {
const expected = new Set(Array.from({ length: GIT_STREAMING_FILE_COUNT }, (_, index) => gitStreamingFilename(index)));
let filenameBytes = 0;
for (const name of names) {
if (!expected.delete(name)) throw new Error("Git copyback contains an unexpected or duplicate generated filename");
filenameBytes += Buffer.byteLength(`${GIT_STREAMING_PARENT}/${name}`) + 1;
}
if (expected.size) throw new Error(`Git copyback is missing ${expected.size} generated files`);
if (filenameBytes <= 32 * 1024 * 1024) throw new Error("Git filename fixture did not cross the 32 MiB boundary");
return { generatedFiles: names.length, filenameBytes };
}
/** Read every copied-back file independently of the agent's verification. */
export async function gitStreamingEvidence(workspacePath: string, completedTurn: number) {
if (![1, 2, 3].includes(completedTurn)) throw new Error("Git copyback requires an exact fixture turn");
const expectedContents = completedTurn === 1 ? "base" : `base-turn-${completedTurn}`;
const parent = path.join(workspacePath, GIT_STREAMING_PARENT);
const names = await readdir(parent);
const inventory = gradeGitStreamingInventory(names);
for (let start = 0; start < names.length; start += 100) {
await Promise.all(names.slice(start, start + 100).map(async name => {
if (await readFile(path.join(parent, name), "utf8") !== expectedContents) throw new Error("Git copyback changed generated file contents or retained an earlier turn");
}));
}
for (const name of unusualNames) {
if (await readFile(path.join(workspacePath, name), "utf8") !== name) throw new Error("Git copyback changed an unusual filename or its contents");
}
return { ...inventory, verifiedContents: names.length, verifiedTurn: completedTurn, unusualNamesPreserved: unusualNames.length };
}
interface FinalizedRun {
id: string;
status: string;
nativePhase?: string | null;
resultJson?: { finalizationPhase?: string; workspaceFinalizeStatus?: string; nextAttemptAt?: string | null; failureCode?: string | null } | null;
runnerProfileJson?: { nativeExecutionInput?: { session?: { lifecyclePolicy?: { mode?: string; idleTimeoutMs?: number | null } } } } | null;
}
interface WorkspaceOperation { heartbeatRunId: string | null; status: string }
interface RecoveryAction { status: string; wakePolicy?: { kind?: string } | null }
interface GitFinalizationObservation {
runs: FinalizedRun[];
operations: WorkspaceOperation[];
recovery: { active: RecoveryAction | null; actions: RecoveryAction[] };
scheduledRetry: unknown;
}
export function gradeGitFinalization(observation: GitFinalizationObservation) {
const failures: string[] = [];
if (observation.runs.length === 0) failures.push("No completed native run evidence");
for (const run of observation.runs) {
const policy = run.runnerProfileJson?.nativeExecutionInput?.session?.lifecyclePolicy;
if (policy?.mode !== "warm" || policy.idleTimeoutMs !== 1_200_000) {
failures.push(`Run ${run.id} did not admit the required 20-minute warm idle policy`);
}
if (run.status !== "succeeded" || run.nativePhase !== "committed" ||
run.resultJson?.finalizationPhase !== "committed" || run.resultJson.workspaceFinalizeStatus !== "succeeded") {
failures.push(`Run ${run.id} is not durably committed after workspace finalization`);
}
if (run.resultJson?.nextAttemptAt != null) failures.push(`Run ${run.id} still schedules a finalization retry`);
if (run.resultJson?.failureCode != null) failures.push(`Run ${run.id} still projects an active finalization failure`);
const operations = observation.operations.filter(operation => operation.heartbeatRunId === run.id);
if (!operations.some(operation => operation.status === "succeeded")) failures.push(`Run ${run.id} has no successful workspace receipt`);
}
// The scoped endpoint also includes workspace cleanup with no run ID.
if (observation.operations.some(operation => ["pending", "queued", "running"].includes(operation.status))) {
failures.push("The workspace still has an active operation");
}
const actions = [observation.recovery.active, ...observation.recovery.actions].filter(action => action !== null);
if (actions.some(action => ["active", "pending"].includes(action.status) && action.wakePolicy?.kind === "resume_native_run")) {
failures.push("A recovery action still schedules a native run retry");
}
if (observation.scheduledRetry != null) failures.push("The task still has a scheduled retry");
return { passed: failures.length === 0, failures };
}
/** Public durable receipts must agree before the next browser turn is sent. */
export async function gitFinalizationEvidence(api: RunnerApi, issueId: string, runIds: string[]) {
const [runs, operationLists, recovery, issue] = await Promise.all([
Promise.all(runIds.map(id => api.get<FinalizedRun>(`/api/heartbeat-runs/${id}`))),
Promise.all(runIds.map(id => api.get<WorkspaceOperation[]>(`/api/heartbeat-runs/${id}/workspace-operations`))),
api.get<GitFinalizationObservation["recovery"]>(`/api/issues/${issueId}/recovery-actions`),
api.get<{ scheduledRetry?: unknown }>(`/api/issues/${issueId}`),
]);
const observation: GitFinalizationObservation = { runs, operations: operationLists.flat(), recovery, scheduledRetry: issue.scheduledRetry };
return { ...observation, ...gradeGitFinalization(observation) };
}