mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can keep one agent process and workspace across human review turns. > - Large workspaces must survive each transfer between Daytona and the host. > - Small fixtures do not cross the previous 32 MiB filename-output limit. > - Correct files alone do not prove that finalization and recovery have settled. > - This pull request adds an explicit three-turn test with 60,000 files and independent host checks. > - The test detects lost files, replaced processes, and stale finalization retries. ## Linked Issues or Issue Description Refs: #14253, #14314, #14315, #14402, #14420. This test covers the workspace Git streaming fix in #14253 and the Daytona archive validation and native finalization fixes consolidated from #14315 and #14314 into #14402. These runtime changes are merged into master. This PR adds regression coverage. ## What Changed - Add one explicit-only native Codex Daytona test. Broad matrix runs do not select it. - Create 60,000 small untracked files through ordinary provider execution. Independently check all host file contents and the 39,828,890-byte filename manifest after each turn. Each continuation changes all generated file contents, so stale host copies fail. - Check spaces, newlines, leading hyphens, Unicode, and glob characters in filenames. - Require committed native finalization, successful workspace receipts, no active transfer or runless cleanup, and no scheduled recovery before each continuation and after the final turn. - Use a fixed external instruction bundle so the same runner PID and process identity can continue across the three browser-driven turns. Managed agent folders intentionally stop the process for file collection after #14420. Set the 20-minute idle window on the environment, then verify the admitted policy on every run. Set a 25-minute Daytona auto-stop window for this large-file test. - Save public workspace-operation evidence when an E2E attempt fails. - Bound this large-file fixture to 15 minutes per turn and 50 minutes total, reserving five minutes outside the turns for setup, host verification, and cleanup. A measured CI continuation succeeded in 11 minutes 5 seconds, exceeding the ordinary warm fixture's 10-minute deadline. The ordinary fixture and all file, process, finalization, and cleanup assertions remain unchanged. ## Verification - Final PR head: `b7eed5517426507f5d7912edb8ae49163d1783b0`. - `pnpm test:e2e:runner:unit` — 56 files and 708 tests passed locally. The catalog retains all 407 existing cells and adds this one explicit-only cell. - `pnpm test:e2e:runner:typecheck` — passed locally. - [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/36640920233) — all gates passed, including repository typechecks, tests, build, browser suites, and canary dry run. The PR has 54 successful checks and two expected skips. Fresh Greptile reviewed all seven files at 5/5; all review threads are resolved. - [Live single-cell Daytona verification](https://github.com/paperclipai/paperclip/actions/runs/36637283902) — passed on the first attempt in 1,822,299 ms (30m 22s) on `227074b7b`, using native Codex `gpt-5.6-sol` and the verified Daytona image. All three turns independently verified every one of the 60,000 file contents, 39,828,890 filename bytes, and five unusual names. All runs committed with one stable runner PID/process fingerprint, native/provider sessions, runner instance, and sandbox. Final browser/download assertions, all nine matchers, explicit cleanup, and report publication passed. Results are published through the [Product E2E history](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/) and [eval hub](https://pages.paperclip.ing/evals/). - The final `b7eed5517` follow-up only extends the total allowance from 45 to 50 minutes, updates its catalog assertion/version, and documents the setup/cleanup margin. Per-turn limits and behavioral assertions are unchanged from the successful live run; that paid run was not repeated for this allowance-only follow-up. - The [first integrated-head live attempt](https://github.com/paperclipai/paperclip/actions/runs/36632127364) is retained: turn 1 passed, then turn 2 hit the old 10-minute deadline while finalizing. Its trace records successful completion after 11m 5s and successful cleanup. Process diagnostics also showed the intentional managed agent-folder stop boundary introduced in #14420. These observations motivated the larger turn budget and fixed external instructions. - The broad local `pnpm test:run` began alongside the build and encountered server setup and port-test failures. The affected setup suites and assertions passed on focused reruns after the build, using canonical macOS temporary paths where needed. The redundant broad local run was stopped; full-suite success is established by final-head CI, not by that local run. ## Risks - The live test makes provider calls and creates a billable Daytona sandbox. It runs only when explicitly selected and deletes its sandbox during cleanup. - The test creates 60,000 files and can take several minutes per transfer. Its longer retention window applies only to this fixture. - The test will fail on a runtime that does not include all three required fixes. ## Model Used OpenAI GPT-6 in Codex assisted with code, terminal tools, and test analysis. The exact serving model suffix and context-window size were not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
158 lines
9.1 KiB
TypeScript
158 lines
9.1 KiB
TypeScript
import { execFile } from "node:child_process";
|
|
import { readFile, readdir } from "node:fs/promises";
|
|
import path from "node:path";
|
|
import { promisify } from "node:util";
|
|
import type { RunnerTaskFixture } from "./types.js";
|
|
import type { RunnerApi } from "./api.js";
|
|
|
|
const exec = promisify(execFile);
|
|
export const GIT_STREAMING_FILE_COUNT = 60_000;
|
|
export const GIT_STREAMING_PARENT = [...Array.from({ length: 2 }, () => "nested-".repeat(30)), "storybook-output"].join("/");
|
|
export const gitStreamingFilename = (index: number) => `${"asset-".repeat(36)}${index}.js`;
|
|
const unusualNames = [" space ", "line\nbreak", "-option", "雪-💾", "wild[?]*"];
|
|
|
|
/** Seed an ordinary empty project repository before its first public task. */
|
|
export async function setupGitStreamingWorkspace(workspacePath: string): Promise<void> {
|
|
await exec("git", ["init", "-q"], { cwd: workspacePath });
|
|
await exec("git", ["-c", "user.name=Runner E2E", "-c", "user.email=runner-e2e@example.test", "commit", "--allow-empty", "-qm", "Empty E2E workspace"], { cwd: workspacePath });
|
|
}
|
|
|
|
export function createGitStreamingTask(base: RunnerTaskFixture): RunnerTaskFixture {
|
|
const setup = [
|
|
"python3 - <<'PY'",
|
|
"from pathlib import Path",
|
|
`parent = Path(${JSON.stringify(GIT_STREAMING_PARENT)})`,
|
|
"parent.mkdir(parents=True, exist_ok=True)",
|
|
`for index in range(${GIT_STREAMING_FILE_COUNT}):`,
|
|
" (parent / ('asset-' * 36 + str(index) + '.js')).write_text('base')",
|
|
`for name in ${JSON.stringify(unusualNames)}:`,
|
|
" Path(name).write_text(name)",
|
|
`print('generated', ${GIT_STREAMING_FILE_COUNT}, 'files')`,
|
|
"PY",
|
|
].join("\n");
|
|
const verify = (previousTurn: number, nextTurn: number) => [
|
|
"Before continuing, run this command and confirm it succeeds:",
|
|
"python3 - <<'PY'",
|
|
"from pathlib import Path",
|
|
`parent = Path(${JSON.stringify(GIT_STREAMING_PARENT)})`,
|
|
`assert len(list(parent.iterdir())) == ${GIT_STREAMING_FILE_COUNT}`,
|
|
`for index in range(${GIT_STREAMING_FILE_COUNT}):`,
|
|
` assert (parent / ('asset-' * 36 + str(index) + '.js')).read_text() == ${JSON.stringify(previousTurn === 1 ? "base" : `base-turn-${previousTurn}`)}`,
|
|
`for name in ${JSON.stringify(unusualNames)}:`,
|
|
" assert Path(name).read_text() == name",
|
|
`for index in range(${GIT_STREAMING_FILE_COUNT}):`,
|
|
` (parent / ('asset-' * 36 + str(index) + '.js')).write_text('base-turn-${nextTurn}')`,
|
|
"print('all generated files preserved')",
|
|
"PY",
|
|
"Keep these generated files untracked. Do not rename, remove, add, commit, or ignore them.",
|
|
].join("\n");
|
|
return {
|
|
...base,
|
|
id: "large-path-three-turn",
|
|
label: "Large Git filename manifest across three Daytona turns",
|
|
// CI needs more than the ordinary warm fixture's ten minutes to prepare
|
|
// and copy back 60,000 files (a measured continuation took eleven minutes).
|
|
turnTimeoutMs: 15 * 60_000,
|
|
// Reserve another five minutes for setup, host checks, and cleanup outside
|
|
// the three per-turn deadlines.
|
|
attemptTimeoutMs: { ...base.attemptTimeoutMs, daytona: 50 * 60_000 },
|
|
buildTitle: nonce => `Runner E2E Git streaming ${nonce}`,
|
|
buildPrompt: nonce => [
|
|
"Execute this bounded fixture command once in the current execution workspace. It creates 60,000 small untracked files whose Git filename list exceeds 32 MiB. Keep all files untracked; do not add, commit, rename, delete, or ignore them. Do not print the filenames or file contents.",
|
|
setup,
|
|
base.buildPrompt(nonce),
|
|
].join("\n"),
|
|
buildFollowupMessages: nonce => {
|
|
const [second, third] = base.buildFollowupMessages!(nonce);
|
|
return [`${verify(1, 2)}\n${second}`, `${verify(2, 3)}\n${third}`];
|
|
},
|
|
};
|
|
}
|
|
|
|
export function gradeGitStreamingInventory(names: string[]): { generatedFiles: number; filenameBytes: number } {
|
|
const expected = new Set(Array.from({ length: GIT_STREAMING_FILE_COUNT }, (_, index) => gitStreamingFilename(index)));
|
|
let filenameBytes = 0;
|
|
for (const name of names) {
|
|
if (!expected.delete(name)) throw new Error("Git copyback contains an unexpected or duplicate generated filename");
|
|
filenameBytes += Buffer.byteLength(`${GIT_STREAMING_PARENT}/${name}`) + 1;
|
|
}
|
|
if (expected.size) throw new Error(`Git copyback is missing ${expected.size} generated files`);
|
|
if (filenameBytes <= 32 * 1024 * 1024) throw new Error("Git filename fixture did not cross the 32 MiB boundary");
|
|
return { generatedFiles: names.length, filenameBytes };
|
|
}
|
|
|
|
/** Read every copied-back file independently of the agent's verification. */
|
|
export async function gitStreamingEvidence(workspacePath: string, completedTurn: number) {
|
|
if (![1, 2, 3].includes(completedTurn)) throw new Error("Git copyback requires an exact fixture turn");
|
|
const expectedContents = completedTurn === 1 ? "base" : `base-turn-${completedTurn}`;
|
|
const parent = path.join(workspacePath, GIT_STREAMING_PARENT);
|
|
const names = await readdir(parent);
|
|
const inventory = gradeGitStreamingInventory(names);
|
|
for (let start = 0; start < names.length; start += 100) {
|
|
await Promise.all(names.slice(start, start + 100).map(async name => {
|
|
if (await readFile(path.join(parent, name), "utf8") !== expectedContents) throw new Error("Git copyback changed generated file contents or retained an earlier turn");
|
|
}));
|
|
}
|
|
for (const name of unusualNames) {
|
|
if (await readFile(path.join(workspacePath, name), "utf8") !== name) throw new Error("Git copyback changed an unusual filename or its contents");
|
|
}
|
|
return { ...inventory, verifiedContents: names.length, verifiedTurn: completedTurn, unusualNamesPreserved: unusualNames.length };
|
|
}
|
|
|
|
interface FinalizedRun {
|
|
id: string;
|
|
status: string;
|
|
nativePhase?: string | null;
|
|
resultJson?: { finalizationPhase?: string; workspaceFinalizeStatus?: string; nextAttemptAt?: string | null; failureCode?: string | null } | null;
|
|
runnerProfileJson?: { nativeExecutionInput?: { session?: { lifecyclePolicy?: { mode?: string; idleTimeoutMs?: number | null } } } } | null;
|
|
}
|
|
interface WorkspaceOperation { heartbeatRunId: string | null; status: string }
|
|
interface RecoveryAction { status: string; wakePolicy?: { kind?: string } | null }
|
|
interface GitFinalizationObservation {
|
|
runs: FinalizedRun[];
|
|
operations: WorkspaceOperation[];
|
|
recovery: { active: RecoveryAction | null; actions: RecoveryAction[] };
|
|
scheduledRetry: unknown;
|
|
}
|
|
|
|
export function gradeGitFinalization(observation: GitFinalizationObservation) {
|
|
const failures: string[] = [];
|
|
if (observation.runs.length === 0) failures.push("No completed native run evidence");
|
|
for (const run of observation.runs) {
|
|
const policy = run.runnerProfileJson?.nativeExecutionInput?.session?.lifecyclePolicy;
|
|
if (policy?.mode !== "warm" || policy.idleTimeoutMs !== 1_200_000) {
|
|
failures.push(`Run ${run.id} did not admit the required 20-minute warm idle policy`);
|
|
}
|
|
if (run.status !== "succeeded" || run.nativePhase !== "committed" ||
|
|
run.resultJson?.finalizationPhase !== "committed" || run.resultJson.workspaceFinalizeStatus !== "succeeded") {
|
|
failures.push(`Run ${run.id} is not durably committed after workspace finalization`);
|
|
}
|
|
if (run.resultJson?.nextAttemptAt != null) failures.push(`Run ${run.id} still schedules a finalization retry`);
|
|
if (run.resultJson?.failureCode != null) failures.push(`Run ${run.id} still projects an active finalization failure`);
|
|
const operations = observation.operations.filter(operation => operation.heartbeatRunId === run.id);
|
|
if (!operations.some(operation => operation.status === "succeeded")) failures.push(`Run ${run.id} has no successful workspace receipt`);
|
|
}
|
|
// The scoped endpoint also includes workspace cleanup with no run ID.
|
|
if (observation.operations.some(operation => ["pending", "queued", "running"].includes(operation.status))) {
|
|
failures.push("The workspace still has an active operation");
|
|
}
|
|
const actions = [observation.recovery.active, ...observation.recovery.actions].filter(action => action !== null);
|
|
if (actions.some(action => ["active", "pending"].includes(action.status) && action.wakePolicy?.kind === "resume_native_run")) {
|
|
failures.push("A recovery action still schedules a native run retry");
|
|
}
|
|
if (observation.scheduledRetry != null) failures.push("The task still has a scheduled retry");
|
|
return { passed: failures.length === 0, failures };
|
|
}
|
|
|
|
/** Public durable receipts must agree before the next browser turn is sent. */
|
|
export async function gitFinalizationEvidence(api: RunnerApi, issueId: string, runIds: string[]) {
|
|
const [runs, operationLists, recovery, issue] = await Promise.all([
|
|
Promise.all(runIds.map(id => api.get<FinalizedRun>(`/api/heartbeat-runs/${id}`))),
|
|
Promise.all(runIds.map(id => api.get<WorkspaceOperation[]>(`/api/heartbeat-runs/${id}/workspace-operations`))),
|
|
api.get<GitFinalizationObservation["recovery"]>(`/api/issues/${issueId}/recovery-actions`),
|
|
api.get<{ scheduledRetry?: unknown }>(`/api/issues/${issueId}`),
|
|
]);
|
|
const observation: GitFinalizationObservation = { runs, operations: operationLists.flat(), recovery, scheduledRetry: issue.scheduledRetry };
|
|
return { ...observation, ...gradeGitFinalization(observation) };
|
|
}
|