Files
PaperClipAI/tests/runner-e2e/web-server-shutdown.test.ts
DottaandPaperclip 04546c82d5 fix(runner): reconnect Daytona sessions after controller restart (#13691)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner can execute a task inside a Daytona sandbox.
> - The sandbox can keep running when the Paperclip controller restarts.
> - Recovery treated sandbox process IDs as local process IDs and
selected the wrong recovery path.
> - Live verification also found races between startup, shutdown, and
queued task cleanup.
> - This pull request verifies the existing remote owner and orders
those transitions.
> - Users can continue the same task and provider session after a
controller restart.

## Linked Issues or Issue Description

**What happened?**

The Daytona `recover-controller` cases failed with
`runner_state_identity_mismatch`. Remote process IDs can be absent on
the controller or collide with unrelated local processes. Recovery then
looked for remote state in the local runner directory. Later turns could
also start before the previous executor released its sandbox resources.

**Expected behavior**

Reconnect to the original sandbox and authenticated runner. Preserve the
task, provider session, and queued comments. Reject a replacement
sandbox or mismatched identity. Do not start another provider during
reattachment.

**Steps to reproduce**

Run the `everyday-workflows` `recover-controller` case for
`runner-codex` or `runner-acpx-claude` in Daytona. The browser creates a
Python tool, requests a revision, restarts the controller during
execution, and queues another revision. It then downloads and tests the
final ZIP.

Related: #13682 is the preceding operational fix. #13291 addresses
legacy sandbox conversation recovery, a different execution path. #13666
includes broader run-capacity work; this change guards cleanup of an
existing native task executor.

## What Changed

- Add remote runner recovery without interpreting sandbox PIDs on the
controller.
- Verify the original provider lease, remote workspace, durable state,
process marker, and authenticated PRP authority before adoption.
- Compare the process marker with live Linux boot identity and start
ticks to reject PID reuse. Read virtual proc files through the
guaranteed Node runtime; unavailable proof blocks adoption without
blocking a fresh launch.
- Make the E2E supervisor own the actual server process so forced
restart cannot leave a late database closer behind.
- Scope the chat delivery lease test to its own fixture instead of
draining other tests’ pending deliveries.
- Preserve provider-attempt counts and recorded evidence during
reattachment.
- Serialize an idle-session checkpoint with admission of the next native
turn.
- Wait for an in-progress startup to acknowledge restart detachment.
Fail after a bounded deadline if it cannot.
- Keep a queued comment waiting until the previous native task executor
releases its resources. Allow unrelated tasks to continue.
- Update the Daytona image's resolved lock digest to match current
dependency manifests.
- Add classifier, ownership, process, startup, checkpoint, and
queued-admission regression tests. Document recovery behavior.

## Verification

- 415 focused tests passed across native execution, restart recovery,
workspace synchronization, queued admission, and real-process restart
tests. The final Node-based fingerprint change passed all 375
native-session tests.
- Runner harness unit tests: 394 passed. Chat integration shard 2: 335
passed after fixture isolation.
- The exact fingerprint command succeeded twice in a disposable Daytona
sandbox and returned the same identity; the sandbox was deleted.
- 11 real-process restart integration tests passed, including absent and
colliding remote PIDs.
- Repository typecheck and final build passed. Broad local checks found
machine-dependent database startup and timing failures; focused retries
passed. The final-revision PR pipeline is green. One unrelated browser
shard hit a five-second blank-page timeout on the first run and passed
its targeted retry.
- Final-revision local headed browser E2E:
`everyday-workflows.runner-acpx-claude.daytona.recover-controller`
passed on attempt 1 in 4.7 minutes, **40/40 checks**. Manual browser
inspection confirmed Done, all three ZIPs, and delivery of the queued
follow-up. All three runs succeeded using the same provider session. The
harness downloaded and independently tested the final artifact.
- Final-revision Daytona campaign:
https://github.com/paperclipai/paperclip/actions/runs/35463999611 —
**Codex passed first attempt (4.8 minutes); ACPX Claude passed first
attempt (6.1 minutes)**. Campaign aggregation/publication is finishing;
both test jobs succeeded.
- Greptile reviewed `beb08d8493b3286f5bb988dead369ff8c96a395d`: **5/5**,
no open findings.
- Staging browser verification is pending selection of a disposable
staging instance and removal of a Chrome extension UI block.

## Risks

- Recovery now depends on the original sandbox remaining available. A
replacement or mismatched identity still blocks adoption.
- Shutdown waits up to 30 seconds for a native startup to reach a safe
detach point. An unfinished startup returns a clear failure instead of a
false detach receipt.
- Queued native work on the same task waits for cleanup. Unrelated tasks
remain eligible.
- The image digest update rebuilds the Daytona runtime image. No
database migration or public API change is included.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code execution, and
browser tools. The runtime does not expose the exact deployed model ID
or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:13:34 -05:00

148 lines
6.5 KiB
TypeScript

import { spawn } from "node:child_process";
import { mkdtemp, readFile, rm, writeFile } from "node:fs/promises";
import { createRequire } from "node:module";
import os from "node:os";
import path from "node:path";
import { expect, it } from "vitest";
import { reserveRunnerE2EDatabasePort } from "./ports.js";
import { runnerE2EWebServerGracefulShutdown, runnerE2ETypeScriptProcessArgs } from "./web-server-command.js";
const require = createRequire(import.meta.url);
it("lets Playwright reap a restarted server through the production bounded shutdown policy", async () => {
const root = await mkdtemp(path.join(os.tmpdir(), "paperclip-playwright-shutdown-"));
const reservation = await reserveRunnerE2EDatabasePort(3100);
const port = reservation.port;
const supervisorPath = path.join(root, "supervisor.cjs");
const workerPath = path.join(root, "worker.cjs");
const stoppedPath = path.join(root, "stopped.json");
const pidsPath = path.join(root, "pids.json");
const configPath = path.join(root, "playwright.config.cjs");
const testModule = require.resolve("@playwright/test");
const cli = require.resolve("@playwright/test/cli");
let child: ReturnType<typeof spawn> | undefined;
let timer: ReturnType<typeof setTimeout> | undefined;
let reservationReleased = false;
try {
await writeFile(workerPath, `
const http = require('node:http');
const server = http.createServer((req, res) => {
res.end(process.env.REVISION);
if (req.url === '/restart') process.send('restart');
});
process.on('SIGTERM', () => server.close(() => process.exit(0)));
server.listen(${port}, '127.0.0.1');
`);
await writeFile(supervisorPath, `
const { spawn } = require('node:child_process');
const fs = require('node:fs');
let stopping = false;
let revision = 0;
let child;
const pids = [];
function start() {
child = spawn(process.execPath, [${JSON.stringify(workerPath)}], {
env: { ...process.env, REVISION: String(++revision) },
stdio: ['ignore', 'inherit', 'inherit', 'ipc'],
});
pids.push(child.pid);
fs.writeFileSync(${JSON.stringify(pidsPath)}, JSON.stringify(pids));
child.once('message', () => child.kill('SIGTERM'));
child.once('exit', () => {
if (!stopping) return start();
fs.writeFileSync(${JSON.stringify(stoppedPath)}, JSON.stringify({ revision, pids }));
});
}
process.on('SIGTERM', () => { stopping = true; child.kill('SIGTERM'); });
start();
`);
await writeFile(configPath, `module.exports = {
testDir: ${JSON.stringify(root)}, testMatch: 'shutdown.spec.cjs', workers: 1,
reporter: 'line', timeout: 10000,
webServer: {
command: ${JSON.stringify(`"${process.execPath}" "${supervisorPath}"`)},
url: 'http://127.0.0.1:${port}', timeout: 10000,
gracefulShutdown: ${JSON.stringify(runnerE2EWebServerGracefulShutdown)},
},
};`);
await writeFile(path.join(root, "shutdown.spec.cjs"), `
const { test, expect } = require(${JSON.stringify(testModule)});
test('restarts before teardown', async ({ request }) => {
await request.get('http://127.0.0.1:${port}/restart');
await expect.poll(async () => {
try { return await (await request.get('http://127.0.0.1:${port}')).text(); }
catch { return ''; }
}).toBe('2');
});
`);
await reservation.close();
reservationReleased = true;
child = spawn(process.execPath, [cli, "test", "--config", configPath], {
stdio: ["ignore", "pipe", "pipe"],
detached: process.platform !== "win32",
env: { ...process.env, CI: "1" },
});
let output = "";
child.stdout?.on("data", (chunk) => { output += chunk; });
child.stderr?.on("data", (chunk) => { output += chunk; });
const exit = await Promise.race([
new Promise<number | null>((resolve, reject) => {
child!.once("error", reject);
child!.once("exit", resolve);
}),
new Promise<never>((_, reject) => {
timer = setTimeout(() => reject(new Error(`Playwright cleanup stalled: ${output}`)), 20_000);
}),
]);
expect(exit, output).toBe(0);
const stopped = JSON.parse(await readFile(stoppedPath, "utf8"));
expect(stopped.revision).toBe(2);
expect(stopped.pids).toHaveLength(2);
for (const pid of stopped.pids) {
expect(() => process.kill(pid, 0)).toThrow();
}
await expect(fetch(`http://127.0.0.1:${port}`, { signal: AbortSignal.timeout(500) })).rejects.toThrow();
} finally {
if (timer) clearTimeout(timer);
if (child?.pid && child.exitCode === null && child.signalCode === null) {
try { process.kill(-child.pid, "SIGKILL"); } catch { child.kill("SIGKILL"); }
}
const pids: number[] = JSON.parse(await readFile(pidsPath, "utf8").catch(() => "[]"));
for (const pid of pids) {
try { process.kill(pid, "SIGKILL"); } catch { /* Already reaped. */ }
}
if (!reservationReleased) await reservation.close();
await rm(root, { recursive: true, force: true });
}
}, 25_000);
it("owns the actual server PID so a forced restart cannot leave a late database closer", async () => {
const root = await mkdtemp(path.join(os.tmpdir(), "paperclip-server-pid-"));
const entry = path.join(root, "server.mts");
let candidate: ReturnType<typeof spawn> | undefined;
let actualPid: number | undefined;
try {
await writeFile(entry, `process.send!({ pid: process.pid }); setInterval(() => {}, 1000);`);
candidate = spawn(process.execPath, runnerE2ETypeScriptProcessArgs(path.resolve(import.meta.dirname, "../.."), entry), { stdio: ["ignore", "pipe", "pipe", "ipc"] });
actualPid = await new Promise<number>((resolve, reject) => {
candidate!.once("message", (message: any) => resolve(message.pid));
candidate!.once("error", reject);
candidate!.once("exit", () => reject(new Error("Server exited before publishing its PID")));
});
expect(actualPid).toBe(candidate.pid);
const exited = new Promise(resolve => candidate!.once("exit", resolve));
candidate.kill("SIGKILL");
await exited;
expect(() => process.kill(actualPid!, 0)).toThrow();
} finally {
if (actualPid) { try { process.kill(actualPid, "SIGKILL"); } catch {} }
if (candidate && candidate.exitCode === null && candidate.signalCode === null) {
const exited = new Promise(resolve => candidate!.once("exit", resolve));
candidate.kill("SIGKILL");
await exited;
}
await rm(root, { recursive: true, force: true });
}
}, 10_000);