Files
PaperClipAI/server/src/dev-server-status.ts
T
Dotta 7b094724e6 fix(runner): recover native sessions across restarts (#12845)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Paperclip Runner keeps durable run and provider state outside
one server process.
> - A server restart can leave that runner alive or can interrupt it
after a provider checkpoint.
> - The old startup path used handoff intent and PID evidence, but it
did not reconstruct native ownership.
> - That gap could block the issue, create a replacement run, or start
duplicate provider work.
> - This pull request adds durable same-run recovery for coordinated and
uncoordinated restarts.
> - The benefit is exact recovery of the run, runner, session, provider,
steering, and finalization state.

## Linked Issues or Issue Description

Refs #9628. That pull request added earlier local-adapter hot-restart
work. This change adds native PRP authority reconstruction and same-run
provider resume.

Refs #10935. That pull request handles missing hot-restart snapshots.
This change also supports hard restarts with no snapshot.

Refs #11624. That pull request prevents unsafe retry after an adopted
legacy process exits. This change reconciles native terminal evidence
before provider recovery.

Refs #12070. That pull request improves process liveness checks. This
change also binds recovery to a process-start fingerprint and fails
closed on ambiguity.

**What happened?**

The server could record hot-restart intent, but startup did not rebuild
native runner ownership. A live runner could not re-register its PRP
authority. A dead runner could not resume the exact native and provider
session on the same heartbeat run. Generic recovery could then block the
issue or create replacement work.

**Expected behavior**

A live native runner must reconnect with the same PID and logical
identities. A dead runner must resume the same durable session and
heartbeat run with only a new operating-system PID. A proposed or
terminal result must finalize once before any provider turn starts.
Ambiguous process or session evidence must stay blocked without a signal
or duplicate spawn.

**Steps to reproduce**

1. Start a Paperclip Runner heartbeat and wait for an active provider
turn.
2. Restart only the Paperclip server, with or without a hot-restart
marker.
3. Observe that the old startup path does not reconstruct the native
control-plane authority.
4. Kill both the server and runner after a provider checkpoint.
5. Observe that the old path cannot resume the exact native session on
the original heartbeat run.

**Paperclip version or commit**

The defect was reproduced from commit
`1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto
the current `master`.

**Deployment mode**

Local development and self-hosted server deployments that use the local
Paperclip Runner.

## What Changed

- Added correlated hot-restart requests and version-compatible native
handoff fields.
- Added controller boot identity, process-start identity, controller
generation, recovery state, request id, and bounded history to the
native finalization ledger.
- Added transactional recovery claims for live-runner reattach,
dead-runner resume, and incomplete bootstrap.
- Added fail-closed ownership takeover rules and process identity
validation.
- Added live runner adoption to the local runner transport without a
duplicate spawn.
- Added same-run provider checkpoint resume and legacy retry-row
compatibility.
- Reconciled proposed and terminal results before runner or provider
recovery.
- Bound the HTTP and PRP listener before startup recovery and delayed
scheduling and generic reapers until classification completes.
- Added restart-aware health diagnostics, run-log recovery transitions,
durable runner diagnostics, and bounded shutdown finalizer draining.
- Moved restart-survivable diagnostics into runner-owned, pre-redacted
bounded writes; raw stdout and stderr are never persisted.
- Added process-start fencing for controller, runner, and provider PIDs;
startup classifies every candidate without an implicit cap.
- Added crash-recoverable, contention-safe development restart-request
coordination and failed-startup listener cleanup.
- Added a credential-free real-process restart suite for eight restart,
scale, and identity scenarios.
- Documented native restart operation, persistence, diagnostics, and
verification.

## Verification

- The documented native restart commands passed. They ran eight
real-process/database recovery scenarios and the live runner adoption
transport test.
- Native executor tests passed: 111 tests.
- Heartbeat recovery tests passed: 124 tests.
- Hot restart, health, and shutdown tests passed: 52 tests.
- The broader affected server suite passed: 350 tests.
- Focused native recovery and startup tests passed: 49 tests.
- Runner transport and control-plane tests passed: 63 tests.
- Runner-owned diagnostic tests passed for write-time bounding,
credential redaction, private file modes, and raw stream
non-persistence.
- Development restart coordination tests passed: 11 tests.
- Database migration checks and the partial-application/replay
regression test passed.
- Server, database, and Paperclip Runner typechecks passed.
- `git diff --check` passed.
- Full Paperclip PR CI passed, including build, canary, all five general
server shards, all five serialized server shards, all three browser E2E
shards, workspace suites, and release-registry verification.
- Greptile completed at 5/5 with no outstanding findings,
recommendations, follow-ups, or open review threads.

## Risks

- Moderate risk. This changes startup ordering and ownership transfer
for active native runs.
- The migration adds nullable columns and does not rewrite existing
rows.
- Recovery fails closed when process or durable session identity is
incomplete or contradictory.
- The first implementation supports the local Paperclip Runner. Remote
targets keep their existing behavior.
- The real-process suite covers cleanup and asserts that no runner or
provider process survives each test.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with GPT-5. The runtime did not expose a more specific
model revision or context-window size. Repository editing, shell
execution, database tests, and real-process test execution were enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-04 15:03:53 -05:00

334 lines
9.8 KiB
TypeScript

import { randomUUID } from "node:crypto";
import {
existsSync,
mkdirSync,
readFileSync,
renameSync,
rmSync,
statSync,
unlinkSync,
writeFileSync,
} from "node:fs";
import path from "node:path";
const MAX_PERSISTED_DEV_SERVER_STATUS_BYTES = 64 * 1024;
const DEV_RESTART_REQUEST_LOCK_STALE_MS = 30_000;
const DEV_RESTART_REQUEST_LOCK_RETRY_COUNT = 50;
const DEV_RESTART_REQUEST_LOCK_RETRY_MS = 2;
const lockWaitBuffer = new Int32Array(new SharedArrayBuffer(4));
export type PersistedDevServerStatus = {
dirty: boolean;
lastChangedAt: string | null;
changedPathCount: number;
changedPathsSample: string[];
pendingMigrations: string[];
lastRestartAt: string | null;
};
export type DevServerHealthStatus = {
enabled: true;
restartRequired: boolean;
reason:
| "backend_changes"
| "pending_migrations"
| "backend_changes_and_pending_migrations"
| null;
lastChangedAt: string | null;
changedPathCount: number;
changedPathsSample: string[];
pendingMigrations: string[];
autoRestartEnabled: boolean;
activeRunCount: number;
waitingForIdle: boolean;
lastRestartAt: string | null;
};
export type DevServerRestartRequest = {
requestedAt: string;
reason: "manual_restart_now";
requestId?: string;
mode?: "hot";
previousServerIdentity?: string;
};
function processIsAlive(pid: number): boolean {
try {
process.kill(pid, 0);
return true;
} catch (error) {
return (error as NodeJS.ErrnoException | undefined)?.code === "EPERM";
}
}
function tryRecoverDevRestartRequestLock(lockPath: string): boolean {
let stale = false;
try {
const lockAgeMs = Date.now() - statSync(lockPath).mtimeMs;
const owner = JSON.parse(
readFileSync(path.join(lockPath, "owner.json"), "utf8"),
) as Record<string, unknown>;
stale =
lockAgeMs >= DEV_RESTART_REQUEST_LOCK_STALE_MS ||
(typeof owner.pid === "number" && !processIsAlive(owner.pid));
} catch {
// Canonical locks are published only after owner.json is durable in a
// private candidate directory. A visible lock without valid ownership is
// therefore abandoned and can be reclaimed immediately.
stale = true;
}
if (!stale) return false;
const stalePath = `${lockPath}.${process.pid}.${randomUUID()}.stale`;
try {
renameSync(lockPath, stalePath);
} catch (error) {
if ((error as NodeJS.ErrnoException | undefined)?.code === "ENOENT") {
return true;
}
return false;
}
rmSync(stalePath, { recursive: true, force: true });
return true;
}
function withDevRestartRequestLock<T>(filePath: string, action: () => T): T {
const lockPath = `${filePath}.lock`;
for (
let attempt = 0;
attempt <= DEV_RESTART_REQUEST_LOCK_RETRY_COUNT;
attempt += 1
) {
const candidateLockPath = `${lockPath}.${process.pid}.${randomUUID()}.candidate`;
mkdirSync(candidateLockPath);
try {
writeFileSync(
path.join(candidateLockPath, "owner.json"),
`${JSON.stringify({ pid: process.pid, acquiredAt: new Date().toISOString() })}\n`,
"utf8",
);
try {
// Publishing a populated directory makes lock ownership visible in one
// rename; there is no canonical owner-less crash window.
renameSync(candidateLockPath, lockPath);
} catch (error) {
const code = (error as NodeJS.ErrnoException | undefined)?.code;
if (code !== "EEXIST" && code !== "ENOTEMPTY" && code !== "EPERM") {
throw error;
}
if (tryRecoverDevRestartRequestLock(lockPath)) continue;
if (attempt === DEV_RESTART_REQUEST_LOCK_RETRY_COUNT) {
throw new Error("dev_server_restart_request_lock_busy");
}
Atomics.wait(
lockWaitBuffer,
0,
0,
DEV_RESTART_REQUEST_LOCK_RETRY_MS,
);
continue;
}
} finally {
if (existsSync(candidateLockPath)) {
rmSync(candidateLockPath, { recursive: true, force: true });
}
}
try {
return action();
} finally {
rmSync(lockPath, { recursive: true, force: true });
}
}
throw new Error("dev_server_restart_request_lock_busy");
}
export function getDevServerRestartRequestFilePath(
env: NodeJS.ProcessEnv = process.env,
): string | null {
const statusFilePath = env.PAPERCLIP_DEV_SERVER_STATUS_FILE?.trim();
if (!statusFilePath) return null;
return path.join(
path.dirname(statusFilePath),
"dev-server-restart-request.json",
);
}
export function writeDevServerRestartRequest(
request: DevServerRestartRequest,
env: NodeJS.ProcessEnv = process.env,
): boolean {
const filePath = getDevServerRestartRequestFilePath(env);
if (!filePath) return false;
mkdirSync(path.dirname(filePath), { recursive: true });
withDevRestartRequestLock(filePath, () => {
const tempPath = `${filePath}.${process.pid}.${randomUUID()}.tmp`;
try {
writeFileSync(tempPath, `${JSON.stringify(request, null, 2)}\n`, "utf8");
renameSync(tempPath, filePath);
} finally {
try {
unlinkSync(tempPath);
} catch (error) {
if ((error as NodeJS.ErrnoException | undefined)?.code !== "ENOENT") {
throw error;
}
}
}
});
return true;
}
function readDevServerRestartRequestAtPath(
filePath: string,
): DevServerRestartRequest | null {
try {
if (statSync(filePath).size > MAX_PERSISTED_DEV_SERVER_STATUS_BYTES)
return null;
const value = JSON.parse(readFileSync(filePath, "utf8")) as Record<
string,
unknown
>;
if (
typeof value.requestedAt !== "string" ||
value.reason !== "manual_restart_now"
) {
return null;
}
return {
requestedAt: value.requestedAt,
reason: "manual_restart_now",
...(typeof value.requestId === "string"
? { requestId: value.requestId }
: {}),
...(value.mode === "hot" ? { mode: "hot" as const } : {}),
...(typeof value.previousServerIdentity === "string"
? { previousServerIdentity: value.previousServerIdentity }
: {}),
};
} catch {
return null;
}
}
export function readDevServerRestartRequest(
env: NodeJS.ProcessEnv = process.env,
): DevServerRestartRequest | null {
const filePath = getDevServerRestartRequestFilePath(env);
if (!filePath || !existsSync(filePath)) return null;
return readDevServerRestartRequestAtPath(filePath);
}
export function removeDevServerRestartRequest(
expected?: Pick<DevServerRestartRequest, "requestId">,
env: NodeJS.ProcessEnv = process.env,
): boolean {
const filePath = getDevServerRestartRequestFilePath(env);
if (!filePath) return false;
try {
withDevRestartRequestLock(filePath, () => {
const current = readDevServerRestartRequestAtPath(filePath);
if (expected?.requestId && current?.requestId !== expected.requestId) return;
rmSync(filePath, { force: true });
});
return true;
} catch (error) {
if (
error instanceof Error &&
error.message === "dev_server_restart_request_lock_busy"
) {
return false;
}
throw error;
}
}
function normalizeStringArray(value: unknown): string[] {
if (!Array.isArray(value)) return [];
return value
.filter((entry): entry is string => typeof entry === "string")
.map((entry) => entry.trim())
.filter((entry) => entry.length > 0);
}
function normalizeTimestamp(value: unknown): string | null {
if (typeof value !== "string") return null;
const trimmed = value.trim();
return trimmed.length > 0 ? trimmed : null;
}
export function readPersistedDevServerStatus(
env: NodeJS.ProcessEnv = process.env,
): PersistedDevServerStatus | null {
const filePath = env.PAPERCLIP_DEV_SERVER_STATUS_FILE?.trim();
if (!filePath || !existsSync(filePath)) return null;
try {
if (statSync(filePath).size > MAX_PERSISTED_DEV_SERVER_STATUS_BYTES) {
return null;
}
const raw = JSON.parse(readFileSync(filePath, "utf8")) as Record<
string,
unknown
>;
const changedPathsSample = normalizeStringArray(
raw.changedPathsSample,
).slice(0, 5);
const pendingMigrations = normalizeStringArray(raw.pendingMigrations);
const changedPathCountRaw = raw.changedPathCount;
const changedPathCount =
typeof changedPathCountRaw === "number" &&
Number.isFinite(changedPathCountRaw)
? Math.max(0, Math.trunc(changedPathCountRaw))
: changedPathsSample.length;
const dirtyRaw = raw.dirty;
const dirty =
typeof dirtyRaw === "boolean"
? dirtyRaw
: changedPathCount > 0 || pendingMigrations.length > 0;
return {
dirty,
lastChangedAt: normalizeTimestamp(raw.lastChangedAt),
changedPathCount,
changedPathsSample,
pendingMigrations,
lastRestartAt: normalizeTimestamp(raw.lastRestartAt),
};
} catch {
return null;
}
}
export function toDevServerHealthStatus(
persisted: PersistedDevServerStatus,
opts: { autoRestartEnabled: boolean; activeRunCount: number },
): DevServerHealthStatus {
const hasPathChanges = persisted.changedPathCount > 0;
const hasPendingMigrations = persisted.pendingMigrations.length > 0;
const reason =
hasPathChanges && hasPendingMigrations
? "backend_changes_and_pending_migrations"
: hasPendingMigrations
? "pending_migrations"
: hasPathChanges
? "backend_changes"
: null;
const restartRequired = persisted.dirty || reason !== null;
return {
enabled: true,
restartRequired,
reason,
lastChangedAt: persisted.lastChangedAt,
changedPathCount: persisted.changedPathCount,
changedPathsSample: persisted.changedPathsSample,
pendingMigrations: persisted.pendingMigrations,
autoRestartEnabled: opts.autoRestartEnabled,
activeRunCount: opts.activeRunCount,
waitingForIdle:
restartRequired && opts.autoRestartEnabled && opts.activeRunCount > 0,
lastRestartAt: persisted.lastRestartAt,
};
}