mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner keeps durable run and provider state outside one server process. > - A server restart can leave that runner alive or can interrupt it after a provider checkpoint. > - The old startup path used handoff intent and PID evidence, but it did not reconstruct native ownership. > - That gap could block the issue, create a replacement run, or start duplicate provider work. > - This pull request adds durable same-run recovery for coordinated and uncoordinated restarts. > - The benefit is exact recovery of the run, runner, session, provider, steering, and finalization state. ## Linked Issues or Issue Description Refs #9628. That pull request added earlier local-adapter hot-restart work. This change adds native PRP authority reconstruction and same-run provider resume. Refs #10935. That pull request handles missing hot-restart snapshots. This change also supports hard restarts with no snapshot. Refs #11624. That pull request prevents unsafe retry after an adopted legacy process exits. This change reconciles native terminal evidence before provider recovery. Refs #12070. That pull request improves process liveness checks. This change also binds recovery to a process-start fingerprint and fails closed on ambiguity. **What happened?** The server could record hot-restart intent, but startup did not rebuild native runner ownership. A live runner could not re-register its PRP authority. A dead runner could not resume the exact native and provider session on the same heartbeat run. Generic recovery could then block the issue or create replacement work. **Expected behavior** A live native runner must reconnect with the same PID and logical identities. A dead runner must resume the same durable session and heartbeat run with only a new operating-system PID. A proposed or terminal result must finalize once before any provider turn starts. Ambiguous process or session evidence must stay blocked without a signal or duplicate spawn. **Steps to reproduce** 1. Start a Paperclip Runner heartbeat and wait for an active provider turn. 2. Restart only the Paperclip server, with or without a hot-restart marker. 3. Observe that the old startup path does not reconstruct the native control-plane authority. 4. Kill both the server and runner after a provider checkpoint. 5. Observe that the old path cannot resume the exact native session on the original heartbeat run. **Paperclip version or commit** The defect was reproduced from commit `1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto the current `master`. **Deployment mode** Local development and self-hosted server deployments that use the local Paperclip Runner. ## What Changed - Added correlated hot-restart requests and version-compatible native handoff fields. - Added controller boot identity, process-start identity, controller generation, recovery state, request id, and bounded history to the native finalization ledger. - Added transactional recovery claims for live-runner reattach, dead-runner resume, and incomplete bootstrap. - Added fail-closed ownership takeover rules and process identity validation. - Added live runner adoption to the local runner transport without a duplicate spawn. - Added same-run provider checkpoint resume and legacy retry-row compatibility. - Reconciled proposed and terminal results before runner or provider recovery. - Bound the HTTP and PRP listener before startup recovery and delayed scheduling and generic reapers until classification completes. - Added restart-aware health diagnostics, run-log recovery transitions, durable runner diagnostics, and bounded shutdown finalizer draining. - Moved restart-survivable diagnostics into runner-owned, pre-redacted bounded writes; raw stdout and stderr are never persisted. - Added process-start fencing for controller, runner, and provider PIDs; startup classifies every candidate without an implicit cap. - Added crash-recoverable, contention-safe development restart-request coordination and failed-startup listener cleanup. - Added a credential-free real-process restart suite for eight restart, scale, and identity scenarios. - Documented native restart operation, persistence, diagnostics, and verification. ## Verification - The documented native restart commands passed. They ran eight real-process/database recovery scenarios and the live runner adoption transport test. - Native executor tests passed: 111 tests. - Heartbeat recovery tests passed: 124 tests. - Hot restart, health, and shutdown tests passed: 52 tests. - The broader affected server suite passed: 350 tests. - Focused native recovery and startup tests passed: 49 tests. - Runner transport and control-plane tests passed: 63 tests. - Runner-owned diagnostic tests passed for write-time bounding, credential redaction, private file modes, and raw stream non-persistence. - Development restart coordination tests passed: 11 tests. - Database migration checks and the partial-application/replay regression test passed. - Server, database, and Paperclip Runner typechecks passed. - `git diff --check` passed. - Full Paperclip PR CI passed, including build, canary, all five general server shards, all five serialized server shards, all three browser E2E shards, workspace suites, and release-registry verification. - Greptile completed at 5/5 with no outstanding findings, recommendations, follow-ups, or open review threads. ## Risks - Moderate risk. This changes startup ordering and ownership transfer for active native runs. - The migration adds nullable columns and does not rewrite existing rows. - Recovery fails closed when process or durable session identity is incomplete or contradictory. - The first implementation supports the local Paperclip Runner. Remote targets keep their existing behavior. - The real-process suite covers cleanup and asserts that no runner or provider process survives each test. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime did not expose a more specific model revision or context-window size. Repository editing, shell execution, database tests, and real-process test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
334 lines
9.8 KiB
TypeScript
334 lines
9.8 KiB
TypeScript
import { randomUUID } from "node:crypto";
|
|
import {
|
|
existsSync,
|
|
mkdirSync,
|
|
readFileSync,
|
|
renameSync,
|
|
rmSync,
|
|
statSync,
|
|
unlinkSync,
|
|
writeFileSync,
|
|
} from "node:fs";
|
|
import path from "node:path";
|
|
|
|
const MAX_PERSISTED_DEV_SERVER_STATUS_BYTES = 64 * 1024;
|
|
const DEV_RESTART_REQUEST_LOCK_STALE_MS = 30_000;
|
|
const DEV_RESTART_REQUEST_LOCK_RETRY_COUNT = 50;
|
|
const DEV_RESTART_REQUEST_LOCK_RETRY_MS = 2;
|
|
const lockWaitBuffer = new Int32Array(new SharedArrayBuffer(4));
|
|
|
|
export type PersistedDevServerStatus = {
|
|
dirty: boolean;
|
|
lastChangedAt: string | null;
|
|
changedPathCount: number;
|
|
changedPathsSample: string[];
|
|
pendingMigrations: string[];
|
|
lastRestartAt: string | null;
|
|
};
|
|
|
|
export type DevServerHealthStatus = {
|
|
enabled: true;
|
|
restartRequired: boolean;
|
|
reason:
|
|
| "backend_changes"
|
|
| "pending_migrations"
|
|
| "backend_changes_and_pending_migrations"
|
|
| null;
|
|
lastChangedAt: string | null;
|
|
changedPathCount: number;
|
|
changedPathsSample: string[];
|
|
pendingMigrations: string[];
|
|
autoRestartEnabled: boolean;
|
|
activeRunCount: number;
|
|
waitingForIdle: boolean;
|
|
lastRestartAt: string | null;
|
|
};
|
|
|
|
export type DevServerRestartRequest = {
|
|
requestedAt: string;
|
|
reason: "manual_restart_now";
|
|
requestId?: string;
|
|
mode?: "hot";
|
|
previousServerIdentity?: string;
|
|
};
|
|
|
|
function processIsAlive(pid: number): boolean {
|
|
try {
|
|
process.kill(pid, 0);
|
|
return true;
|
|
} catch (error) {
|
|
return (error as NodeJS.ErrnoException | undefined)?.code === "EPERM";
|
|
}
|
|
}
|
|
|
|
function tryRecoverDevRestartRequestLock(lockPath: string): boolean {
|
|
let stale = false;
|
|
try {
|
|
const lockAgeMs = Date.now() - statSync(lockPath).mtimeMs;
|
|
const owner = JSON.parse(
|
|
readFileSync(path.join(lockPath, "owner.json"), "utf8"),
|
|
) as Record<string, unknown>;
|
|
stale =
|
|
lockAgeMs >= DEV_RESTART_REQUEST_LOCK_STALE_MS ||
|
|
(typeof owner.pid === "number" && !processIsAlive(owner.pid));
|
|
} catch {
|
|
// Canonical locks are published only after owner.json is durable in a
|
|
// private candidate directory. A visible lock without valid ownership is
|
|
// therefore abandoned and can be reclaimed immediately.
|
|
stale = true;
|
|
}
|
|
if (!stale) return false;
|
|
|
|
const stalePath = `${lockPath}.${process.pid}.${randomUUID()}.stale`;
|
|
try {
|
|
renameSync(lockPath, stalePath);
|
|
} catch (error) {
|
|
if ((error as NodeJS.ErrnoException | undefined)?.code === "ENOENT") {
|
|
return true;
|
|
}
|
|
return false;
|
|
}
|
|
rmSync(stalePath, { recursive: true, force: true });
|
|
return true;
|
|
}
|
|
|
|
function withDevRestartRequestLock<T>(filePath: string, action: () => T): T {
|
|
const lockPath = `${filePath}.lock`;
|
|
for (
|
|
let attempt = 0;
|
|
attempt <= DEV_RESTART_REQUEST_LOCK_RETRY_COUNT;
|
|
attempt += 1
|
|
) {
|
|
const candidateLockPath = `${lockPath}.${process.pid}.${randomUUID()}.candidate`;
|
|
mkdirSync(candidateLockPath);
|
|
try {
|
|
writeFileSync(
|
|
path.join(candidateLockPath, "owner.json"),
|
|
`${JSON.stringify({ pid: process.pid, acquiredAt: new Date().toISOString() })}\n`,
|
|
"utf8",
|
|
);
|
|
try {
|
|
// Publishing a populated directory makes lock ownership visible in one
|
|
// rename; there is no canonical owner-less crash window.
|
|
renameSync(candidateLockPath, lockPath);
|
|
} catch (error) {
|
|
const code = (error as NodeJS.ErrnoException | undefined)?.code;
|
|
if (code !== "EEXIST" && code !== "ENOTEMPTY" && code !== "EPERM") {
|
|
throw error;
|
|
}
|
|
if (tryRecoverDevRestartRequestLock(lockPath)) continue;
|
|
if (attempt === DEV_RESTART_REQUEST_LOCK_RETRY_COUNT) {
|
|
throw new Error("dev_server_restart_request_lock_busy");
|
|
}
|
|
Atomics.wait(
|
|
lockWaitBuffer,
|
|
0,
|
|
0,
|
|
DEV_RESTART_REQUEST_LOCK_RETRY_MS,
|
|
);
|
|
continue;
|
|
}
|
|
} finally {
|
|
if (existsSync(candidateLockPath)) {
|
|
rmSync(candidateLockPath, { recursive: true, force: true });
|
|
}
|
|
}
|
|
try {
|
|
return action();
|
|
} finally {
|
|
rmSync(lockPath, { recursive: true, force: true });
|
|
}
|
|
}
|
|
throw new Error("dev_server_restart_request_lock_busy");
|
|
}
|
|
|
|
export function getDevServerRestartRequestFilePath(
|
|
env: NodeJS.ProcessEnv = process.env,
|
|
): string | null {
|
|
const statusFilePath = env.PAPERCLIP_DEV_SERVER_STATUS_FILE?.trim();
|
|
if (!statusFilePath) return null;
|
|
return path.join(
|
|
path.dirname(statusFilePath),
|
|
"dev-server-restart-request.json",
|
|
);
|
|
}
|
|
|
|
export function writeDevServerRestartRequest(
|
|
request: DevServerRestartRequest,
|
|
env: NodeJS.ProcessEnv = process.env,
|
|
): boolean {
|
|
const filePath = getDevServerRestartRequestFilePath(env);
|
|
if (!filePath) return false;
|
|
|
|
mkdirSync(path.dirname(filePath), { recursive: true });
|
|
withDevRestartRequestLock(filePath, () => {
|
|
const tempPath = `${filePath}.${process.pid}.${randomUUID()}.tmp`;
|
|
try {
|
|
writeFileSync(tempPath, `${JSON.stringify(request, null, 2)}\n`, "utf8");
|
|
renameSync(tempPath, filePath);
|
|
} finally {
|
|
try {
|
|
unlinkSync(tempPath);
|
|
} catch (error) {
|
|
if ((error as NodeJS.ErrnoException | undefined)?.code !== "ENOENT") {
|
|
throw error;
|
|
}
|
|
}
|
|
}
|
|
});
|
|
return true;
|
|
}
|
|
|
|
function readDevServerRestartRequestAtPath(
|
|
filePath: string,
|
|
): DevServerRestartRequest | null {
|
|
try {
|
|
if (statSync(filePath).size > MAX_PERSISTED_DEV_SERVER_STATUS_BYTES)
|
|
return null;
|
|
const value = JSON.parse(readFileSync(filePath, "utf8")) as Record<
|
|
string,
|
|
unknown
|
|
>;
|
|
if (
|
|
typeof value.requestedAt !== "string" ||
|
|
value.reason !== "manual_restart_now"
|
|
) {
|
|
return null;
|
|
}
|
|
return {
|
|
requestedAt: value.requestedAt,
|
|
reason: "manual_restart_now",
|
|
...(typeof value.requestId === "string"
|
|
? { requestId: value.requestId }
|
|
: {}),
|
|
...(value.mode === "hot" ? { mode: "hot" as const } : {}),
|
|
...(typeof value.previousServerIdentity === "string"
|
|
? { previousServerIdentity: value.previousServerIdentity }
|
|
: {}),
|
|
};
|
|
} catch {
|
|
return null;
|
|
}
|
|
}
|
|
|
|
export function readDevServerRestartRequest(
|
|
env: NodeJS.ProcessEnv = process.env,
|
|
): DevServerRestartRequest | null {
|
|
const filePath = getDevServerRestartRequestFilePath(env);
|
|
if (!filePath || !existsSync(filePath)) return null;
|
|
return readDevServerRestartRequestAtPath(filePath);
|
|
}
|
|
|
|
export function removeDevServerRestartRequest(
|
|
expected?: Pick<DevServerRestartRequest, "requestId">,
|
|
env: NodeJS.ProcessEnv = process.env,
|
|
): boolean {
|
|
const filePath = getDevServerRestartRequestFilePath(env);
|
|
if (!filePath) return false;
|
|
try {
|
|
withDevRestartRequestLock(filePath, () => {
|
|
const current = readDevServerRestartRequestAtPath(filePath);
|
|
if (expected?.requestId && current?.requestId !== expected.requestId) return;
|
|
rmSync(filePath, { force: true });
|
|
});
|
|
return true;
|
|
} catch (error) {
|
|
if (
|
|
error instanceof Error &&
|
|
error.message === "dev_server_restart_request_lock_busy"
|
|
) {
|
|
return false;
|
|
}
|
|
throw error;
|
|
}
|
|
}
|
|
|
|
function normalizeStringArray(value: unknown): string[] {
|
|
if (!Array.isArray(value)) return [];
|
|
return value
|
|
.filter((entry): entry is string => typeof entry === "string")
|
|
.map((entry) => entry.trim())
|
|
.filter((entry) => entry.length > 0);
|
|
}
|
|
|
|
function normalizeTimestamp(value: unknown): string | null {
|
|
if (typeof value !== "string") return null;
|
|
const trimmed = value.trim();
|
|
return trimmed.length > 0 ? trimmed : null;
|
|
}
|
|
|
|
export function readPersistedDevServerStatus(
|
|
env: NodeJS.ProcessEnv = process.env,
|
|
): PersistedDevServerStatus | null {
|
|
const filePath = env.PAPERCLIP_DEV_SERVER_STATUS_FILE?.trim();
|
|
if (!filePath || !existsSync(filePath)) return null;
|
|
|
|
try {
|
|
if (statSync(filePath).size > MAX_PERSISTED_DEV_SERVER_STATUS_BYTES) {
|
|
return null;
|
|
}
|
|
const raw = JSON.parse(readFileSync(filePath, "utf8")) as Record<
|
|
string,
|
|
unknown
|
|
>;
|
|
const changedPathsSample = normalizeStringArray(
|
|
raw.changedPathsSample,
|
|
).slice(0, 5);
|
|
const pendingMigrations = normalizeStringArray(raw.pendingMigrations);
|
|
const changedPathCountRaw = raw.changedPathCount;
|
|
const changedPathCount =
|
|
typeof changedPathCountRaw === "number" &&
|
|
Number.isFinite(changedPathCountRaw)
|
|
? Math.max(0, Math.trunc(changedPathCountRaw))
|
|
: changedPathsSample.length;
|
|
const dirtyRaw = raw.dirty;
|
|
const dirty =
|
|
typeof dirtyRaw === "boolean"
|
|
? dirtyRaw
|
|
: changedPathCount > 0 || pendingMigrations.length > 0;
|
|
|
|
return {
|
|
dirty,
|
|
lastChangedAt: normalizeTimestamp(raw.lastChangedAt),
|
|
changedPathCount,
|
|
changedPathsSample,
|
|
pendingMigrations,
|
|
lastRestartAt: normalizeTimestamp(raw.lastRestartAt),
|
|
};
|
|
} catch {
|
|
return null;
|
|
}
|
|
}
|
|
|
|
export function toDevServerHealthStatus(
|
|
persisted: PersistedDevServerStatus,
|
|
opts: { autoRestartEnabled: boolean; activeRunCount: number },
|
|
): DevServerHealthStatus {
|
|
const hasPathChanges = persisted.changedPathCount > 0;
|
|
const hasPendingMigrations = persisted.pendingMigrations.length > 0;
|
|
const reason =
|
|
hasPathChanges && hasPendingMigrations
|
|
? "backend_changes_and_pending_migrations"
|
|
: hasPendingMigrations
|
|
? "pending_migrations"
|
|
: hasPathChanges
|
|
? "backend_changes"
|
|
: null;
|
|
const restartRequired = persisted.dirty || reason !== null;
|
|
|
|
return {
|
|
enabled: true,
|
|
restartRequired,
|
|
reason,
|
|
lastChangedAt: persisted.lastChangedAt,
|
|
changedPathCount: persisted.changedPathCount,
|
|
changedPathsSample: persisted.changedPathsSample,
|
|
pendingMigrations: persisted.pendingMigrations,
|
|
autoRestartEnabled: opts.autoRestartEnabled,
|
|
activeRunCount: opts.activeRunCount,
|
|
waitingForIdle:
|
|
restartRequired && opts.autoRestartEnabled && opts.activeRunCount > 0,
|
|
lastRestartAt: persisted.lastRestartAt,
|
|
};
|
|
}
|