mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 14:52:16 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner keeps durable run and provider state outside one server process. > - A server restart can leave that runner alive or can interrupt it after a provider checkpoint. > - The old startup path used handoff intent and PID evidence, but it did not reconstruct native ownership. > - That gap could block the issue, create a replacement run, or start duplicate provider work. > - This pull request adds durable same-run recovery for coordinated and uncoordinated restarts. > - The benefit is exact recovery of the run, runner, session, provider, steering, and finalization state. ## Linked Issues or Issue Description Refs #9628. That pull request added earlier local-adapter hot-restart work. This change adds native PRP authority reconstruction and same-run provider resume. Refs #10935. That pull request handles missing hot-restart snapshots. This change also supports hard restarts with no snapshot. Refs #11624. That pull request prevents unsafe retry after an adopted legacy process exits. This change reconciles native terminal evidence before provider recovery. Refs #12070. That pull request improves process liveness checks. This change also binds recovery to a process-start fingerprint and fails closed on ambiguity. **What happened?** The server could record hot-restart intent, but startup did not rebuild native runner ownership. A live runner could not re-register its PRP authority. A dead runner could not resume the exact native and provider session on the same heartbeat run. Generic recovery could then block the issue or create replacement work. **Expected behavior** A live native runner must reconnect with the same PID and logical identities. A dead runner must resume the same durable session and heartbeat run with only a new operating-system PID. A proposed or terminal result must finalize once before any provider turn starts. Ambiguous process or session evidence must stay blocked without a signal or duplicate spawn. **Steps to reproduce** 1. Start a Paperclip Runner heartbeat and wait for an active provider turn. 2. Restart only the Paperclip server, with or without a hot-restart marker. 3. Observe that the old startup path does not reconstruct the native control-plane authority. 4. Kill both the server and runner after a provider checkpoint. 5. Observe that the old path cannot resume the exact native session on the original heartbeat run. **Paperclip version or commit** The defect was reproduced from commit `1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto the current `master`. **Deployment mode** Local development and self-hosted server deployments that use the local Paperclip Runner. ## What Changed - Added correlated hot-restart requests and version-compatible native handoff fields. - Added controller boot identity, process-start identity, controller generation, recovery state, request id, and bounded history to the native finalization ledger. - Added transactional recovery claims for live-runner reattach, dead-runner resume, and incomplete bootstrap. - Added fail-closed ownership takeover rules and process identity validation. - Added live runner adoption to the local runner transport without a duplicate spawn. - Added same-run provider checkpoint resume and legacy retry-row compatibility. - Reconciled proposed and terminal results before runner or provider recovery. - Bound the HTTP and PRP listener before startup recovery and delayed scheduling and generic reapers until classification completes. - Added restart-aware health diagnostics, run-log recovery transitions, durable runner diagnostics, and bounded shutdown finalizer draining. - Moved restart-survivable diagnostics into runner-owned, pre-redacted bounded writes; raw stdout and stderr are never persisted. - Added process-start fencing for controller, runner, and provider PIDs; startup classifies every candidate without an implicit cap. - Added crash-recoverable, contention-safe development restart-request coordination and failed-startup listener cleanup. - Added a credential-free real-process restart suite for eight restart, scale, and identity scenarios. - Documented native restart operation, persistence, diagnostics, and verification. ## Verification - The documented native restart commands passed. They ran eight real-process/database recovery scenarios and the live runner adoption transport test. - Native executor tests passed: 111 tests. - Heartbeat recovery tests passed: 124 tests. - Hot restart, health, and shutdown tests passed: 52 tests. - The broader affected server suite passed: 350 tests. - Focused native recovery and startup tests passed: 49 tests. - Runner transport and control-plane tests passed: 63 tests. - Runner-owned diagnostic tests passed for write-time bounding, credential redaction, private file modes, and raw stream non-persistence. - Development restart coordination tests passed: 11 tests. - Database migration checks and the partial-application/replay regression test passed. - Server, database, and Paperclip Runner typechecks passed. - `git diff --check` passed. - Full Paperclip PR CI passed, including build, canary, all five general server shards, all five serialized server shards, all three browser E2E shards, workspace suites, and release-registry verification. - Greptile completed at 5/5 with no outstanding findings, recommendations, follow-ups, or open review threads. ## Risks - Moderate risk. This changes startup ordering and ownership transfer for active native runs. - The migration adds nullable columns and does not rewrite existing rows. - Recovery fails closed when process or durable session identity is incomplete or contradictory. - The first implementation supports the local Paperclip Runner. Remote targets keep their existing behavior. - The real-process suite covers cleanup and asserts that no runner or provider process survives each test. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime did not expose a more specific model revision or context-window size. Repository editing, shell execution, database tests, and real-process test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
210 lines
6.2 KiB
TypeScript
210 lines
6.2 KiB
TypeScript
import {
|
|
existsSync,
|
|
mkdirSync,
|
|
mkdtempSync,
|
|
readFileSync,
|
|
rmSync,
|
|
writeFileSync,
|
|
} from "node:fs";
|
|
import os from "node:os";
|
|
import path from "node:path";
|
|
import { afterEach, describe, expect, it } from "vitest";
|
|
import {
|
|
getDevServerRestartRequestFilePath,
|
|
readDevServerRestartRequest,
|
|
readPersistedDevServerStatus,
|
|
removeDevServerRestartRequest,
|
|
toDevServerHealthStatus,
|
|
writeDevServerRestartRequest,
|
|
} from "../dev-server-status.js";
|
|
|
|
const tempDirs = [];
|
|
|
|
function createTempStatusFile(payload: unknown) {
|
|
const dir = mkdtempSync(path.join(os.tmpdir(), "paperclip-dev-status-"));
|
|
tempDirs.push(dir);
|
|
const filePath = path.join(dir, "dev-server-status.json");
|
|
writeFileSync(filePath, `${JSON.stringify(payload)}\n`, "utf8");
|
|
return filePath;
|
|
}
|
|
|
|
afterEach(() => {
|
|
for (const dir of tempDirs.splice(0)) {
|
|
rmSync(dir, { recursive: true, force: true });
|
|
}
|
|
});
|
|
|
|
describe("dev server status helpers", () => {
|
|
it("reads and normalizes persisted supervisor state", () => {
|
|
const filePath = createTempStatusFile({
|
|
dirty: true,
|
|
lastChangedAt: "2026-03-20T12:00:00.000Z",
|
|
changedPathCount: 4,
|
|
changedPathsSample: ["server/src/app.ts", "packages/shared/src/index.ts"],
|
|
pendingMigrations: ["0040_restart_banner.sql"],
|
|
lastRestartAt: "2026-03-20T11:30:00.000Z",
|
|
});
|
|
|
|
expect(
|
|
readPersistedDevServerStatus({
|
|
PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath,
|
|
}),
|
|
).toEqual({
|
|
dirty: true,
|
|
lastChangedAt: "2026-03-20T12:00:00.000Z",
|
|
changedPathCount: 4,
|
|
changedPathsSample: ["server/src/app.ts", "packages/shared/src/index.ts"],
|
|
pendingMigrations: ["0040_restart_banner.sql"],
|
|
lastRestartAt: "2026-03-20T11:30:00.000Z",
|
|
});
|
|
});
|
|
|
|
it("derives waiting-for-idle health state", () => {
|
|
const health = toDevServerHealthStatus(
|
|
{
|
|
dirty: true,
|
|
lastChangedAt: "2026-03-20T12:00:00.000Z",
|
|
changedPathCount: 2,
|
|
changedPathsSample: ["server/src/app.ts"],
|
|
pendingMigrations: [],
|
|
lastRestartAt: "2026-03-20T11:30:00.000Z",
|
|
},
|
|
{ autoRestartEnabled: true, activeRunCount: 3 },
|
|
);
|
|
|
|
expect(health).toMatchObject({
|
|
enabled: true,
|
|
restartRequired: true,
|
|
reason: "backend_changes",
|
|
autoRestartEnabled: true,
|
|
activeRunCount: 3,
|
|
waitingForIdle: true,
|
|
});
|
|
});
|
|
|
|
it("ignores oversized persisted status files", () => {
|
|
const filePath = createTempStatusFile({
|
|
dirty: true,
|
|
changedPathsSample: ["x".repeat(70 * 1024)],
|
|
pendingMigrations: [],
|
|
});
|
|
|
|
expect(
|
|
readPersistedDevServerStatus({
|
|
PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath,
|
|
}),
|
|
).toBeNull();
|
|
});
|
|
|
|
it("writes restart requests next to the persisted status file", () => {
|
|
const filePath = createTempStatusFile({
|
|
dirty: true,
|
|
changedPathsSample: ["server/src/app.ts"],
|
|
pendingMigrations: [],
|
|
});
|
|
|
|
const env = { PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath };
|
|
expect(
|
|
writeDevServerRestartRequest(
|
|
{
|
|
requestedAt: "2026-03-20T12:05:00.000Z",
|
|
reason: "manual_restart_now",
|
|
},
|
|
env,
|
|
),
|
|
).toBe(true);
|
|
|
|
const requestPath = getDevServerRestartRequestFilePath(env);
|
|
expect(requestPath).toBe(
|
|
path.join(path.dirname(filePath), "dev-server-restart-request.json"),
|
|
);
|
|
expect(requestPath && existsSync(requestPath)).toBe(true);
|
|
expect(JSON.parse(readFileSync(requestPath!, "utf8"))).toEqual({
|
|
requestedAt: "2026-03-20T12:05:00.000Z",
|
|
reason: "manual_restart_now",
|
|
});
|
|
});
|
|
|
|
it("correlates restart request cleanup so stale consumers cannot remove a replacement", () => {
|
|
const filePath = createTempStatusFile({ dirty: true });
|
|
const env = { PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath };
|
|
expect(
|
|
writeDevServerRestartRequest(
|
|
{
|
|
requestedAt: "2026-09-04T12:00:00.000Z",
|
|
reason: "manual_restart_now",
|
|
requestId: "restart-new",
|
|
mode: "hot",
|
|
previousServerIdentity: "server-start-new",
|
|
},
|
|
env,
|
|
),
|
|
).toBe(true);
|
|
|
|
removeDevServerRestartRequest({ requestId: "restart-stale" }, env);
|
|
expect(readDevServerRestartRequest(env)).toEqual({
|
|
requestedAt: "2026-09-04T12:00:00.000Z",
|
|
reason: "manual_restart_now",
|
|
requestId: "restart-new",
|
|
mode: "hot",
|
|
previousServerIdentity: "server-start-new",
|
|
});
|
|
|
|
removeDevServerRestartRequest({ requestId: "restart-new" }, env);
|
|
expect(readDevServerRestartRequest(env)).toBeNull();
|
|
});
|
|
|
|
it("immediately recovers an owner-less lock from an interrupted publisher", () => {
|
|
const filePath = createTempStatusFile({ dirty: true });
|
|
const env = { PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath };
|
|
const requestPath = getDevServerRestartRequestFilePath(env)!;
|
|
const lockPath = `${requestPath}.lock`;
|
|
mkdirSync(lockPath);
|
|
|
|
writeDevServerRestartRequest(
|
|
{
|
|
requestedAt: "2026-09-04T12:00:01.000Z",
|
|
reason: "manual_restart_now",
|
|
requestId: "restart-after-crash",
|
|
mode: "hot",
|
|
},
|
|
env,
|
|
);
|
|
|
|
expect(readDevServerRestartRequest(env)).toMatchObject({
|
|
requestId: "restart-after-crash",
|
|
requestedAt: "2026-09-04T12:00:01.000Z",
|
|
});
|
|
expect(existsSync(lockPath)).toBe(false);
|
|
});
|
|
|
|
it("preserves the request instead of throwing when a live writer holds the lock", () => {
|
|
const filePath = createTempStatusFile({ dirty: true });
|
|
const env = { PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath };
|
|
writeDevServerRestartRequest(
|
|
{
|
|
requestedAt: "2026-09-04T12:00:02.000Z",
|
|
reason: "manual_restart_now",
|
|
requestId: "restart-contended",
|
|
mode: "hot",
|
|
},
|
|
env,
|
|
);
|
|
const requestPath = getDevServerRestartRequestFilePath(env)!;
|
|
const lockPath = `${requestPath}.lock`;
|
|
mkdirSync(lockPath);
|
|
writeFileSync(
|
|
path.join(lockPath, "owner.json"),
|
|
JSON.stringify({ pid: process.pid, acquiredAt: new Date().toISOString() }),
|
|
"utf8",
|
|
);
|
|
|
|
expect(
|
|
removeDevServerRestartRequest({ requestId: "restart-contended" }, env),
|
|
).toBe(false);
|
|
expect(readDevServerRestartRequest(env)).toMatchObject({
|
|
requestId: "restart-contended",
|
|
});
|
|
});
|
|
});
|