Files
PaperClipAI/server/src/__tests__/dev-server-status.test.ts
T
Dotta 7b094724e6 fix(runner): recover native sessions across restarts (#12845)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Paperclip Runner keeps durable run and provider state outside
one server process.
> - A server restart can leave that runner alive or can interrupt it
after a provider checkpoint.
> - The old startup path used handoff intent and PID evidence, but it
did not reconstruct native ownership.
> - That gap could block the issue, create a replacement run, or start
duplicate provider work.
> - This pull request adds durable same-run recovery for coordinated and
uncoordinated restarts.
> - The benefit is exact recovery of the run, runner, session, provider,
steering, and finalization state.

## Linked Issues or Issue Description

Refs #9628. That pull request added earlier local-adapter hot-restart
work. This change adds native PRP authority reconstruction and same-run
provider resume.

Refs #10935. That pull request handles missing hot-restart snapshots.
This change also supports hard restarts with no snapshot.

Refs #11624. That pull request prevents unsafe retry after an adopted
legacy process exits. This change reconciles native terminal evidence
before provider recovery.

Refs #12070. That pull request improves process liveness checks. This
change also binds recovery to a process-start fingerprint and fails
closed on ambiguity.

**What happened?**

The server could record hot-restart intent, but startup did not rebuild
native runner ownership. A live runner could not re-register its PRP
authority. A dead runner could not resume the exact native and provider
session on the same heartbeat run. Generic recovery could then block the
issue or create replacement work.

**Expected behavior**

A live native runner must reconnect with the same PID and logical
identities. A dead runner must resume the same durable session and
heartbeat run with only a new operating-system PID. A proposed or
terminal result must finalize once before any provider turn starts.
Ambiguous process or session evidence must stay blocked without a signal
or duplicate spawn.

**Steps to reproduce**

1. Start a Paperclip Runner heartbeat and wait for an active provider
turn.
2. Restart only the Paperclip server, with or without a hot-restart
marker.
3. Observe that the old startup path does not reconstruct the native
control-plane authority.
4. Kill both the server and runner after a provider checkpoint.
5. Observe that the old path cannot resume the exact native session on
the original heartbeat run.

**Paperclip version or commit**

The defect was reproduced from commit
`1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto
the current `master`.

**Deployment mode**

Local development and self-hosted server deployments that use the local
Paperclip Runner.

## What Changed

- Added correlated hot-restart requests and version-compatible native
handoff fields.
- Added controller boot identity, process-start identity, controller
generation, recovery state, request id, and bounded history to the
native finalization ledger.
- Added transactional recovery claims for live-runner reattach,
dead-runner resume, and incomplete bootstrap.
- Added fail-closed ownership takeover rules and process identity
validation.
- Added live runner adoption to the local runner transport without a
duplicate spawn.
- Added same-run provider checkpoint resume and legacy retry-row
compatibility.
- Reconciled proposed and terminal results before runner or provider
recovery.
- Bound the HTTP and PRP listener before startup recovery and delayed
scheduling and generic reapers until classification completes.
- Added restart-aware health diagnostics, run-log recovery transitions,
durable runner diagnostics, and bounded shutdown finalizer draining.
- Moved restart-survivable diagnostics into runner-owned, pre-redacted
bounded writes; raw stdout and stderr are never persisted.
- Added process-start fencing for controller, runner, and provider PIDs;
startup classifies every candidate without an implicit cap.
- Added crash-recoverable, contention-safe development restart-request
coordination and failed-startup listener cleanup.
- Added a credential-free real-process restart suite for eight restart,
scale, and identity scenarios.
- Documented native restart operation, persistence, diagnostics, and
verification.

## Verification

- The documented native restart commands passed. They ran eight
real-process/database recovery scenarios and the live runner adoption
transport test.
- Native executor tests passed: 111 tests.
- Heartbeat recovery tests passed: 124 tests.
- Hot restart, health, and shutdown tests passed: 52 tests.
- The broader affected server suite passed: 350 tests.
- Focused native recovery and startup tests passed: 49 tests.
- Runner transport and control-plane tests passed: 63 tests.
- Runner-owned diagnostic tests passed for write-time bounding,
credential redaction, private file modes, and raw stream
non-persistence.
- Development restart coordination tests passed: 11 tests.
- Database migration checks and the partial-application/replay
regression test passed.
- Server, database, and Paperclip Runner typechecks passed.
- `git diff --check` passed.
- Full Paperclip PR CI passed, including build, canary, all five general
server shards, all five serialized server shards, all three browser E2E
shards, workspace suites, and release-registry verification.
- Greptile completed at 5/5 with no outstanding findings,
recommendations, follow-ups, or open review threads.

## Risks

- Moderate risk. This changes startup ordering and ownership transfer
for active native runs.
- The migration adds nullable columns and does not rewrite existing
rows.
- Recovery fails closed when process or durable session identity is
incomplete or contradictory.
- The first implementation supports the local Paperclip Runner. Remote
targets keep their existing behavior.
- The real-process suite covers cleanup and asserts that no runner or
provider process survives each test.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with GPT-5. The runtime did not expose a more specific
model revision or context-window size. Repository editing, shell
execution, database tests, and real-process test execution were enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-04 15:03:53 -05:00

210 lines
6.2 KiB
TypeScript

import {
existsSync,
mkdirSync,
mkdtempSync,
readFileSync,
rmSync,
writeFileSync,
} from "node:fs";
import os from "node:os";
import path from "node:path";
import { afterEach, describe, expect, it } from "vitest";
import {
getDevServerRestartRequestFilePath,
readDevServerRestartRequest,
readPersistedDevServerStatus,
removeDevServerRestartRequest,
toDevServerHealthStatus,
writeDevServerRestartRequest,
} from "../dev-server-status.js";
const tempDirs = [];
function createTempStatusFile(payload: unknown) {
const dir = mkdtempSync(path.join(os.tmpdir(), "paperclip-dev-status-"));
tempDirs.push(dir);
const filePath = path.join(dir, "dev-server-status.json");
writeFileSync(filePath, `${JSON.stringify(payload)}\n`, "utf8");
return filePath;
}
afterEach(() => {
for (const dir of tempDirs.splice(0)) {
rmSync(dir, { recursive: true, force: true });
}
});
describe("dev server status helpers", () => {
it("reads and normalizes persisted supervisor state", () => {
const filePath = createTempStatusFile({
dirty: true,
lastChangedAt: "2026-03-20T12:00:00.000Z",
changedPathCount: 4,
changedPathsSample: ["server/src/app.ts", "packages/shared/src/index.ts"],
pendingMigrations: ["0040_restart_banner.sql"],
lastRestartAt: "2026-03-20T11:30:00.000Z",
});
expect(
readPersistedDevServerStatus({
PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath,
}),
).toEqual({
dirty: true,
lastChangedAt: "2026-03-20T12:00:00.000Z",
changedPathCount: 4,
changedPathsSample: ["server/src/app.ts", "packages/shared/src/index.ts"],
pendingMigrations: ["0040_restart_banner.sql"],
lastRestartAt: "2026-03-20T11:30:00.000Z",
});
});
it("derives waiting-for-idle health state", () => {
const health = toDevServerHealthStatus(
{
dirty: true,
lastChangedAt: "2026-03-20T12:00:00.000Z",
changedPathCount: 2,
changedPathsSample: ["server/src/app.ts"],
pendingMigrations: [],
lastRestartAt: "2026-03-20T11:30:00.000Z",
},
{ autoRestartEnabled: true, activeRunCount: 3 },
);
expect(health).toMatchObject({
enabled: true,
restartRequired: true,
reason: "backend_changes",
autoRestartEnabled: true,
activeRunCount: 3,
waitingForIdle: true,
});
});
it("ignores oversized persisted status files", () => {
const filePath = createTempStatusFile({
dirty: true,
changedPathsSample: ["x".repeat(70 * 1024)],
pendingMigrations: [],
});
expect(
readPersistedDevServerStatus({
PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath,
}),
).toBeNull();
});
it("writes restart requests next to the persisted status file", () => {
const filePath = createTempStatusFile({
dirty: true,
changedPathsSample: ["server/src/app.ts"],
pendingMigrations: [],
});
const env = { PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath };
expect(
writeDevServerRestartRequest(
{
requestedAt: "2026-03-20T12:05:00.000Z",
reason: "manual_restart_now",
},
env,
),
).toBe(true);
const requestPath = getDevServerRestartRequestFilePath(env);
expect(requestPath).toBe(
path.join(path.dirname(filePath), "dev-server-restart-request.json"),
);
expect(requestPath && existsSync(requestPath)).toBe(true);
expect(JSON.parse(readFileSync(requestPath!, "utf8"))).toEqual({
requestedAt: "2026-03-20T12:05:00.000Z",
reason: "manual_restart_now",
});
});
it("correlates restart request cleanup so stale consumers cannot remove a replacement", () => {
const filePath = createTempStatusFile({ dirty: true });
const env = { PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath };
expect(
writeDevServerRestartRequest(
{
requestedAt: "2026-09-04T12:00:00.000Z",
reason: "manual_restart_now",
requestId: "restart-new",
mode: "hot",
previousServerIdentity: "server-start-new",
},
env,
),
).toBe(true);
removeDevServerRestartRequest({ requestId: "restart-stale" }, env);
expect(readDevServerRestartRequest(env)).toEqual({
requestedAt: "2026-09-04T12:00:00.000Z",
reason: "manual_restart_now",
requestId: "restart-new",
mode: "hot",
previousServerIdentity: "server-start-new",
});
removeDevServerRestartRequest({ requestId: "restart-new" }, env);
expect(readDevServerRestartRequest(env)).toBeNull();
});
it("immediately recovers an owner-less lock from an interrupted publisher", () => {
const filePath = createTempStatusFile({ dirty: true });
const env = { PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath };
const requestPath = getDevServerRestartRequestFilePath(env)!;
const lockPath = `${requestPath}.lock`;
mkdirSync(lockPath);
writeDevServerRestartRequest(
{
requestedAt: "2026-09-04T12:00:01.000Z",
reason: "manual_restart_now",
requestId: "restart-after-crash",
mode: "hot",
},
env,
);
expect(readDevServerRestartRequest(env)).toMatchObject({
requestId: "restart-after-crash",
requestedAt: "2026-09-04T12:00:01.000Z",
});
expect(existsSync(lockPath)).toBe(false);
});
it("preserves the request instead of throwing when a live writer holds the lock", () => {
const filePath = createTempStatusFile({ dirty: true });
const env = { PAPERCLIP_DEV_SERVER_STATUS_FILE: filePath };
writeDevServerRestartRequest(
{
requestedAt: "2026-09-04T12:00:02.000Z",
reason: "manual_restart_now",
requestId: "restart-contended",
mode: "hot",
},
env,
);
const requestPath = getDevServerRestartRequestFilePath(env)!;
const lockPath = `${requestPath}.lock`;
mkdirSync(lockPath);
writeFileSync(
path.join(lockPath, "owner.json"),
JSON.stringify({ pid: process.pid, acquiredAt: new Date().toISOString() }),
"utf8",
);
expect(
removeDevServerRestartRequest({ requestId: "restart-contended" }, env),
).toBe(false);
expect(readDevServerRestartRequest(env)).toMatchObject({
requestId: "restart-contended",
});
});
});