## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner keeps durable run and provider state outside one server process. > - A server restart can leave that runner alive or can interrupt it after a provider checkpoint. > - The old startup path used handoff intent and PID evidence, but it did not reconstruct native ownership. > - That gap could block the issue, create a replacement run, or start duplicate provider work. > - This pull request adds durable same-run recovery for coordinated and uncoordinated restarts. > - The benefit is exact recovery of the run, runner, session, provider, steering, and finalization state. ## Linked Issues or Issue Description Refs #9628. That pull request added earlier local-adapter hot-restart work. This change adds native PRP authority reconstruction and same-run provider resume. Refs #10935. That pull request handles missing hot-restart snapshots. This change also supports hard restarts with no snapshot. Refs #11624. That pull request prevents unsafe retry after an adopted legacy process exits. This change reconciles native terminal evidence before provider recovery. Refs #12070. That pull request improves process liveness checks. This change also binds recovery to a process-start fingerprint and fails closed on ambiguity. **What happened?** The server could record hot-restart intent, but startup did not rebuild native runner ownership. A live runner could not re-register its PRP authority. A dead runner could not resume the exact native and provider session on the same heartbeat run. Generic recovery could then block the issue or create replacement work. **Expected behavior** A live native runner must reconnect with the same PID and logical identities. A dead runner must resume the same durable session and heartbeat run with only a new operating-system PID. A proposed or terminal result must finalize once before any provider turn starts. Ambiguous process or session evidence must stay blocked without a signal or duplicate spawn. **Steps to reproduce** 1. Start a Paperclip Runner heartbeat and wait for an active provider turn. 2. Restart only the Paperclip server, with or without a hot-restart marker. 3. Observe that the old startup path does not reconstruct the native control-plane authority. 4. Kill both the server and runner after a provider checkpoint. 5. Observe that the old path cannot resume the exact native session on the original heartbeat run. **Paperclip version or commit** The defect was reproduced from commit `1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto the current `master`. **Deployment mode** Local development and self-hosted server deployments that use the local Paperclip Runner. ## What Changed - Added correlated hot-restart requests and version-compatible native handoff fields. - Added controller boot identity, process-start identity, controller generation, recovery state, request id, and bounded history to the native finalization ledger. - Added transactional recovery claims for live-runner reattach, dead-runner resume, and incomplete bootstrap. - Added fail-closed ownership takeover rules and process identity validation. - Added live runner adoption to the local runner transport without a duplicate spawn. - Added same-run provider checkpoint resume and legacy retry-row compatibility. - Reconciled proposed and terminal results before runner or provider recovery. - Bound the HTTP and PRP listener before startup recovery and delayed scheduling and generic reapers until classification completes. - Added restart-aware health diagnostics, run-log recovery transitions, durable runner diagnostics, and bounded shutdown finalizer draining. - Moved restart-survivable diagnostics into runner-owned, pre-redacted bounded writes; raw stdout and stderr are never persisted. - Added process-start fencing for controller, runner, and provider PIDs; startup classifies every candidate without an implicit cap. - Added crash-recoverable, contention-safe development restart-request coordination and failed-startup listener cleanup. - Added a credential-free real-process restart suite for eight restart, scale, and identity scenarios. - Documented native restart operation, persistence, diagnostics, and verification. ## Verification - The documented native restart commands passed. They ran eight real-process/database recovery scenarios and the live runner adoption transport test. - Native executor tests passed: 111 tests. - Heartbeat recovery tests passed: 124 tests. - Hot restart, health, and shutdown tests passed: 52 tests. - The broader affected server suite passed: 350 tests. - Focused native recovery and startup tests passed: 49 tests. - Runner transport and control-plane tests passed: 63 tests. - Runner-owned diagnostic tests passed for write-time bounding, credential redaction, private file modes, and raw stream non-persistence. - Development restart coordination tests passed: 11 tests. - Database migration checks and the partial-application/replay regression test passed. - Server, database, and Paperclip Runner typechecks passed. - `git diff --check` passed. - Full Paperclip PR CI passed, including build, canary, all five general server shards, all five serialized server shards, all three browser E2E shards, workspace suites, and release-registry verification. - Greptile completed at 5/5 with no outstanding findings, recommendations, follow-ups, or open review threads. ## Risks - Moderate risk. This changes startup ordering and ownership transfer for active native runs. - The migration adds nullable columns and does not rewrite existing rows. - Recovery fails closed when process or durable session identity is incomplete or contradictory. - The first implementation supports the local Paperclip Runner. Remote targets keep their existing behavior. - The real-process suite covers cleanup and asserts that no runner or provider process survives each test. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime did not expose a more specific model revision or context-window size. Repository editing, shell execution, database tests, and real-process test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
5.4 KiB
Run-Log Events
Run-log events write to the heartbeat_run_events table
(packages/db/src/schema/heartbeat_run_events.ts:6-20). They are not
Paperclip Telemetry events, and they are not OpenTelemetry exports. A run-log
event needs no operator endpoint.
Native PRP Run-Log Events
The hidden native coordinator writes each validated PRP event to the bound
run's existing event stream before it acknowledges the runner. The row keeps
the PRP eventType, source instance, source event ID, source sequence, protocol
schema version, and a SHA-256 digest of the canonical source envelope. Its
payload is { "prpEvent": <canonical PRP event> }.
The writer locks the native heartbeat_runs row and allocates the existing
per-run seq cursor. A byte-equivalent retry reuses the first row; a changed
retry or source-sequence gap is rejected. Company, issue, agent, run, session,
and runner-source bindings must match the persisted native run. Bootstrap
tickets, reconnect leases, authentication proofs, encryption keys, and raw
credential material are never written to the run log.
These records remain run-log events. They do not create an OpenTelemetry or Paperclip Telemetry export, and legacy adapters do not use this writer.
Native Restart Recovery Run-Log Event
Paperclip writes a native.recovery.transition event for every native restart
classification and for graceful restart suspension. This immutable run-log
record lets operators reconstruct recovery decisions without exporting data to
Paperclip Telemetry or OpenTelemetry.
The payload contains the restart kind, recovery request id when one exists, runner disposition, and the controller generation and provider attempt for a claimed recovery. Live-runner adoption also records the runner PID, process group, and process-start fingerprint. A non-claim disposition records a bounded reason instead. Graceful suspension records the signal and confirms that it did not create a retry run.
The event never includes bootstrap tickets, reconnect leases, authentication
proofs, encryption keys, environment variables, provider credentials, command
arguments, or an unsanitized stderr stream. Detailed failed-attempt diagnostics
remain in the bounded native_run_finalizations.recovery_history ledger.
Sandbox Startup Run-Log Event
Paperclip writes one run.startup.step event to the run log for each bring-up
step. This event is a run-log record, not a first-party telemetry event. The
generated telemetry contract does not cover it, so this section is its canonical
contract.
The event payload carries only three fields.
| Field | Type | Meaning |
|---|---|---|
step |
string | The bring-up step name, for example stage.sync. |
durationMs |
number | The wall time of the step. A skipped step reports 0. |
outcome |
string | The step outcome (ok, skipped, or failed). |
The event no longer carries the per-step round-trip count or the provider
duration fields. It dropped roundTrips, providerExecMs, providerGetMs,
createRuntimeMs, and ensureSessionMs. The startup spans in
doc/observability.md carry that detail now. The
sandbox.exec child spans hold the round-trip and provider durations. The
acp.handshake step span holds the create-runtime and ensure-session
sub-times.
To read the detailed timing, use the startup spans. The spans need an OTLP endpoint. A run with no endpoint keeps only the three run-log fields above.
Run Phase Timing Run-Log Event
Paperclip writes one run.phase.timing event to the run log for each
run-lifecycle phase. This event is a run-log record, not a first-party telemetry
event. The generated telemetry contract does not cover it, so this section is its
canonical contract. The producer is emitRunPhaseTiming in
packages/adapter-utils/src/acpx-engine/startup-timing.ts.
The event payload carries only three fields.
| Field | Type | Meaning |
|---|---|---|
phase |
string | The run-lifecycle phase name from the closed allowlist below. |
durationMs |
number | The wall time of the phase. A negative or a non-finite value clamps to 0. |
outcome |
string | The phase outcome (ok or failed). |
The phase field is one member of a closed, low-cardinality allowlist. The
producer drops any event whose phase name is outside this allowlist, so a
free-form label never reaches the run log. The allowlist has twelve phase names.
| Phase | Meaning |
|---|---|
place_workspace |
Place the run workspace. |
start_transport |
Start the agent transport. |
create_runtime |
Create the agent runtime. |
ensure_session |
Ensure the agent session exists. |
configure_session |
Configure the agent session. |
prepare_turn |
Prepare the turn. |
turn |
Run the turn. |
end_session |
End the agent session. |
settle_reuse |
Settle the session for reuse. |
stop_transport |
Stop the agent transport. |
sync_back |
Sync the workspace back. |
release_staging_lease |
Release the staging lease. |
The payload never carries a command, an argument, a path, an environment value,
or a raw identifier. The event rides the ctx.onEvent run-event bridge and is
run-log-only. It needs no OTLP endpoint.
Related instrumentation
The sandbox duplex transport also writes one run-log event as one of its three sinks. See the Sandbox Duplex Transport Instrumentation section in the Observability contract.