Files
PaperClipAI/doc/run-log-events.md
T
DottaandPaperclip b54b2dc35c fix: preserve warm Codex turns with incremental managed file checkpoints (#14735)
## Thinking Path

> - Paperclip manages AI agents and keeps their instructions and files
durable.
> - Native Codex runners can keep a process alive between compatible
turns.
> - Managed file collection stopped that process after each turn, which
defeated warm reuse.
> - Agent folders can contain large images and other files, so full
copies on every turn are expensive.
> - This change keeps one managed directory for the live session and
saves only file changes after each turn.
> - Ownership, authorization, instruction changes, and process
retirement still control when reuse is safe.

## Linked Issues or Issue Description

Related: #13710 introduced native warm session reuse. This fixes managed
file collection that still forced those sessions to stop. No duplicate
open PR or issue was found.

**What happened?**

With managed instructions and warm native Codex enabled, consecutive
turns reused a Daytona sandbox but started a new runner process each
time. The managed directory collector required process termination
before saving files.

**Expected behavior**

Compatible turns keep the same process and managed `AGENT_HOME`. Each
completed turn saves added, changed, and deleted files before the next
turn starts. Unchanged large files do not transfer again.

**Steps to reproduce**

1. Use a native Codex agent with managed instructions and a reusable
Daytona environment.
2. Enable warm session reuse and run three turns on the same task.
3. Write a large binary on the first turn, edit a small note on each
turn, and delete a file on the second turn.
4. Compare process identity across turns and read the canonical files
through the public agent-files API.

**Paperclip version or commit**

Reproduced on `d30b03bd8c17604cdab1533eeeeb087aba30e8b1`.

**Deployment mode**

Local server with remote Daytona execution; cloud native runner uses the
same path.

## What Changed

- Retain the managed directory only for the verified owner of a live
native Codex session.
- Checkpoint each completed turn before releasing the session for reuse.
Retry unstable captures, then stop and collect when a warm checkpoint
cannot be validated.
- Compare metadata and cached hashes, stream only changed file payloads,
record deletions, and validate path, content, quota, and authorization
before saving.
- Rotate sessions when canonical files, loaded instructions,
credentials, or launch policy change. Fence stale collection and cleanup
callbacks from later owners.
- Keep cleanup and recovery aware of the current session owner. Recheck
canonical files under the writer lock at handoff, attach the successor
collector before fallible bookkeeping, and emit one final save receipt
on checkpoint fallback. Preserve storage warnings across unchanged
checkpoints.
- Add regression coverage and a three-turn Daytona test with independent
public API file checks, an unchanged 8 MiB binary, deletion checks, and
strict process identity checks.
- Document checkpoint consistency, lifecycle behavior, and local run-log
counters.
- Replace a timing assumption in the Daytona teardown test with explicit
transfer-arrival gates after CI exposed an unset release callback.

## Verification

- Full local `pnpm -r typecheck` and `pnpm build` passed. Server checks
were repeated after the final storage-warning fix.
- Runner E2E typecheck and 749 runner E2E unit tests passed.
- Focused file checkpoint, directory ownership, instruction collection,
native session, and merge tests passed. After review fixes, the
managed-directory and native-session suites passed 550 tests, including
intervening canonical edits, same-run fresh restore, failed handoff
collection, and one-call fallback collection. Server typecheck passed
again. The Daytona plugin suite passed 218 tests. The quota-warning
regression failed before the fix and passed afterward.
- Three real Daytona campaigns passed before the final handoff review
fixes. The latest kept PID 547 across all three turns. The first
checkpoint copied 8,388,635 bytes; the next two copied 36 and 54 bytes.
Public API reads verified the binary, note contents, and deletion after
every turn. Test cleanup deleted the sandbox.
- The final head was also deployed to an isolated cloud staging instance
and passed three UI-triggered native Codex turns with managed
instructions. All three retained the same process ID/start time, native
session, provider session, runner instance, and Daytona sandbox.
Checkpoints copied 8,388,643 bytes on turn 1, then only 52 and 78 bytes
on turns 2 and 3; those warm captures also hashed only 52 and 78 bytes.
Independent canonical API reads verified every byte of the unchanged 8
MiB binary and the exact note contents after every turn; the deleted
file returned 404 after turns 2 and 3. After restoring the original
lifecycle and agent-auth configuration, removing the temporary secret,
pausing the test agent, and deleting both test sandboxes, independent
canonical API reads still verified the entire binary, the final 78-byte
three-line note, and the deletion. The native runner flag remained
enabled and the final serving revision remained the PR head.
- Two earlier staging attempts are preserved as failures and are
excluded from the acceptance result: a saved ChatGPT login failed with a
provider routing 401, and its subsequent stopped-sandbox retry failed
before provider startup with a closed-lease admission error. The
successful campaign used a fresh sandbox and a temporary encrypted
API-key binding. The stopped-lease retry remains unexplained; this
campaign does not establish recovery of that failed sandbox.
- All [Paperclip CI
gates](https://github.com/paperclipai/paperclip/actions/runs/36750397355)
pass on `26ef2ef56a389259246809805c0b34a4747eb86b`, including full test
partitions, build, typecheck, runner verification, E2E shards, and the
Canary clean public-npm install. Greptile reviewed that exact head at
5/5 with no unresolved review threads or outstanding findings.
- Full local repository coverage used the existing CI partitions, but
the 40,000-file Git streaming stress test timed out and its local retry
was interrupted by macOS thermal emergency sleep; this is not a green
full local suite claim. The exact stress test passed on the final head
in [CI server shard
2/12](https://github.com/paperclipai/paperclip/actions/runs/36750397355/job/110008294290),
in 111.9 seconds.
- Repeat the live test with configured credentials and a Linux runner
artifact: `pnpm test:e2e:runner -- --id
daytona-warm-continuity.runner-codex.daytona.warm-three-turn`.

## Risks

- This is a file-level checkpoint, not an atomic snapshot of the whole
folder. Background writes after a capture are saved by the next
checkpoint or final stopped collection.
- Metadata scans still visit all paths. Modified files transfer in full;
unchanged files do not rehash or transfer.
- Incorrect ownership or reuse could collect the wrong directory. Run
ownership fences, current authorization, stable capture validation, and
stopped collection fallbacks are covered by tests.
- Warm reuse remains opt-in. No database migration or fleet default
changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code editing, tool use, and
test execution. The exact serving model ID and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:33:37 -05:00

19 KiB
Raw Blame History

Run-Log Events

Run-log events write to the heartbeat_run_events table (packages/db/src/schema/heartbeat_run_events.ts:6-20). They are not Paperclip Telemetry events, and they are not OpenTelemetry exports. A run-log event needs no operator endpoint.

Native PRP Run-Log Events

The hidden native coordinator writes each validated PRP event to the bound run's existing event stream before it acknowledges the runner. The row keeps the PRP eventType, source instance, source event ID, source sequence, protocol schema version, and a SHA-256 digest of the canonical source envelope. Its payload is { "prpEvent": <canonical PRP event> }.

PostgreSQL JSONB cannot represent NUL (U+0000), which can occur in command output such as Vite virtual-module paths. The run-event payload column uses a lossless storage codec for these events: the JSONB projection renders NUL as the literal \u0000, and the reserved $paperclipRunEventJsonV1 field contains the original serialized JSON as a doubly escaped string. Ordinary payloads retain their existing representation. Drizzle reads restore the exact original payload before replay, hash validation, redaction, or API presentation. SQL queries can still inspect ordinary routing fields in the projection; raw SQL readers of the whole payload must apply decodeRunEventPayload. The column remains JSONB and requires no schema migration.

The writer locks the native heartbeat_runs row and allocates the existing per-run seq cursor. A byte-equivalent retry reuses the first row; a changed retry or source-sequence gap is rejected. Company, issue, agent, run, session, and runner-source bindings must match the persisted native run. Bootstrap tickets, reconnect leases, authentication proofs, encryption keys, and raw credential material are never written to the run log.

These records remain run-log events. They do not create an OpenTelemetry or Paperclip Telemetry export, and legacy adapters do not use this writer.

Semantic Settlement Diagnostics and Retry Logs

Before interrupting a Codex turn, the runner durably records a harness.diagnostic event with code provider_interrupt_requested, the stop reason, provider turn ID, and pending tool call IDs. A harmless identical result replay records semantic_tool_result_duplicate with warning severity, call ID, operation ID, and result digest. It does not fail the task or create a user attention request. A true conflict includes the call and operation IDs and both result digests in its error; full arguments belong to the canonical semantic input record, not the error message.

Incomplete close emits the bounded native_session_settlement_incomplete runner diagnostic. Its evidence includes runner suspension, provider drain, final provider state, pending call/operation/source-event IDs and input digests, and incomplete result-delivery command IDs and statuses. If execution and cleanup both fail, execution retains its original error identity and cleanup is attached as cleanupError.

Instruction writes also commit an agent.instruction_write_attempted activity row and a run-scoped instructionToolAttempts entry before permitting the filesystem effect. They retain the call ID, operation ID, and input digest, not instruction text. A later transaction rollback cannot erase this attempt proof. A missing success or definite pre-write failure receipt means the outcome is unknown, even if the current file contains the requested text. Known validation and stale-base failures are saved under the attempt's failure entry and replay their original status, message, and details without a new write. Unknown outcomes remain blocked until reconciled; a committed receipt with a lost acknowledgement can replay its exact result.

The NDJSON output log appends across repeated begin calls for one run. Each new handle adds an attemptId to its lines. If the local file is absent and a durable S3 mirror exists, begin restores that prefix before appending. A failed restore must not replace the mirror with an empty or partial attempt. Publishing a restored prefix uses an atomic create-if-absent operation, so a concurrent restore cannot overwrite lines another attempt has already appended. Earlier attempts therefore remain available for incident diagnosis. Existing records without attemptId remain readable.

These records use the instance run log and its configured storage. They add no Paperclip Telemetry or OpenTelemetry export.

Omitted Unsafe Workspace Export

workspace_export_omitted is an informational system event in the local run log. Its payload is { "reason": "restore_unsafe_archive" }, with "legacy": true when recovering an unsafe failure from an older controller. It records that native finalization discarded an unsafe export and continued with the accepted result. It contains no archive names, link targets, or raw error details. It does not create a task warning, recovery action, Telemetry event, or OpenTelemetry export.

Native Restart Recovery Run-Log Event

Paperclip writes a native.recovery.transition event for every native restart classification and for graceful restart suspension. This immutable run-log record lets operators reconstruct recovery decisions without exporting data to Paperclip Telemetry or OpenTelemetry.

The payload contains the restart kind, recovery request id when one exists, runner disposition, and the controller generation and provider attempt for a claimed recovery. Live-runner adoption also records the runner PID, process group, and process-start fingerprint. A non-claim disposition records a bounded reason instead. Graceful suspension records the signal and confirms that it did not create a retry run.

The event never includes bootstrap tickets, reconnect leases, authentication proofs, encryption keys, environment variables, provider credentials, command arguments, or an unsanitized stderr stream. Detailed failed-attempt diagnostics remain in the bounded native_run_finalizations.recovery_history ledger.

Native Local Process Stop Evidence

The server writes native.local_process_stopped in the same transaction that clears a local run's process identity, after it verifies that its PID and process group are absent. The payload contains only those process IDs. Remote process IDs are never checked against the control-plane host.

The server writes native.process_start_requested before a backend can spawn, and native.process_identity_recorded when it stores a new native process identity. Either invalidates an earlier local stop receipt, including a crash before the new PID callback. Continuation admission accepts only the latest server-authored event among these three types; provider source events cannot supply stop authority. These records stay in the local run log and do not add Telemetry or OpenTelemetry data.

Verified Local Codex Replacement Evidence

The server writes native.stopped_text_turn_verified in the same transaction that schedules a fresh successor for a stopped local Codex run. It records the runner and provider process identities, retained-state digests, provider thread and turn IDs, and IDs of exactly receipted task-completion calls. The server first checks the complete turn inventory, process-stop receipt, and execution binding. Unknown actions or changed retained state prevent this event and replacement.

The record documents why the old execution can be retired. It does not make the old session resumable, rewrite provider files, or authorize replay on its own. It remains in the local run log and adds no Telemetry or OpenTelemetry export.

Sandbox Startup Run-Log Event

Paperclip writes one run.startup.step event to the run log for each bring-up step. This event is a run-log record, not a first-party telemetry event. The generated telemetry contract does not cover it, so this section is its canonical contract.

The event payload carries only three fields.

Field Type Meaning
step string The bring-up step name, for example stage.sync.
durationMs number The wall time of the step. A skipped step reports 0.
outcome string The step outcome (ok, skipped, or failed).

The event no longer carries the per-step round-trip count or the provider duration fields. It dropped roundTrips, providerExecMs, providerGetMs, createRuntimeMs, and ensureSessionMs. The startup spans in doc/observability.md carry that detail now. The sandbox.exec child spans hold the round-trip and provider durations. The acp.handshake step span holds the create-runtime and ensure-session sub-times.

To read the detailed timing, use the startup spans. The spans need an OTLP endpoint. A run with no endpoint keeps only the three run-log fields above.

Run Phase Timing Run-Log Event

Paperclip writes one run.phase.timing event to the run log for each run-lifecycle phase. This event is a run-log record, not a first-party telemetry event. The generated telemetry contract does not cover it, so this section is its canonical contract. The producer is emitRunPhaseTiming in packages/adapter-utils/src/acpx-engine/startup-timing.ts.

The event payload carries only three fields.

Field Type Meaning
phase string The run-lifecycle phase name from the closed allowlist below.
durationMs number The wall time of the phase. A negative or a non-finite value clamps to 0.
outcome string The phase outcome (ok or failed).

The phase field is one member of a closed, low-cardinality allowlist. The producer drops any event whose phase name is outside this allowlist, so a free-form label never reaches the run log. The allowlist has twelve phase names.

Phase Meaning
place_workspace Place the run workspace.
start_transport Start the agent transport.
create_runtime Create the agent runtime.
ensure_session Ensure the agent session exists.
configure_session Configure the agent session.
prepare_turn Prepare the turn.
turn Run the turn.
end_session End the agent session.
settle_reuse Settle the session for reuse.
stop_transport Stop the agent transport.
sync_back Sync the workspace back.
release_staging_lease Release the staging lease.

The payload never carries a command, an argument, a path, an environment value, or a raw identifier. The event rides the ctx.onEvent run-event bridge and is run-log-only. It needs no OTLP endpoint.

ACP terminal failure diagnostics

The shared ACP adapter engine preserves typed terminal session failures in the run error, the acpx.error transcript record, and heartbeat_runs.result_json.terminalSessionFailure. The structured diagnostic contains the provider category, title, and details. It works when raw provider tracing is disabled. The existing UI and CLI render the diagnostic as an error, not as assistant output or an automatic task response. Issue continuation summaries and session-compaction handoffs retain only the generic failure category; provider diagnostic prose is not copied into prompts.

Both pinned ACPX patches pass complete title and detail strings to the in-memory callback. The engine redacts configured environment values (including resolved secrets with arbitrary variable names), launch environment values outside a closed allowlist of public process settings, known boolean flags, and run identifiers, connection URL passwords, the run API key, and common credential forms before truncation. It removes control characters, retains line breaks for JSON and stack traces, and preserves up to 4,096 title characters and 24,576 detail characters. These bounds also keep the escaped transcript JSON below the server's 64 KiB chunk limit. Longer fields end with an explicit omission count and appear in truncatedFields. The error message includes the same sanitized text. Ordinary run retrieval preserves the bounded structured diagnostic even when multibyte text or other result fields exceed the result byte budget. In that reduced response, title and details have 1 KiB and 8 KiB byte budgets, including truncation markers. retrievalTruncated directs callers to the full adapter-bounded text in the run error or transcript. Other provider metadata and action payloads are not copied.

Recovery still uses the typed failure category and the adapter's existing classifier. Provider warnings do not become failures, and timeouts or lost control channels keep their authoritative failure messages. These diagnostics stay in the instance's run records and configured run-log storage. They add no Paperclip Telemetry or OpenTelemetry export.

The sandbox duplex transport also writes one run-log event as one of its three sinks. See the Sandbox Duplex Transport Instrumentation section in the Observability contract.

Execution recovery

Provider identity diagnostics remain in the local run log. They record the notification method, expected and received thread/turn identifiers, and the classification (root, verified descendant, stale, unrelated informational, or invalid authoritative). They omit the original provider payload and credentials. Repeated informational notices are bounded.

Recovery lifecycle events retain the original structured failure code, retry attempt, next retry time, and predecessor/successor identifiers. Durable status delivery uses an idempotency marker; delivery grants no provider authority. Failed publication is retried without repeating provider work. These records are not first-party Telemetry.

If execution-continuation setup finds that a task no longer exists, is closed, or its owner changed, the existing cancellation settlement records continuation_task_ownership_changed. The run and wake request become cancelled before adapter dispatch, with the run-log message stale execution continuation cancelled before dispatch. Immediate recovery is suppressed. Missing source context, authorization failures, and other setup errors retain their failure classification. An untyped error with the same message is also still a failure; cancellation requires the typed ownership guard.

Bounded retry exhaustion writes one lifecycle receipt per run, retry reason, scheduled attempt, and retry limit. Repeated or concurrent recovery checks reuse that receipt, including receipts from earlier builds, without advancing the event sequence or publishing another live event. Attention reads select the latest matching receipt in PostgreSQL and project only the run's issue/task identifiers from its context, so historical duplicate receipts cannot multiply run contexts in server memory. Existing duplicate events do not require deletion or migration.

Workspace restore failures

Legacy adapter results can carry workspaceRestoreFailure with the code restore_permission_denied, restore_lock_timeout, restore_unsafe_archive, or restore_failed. The heartbeat records workspace_restore_failed and keeps the run failed and the workspace-finalization barrier closed. Available output, usage, session metadata, and the previous execution outcome survive settlement. executionBeforeRestore retains the earlier error code, exit code, signal, and timeout flag. The ordinary redacted error field retains an earlier error message.

The chat reports the restore phase separately from a missing final response. A saved-plan link requires a stored document and its run-bound revision or a matching run-bound review record. Older unclassified failures use neutral wording. Diagnostics show only a validated relative member path, never an archive link target or host path.

An unsafe archive or an outbound confinement refusal keeps the existing execution recovery hold, including across conversation resets. It cannot start another model turn until an operator uses the existing recovery action to record executionReconciliation with workspaceRepairEvidence (20–12000 characters). This evidence must describe verified safe staging or repair for the referenced failed run. It does not grant plan approval. Saved comments, document revisions, and confirmation IDs and states stay unchanged. Recovery uses the existing delivery identity and links the successor to the original failed run. Repair does not reset the automatic retry budget. Transient failures retain the existing bounded retry policy. Archive confinement remains required.

Sandbox restore tasks also write a Workspace restore diagnostic line to the run log on failure. phase is workspace or asset, so a failed staged-asset copy-back (such as credentials) can be distinguished from workspace restoration. The line contains only an allowlisted OS/transport errorCode (otherwise unknown), an optional numeric HTTP error status, and an optional bounded process exit code. Up to four nested causes are inspected. Messages, URLs, filesystem paths, asset names, credentials, and response bodies are excluded. Every failed outbound task emits its own diagnostic; nested repository failures are logged once by the enclosing workspace task. The original error and restore safety policy are unchanged. These lines stay in the instance run log and its configured durable storage, and are not new first-party telemetry events.

Codex resume usage snapshot

The native runner retains a bounded local harness.diagnostic event with code codex_resume_usage_snapshot. It identifies thread/tokenUsage/updated as resume_usage_snapshot, retains the reported thread and completed-turn IDs, and records cumulative usage counters. It does not include provider credentials or message content. The event establishes the accounting baseline; it is not a new billable usage receipt or a user-facing provider warning. Other provider identity checks remain in force.

AI subscription contention

A fresh task execution cannot enter this wait. A run that already entered this wait writes an informational lifecycle event to the local run log. Its payload contains only retryScheduled, a boolean that reports whether the scheduler created a retry. The message distinguishes an automatic retry from work that is no longer eligible. This pre-provider wait records ai_connection_busy on the cancelled run and does not consume the provider-failure retry allowance. The event contains no credentials and creates no Telemetry or OpenTelemetry export.

Managed Agent File Save Receipts

The server writes instruction_save after managed file collection or a warm turn checkpoint. The payload includes the save state, instruction entry path, storage warning, and error code/message. Agent-directory receipts identify contract: "agent_files" and the applied candidate hash. Legacy instruction receipts instead identify the saved revision.

A validated warm checkpoint reports saved or unchanged, even though its working directory remains owned by the live session. An unstable checkpoint reports pending_collection until stopped collection produces a final receipt. Successful checkpoints can include checkpointStats: scannedEntries, hashedBytes, copiedFiles, and copiedBytes. These counts describe that capture, not cumulative traffic or an atomic snapshot of background writers. They contain no file contents. The receipt remains in the instance run log; it adds no Paperclip Telemetry or OpenTelemetry export.