mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A comment can queue execution before the task changes state. > - The task may be completed, cancelled, deleted, or reassigned before setup checks ownership. > - The ownership guard must prevent that obsolete execution from starting. > - This pull request records that guard outcome as cancellation and settles the wake request. > - Missing history and authorization errors remain failures that operators can investigate. ## Linked Issues or Issue Description **What happened?** A comment can queue a run just before a user marks its task Done. The ownership guard stops setup before adapter dispatch, but the run becomes a setup failure and can put the agent into an error state. **Expected behavior** Cancel obsolete work when the typed task ownership guard rejects it. Keep genuine setup errors visible. Do not change the task's terminal state or current owner. **Steps to reproduce** Queue a task continuation, then mark the task Done or Cancelled before continuation setup. The run should settle as cancelled without invoking the adapter. Deleting or reassigning the task must also prevent dispatch. Missing source history or explicit user authorization must still fail. **Paperclip version or commit** Refreshed against master `b2e9e82f053af14777fd68e8834fff6bb64c84c4`. Related: #13546 covers lost issue-lock claims, and #13888 covers persisted continuation decisions and retry budgets. This change handles the final continuation ownership guard during setup. Thanks to @MrBlackTongue for the original typed cancellation implementation and regression coverage. This update preserves the contributor's commits, resolves the master conflict, and narrows cancellation to task ownership invalidation. ## What Changed - Add a typed `StaleExecutionContinuationError` for `continuation_task_ownership_changed`. - Use existing cancellation settlement for that typed guard, including the run, wake request, issue execution ownership, and agent state. Suppress immediate recovery of obsolete work. - Preserve failure classification for missing source context, missing user authorization, and untyped errors, including an untyped error with identical text. - Cover Done and Cancelled tasks with real database checks. Retain company and assignment guard coverage. Verify cancellation performs no adapter dispatch or automatic replay. - Document the cancellation event and its error classification in the run-log guide. ## Verification Current head: `b875486e81`. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/services/execution-continuation.test.ts --maxWorkers=1 --no-file-parallelism -t 'authorized continuation context|obsolete continuation setup|untyped continuation setup failures'`: 17 passed; 316 unrelated tests excluded by the name filter. - `pnpm test:run`: local full run is still finishing. It is not a clean pass: the observed failures are three company-skills cache checks (`EACCES` while renaming read-only cache directories on macOS), one tool-access assertion, and one comment-wake timeout. The three cache failures reproduce in isolation in unchanged code. The tool-access check passes in isolation, and the entire comment-wake file passes (29 tests), including explicit feedback after completion. Final full-run totals will be added when available. - [CI](https://github.com/paperclipai/paperclip/actions/runs/36282532652): all gates green on this head. An unchanged Runner workspace-diff test initially returned no diff; an isolated check passed, and the failed job plus its dependent gate passed on retry with no code change. Original failed job: 108517010989. - Greptile: 5/5 on the full current head, with no unresolved review threads. Branch is current with master and conflict-free. - Diff check and local secret/PII scan passed. No customer data or private deployment identifiers are included. ## Risks The classification change is limited to the typed ownership guard. A task with missing history remains a failure rather than being treated as an expected cancellation. Existing company, task-owner, terminal-state, and authorization checks still prevent dispatch. Cancellation retains its specific reason in the run log and does not grant a retry. No schema, dependency, or API changes; no migration is required. ## Model Used OpenAI GPT-6 via Codex assisted investigation, implementation, code review, and test execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the bug issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused regression tests locally and they pass; full-suite results are tracked above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented the risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before merge Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Devin Foley <devin@paperclip.ing>
237 lines
13 KiB
Markdown
237 lines
13 KiB
Markdown
# Run-Log Events
|
||
|
||
Run-log events write to the `heartbeat_run_events` table
|
||
(`packages/db/src/schema/heartbeat_run_events.ts:6-20`). They are not
|
||
Paperclip Telemetry events, and they are not OpenTelemetry exports. A run-log
|
||
event needs no operator endpoint.
|
||
|
||
## Native PRP Run-Log Events
|
||
|
||
The hidden native coordinator writes each validated PRP event to the bound
|
||
run's existing event stream before it acknowledges the runner. The row keeps
|
||
the PRP `eventType`, source instance, source event ID, source sequence, protocol
|
||
schema version, and a SHA-256 digest of the canonical source envelope. Its
|
||
payload is `{ "prpEvent": <canonical PRP event> }`.
|
||
|
||
PostgreSQL JSONB cannot represent NUL (U+0000), which can occur in command
|
||
output such as Vite virtual-module paths. The run-event payload column uses a
|
||
lossless storage codec for these events: the JSONB projection renders NUL as
|
||
the literal `\u0000`, and the reserved `$paperclipRunEventJsonV1` field contains
|
||
the original serialized JSON as a doubly escaped string. Ordinary payloads
|
||
retain their existing representation. Drizzle reads restore the exact original
|
||
payload before replay, hash validation, redaction, or API presentation. SQL
|
||
queries can still inspect ordinary routing fields in the projection; raw SQL
|
||
readers of the whole payload must apply `decodeRunEventPayload`. The column
|
||
remains JSONB and requires no schema migration.
|
||
|
||
The writer locks the native `heartbeat_runs` row and allocates the existing
|
||
per-run `seq` cursor. A byte-equivalent retry reuses the first row; a changed
|
||
retry or source-sequence gap is rejected. Company, issue, agent, run, session,
|
||
and runner-source bindings must match the persisted native run. Bootstrap
|
||
tickets, reconnect leases, authentication proofs, encryption keys, and raw
|
||
credential material are never written to the run log.
|
||
|
||
These records remain run-log events. They do not create an OpenTelemetry or
|
||
Paperclip Telemetry export, and legacy adapters do not use this writer.
|
||
|
||
## Native Restart Recovery Run-Log Event
|
||
|
||
Paperclip writes a `native.recovery.transition` event for every native restart
|
||
classification and for graceful restart suspension. This immutable run-log
|
||
record lets operators reconstruct recovery decisions without exporting data to
|
||
Paperclip Telemetry or OpenTelemetry.
|
||
|
||
The payload contains the restart kind, recovery request id when one exists,
|
||
runner disposition, and the controller generation and provider attempt for a
|
||
claimed recovery. Live-runner adoption also records the runner PID, process
|
||
group, and process-start fingerprint. A non-claim disposition records a bounded
|
||
reason instead. Graceful suspension records the signal and confirms that it did
|
||
not create a retry run.
|
||
|
||
The event never includes bootstrap tickets, reconnect leases, authentication
|
||
proofs, encryption keys, environment variables, provider credentials, command
|
||
arguments, or an unsanitized stderr stream. Detailed failed-attempt diagnostics
|
||
remain in the bounded `native_run_finalizations.recovery_history` ledger.
|
||
|
||
## Native Local Process Stop Evidence
|
||
|
||
The server writes `native.local_process_stopped` in the same transaction that
|
||
clears a local run's process identity, after it verifies that its PID and process
|
||
group are absent. The payload contains only those process IDs. Remote process
|
||
IDs are never checked against the control-plane host.
|
||
|
||
The server writes `native.process_start_requested` before a backend can spawn,
|
||
and `native.process_identity_recorded` when it stores a new native process
|
||
identity. Either invalidates an earlier local stop receipt, including a crash
|
||
before the new PID callback. Continuation admission accepts only the latest
|
||
server-authored event among these three types; provider
|
||
source events cannot supply stop authority. These records stay in the local run
|
||
log and do not add Telemetry or OpenTelemetry data.
|
||
|
||
## Verified Local Codex Replacement Evidence
|
||
|
||
The server writes `native.stopped_text_turn_verified` in the same transaction
|
||
that schedules a fresh successor for a stopped local Codex run. It records the
|
||
runner and provider process identities, retained-state digests, provider thread
|
||
and turn IDs, and IDs of exactly receipted task-completion calls. The server first
|
||
checks the complete turn inventory, process-stop receipt, and execution binding.
|
||
Unknown actions or changed retained state prevent this event and replacement.
|
||
|
||
The record documents why the old execution can be retired. It does not make the
|
||
old session resumable, rewrite provider files, or authorize replay on its own.
|
||
It remains in the local run log and adds no Telemetry or OpenTelemetry export.
|
||
|
||
## Sandbox Startup Run-Log Event
|
||
|
||
Paperclip writes one `run.startup.step` event to the run log for each bring-up
|
||
step. This event is a run-log record, not a first-party telemetry event. The
|
||
generated telemetry contract does not cover it, so this section is its canonical
|
||
contract.
|
||
|
||
The event payload carries only three fields.
|
||
|
||
| Field | Type | Meaning |
|
||
| --- | --- | --- |
|
||
| `step` | string | The bring-up step name, for example `stage.sync`. |
|
||
| `durationMs` | number | The wall time of the step. A skipped step reports `0`. |
|
||
| `outcome` | string | The step outcome (`ok`, `skipped`, or `failed`). |
|
||
|
||
The event no longer carries the per-step round-trip count or the provider
|
||
duration fields. It dropped `roundTrips`, `providerExecMs`, `providerGetMs`,
|
||
`createRuntimeMs`, and `ensureSessionMs`. The startup spans in
|
||
[`doc/observability.md`](observability.md) carry that detail now. The
|
||
`sandbox.exec` child spans hold the round-trip and provider durations. The
|
||
`acp.handshake` step span holds the create-runtime and ensure-session
|
||
sub-times.
|
||
|
||
To read the detailed timing, use the startup spans. The spans need an OTLP
|
||
endpoint. A run with no endpoint keeps only the three run-log fields above.
|
||
|
||
## Run Phase Timing Run-Log Event
|
||
|
||
Paperclip writes one `run.phase.timing` event to the run log for each
|
||
run-lifecycle phase. This event is a run-log record, not a first-party telemetry
|
||
event. The generated telemetry contract does not cover it, so this section is its
|
||
canonical contract. The producer is `emitRunPhaseTiming` in
|
||
`packages/adapter-utils/src/acpx-engine/startup-timing.ts`.
|
||
|
||
The event payload carries only three fields.
|
||
|
||
| Field | Type | Meaning |
|
||
| --- | --- | --- |
|
||
| `phase` | string | The run-lifecycle phase name from the closed allowlist below. |
|
||
| `durationMs` | number | The wall time of the phase. A negative or a non-finite value clamps to `0`. |
|
||
| `outcome` | string | The phase outcome (`ok` or `failed`). |
|
||
|
||
The `phase` field is one member of a closed, low-cardinality allowlist. The
|
||
producer drops any event whose phase name is outside this allowlist, so a
|
||
free-form label never reaches the run log. The allowlist has twelve phase names.
|
||
|
||
| Phase | Meaning |
|
||
| --- | --- |
|
||
| `place_workspace` | Place the run workspace. |
|
||
| `start_transport` | Start the agent transport. |
|
||
| `create_runtime` | Create the agent runtime. |
|
||
| `ensure_session` | Ensure the agent session exists. |
|
||
| `configure_session` | Configure the agent session. |
|
||
| `prepare_turn` | Prepare the turn. |
|
||
| `turn` | Run the turn. |
|
||
| `end_session` | End the agent session. |
|
||
| `settle_reuse` | Settle the session for reuse. |
|
||
| `stop_transport` | Stop the agent transport. |
|
||
| `sync_back` | Sync the workspace back. |
|
||
| `release_staging_lease` | Release the staging lease. |
|
||
|
||
The payload never carries a command, an argument, a path, an environment value,
|
||
or a raw identifier. The event rides the `ctx.onEvent` run-event bridge and is
|
||
run-log-only. It needs no OTLP endpoint.
|
||
|
||
## Related instrumentation
|
||
|
||
The sandbox duplex transport also writes one run-log event as one of its three
|
||
sinks. See the
|
||
[Sandbox Duplex Transport Instrumentation](observability.md#sandbox-duplex-transport-instrumentation)
|
||
section in the Observability contract.
|
||
|
||
## Execution recovery
|
||
|
||
Provider identity diagnostics remain in the local run log. They record the notification method, expected and received thread/turn identifiers, and the classification (root, verified descendant, stale, unrelated informational, or invalid authoritative). They omit the original provider payload and credentials. Repeated informational notices are bounded.
|
||
|
||
Recovery lifecycle events retain the original structured failure code, retry attempt, next retry time, and predecessor/successor identifiers. Durable status delivery uses an idempotency marker; delivery grants no provider authority. Failed publication is retried without repeating provider work. These records are not first-party Telemetry.
|
||
|
||
If execution-continuation setup finds that a task no longer exists, is closed,
|
||
or its owner changed, the existing cancellation settlement records
|
||
`continuation_task_ownership_changed`. The run and wake request become cancelled
|
||
before adapter dispatch, with the run-log message
|
||
`stale execution continuation cancelled before dispatch`. Immediate recovery is
|
||
suppressed. Missing source context, authorization failures, and other setup
|
||
errors retain their failure classification. An untyped error with the same
|
||
message is also still a failure; cancellation requires the typed ownership guard.
|
||
|
||
Bounded retry exhaustion writes one lifecycle receipt per run, retry reason,
|
||
scheduled attempt, and retry limit. Repeated or concurrent recovery checks reuse
|
||
that receipt, including receipts from earlier builds, without advancing the event
|
||
sequence or publishing another live event. Attention reads select the latest
|
||
matching receipt in PostgreSQL and project only the run's issue/task identifiers
|
||
from its context, so historical duplicate receipts cannot multiply run contexts
|
||
in server memory. Existing duplicate events do not require deletion or migration.
|
||
|
||
### Workspace restore failures
|
||
|
||
Legacy adapter results can carry `workspaceRestoreFailure` with the code
|
||
`restore_permission_denied`, `restore_lock_timeout`, `restore_unsafe_archive`,
|
||
or `restore_failed`. The heartbeat records `workspace_restore_failed` and keeps
|
||
the run failed and the workspace-finalization barrier closed. Available output,
|
||
usage, session metadata, and the previous execution outcome survive settlement.
|
||
`executionBeforeRestore` retains the earlier error code, exit code, signal, and
|
||
timeout flag. The ordinary redacted error field retains an earlier error message.
|
||
|
||
The chat reports the restore phase separately from a missing final response.
|
||
A saved-plan link requires a stored document and its run-bound revision or a
|
||
matching run-bound review record. Older unclassified failures use neutral wording. Diagnostics show only
|
||
a validated relative member path, never an archive link target or host path.
|
||
|
||
An unsafe archive or an outbound confinement refusal keeps the existing execution recovery hold, including across
|
||
conversation resets. It cannot start another model turn until an operator uses
|
||
the existing recovery action to record `executionReconciliation` with
|
||
`workspaceRepairEvidence` (20–12000 characters). This evidence must describe
|
||
verified safe staging or repair for the referenced failed run. It does not grant
|
||
plan approval. Saved comments, document revisions, and confirmation IDs and
|
||
states stay unchanged. Recovery uses the existing delivery identity and links
|
||
the successor to the original failed run. Repair does not reset the automatic
|
||
retry budget. Transient failures retain the existing
|
||
bounded retry policy. Archive confinement remains required.
|
||
|
||
Sandbox restore tasks also write a `Workspace restore diagnostic` line to the
|
||
run log on failure. `phase` is `workspace` or `asset`, so a failed staged-asset
|
||
copy-back (such as credentials) can be distinguished from workspace restoration. The line
|
||
contains only an allowlisted OS/transport `errorCode` (otherwise `unknown`), an
|
||
optional numeric HTTP error status, and an optional bounded process exit code.
|
||
Up to four nested causes are inspected. Messages, URLs, filesystem paths, asset
|
||
names, credentials, and response bodies are excluded. Every failed outbound
|
||
task emits its own diagnostic; nested repository failures are logged once by
|
||
the enclosing workspace task. The original error and restore safety policy are
|
||
unchanged. These lines stay in the instance run log and its configured durable
|
||
storage, and are not new first-party telemetry events.
|
||
|
||
## Codex resume usage snapshot
|
||
|
||
The native runner retains a bounded local `harness.diagnostic` event with code
|
||
`codex_resume_usage_snapshot`. It identifies `thread/tokenUsage/updated` as
|
||
`resume_usage_snapshot`, retains the reported thread and completed-turn IDs,
|
||
and records cumulative usage counters. It does not include provider credentials
|
||
or message content. The event establishes the accounting baseline; it is not a
|
||
new billable usage receipt or a user-facing provider warning. Other provider
|
||
identity checks remain in force.
|
||
|
||
## AI subscription contention
|
||
|
||
A fresh task execution cannot enter this wait. A run that already entered this
|
||
wait writes an informational `lifecycle` event to the local run log. Its
|
||
payload contains only `retryScheduled`, a boolean that reports
|
||
whether the scheduler created a retry.
|
||
The message distinguishes an automatic retry from work that is no longer eligible.
|
||
This pre-provider wait records `ai_connection_busy` on the cancelled run and does
|
||
not consume the provider-failure retry allowance. The event contains no credentials
|
||
and creates no Telemetry or OpenTelemetry export.
|