Files
PaperClipAI/doc/architecture/durable-continuation-scheduler.md
T
DottaandDev Agent 25cf079ec5 feat(runner): add Codex-native application integration (#12591)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner package is useful only when the application can start,
observe, and recover a native Codex run safely.
> - Existing direct adapters must keep their current execution and
finalization paths.
> - The application boundary therefore needs additive persistence,
authorization, coordination, and recovery behind an explicit
experimental adapter.
> - This pull request adds that Codex-only boundary without activating
generalized providers, remote environments, or the later task/SDK
surfaces.

## Linked Issues or Issue Description

**Subsystem affected**

Shared contracts, database persistence, adapter utilities, server
native-runtime services, and the experimental Paperclip Runner adapter.

**Problem or motivation**

The already-landed runner package has a qualified Codex path, but the
application needs durable native-run state, guarded runtime selection,
authenticated coordination, tool security, finalization, and recovery
before the experimental adapter can be exercised safely.

**Proposed solution**

Add a Codex-only `paperclip_runner` application path behind the existing
default-off native-runner setting. Bind native state and coordination to
company/run identity, preserve persisted-run recovery, and leave every
direct adapter on its existing legacy execution path.

**Alternatives considered**

The earlier stack boundary introduced a generalized executor and
remote-environment lifecycle here. That made this PR depend on
implementations in higher PRs and changed reusable sandbox behavior
globally. Those pieces are now deferred together to #12592.

**Roadmap alignment**

ROADMAP.md does not list a conflicting native-runner integration
project. This change adds the application boundary for the existing
Runner architecture.

## What Changed

- Added native run/result/finalization/provider-trace persistence,
shared validators, and idempotent migration/replay coverage.
- Added guarded Codex-only runtime selection, authenticated PRP
coordination, recovery, finalization, and interaction services.
- Added run/company-bound tool-gateway authorization, credential
redaction, SSRF protections, and replay-safe behavior.
- Added the explicit `paperclip_runner` adapter behind the default-off
rollout setting.
- Preserved legacy answered-question wake projection and direct-adapter
execution/finalization paths.
- Hardened cancellation so only owned in-memory child processes are
signaled; persisted recycled PIDs/process groups are never trusted.
- Retained the narrow Claude ACPX isolated-context security follow-up
discovered after #12590.
- Deferred the generalized executor, provider ingress, remote lifecycle,
SDK/lab/eval work, release-process changes, and lockfile.

## Verification

- Changed-file delta against `master`: 133 files.
- GitHub Actions is the authoritative verification environment for this
PR.
- Full CI, security, and Greptile review will run on this lowest
unmerged stack PR.
- Local tests/build/typecheck were not run because this checkout is
resource constrained.
- Static diff/reference checks pass, and `pnpm-lock.yaml` is unchanged.

## Risks

- This touches central heartbeat and agent-route code, so legacy
compatibility is the primary risk.
- Runtime selection remains Codex-only and explicit; direct Codex,
Claude, OpenCode, process, HTTP, and plugin adapters remain on their
existing paths.
- Fresh native starts fail closed while the rollout flag is off;
persisted native records remain readable and recoverable.
- Cancellation, company/run binding, tool calls, status decisions, and
completion writes are guarded or replay-safe.

> For core feature work, check [ROADMAP.md](ROADMAP.md) first and
discuss it in #dev before opening the PR. Feature PRs that overlap with
planned core work may need to be redirected.

## Model Used

OpenAI Codex, GPT-5.6, with repository tools, code execution, and
parallel agent review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
following the relevant issue template
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass — GitHub Actions is
authoritative for this resource-constrained checkout
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI and security gates are green
- [ ] Greptile is 5/5 with no open actionable findings
- [x] I will address all Greptile and reviewer comments before merge

## Stack

- Position: 3 of 5 overall; lowest of 3 currently unmerged
- Base: `master`
- Previous:
[#12590](https://github.com/paperclipai/paperclip/pull/12590), qualified
Claude ACPX runtime — merged
- Next: [#12592](https://github.com/paperclipai/paperclip/pull/12592),
generalized Codex executor, task experience, and developer SDKs

---------

Co-authored-by: Dev Agent <dev@paperclip.ing>
2026-08-31 14:38:38 -05:00

6.7 KiB

Durable continuation scheduling

Paperclip does not keep an agent process alive between turns. A heartbeat run is finite: it starts, performs work, records a terminal result, and exits. If work must continue later, Paperclip represents that intent in database state and creates another heartbeat run when the continuation becomes eligible.

“Durable continuation scheduler” is a useful umbrella term, but it is not the name of one class or queue. Two related mechanisms provide the behavior:

  1. Explicit continuation effects are written by native status arbitration.
  2. Stranded-issue reconciliation is a startup and periodic safety net for an assigned open issue that has no live execution or durable wait path.

The second mechanism produced the repeated issue_continuation_needed runs on DOT-2.

When it runs

The heartbeat scheduler is enabled unless HEARTBEAT_SCHEDULER_ENABLED=false. Its interval is configured with HEARTBEAT_SCHEDULER_INTERVAL_MS and defaults to 30,000 ms. The configured value is clamped to a minimum of 10,000 ms.

Continuation recovery runs:

  • once during server startup, after orphaned runs are reaped, due retries are promoted, and already-queued work is resumed; and
  • on every heartbeat scheduler tick, after the same orphan/retry/queue cleanup.

The periodic tick calls reconcileStrandedAssignedIssues(). Consequently, a new recovery run normally appears within one scheduler interval. The interval is polling cadence, not a promise that every continuation waits exactly that long.

Explicit native continuation effects do not need to wait for this scan. The native status-decision committer writes an idempotent agent_wakeup_requests row directly. The normal queued-run machinery then claims it.

What is durable

The system reconstructs intent from persisted control-plane records rather than an in-memory timer owned by an agent:

  • the issue status and assignee;
  • the issue execution lock/run identity;
  • heartbeat run status, context, retry ancestry, and terminal timestamps;
  • agent_wakeup_requests rows;
  • native status decisions and their materialized effects;
  • pending interactions, approvals, monitors, blockers, and execution stages;
  • scheduled retry timestamps and recovery actions.

Because these records survive a process restart, startup reconciliation can resume queued work or repair an issue whose previous execution disappeared.

Explicit continuations

A native result can report yielded with a continuation containing:

  • a kind: same_agent, retry, delegated_issue, response_wake, or monitor;
  • a summary; and
  • an idempotency key.

The native status arbiter keeps the issue in_progress and emits an enqueue_continuation effect. The status-decision committer materializes that effect as an idempotent wake request. A monitor continuation uses monitor_due; other continuation kinds use issue_status_changed at this boundary.

This is intentional continuation: the result explicitly says more work is needed and leaves a persisted execution path.

Stranded-issue reconciliation

reconcileStrandedAssignedIssues() scans agent-owned issues in todo, in_progress, and relevant in_review states. Before creating work, it checks for reasons not to wake the agent, including:

  • an active or queued execution path already exists;
  • a pending interaction or durable wait path exists;
  • the issue or its tree is paused;
  • a provider-quota monitor is pending;
  • the run was explicitly cancelled by an operator;
  • the assigned agent is not invokable;
  • invocation budget or retry/backoff policy prevents a run; or
  • recovery has failed enough times that the issue should be escalated instead.

For an eligible in_progress issue with no live path, it creates an automated wake/run with:

wakeReason:  issue_continuation_needed
retryReason: issue_continuation_needed
source:      issue.continuation_recovery

This is a liveness repair, not evidence that the previous model requested another turn. Its invariant is: an assigned open issue should have either a live execution path, a durable reason to wait, or a visible terminal/blocking disposition.

Why DOT-2 looped

DOT-2's runner repeatedly produced a successful result that reported done. The native evidence classifier did not accept the model's evidence references, so status arbitration preserved in_progress and created another continuation path. After each run exited, the periodic reconciler saw:

assigned + in_progress + successful terminal run + no live/wait path

It therefore queued issue_continuation_needed again on the next approximately 30-second tick. The new run received the original task title because the native model envelope also omitted the wake-comment text, so it repeated the same answer.

The fix has two parts:

  • native model input now contains the redacted server-authored task prompt, including the current wake comment; and
  • ordinary low-risk issue completion can accept a schema-valid completion claim, while tools, governed effects, interactions, and approvals keep their independent authorization gates.

Thus a completed one-shot task becomes done; it is no longer a candidate for stranded-issue reconciliation.

How to diagnose a suspected loop

Inspect the latest runs for the issue and compare these fields:

  • invocationSource — recovery-created runs use automation;
  • contextSnapshot.wakeReason;
  • contextSnapshot.retryReason;
  • retryOfRunId;
  • terminal run status and resultJson.authoritativeDecision;
  • resultJson.issueStatusAfter;
  • issue status, executionRunId, and pending interactions/monitors;
  • lifecycle events explaining enqueue, suppression, escalation, or cleanup.

Repeated successful runs with wakeReason: issue_continuation_needed, an authoritative decision of in_progress, and no durable wait path usually mean the result/disposition contract is failing to converge. Fix the status or continuation decision; increasing the polling interval only hides the bug.

Primary implementation locations

  • server/src/index.ts — startup recovery and periodic heartbeat scheduler.
  • server/src/config.ts — scheduler enable flag and interval.
  • server/src/services/recovery/service.ts — reconcileStrandedAssignedIssues() and recovery wake creation.
  • server/src/services/issue-rewake-throttle.ts — repeated state-wake throttling and progress detection.
  • server/src/services/native-runtime/status-arbiter.ts — native terminal disposition and explicit continuation decisions.
  • server/src/services/native-runtime/status-decision-committer.ts — durable, idempotent materialization of native continuation effects.
  • server/src/services/heartbeat.ts — run lifecycle, immediate recovery, queue promotion, and wake execution.