mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package is useful only when the application can start, observe, and recover a native Codex run safely. > - Existing direct adapters must keep their current execution and finalization paths. > - The application boundary therefore needs additive persistence, authorization, coordination, and recovery behind an explicit experimental adapter. > - This pull request adds that Codex-only boundary without activating generalized providers, remote environments, or the later task/SDK surfaces. ## Linked Issues or Issue Description **Subsystem affected** Shared contracts, database persistence, adapter utilities, server native-runtime services, and the experimental Paperclip Runner adapter. **Problem or motivation** The already-landed runner package has a qualified Codex path, but the application needs durable native-run state, guarded runtime selection, authenticated coordination, tool security, finalization, and recovery before the experimental adapter can be exercised safely. **Proposed solution** Add a Codex-only `paperclip_runner` application path behind the existing default-off native-runner setting. Bind native state and coordination to company/run identity, preserve persisted-run recovery, and leave every direct adapter on its existing legacy execution path. **Alternatives considered** The earlier stack boundary introduced a generalized executor and remote-environment lifecycle here. That made this PR depend on implementations in higher PRs and changed reusable sandbox behavior globally. Those pieces are now deferred together to #12592. **Roadmap alignment** ROADMAP.md does not list a conflicting native-runner integration project. This change adds the application boundary for the existing Runner architecture. ## What Changed - Added native run/result/finalization/provider-trace persistence, shared validators, and idempotent migration/replay coverage. - Added guarded Codex-only runtime selection, authenticated PRP coordination, recovery, finalization, and interaction services. - Added run/company-bound tool-gateway authorization, credential redaction, SSRF protections, and replay-safe behavior. - Added the explicit `paperclip_runner` adapter behind the default-off rollout setting. - Preserved legacy answered-question wake projection and direct-adapter execution/finalization paths. - Hardened cancellation so only owned in-memory child processes are signaled; persisted recycled PIDs/process groups are never trusted. - Retained the narrow Claude ACPX isolated-context security follow-up discovered after #12590. - Deferred the generalized executor, provider ingress, remote lifecycle, SDK/lab/eval work, release-process changes, and lockfile. ## Verification - Changed-file delta against `master`: 133 files. - GitHub Actions is the authoritative verification environment for this PR. - Full CI, security, and Greptile review will run on this lowest unmerged stack PR. - Local tests/build/typecheck were not run because this checkout is resource constrained. - Static diff/reference checks pass, and `pnpm-lock.yaml` is unchanged. ## Risks - This touches central heartbeat and agent-route code, so legacy compatibility is the primary risk. - Runtime selection remains Codex-only and explicit; direct Codex, Claude, OpenCode, process, HTTP, and plugin adapters remain on their existing paths. - Fresh native starts fail closed while the rollout flag is off; persisted native records remain readable and recoverable. - Cancellation, company/run binding, tool calls, status decisions, and completion writes are guarded or replay-safe. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI and security gates are green - [ ] Greptile is 5/5 with no open actionable findings - [x] I will address all Greptile and reviewer comments before merge ## Stack - Position: 3 of 5 overall; lowest of 3 currently unmerged - Base: `master` - Previous: [#12590](https://github.com/paperclipai/paperclip/pull/12590), qualified Claude ACPX runtime — merged - Next: [#12592](https://github.com/paperclipai/paperclip/pull/12592), generalized Codex executor, task experience, and developer SDKs --------- Co-authored-by: Dev Agent <dev@paperclip.ing>
162 lines
6.7 KiB
Markdown
162 lines
6.7 KiB
Markdown
# Durable continuation scheduling
|
|
|
|
Paperclip does not keep an agent process alive between turns. A heartbeat run is
|
|
finite: it starts, performs work, records a terminal result, and exits. If work
|
|
must continue later, Paperclip represents that intent in database state and
|
|
creates another heartbeat run when the continuation becomes eligible.
|
|
|
|
“Durable continuation scheduler” is a useful umbrella term, but it is not the
|
|
name of one class or queue. Two related mechanisms provide the behavior:
|
|
|
|
1. **Explicit continuation effects** are written by native status arbitration.
|
|
2. **Stranded-issue reconciliation** is a startup and periodic safety net for an
|
|
assigned open issue that has no live execution or durable wait path.
|
|
|
|
The second mechanism produced the repeated `issue_continuation_needed` runs on
|
|
DOT-2.
|
|
|
|
## When it runs
|
|
|
|
The heartbeat scheduler is enabled unless
|
|
`HEARTBEAT_SCHEDULER_ENABLED=false`. Its interval is configured with
|
|
`HEARTBEAT_SCHEDULER_INTERVAL_MS` and defaults to 30,000 ms. The configured
|
|
value is clamped to a minimum of 10,000 ms.
|
|
|
|
Continuation recovery runs:
|
|
|
|
- once during server startup, after orphaned runs are reaped, due retries are
|
|
promoted, and already-queued work is resumed; and
|
|
- on every heartbeat scheduler tick, after the same orphan/retry/queue cleanup.
|
|
|
|
The periodic tick calls `reconcileStrandedAssignedIssues()`. Consequently, a
|
|
new recovery run normally appears within one scheduler interval. The interval
|
|
is polling cadence, not a promise that every continuation waits exactly that
|
|
long.
|
|
|
|
Explicit native continuation effects do not need to wait for this scan. The
|
|
native status-decision committer writes an idempotent `agent_wakeup_requests`
|
|
row directly. The normal queued-run machinery then claims it.
|
|
|
|
## What is durable
|
|
|
|
The system reconstructs intent from persisted control-plane records rather
|
|
than an in-memory timer owned by an agent:
|
|
|
|
- the issue status and assignee;
|
|
- the issue execution lock/run identity;
|
|
- heartbeat run status, context, retry ancestry, and terminal timestamps;
|
|
- `agent_wakeup_requests` rows;
|
|
- native status decisions and their materialized effects;
|
|
- pending interactions, approvals, monitors, blockers, and execution stages;
|
|
- scheduled retry timestamps and recovery actions.
|
|
|
|
Because these records survive a process restart, startup reconciliation can
|
|
resume queued work or repair an issue whose previous execution disappeared.
|
|
|
|
## Explicit continuations
|
|
|
|
A native result can report `yielded` with a continuation containing:
|
|
|
|
- a kind: `same_agent`, `retry`, `delegated_issue`, `response_wake`, or
|
|
`monitor`;
|
|
- a summary; and
|
|
- an idempotency key.
|
|
|
|
The native status arbiter keeps the issue `in_progress` and emits an
|
|
`enqueue_continuation` effect. The status-decision committer materializes that
|
|
effect as an idempotent wake request. A monitor continuation uses
|
|
`monitor_due`; other continuation kinds use `issue_status_changed` at this
|
|
boundary.
|
|
|
|
This is intentional continuation: the result explicitly says more work is
|
|
needed and leaves a persisted execution path.
|
|
|
|
## Stranded-issue reconciliation
|
|
|
|
`reconcileStrandedAssignedIssues()` scans agent-owned issues in `todo`,
|
|
`in_progress`, and relevant `in_review` states. Before creating work, it checks
|
|
for reasons not to wake the agent, including:
|
|
|
|
- an active or queued execution path already exists;
|
|
- a pending interaction or durable wait path exists;
|
|
- the issue or its tree is paused;
|
|
- a provider-quota monitor is pending;
|
|
- the run was explicitly cancelled by an operator;
|
|
- the assigned agent is not invokable;
|
|
- invocation budget or retry/backoff policy prevents a run; or
|
|
- recovery has failed enough times that the issue should be escalated instead.
|
|
|
|
For an eligible `in_progress` issue with no live path, it creates an automated
|
|
wake/run with:
|
|
|
|
```text
|
|
wakeReason: issue_continuation_needed
|
|
retryReason: issue_continuation_needed
|
|
source: issue.continuation_recovery
|
|
```
|
|
|
|
This is a liveness repair, not evidence that the previous model requested
|
|
another turn. Its invariant is: an assigned open issue should have either a
|
|
live execution path, a durable reason to wait, or a visible terminal/blocking
|
|
disposition.
|
|
|
|
## Why DOT-2 looped
|
|
|
|
DOT-2's runner repeatedly produced a successful result that reported `done`.
|
|
The native evidence classifier did not accept the model's evidence references,
|
|
so status arbitration preserved `in_progress` and created another continuation
|
|
path. After each run exited, the periodic reconciler saw:
|
|
|
|
```text
|
|
assigned + in_progress + successful terminal run + no live/wait path
|
|
```
|
|
|
|
It therefore queued `issue_continuation_needed` again on the next approximately
|
|
30-second tick. The new run received the original task title because the native
|
|
model envelope also omitted the wake-comment text, so it repeated the same
|
|
answer.
|
|
|
|
The fix has two parts:
|
|
|
|
- native model input now contains the redacted server-authored task prompt,
|
|
including the current wake comment; and
|
|
- ordinary low-risk issue completion can accept a schema-valid completion
|
|
claim, while tools, governed effects, interactions, and approvals keep their
|
|
independent authorization gates.
|
|
|
|
Thus a completed one-shot task becomes `done`; it is no longer a candidate for
|
|
stranded-issue reconciliation.
|
|
|
|
## How to diagnose a suspected loop
|
|
|
|
Inspect the latest runs for the issue and compare these fields:
|
|
|
|
- `invocationSource` — recovery-created runs use `automation`;
|
|
- `contextSnapshot.wakeReason`;
|
|
- `contextSnapshot.retryReason`;
|
|
- `retryOfRunId`;
|
|
- terminal run status and `resultJson.authoritativeDecision`;
|
|
- `resultJson.issueStatusAfter`;
|
|
- issue `status`, `executionRunId`, and pending interactions/monitors;
|
|
- lifecycle events explaining enqueue, suppression, escalation, or cleanup.
|
|
|
|
Repeated successful runs with `wakeReason: issue_continuation_needed`, an
|
|
authoritative decision of `in_progress`, and no durable wait path usually mean
|
|
the result/disposition contract is failing to converge. Fix the status or
|
|
continuation decision; increasing the polling interval only hides the bug.
|
|
|
|
## Primary implementation locations
|
|
|
|
- `server/src/index.ts` — startup recovery and periodic heartbeat scheduler.
|
|
- `server/src/config.ts` — scheduler enable flag and interval.
|
|
- `server/src/services/recovery/service.ts` —
|
|
`reconcileStrandedAssignedIssues()` and recovery wake creation.
|
|
- `server/src/services/issue-rewake-throttle.ts` — repeated state-wake
|
|
throttling and progress detection.
|
|
- `server/src/services/native-runtime/status-arbiter.ts` — native terminal
|
|
disposition and explicit continuation decisions.
|
|
- `server/src/services/native-runtime/status-decision-committer.ts` — durable,
|
|
idempotent materialization of native continuation effects.
|
|
- `server/src/services/heartbeat.ts` — run lifecycle, immediate recovery, queue
|
|
promotion, and wake execution.
|