Files
PaperClipAI/doc/architecture/native-status-arbitration.md
DottaandPaperclip 1778075155 fix(server): continue unfinished tasks after status replies (#14626)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs report their task outcome through a structured finish
result.
> - A Board comment can permit a passive wait for the next response.
> - That exception accepted reports that also admitted blocking
unfinished work.
> - The task then stayed In Progress without a runner, and recovery
treated the wait as healthy.
> - This pull request rejects that contradiction and uses the existing
bounded continuation path.
> - Real wait conditions and protection against obsolete requests remain
in force.

## Linked Issues or Issue Description

**What happened?**

An agent answered a status inquiry with `yielded` and `response_wake`.
The same result listed blocking remaining work. The server accepted an
indefinite wait without a question, approval, dependency, or pause. No
further run was queued.

**Expected behavior**

Unfinished ordinary tasks must continue or have a recorded reason to
wait. A status reply alone must not suspend the work.

**Steps to reproduce**

1. Add a Board status inquiry to an unfinished assigned task.
2. Submit a successful native result with `yielded`, `response_wake`,
and `remainingWork[].blocksCompletion: true`.
3. Leave the task without any real wait condition.
4. Observe that the old policy preserves In Progress with no
continuation and suppresses recovery.

**Paperclip version or commit**

Reproduced in the native status policy at `da887ea3e`. The branch is
based on current master.

**Deployment mode**

Authenticated server with the native Paperclip Runner.

Related work: Refs #13338 (native response waits and recovery). Refs
#12071 (separate legacy recovery and retry-state work). This change
fixes the native unfinished-response-wait exception.

## What Changed

- Reject contradictory finish reports while the provider can still
correct them.
- Route accepted unfinished response waits through the existing
one-follow-up continuation budget. Repeated incomplete results create a
visible recovery action.
- Preserve questions, approvals, dependencies, pauses, conversation
lifecycles, and superseded Board requests.
- Recheck the current Board source in the decision transaction before
queuing repair.
- Let normal recovery reconsider old committed waits that report
blocking work and still have a current source.
- Add policy and database regression tests. Document the rule.

## Verification

- `pnpm exec vitest run
server/src/services/native-runtime/status-arbiter.test.ts
server/src/__tests__/heartbeat-process-recovery.test.ts
--no-file-parallelism`: 346 tests passed, no skips.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `git diff --check origin/master...HEAD`: passed.
- Final head `8f0824d9bb4d631f2347f8afffe5fff6d574cb69`: all CI checks
green (54 passed, 2 intentionally skipped), including all general tests,
serialized server suites, runner tests, eight E2E shards, and the canary
dry run.
- Greptile: 5/5 on the final head, with no unresolved review threads.
- The broad local `pnpm test:run` encountered an unrelated timing
failure in `workspace-runtime.test.ts` (waiting for managed process-tree
listeners). That test passed on an isolated rerun. The duplicate broad
run was stopped; the complete CI suite passed on the final commit.
- The transaction-race regression reproduced an obsolete continuation
before the fix. All 12 source-change cases now pass, covering passive
waits, corrective continuations, and exhausted-repair decisions.

## Risks

- Agents that previously parked unfinished ordinary tasks must now
continue or record an actual wait condition.
- Old contradictory waits become eligible for normal recovery. Existing
ownership, budget, pause, and supersession checks still apply.
- The rule uses the structured blocking-work flag. It does not infer
omitted work from prose.
- No schema, dependency, or UI change.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
editing, and test execution. The session does not expose a more specific
deployment identifier or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 16:51:50 -05:00

17 KiB

Native status arbitration

Native runner results do not directly mutate an issue's status. The model may report that work is done, blocked, ready for review, or yielded, but Paperclip's server remains the authority that decides and commits the resulting workflow state.

The native status pipeline is:

structured runner result
        |
        v
schema and terminal validation
        |
        v
evidence classification
        |
        v
pure status arbitration
        |
        v
transactional decision commit
        |
        +--> issue status/version
        +--> durable side effects
        +--> audit and recovery records

This separation prevents model prose from acting as a privileged status command, protects newer issue state from stale runs, and makes every decision replayable and auditable.

Inputs to finalization

The runner returns a paperclip.run_result.v1 result and a matching terminal record. Important result fields include:

  • reportedWorkDisposition: done, blocked, needs_review, or yielded;
  • completionClaim, including the contract revision, criterion claims, and remaining work;
  • verification claims;
  • evidence references;
  • an optional blocker; and
  • an optional continuation.

The server also owns facts the runner cannot choose:

  • the persisted completion contract;
  • the run's actual terminal state;
  • whether workspace finalization succeeded;
  • the issue's current status and status version;
  • pending approvals, interactions, and execution-policy stages; and
  • the completion-authority policy recorded on the contract.

Evidence classification

classifyNativeEvidence() compares model claims with durable Paperclip records. It recognizes these evidence families:

  • run events with an authoritative control-plane evidence verdict;
  • issue work products;
  • approvals;
  • issue-thread interactions; and
  • attachments.

Each evidence reference becomes one of:

Outcome Meaning
accepted A matching durable record exists and authoritatively supports the claim.
missing The required claim or referenced record does not exist or is still pending.
rejected The record or claim explicitly contradicts completion.
unverifiable A record exists, or a string was supplied, but it is not authoritative evidence.

For example, a model-authored reference such as task-response is not trusted merely because it looks descriptive. Likewise, the model's own run.result.proposed event is a claim, not independent proof.

The classifier produces a NativeEvidenceAssessment containing:

  • contract-revision validity;
  • criterion claim and evidence outcomes;
  • verification claim and evidence outcomes;
  • accepted, missing, rejected, and unverifiable references;
  • blocking remaining work;
  • normalized blocker or continuation data; and
  • pending attention requests.

The classifier does not update the issue.

Completion authority

The persisted completion contract controls how a done claim may be accepted.

Durable-evidence completion

The strongest completion path requires all of the following:

  • the runner reports done;
  • the objective is satisfied;
  • every contract criterion has accepted durable evidence;
  • every verification has accepted durable evidence; and
  • no remaining work blocks completion.

This produces completion_contract_satisfied.

Low-risk claim-policy completion

Ordinary issue completion changes Paperclip workflow state, but it does not by itself authorize deployments, spending, secret access, approvals, or arbitrary API calls. Default native completion contracts therefore use low-risk agent_claim_policy authority.

Under that policy, a result can complete the issue when:

  • it reports done;
  • its contract revision matches;
  • every contract criterion is claimed satisfied;
  • every verification is claimed passed; and
  • no remaining work blocks completion.

This produces completion_claim_policy_accepted. Independently governed tools and effects still enforce their own authorization and approval rules.

Decision order

arbitrateNativeStatus() is a pure function. It evaluates higher-authority conditions before model disposition:

Condition Status decision Important effects/reason
Issue is already done or cancelled Preserve terminal_status_preserved
Workspace finalization failed Preserve Record a retryable finalization error
Run performs a named native completion review Preserve The recorded review decision controls the child; unresolved review records a reviewer recovery action
Run was cancelled Preserve Release run resources
Run failed Preserve Schedule recovery
Approval, interaction, or execution stage is pending in_review Materialize/bind the governance gate and notify its owner
Completion satisfies its authority policy done Release checkout
Runner reports a concrete attention request with a reviewer and decision in_review Bind the requested reviewer
Runner reports needs_review without a decision, or an incomplete completion claim Keep work with the agent No automatic human approval; at most one corrective continuation, then a visible recovery action
Runner reports a task-wide blocker blocked Persist blocker owner and unblock action
Runner reports a current-track blocker in_progress Enqueue another productive track
Runner reports yielded with a valid continuation in_progress Enqueue the declared continuation
Ordinary task reports response_wake and blocking remaining work, without a real wait in_progress Use the bounded incomplete-work continuation, then a visible recovery error if repair was already used
Completion evidence is incomplete and continuation is forbidden Preserve Record a finalization error and named next action
Completion evidence is otherwise incomplete in_progress Enqueue a bounded, idempotent continuation

The output is a NativeStatusDecision containing:

  • the arbiter policy version;
  • statusAction and toStatus;
  • a stable reason code;
  • an optional unblock descriptor; and
  • declarative side effects.

The pure arbiter performs no database or network writes.

A response to a Board comment does not by itself justify parking unfinished work. When a yielded / response_wake result contains remainingWork[].blocksCompletion: true, paperclip_finish rejects the report unless a real wait owns the next action. The agent can correct it in the same turn. Finalization also checks accepted reports and uses the existing native-completion-incomplete one-follow-up budget. Model-selected retry keys cannot reset that budget. Pending interactions, approvals, dependencies, pause holds, and conversation lifecycles retain their authority.

The current Board source is rechecked in the status transaction before a repair is queued. Edited, deleted, or superseded comments do not authorize replay of an old response. Recovery no longer treats an old committed passive wait as healthy when its report admits blocking work and its Board source is still current. Existing recovery gates continue to protect pauses, budgets, and ownership.

Transactional decision commit

commitNativeStatusDecision() applies the decision against authoritative issue state. It uses the issue's prior status, status version, and prior decision ID as compare-and-swap inputs. If another actor changed the issue first, the commit raises NativeStatusRaceError; finalization reloads current state, reassesses, and retries a bounded number of times.

Within the transaction, the committer:

  1. validates that the assessment and issue bindings still match;
  2. records the status decision and reason;
  3. updates the issue status and increments its version when required;
  4. materializes declared effects;
  5. writes an effect ledger with deterministic idempotency keys;
  6. updates the native-finalization coordinator; and
  7. persists audit activity.

After commit, activity publications are emitted. Reconciliation may redeliver a pending effect, but it resumes the recorded ledger decision; it does not ask the model or arbiter to invent a new decision.

Durable side effects

Depending on the decision, materialized effects may include:

  • an idempotent agent wake request;
  • an issue-thread interaction;
  • a reviewer binding or owner notification;
  • a persisted blocker and unblock action;
  • a scheduled retry or recovery action;
  • a delegated child issue;
  • checkout/resource release; or
  • finalization/reconciliation records.

Effect rows bind company, issue, decision, target, ordinal, delivery state, and idempotency key. This is what makes a status transition with follow-up work recoverable after a process crash.

Governance precedence

A model cannot bypass a pending governance gate by reporting done. Before completion is considered, finalization checks:

  • an active execution-policy stage;
  • a pending issue-thread interaction; and
  • a pending or revision-requested approval linked to the issue.

If one exists, the issue goes to in_review and the durable gate remains the path forward.

Failure and recovery behavior

Status finalization is coordinated by native_run_finalizations. Important phases include observation, workspace finalization, assessment, arbitration, commit, and retryable or terminal failure.

Examples:

  • an unsafe workspace archive triggers one automatic export of confined entries, omitting unsafe links; if it remains unsafe, discard the export and settle the saved result with an informational run-log entry and no task warning or repair action;
  • other workspace-finalization failures preserve the claim and record a retryable error rather than falsely completing the issue;
  • a failed provider run preserves partial evidence and schedules recovery;
  • a status-version race causes bounded reassessment against current issue state; and
  • a materialization failure records the failed phase and next retry time rather than silently dropping the side effect.

Policy upgrades

The policy version on an assessment is audit metadata. New runs use the current rules. A version change alone does not reassess an old run, change task status, or ask a person to review completion. New evidence and explicit status changes still use the existing reconciliation paths.

Reconciliation also withdraws pending review cards created solely by the old policy-version check. It restores the previous status only if that exact decision and status version are still current and no other review gate is pending. A later user or agent decision takes precedence. The old assessments and decisions remain in the audit history; cleanup does not accept or reject the agent's work.

Diagnosing an unexpected status

Start with the terminal heartbeat run and inspect:

  1. resultJson.nativeResult.reportedWorkDisposition;
  2. completion-claim contract revision and criterion statuses;
  3. verification statuses and evidence references;
  4. resultJson.assessmentId and decisionId;
  5. resultJson.authoritativeDecision;
  6. resultJson.issueStatusBefore and issueStatusAfter;
  7. finalizationPhase and workspaceFinalizeStatus; and
  8. pending approvals, interactions, execution stages, blockers, or continuations on the issue.

Common patterns:

  • done claim + in_progress decision: evidence/claim policy did not accept completion, or blocking work remained;
  • done claim + in_review: a governance gate took precedence;
  • successful run + unchanged status + retryable finalization: workspace or status-effect commit failed;
  • repeated issue_continuation_needed runs: the issue remains open without a converging completion or durable wait path; and
  • terminal issue unchanged by a stale run: terminal-state preservation or status-version race protection worked as designed.

Primary implementation locations

  • server/src/services/native-runtime/evidence-classifier.ts — validates claims against durable records.
  • server/src/services/native-runtime/status-arbiter.ts — pure policy and decision table.
  • server/src/services/native-runtime/status-decision-committer.ts — CAS commit, effect materialization, ledger, and audit.
  • server/src/services/native-runtime/native-run-finalizer.ts — orchestrates assessment, governance checks, arbitration, bounded race retry, and result projection.
  • server/src/services/native-runtime/completion-contracts.ts — creates and versions default native completion contracts.
  • server/src/services/native-runtime/native-finalization-reconciler.ts — resumes interrupted finalization without re-inventing committed decisions.
  • server/src/services/recovery/service.ts — repairs open issues that finish without a live or durable wait path.

See also durable-continuation-scheduler.md for the scheduler and recovery behavior that follows an in_progress decision.

Explicit completion reviews

Ordinary task completion uses the agent's structured done claim under the contract's low-risk claim policy. Unknown evidence references remain diagnostic information; they do not create human approval requirements. Cancellation, newer task state, unresolved dependencies, and explicit governance still win.

Paperclip no longer creates a generic "Native completion review" because a report is incomplete, verification failed, or the agent says needs_review. A new review interaction requires an explicit attention request naming the reviewer's responsibility and the decision. The card displays that request. Waiting for CI remains agent work, not a human completion approval.

The native runner returns current approval/dependency constraints to the agent when it calls paperclip_finish. An empty needs_review report without an existing gate is rejected with instructions to correct it. The final reply must explain any required user action and link to the relevant task or approval. The tool acknowledges receipt, not a premature status commit: final status is committed only after the provider turn and workspace finalization settle.

On upgrade, bounded cleanup withdraws only unanswered, system-created fallback cards proven by their decision/effect ledger, original prompt/target, empty attention request list, and low-risk claim policy. Explicit or answered reviews and stronger completion policies are preserved. Withdrawal has audit history and retires chat actions. Reconciliation reassesses only the current successful run's result, with the same task status/version and completion contract and no newer execution owner. It applies normal governance and dependency checks and appends a decision; it never marks every affected task done blindly. A persisted withdrawal marker makes restart between cleanup and reassessment retryable.

Agent review handoff

A review request and its next action must be saved together. For a native completion review addressed to an agent, the status transaction saves both the review card and one durable wake for that agent. The wake identifies the card and the status decision. Retrying the transaction does not create another wake.

The reviewer runs on the child task, even when the reviewer's own parent task is blocked on that child. The worker remains the child task's assignee. This review role gives access to the child context, history, and documents, plus resolve_review. It does not give access to ordinary task mutation tools. The usual company, invokability, budget, and workspace gates still apply. Human-only requests and governed actions do not acquire this review role.

This restriction applies to Paperclip tools. It is not a read-only filesystem boundary. The reviewer retains the configured agent and environment permissions, including the ability to run tests and create temporary files. An operator who needs filesystem isolation must configure it in the execution environment. The review role does not raise the provider permission mode or bypass a sandbox.

Admission claims the reviewer run, its wake, and the child execution lock in one transaction. If another run holds the lock, the reviewer stays queued. The provider does not start without this claim.

resolve_review accepts or rejects the one card assigned to the run. Rejection requires specific requested changes. The server checks the reviewer, report, and current task version again before it accepts a decision. It also checks that the acting run is running, holds the execution lock, and names this exact card and decision. Another run for the same agent cannot resolve the card. An old or reassigned review does not grant access to newer work.

The final required acceptance completes the child through the existing review resolution path and makes its blocked dependents eligible to run. Rejection returns the requested changes to the worker. A reviewer must record that decision before calling paperclip_finish. Finishing the review run cannot turn rejected work into Done. If the reviewer cannot act, paperclip_block preserves the task and records a recovery action for the reviewer without an automatic retry loop.

Both Codex and ACPX providers wait for server completion feedback before their terminal tool call succeeds. A rejected completion report stays correctable in the same provider turn. The runner does not latch it as an accepted result.