Files
PaperClipAI/tests/runner-e2e/continuation-readiness.ts
DottaandPaperclip 24a15c5afa test(evals): verify durable waiting across continuation checkpoints (#15564)
## Thinking Path

> - Paperclip manages agents and durable tasks across provider runs.
> - Human questions must preserve task ownership and stop work until a
real answer arrives.
> - The continuation eval checks the final work, but its lifecycle
oracle misses several broken waiting states.
> - A correct final answer can hide a stale execution lock or a lost
intermediate answer receipt.
> - This pull request checks each wait and retains every question and
run identity.
> - A source audit records which waiting operations belong to native
runtimes and which still require legacy API calls.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The existing Product E2E continuation oracle for human question and
approval journeys.

**Current behavior**

A pending interaction can pass the waiting check even when its task has
the wrong status, a stale lock, or a retry. Only the first pending
question and first run receipts are checked at the final checkpoint.

**Proposed behavior**

Require a healthy wait on the same task and assignee. Preserve all
intermediate run receipts and every pending question's answered
identity. Allow a paused native provider question only when its pending
runtime request identifies the running native run and execution lock.

Related: #15548 and #15554 cover earlier bookkeeping slices. #15544
changes production continuation summaries; this PR changes the eval
oracle and does not overlap that fix.

## What Changed

- Check waiting state, execution ownership, and all answer/run
identities in the existing lifecycle oracle.
- Add negative calibrations for broken records and positive coverage for
both semantic waits and paused provider questions.
- Retain task activity at each checkpoint for inspection of successful
persisted mutations.
- Verify 1,000-cent company and agent budget limits before continuation
work. Disable automatic cell rerolls.
- Document runtime ownership, existing coverage, and the remaining
instruction decision. Keep production instructions unchanged.

## Verification

- `pnpm test:e2e:runner:unit`: 1,867 Vitest tests pass, one skips; 128
Node tests pass.
- `pnpm test:e2e:runner:typecheck`: passes.
- `pnpm -r typecheck`: passes.
- `pnpm build`: passes.
- Negative calibration: 27 added cases fail against the prior oracle and
pass with these checks.
- Four existing local continuation cells are selected for a separate
bounded live canary. Live results are pending; no behavioral pass is
claimed here.
- The full local `pnpm test:run` suite was not repeated because embedded
PostgreSQL was unavailable in the preceding workspace verification.
Required Linux PR CI must pass before readiness.

## Risks

The stronger oracle can expose existing product or fixture defects. A
paused provider question and a terminal semantic wait have distinct
valid states. The four-cell canary does not qualify approval/review,
dependency unblock, crash races, remote execution, or general task
quality. Task activity records successful persisted writes, not failed
API attempts; repeated progress comments are not automatically defects.
Historical eval grades remain unchanged. No production scheduling,
prompt, tool, schema or migration changes.

## Model Used

OpenAI Codex, GPT-6-based, with tool use and code execution. The exact
serving model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 08:45:36 -05:00

35 lines
2.1 KiB
TypeScript

/** A terminal child does not mean the parent has processed its completion wake. */
export function continuationInitialReady(interactions: ReadonlyArray<{ kind?: unknown; status?: unknown }>): boolean {
return interactions.some((i) => i.kind === "ask_user_questions" && i.status === "pending");
}
/** A successful click can return before the form POST commits. Do not accept
* the original paused run as the result of the answer we just submitted. */
export function continuationAnswerCommitted(interactions: ReadonlyArray<{ id?: unknown; status?: unknown }>, interactionId?: string): boolean {
return !interactionId || interactions.some(i => i.id === interactionId && i.status === "answered");
}
/** Run completion precedes task-lock release. Observe both before taking a
* checkpoint, even when the terminal run status is stable across many polls. */
export function continuationCheckpointReady(input: {
issue: { status: string; executionRunId?: string | null };
runs: ReadonlyArray<{ id: string; status: string; runtimeMode?: string }>;
interactions: ReadonlyArray<{ id?: unknown; kind?: unknown; status?: unknown;
sourceRunId?: unknown; payload?: { runtimeRequestId?: unknown } }>;
}): boolean {
if (!input.runs.length) return false;
const active = input.runs.filter(run => ["running", "queued"].includes(run.status));
const pending = input.interactions.filter(i => i.status === "pending");
if (active.length === 0) {
const waiting = pending.some(i => ["ask_user_questions", "request_confirmation", "request_approval"].includes(String(i.kind)));
return input.issue.executionRunId === null && (!waiting || input.issue.status === "in_review");
}
if (active.length !== 1) return false;
const run = active[0];
return run.status === "running" && run.runtimeMode === "native" &&
input.issue.executionRunId === run.id && ["in_progress", "in_review"].includes(input.issue.status) &&
pending.some(i => i.kind === "ask_user_questions" && typeof i.id === "string" && !!i.id.trim() &&
i.sourceRunId === run.id && typeof i.payload?.runtimeRequestId === "string" && !!i.payload.runtimeRequestId.trim());
}