mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
## Thinking Path > - Paperclip manages agents and durable tasks across provider runs. > - Human questions must preserve task ownership and stop work until a real answer arrives. > - The continuation eval checks the final work, but its lifecycle oracle misses several broken waiting states. > - A correct final answer can hide a stale execution lock or a lost intermediate answer receipt. > - This pull request checks each wait and retains every question and run identity. > - A source audit records which waiting operations belong to native runtimes and which still require legacy API calls. ## Linked Issues or Issue Description **What existing behavior does this improve?** The existing Product E2E continuation oracle for human question and approval journeys. **Current behavior** A pending interaction can pass the waiting check even when its task has the wrong status, a stale lock, or a retry. Only the first pending question and first run receipts are checked at the final checkpoint. **Proposed behavior** Require a healthy wait on the same task and assignee. Preserve all intermediate run receipts and every pending question's answered identity. Allow a paused native provider question only when its pending runtime request identifies the running native run and execution lock. Related: #15548 and #15554 cover earlier bookkeeping slices. #15544 changes production continuation summaries; this PR changes the eval oracle and does not overlap that fix. ## What Changed - Check waiting state, execution ownership, and all answer/run identities in the existing lifecycle oracle. - Add negative calibrations for broken records and positive coverage for both semantic waits and paused provider questions. - Retain task activity at each checkpoint for inspection of successful persisted mutations. - Verify 1,000-cent company and agent budget limits before continuation work. Disable automatic cell rerolls. - Document runtime ownership, existing coverage, and the remaining instruction decision. Keep production instructions unchanged. ## Verification - `pnpm test:e2e:runner:unit`: 1,867 Vitest tests pass, one skips; 128 Node tests pass. - `pnpm test:e2e:runner:typecheck`: passes. - `pnpm -r typecheck`: passes. - `pnpm build`: passes. - Negative calibration: 27 added cases fail against the prior oracle and pass with these checks. - Four existing local continuation cells are selected for a separate bounded live canary. Live results are pending; no behavioral pass is claimed here. - The full local `pnpm test:run` suite was not repeated because embedded PostgreSQL was unavailable in the preceding workspace verification. Required Linux PR CI must pass before readiness. ## Risks The stronger oracle can expose existing product or fixture defects. A paused provider question and a terminal semantic wait have distinct valid states. The four-cell canary does not qualify approval/review, dependency unblock, crash races, remote execution, or general task quality. Task activity records successful persisted writes, not failed API attempts; repeated progress comments are not automatically defects. Historical eval grades remain unchanged. No production scheduling, prompt, tool, schema or migration changes. ## Model Used OpenAI Codex, GPT-6-based, with tool use and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
60 lines
3.4 KiB
TypeScript
60 lines
3.4 KiB
TypeScript
import { describe, expect, it } from "vitest";
|
|
import { continuationAnswerCommitted, continuationInitialReady, continuationCheckpointReady } from "./continuation-readiness.js";
|
|
import { runnerMatrix } from "./catalog.js";
|
|
|
|
describe("continuation readiness", () => {
|
|
it("waits through the idle gap between child completion and its parent's question", () => {
|
|
const polls = [[], [], [{ kind: "request_confirmation", status: "pending" }],
|
|
[{ kind: "ask_user_questions", status: "answered" }],
|
|
[{ kind: "ask_user_questions", status: "pending" }]];
|
|
expect(polls.map(continuationInitialReady)).toEqual([false, false, false, false, true]);
|
|
});
|
|
it("never instructs a resumed provider to use a hardcoded completion revision", () => {
|
|
for (const execution of runnerMatrix) {
|
|
expect(execution.task.buildPrompt("revision-test")).not.toMatch(/contractRevision\s*:\s*["']1["']/);
|
|
}
|
|
});
|
|
});
|
|
|
|
it("waits for the clicked answer to commit instead of grading the original paused state", () => {
|
|
const card = { id: "submitted", status: "pending" };
|
|
expect(continuationAnswerCommitted([card], card.id)).toBe(false);
|
|
expect(continuationAnswerCommitted([{ ...card, id: "different", status: "answered" }], card.id)).toBe(false);
|
|
expect(continuationAnswerCommitted([{ ...card, status: "cancelled" }], card.id)).toBe(false);
|
|
expect(continuationAnswerCommitted([{ ...card, status: "answered" }], card.id)).toBe(true);
|
|
});
|
|
|
|
|
|
describe("checkpoint finalization boundary", () => {
|
|
const pending = [{ id: "question", kind: "ask_user_questions", status: "pending" }];
|
|
const run = { id: "run", status: "succeeded", runtimeMode: "native" };
|
|
it("waits across repeated terminal polls until review state and lock release are both visible", () => {
|
|
const polls = [
|
|
{ status: "in_progress", executionRunId: "run" },
|
|
{ status: "in_progress", executionRunId: "run" },
|
|
{ status: "in_review", executionRunId: "run" },
|
|
{ status: "in_review", executionRunId: null },
|
|
];
|
|
expect(polls.map(issue => continuationCheckpointReady({ issue, runs: [run], interactions: pending })))
|
|
.toEqual([false, false, false, true]);
|
|
expect(continuationCheckpointReady({ issue: { status: "in_progress", executionRunId: null },
|
|
runs: [run], interactions: pending })).toBe(false);
|
|
});
|
|
it("also waits for final completion to release its execution lock", () => {
|
|
for (const executionRunId of ["run", null]) {
|
|
expect(continuationCheckpointReady({ issue: { status: "done", executionRunId },
|
|
runs: [run], interactions: [{ ...pending[0], status: "answered" }] })).toBe(executionRunId === null);
|
|
}
|
|
});
|
|
it("admits only the bound running native provider question without releasing its lock", () => {
|
|
const input = { issue: { status: "in_progress", executionRunId: "run" },
|
|
runs: [{ ...run, status: "running" }], interactions: [{ ...pending[0], sourceRunId: "run",
|
|
payload: { runtimeRequestId: "provider-request" } }] };
|
|
expect(continuationCheckpointReady(input)).toBe(true);
|
|
expect(continuationCheckpointReady({ ...input, interactions: pending })).toBe(false);
|
|
expect(continuationCheckpointReady({ ...input, issue: { ...input.issue, executionRunId: "other" } })).toBe(false);
|
|
expect(continuationCheckpointReady({ ...input, runs: [{ ...input.runs[0], runtimeMode: "legacy" }] })).toBe(false);
|
|
expect(continuationCheckpointReady({ ...input, runs: [...input.runs, { ...run, id: "extra", status: "queued" }] })).toBe(false);
|
|
});
|
|
});
|