Files
PaperClipAI/tests/runner-e2e/workflow-timeout.test.ts
DottaandPaperclip d0db8820db fix(connections): repair native baseline and approval continuations (#15420)
## Thinking Path

> - Paperclip manages AI agents and the tools they may use.
> - Connection setup separates provider preference from permission to
use a tool.
> - The first native connection baseline could not exercise its intended
decisions.
> - The browser used mutable task titles, and the provider fixture
already granted access.
> - Native provider-choice instructions also disagreed with the
preferred question format. Schema rejection gave no field guidance.
> - This pull request repairs those test preconditions and native
guidance, then fixes restart/approval defects exposed by the corrected
baseline. It also restores missing OpenCode tool-error evidence.
> - The benefit is an inspectable baseline before any further
instruction reduction.

## Linked Issues or Issue Description

Refs #15407.

The original 15-cell baseline remains 0 PASS / 15 FAIL. Ten cells
stopped on stale titles, two Codex cells had schema denials, two
OpenCode cells used already-granted tools, and one Claude cell returned
no native result. No intended user decisions were submitted. The exact
invalid Codex field and underlying Claude failure cause remain unknown.

[Original
campaign](https://github.com/paperclipai/paperclip/actions/runs/37562577199)
· [Original
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37562577199-1/index.html)

## What Changed

- Match the browser's task route and visible identifier instead of a
title the agent can change.
- Start native provider-choice fixtures with no agent tool access.
Verify the public effective-access records.
- In the positive case, select Arcade, then grant its exact HubSpot tool
through the real access card. Require both saved decisions and exactly
one observed call.
- Return a canonical `providerQuestionSet` for native input and retain
the equivalent legacy `providerQuestion`.
- Keep invalid input rejected. Return bounded schema locations and
required field names without submitted values.
- Preserve Claude's exact session/content identity while allowing
authenticated registered instruction-copy paths to rotate on a new run.
- Reject duplicate approval reports for an existing exact tool-action
card before they create another human review.
- Wait for a recorded service-approval continuation within the existing
deadline; retain missing or failed continuation grades.
- Forward OpenCode tool activity through the runner facade, preserving
bounded errors and execution-part identity without inventing host-call
joins or exposing arguments.
- Preserve original grades, costs, scope limits and diagnoses in the
dated repair report.

## Verification

- Eval typecheck passes. Support suite: 1,801 PASS, one intentional
skip; Node checks: 128 PASS.
- Connection/schema tests: 51 PASS. Real-server public fixture setup:
one PASS with zero providers.
- Browser support regression: five PASS, including renamed and wrong
tasks.
- Focused Rust safe-feedback test: one PASS.
- Repository typecheck and build pass before the latest master replay.
Post-replay connection/shared/real-server fixture checks: 52 PASS; eval
typecheck passes. The browser review fix additionally passes all five
browser checks and seven suite checks.
- The full local repository run was interrupted incomplete after about
45 minutes, with five integration failures retained. All five pass in a
separate targeted invocation (1,250 unrelated tests skipped). No full
local-suite pass or root cause for the initial local failures is
claimed.
- Corrected frozen source `162cc90fdabe7f505b88ae095044531b82784c92`:
**10 PASS / 5 FAIL** across the [passing Codex
canary](https://github.com/paperclipai/paperclip/actions/runs/37575158761)
and [remaining 14
cells](https://github.com/paperclipai/paperclip/actions/runs/37576261807).
The canary passes all 17 checks. Claude's two provider-choice
continuations fail on restart, Claude service approval exposes an early
evaluator rejection, Codex service approval creates a duplicate
approval, and OpenCode provider-second times out after both decisions
with no HubSpot call. No original result is regraded.
- Final ledgers count 31 actual runs: 27 succeeded, two failed, two
cancelled during cleanup. All 15 cleanup/budget checks pass. The late
Claude continuation is absent from its earlier workflow snapshot; it
remains in the result/API/final ledger. Original evidence retains 279
hashes. Recorded LLM subtotal $0.04553787 is incomplete billing, not
actual total cost; local runtime is unmetered.
- New repair regressions reproduce the Claude attach failure, duplicate
approval acceptance and dropped OpenCode tool events before their
respective fixes. Nine Rust attachment checks, 127 ACPX host/adapter
tests, 33 completion/control-plane checks, nine eval deadline tests, 59
OpenCode proxy/driver tests, one Rust tool-error/redaction check, and
TypeScript/Rust composer parity pass. Eval typecheck, repository
typecheck and build pass. Existing support coverage is 1,802 PASS plus
128 Node PASS, one intentional support skip; two additional deadline
tests also pass.
- New-source full CI/review and live canaries are pending. The next
bounded selection is Claude provider-decline, Codex service-approve and
one OpenCode provider-second diagnostic with repaired event evidence. No
broader campaign or instruction-reduction qualification is claimed.
- Initial corrected campaign
[37574251834](https://github.com/paperclipai/paperclip/actions/runs/37574251834)
was cancelled during shared build after review found the breadcrumb
whitespace assumption. Its matrix job has zero steps and no provider
execution. The real adjacent-span browser regression now reproduces the
old failure and passes after the fix.

## Risks

- The corrected baseline remains 10/15. The new restart/approval fixes
require live qualification; OpenCode evidence forwarding does not itself
establish or fix its prior behavioral failure.
- The positive provider case now expects three runs, including separate
access approval. Its new results are distinct from the original invalid
fixture.
- The old Claude missing-result cause and rejected Codex field are
unknown. These repairs do not retroactively explain or erase either
failure.
- Path rotation must preserve prompt, custom instruction, skill/content
identity and protected provider settings; regression checks reject stale
or changed content. No connection authorization, JSON schema, budget,
cleanup, or final-result requirement is relaxed. Historical Everyday
prompts and gateway setup remain unchanged.

## Model Used

OpenAI Codex, GPT-6. The exact deployment variant and context window are
not exposed in this session. Used repository inspection, code editing,
test execution and retained-evidence analysis.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 06:18:51 -05:00

185 lines
9.4 KiB
TypeScript

import { describe, expect, it, vi } from "vitest";
import { pollUntil } from "./api.js";
import { classifyFailure } from "./failure-classifier.js";
import { storyHasDurableServiceContinuation, storyHasDurableAgentReviewContinuation, storyHasStrandedBlockedLeaf, storyReviewContinuationTimeoutDetail, type StoryIssue } from "./everyday-observations.js";
describe("workflow timeout classification", () => {
it("does not classify observed task data as an infrastructure error", async () => {
vi.useFakeTimers();
try {
const pending = pollUntil({
label: "everyday hire-reuse settled",
deadlineAt: Date.now() + 10,
intervalMs: 10,
load: async () => ({
status: "in_progress",
connection: "server unavailable",
secret: "plaintext in an ordinary task description",
}),
accept: () => false,
});
const caught = pending.catch((error: unknown) => error);
await vi.advanceTimersByTimeAsync(11);
const error = await caught;
expect(classifyFailure(error)).toBe("candidate_failure");
expect((error as Error).message).not.toContain(
"ordinary task description",
);
} finally {
vi.useRealTimers();
}
});
it("keeps a failed network read retryable", async () => {
vi.useFakeTimers();
try {
const caught = pollUntil({
label: "task state",
deadlineAt: Date.now() + 10,
intervalMs: 10,
load: async () => {
throw new Error("ECONNRESET");
},
accept: () => false,
}).catch((error: unknown) => error);
await vi.advanceTimersByTimeAsync(11);
expect(classifyFailure(await caught)).toBe("transient_infrastructure");
} finally {
vi.useRealTimers();
}
});
});
describe("review continuation deadline", () => {
function fixture(now: number) {
const issues: StoryIssue[] = [
{ id: "parent", companyId: "company", title: "parent", status: "blocked", blockedTransitionAt: new Date(now - 1000).toISOString() },
{ id: "child", parentId: "parent", companyId: "company", title: "child", status: "done", interactions: [{
id: "review", issueId: "child", kind: "request_confirmation", status: "accepted",
addresseeAgentId: "lead", resolvedByAgentId: "lead", resolvedByRunId: "review-run",
resolvedAt: new Date(now).toISOString(), result: { version: 1, outcome: "accepted" },
payload: { target: { type: "custom", key: "native_completion_review", revisionId: "decision" } },
}] },
];
const runs = [{ id: "review-run", companyId: "company", agentId: "lead", status: "succeeded", finishedAt: new Date(now).toISOString() }];
return { issues, runs };
}
it("uses the review timeout detail only with stranded work and durable review evidence", () => {
const state = fixture(Date.now());
expect(
storyReviewContinuationTimeoutDetail(state.issues, "parent", "lead", state.runs, ["lead"]),
).toBe("task is Blocked without an active continuation after accepted review");
state.issues[1]!.interactions = [];
expect(
storyReviewContinuationTimeoutDetail(state.issues, "parent", "lead", state.runs, ["lead"]),
).toBeUndefined();
const completed = fixture(Date.now());
completed.issues[0]!.status = "done";
expect(
storyReviewContinuationTimeoutDetail(
completed.issues, "parent", "lead", completed.runs, ["lead"],
),
).toBeUndefined();
});
it("permits a delayed parent continuation within the existing deadline", async () => {
vi.useFakeTimers();
try {
const start = Date.now();
const state = fixture(start);
const pending = pollUntil({
label: "review handoff", deadlineAt: start + 30_000, intervalMs: 1000,
load: async () => {
if (Date.now() >= start + 20_000) state.issues[0]!.status = "done";
return state;
},
accept: ({ issues }) => issues.every((issue) => issue.status === "done"),
reject: ({ issues, runs }) => storyHasStrandedBlockedLeaf(issues, "lead") &&
!storyHasDurableAgentReviewContinuation(issues, "parent", "lead", runs)
? "task is Blocked without an active continuation" : undefined,
});
const caught = pending.catch((error: unknown) => error);
await vi.advanceTimersByTimeAsync(20_001);
expect(await caught).toBe(state);
} finally { vi.useRealTimers(); }
});
it("reports a missing wake specifically when the existing deadline expires", async () => {
vi.useFakeTimers();
try {
const state = fixture(Date.now());
const pending = pollUntil({
label: "review handoff", deadlineAt: Date.now() + 30_000, intervalMs: 1000,
load: async () => state, accept: () => false,
timeoutDetail: (last) => last && storyReviewContinuationTimeoutDetail(
last.issues, "parent", "lead", last.runs, ["lead"],
),
});
const caught = pending.catch((error: unknown) => error);
await vi.advanceTimersByTimeAsync(30_001);
const error = await caught;
expect((error as Error).message).toContain("Blocked without an active continuation after accepted review");
expect(classifyFailure(error)).toBe("candidate_failure");
} finally { vi.useRealTimers(); }
});
it("uses the generic timeout when durable review evidence is absent", async () => {
vi.useFakeTimers();
try {
const state = fixture(Date.now());
state.issues[1]!.interactions = [];
const pending = pollUntil({
label: "review handoff", deadlineAt: Date.now() + 30_000, intervalMs: 1000,
load: async () => state, accept: () => false,
timeoutDetail: (last) => last && storyReviewContinuationTimeoutDetail(
last.issues, "parent", "lead", last.runs, ["lead"],
),
});
const caught = pending.catch((error: unknown) => error);
await vi.advanceTimersByTimeAsync(30_001);
const error = await caught;
expect(error).toBeInstanceOf(Error);
expect((error as Error).message).not.toContain("after accepted review");
expect(classifyFailure(error)).toBe("candidate_failure");
} finally { vi.useRealTimers(); }
});
});
describe("approved service continuation deadline", () => {
it.each([true, false])("keeps the original deadline for a delayed dispatch (arrives: %s)", async arrives => {
vi.useFakeTimers();
try {
const start = Date.now();
const issues: StoryIssue[] = [{ id: "task", companyId: "company", title: "Briefing", status: "blocked", assigneeAgentId: "agent", interactions: [{
id: "approval", issueId: "task", sourceRunId: "first", createdByAgentId: "agent", kind: "request_confirmation", status: "accepted", continuationPolicy: "wake_assignee",
resolvedAt: new Date(start - 1000).toISOString(), payload: { toolAction: { version: 1, actionRequestId: "action" } }, result: { outcome: "accepted", toolAction: { status: "executed" } },
}] }];
const runs = [{ id: "first", companyId: "company", agentId: "agent", status: "succeeded", finishedAt: new Date(start).toISOString() }];
const caught = pollUntil({ label: "service continuation", deadlineAt: start + 30_000, intervalMs: 1000,
load: async () => { if (arrives && Date.now() >= start + 20_000) issues[0]!.status = "done"; return { issues, runs }; },
accept: state => state.issues[0]!.status === "done",
reject: state => storyHasStrandedBlockedLeaf(state.issues, "agent") && !storyHasDurableServiceContinuation(state.issues, "task", "agent", state.runs) ? "stranded" : undefined,
timeoutDetail: state => state && storyHasDurableServiceContinuation(state.issues, "task", "agent", state.runs) ? "missing service continuation" : undefined,
}).catch((error: unknown) => error);
await vi.advanceTimersByTimeAsync(30_001);
const result = await caught;
if (arrives) expect(result).toEqual({ issues, runs });
else { expect((result as Error).message).toContain("missing service continuation"); expect(classifyFailure(result)).toBe("candidate_failure"); }
} finally { vi.useRealTimers(); }
});
it("waits for dispatch only until the same response has been consumed", () => {
const issues: StoryIssue[] = [{ id: "task", companyId: "company", title: "Briefing", status: "blocked", assigneeAgentId: "agent", interactions: [{
id: "approval", issueId: "task", sourceRunId: "first", createdByAgentId: "agent", kind: "request_confirmation", status: "accepted", continuationPolicy: "wake_assignee",
resolvedAt: "2026-10-07T05:37:19Z", payload: { toolAction: { version: 1, actionRequestId: "action" } }, result: { outcome: "accepted", toolAction: { status: "executed" } },
}] }];
const runs = [{ id: "first", companyId: "company", agentId: "agent", status: "succeeded", finishedAt: "2026-10-07T05:37:28Z" }];
expect(storyHasDurableServiceContinuation(issues, "task", "agent", runs)).toBe(true);
expect(storyHasDurableServiceContinuation(issues, "task", "other", runs)).toBe(false);
const continued = [...runs, { id: "next", companyId: "company", agentId: "agent", status: "succeeded", runnerProfileJson: { nativeExecutionInput: { interactionResponses: [{ interactionId: "approval" }] } } }];
expect(storyHasDurableServiceContinuation(issues, "task", "agent", continued)).toBe(false);
issues[0]!.interactions![0]!.result = { outcome: "accepted", toolAction: { status: "failed" } };
expect(storyHasDurableServiceContinuation(issues, "task", "agent", runs)).toBe(false);
});
});