Files
PaperClipAI/tests/runner-e2e/connection-guidance.test.ts
DottaandPaperclip b5342febe5 fix(runner): require current-turn completion after connection continuations (#15514)
## Thinking Path

> - Paperclip manages agents, tasks, permissions, and execution budgets.
> - Native tasks can resume after a connection decision in the same
provider conversation.
> - Each turn still needs an accepted completion report.
> - Compact continuation messages did not explain that reports from
earlier turns cannot finish the new turn.
> - The connection evaluator could also grade before the final reply was
stored or reject valid unavailable-access wording.
> - This PR clarifies the current-turn report requirement and fixes
those observation boundaries.
> - The original connection instructions and strict native completion
gate stay in place.

## Linked Issues or Issue Description

Refs #15489. The reduction remains draft while this separate repair is
qualified. Refs #15471 for the earlier connection continuation work.

## What Changed

- Add a current-turn completion reminder to compact continuation inputs.
- Keep final prose insufficient for completion. Preserve permissions and
retry policy.
- Wait for the final successful task run's saved, attributed decline
reply within the existing deadline.
- Use one bounded explanation matcher for both decline checks.
- Wait for a recorded tool-action rejection to dispatch its bound
continuation, with strict company, task, agent and source-run checks.
- Retain the exact grading input before later API refreshes.
- Add failure and delay regressions and update the Runner and evaluator
docs.

## Verification

The fresh comparison has **15/15 original passes on each variant**: 15
unchanged pass pairs, zero new failures, and no pending pair. There are
30 case attempts and **65 actual agent runs** (baseline 33; candidate
32). All runs succeeded. All 30 cleanups passed. No model attempt was
retried.

| Profile | Baseline | Candidate |
| --- | --- | --- |
| Native Codex `gpt-5.6-sol` | 5/5 | 5/5 |
| ACPX Claude `claude-sonnet-5` | 5/5 | 5/5 |
| OpenCode `openrouter/deepseek/deepseek-v4-flash-0731` | 5/5 | 5/5 |

Each profile covers service approval, service decline, connection
decline, provider decline, and selection of the second provider. The
saved replies, approved briefings, decisions, fixture observations,
final task states, and native completion records were inspected. All 18
saved decline-grade snapshots match their original captured inputs and
checks. Result, API snapshot, and final ledger run sets agree.

- [Candidate
campaign](https://github.com/paperclipai/paperclip/actions/runs/37711658378)
· [public candidate
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711658378-1/index.html)
- [Baseline
campaign](https://github.com/paperclipai/paperclip/actions/runs/37711675579)
· [public baseline
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711675579-1/index.html)
- Candidate source: `79905343bba280d462765faad19a26e7f179259e`. Baseline
source: `7c5e120158f1385a1fc5f20be41f66c58a605534`. Both use master
context `fc6304dfe5f2e446e09bd052a7b45f51e930f250`. The trusted workflow
source is separately frozen at
`dd777f4b7343305c4e6f44c422f44a1d78e12e4f`.
- Both variants use the same evaluator, fixtures, models, permissions,
720-second cell deadline, and 1,000-cent company and agent hard stops.
The only production difference is the compact continuation reminder.
- Suite hash:
`aee30b74b4d38ada08777798db0932fbc64e368bb427d29caa8cfd86c7f59747`.
Definition hash:
`ea9e17f54fe0af1acbb2337ddfab8a7e92db8488bd0f59510a63095ca0229b60`.
- Provider-free transport capture: startup/resume input stays at 55,726
bytes. Compact continuation input grows from 53,334 to 53,535 bytes. The
42-tool catalog stays unchanged. These are Paperclip input bytes, not
complete vendor prompt tokens.
- Focused evaluator tests: 31 pass. Native contract, transport delivery,
and session tests: 194 pass. Evaluator support: 1,819 TypeScript tests
and 128 Node tests pass, with one intentional skip.
- Full build, workspace typecheck, and evaluator typecheck pass.
Current-head CI passes all required gates. The current rollup has 51
successful check runs, two intentional Storybook skips, and a successful
Snyk status. Review is 5/5 with zero unresolved threads.
- CI attempt 1 had one initial runtime-fixture health timeout. The exact
test and its full 164-test file pass locally. One CI shard retry passed.
The original CI failure, its dependent verify failure, and the retry
remain visible in [CI
history](https://github.com/paperclipai/paperclip/actions/runs/37711235060).
- The broad local `pnpm test:run` attempt was interrupted after about 49
minutes (exit 130). It recorded one failure in the unchanged Zep
memory-connector disabled-setting test. That test and the full 388-test
tool-access file pass in separate local checks; current-head CI also
passes. The local cause is not established, and this broad local attempt
is **not** claimed as passing. Three earlier local failures also pass in
their isolated checks; their original logs remain retained.
- The first two setup admissions were cancelled before provider jobs to
include the review correction. They made no provider calls. The
completed campaigns above are the first and only model attempts for
these corrected variants.

Cost evidence stays separate from behavior. Original result summaries
report only OpenCode amounts: baseline $0.039192096 and candidate
$0.063221620. Final run ledgers also retain estimates for Claude
(baseline $1.232439000; candidate $1.333611200) and Codex (baseline
$2.340324400; candidate $2.129335600). These estimates do not replace
the original summaries. Local and GitHub runtime are unmetered here.
Invoices are unknown. This is not a cheaper or faster claim.

## Risks

- A single matched trial cannot prove general equivalence or causation.
The reminder is an instruction change, not a new completion enforcement
rule.
- Candidate OpenCode service-decline finished within one continuous run;
its baseline used two. That pair passed the task outcome, but it does
not qualify the reminder on a resumed decline turn. No extra paid run
was used to replace it.
- The text matcher is bounded evidence of an explanation. It does not
prove reasoning or consumption of feedback. Bounded stdout excerpts do
not prove that every extra attempted tool call is absent.
- Missing saved replies or continuations still fail at the original
deadline. Failed native completion remains a failure even when final
prose is correct.
- The unresolved local-suite discrepancy above remains a validation
limit. Full remote CI and both focused local reproductions pass.
- The original connection instructions stay in place. These results do
not qualify the reduction in #15489. Its original 11/15 versus 12/15
grades and two new failing pairs remain unchanged.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, test inspection and eval analysis.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run the focused local tests listed above and they pass; the
interrupted broad local run is disclosed above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 04:52:12 -05:00

166 lines
11 KiB
TypeScript
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
import { describe, expect, it, vi } from "vitest";
import { runnerMatrix, runnerSuites } from "./catalog.js";
import { everydayTasks } from "./everyday-cases.js";
import { CONNECTION_GUIDANCE_SUITE, connectionGuidanceTasks, connectionGuidanceDefinitionDigest } from "./connection-guidance-cases.js";
import { gradeConnectionGuidanceDecline, hasConnectionGuidanceDeclineReply } from "./connection-guidance-evidence.js";
import { pollUntil } from "./api.js";
import { parseRunnerSelectors, selectRunnerExecutions } from "./selectors.js";
describe("neutral connection guidance selection", () => {
it("requires explicit selection and admits five stories on exactly three local native profiles", () => {
const selected = selectRunnerExecutions(parseRunnerSelectors(["--suite", CONNECTION_GUIDANCE_SUITE]));
expect(selected).toHaveLength(15);
expect(new Set(selected.map(row => row.profile.id))).toEqual(new Set(["runner-codex", "runner-acpx-claude", "runner-opencode"]));
for (const row of selected) {
expect(row.environment.id).toBe("local");
expect(row.task.automaticRetryPolicy).toBe("single_attempt");
expect(row.task.attemptTimeoutMs.local).toBe(720_000);
expect(row.task.expectedRunCount).toBe(row.task.id === "provider-second" ? 3 : 2);
}
for (const args of [["--all"], ["--profile", "runner-opencode"]]) {
expect(selectRunnerExecutions(parseRunnerSelectors(args), runnerMatrix).some(row => row.suite.id === CONNECTION_GUIDANCE_SUITE)).toBe(false);
}
expect(runnerSuites.find(suite => suite.id === CONNECTION_GUIDANCE_SUITE)?.definitionMetadata)
.toMatchObject({ fixtureDigest: connectionGuidanceDefinitionDigest(), companyAndAgentBudgetCents: 1_000, maximumAttemptsPerCell: 1 });
});
it("keeps workflow instructions out of new decline prompts and preserves the historical prompts", () => {
for (const task of connectionGuidanceTasks) {
const original = everydayTasks.find(row => row.id === task.id)!;
if (task.id.endsWith("decline")) {
expect(task.buildPrompt("nonce")).toMatch(/brief explanation is enough/);
expect(task.buildPrompt("nonce")).not.toMatch(/declin|Not now|None for now|do not|retry|try again|connection_request|connections_search|paperclip_|yield|poll/i);
expect(task.buildPrompt("nonce")).not.toBe(original.buildPrompt("nonce"));
} else expect(task.buildPrompt("nonce")).toBe(original.buildPrompt("nonce"));
}
expect(everydayTasks.find(task => task.id === "service-decline")!.buildPrompt("nonce")).toContain("do not try again");
expect(everydayTasks.find(task => task.id === "connection-decline")!.buildPrompt("nonce")).toContain("do not try again");
});
});
const valid = {
caseId: "provider-decline" as const, decisionId: "decision", leadAgentId: "lead", issueId: "task",
decisions: [{ id: "decision", kind: "ask_user_questions", status: "answered", resolvedAt: "2026-10-06T22:00:00Z",
result: { answers: [{ questionId: "connection-provider:hubspot", optionIds: ["none"] }] } }],
replies: [{ authorAgentId: "lead", createdByRunId: "run", createdAt: "2026-10-06T22:00:02Z", body: "Contacts are unavailable after your decision." }],
runs: [{ id: "run", nativeIssueId: "task", agentId: "lead", status: "succeeded", finishedAt: "2026-10-06T22:00:03Z" }],
calls: 0, marker: "PRIVATE_MARKER", sameConnections: true,
};
const passes = (input: Parameters<typeof gradeConnectionGuidanceDecline>[0]) =>
gradeConnectionGuidanceDecline(input).every(check => check.passed);
describe("neutral decline evidence", () => {
it("accepts an attributed explanation after each kind of saved decline", () => {
expect(passes(valid)).toBe(true);
for (const [caseId, kind] of [["service-decline", "request_confirmation"], ["connection-decline", "connection_intent"]] as const) {
expect(passes({ ...valid, caseId, decisions: [{ ...valid.decisions[0], kind, status: "rejected" }] })).toBe(true);
}
});
it("recognizes unavailable-connection contractions without accepting positive or unrelated claims", () => {
for (const body of [
"HubSpot isn't connected because you chose None for now.",
"HubSpot isn’t connected after the provider decision.",
"I wasn't able to pull your recent HubSpot contacts. You chose None for now, so no connection was made.",
"The page service approval was rejected, so the data was unavailable.",
"The pages aren't retrievable after the connection decision.",
]) expect(passes({ ...valid, replies: [{ ...valid.replies[0], body }] })).toBe(true);
for (const body of ["HubSpot is connected.", "HubSpot isn't disconnected.", "HubSpot isn't slow.", "Contacts are available.", "The request was not declined.", "The request was not rejected.", "Contacts are not unavailable.", "I wasn't able to finish my coffee."]) {
expect(passes({ ...valid, replies: [{ ...valid.replies[0], body }] })).toBe(false);
}
});
it("rejects missing, stale, unattributed, wrong-agent, or unsuccessful explanations", () => {
for (const input of [
{ ...valid, replies: [] },
{ ...valid, replies: [{ ...valid.replies[0], body: "Done." }] },
{ ...valid, replies: [{ ...valid.replies[0], body: { text: "Unavailable" } }] },
{ ...valid, replies: [{ ...valid.replies[0], createdAt: "2026-10-06T21:59:00Z" }] },
{ ...valid, replies: [{ ...valid.replies[0], createdByRunId: undefined }] },
{ ...valid, replies: [{ ...valid.replies[0], createdByRunId: null, runId: "run" }] },
{ ...valid, replies: [{ ...valid.replies[0], createdByRunId: "other-run" }] },
{ ...valid, replies: [{ ...valid.replies[0], authorAgentId: "worker" }] },
{ ...valid, replies: [{ ...valid.replies[0], createdAt: "invalid" }] },
{ ...valid, runs: [{ ...valid.runs[0], agentId: "worker" }] },
{ ...valid, runs: [{ ...valid.runs[0], nativeIssueId: "other-task" }] },
{ ...valid, runs: [{ ...valid.runs[0], nativeIssueId: undefined }] },
{ ...valid, runs: [{ ...valid.runs[0], status: "failed" }] },
{ ...valid, runs: [{ ...valid.runs[0], finishedAt: "2026-10-06T21:59:00Z" }] },
]) expect(passes(input)).toBe(false);
});
it("requires output from the final successful task run, not a late comment from an earlier wait", () => {
const runs = [...valid.runs, { ...valid.runs[0], id: "final", finishedAt: "2026-10-06T22:00:06Z" }];
expect(hasConnectionGuidanceDeclineReply({ ...valid, runs })).toBe(false);
expect(passes({ ...valid, runs })).toBe(false);
const input = { ...valid, runs, replies: [{ ...valid.replies[0], createdByRunId: "final", createdAt: "2026-10-06T22:00:07Z" }] };
expect(hasConnectionGuidanceDeclineReply(input)).toBe(true);
expect(passes(input)).toBe(true);
});
it.each(["delayed", "missing", "wrong-run", "wrong-content"])("waits for durable output within the original deadline: %s", async variant => {
vi.useFakeTimers();
try {
const start = Date.now();
const input = structuredClone({ ...valid, replies: [] as typeof valid.replies });
const pending = pollUntil({ label: "final decline reply", deadlineAt: start + 30_000, intervalMs: 1000,
load: async () => {
if (variant !== "missing" && Date.now() >= start + 20_000) input.replies = [{ ...valid.replies[0],
createdByRunId: variant === "wrong-run" ? "other" : "run",
body: variant === "wrong-content" ? "Done." : valid.replies[0].body }];
return input;
},
accept: hasConnectionGuidanceDeclineReply,
}).catch((error: unknown) => error);
await vi.advanceTimersByTimeAsync(30_001);
const result = await pending;
if (variant === "missing" || variant === "wrong-run") expect(result).toBeInstanceOf(Error);
else {
expect(result).toEqual(input);
// Readiness must not wait for favorable wording or silently convert a bad answer into a pass.
expect(passes(input)).toBe(variant === "delayed");
}
} finally { vi.useRealTimers(); }
});
it("rejects wrong or repeated decisions, early use, missing call evidence and connection changes", () => {
for (const input of [
{ ...valid, decisionId: "other" },
{ ...valid, decisions: [] },
{ ...valid, decisions: [...valid.decisions, ...valid.decisions] },
{ ...valid, decisions: [{ ...valid.decisions[0], resolvedAt: undefined }] },
{ ...valid, decisions: [{ ...valid.decisions[0], kind: "request_confirmation" }] },
{ ...valid, decisions: [{ ...valid.decisions[0], status: "pending" }] },
{ ...valid, decisions: [{ ...valid.decisions[0], result: { answers: [{ questionId: "connection-provider:hubspot", optionIds: ["via:arcade:hubspot"] }] } }] },
{ ...valid, calls: 1 },
{ ...valid, calls: undefined },
{ ...valid, sameConnections: false },
{ ...valid, replies: [{ ...valid.replies[0], body: "Unavailable, but PRIVATE_MARKER" }] },
]) expect(passes(input)).toBe(false);
});
});
import { requireProviderAccessCard, gradeProviderOutcome } from "./connection-routing-evidence.js";
const access = { id: "access", kind: "connection_intent", status: "pending", payload: {
serviceSlug: "arcade", requestingAgentId: "lead", upstreamService: { slug: "hubspot", selectionInteractionId: "choice" },
accessRequest: { connectionId: "connection", tools: [{ catalogEntryId: "tool", toolName: "Hubspot_ListContacts", permission: "allowed" }] },
} };
const accessInput = { rows: [{ id: "choice", kind: "ask_user_questions", status: "answered" }, access],
decisionId: "choice", connectionId: "connection", agentId: "lead", catalogEntryIds: ["tool"], calls: 0 };
it("requires the exact scoped access card and rejects early calls, swapped identities and duplicates", () => {
expect(requireProviderAccessCard(accessInput)).toBe(access);
for (const input of [
{ ...accessInput, rows: [] }, { ...accessInput, calls: 1 }, { ...accessInput, catalogEntryIds: [] },
{ ...accessInput, connectionId: "other" }, { ...accessInput, agentId: "other" }, { ...accessInput, decisionId: "other" },
{ ...accessInput, rows: [...accessInput.rows, access] },
{ ...accessInput, rows: [accessInput.rows[0], { ...access, status: "accepted" }] },
{ ...accessInput, rows: [accessInput.rows[0], { ...access, payload: { ...access.payload, upstreamService: { slug: "hubspot", selectionInteractionId: "other" } } }] },
]) expect(() => requireProviderAccessCard(input)).toThrow();
});
it("requires the real accepted access result without allowing repeated provider questions", () => {
const choice = { id: "choice", kind: "ask_user_questions", status: "answered", result: { answers: [{ questionId: "connection-provider:hubspot", optionIds: ["via:arcade:hubspot"] }] } };
const granted = { ...access, status: "accepted", result: { outcome: "connected", connectionId: "connection" } };
const input = { rows: [choice, granted], decisionId: "choice", selected: "via:arcade:hubspot", calls: 1,
response: "Ada Fixture MARKER", marker: "MARKER", sameConnections: true, accessDecision: { id: "access", connectionId: "connection" } };
expect(gradeProviderOutcome(input).every(c => c.passed)).toBe(true);
for (const bad of [{ ...input, rows: [choice] }, { ...input, rows: [choice, granted, choice] },
{ ...input, rows: [choice, { ...granted, status: "pending" }] }, { ...input, accessDecision: { id: "access", connectionId: "other" } },
{ ...input, calls: 0 }, { ...input, calls: 2 }, { ...input, response: "Done" }])
expect(gradeProviderOutcome(bad).every(c => c.passed)).toBe(false);
});