mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents access to external services. > - AgentMail needs both a saved key and an inbox assigned to the agent. > - Chat requests offered a setup link instead of an inline card and could treat a saved key as complete. > - Inbox setup also hid address conflicts behind a generic server error and a separate review step. > - This pull request makes AgentMail a default connection, adds the inline card, reduces setup to two steps, and shows conflicts beside the address. > - Shared native dropdown styles also give every caret a consistent inset. ## Linked Issues or Issue Description **What happened?** AgentMail requests in chat did not show a usable inline connection card. Manual setup required extra screens, ignored saved account keys, and could trap new-address setup in a locked inbox dropdown. Agent selectors omitted the avatar from the selected value. A taken address could produce an HTTP 403 from AgentMail and appear as an internal server error. Native dropdown arrows also touched the right edge of their fields. **Expected behavior** Make AgentMail available as a default connection. Ask for the API key inline, with a direct link to its provider page. Default human access to the company and agent access to the requesting agent. Resume the agent only after an assigned inbox is active. Manual setup should ask for an agent and email address, then finish. Address checks should run as the user types. Taken addresses should show clickable alternatives. A domain dropdown beside the name should prefer a verified custom domain. Setup should suggest authorized saved AgentMail keys and show agent avatars in the picker and selected value. **Steps to reproduce** 1. Ask an agent to connect AgentMail when it has no assigned inbox. 2. Check that an inline API-key card appears and links to the provider's API-key page. 3. Open AgentMail setup, choose an agent, and request an address that is already taken. 4. Correct the inline error, refresh, and finish setup with the same request ID. 5. Inspect native dropdown carets in light, dark, disabled, and right-to-left states. Uses the bounded provider-error parser merged in #14768. Related work: #13256 introduced AgentMail; #14725 expanded connection search. ## What Changed - Stop recurring email queries for tasks that have no email thread. Share the query between the thread provider and activity view. Keep email-task updates and invalidation-based discovery. - Make AgentMail available without the experimental chat setting. Keep the catalog, setup and management routes, agent Channels tab, task email feed, receiving worker, and agent tools available by default. Other experimental chat providers stay gated. - Make the email address and copy icon a single clickable action with the shared Copied! confirmation. Add View inbox linking directly to the matching AgentMail console inbox, with the address encoded as one URL path segment. - Reorganize inbox Settings around the copyable email address, usage instructions, and receiving status. Move reconnect credentials into a disclosure and separate the Disconnect action. Add production Settings stories for active, paused, unassigned-address, revoked, webhook, long-address, mobile, and reconnect states. Show repair controls when the inbox has an error. Keep usage instructions tied to an active inbox with an address. - Add AgentMail channel intents and an inline key field with the direct API-key URL. - Keep setup and retry state tied to the interaction. Require an active inbox for completion. Preserve company and agent access checks. - Reduce manual setup to agent selection and email selection. Put the domain dropdown beside the address and default to a verified custom domain. Preserve explicit choices across reloads. Keep receiving settings under Advanced options. - Check the initial address and edits after a 350 ms pause. Abort superseded requests and ignore stale responses. Show clickable suggestions and retain known creation conflicts across reloads. - Add a company-scoped, manager-only address check using the saved credential. Search the visible inbox list instead of fetching an uncreated inbox: live AgentMail retains negative lookups that can break subsequent access-key creation. Unlisted addresses remain unknown; creation is authoritative. - Suggest labeled saved AgentMail keys in both manual setup and the inline card. Filter by company, provider, active credential, and current-user grants on the server. Prefer an account key and preserve the selected key or an explicit new-key choice across refresh. Use verified scope metadata and bounded concurrent checks for legacy keys. Never return secret values. - Catch an inbox-only key before the email step. Allow its existing inbox only after an explicit choice. Recover old locked drafts at the key picker. Save the replacement key before retiring an empty draft, then use a new setup URL so refresh preserves the switched account; stop if cleanup fails. Preserve already allocated addresses and their original accounts. - Use the shared AgentSelect in email setup. Show the canonical agent avatar in each option and the selected value, including other consumers of the shared component. Add regression coverage for legacy and current Lucide agent-mention icon formats. - Start each catalog Add connection with a fresh setup identity. Honor Finish setup's exact draft/account/address instead of resuming an unrelated browser draft. Return Cancel and Done to Connectors and Email settings to the inbox. Group the task/thread explanation in a How it Works card. - Route AgentMail catalog removal through the email inbox control API, including unfinished drafts. Refresh both the catalog and inbox views. - Render each inbox management tab separately. Access uses the saved account grants and agent controls; Conversations and Activity use the shared persisted email feed. Activity lifecycle actions use the email API. Reconnect returns to inbox Settings. Conversation failures show a retry instead of a false empty state. Email delivery recovery stays in the task. - Map documented provider address conflicts to a field error. Preserve actionable messages for other failures. - Preserve non-secret draft fields across refresh, scoped to the requested agent. Never save API keys in browser storage. Resume partial inbox creation with the original agent, address, and request ID. - Show an already-created address with explicit retry and new-address recovery instead of locked inputs. Preserve the original inbox and resumable draft when choosing another address. Distinguish runtime-key 404 errors and log safe provider status/operation/code. - Apply final agent access once within email setup authorization for a new account whose original installs are unchanged. Preserve later permission edits and reused account installs. Support in-place retry of progress loading. - Let a failed inline setup change keys after retiring an empty draft. Persist its replacement setup identity without storing secrets. Recover a server-saved account when refresh interrupts the save response, while preserving intentional account changes. - Render the production setup in Storybook and add error, recovery, and mobile states. - Inset native select carets in shared CSS. Preserve custom icons, listboxes, keyboard behavior, and forced-color controls. - Add browser regression coverage and an AgentMail Product E2E case with persisted-state and rendered-card evidence. ## Verification - Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `git diff --check` passed after the default-availability change. - All 485 focused tests passed. These cover setup, management, catalog and route gates, connection intents, email authorization, Cursor execution, and the OpenAPI contract. All 39 email integration tests run with the experimental chat setting off. - The shared polling change passed four behavioral tests, UI typecheck and build, and token gates. - `tests/e2e/agentmail.spec.ts` passed with the actual server setting off. This full-stack browser test uses simulated provider responses. It covers catalog entry, saved keys, editable address and domain controls, creation, conflicts, retry, all management tabs, clipboard feedback, the provider link, and task email rendering. - In the live local browser, Add connection reached the editable email step with the saved account key. The verified custom domain was selected by default. Both domain choices worked. The existing inbox Settings page remained available. Both active inboxes completed new mail checks with the setting off. No new provider inbox or email message was created for this pass. - Earlier live provider acceptance covered creation on a verified custom domain, Finish connecting on the reported draft, successful mail checks after refresh, and catalog removal of disposable draft and active connections. Clicking the email address copied the exact address and showed Copied!. View inbox opened the same inbox in AgentMail’s console. No email messages were sent. - Production setup and Settings Storybook builds and interactions passed. Settings states include active, paused, unassigned, revoked, webhook, long-address, mobile, and reconnect. Receiving and revoked-access stories had zero accessibility violations. - Full local `pnpm test:run` on an earlier revision completed with 14,709 passing, 87 skipped, and four transient failures. All four failed cases passed in focused reruns without product changes. That serial full local command was not repeated after each follow-up. The latest-head full CI suite is the final test gate. - CI found an obsolete browser assertion that hid every channel when the flag was off. Updated it to keep AgentMail and the Channels surface visible while preserving the GitHub chat route gates. All 11 provider browser tests passed locally after scoping the Channels selector to the agent sidebar. Two initial local attempts stopped at temporary Postgres initialization. The passing run used a separate disposable database on the existing local Postgres server; it was removed after the test. - Updated the remaining sidebar and aggregator discovery assertions for default AgentMail availability. Ordinary task fixtures now return no email thread. All 128 sidebar/task-page tests and all 42 aggregator tests passed locally. - Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI passed, with 54 successful checks including Snyk and two intentional Storybook skips. The CI run is https://github.com/paperclipai/paperclip/actions/runs/37020833647. A fresh Greptile review scored 5/5 with no unresolved threads. Live model evaluations and inbound/outbound email delivery were not run. ## Risks - AgentMail no longer needs experimental opt-in. Setup still requires a human to connect an account and assign an inbox. Inline setup creates an inbox after a human submits a new or saved key. Company access, agent access, inbox assignment, and completion checks remain enforced. - AgentMail read APIs cannot prove global address availability. The visible-list check is bounded to 100 entries and cannot see inboxes outside the key’s scope. The UI reports this limitation, suggests alternatives without claiming they are free, and keeps final creation conflicts inline. Lookup outages show an error without preventing the authoritative creation attempt. - Native select CSS affects the whole app. Custom-icon selects and multi-row lists are excluded. Forced-color mode keeps the browser caret. - Saved-key discovery uses stored verified scope metadata and checks authorized legacy credentials concurrently within a shared three-second deadline. Provider outages mark legacy choices unavailable; users can still enter another key. Final use rechecks authorization and provider access. - No database migration or transport default change. Live connection remains the default. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The exact served model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
810 lines
26 KiB
TypeScript
810 lines
26 KiB
TypeScript
import { describe, it, expect } from "vitest";
|
|
import { runnerMatrix, runnerSuites } from "./catalog.js";
|
|
import { parseRunnerSelectors, selectRunnerExecutions } from "./selectors.js";
|
|
import { everydayTasks, productionStoryProfile } from "./everyday-cases.js";
|
|
import {
|
|
isStoryWorkspaceDeferral,
|
|
storyUnexpectedRunFailure,
|
|
storyUnexercisedReviewBoundary,
|
|
storyLifecycleChecks,
|
|
storyRepliesConsumed,
|
|
storyHasAgentReply,
|
|
storyHasPendingAgentReview,
|
|
storyHasPendingAgentWake,
|
|
storyHasPendingAgentWork,
|
|
storyHasPendingHumanInteraction,
|
|
storyHasStrandedBlockedLeaf,
|
|
storyHasDurableAgentReviewContinuation,
|
|
storyIssueHasBlockedTimeline,
|
|
storyIssueHasBlockedTimelineBefore,
|
|
storyIssueHasUnresolvedDependency,
|
|
storyRunReportsDependencyBlock,
|
|
storyParentFinishedAfterChildren,
|
|
artifactGradeModeForPhase,
|
|
storyParentCompletionPrecedesReview,
|
|
storyAcceptedAgentReview,
|
|
type StoryRun,
|
|
type StoryIssue,
|
|
} from "./everyday-observations.js";
|
|
import { hasPersistedSource, isSavedSourceCheckpoint } from "./everyday-interruption.js";
|
|
|
|
describe("everyday workflow grader and review timing", () => {
|
|
it("requires nonempty persisted source evidence", () => {
|
|
expect(hasPersistedSource(undefined)).toBe(false);
|
|
expect(hasPersistedSource(Buffer.alloc(0))).toBe(false);
|
|
expect(hasPersistedSource(Buffer.from("source"))).toBe(true);
|
|
});
|
|
|
|
it("requires an active run and saved source at the same Stop checkpoint", () => {
|
|
expect(isSavedSourceCheckpoint(false, Buffer.from("source"))).toBe(false);
|
|
expect(isSavedSourceCheckpoint(true, Buffer.alloc(0))).toBe(false);
|
|
expect(isSavedSourceCheckpoint(true, Buffer.from("source"))).toBe(true);
|
|
});
|
|
it("uses base grading for review delivery and preserves late max-length cases", () => {
|
|
expect(artifactGradeModeForPhase("reviewed-delivery")).toBe("base");
|
|
expect(artifactGradeModeForPhase("delegated-delivery")).toBe("max-length");
|
|
expect(artifactGradeModeForPhase("recovered-delivery")).toBe("max-length");
|
|
});
|
|
|
|
it("compares parent completion with review start rather than card creation", () => {
|
|
const interactionCreatedAt = "2026-09-17T10:00:00.000Z";
|
|
const parentFinishedAt = "2026-09-17T10:05:00.000Z";
|
|
const reviewStartedAt = "2026-09-17T10:06:00.000Z";
|
|
expect(interactionCreatedAt < parentFinishedAt).toBe(true);
|
|
expect(
|
|
storyParentCompletionPrecedesReview(
|
|
parentFinishedAt,
|
|
reviewStartedAt,
|
|
false,
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
storyParentCompletionPrecedesReview(
|
|
"2026-09-17T10:07:00.000Z",
|
|
reviewStartedAt,
|
|
false,
|
|
),
|
|
).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe("unexercised review boundary diagnostics", () => {
|
|
const done = { id: "parent", companyId: "company", title: "task", status: "done" };
|
|
const completed = {
|
|
id: "run", companyId: "company", agentId: "lead", status: "succeeded",
|
|
nativeIssueId: "parent", finishedAt: "2026-09-19T14:47:48.134Z",
|
|
};
|
|
const child = {
|
|
...done, id: "child", parentId: "parent", interactions: [{
|
|
id: "card", issueId: "child", kind: "request_confirmation", status: "accepted",
|
|
addresseeAgentId: "lead", resolvedByAgentId: "lead", resolvedByRunId: "review",
|
|
createdAt: "2026-09-19T14:47:04.000Z", resolvedAt: "2026-09-19T14:47:21.876Z",
|
|
payload: { target: { type: "custom", key: "native_completion_review", revisionId: "decision" } },
|
|
result: { version: 1, outcome: "accepted" },
|
|
}],
|
|
};
|
|
const review = {
|
|
...completed, id: "review", nativeIssueId: "child", startedAt: "2026-09-19T14:47:05.802Z",
|
|
contextSnapshot: { nativeReviewInteractionId: "card", nativeReviewDecisionId: "decision" },
|
|
};
|
|
const diagnose = (issues: StoryIssue[] = [done, child], runs: StoryRun[] = [completed, review]) =>
|
|
storyUnexercisedReviewBoundary(issues, runs, "parent", "lead");
|
|
it("fails promptly with persisted proof that review preceded parent completion", () => {
|
|
expect(diagnose()).toContain("not exercised");
|
|
});
|
|
it("does not preempt in-flight finalization, recovery, or unfinished work", () => {
|
|
expect(diagnose([done, child], [{ ...completed, status: "running" }, review])).toBeUndefined();
|
|
expect(storyUnexercisedReviewBoundary([{ ...done, scheduledRetry: {} }, child], [completed, review], "parent", "lead")).toBeUndefined();
|
|
expect(storyUnexercisedReviewBoundary([{ ...done, activeRecoveryAction: {} }, child], [completed, review], "parent", "lead")).toBeUndefined();
|
|
expect(diagnose([{ ...done, status: "blocked" }, child])).toBeUndefined();
|
|
expect(diagnose([], [])).toBeUndefined();
|
|
});
|
|
it("waits when separately fetched snapshots lack review evidence or timestamps", () => {
|
|
expect(diagnose([done, { ...child, interactions: [] }])).toBeUndefined();
|
|
expect(diagnose([done, child], [completed])).toBeUndefined();
|
|
expect(diagnose([done, child], [{ ...completed, finishedAt: "" }, review])).toBeUndefined();
|
|
expect(diagnose([done, child], [completed, { ...review, startedAt: "" }])).toBeUndefined();
|
|
expect(diagnose([done, child], [review])).toBeUndefined();
|
|
});
|
|
it("does not reject a completed workflow with the required timing", () => {
|
|
expect(diagnose([done, child], [{ ...completed, finishedAt: review.startedAt }, review])).toBeUndefined();
|
|
});
|
|
});
|
|
|
|
describe("multi-round agent review handoff", () => {
|
|
const child = {
|
|
id: "child",
|
|
companyId: "company",
|
|
title: "child",
|
|
status: "done",
|
|
interactions: [
|
|
{
|
|
id: "first",
|
|
issueId: "child",
|
|
kind: "request_confirmation",
|
|
status: "rejected",
|
|
addresseeAgentId: "lead",
|
|
createdAt: "2026-09-17T19:07:19.277Z",
|
|
payload: { target: { type: "custom", key: "native_completion_review", revisionId: "decision-1" } },
|
|
result: { version: 1, outcome: "rejected" },
|
|
},
|
|
{
|
|
id: "second",
|
|
issueId: "child",
|
|
kind: "request_confirmation",
|
|
status: "accepted",
|
|
addresseeAgentId: "lead",
|
|
resolvedByAgentId: "lead",
|
|
resolvedByRunId: "review-run",
|
|
createdAt: "2026-09-17T19:08:54.510Z",
|
|
resolvedAt: "2026-09-17T19:09:33.722Z",
|
|
payload: { target: { type: "custom", key: "native_completion_review", revisionId: "decision-2" } },
|
|
result: {
|
|
version: 1,
|
|
outcome: "accepted",
|
|
},
|
|
},
|
|
],
|
|
};
|
|
const runs = [{
|
|
id: "review-run",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "succeeded",
|
|
contextSnapshot: {
|
|
nativeReviewInteractionId: "second",
|
|
nativeReviewDecisionId: "decision-2",
|
|
},
|
|
}];
|
|
|
|
it("follows a later accepted repair review for the same child and lead", () => {
|
|
expect(storyAcceptedAgentReview(child, "first", "lead", runs)?.id).toBe("second");
|
|
});
|
|
|
|
it("accepts a single first review when it is already accepted", () => {
|
|
expect(
|
|
storyAcceptedAgentReview(
|
|
{ ...child, interactions: [child.interactions[1]] },
|
|
"second",
|
|
"lead",
|
|
runs,
|
|
)?.id,
|
|
).toBe("second");
|
|
});
|
|
|
|
it("does not pass with only a rejected review", () => {
|
|
expect(
|
|
storyAcceptedAgentReview(
|
|
{ ...child, interactions: [child.interactions[0]] },
|
|
"first",
|
|
"lead",
|
|
runs,
|
|
),
|
|
).toBeUndefined();
|
|
});
|
|
|
|
it("rejects accepted cards for another child, agent, or interaction kind", () => {
|
|
const unrelated = [
|
|
{ ...child.interactions[1], issueId: "other" },
|
|
{ ...child.interactions[1], addresseeAgentId: "other", resolvedByAgentId: "other" },
|
|
{ ...child.interactions[1], payload: { target: { type: "custom", key: "other", revisionId: "decision-2" } } },
|
|
{ ...child.interactions[1], resolvedByRunId: "wrong-run" },
|
|
];
|
|
for (const interaction of unrelated)
|
|
expect(
|
|
storyAcceptedAgentReview({ ...child, interactions: [interaction] }, "first", "lead", runs),
|
|
).toBeUndefined();
|
|
});
|
|
|
|
it("rejects a review whose decision revision does not match the reviewer run", () => {
|
|
const bad = {
|
|
...child.interactions[1],
|
|
payload: { target: { type: "custom", key: "native_completion_review", revisionId: "wrong" } },
|
|
};
|
|
expect(storyAcceptedAgentReview({ ...child, interactions: [bad] }, "second", "lead", runs)).toBeUndefined();
|
|
});
|
|
});
|
|
|
|
describe("manual everyday workflow catalog", () => {
|
|
it("is discoverable and explicitly selected without changing scheduled --all", () => {
|
|
const all = selectRunnerExecutions(parseRunnerSelectors(["--all"]));
|
|
expect(all.some((e) => e.suite.id === "everyday-workflows")).toBe(false);
|
|
const selected = selectRunnerExecutions(
|
|
parseRunnerSelectors(["--suite", "everyday-workflows"]),
|
|
);
|
|
expect(selected).toHaveLength(50);
|
|
expect(selected.every((e) => e.profile.generation === "native")).toBe(true);
|
|
expect(
|
|
selectRunnerExecutions(
|
|
parseRunnerSelectors(["--profile", "runner-codex"]),
|
|
).some((e) => e.suite.id === "everyday-workflows"),
|
|
).toBe(false);
|
|
const listed = selectRunnerExecutions(parseRunnerSelectors(["--list"]));
|
|
expect(listed.some((e) => e.suite.id === "everyday-workflows")).toBe(true);
|
|
});
|
|
it("does not schedule arbitrary runner-crash probes as model evals", () => {
|
|
const selected = selectRunnerExecutions(
|
|
parseRunnerSelectors(["--suite", "everyday-workflows"]),
|
|
);
|
|
expect(selected.some((e) => e.task.id.startsWith("recover-runner"))).toBe(false);
|
|
for (const id of ["recover-runner", "recover-runner-safe", "recover-runner-uncertain"])
|
|
expect(() => selectRunnerExecutions(parseRunnerSelectors([
|
|
"--id", `everyday-workflows.runner-codex.local.${id}`,
|
|
]))).toThrow();
|
|
expect(selected.some((e) => e.task.id === "recover-controller")).toBe(true);
|
|
expect(selected.some((e) => e.task.id === "stop-redirect")).toBe(true);
|
|
});
|
|
it("keeps ordinary prompts free of completion/API instructions", () => {
|
|
for (const task of everydayTasks.filter(
|
|
(candidate) => candidate.id !== "agent-review-handoff",
|
|
))
|
|
expect(task.buildPrompt("sample")).not.toMatch(
|
|
/finish_task|paperclip_finish|PATCH|mark .*done|idempotencyKey/i,
|
|
);
|
|
expect(
|
|
everydayTasks
|
|
.find((task) => task.id === "agent-review-handoff")!
|
|
.buildPrompt("sample"),
|
|
).toContain("resolve_review");
|
|
const profile = runnerSuites.find((s) => s.id === "everyday-workflows")!
|
|
.profiles[0]!;
|
|
const value = productionStoryProfile(profile).buildAgent({
|
|
environmentId: "env",
|
|
environmentFixtureId: "local",
|
|
workspacePath: "/tmp/test",
|
|
secretRefs: {
|
|
OPENAI_API_KEY: {
|
|
type: "secret_ref",
|
|
secretId: "secret",
|
|
version: "latest",
|
|
},
|
|
},
|
|
executionId: "case",
|
|
});
|
|
expect(JSON.stringify(value.instructionsBundle)).not.toMatch(
|
|
/fixture|mark .*done|finish_task|api\/issues/i,
|
|
);
|
|
});
|
|
it("does not claim unsupported remote crash/hiring coverage", () => {
|
|
const remote = runnerMatrix.filter(
|
|
(e) =>
|
|
e.suite.id === "everyday-workflows" && e.environment.id === "daytona",
|
|
);
|
|
expect(remote).toHaveLength(8);
|
|
expect(new Set(remote.map((e) => e.task.id))).toEqual(
|
|
new Set(["build-revise", "delegate-feedback", "recover-controller", "create-skill-studio"]),
|
|
);
|
|
});
|
|
});
|
|
|
|
describe("lifecycle oracle calibrated failures", () => {
|
|
const parent = {
|
|
id: "parent",
|
|
companyId: "company",
|
|
title: "Project",
|
|
status: "done",
|
|
assigneeAgentId: "lead",
|
|
};
|
|
const run: StoryRun = {
|
|
id: "run",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "succeeded",
|
|
runtimeMode: "native",
|
|
runnerInstanceId: "runner",
|
|
contextSnapshot: { issueId: "parent" },
|
|
};
|
|
const score = (
|
|
runs: StoryRun[],
|
|
issues = [parent],
|
|
allowedInterruptedRuns: string[] = [],
|
|
) =>
|
|
storyLifecycleChecks({
|
|
issues,
|
|
runs,
|
|
parentId: parent.id,
|
|
leadId: "lead",
|
|
allowedInterruptedRuns,
|
|
});
|
|
it("accepts successful owned native work", () =>
|
|
expect(score([run]).every((c) => c.passed)).toBe(true));
|
|
it.each([
|
|
["legacy execution", { ...run, runtimeMode: "legacy" }, "native-runtime"],
|
|
[
|
|
"no runner identity",
|
|
{ ...run, runnerInstanceId: null },
|
|
"native-runtime",
|
|
],
|
|
["worker on parent", { ...run, agentId: "worker" }, "parent-owned-by-lead"],
|
|
["crash without recovery", { ...run, status: "failed" }, "successful-runs"],
|
|
["unsettled run", { ...run, status: "running" }, "settled"],
|
|
] as const)("rejects %s", (_label, bad, id) =>
|
|
expect(score([bad]).find((c) => c.id === id)?.passed).toBe(false),
|
|
);
|
|
it("does not exempt unrelated failures because another run was intentionally stopped", () => {
|
|
const checks = score(
|
|
[
|
|
{ ...run, status: "cancelled" },
|
|
{ ...run, id: "other", status: "failed" },
|
|
],
|
|
[parent],
|
|
["run"],
|
|
);
|
|
expect(checks.find((c) => c.id === "successful-runs")?.passed).toBe(false);
|
|
});
|
|
it.each(["adapter_failed", "tool_validation_error", "timed_out"])(
|
|
"does not excuse a later %s on the same deliberately interrupted run",
|
|
(errorCode) => {
|
|
const checks = score(
|
|
[{ ...run, status: "failed", errorCode }],
|
|
[parent],
|
|
[run.id],
|
|
);
|
|
expect(checks.find((c) => c.id === "successful-runs")?.passed).toBe(false);
|
|
const failed = { ...run, status: "failed", errorCode };
|
|
expect(storyUnexpectedRunFailure([failed], [run.id])).toBe(failed);
|
|
},
|
|
);
|
|
it.each([
|
|
["cancelled", "cancelled"],
|
|
["interrupted", "server_shutdown_interrupted"],
|
|
["failed", "process_lost"],
|
|
])("accepts an injected interruption with %s / %s", (status, errorCode) => {
|
|
const checks = score([{ ...run, status, errorCode }], [parent], [run.id]);
|
|
expect(checks.find((c) => c.id === "successful-runs")?.passed).toBe(true);
|
|
expect(storyUnexpectedRunFailure([{ ...run, status, errorCode }], [run.id])).toBeUndefined();
|
|
});
|
|
it("excludes a proven pre-dispatch workspace deferral without hiding executed failures", () => {
|
|
const deferred: StoryRun = {
|
|
id: "deferred",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "cancelled",
|
|
errorCode: "workspace_busy",
|
|
resultJson: {
|
|
executionRecovery: {
|
|
kind: "workspace_wait",
|
|
providerWorkStarted: false,
|
|
},
|
|
},
|
|
};
|
|
expect(score([run, deferred]).every((c) => c.passed)).toBe(true);
|
|
expect(
|
|
score([run, { ...deferred, processPid: 123 }]).every((c) => c.passed),
|
|
).toBe(false);
|
|
expect(
|
|
score([run, { ...deferred, resultJson: {} }]).every((c) => c.passed),
|
|
).toBe(false);
|
|
});
|
|
it("rejects a finished answer left in review", () =>
|
|
expect(
|
|
score([run], [{ ...parent, status: "in_review" }]).find(
|
|
(c) => c.id === "tasks-done",
|
|
)?.passed,
|
|
).toBe(false));
|
|
});
|
|
|
|
describe("reply completion boundary", () => {
|
|
const run: StoryRun = {
|
|
id: "old",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "succeeded",
|
|
runnerProfileJson: {
|
|
nativeExecutionInput: { task: { prompt: "Original task" } },
|
|
},
|
|
};
|
|
it("does not accept Done from the first run while a later request is still queued", () => {
|
|
expect(storyRepliesConsumed([run], ["later-comment"])).toBe(false);
|
|
const next = {
|
|
...run,
|
|
id: "new",
|
|
runnerProfileJson: {
|
|
nativeExecutionInput: { task: { prompt: "User reply later-comment" } },
|
|
},
|
|
};
|
|
expect(
|
|
storyRepliesConsumed(
|
|
[run, { ...next, status: "running" }],
|
|
["later-comment"],
|
|
),
|
|
).toBe(false);
|
|
expect(storyRepliesConsumed([run, next], ["later-comment"])).toBe(true);
|
|
});
|
|
|
|
it("waits for the expected agent reply after terminal state is visible", () => {
|
|
const issue = {
|
|
id: "parent",
|
|
companyId: "company",
|
|
title: "Stop and redirect",
|
|
status: "done",
|
|
comments: [
|
|
{
|
|
id: "user-reply",
|
|
authorAgentId: null,
|
|
body: "Change direction. Reference nonce.",
|
|
},
|
|
],
|
|
};
|
|
expect(storyHasAgentReply(issue, "lead", "Reference nonce")).toBe(false);
|
|
expect(
|
|
storyHasAgentReply(
|
|
{
|
|
...issue,
|
|
comments: [
|
|
...issue.comments,
|
|
{
|
|
id: "agent-reply",
|
|
authorAgentId: "lead",
|
|
body: "The studio is ready. Reference nonce.",
|
|
},
|
|
],
|
|
},
|
|
"lead",
|
|
"Reference nonce",
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
storyHasAgentReply(
|
|
{
|
|
...issue,
|
|
comments: [
|
|
...issue.comments,
|
|
{
|
|
id: "other-agent-reply",
|
|
authorAgentId: "worker",
|
|
body: "The studio is ready. Reference nonce.",
|
|
},
|
|
],
|
|
},
|
|
"lead",
|
|
"Reference nonce",
|
|
),
|
|
).toBe(false);
|
|
});
|
|
|
|
it("keeps a blocked lead alive when its child has a durable agent wake", () => {
|
|
const child = {
|
|
id: "child",
|
|
companyId: "company",
|
|
title: "Worker delivery",
|
|
status: "in_progress",
|
|
interactions: [
|
|
{
|
|
id: "review",
|
|
status: "pending",
|
|
continuationPolicy: "wake_assignee",
|
|
addresseeAgentId: "lead",
|
|
},
|
|
],
|
|
wakeDiagnostics: {
|
|
events: [
|
|
{
|
|
kind: "wake_request",
|
|
agentId: "lead",
|
|
status: "queued",
|
|
},
|
|
],
|
|
},
|
|
};
|
|
expect(storyHasPendingAgentWake(child, "lead")).toBe(true);
|
|
expect(
|
|
storyHasPendingAgentWork(
|
|
[
|
|
{ ...child, status: "blocked", wakeDiagnostics: { events: [] } },
|
|
child,
|
|
],
|
|
"lead",
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
storyHasPendingAgentWork(
|
|
[{ ...child, status: "blocked", wakeDiagnostics: { events: [] } }],
|
|
"lead",
|
|
),
|
|
).toBe(false);
|
|
expect(
|
|
storyHasStrandedBlockedLeaf(
|
|
[{ ...child, status: "blocked", wakeDiagnostics: { events: [] } }],
|
|
"lead",
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
storyHasStrandedBlockedLeaf(
|
|
[
|
|
{ ...child, status: "blocked", wakeDiagnostics: { events: [] } },
|
|
child,
|
|
],
|
|
"lead",
|
|
),
|
|
).toBe(false);
|
|
expect(storyHasPendingHumanInteraction(child, "lead")).toBe(false);
|
|
expect(storyHasPendingAgentReview(child, "lead")).toBe(true);
|
|
expect(storyHasPendingAgentWake(child, "worker")).toBe(false);
|
|
expect(
|
|
storyHasPendingHumanInteraction(
|
|
{
|
|
...child,
|
|
wakeDiagnostics: { events: [] },
|
|
},
|
|
"lead",
|
|
),
|
|
).toBe(true);
|
|
const sameAgentCards = {
|
|
...child,
|
|
interactions: [
|
|
{
|
|
id: "review-a",
|
|
status: "pending",
|
|
continuationPolicy: "wake_assignee",
|
|
addresseeAgentId: "lead",
|
|
},
|
|
{
|
|
id: "review-b",
|
|
status: "pending",
|
|
continuationPolicy: "wake_assignee",
|
|
addresseeAgentId: "lead",
|
|
},
|
|
],
|
|
wakeDiagnostics: {
|
|
events: [
|
|
{
|
|
kind: "wake_request",
|
|
agentId: "lead",
|
|
status: "queued",
|
|
payload: { nativeReviewInteractionId: "review-a" },
|
|
},
|
|
],
|
|
},
|
|
};
|
|
expect(storyHasPendingAgentReview(sameAgentCards, "lead")).toBe(true);
|
|
expect(storyHasPendingHumanInteraction(sameAgentCards, "lead")).toBe(
|
|
true,
|
|
);
|
|
expect(
|
|
storyHasPendingHumanInteraction(
|
|
{
|
|
...child,
|
|
interactions: [
|
|
{
|
|
id: "review",
|
|
status: "pending",
|
|
continuationPolicy: "wake_assignee",
|
|
effectiveResolverPolicy: "human_only",
|
|
addresseeAgentId: "lead",
|
|
},
|
|
],
|
|
},
|
|
"lead",
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
storyHasPendingHumanInteraction(
|
|
{
|
|
...child,
|
|
interactions: [
|
|
...(child.interactions ?? []),
|
|
{
|
|
id: "human-review",
|
|
status: "pending",
|
|
continuationPolicy: "wake_assignee",
|
|
resolverPolicy: "human_only",
|
|
addresseeAgentId: null,
|
|
},
|
|
],
|
|
},
|
|
["lead", "hired-worker"],
|
|
),
|
|
).toBe(true);
|
|
expect(storyHasPendingAgentWake(child, ["hired-worker", "lead"])).toBe(
|
|
true,
|
|
);
|
|
expect(
|
|
storyIssueHasBlockedTimeline({
|
|
...child,
|
|
status: "done",
|
|
activity: [
|
|
{
|
|
action: "issue.updated",
|
|
details: { fromStatus: "in_progress", toStatus: "blocked" },
|
|
createdAt: "2026-09-17T08:00:00.000Z",
|
|
},
|
|
],
|
|
}),
|
|
).toBe(true);
|
|
expect(
|
|
storyIssueHasBlockedTimelineBefore(
|
|
{
|
|
...child,
|
|
status: "done",
|
|
activity: [
|
|
{
|
|
action: "issue.updated",
|
|
details: { status: "blocked" },
|
|
createdAt: "2026-09-17T08:00:00.000Z",
|
|
},
|
|
],
|
|
},
|
|
"2026-09-17T08:01:00.000Z",
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
storyIssueHasUnresolvedDependency({
|
|
...child,
|
|
status: "blocked",
|
|
wakeDiagnostics: {
|
|
events: [],
|
|
blockerDiagnostics: {
|
|
readiness: { unresolvedBlockerCount: 1 },
|
|
},
|
|
},
|
|
}),
|
|
).toBe(true);
|
|
expect(
|
|
storyRunReportsDependencyBlock({
|
|
id: "parent-run",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "succeeded",
|
|
resultJson: {
|
|
nativeResult: {
|
|
reportedWorkDisposition: "blocked",
|
|
blocker: { reasonCode: "dependency_unresolved" },
|
|
},
|
|
},
|
|
}),
|
|
).toBe(true);
|
|
});
|
|
|
|
it("keeps polling when review acceptance is durable but the parent wake is not projected yet", () => {
|
|
const issues = [
|
|
{
|
|
id: "parent", companyId: "company", title: "parent", status: "blocked",
|
|
blockedTransitionAt: "2026-09-17T19:07:30.000Z",
|
|
},
|
|
{
|
|
id: "child",
|
|
parentId: "parent",
|
|
companyId: "company",
|
|
title: "child",
|
|
status: "done",
|
|
interactions: [{
|
|
id: "review",
|
|
issueId: "child",
|
|
kind: "request_confirmation",
|
|
status: "accepted",
|
|
addresseeAgentId: "lead",
|
|
resolvedByAgentId: "lead",
|
|
resolvedByRunId: "review-run",
|
|
resolvedAt: "2026-09-17T19:08:10.000Z",
|
|
payload: { target: { type: "custom", key: "native_completion_review", revisionId: "decision" } },
|
|
result: { version: 1, outcome: "accepted" },
|
|
}],
|
|
},
|
|
];
|
|
const runs = [{
|
|
id: "review-run",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "succeeded",
|
|
finishedAt: "2026-09-17T19:08:11.000Z",
|
|
}];
|
|
expect(storyHasStrandedBlockedLeaf(issues, "lead")).toBe(true);
|
|
expect(storyHasDurableAgentReviewContinuation(
|
|
issues, "parent", "lead", runs,
|
|
)).toBe(true);
|
|
for (const status of ["completed", "failed", "cancelled"]) {
|
|
expect(storyHasDurableAgentReviewContinuation([
|
|
{ ...issues[0]!, wakeDiagnostics: { events: [{
|
|
kind: "wake_request", agentId: "lead", reason: "issue_blockers_resolved", status,
|
|
}] } },
|
|
issues[1]!,
|
|
], "parent", "lead", runs)).toBe(false);
|
|
}
|
|
expect(storyHasDurableAgentReviewContinuation([
|
|
{ ...issues[0]!, blockedTransitionAt: "2026-09-17T19:09:00.000Z" }, issues[1]!,
|
|
], "parent", "lead", runs)).toBe(false);
|
|
});
|
|
|
|
it("still identifies a blocked parent as stranded without accepted review evidence", () => {
|
|
const issues = [
|
|
{ id: "parent", companyId: "company", title: "parent", status: "blocked" },
|
|
{ id: "child", parentId: "parent", companyId: "company", title: "child", status: "done", interactions: [] },
|
|
];
|
|
expect(storyHasDurableAgentReviewContinuation(issues, "parent", "lead", [])).toBe(false);
|
|
});
|
|
|
|
it.each(["deferred_issue_execution", "claimed"])(
|
|
"keeps a %s agent wake actionable",
|
|
(status) =>
|
|
expect(
|
|
storyHasPendingAgentWake(
|
|
{
|
|
id: "child",
|
|
companyId: "company",
|
|
title: "Worker delivery",
|
|
status: "in_progress",
|
|
wakeDiagnostics: {
|
|
events: [{ kind: "wake_request", agentId: "lead", status }],
|
|
},
|
|
},
|
|
"lead",
|
|
),
|
|
).toBe(true),
|
|
);
|
|
});
|
|
|
|
describe("delegation completion order", () => {
|
|
it("rejects a parent closed before its worker finishes", () => {
|
|
const base: StoryRun = {
|
|
id: "lead-run",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "succeeded",
|
|
nativeIssueId: "parent",
|
|
finishedAt: "2026-09-14T12:00:00Z",
|
|
};
|
|
const child = {
|
|
...base,
|
|
id: "worker-run",
|
|
agentId: "worker",
|
|
nativeIssueId: "child",
|
|
finishedAt: "2026-09-14T12:01:00Z",
|
|
};
|
|
expect(
|
|
storyParentFinishedAfterChildren([base, child], "parent", "lead", [
|
|
"child",
|
|
]),
|
|
).toBe(false);
|
|
expect(
|
|
storyParentFinishedAfterChildren(
|
|
[{ ...base, finishedAt: "2026-09-14T12:02:00Z" }, child],
|
|
"parent",
|
|
"lead",
|
|
["child"],
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
storyParentFinishedAfterChildren([base], "parent", "lead", ["child"]),
|
|
).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe("cancelled queued workspace retry", () => {
|
|
const queued: StoryRun = {
|
|
id: "retry",
|
|
companyId: "company",
|
|
agentId: "agent",
|
|
status: "cancelled",
|
|
runtimeMode: "legacy",
|
|
runtimeModeResolvedAt: null,
|
|
startedAt: null,
|
|
retryOfRunId: "workspace-deferral",
|
|
scheduledRetryReason: "workspace_busy",
|
|
errorCode: "cancelled",
|
|
lastOutputSeq: 0,
|
|
resultJson: {
|
|
startupCancellation: {
|
|
requestedAt: "2026-09-14T19:10:23.057Z",
|
|
beforeNativeSelection: false,
|
|
},
|
|
},
|
|
};
|
|
it("does not mistake an unstarted cancelled retry for provider execution", () => {
|
|
expect(isStoryWorkspaceDeferral(queued)).toBe(true);
|
|
});
|
|
it.each([
|
|
{ startedAt: "2026-09-14T19:10:22Z" },
|
|
{ processPid: 42 },
|
|
{ runnerInstanceId: "runner" },
|
|
{ nativeSessionId: "session" },
|
|
{ lastOutputSeq: 1 },
|
|
{ usageJson: { outputTokens: 1 } },
|
|
{ runtimeModeResolvedAt: "2026-09-14T19:10:22Z" },
|
|
{ scheduledRetryReason: "other" },
|
|
{ resultJson: null },
|
|
])("retains a contradictory or unproven cancellation %j", (change) => {
|
|
expect(isStoryWorkspaceDeferral({ ...queued, ...change })).toBe(false);
|
|
});
|
|
});
|