Files
DottaandPaperclip 6c1a75da49 feat(connections): make AgentMail a default connection with inline setup (#14772)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents access to external services.
> - AgentMail needs both a saved key and an inbox assigned to the agent.
> - Chat requests offered a setup link instead of an inline card and
could treat a saved key as complete.
> - Inbox setup also hid address conflicts behind a generic server error
and a separate review step.
> - This pull request makes AgentMail a default connection, adds the
inline card, reduces setup to two steps, and shows conflicts beside the
address.
> - Shared native dropdown styles also give every caret a consistent
inset.

## Linked Issues or Issue Description

**What happened?**

AgentMail requests in chat did not show a usable inline connection card.
Manual setup required extra screens, ignored saved account keys, and
could trap new-address setup in a locked inbox dropdown. Agent selectors
omitted the avatar from the selected value. A taken address could
produce an HTTP 403 from AgentMail and appear as an internal server
error. Native dropdown arrows also touched the right edge of their
fields.

**Expected behavior**

Make AgentMail available as a default connection. Ask for the API key
inline, with a direct link to its provider page. Default human access to
the company and agent access to the requesting agent. Resume the agent
only after an assigned inbox is active. Manual setup should ask for an
agent and email address, then finish. Address checks should run as the
user types. Taken addresses should show clickable alternatives. A domain
dropdown beside the name should prefer a verified custom domain. Setup
should suggest authorized saved AgentMail keys and show agent avatars in
the picker and selected value.

**Steps to reproduce**

1. Ask an agent to connect AgentMail when it has no assigned inbox.
2. Check that an inline API-key card appears and links to the provider's
API-key page.
3. Open AgentMail setup, choose an agent, and request an address that is
already taken.
4. Correct the inline error, refresh, and finish setup with the same
request ID.
5. Inspect native dropdown carets in light, dark, disabled, and
right-to-left states.

Uses the bounded provider-error parser merged in #14768. Related work:
#13256 introduced AgentMail; #14725 expanded connection search.

## What Changed

- Stop recurring email queries for tasks that have no email thread.
Share the query between the thread provider and activity view. Keep
email-task updates and invalidation-based discovery.
- Make AgentMail available without the experimental chat setting. Keep
the catalog, setup and management routes, agent Channels tab, task email
feed, receiving worker, and agent tools available by default. Other
experimental chat providers stay gated.

- Make the email address and copy icon a single clickable action with
the shared Copied! confirmation. Add View inbox linking directly to the
matching AgentMail console inbox, with the address encoded as one URL
path segment.

- Reorganize inbox Settings around the copyable email address, usage
instructions, and receiving status. Move reconnect credentials into a
disclosure and separate the Disconnect action. Add production Settings
stories for active, paused, unassigned-address, revoked, webhook,
long-address, mobile, and reconnect states. Show repair controls when
the inbox has an error. Keep usage instructions tied to an active inbox
with an address.

- Add AgentMail channel intents and an inline key field with the direct
API-key URL.
- Keep setup and retry state tied to the interaction. Require an active
inbox for completion. Preserve company and agent access checks.
- Reduce manual setup to agent selection and email selection. Put the
domain dropdown beside the address and default to a verified custom
domain. Preserve explicit choices across reloads. Keep receiving
settings under Advanced options.
- Check the initial address and edits after a 350 ms pause. Abort
superseded requests and ignore stale responses. Show clickable
suggestions and retain known creation conflicts across reloads.
- Add a company-scoped, manager-only address check using the saved
credential. Search the visible inbox list instead of fetching an
uncreated inbox: live AgentMail retains negative lookups that can break
subsequent access-key creation. Unlisted addresses remain unknown;
creation is authoritative.
- Suggest labeled saved AgentMail keys in both manual setup and the
inline card. Filter by company, provider, active credential, and
current-user grants on the server. Prefer an account key and preserve
the selected key or an explicit new-key choice across refresh. Use
verified scope metadata and bounded concurrent checks for legacy keys.
Never return secret values.
- Catch an inbox-only key before the email step. Allow its existing
inbox only after an explicit choice. Recover old locked drafts at the
key picker. Save the replacement key before retiring an empty draft,
then use a new setup URL so refresh preserves the switched account; stop
if cleanup fails. Preserve already allocated addresses and their
original accounts.
- Use the shared AgentSelect in email setup. Show the canonical agent
avatar in each option and the selected value, including other consumers
of the shared component. Add regression coverage for legacy and current
Lucide agent-mention icon formats.
- Start each catalog Add connection with a fresh setup identity. Honor
Finish setup's exact draft/account/address instead of resuming an
unrelated browser draft. Return Cancel and Done to Connectors and Email
settings to the inbox. Group the task/thread explanation in a How it
Works card.
- Route AgentMail catalog removal through the email inbox control API,
including unfinished drafts. Refresh both the catalog and inbox views.
- Render each inbox management tab separately. Access uses the saved
account grants and agent controls; Conversations and Activity use the
shared persisted email feed. Activity lifecycle actions use the email
API. Reconnect returns to inbox Settings. Conversation failures show a
retry instead of a false empty state. Email delivery recovery stays in
the task.
- Map documented provider address conflicts to a field error. Preserve
actionable messages for other failures.
- Preserve non-secret draft fields across refresh, scoped to the
requested agent. Never save API keys in browser storage. Resume partial
inbox creation with the original agent, address, and request ID.
- Show an already-created address with explicit retry and new-address
recovery instead of locked inputs. Preserve the original inbox and
resumable draft when choosing another address. Distinguish runtime-key
404 errors and log safe provider status/operation/code.
- Apply final agent access once within email setup authorization for a
new account whose original installs are unchanged. Preserve later
permission edits and reused account installs. Support in-place retry of
progress loading.
- Let a failed inline setup change keys after retiring an empty draft.
Persist its replacement setup identity without storing secrets. Recover
a server-saved account when refresh interrupts the save response, while
preserving intentional account changes.
- Render the production setup in Storybook and add error, recovery, and
mobile states.
- Inset native select carets in shared CSS. Preserve custom icons,
listboxes, keyboard behavior, and forced-color controls.
- Add browser regression coverage and an AgentMail Product E2E case with
persisted-state and rendered-card evidence.

## Verification

- Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and
`git diff --check` passed after the default-availability change.
- All 485 focused tests passed. These cover setup, management, catalog
and route gates, connection intents, email authorization, Cursor
execution, and the OpenAPI contract. All 39 email integration tests run
with the experimental chat setting off.
- The shared polling change passed four behavioral tests, UI typecheck
and build, and token gates.
- `tests/e2e/agentmail.spec.ts` passed with the actual server setting
off. This full-stack browser test uses simulated provider responses. It
covers catalog entry, saved keys, editable address and domain controls,
creation, conflicts, retry, all management tabs, clipboard feedback, the
provider link, and task email rendering.
- In the live local browser, Add connection reached the editable email
step with the saved account key. The verified custom domain was selected
by default. Both domain choices worked. The existing inbox Settings page
remained available. Both active inboxes completed new mail checks with
the setting off. No new provider inbox or email message was created for
this pass.
- Earlier live provider acceptance covered creation on a verified custom
domain, Finish connecting on the reported draft, successful mail checks
after refresh, and catalog removal of disposable draft and active
connections. Clicking the email address copied the exact address and
showed Copied!. View inbox opened the same inbox in AgentMail’s console.
No email messages were sent.
- Production setup and Settings Storybook builds and interactions
passed. Settings states include active, paused, unassigned, revoked,
webhook, long-address, mobile, and reconnect. Receiving and
revoked-access stories had zero accessibility violations.
- Full local `pnpm test:run` on an earlier revision completed with
14,709 passing, 87 skipped, and four transient failures. All four failed
cases passed in focused reruns without product changes. That serial full
local command was not repeated after each follow-up. The latest-head
full CI suite is the final test gate.
- CI found an obsolete browser assertion that hid every channel when the
flag was off. Updated it to keep AgentMail and the Channels surface
visible while preserving the GitHub chat route gates. All 11 provider
browser tests passed locally after scoping the Channels selector to the
agent sidebar. Two initial local attempts stopped at temporary Postgres
initialization. The passing run used a separate disposable database on
the existing local Postgres server; it was removed after the test.
- Updated the remaining sidebar and aggregator discovery assertions for
default AgentMail availability. Ordinary task fixtures now return no
email thread. All 128 sidebar/task-page tests and all 42 aggregator
tests passed locally.
- Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI
passed, with 54 successful checks including Snyk and two intentional
Storybook skips. The CI run is
https://github.com/paperclipai/paperclip/actions/runs/37020833647. A
fresh Greptile review scored 5/5 with no unresolved threads. Live model
evaluations and inbound/outbound email delivery were not run.

## Risks

- AgentMail no longer needs experimental opt-in. Setup still requires a
human to connect an account and assign an inbox. Inline setup creates an
inbox after a human submits a new or saved key. Company access, agent
access, inbox assignment, and completion checks remain enforced.
- AgentMail read APIs cannot prove global address availability. The
visible-list check is bounded to 100 entries and cannot see inboxes
outside the key’s scope. The UI reports this limitation, suggests
alternatives without claiming they are free, and keeps final creation
conflicts inline. Lookup outages show an error without preventing the
authoritative creation attempt.
- Native select CSS affects the whole app. Custom-icon selects and
multi-row lists are excluded. Forced-color mode keeps the browser caret.
- Saved-key discovery uses stored verified scope metadata and checks
authorized legacy credentials concurrently within a shared three-second
deadline. Provider outages mark legacy choices unavailable; users can
still enter another key. Final use rechecks authorization and provider
access.
- No database migration or transport default change. Live connection
remains the default.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
exact served model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (latest head
`b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 10:01:15 -05:00

1617 lines
62 KiB
TypeScript

import { gradeAgentmailSetup } from "./agentmail-setup-evidence.js";
import { expect, type Page } from "@playwright/test";
import { runnerApiToolsEnabled } from "../../server/src/services/native-runtime/runner-api-rollout.js";
import { spawn } from "node:child_process";
import { createHash } from "node:crypto";
import { mkdir, readFile, readdir } from "node:fs/promises";
import path from "node:path";
import { isDeepStrictEqual } from "node:util";
import { pollUntil, type RunnerApi } from "./api.js";
import {
latestStoryDelivery,
type StoryDelivery,
} from "./everyday-delivery.js";
import type { LiveFixtureValues } from "./live-fixtures.js";
import type { MatrixExecution } from "./types.js";
import { createTaskThroughUi, submitTaskReply } from "./user-actions.js";
import { waitForTaskChatRendered } from "./continuation-screenshot.js";
import { hasPersistedSource, isSavedSourceCheckpoint } from "./everyday-interruption.js";
import { setupAggregatorFixture } from "./aggregator-fixture.js";
import { gradeProviderChoice, gradeProviderOutcome } from "./connection-routing-evidence.js";
import { setupConnectionReview } from "./connection-reviews.js";
import {
pendingStoryDecision,
StoryDecisionError,
type StoryInteraction,
} from "./everyday-decisions.js";
import { LATE_REQUIREMENT, SLUGIFY_REVISION, requiresEverydayArtifactOracle } from "./everyday-cases.js";
import {
isActiveStoryRun,
isStoryWorkspaceDeferral,
isExpectedStoryInterruption,
storyUnexpectedRunFailure,
storyUnexercisedReviewBoundary,
storyLifecycleChecks,
storyRepliesConsumed,
storyHasAgentReply,
storyHasPendingHumanInteraction,
storyReviewContinuationTimeoutDetail,
storyHasStrandedBlockedLeaf,
storyHasDurableAgentReviewContinuation,
storyIssueHasBlockedTimelineBefore,
storyIssueHasUnresolvedDependency,
storyRunReportsDependencyBlock,
storyParentFinishedAfterChildren,
storyAcceptedAgentReview,
artifactGradeModeForPhase,
storyParentCompletionPrecedesReview,
type StoryCheck,
type StoryIssue,
type StoryRun,
} from "./everyday-observations.js";
type Row = Record<string, any>;
export interface EverydayEvidence {
schema: "paperclip.everyday-workflow.v1";
caseId: string;
prompt: string;
harnessDigest?: string;
sourceRevision?: string;
providerVersion?: string;
fixtureConfiguration?: {
apiToolsEnabled: boolean;
aiConnection?: LiveFixtureValues["aiConnection"];
};
documents?: Row[];
checks: StoryCheck[];
timeline: Array<{ at: string; action: string; detail?: unknown }>;
issues: Array<StoryIssue & Row>;
runs: StoryRun[];
agents: Row[];
downloads: Row[];
allowedInterruptedRuns: string[];
}
interface Input {
page: Page;
api: RunnerApi;
fixtures: LiveFixtureValues;
execution: MatrixExecution;
nonce: string;
workspacePath: string;
privateDir: string;
deadlineAt: number;
restart(): Promise<void>;
observe(issue: StoryIssue, runs: StoryRun[]): void;
capture(id: string, label: string, file: string): Promise<void>;
evidence(name: string, value: unknown): Promise<void>;
}
function runCommand(
command: string,
args: string[],
timeout = 20_000,
): Promise<{ code: number; stdout: string }> {
return new Promise((resolve, reject) => {
const env = Object.fromEntries(
Object.entries(process.env).filter(([key]) =>
["PATH", "SYSTEMROOT", "TMPDIR"].includes(key),
),
);
const child = spawn(command, args, {
env,
stdio: ["ignore", "pipe", "pipe"],
});
let stdout = "";
let stderr = "";
const timer = setTimeout(() => child.kill("SIGKILL"), timeout);
child.stdout.on("data", (chunk) => {
stdout += String(chunk);
if (stdout.length > 1_000_000) child.kill("SIGKILL");
});
child.stderr.on("data", (chunk) => {
stderr += String(chunk);
});
child.on("error", (error) => {
clearTimeout(timer);
reject(error);
});
child.on("exit", (code) => {
clearTimeout(timer);
if (code === null)
reject(new Error("Bounded evaluator command timed out"));
else resolve({ code, stdout: stdout || stderr });
});
});
}
/** No DB writes, fabricated tool receipts, or corrective messages after a failed check. */
export async function runEverydayFlow(input: Input) {
const { page, api, fixtures, execution, nonce } = input;
const prefix = fixtures.company.issuePrefix!;
const ev: EverydayEvidence = {
schema: "paperclip.everyday-workflow.v1",
caseId: execution.task.id,
prompt: execution.task.buildPrompt(nonce),
fixtureConfiguration: {
apiToolsEnabled: runnerApiToolsEnabled(fixtures.company.id),
aiConnection: fixtures.aiConnection,
},
checks: [],
timeline: [],
issues: [],
runs: [],
agents: [],
downloads: [],
allowedInterruptedRuns: [],
};
const note = (action: string, detail?: unknown) =>
ev.timeline.push({
at: new Date().toISOString(),
action,
...(detail === undefined ? {} : { detail }),
});
const check = (id: string, passed: boolean, detail: string) => {
ev.checks.push({ id, passed, detail });
};
let parent: StoryIssue | undefined;
let lastSubmissionAt = 0;
const submittedCommentIds: string[] = [];
let review: Awaited<ReturnType<typeof setupConnectionReview>> | undefined;
let project = fixtures.project;
const caseId = execution.task.id;
const providerChoice = caseId === "provider-decline" || caseId === "provider-second";
const nativeProviderCase = caseId === "provider-native";
let aggregatorFixture: Awaited<ReturnType<typeof setupAggregatorFixture>> | undefined;
const agentmailSetup = caseId === "agentmail-setup";
const decliningConnection = caseId === "connection-decline" || nativeProviderCase || agentmailSetup;
const declining = decliningConnection || caseId === "service-decline";
let decisionId: string | undefined;
let decisionResolvedAt: string | undefined;
let initialConnections: string[] = [];
let stoppedWorkspace: Record<string, string> | undefined;
let settledAgentReply: Row | undefined;
let reviewHandoffBoundary: {
childId: string;
assigneeAgentId: string | null | undefined;
interactionId: string;
interactionCreatedAt?: string;
parentRunId?: string;
parentBlockedObserved: boolean;
} | undefined;
async function workspaceFiles() {
const workspaceRoot = input.workspacePath;
const files: Record<string, string> = {};
for (const entry of await readdir(workspaceRoot, {
withFileTypes: true,
})) {
if (entry.isFile() && /\.(py|md|zip)$/.test(entry.name))
files[entry.name] = createHash("sha256")
.update(await readFile(path.join(workspaceRoot, entry.name)))
.digest("hex");
}
return files;
}
async function runnerStopped(run: StoryRun) {
if (!run.processPid) return false;
const identity = await runCommand("ps", [
"-p",
String(run.processPid),
"-o",
"command=",
]);
return (
identity.code !== 0 || !identity.stdout.includes(`--run-id ${run.id} `)
);
}
async function refresh() {
const listed = await api.get<StoryRun[]>(
`/api/companies/${fixtures.company.id}/heartbeat-runs?limit=100`,
);
ev.runs = await Promise.all(
listed.map((r) => api.get<StoryRun>(`/api/heartbeat-runs/${r.id}`)),
);
const issues = await api.get<StoryIssue[]>(
`/api/companies/${fixtures.company.id}/issues`,
);
ev.issues = await Promise.all(
issues.map(async (issue) => ({
...issue,
comments: await api.get<Row[]>(
`/api/issues/${issue.id}/comments?order=asc`,
),
queuedComments: await api.get<Row>(
`/api/issues/${issue.id}/queued-comments`,
),
interactions: await api.get<Row[]>(
`/api/issues/${issue.id}/interactions`,
),
wakeDiagnostics: await api.get<Row>(
`/api/issues/${issue.id}/diagnostics/wakes`,
),
activity: await api.get<Row[]>(`/api/issues/${issue.id}/activity`),
})),
);
if (parent) {
parent = ev.issues.find((i) => i.id === parent!.id) as StoryIssue;
input.observe(parent!, ev.runs);
}
return ev;
}
const taskUrl = (issue: StoryIssue) =>
`/${prefix}/issues/${issue.identifier ?? issue.id}`;
async function openTask(issue: StoryIssue) {
await page.goto(taskUrl(issue), { waitUntil: "domcontentloaded" });
if (providerChoice || nativeProviderCase) {
await expect(page.locator('[data-testid="task-chat-thread"], [data-testid="thread-root"]').first()).toBeVisible({timeout:30_000});
await expect(page.getByRole("heading", {name:String(issue.title),exact:true})).toBeVisible();
await expect(page.getByTestId("issue-chat-skeleton")).toHaveCount(0);
} else await waitForTaskChatRendered(page, String(issue.title));
}
async function openParent() {
await openTask(parent!);
}
function observableAgentIds(state: EverydayEvidence) {
return [
fixtures.agent.id,
...state.issues.flatMap((issue) =>
[
issue.assigneeAgentId,
...(issue.interactions ?? []).map(
(interaction) => interaction.addresseeAgentId,
),
].filter((agentId): agentId is string => Boolean(agentId)),
),
];
}
async function reply(message: string, target: StoryIssue = parent!) {
const priorFailures = ev.checks.filter((c) => !c.passed);
if (priorFailures.length)
throw new Error(
`Story prerequisite failed before the next user request: ${priorFailures.map((c) => c.id).join(", ")}`,
);
const before = await api.get<Row[]>(`/api/issues/${target.id}/comments`);
lastSubmissionAt = Date.now();
await submitTaskReply(page, message);
const after = await pollUntil({
label: "one persisted user reply",
deadlineAt: Date.now() + 30_000,
load: () => api.get<Row[]>(`/api/issues/${target.id}/comments`),
accept: (rows) =>
rows.filter(
(c) => !c.authorAgentId && !before.some((old) => old.id === c.id),
).length > 0,
});
const added = after.filter(
(c) => !c.authorAgentId && !before.some((old) => old.id === c.id),
);
check(
`reply-${ev.timeline.length}-stored-once`,
added.length === 1,
"A composer submission creates exactly one user message.",
);
submittedCommentIds.push(...added.map((c) => c.id));
note("composer-message-persisted", {
issueId: target.id,
commentIds: added.map((c) => c.id),
});
}
async function settled(expectedAgentReply?: string) {
const settledState = await pollUntil({
label: `everyday ${caseId} settled`,
deadlineAt: input.deadlineAt,
timeoutDetail: (state) => state &&
storyReviewContinuationTimeoutDetail(
state.issues, parent?.id ?? "", fixtures.agent.id, state.runs,
observableAgentIds(state),
),
intervalMs: 1000,
load: refresh,
accept: (state) =>
state.issues.length > 0 &&
state.issues.every(
(i) => i.status === "done" && i.queuedComments.entries.length === 0,
) &&
storyRepliesConsumed(state.runs, submittedCommentIds) &&
state.runs.length > 0 &&
!state.runs.some(isActiveStoryRun) &&
state.runs.some(
(r) => Date.parse(r.finishedAt ?? "") >= lastSubmissionAt,
) &&
(!expectedAgentReply ||
storyHasAgentReply(
state.issues.find((issue) => issue.id === parent?.id),
fixtures.agent.id,
expectedAgentReply,
)),
reject: (state) => {
if (state.runs.length > 12) return "bounded execution count exceeded";
const bad = storyUnexpectedRunFailure(state.runs, ev.allowedInterruptedRuns);
if (bad)
return `native execution failed ${bad.errorCode ?? ""}: ${bad.error ?? bad.status}`;
if (
state.runs.some(isActiveStoryRun) ||
state.issues.some((i) => i.scheduledRetry || i.activeRecoveryAction)
)
return;
if (
state.runs.length &&
storyHasStrandedBlockedLeaf(state.issues, observableAgentIds(state)) &&
!storyHasDurableAgentReviewContinuation(
state.issues,
parent?.id ?? "",
fixtures.agent.id,
state.runs,
)
)
return "task is Blocked without an active continuation";
if (
state.issues.some(
(i) =>
i.status === "in_review" &&
storyHasPendingHumanInteraction(i, observableAgentIds(state)),
)
)
return "unexpected human interaction: task did not finish autonomously";
},
});
if (expectedAgentReply) {
const issue = settledState.issues.find((candidate) => candidate.id === parent?.id);
const reply = issue?.comments?.find(
(comment) =>
comment.authorAgentId === fixtures.agent.id &&
String(comment.body ?? "").includes(expectedAgentReply),
);
if (reply) {
// Keep the exact snapshot that satisfied the readiness predicate. A
// later refresh may return a newer projection and must not change the
// evidence used by the assertion below.
settledAgentReply = { ...reply };
note("expected-agent-reply-visible-at-settlement", {
issueId: issue?.id,
commentId: reply.id,
authorAgentId: reply.authorAgentId,
createdAt: reply.createdAt,
body: reply.body,
});
}
}
await openParent();
await expect(
page.getByTestId("issue-detail-header").getByRole("button", {
name: "Change status (current: Done)",
exact: true,
}),
).toBeVisible();
note("all-tasks-done");
}
async function downloadDelegated(
mode: "base" | "separator" | "max-length",
phase: string,
after?: number,
) {
const related = ev.issues.filter(
(issue) => issue.id === parent!.id || issue.parentId === parent!.id,
);
const attachments = (
await Promise.all(
related.map(async (issue) =>
(await api.get<Row[]>(`/api/issues/${issue.id}/attachments`)).map(
(a) => ({ ...a, issueId: issue.id }) as StoryDelivery,
),
),
)
).flat();
const delivery = latestStoryDelivery(
attachments,
related.map((issue) => issue.id),
after,
after === undefined ? [] : ev.downloads.map((d) => d.sha256),
);
if (!delivery) {
check(
`${phase}.zip-delivered`,
false,
"No new downloadable ZIP was delivered on the parent or its child tasks.",
);
return;
}
note("delivery-selected", { attachmentId: delivery.id, issueId: delivery.issueId, sha256: delivery.sha256, after });
await download(delivery.issueId, mode, phase, delivery.id);
}
async function download(
issueId: string,
mode: "base" | "separator" | "max-length",
phase: string,
selectedAttachmentId?: string,
) {
const attachments = await api.get<Row[]>(
`/api/issues/${issueId}/attachments`,
);
const zips = attachments
.filter(
(a) =>
String(a.originalFilename ?? a.filename ?? "").endsWith(".zip") ||
a.contentType === "application/zip",
)
.sort((a, b) => String(a.createdAt).localeCompare(String(b.createdAt)));
if (!zips.length) {
check(
`${phase}.zip-delivered`,
false,
"No downloadable ZIP attachment was delivered.",
);
return;
}
const attachment = selectedAttachmentId
? zips.find((a) => a.id === selectedAttachmentId)
: zips[zips.length - 1];
if (!attachment) throw new Error("Selected delivery is no longer available");
const issue = ev.issues.find((i) => i.id === issueId)!;
await openTask(issue);
const links = page.locator(
`a[href*="/api/attachments/${attachment.id}/content"]`,
);
const link = links.filter({ hasText: /Download/i }).first();
await expect(link).toBeVisible({ timeout: 15_000 });
const pending = page.waitForEvent("download", { timeout: 30_000 });
await link.click();
const file = await pending;
const target = path.join(input.privateDir, "snapshots", `${phase}.zip`);
await file.saveAs(target);
const bytes = await readFile(target);
const digest = createHash("sha256").update(bytes).digest("hex");
const result = await runCommand(process.env.PYTHON ?? "python3", [
path.join(import.meta.dirname, "everyday-artifact.py"),
target,
"--mode",
mode,
], Math.max(1, Math.min(60_000, input.deadlineAt - Date.now())));
const oracle = JSON.parse(result.stdout) as { checks: StoryCheck[] };
ev.checks.push(
...oracle.checks.map((c) => ({ ...c, id: `${phase}.${c.id}` })),
);
ev.downloads.push({
phase,
attachmentId: attachment.id,
issueId,
sha256: digest,
bytes: bytes.length,
filename: file.suggestedFilename(),
mode,
});
note("download-verified", {
phase,
attachmentId: attachment.id,
sha256: digest,
});
await openParent();
}
async function recordSource(
label = "source-saved-before-interruption",
evidenceName = "source-before-interruption.json",
requireActive = false,
maxWaitMs = 180_000,
) {
const filePath = path.join(input.workspacePath, "slugify.py");
const bytes = await pollUntil({
label: requireActive ? "active run with saved source" : "saved source after interruption",
deadlineAt: Math.min(input.deadlineAt, Date.now() + maxWaitMs),
intervalMs: 500,
load: async () => {
await refresh();
const source = await readFile(filePath).catch(() => undefined);
return { active: ev.runs.some(isActiveStoryRun), source };
},
accept: ({ active, source }) =>
requireActive
? isSavedSourceCheckpoint(active, source)
: hasPersistedSource(source),
});
if (!bytes.source) throw new Error("Saved source was not available at the controlled boundary");
await input.evidence(evidenceName, {
body: bytes.source.toString("utf8"),
sha256: createHash("sha256").update(bytes.source).digest("hex"),
});
note(label, {
sha256: createHash("sha256").update(bytes.source).digest("hex"),
bytes: bytes.source.length,
active: bytes.active,
});
}
async function sourceReady() {
await recordSource("source-saved-before-interruption", "source-before-interruption.json", true);
}
async function prepareStopBoundary() {
// Providers can finish a short first turn before the browser can click
// Stop. If that happens, submit one ordinary user follow-up through the
// composer and use that fresh run as the controlled interruption boundary.
// Save the checkpoint even if a fast provider has already finished. Do
// not shorten the normal source-creation budget to manufacture a timeout.
await recordSource("source-saved-before-interruption", "source-before-interruption.json");
await refresh();
if (ev.runs.some(isActiveStoryRun)) return;
await submitTaskReply(page, `${SLUGIFY_REVISION}\nContinue working until the source file is saved.`);
note("stop-boundary-continuation-submitted");
await recordSource(
"source-saved-before-interruption",
"source-before-interruption.json",
true,
);
}
try {
await mkdir(path.join(input.privateDir, "snapshots"), { recursive: true });
const revision = await runCommand("git", ["rev-parse", "HEAD"]);
if (revision.code === 0) ev.sourceRevision = revision.stdout.trim();
// Native providers run the packaged runtime (possibly remotely). A host
// `claude`/`codex` binary is neither required nor its observed version.
const harnessFiles = [
"everyday-flow.ts",
"everyday-cases.ts",
"everyday-decisions.ts",
"aggregator-fixture.ts",
"connection-routing-evidence.ts",
"agentmail-setup-evidence.ts",
"everyday-delivery.ts",
"everyday-observations.ts",
"everyday-artifact.py",
"user-actions.ts",
"runner.spec.ts",
"api.ts",
"harness-env.ts",
"failure-classifier.ts",
"live-fixtures.ts",
"connection-reviews.ts",
"catalog.ts",
];
ev.harnessDigest = createHash("sha256")
.update(
(
await Promise.all(
harnessFiles.map((f) =>
readFile(path.join(import.meta.dirname, f)),
),
)
)
.map((b) => b.toString())
.join("\n"),
)
.digest("hex");
if (requiresEverydayArtifactOracle(caseId)) {
try {
const sandbox = await runCommand(process.env.PYTHON ?? "python3", [
path.join(import.meta.dirname, "everyday-artifact.py"), "--preflight",
]);
if (sandbox.code !== 0) throw new Error(sandbox.stdout);
note("artifact-sandbox-qualified", { isolation: "docker", network: "none" });
} catch (error) {
throw new Error(`Artifact sandbox qualification failed before task creation: ${error instanceof Error ? error.message : String(error)}`, { cause: error });
}
}
if (!project && !caseId.startsWith("service-") && !decliningConnection && !providerChoice) {
project = await api.post(
`/api/companies/${fixtures.company.id}/projects`,
{
name: `Studio project ${nonce}`,
description: "Small software project",
executionWorkspacePolicy: {
enabled: true,
defaultMode: "shared_workspace",
sharedWorkspaceConcurrency: "serialize",
allowIssueOverride: false,
environmentId: fixtures.environment.id,
workspaceStrategy: { type: "project_primary" },
},
workspace: {
name: "Primary",
sourceType: "local_path",
cwd: input.workspacePath,
isPrimary: true,
},
},
);
}
await api.patch(`/api/agents/${fixtures.agent.id}/permissions`, {
canCreateAgents: true,
canAssignTasks: true,
});
if (caseId === "delegate-feedback" || caseId === "agent-review-handoff") {
const config = execution.profile.buildAgent({
environmentId: fixtures.environment.id,
environmentFixtureId: execution.environment.id,
workspacePath: input.workspacePath,
secretRefs: fixtures.secretRefs,
executionId: nonce,
});
await api.post(`/api/companies/${fixtures.company.id}/agents`, {
...config,
name: "Riley Builder",
role: "engineer",
title: "Engineer",
reportsTo: fixtures.agent.id,
instructionsBundle: {
entryFile: "AGENTS.md",
files: {
"AGENTS.md":
"You implement small software projects and verify your work. Follow task feedback and deliver usable files.",
},
},
});
}
if (caseId.startsWith("service-"))
review = await setupConnectionReview({
page,
api,
prefix,
companyId: fixtures.company.id,
agentId: fixtures.agent.id,
marker: `Pages: Roadmap, Meeting notes. Verification code: SERVICE_${nonce}`,
authenticated: true,
});
if (providerChoice || nativeProviderCase) {
if (caseId === "provider-second") aggregatorFixture = await setupAggregatorFixture(api, fixtures.company.id, fixtures.agent.id, `CONTACTS_${nonce}`);
const state = await api.get<{connections:Row[]}>(`/api/companies/${fixtures.company.id}/tools/connections`);
initialConnections = state.connections.map(c=>c.id);
}
if (agentmailSetup) await api.patch("/api/instance/settings/experimental", { enableChatConnectors: true });
if (decliningConnection) {
const state = await api.get<{ connections: Row[] }>(
`/api/companies/${fixtures.company.id}/tools/connections`,
);
initialConnections = state.connections.map((c) => c.id);
check(
"connection-starts-unconfigured",
state.connections.length === 0,
"This isolated company has no service connection before the request.",
);
if (state.connections.length)
throw new Error("New-connection story requires an unconnected company");
}
await createTaskThroughUi({
page,
issuePrefix: prefix,
agentName: fixtures.agent.name,
title: execution.task.buildTitle(nonce),
prompt: ev.prompt,
workMode: "standard",
projectName: project?.name,
});
parent = await pollUntil({
label: "browser-created story task",
deadlineAt: Date.now() + 30_000,
load: () =>
api.get<StoryIssue[]>(`/api/companies/${fixtures.company.id}/issues`),
accept: (rows) =>
rows.some((i) => i.title === execution.task.buildTitle(nonce)),
}).then((rows) =>
rows.find((i) => i.title === execution.task.buildTitle(nonce))!,
);
input.observe(parent!, []);
note("task-submitted", { issueId: parent!.id });
await openParent();
if (caseId === "agent-review-handoff") {
const boundary = await pollUntil({
label: "blocked parent with agent review wake",
deadlineAt: input.deadlineAt,
load: refresh,
accept: (state) => {
const child = state.issues.find((issue) => issue.parentId === parent!.id);
const interaction = child?.interactions?.find(
(candidate) =>
["pending", "accepted"].includes(String(candidate.status)) &&
candidate.addresseeAgentId === fixtures.agent.id &&
candidate.effectiveResolverPolicy !== "human_only",
);
const parentIssue = state.issues.find((issue) => issue.id === parent!.id);
const reviewRun = interaction?.resolvedByRunId
? state.runs.find((run) => run.id === interaction.resolvedByRunId)
: undefined;
const parentRun = state.runs.find(
(run) =>
(run.nativeIssueId === parent!.id ||
run.contextSnapshot?.issueId === parent!.id ||
run.contextSnapshot?.taskId === parent!.id) &&
run.status === "succeeded" &&
run.finishedAt &&
(reviewRun?.startedAt
? Date.parse(run.finishedAt) <= Date.parse(reviewRun.startedAt)
: parentIssue?.status === "blocked"),
);
const parentBlockedEvidence =
Boolean(
parentIssue && storyIssueHasUnresolvedDependency(parentIssue),
) ||
Boolean(
parentIssue &&
storyIssueHasBlockedTimelineBefore(
parentIssue,
reviewRun?.startedAt ?? undefined,
),
) ||
Boolean(parentRun && storyRunReportsDependencyBlock(parentRun));
return Boolean(
parentIssue &&
["blocked", "done"].includes(parentIssue.status) &&
["in_review", "done"].includes(String(child?.status)) &&
interaction &&
parentBlockedEvidence &&
parentRun &&
(parentIssue.status === "blocked" || reviewRun),
);
},
reject: (state) => {
const failed = state.runs.find((run) =>
["failed", "timed_out"].includes(run.status),
);
return failed
? `Review handoff prerequisite failed: ${failed.errorCode}: ${failed.error}`
: storyUnexercisedReviewBoundary(state.issues, state.runs, parent!.id, fixtures.agent.id);
},
});
const child = boundary.issues.find((issue) => issue.parentId === parent!.id)!;
const interaction = child.interactions!.find(
(candidate) =>
["pending", "accepted"].includes(String(candidate.status)) &&
candidate.addresseeAgentId === fixtures.agent.id,
)!;
const reviewRun = interaction.resolvedByRunId
? boundary.runs.find((run) => run.id === interaction.resolvedByRunId)
: undefined;
const parentAtBoundary = boundary.issues.find(
(issue) => issue.id === parent!.id,
);
const parentRun = boundary.runs.find(
(run) =>
(run.nativeIssueId === parent!.id ||
run.contextSnapshot?.issueId === parent!.id ||
run.contextSnapshot?.taskId === parent!.id) &&
run.status === "succeeded" &&
run.finishedAt &&
storyParentCompletionPrecedesReview(
run.finishedAt,
reviewRun?.startedAt,
parentAtBoundary?.status === "blocked",
),
);
reviewHandoffBoundary = {
childId: child.id,
assigneeAgentId: child.assigneeAgentId,
interactionId: interaction.id!,
interactionCreatedAt: interaction.createdAt,
parentRunId: parentRun?.id,
parentBlockedObserved: boundary.issues.some(
(issue) =>
issue.id === parent!.id && storyIssueHasUnresolvedDependency(issue),
),
};
check(
"parent-blocked-before-agent-review",
reviewHandoffBoundary.parentBlockedObserved || Boolean(parentRun),
"The completed lead run and durable dependency projection precede the child review card.",
);
note("agent-review-requested", {
childId: child.id,
childAssigneeAgentId: child.assigneeAgentId,
interactionId: interaction.id,
parentStatus: boundary.issues.find((issue) => issue.id === parent!.id)?.status,
parentRunId: parentRun?.id,
parentRunFinishedAt: parentRun?.finishedAt,
interactionCreatedAt: interaction.createdAt,
});
await openParent();
}
if (caseId === "delegate-feedback") {
await pollUntil({
label: "active delegated child",
deadlineAt: input.deadlineAt,
load: refresh,
accept: (state) =>
state.issues.some(
(i) =>
i.parentId === parent!.id &&
state.runs.some(
(r) =>
isActiveStoryRun(r) &&
(r.contextSnapshot?.issueId === i.id ||
r.contextSnapshot?.taskId === i.id),
),
),
reject: (state) => {
const failed = state.runs.find((run) =>
["failed", "timed_out"].includes(run.status),
);
return failed
? `Delegation prerequisite failed before feedback: ${failed.errorCode}: ${failed.error}`
: undefined;
},
});
const child = ev.issues.find((i) => i.parentId === parent!.id)!;
note("late-feedback-boundary", {
childId: child.id,
activeRunIds: ev.runs.filter(isActiveStoryRun).map((r) => r.id),
});
await openTask(child);
await reply(LATE_REQUIREMENT, child);
note("late-feedback-delivered-to-child", { childId: child.id });
await openParent();
}
if (
caseId === "recover-controller" ||
caseId === "stop-redirect"
) {
if (execution.environment.id === "daytona") {
// A completed, downloaded first version is an observable remote persistence checkpoint.
await settled();
await download(parent!.id, "base", "before-restart");
if (!ev.downloads.length || ev.checks.some((c) => !c.passed))
throw new Error(
"Remote fault boundary unexercised: no verified saved project",
);
await reply(SLUGIFY_REVISION);
note("remote-revision-submitted");
await pollUntil({
label: "remote revision executing",
deadlineAt: input.deadlineAt,
load: refresh,
accept: (s) => s.runs.some(isActiveStoryRun),
});
} else if (caseId === "stop-redirect") await prepareStopBoundary();
else await sourceReady();
const active = ev.runs.find((r) => r.status === "running");
if (!active)
throw new Error(
"Fault boundary was not exercised: no active execution",
);
if (caseId === "stop-redirect") {
await page.getByTestId("task-chat-composer-stop").last().click();
ev.allowedInterruptedRuns.push(active.id);
note("stop-clicked", { runId: active.id });
await pollUntil({
label: "owned runner stopped",
deadlineAt: input.deadlineAt,
load: () => runnerStopped(active),
accept: Boolean,
intervalMs: 250,
});
await recordSource("source-saved-after-interruption", "source-after-interruption.json");
stoppedWorkspace = await workspaceFiles();
note("stopped-workspace-snapshot", stoppedWorkspace);
await reply(
`Change direction. Leave the project as it is. Reply with just this short note: "The studio is ready. Reference ${nonce}."`,
);
note("new-direction-submitted");
await page.reload();
} else {
await reply(LATE_REQUIREMENT);
note("followup-submitted-before-interruption");
ev.allowedInterruptedRuns.push(active.id);
await input.restart();
note("controller-restarted");
await openParent();
}
}
if (providerChoice) {
const rows = await pollUntil({
label: "external-provider choice", deadlineAt: input.deadlineAt,
load: async () => {
const runs = await api.get<StoryRun[]>(`/api/issues/${parent!.id}/runs`);
const failure = runs.find(run => ["failed", "timed_out"].includes(run.status));
if (failure) throw new Error(`Stopped waiting for external-provider choice: agent failed before selection: ${failure.error ?? failure.status}`);
const rows = await api.get<Row[]>(`/api/issues/${parent!.id}/interactions`);
const issue = await api.get<StoryIssue>(`/api/issues/${parent!.id}`);
if (!rows.some(row=>row.status==="pending") && ["done", "blocked", "cancelled"].includes(issue.status)) throw new Error(`Stopped waiting for external-provider choice: task reached ${issue.status} without asking the user`);
return rows;
},
accept: rows => rows.some(row=>row.status==="pending"),
});
const decision = gradeProviderChoice(rows as any, aggregatorFixture?.invocationCount() ?? 0);
decisionId = decision.interaction.id;
check("provider-disclosed-before-choice", true, "Ranked external providers and None were offered before any call.");
// Exercise durable selection across a real controller restart and browser reload.
await input.restart();
await openParent();
const choice = page.getByRole("radio", {name: caseId === "provider-decline" ? /None for now/ : /^Arcade/});
await expect(choice).toBeVisible();
await input.capture("provider-choice", "External service choice after restart", "provider-choice.png");
await choice.click();
await page.getByRole("button", {name:"Submit answers",exact:true}).click();
await pollUntil({label:"provider choice saved", deadlineAt:input.deadlineAt,
load:()=>api.get<Row[]>(`/api/issues/${parent!.id}/interactions`),
accept:rows=>rows.some(row=>row.id===decisionId && row.status==="answered"),
});
note("provider-choice-submitted", {interactionId:decisionId, selected:caseId === "provider-decline" ? "none" : "via:arcade:hubspot"});
}
if (review || decliningConnection) {
const interactions = await pollUntil({
label: "story decision request",
deadlineAt: input.deadlineAt,
load: async () => {
const [interactions, issue] = await Promise.all([
api.get<StoryInteraction[]>(
`/api/issues/${parent!.id}/interactions`,
),
api.get<StoryIssue>(`/api/issues/${parent!.id}`),
]);
return { interactions, issue, calls: review?.invocationCount() ?? 0 };
},
accept: (state) =>
state.interactions.some((i) => i.status === "pending"),
reject: (state) =>
state.calls > 0
? `The provider received ${state.calls} call(s) before approval.`
: ["done", "blocked", "cancelled"].includes(state.issue.status)
? `Task reached ${state.issue.status} without requesting the expected decision.`
: undefined,
});
await openParent();
await expect(
page
.locator(
'[data-testid="task-chat-thread"], [data-testid="thread-root"]',
)
.first(),
).toBeVisible();
await expect(page.getByTestId("issue-chat-skeleton")).toHaveCount(0);
const pendingInteraction = pendingStoryDecision(
interactions.interactions,
review
? { kind: "tool", connectionId: review.connectionId }
: { kind: "connection", serviceSlug: agentmailSetup ? "agentmail" : nativeProviderCase ? "jira" : "notion" },
);
await expect(
page.getByRole("button", {
name: decliningConnection ? "Not now" : "Review request",
exact: true,
}),
).toBeVisible();
await input.capture(
"decision-pending",
"Request before the user decision",
"decision-pending.png",
);
decisionId = pendingInteraction.id;
check(
"decision-request-matches-story",
true,
review
? "Tool approval belongs to the installed page service."
: `New connection request is for ${agentmailSetup ? "AgentMail" : nativeProviderCase ? "Jira" : "Notion"}, without an external-provider question.`,
);
if (agentmailSetup) {
// Reload proves the card is durable, rather than a transient model UI.
await page.reload();
const form = page.getByTestId("agentmail-inline-setup");
await expect(form).toBeVisible();
const card = page.getByTestId("connection-intent-focus-target").filter({ has: form });
const rows = await api.get<Parameters<typeof gradeAgentmailSetup>[0]["interactions"]>(`/api/issues/${parent!.id}/interactions`);
const checks = gradeAgentmailSetup({
// This harness boots a dedicated local_trusted instance.
interactions: rows, agentId: fixtures.agent.id, userId: "local-board",
visible: await form.isVisible(),
inputTypes: await card.locator("input").evaluateAll(inputs => inputs.map(input => input.getAttribute("type") ?? "text")),
keyLink: await form.getByRole("link", { name: "Get an AgentMail API key" }).getAttribute("href"),
accessSelectorCount: await card.locator('[role="radiogroup"], [role="combobox"], select').count(),
dialogCount: await page.getByRole("dialog").count(),
});
ev.checks.push(...checks);
await input.capture("agentmail-inline-key", "AgentMail API-key card after reload", "agentmail-inline-key.png");
if (checks.some(result => !result.passed)) throw new Error("AgentMail inline setup evidence failed");
}
if (review)
check(
"no-call-before-approval",
review.invocationCount() === 0,
"Service must not execute before the user decides.",
);
if (decliningConnection) {
await page
.getByRole("button", { name: "Not now", exact: true })
.click();
} else {
const dismiss = page.getByRole("button", {
name: "Dismiss Approve tool action",
});
if (await dismiss.isVisible()) await dismiss.click();
await page
.getByRole("button", { name: "Review request", exact: true })
.click();
await page
.getByRole("button", {
name: declining ? "Decline" : "Approve & run",
exact: true,
})
.click();
}
const decisionStatus = declining ? "rejected" : "accepted";
const decided = await pollUntil({
label: "connection decision persisted",
deadlineAt: Math.min(input.deadlineAt, Date.now() + 30_000),
load: () => api.get<Row[]>(`/api/issues/${parent!.id}/interactions`),
accept: (rows) =>
rows.some(
(interaction) =>
interaction.id === pendingInteraction.id &&
interaction.status === decisionStatus,
),
});
decisionResolvedAt = decided.find((i) => i.id === decisionId)?.resolvedAt;
check(
"user-decision-persisted",
true,
`Interaction ${decisionId} saved as ${decisionStatus}.`,
);
note("connection-decision", {
decision: caseId,
interactionId: pendingInteraction.id,
status: decisionStatus,
});
}
await settled(
caseId === "stop-redirect" ? `Reference ${nonce}`
: caseId === "provider-second" ? `CONTACTS_${nonce}` : undefined,
);
if (caseId === "create-skill-studio") {
const createdSkills = await api.get<Row[]>(
`/api/companies/${fixtures.company.id}/skills`,
);
const created = createdSkills.find(
(skill) => skill.slug === "release-readiness-checklist",
);
check(
"skill-persisted",
Boolean(created && String(created.name) === "release-readiness-checklist"),
"The runner-created skill is present in the company library after the run.",
);
if (created) {
await openParent();
const card = page.getByRole("article", {
name: "Skill created: release-readiness-checklist",
});
await expect(card).toHaveCount(1);
await expect(card).toBeVisible();
await card.getByRole("button").click();
await expect(page.getByRole("heading", { name: "release-readiness-checklist" })).toBeVisible();
await expect(page.getByText("Verify checks.", { exact: true })).toBeVisible();
check("feed-card-opened", true, "The task thread card opened the created skill sidebar.");
const openStudio = page.getByRole("button", { name: "Open in Skill Studio", exact: true });
await expect(openStudio).toBeVisible();
await openStudio.click();
await expect(page).toHaveURL(new RegExp(`/skills/studio/${created.id}$`));
// Studio's skill selector identifies the resource. Headings inside the
// authored document can differ from its canonical skill name.
await expect(page.getByRole("combobox").filter({ hasText: "release-readiness-checklist" })).toBeVisible();
const editor = page.getByRole("textbox", { name: "editable markdown", exact: true });
await editor.click();
await editor.press("ControlOrMeta+End");
await editor.press("Enter");
await editor.press("Enter");
await editor.pressSequentially("Studio edit marker: verified");
await page.getByRole("button", { name: /^Save$/ }).click();
await pollUntil({
label: "Skill Studio edit persisted",
deadlineAt: Math.min(input.deadlineAt, Date.now() + 30_000),
load: () => api.get<Row>(
`/api/companies/${fixtures.company.id}/skills/${encodeURIComponent(String(created.id))}`,
),
accept: (detail) => JSON.stringify(detail).includes("Studio edit marker: verified"),
});
const detail = await api.get<Row>(
`/api/companies/${fixtures.company.id}/skills/${encodeURIComponent(String(created.id))}`,
);
check(
"studio-edit-persisted",
JSON.stringify(detail).includes("Studio edit marker: verified"),
"The Skill Studio edit remains in the persisted skill after returning to the page.",
);
check("studio-opened", true, "The skill detail opened in Skill Studio.");
await page.goBack();
await expect(page).toHaveURL(new RegExp(`/issues/`));
const returnedCard = page.getByRole("article", {
name: "Skill created: release-readiness-checklist",
});
await expect(returnedCard).toBeVisible();
await returnedCard.getByRole("button").click();
await expect(page.getByRole("heading", { name: "release-readiness-checklist" })).toBeVisible();
await expect(page.getByText("Studio edit marker: verified", { exact: true })).toBeVisible();
check("return-content-persisted", true, "Returning to the task shows the saved Skill Studio edit.");
}
}
if (providerChoice) {
const issue = ev.issues.find(i=>i.id===parent!.id)!;
const state = await api.get<{connections:Row[]}>(`/api/companies/${fixtures.company.id}/tools/connections`);
ev.checks.push(...gradeProviderOutcome({rows:issue.interactions as any, decisionId:decisionId!,
selected:caseId === "provider-decline" ? "none" : "via:arcade:hubspot", calls:aggregatorFixture?.invocationCount() ?? 0,
response: (issue.comments ?? []).filter((c:Row)=>c.authorAgentId).map((c:Row)=>c.body).join("\n"), marker:`CONTACTS_${nonce}`,
sameConnections:isDeepStrictEqual(state.connections.map(c=>c.id).sort(), initialConnections.sort()),
}));
}
if (declining) {
const issue = ev.issues.find((i) => i.id === parent!.id)!;
const requests = issue.interactions as Row[];
check(
"decline-not-repeated",
requests.length === 1 &&
requests[0]?.id === decisionId &&
requests[0]?.status === "rejected",
"The saved decline remains rejected and no replacement request appears.",
);
const replies = (issue.comments ?? []).filter(
(c: Row) =>
c.authorAgentId &&
Date.parse(c.createdAt) >= Date.parse(decisionResolvedAt ?? ""),
);
const text = replies.map((c: Row) => String(c.body ?? "")).join("\n");
check(
"decline-visible-explanation",
replies.length > 0 &&
/declin|not now|could(?:n.t| not)|cannot|can.t|unable|not (?:connect|retriev)|without (?:access|connect)/i.test(
text,
),
"A new agent response explains the missing access after the saved decline.",
);
check(
"decline-no-fabricated-result",
!text.includes(`SERVICE_${nonce}`),
"The fallback does not claim the private verification code.",
);
if (decliningConnection) {
const state = await api.get<{ connections: Row[] }>(
`/api/companies/${fixtures.company.id}/tools/connections`,
);
check(
"decline-no-connection-created",
isDeepStrictEqual(
state.connections.map((c) => c.id).sort(),
initialConnections.sort(),
),
"Not now did not create a service connection.",
);
}
}
if (caseId === "build-revise") {
await download(parent!.id, "base", "initial");
await reply(SLUGIFY_REVISION);
note("revision-requested");
await settled();
await download(parent!.id, "separator", "revised");
check(
"new-artifact-revision",
ev.downloads.length === 2 &&
ev.downloads[0]!.sha256 !== ev.downloads[1]!.sha256,
"The follow-up must deliver a new version.",
);
if (ev.downloads[0]) {
const response = await api.request.get(
`/api/attachments/${ev.downloads[0].attachmentId}/content`,
);
check(
"prior-download-preserved",
response.ok() &&
createHash("sha256")
.update(await response.body())
.digest("hex") === ev.downloads[0].sha256,
"The first delivered version remains retrievable.",
);
}
} else if (caseId === "hire-reuse") {
let agents = await api.get<Row[]>(
`/api/companies/${fixtures.company.id}/agents`,
);
const hires = agents.filter((a) => a.name === "Morgan QA");
check(
"exactly-one-hire",
hires.length === 1,
"One Morgan QA must be hired.",
);
check(
"manager-correct",
hires.length === 1 && hires[0]!.reportsTo === fixtures.agent.id,
"The hired agent reports to the lead.",
);
const lead = agents.find((a) => a.id === fixtures.agent.id);
check(
"hire-native-connection",
hires.length === 1 &&
hires[0]!.adapterType === "paperclip_runner" &&
hires[0]!.adapterConfig?.model === lead?.adapterConfig?.model &&
isDeepStrictEqual(
hires[0]!.runtimeConfig?.aiConnection,
fixtures.aiConnection?.binding,
),
"The hire keeps the native model and inherits the managed AI account binding.",
);
const children = ev.issues.filter((i) => i.parentId === parent!.id);
check(
"hired-agent-executed",
hires.length === 1 &&
ev.runs.some(
(r) =>
r.agentId === hires[0]!.id &&
r.status === "succeeded" &&
(r.contextSnapshot?.aiConnection as Row | undefined)
?.connectionId === fixtures.aiConnection?.connectionId,
),
"The new hire must complete a run using the fixture managed account.",
);
await downloadDelegated("base", "hired-delivery");
const reuseRequestedAt = Date.now();
await reply(
`Have the existing Morgan QA add --separator support to the delivered project. Use the same agent; do not hire another. ${SLUGIFY_REVISION}`,
);
note("reuse-requested");
await settled();
agents = await api.get<Row[]>(
`/api/companies/${fixtures.company.id}/agents`,
);
check(
"hire-reused",
agents.filter((a) => a.name === "Morgan QA").length === 1 &&
agents.some((a) => a.id === hires[0]?.id),
"The original hired identity remains unique.",
);
check(
"hired-agent-executed-revision",
ev.runs.some(
(r) =>
r.agentId === hires[0]?.id &&
r.status === "succeeded" &&
Date.parse(r.finishedAt ?? "") >= reuseRequestedAt,
),
"The same hired agent performs the follow-up work.",
);
await downloadDelegated("separator", "reused-delivery", reuseRequestedAt);
} else if (caseId === "delegate-feedback") {
const children = ev.issues.filter((i) => i.parentId === parent!.id);
check(
"one-child",
children.length === 1,
"Exactly one delegated child task.",
);
check(
"child-consumed-feedback",
Boolean(children[0]) &&
storyRepliesConsumed(
ev.runs.filter(
(run) =>
run.contextSnapshot?.issueId === children[0]?.id ||
run.contextSnapshot?.taskId === children[0]?.id,
),
submittedCommentIds,
),
"A completed child execution consumed the delivered user feedback.",
);
if (children[0])
await downloadDelegated(
artifactGradeModeForPhase("delegated-delivery"),
"delegated-delivery",
);
check(
"feedback-delivered-to-child",
Boolean(
children[0]?.comments?.some((c: Row) =>
String(c.body).includes("--max-length"),
),
),
"The child history contains the late requirement.",
);
} else if (caseId === "agent-review-handoff") {
const children = ev.issues.filter((i) => i.parentId === parent!.id);
const child = children.find((candidate) =>
candidate.id === reviewHandoffBoundary?.childId,
);
const interaction = storyAcceptedAgentReview(
child,
reviewHandoffBoundary?.interactionId,
fixtures.agent.id,
ev.runs,
) as Row | undefined;
const reviewRun = interaction?.resolvedByRunId
? ev.runs.find((run) => run.id === interaction.resolvedByRunId)
: undefined;
const parentContinuationRun = ev.runs.find(
(run) =>
run.agentId === fixtures.agent.id &&
run.status === "succeeded" &&
run.id !== reviewHandoffBoundary?.parentRunId &&
(run.nativeIssueId === parent!.id ||
run.contextSnapshot?.issueId === parent!.id ||
run.contextSnapshot?.taskId === parent!.id) &&
interaction?.resolvedAt &&
run.startedAt &&
Date.parse(run.startedAt) >= Date.parse(interaction.resolvedAt),
);
check(
"agent-review-one-child",
children.length === 1 && Boolean(child),
"Exactly one child remains attached to the lead task.",
);
check(
"agent-review-child-completed",
child?.status === "done" && interaction?.status === "accepted",
"The named agent review is accepted and the child reaches Done.",
);
check(
"agent-review-assignee-preserved",
child?.assigneeAgentId === reviewHandoffBoundary?.assigneeAgentId,
"Review resolution preserves the child task assignee.",
);
check(
"agent-review-run-scoped",
reviewRun?.status === "succeeded" &&
reviewRun.agentId === fixtures.agent.id &&
reviewRun.contextSnapshot?.nativeReviewInteractionId ===
interaction?.id &&
typeof reviewRun.contextSnapshot?.nativeReviewDecisionId === "string",
"The lead resolves the review from a successful review-scoped native run.",
);
check(
"agent-review-parent-resumed",
parent?.status === "done" &&
Boolean(parentContinuationRun) &&
Boolean(
parentContinuationRun?.finishedAt &&
parentContinuationRun.startedAt &&
Date.parse(parentContinuationRun.finishedAt) >=
Date.parse(parentContinuationRun.startedAt),
),
"The parent continuation starts after review acceptance and finishes successfully.",
);
note("agent-review-accepted", {
childId: child?.id,
childAssigneeAgentId: child?.assigneeAgentId,
interactionId: interaction?.id,
interactionStatus: interaction?.status,
resolvedByRunId: interaction?.resolvedByRunId,
reviewRunId: reviewRun?.id,
reviewRunAgentId: reviewRun?.agentId,
parentContinuationRunId: parentContinuationRun?.id,
parentContinuationStartedAt: parentContinuationRun?.startedAt,
nativeReviewInteractionId:
reviewRun?.contextSnapshot?.nativeReviewInteractionId,
nativeReviewDecisionId:
reviewRun?.contextSnapshot?.nativeReviewDecisionId,
});
// The review handoff case uses the base slugify requirements. The
// max-length grader belongs to the later follow-up requirement cases.
if (child)
await downloadDelegated(
artifactGradeModeForPhase("reviewed-delivery"),
"reviewed-delivery",
);
} else if (caseId === "recover-controller")
await download(
parent!.id,
artifactGradeModeForPhase("recovered-delivery"),
"recovered-delivery",
);
else if (caseId === "stop-redirect")
check(
"new-direction-delivered",
settledAgentReply?.authorAgentId === fixtures.agent.id &&
String(settledAgentReply.body ?? "").includes(`Reference ${nonce}`),
"The new request is answered after Stop and reload.",
);
if (stoppedWorkspace) {
check(
"stop.workspace-unchanged",
isDeepStrictEqual(stoppedWorkspace, await workspaceFiles()),
"Project files remain unchanged from verified runner stop through the new response.",
);
check(
"stop.no-old-run-active",
ev.runs
.filter((r) => ev.allowedInterruptedRuns.includes(r.id))
.every((r) => !isActiveStoryRun(r)),
"The stopped run is terminal after the new response.",
);
}
if (review) {
check(
"service-call-count",
review.invocationCount() === (caseId === "service-decline" ? 0 : 1),
"Exactly one approved service call; none after decline.",
);
if (caseId === "service-approve") {
const docs = await api.get<Row[]>(
`/api/issues/${parent!.id}/documents`,
);
const bodies = await Promise.all(
docs.map((d) =>
api.get<Row>(
`/api/issues/${parent!.id}/documents/${encodeURIComponent(d.key)}`,
),
),
);
const attachments = await api.get<Row[]>(
`/api/issues/${parent!.id}/attachments`,
);
for (const attachment of attachments.filter(
(a) =>
/\.(?:md|txt)$/i.test(a.originalFilename ?? a.filename ?? "") ||
["text/markdown", "text/plain"].includes(a.contentType),
)) {
const url = `/api/attachments/${attachment.id}/content`;
const response = await api.request.get(url);
if (!response.ok()) continue;
await expect(page.locator(`a[href*="${url}"]`).first()).toBeVisible();
bodies.push({
id: attachment.id,
title: attachment.originalFilename ?? attachment.filename,
body: await response.text(),
source: "delivered-attachment",
});
}
const text = bodies
.map((d) => String(d.body ?? d.revision?.body ?? ""))
.join("\n");
check(
"briefing-uses-real-result",
text.includes(`SERVICE_${nonce}`),
"A delivered issue document or Markdown attachment contains the actual service verification code.",
);
check(
"briefing-includes-page-titles",
text.includes("Roadmap") && text.includes("Meeting notes"),
"The briefing includes both page titles returned by the service.",
);
ev.documents = bodies;
ev.issues.find((i) => i.id === parent!.id)!.documents = bodies;
}
}
await page.reload();
await refresh();
check(
"submitted-replies-consumed",
storyRepliesConsumed(ev.runs, submittedCommentIds),
"Every submitted user message appears in a successfully completed native execution input.",
);
if (
caseId === "delegate-feedback" ||
caseId === "hire-reuse" ||
caseId === "agent-review-handoff"
) {
check(
"parent-finishes-after-child",
storyParentFinishedAfterChildren(
ev.runs,
parent!.id,
fixtures.agent.id,
ev.issues.filter((i) => i.parentId === parent!.id).map((i) => i.id),
),
"The lead completes only after the final child execution.",
);
}
const expectedModel = execution.profile.model;
check(
"native-model-config",
ev.runs.length > 0 &&
ev.runs
.filter((r) => !isStoryWorkspaceDeferral(r))
.every(
(r) =>
(r.runnerProfileJson?.nativeExecutionInput as Row | undefined)
?.provider?.model === expectedModel,
),
"Persisted native execution inputs use the selected model; this does not claim provider-side model identity.",
);
check(
"native-terminal-contract",
ev.runs
.filter(
(r) =>
!isStoryWorkspaceDeferral(r) &&
!isExpectedStoryInterruption(r, ev.allowedInterruptedRuns),
)
.every(
(r) =>
(r.resultJson?.nativeTerminal as Row | undefined)?.schema ===
"paperclip.prp.terminal.v1",
),
"Successful runs retain the native terminal contract.",
);
ev.checks.push(
...storyLifecycleChecks({
issues: ev.issues as StoryIssue[],
runs: ev.runs,
parentId: parent!.id,
leadId: fixtures.agent.id,
allowedInterruptedRuns: ev.allowedInterruptedRuns,
}),
);
check(
"no-pending-bookkeeping",
ev.issues.every((i) =>
(i.interactions ?? []).every((x: Row) => x.status !== "pending"),
),
"No completion confirmation or unanswered interaction remains.",
);
if (providerChoice || nativeProviderCase) await openParent();
else await waitForTaskChatRendered(page, String(parent!.title));
const latestAgentComment = ev.issues
.find((i) => i.id === parent!.id)
?.comments?.filter((c: Row) => c.authorAgentId)
.at(-1);
if (latestAgentComment && !providerChoice && !nativeProviderCase) {
const response = page.locator(`[id="comment-${latestAgentComment.id}"]`);
await expect(response).toBeVisible();
await response.scrollIntoViewIfNeeded();
}
await input.capture(
"final-state",
"Finished everyday workflow",
"final-state.png",
);
const failed = ev.checks.filter((c) => !c.passed);
if (failed.length)
throw new Error(
`Everyday outcome checks failed: ${failed.map((c) => c.id).join(", ")}`,
);
return { issue: parent!, runs: ev.runs, evidence: ev };
} catch (error) {
check(
error instanceof StoryDecisionError
? error.checkId
: "workflow-completed",
false,
error instanceof Error ? error.message : String(error),
);
throw error;
} finally {
try {
await refresh();
ev.agents = await api.get<Row[]>(
`/api/companies/${fixtures.company.id}/agents`,
);
if (ev.documents && parent)
ev.issues.find((i) => i.id === parent!.id)!.documents = ev.documents;
} catch (error) {
note("evidence-capture-error", String(error));
}
if (caseId === "recover-controller") {
note("recovery-final-observation", {
taskStatus: parent?.status,
pendingCommentIds: ev.issues.flatMap((i) =>
(i.queuedComments?.entries ?? []).map((e: Row) => e.comment.id),
),
failureCodes: ev.runs
.filter((r) => r.errorCode)
.map((r) => ({ runId: r.id, code: r.errorCode })),
});
if (execution.environment.id === "local") {
try {
const source = await readFile(
path.join(input.workspacePath, "slugify.py"),
"utf8",
);
await input.evidence("source-after-interruption.json", {
body: source,
sha256: createHash("sha256").update(source).digest("hex"),
});
} catch {
note("saved-source-unavailable-at-final-capture");
}
}
}
if (review) {
note("service-final-observation", {
invocationCount: review.invocationCount(),
requests: review.captures,
decisionTaken: ev.timeline.some(
(entry) => entry.action === "connection-decision",
),
});
}
try {
await input.evidence("everyday-workflow.json", ev);
await input.evidence("api-state.json", {
capturePhase: "everyday-final", issue: parent, runs: ev.runs,
issues: ev.issues, checks: ev.checks,
});
if (aggregatorFixture) await input.evidence("aggregator-provider-calls.json", { calls: aggregatorFixture.captures });
} finally {
try { await aggregatorFixture?.close(); }
finally { await review?.close(); }
}
}
}