mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users guide running agents through messages, questions, and approval cards. > - Messages already wait in a queue when an agent is running. > - Card responses did not appear in that queue. Some question answers also steered a later run without a user click. > - A fast approval could invalidate the agent's review handoff and cause it to stop its own run. > - This pull request gives card responses the same queue controls and preserves the exact response during delivery. > - Users can wait for completion or explicitly send the response with Interrupt or Steer. ## Linked Issues or Issue Description Refs #13517, which is merged. This PR targets master and adds queued interaction responses on top of the onboarding changes. Related continuation work: #10519 and #12866. **What happened?** Accepting a proposal while its source run was active left a saved response outside the message queue. The agent could then lose its review path, reassign the task, and cancel itself. Answers to older questions could also steer another active turn without a click. **Expected behavior** Save the response immediately. Queue its continuation behind the active run. Deliver it after completion, or when the user explicitly chooses Interrupt or Steer. Preserve approval revisions and answer choices. **Steps to reproduce** 1. Let an agent publish a confirmation card while its run is still active. 2. Accept the card before the agent finishes its review handoff. 3. Inspect the message queue and the task's next run. **Paperclip version or commit** Reproduced on da8a3876c with the onboarding changes from #13517. **Deployment mode** Local development from source. The fix covers legacy adapters and native Runner turns. ## What Changed - Project resolved cards into the existing queue as immutable responses. Keep answers and exact approval revisions. - Require an explicit click to steer a response into a compatible native turn. Use Interrupt when a fresh session is required. - Preserve typed response context through interruption, cleanup waits, and normal queue promotion. Keep the direct answer channel for a provider blocked on its original question request. - Accept the source run's review handoff after its card resolves. Reject stale agent reassignment that would orphan a queued response. - Add deterministic regression tests and an `accept-while-running` case to the first-task suite. Require recorded timestamp overlap before that case can pass. - Keep the first-task skill name out of user-facing messages. ## Verification - Red-green: the original route failed the queue regression; the changed route passes it. - Focused server/UI tests: 139 passed, including 64 queue-route tests. - Runner harness unit tests: 314 passed. - Server, UI, and Runner E2E typechecks passed. UI token gates passed. - Full repository typecheck and build passed. Server typecheck passed again after review fixes. - Review regressions: 165 queue/reopen route tests, 53 wake admission tests, and 18 run identity tests passed. Approval acknowledgement recovery and both message/approval arrival orders are covered. - Full local test run: 12,401 passed; three new admission regressions ran against a cached pre-fix module. A fresh run of that entire suite passed (53 tests). The complete CI suite passed on the final commit. - Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional Storybook checks skipped. Every server/workspace/browser shard, Runner verification, build, typecheck/release registry, canary, policy, and security check passed. Greptile: 5/5, no unresolved threads. Earlier interrupted CI workers were replaced by this fresh complete run. - After integrating the updated parent: 314 harness tests, 119 queue/admission tests, 44 onboarding/question-delivery tests, and 13 native recovery tests passed locally. Full repository typecheck and build passed. - Clarified the skill wording preference: routine replies describe the action without announcing the internal skill; direct questions and permission/security/execution disclosures remain truthful. - The paid `accept-while-running` scenario is registered for all four local first-task profiles. It has not been run against a model in this change. - Rebased onto the merged parent at `11921075a`; the resulting tree exactly matches the locally verified integration tree. Final-head CI on `b53054807` passed: 54 successful checks, 2 optional Storybook checks skipped, no failed checks. Every new server/browser shard, aggregate verify/e2e gate, Runner, typecheck, build, canary, and security check passed on the first attempt. Greptile reviewed this exact head at 5/5 with no unresolved threads. ## Risks - Responses now wait instead of implicitly steering another active turn. A provider blocked on the original question still receives its answer directly. - Approval receipts cannot be edited, discarded, or reordered as comments. This preserves the recorded decision. - Interruption must still prove that the prior execution stopped. The tests cover cleanup waits and duplicate delivery. - The new paid overlap case can be unexercised if the model finishes before the click lands. It cannot pass without evidence of overlap. - No database migration is required. This repairs the existing approvals and execution controls; it does not implement the roadmap's work-stream queues. ## Model Used OpenAI GPT-6 through Codex. The exact deployed model ID and context-window size were not exposed in this session. Capabilities used: agentic reasoning, repository inspection, code editing, terminal commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
625 lines
22 KiB
TypeScript
625 lines
22 KiB
TypeScript
import { captureFirstTaskAttachments } from "./first-task-attachments.js";
|
|
import { waitForFirstTaskReply } from "./first-task-replies.js";
|
|
import {
|
|
firstTaskNativeRuntimePatch,
|
|
provisionFirstTaskFixtures,
|
|
} from "./first-task-fixtures.js";
|
|
import { execFileSync } from "node:child_process";
|
|
import { expect, type Page } from "@playwright/test";
|
|
import { readFile } from "node:fs/promises";
|
|
import { pollUntil, type RunnerApi } from "./api.js";
|
|
import {
|
|
sendChatMessage,
|
|
chatQuestionPresentation,
|
|
collectChatRunEvidence,
|
|
} from "./chat-flow.js";
|
|
import type { LiveFixtureValues } from "./live-fixtures.js";
|
|
import type { CredentialName, MatrixExecution } from "./types.js";
|
|
import { firstTaskScenario } from "./first-task-cases.js";
|
|
import {
|
|
activeRuns,
|
|
snapshotInstruction,
|
|
gradeFirstTask,
|
|
firstTaskCompletionSettled,
|
|
type FirstTaskEvidence,
|
|
type FirstTaskCheckpoint,
|
|
type Row,
|
|
} from "./first-task-scoring.js";
|
|
|
|
/** The wizard creates the agent/task. Only credential provisioning uses Node fetch,
|
|
* keeping plaintext keys out of Playwright's trace and recorded form values. */
|
|
export async function setupFirstTaskFixtures(input: {
|
|
page: Page;
|
|
api: RunnerApi;
|
|
execution: MatrixExecution;
|
|
nonce: string;
|
|
credentials: Partial<Record<CredentialName, string>>;
|
|
observe: (fixtures: LiveFixtureValues) => void;
|
|
}): Promise<LiveFixtureValues> {
|
|
const { page, api, execution, nonce } = input;
|
|
await api.patch("/api/instance/settings/experimental", {
|
|
enableClassicTaskInterface: false,
|
|
});
|
|
await page.goto("/onboarding", { waitUntil: "domcontentloaded" });
|
|
const launcher = page.getByRole("button", {
|
|
name: /Start Onboarding|New Organization|Add Agent/,
|
|
});
|
|
if (await launcher.count()) await launcher.first().click();
|
|
const create = page.getByRole("button", { name: /Build a new organization/ });
|
|
if (await create.count()) await create.first().click();
|
|
await page
|
|
.getByPlaceholder("e.g. Northwind Labs")
|
|
.fill(`First task ${nonce}`);
|
|
await page.getByRole("button", { name: "Continue", exact: true }).click();
|
|
await expect(page.locator("#onboarding-agent-name")).toBeVisible();
|
|
const companies = await api.get<Row[]>("/api/companies");
|
|
const company = companies.find((c) => c.name === `First task ${nonce}`);
|
|
if (!company) throw new Error("Onboarding fixture company missing");
|
|
const fixtures = await provisionFirstTaskFixtures({
|
|
api,
|
|
execution,
|
|
nonce,
|
|
credentials: input.credentials,
|
|
company: {
|
|
id: company.id,
|
|
name: company.name,
|
|
issuePrefix: company.issuePrefix,
|
|
},
|
|
});
|
|
const secret = fixtures.secretRefs[execution.profile.credential]!;
|
|
input.observe(fixtures);
|
|
await page.locator("#onboarding-agent-name").fill(fixtures.agent.name);
|
|
await page.getByRole("button", { name: "Next", exact: true }).click();
|
|
await expect(
|
|
page.getByRole("heading", { name: "Connect a model" }),
|
|
).toBeVisible();
|
|
await page
|
|
.getByRole("radio", {
|
|
name:
|
|
execution.profile.credential === "OPENAI_API_KEY"
|
|
? /^OpenAI/
|
|
: /^Claude/,
|
|
})
|
|
.click();
|
|
const useKey = page.getByRole("button", {
|
|
name: "Use API key instead",
|
|
exact: true,
|
|
});
|
|
const savedKey = page.getByRole("combobox", { name: "Saved API key" });
|
|
// Credential mode depends on the selected provider and its asynchronous key lookup.
|
|
await expect(savedKey.or(useKey).first()).toBeVisible();
|
|
if (await useKey.isVisible()) await useKey.click();
|
|
await page
|
|
.getByRole("combobox", { name: "Saved API key" })
|
|
.selectOption(`company:${secret.secretId}`);
|
|
await page
|
|
.getByRole("button", { name: /^(Connect|Next)$/, exact: true })
|
|
.last()
|
|
.click();
|
|
await expect(
|
|
page.getByRole("button", { name: "Get started", exact: true }),
|
|
).toBeVisible({ timeout: 120_000 });
|
|
const agents = await api.get<Row[]>(`/api/companies/${company.id}/agents`);
|
|
expect(agents).toHaveLength(1);
|
|
const wizardAdapter =
|
|
execution.profile.credential === "OPENAI_API_KEY"
|
|
? "codex_local"
|
|
: "claude_local";
|
|
expect(agents[0].adapterType).toBe(wizardAdapter);
|
|
fixtures.onboardingRuntime = {
|
|
mode: "production-wizard",
|
|
originalAdapterType: wizardAdapter,
|
|
testedAdapterType: wizardAdapter,
|
|
originalModel: agents[0].adapterConfig?.model ?? null,
|
|
};
|
|
if (execution.profile.generation === "native") {
|
|
const runtimePatch = firstTaskNativeRuntimePatch(
|
|
execution,
|
|
fixtures,
|
|
agents[0],
|
|
);
|
|
const migrated = await api.patch<Row>(
|
|
`/api/agents/${agents[0].id}`,
|
|
runtimePatch,
|
|
);
|
|
expect(migrated.adapterType).toBe("paperclip_runner");
|
|
expect(migrated.adapterConfig?.provider).toBe(execution.profile.provider);
|
|
expect(migrated.adapterConfig?.instructionsFilePath).toBe(
|
|
agents[0].adapterConfig?.instructionsFilePath,
|
|
);
|
|
expect(migrated.adapterConfig?.paperclipSkillSync?.desiredSkills).toEqual(
|
|
(
|
|
runtimePatch.adapterConfig.paperclipSkillSync as {
|
|
desiredSkills?: unknown[];
|
|
}
|
|
)?.desiredSkills,
|
|
);
|
|
fixtures.onboardingRuntime = {
|
|
mode: "post-onboarding-runtime-switch",
|
|
originalAdapterType: wizardAdapter,
|
|
testedAdapterType: migrated.adapterType,
|
|
originalModel: agents[0].adapterConfig?.model ?? null,
|
|
};
|
|
}
|
|
fixtures.agent = {
|
|
id: agents[0].id,
|
|
companyId: company.id,
|
|
name: agents[0].name,
|
|
};
|
|
input.observe(fixtures);
|
|
await page.getByRole("button", { name: "Get started", exact: true }).click();
|
|
return fixtures;
|
|
}
|
|
|
|
export async function runFirstTaskFlow(input: {
|
|
page: Page;
|
|
api: RunnerApi;
|
|
fixtures: LiveFixtureValues;
|
|
execution: MatrixExecution;
|
|
nonce: string;
|
|
observe: (issue: any, runs: any[], evidence: FirstTaskEvidence) => void;
|
|
secrets: readonly string[];
|
|
createOrdinary: (title: string, prompt: string) => Promise<Row>;
|
|
capture: (id: string, label: string, file: string) => Promise<void>;
|
|
evidence: (name: string, value: unknown) => Promise<void>;
|
|
}) {
|
|
const { page, api, fixtures, execution, nonce } = input;
|
|
const scenario = firstTaskScenario(execution.task.id, nonce);
|
|
const deadlineAt =
|
|
Date.now() + execution.task.attemptTimeoutMs.local - 120_000;
|
|
const tasksPath = `/api/companies/${fixtures.company.id}/issues?limit=100`;
|
|
const allRuns = async () => {
|
|
const runs = await api.get<Row[]>(
|
|
`/api/companies/${fixtures.company.id}/heartbeat-runs?limit=100`,
|
|
);
|
|
return Promise.all(
|
|
runs.map((r) => api.get<Row>(`/api/heartbeat-runs/${r.id}`)),
|
|
);
|
|
};
|
|
const initial = await pollUntil({
|
|
label: "onboarding first task",
|
|
deadlineAt,
|
|
load: () => api.get<Row[]>(tasksPath),
|
|
accept: (rows) => rows.length === 1,
|
|
});
|
|
let issue = initial[0];
|
|
expect(issue.assigneeAgentId).toBe(fixtures.agent.id);
|
|
expect(issue.description).toContain("/first-task");
|
|
expect(await allRuns()).toHaveLength(0);
|
|
const e: FirstTaskEvidence = {
|
|
caseId: scenario.id,
|
|
nonce,
|
|
onboardingIssueId: issue.id,
|
|
agentId: fixtures.agent.id,
|
|
initialTaskIds: initial.map((t) => t.id),
|
|
instructions: [],
|
|
configuredModel: null,
|
|
observedModels: [],
|
|
checkpoints: [],
|
|
checks: [],
|
|
};
|
|
input.observe(issue, [], e);
|
|
const snapshot = async (
|
|
phase: FirstTaskCheckpoint["phase"],
|
|
at = new Date().toISOString(),
|
|
) => {
|
|
const [tasks, agents, comments, interactions, runs] = await Promise.all([
|
|
api.get<Row[]>(tasksPath),
|
|
api.get<Row[]>(`/api/companies/${fixtures.company.id}/agents`),
|
|
api.get<Row[]>(`/api/issues/${issue.id}/comments?order=asc`),
|
|
api.get<Row[]>(`/api/issues/${issue.id}/interactions`),
|
|
allRuns(),
|
|
]);
|
|
const documents = (
|
|
await Promise.all(
|
|
tasks.map(async (t) => {
|
|
const summaries = await api.get<Row[]>(
|
|
`/api/issues/${t.id}/documents`,
|
|
);
|
|
return Promise.all(
|
|
summaries.map(async (d) => ({
|
|
...(await api.get<Row>(
|
|
`/api/issues/${t.id}/documents/${encodeURIComponent(d.key)}`,
|
|
)),
|
|
key: d.key,
|
|
issueId: t.id,
|
|
})),
|
|
);
|
|
}),
|
|
)
|
|
).flat();
|
|
issue = tasks.find((t) => t.id === issue.id) ?? issue;
|
|
e.observedModels = [
|
|
...new Set(
|
|
runs
|
|
.map((r) => r.resultJson?.model ?? r.usageJson?.model)
|
|
.filter((m): m is string => typeof m === "string"),
|
|
),
|
|
];
|
|
const checkpoint: FirstTaskCheckpoint = {
|
|
id: `${phase}-${e.checkpoints.length}`,
|
|
at,
|
|
phase,
|
|
issueId: issue.id,
|
|
tasks,
|
|
agents,
|
|
comments,
|
|
interactions,
|
|
documents,
|
|
attachments: await captureFirstTaskAttachments(api, tasks, input.secrets),
|
|
runs,
|
|
};
|
|
e.checkpoints.push(checkpoint);
|
|
e.checks = gradeFirstTask(e);
|
|
input.observe(issue, runs, e);
|
|
await input.evidence("first-task.json", e);
|
|
await input.evidence("api-state.json", checkpoint);
|
|
return checkpoint;
|
|
};
|
|
const settle = async (priorRunIds: Set<string>, completion = false) => {
|
|
let stable = 0;
|
|
await pollUntil({
|
|
label: "first-task response and durable outcome",
|
|
deadlineAt: Math.min(deadlineAt, Date.now() + 300_000),
|
|
intervalMs: 1000,
|
|
load: async () => ({
|
|
runs: await allRuns(),
|
|
tasks: await api.get<Row[]>(tasksPath),
|
|
}),
|
|
reject: ({ runs }) => {
|
|
const bad = runs.find((r) =>
|
|
["failed", "timed_out", "cancelled"].includes(r.status),
|
|
);
|
|
if (bad)
|
|
return `run status ${bad.status}: ${bad.errorCode ?? ""} ${bad.error ?? ""}`;
|
|
if (runs.length > 12) return "first-task run count exceeded 12";
|
|
},
|
|
accept: ({ runs, tasks }) => {
|
|
const settled =
|
|
runs.some((r) => !priorRunIds.has(r.id)) &&
|
|
activeRuns(runs).length === 0;
|
|
const done =
|
|
!completion ||
|
|
firstTaskCompletionSettled(
|
|
tasks,
|
|
e.initialTaskIds,
|
|
e.onboardingIssueId,
|
|
);
|
|
stable = settled && done ? stable + 1 : 0;
|
|
return stable >= 3;
|
|
},
|
|
});
|
|
};
|
|
const turn = async (
|
|
message: string,
|
|
phase: FirstTaskCheckpoint["phase"],
|
|
complete = false,
|
|
) => {
|
|
const before = new Set((await allRuns()).map((r) => r.id));
|
|
const loadComments = () =>
|
|
api.get<Row[]>(`/api/issues/${issue.id}/comments?order=asc`);
|
|
const previousIds = new Set(
|
|
(await loadComments()).map((comment) => comment.id),
|
|
);
|
|
const at = new Date().toISOString();
|
|
await sendChatMessage(page, message);
|
|
await waitForFirstTaskReply({
|
|
load: loadComments,
|
|
previousIds,
|
|
message,
|
|
deadlineAt: Math.min(deadlineAt, Date.now() + 30_000),
|
|
});
|
|
if (phase === "accepted") await snapshot("accepted", at);
|
|
await settle(before, complete);
|
|
return snapshot(phase === "accepted" ? "finished" : phase);
|
|
};
|
|
let failure: unknown;
|
|
try {
|
|
const git = (args: string[]) =>
|
|
execFileSync("git", args, {
|
|
cwd: new URL("../../", import.meta.url),
|
|
encoding: "utf8",
|
|
}).trim();
|
|
e.source = {
|
|
sha: git(["rev-parse", "HEAD"]),
|
|
ref: git(["branch", "--show-current"]),
|
|
dirty: Boolean(git(["status", "--porcelain"])),
|
|
};
|
|
const agent = await api.get<Row>(`/api/agents/${fixtures.agent.id}`);
|
|
e.configuredModel = agent.adapterConfig?.model ?? null;
|
|
e.runtimeSettings = {
|
|
onboardingRuntime: fixtures.onboardingRuntime,
|
|
adapterType: agent.adapterType,
|
|
adapterConfig: agent.adapterConfig,
|
|
runtimeConfig: agent.runtimeConfig,
|
|
permissions: agent.permissions,
|
|
};
|
|
const desired =
|
|
agent.adapterConfig?.paperclipSkillSync?.desiredSkills ?? [];
|
|
expect(
|
|
desired.map((s: string | { key: string }) =>
|
|
typeof s === "string" ? s : s.key,
|
|
),
|
|
).toContain("paperclipai/paperclip/first-task");
|
|
const bundle = await api.get<{ files: Array<{ path: string }> }>(
|
|
`/api/agents/${fixtures.agent.id}/instructions-bundle`,
|
|
);
|
|
for (const file of bundle.files) {
|
|
const detail = await api.get<{ content: string }>(
|
|
`/api/agents/${fixtures.agent.id}/instructions-bundle/file?path=${encodeURIComponent(file.path)}`,
|
|
);
|
|
e.instructions.push(
|
|
snapshotInstruction(file.path, detail.content, input.secrets),
|
|
);
|
|
}
|
|
const skills = await api.get<Row[]>(
|
|
`/api/companies/${fixtures.company.id}/skills`,
|
|
);
|
|
for (const selection of desired) {
|
|
const key = typeof selection === "string" ? selection : selection.key;
|
|
const skill = skills.find((s) => s.key === key);
|
|
if (!skill) throw new Error(`Missing assigned skill ${key}`);
|
|
const file = await api.get<{ content: string }>(
|
|
`/api/companies/${fixtures.company.id}/skills/${skill.id}/files?path=SKILL.md`,
|
|
);
|
|
e.instructions.push(
|
|
snapshotInstruction(
|
|
`${skill.slug}/SKILL.md`,
|
|
file.content,
|
|
input.secrets,
|
|
),
|
|
);
|
|
}
|
|
// The execution contract is appended by production runners outside the managed persona.
|
|
const contract = await readFile(
|
|
new URL(
|
|
"../../server/src/onboarding-assets/default/AGENTS.md",
|
|
import.meta.url,
|
|
),
|
|
"utf8",
|
|
);
|
|
e.instructions.push(
|
|
snapshotInstruction("runtime/default/AGENTS.md", contract, input.secrets),
|
|
);
|
|
await snapshot("opening");
|
|
await page.goto(
|
|
`/${fixtures.company.issuePrefix}/issues/${issue.identifier ?? issue.id}`,
|
|
);
|
|
if (scenario.opening === "ordinary") {
|
|
const before = new Set((await allRuns()).map((r) => r.id));
|
|
issue = await input.createOrdinary(
|
|
execution.task.buildTitle(nonce),
|
|
scenario.prompt,
|
|
);
|
|
e.initialTaskIds.push(issue.id);
|
|
input.observe(issue, [], e);
|
|
expect(issue.description).not.toContain("/first-task");
|
|
await page.goto(
|
|
`/${fixtures.company.issuePrefix}/issues/${issue.identifier ?? issue.id}`,
|
|
);
|
|
await settle(before);
|
|
await snapshot("response");
|
|
} else if (scenario.opening === "message") {
|
|
await turn(scenario.prompt, "response");
|
|
} else {
|
|
const opening = e.checkpoints[0].interactions.find(
|
|
(i) => i.kind === "ask_user_questions" && i.status === "pending",
|
|
);
|
|
expect(opening, "deterministic opening card").toBeTruthy();
|
|
const option = opening!.payload.questions[0].options.find(
|
|
(o: any) =>
|
|
o.id === (scenario.opening === "interview" ? "interview" : "task"),
|
|
);
|
|
await page
|
|
// Paperclip includes the option description in the accessible name.
|
|
.getByRole("radio", { name: option.label })
|
|
.last()
|
|
.click();
|
|
if (scenario.opening !== "interview")
|
|
await page
|
|
.getByTestId("question-other-answer-composer")
|
|
.last()
|
|
.locator('[contenteditable="true"],textarea')
|
|
.first()
|
|
.fill(scenario.prompt);
|
|
const before = new Set((await allRuns()).map((r) => r.id));
|
|
await page
|
|
.getByRole("button", {
|
|
name: opening!.payload.submitLabel ?? "Continue",
|
|
exact: true,
|
|
})
|
|
.last()
|
|
.click();
|
|
if (scenario.id === "accept-while-running") {
|
|
await pollUntil({
|
|
label: "approval card published before the source run finishes",
|
|
deadlineAt: Math.min(deadlineAt, Date.now() + 300_000), intervalMs: 100,
|
|
load: async () => ({ interactions: await api.get<Row[]>(`/api/issues/${issue.id}/interactions`), runs: await allRuns() }),
|
|
accept: ({ interactions, runs }) => interactions.some(i => i.status === "pending" &&
|
|
["request_confirmation", "request_checkbox_confirmation"].includes(i.kind) &&
|
|
runs.some(r => r.id === i.sourceRunId && r.status === "running")),
|
|
reject: ({ runs }) => runs.some(r => !before.has(r.id)) && !activeRuns(runs).length
|
|
? "Acceptance overlap was not exercised: source run finished before a live approval card was observed" : undefined,
|
|
});
|
|
} else await settle(before);
|
|
await snapshot("response");
|
|
}
|
|
// Screenshots can take longer than the source turn's final handoff. In the
|
|
// overlap case, accept first and retain the response checkpoint as evidence.
|
|
if (scenario.id !== "accept-while-running") await input.capture(
|
|
"first-task-response",
|
|
"First onboarding response",
|
|
"first-task-response.png",
|
|
);
|
|
const assertBeforeAcceptance = () => {
|
|
const check = gradeFirstTask(e).find((c) => c.id === "no-premature-work");
|
|
expect(
|
|
check?.passed ?? true,
|
|
"Behavior failure: work executed before acceptance",
|
|
).toBe(true);
|
|
};
|
|
assertBeforeAcceptance();
|
|
if (!scenario.firstResponseOnly) {
|
|
if (["interview", "ambiguous"].includes(scenario.opening)) {
|
|
const facts =
|
|
scenario.id === "interview-plan-accept"
|
|
? `${scenario.facts} Please write a plan for me to review, before doing the work.`
|
|
: scenario.facts;
|
|
const pending = (
|
|
await api.get<Row[]>(`/api/issues/${issue.id}/interactions`)
|
|
).find(
|
|
(i) => i.status === "pending" && i.kind === "ask_user_questions",
|
|
);
|
|
if (!pending) await turn(facts, "clarified");
|
|
else {
|
|
const set = chatQuestionPresentation(pending.payload);
|
|
const before = new Set((await allRuns()).map((r) => r.id));
|
|
for (const [index, question] of set.questions.entries()) {
|
|
const text = page
|
|
.getByTestId("question-text-answer-composer")
|
|
.last();
|
|
if (await text.isVisible())
|
|
await text
|
|
.locator('[contenteditable="true"],textarea')
|
|
.first()
|
|
.fill(facts);
|
|
else {
|
|
await page
|
|
.getByRole(
|
|
question.answerMode === "multi_select" ? "checkbox" : "radio",
|
|
{
|
|
name: question.customAnswer?.label ?? "Other",
|
|
exact: true,
|
|
},
|
|
)
|
|
.last()
|
|
.click();
|
|
await page
|
|
.getByTestId("question-other-answer-composer")
|
|
.last()
|
|
.locator('[contenteditable="true"],textarea')
|
|
.first()
|
|
.fill(facts);
|
|
}
|
|
await page
|
|
.getByRole("button", {
|
|
name:
|
|
index === set.questions.length - 1
|
|
? (set.submitLabel ?? "Submit answers")
|
|
: "Next",
|
|
exact: true,
|
|
})
|
|
.last()
|
|
.click();
|
|
}
|
|
await settle(before);
|
|
await snapshot("clarified");
|
|
}
|
|
}
|
|
if (scenario.id === "revise-accept")
|
|
await turn(scenario.revision, "revised");
|
|
assertBeforeAcceptance();
|
|
if (scenario.id === "reject-no-execution")
|
|
await turn(scenario.rejection, "rejected");
|
|
else if (["task-card-accept", "accept-while-running"].includes(scenario.id)) {
|
|
const pending = (
|
|
await api.get<Row[]>(`/api/issues/${issue.id}/interactions`)
|
|
).find(
|
|
(i) =>
|
|
i.status === "pending" &&
|
|
["request_confirmation", "request_checkbox_confirmation"].includes(
|
|
i.kind,
|
|
),
|
|
);
|
|
expect(
|
|
pending,
|
|
"proposal must offer an acceptance card in this case",
|
|
).toBeTruthy();
|
|
const before = new Set((await allRuns()).map((r) => r.id));
|
|
const at = new Date().toISOString();
|
|
if (pending!.kind === "request_checkbox_confirmation") {
|
|
for (const item of pending!.payload.options ?? [])
|
|
await page
|
|
.getByRole("checkbox", { name: item.label, exact: true })
|
|
.last()
|
|
.check();
|
|
}
|
|
await page
|
|
.getByRole("button", {
|
|
name:
|
|
pending!.payload.acceptLabel ??
|
|
(pending!.kind === "request_confirmation"
|
|
? "Approve"
|
|
: "Confirm selection"),
|
|
exact: true,
|
|
})
|
|
.last()
|
|
.click();
|
|
await pollUntil({
|
|
label: "first-task confirmation acceptance",
|
|
deadlineAt: Math.min(deadlineAt, Date.now() + 30_000),
|
|
intervalMs: 250,
|
|
load: () => api.get<Row[]>(`/api/issues/${issue.id}/interactions`),
|
|
accept: (interactions) =>
|
|
interactions.find((i) => i.id === pending!.id)?.status ===
|
|
"accepted",
|
|
reject: (interactions) => {
|
|
const status = interactions.find(
|
|
(i) => i.id === pending!.id,
|
|
)?.status;
|
|
return status && !["pending", "accepted"].includes(status)
|
|
? `Confirmation ended as ${status}`
|
|
: undefined;
|
|
},
|
|
});
|
|
await snapshot("accepted", at);
|
|
await settle(before, true);
|
|
await snapshot("finished");
|
|
} else
|
|
await turn(
|
|
scenario.acceptance,
|
|
"accepted",
|
|
scenario.id !== "interview-plan-accept",
|
|
);
|
|
}
|
|
e.checks = gradeFirstTask(e);
|
|
await input.evidence("first-task.json", e);
|
|
await input.capture(
|
|
"final-state",
|
|
"First-task final state",
|
|
"final-state.png",
|
|
);
|
|
const failures = e.checks.filter((c) => !c.passed);
|
|
expect(
|
|
failures,
|
|
failures.map((c) => `${c.id}: ${c.detail}`).join("\n"),
|
|
).toEqual([]);
|
|
return { issue, runs: e.checkpoints.at(-1)!.runs, evidence: e };
|
|
} catch (error) {
|
|
failure = error;
|
|
throw error;
|
|
} finally {
|
|
// Save failed journeys without replacing the original assertion/transport error.
|
|
try {
|
|
await snapshot("finished");
|
|
const runs = await allRuns();
|
|
await input.evidence(
|
|
"first-task-run-evidence.json",
|
|
await Promise.all(
|
|
runs.map((r) =>
|
|
collectChatRunEvidence(
|
|
api,
|
|
r as Parameters<typeof collectChatRunEvidence>[1],
|
|
),
|
|
),
|
|
),
|
|
);
|
|
} catch (error) {
|
|
if (!failure) throw error;
|
|
await input.evidence("first-task-evidence-error.json", {
|
|
captureFailed: true,
|
|
});
|
|
}
|
|
}
|
|
}
|