Files
PaperClipAI/tests/runner-e2e/first-task-scoring.ts
T
DottaandPaperclip cea8dda472 test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can delegate work through onboarding and Agent Chat.
> - A completed task does not prove that its result reached the original
conversation.
> - Existing tests do not isolate completion after the source chat
becomes idle.
> - This pull request adds four explicit native-runner probes across
Claude and Codex.
> - The probes preserve the result and reply so we can separate delivery
failures from inaccurate answers.

## Linked Issues or Issue Description

Refs #13775. Refs #13813.

These evals extend native-runner qualification. They measure completion
updates before we choose a product change.

## What Changed

- Add the opt-in `completion-updates` suite with two stories for each
native provider.
- Test completion in the existing onboarding task flow and after an
Agent Chat handoff becomes idle.
- Gate the chat worker on a brief inside its managed project workspace.
Prove the source is idle before releasing the worker.
- Check durable task completion, saved output, a subsequent source
reply, and rendered access to the result.
- Preserve replies, task state, screenshots, run events, and a separate
semantic review rubric.
- Add grader regression tests and update the documented eval contract.
- Preserve the suites added on master and include four completion cases
in the 306-cell catalog.

Production behavior and prompts are unchanged.

## Verification

- Passed all 565 eval support tests across 45 files after merging
current master: `node node_modules/vitest/vitest.mjs run --config
tests/runner-e2e/vitest.config.ts`.
- Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p
tests/runner-e2e/tsconfig.json`.
- Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs
tests/runner-e2e/launch.ts --list --suite completion-updates`.
- Four-cell behavior campaign on source
`ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`:
https://github.com/paperclipai/paperclip/actions/runs/36072337485.
- A screenshot-only follow-up waits for the restored source reply to
render after result-link navigation. Its one-cell Claude onboarding
verification passed on final head:
https://github.com/paperclipai/paperclip/actions/runs/36075716141.
Corrected report:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/.
The original four-cell onboarding screenshots caught navigation loading;
its saved reply evidence remains valid. The follow-up again found stale
wording: "That work will run next" was posted 38 seconds after the child
was Done. The four-cell campaign keeps its original source and
measurements.
- Published evidence:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/.
- Suite definition:
`afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`,
version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local
execution, one attempt per cell. All four cleanup checks passed.
Onboarding billing coverage is partial; reported zero cost must not be
read as a free run.

| Story | Automated delivery/access | Separate semantic review |
| --- | --- | --- |
| Codex onboarding | Pass | Pass: accurate completion reply with an
accessible result |
| Claude onboarding | Pass | Fail: reply says it will save the note once
the task runs, after the note is already saved and the task is Done |
| Codex idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |
| Claude idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |

Both chat cases positively recorded the source waiting and the worker at
the brief gate before release. Both saved outputs include the brief-only
start time. The opt-in campaign is red because it exposes current
behavior. It is not a required merge gate. The PR does not fix that
product behavior. Semantic review is a recorded human/agent assessment
of retained evidence; it is not an automated prose-quality judge.
- Second campaign:
https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex
chat reached the idle boundary and completed its task, then received no
completion reply during the full window. Claude onboarding again
returned a stale handoff answer. Claude chat exceeded the prior
110-second handoff setup budget; this revision raises that bounded setup
window to 180 seconds.
- Retained baseline:
https://github.com/paperclipai/paperclip/actions/runs/36069427676.
Onboarding passed delivery/access for both providers, but Claude gave a
stale handoff answer. Chat cases stopped at fixture problems; they do
not establish a completion-delivery failure. This revision fixes the
workspace path and competing reference requirements.
- On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR
checks passed and two were skipped, including typecheck, tests, and
build. Broad checks ran in CI, not locally. That head received Greptile
5/5 with no unresolved findings. The unchanged mobile
repository-settings browser test passed on one targeted retry after a
detached/disabled Save-button timeout.

- Merged current master in `9b4491e1f` and resolved the catalog-count
conflict. Eval support tests and eval TypeScript pass locally. All
individual CI jobs passed on this merge commit, including build,
typecheck, server tests, runner checks, and browser shards. The final
aggregate check also passed: 54 checks passed and two were skipped.
Greptile reviewed this exact commit at 5/5 with no unresolved findings.

## Risks

- These explicit probes can expose current product failures. They do not
change the default paid test selection.
- Mechanical delivery and result access do not establish answer
accuracy. The preserved reply still requires semantic review.
- A fixture failure before the idle boundary or worker completion cannot
establish a completion-update failure.
- The handoff setup window lasts three minutes. The worker brief wait is
bounded at four minutes. The observation window lasts two minutes after
worker completion. It retains later replies without erasing earlier
accessible delivery.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository
inspection, code execution, and GitHub tool use. The runtime does not
expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:58:14 -05:00

477 lines
18 KiB
TypeScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
import { isBlockedUnstartedWake } from "./non-execution-wake.js";
import { answerableRuntimeRunIds } from "./runtime-question-readiness.js";
import { sanitizeJson } from "./redaction.js";
import { createHash } from "node:crypto";
import { firstTaskScenario } from "./first-task-cases.js";
import { completionDelivery } from "./completion-updates.js";
export type Row = { id: string; [key: string]: any };
export interface FirstTaskCheckpoint {
id: string;
at: string;
phase:
| "opening"
| "response"
| "clarified"
| "revised"
| "accepted"
| "rejected"
| "finished";
issueId: string;
tasks: Row[];
agents: Row[];
comments: Row[];
interactions: Row[];
documents: Row[];
attachments?: Row[];
runs: Row[];
}
export interface FirstTaskEvidence {
caseId: string;
nonce: string;
onboardingIssueId: string;
agentId: string;
initialTaskIds: string[];
instructions: Array<{
path: string;
content: string;
sha256: string;
contentSha256?: string;
redacted?: boolean;
}>;
source?: { sha: string; ref: string; dirty: boolean };
runtimeSettings?: Record<string, unknown>;
configuredModel: string | null;
observedModels: string[];
checkpoints: FirstTaskCheckpoint[];
checks: FirstTaskCheck[];
}
export interface FirstTaskCheck {
id: string;
passed: boolean;
/** Unexercised checks remain non-passing, but are not behavioral failures. */
notReached?: string;
evidence: string[];
detail: string;
}
export const digestText = (text: string) =>
createHash("sha256").update(text).digest("hex");
export function snapshotInstruction(
path: string,
content: string,
secrets: readonly string[] = [],
) {
const safe = sanitizeJson({ content }, secrets) as { content: string };
return {
path,
content: safe.content,
sha256: digestText(content),
contentSha256: digestText(safe.content),
redacted: safe.content !== content,
};
}
/** Only user-visible affirmative proposals count; incidental task nouns do not. */
function hasTaskProposal(text: string, confirmationCard = false) {
const affirmative = text.split(/(?<=[.!?])\s+/).filter((sentence) =>
!/\b(?:no\s+(?:(?:new|child)\s+)?(?:subtask|task)|(?:not|never|won['’]t|don['’]t|without)\b[^.!?]{0,80}\b(?:subtask|task))\b/i.test(sentence),
).join("\n");
return /\b(?:approve|accept|propose|proposed|proposing|proposal|suggest|suggested|suggesting|recommend|recommended|recommending|create|creating|set up)\b[\s\S]{0,160}\b(?:subtask|task)\b/i.test(affirmative) ||
/\b(?:subtask|task)\b[^.!?\n]{0,30}\bproposal\b/i.test(affirmative) ||
(confirmationCard && (
/\b(?:I|we)(?:['’]ll| will)\s+(?:save|attach|write|make|open|create)\b[^.!?]{0,160}\b(?:(?:new|child)\s+task|subtask)\b/i.test(affirmative) ||
/\b(?:subtask|task)\s+(?:I|we)(?:['’]ll| will)\s+(?:create|open|set up)\b/i.test(affirmative)
));
}
export const activeRuns = (runs: Row[]) =>
runs.filter((run) => ["queued", "running"].includes(run.status));
export function questionCount(interaction: Row) {
return (
interaction.payload?.questionSet?.questions ??
interaction.payload?.questions ??
[]
).length;
}
/** Inspect the presentation the user saw, including superseded cards at earlier checkpoints. */
function gradeQuestionChoices(
checkpoints: FirstTaskCheckpoint[],
): FirstTaskCheck {
const seen = new Set<string>();
const failures: string[] = [];
const refs = new Set<string>();
for (const checkpoint of checkpoints) {
for (const interaction of checkpoint.interactions) {
if (interaction.kind !== "ask_user_questions") continue;
const questions =
interaction.payload?.questionSet?.questions ??
interaction.payload?.questions ??
[];
for (const question of questions) {
const key = JSON.stringify([interaction.id, question]);
if (seen.has(key)) continue;
seen.add(key);
// A canonical text question has no visible choice controls. Its legacy
// compatibility entry may contain a single freeText option; ignore that.
if (question?.answerMode === "text") continue;
const labels = new Set<string>(
(question?.options ?? [])
.map((option: any) =>
String(option.label ?? "")
.trim()
.toLowerCase(),
)
.filter(Boolean),
);
if (labels.size >= 2) continue;
refs.add(checkpoint.id);
failures.push(
`${interaction.id} / ${question?.id ?? "unknown question"}: "${question?.prompt ?? ""}" has ${labels.size} distinct choice option(s); expected at least 2, or a text question`,
);
}
}
}
return {
id: "question-choice-options",
passed: failures.length === 0,
detail:
failures.length > 0
? failures.join("; ")
: "Every presented choice question offers at least two distinct options; open-ended text questions are allowed",
evidence: failures.length > 0 ? [...refs] : checkpoints.map((c) => c.id),
};
}
/** A terminal parent with no child is an assessable wrong outcome, not a transport timeout. */
export function firstTaskCompletionSettled(
tasks: Row[],
initialTaskIds: string[],
issueId: string,
): boolean {
const children = tasks.filter((task) => !initialTaskIds.includes(task.id));
return children.length > 0
? children.every((task) => task.status === "done")
: tasks.some((task) => task.id === issueId && task.status === "done");
}
function isPlanDocument(document: Row): boolean {
return (
document.key === "plan" ||
(/(?:^|[-_])plan(?:$|[-_])/i.test(String(document.key)) &&
(/\bplan\b/i.test(String(document.title ?? "")) ||
/^#+\s+plan\b/im.test(String(document.body ?? ""))))
);
}
/** Plan/proposal documents are planning evidence, never durable completion output. */
function isPlanningDocument(document: Row): boolean {
return (
isPlanDocument(document) ||
(/(?:^|[-_])proposal(?:$|[-_])/i.test(String(document.key)) &&
/(?:^|\n)(?:#+\s*)?(?:proposed (?:(?:child|single|first) )?task|(?:(?:first|single)[- ])?task proposal|proposal)\b/i.test(
`${document.title ?? ""}\n${document.body ?? ""}`,
))
);
}
function isVerifiedAttachment(a: Row): boolean {
return a.contentVerified === true && typeof a.body === "string" && a.contentSha256 === digestText(a.body);
}
function attachmentDocument(a: Row): Row {
return { ...a, key: String(a.originalFilename ?? a.filename ?? "").replace(/\.(?:md|txt)$/i, ""), title: a.title ?? a.originalFilename ?? a.filename };
}
function isPlanningAttachment(a: Row): boolean {
return isVerifiedAttachment(a) && isPlanningDocument(attachmentDocument(a));
}
function verifiedFirstTaskOutputs(checkpoint: FirstTaskCheckpoint): Row[] {
return [...checkpoint.documents, ...(checkpoint.attachments ?? []).filter(isVerifiedAttachment).map(attachmentDocument)];
}
/** Provider identities, not generic sessionReused flags, prove continuity. */
export function gradeNativeSessionContinuity(runs: Row[], issueId: string): FirstTaskCheck {
const parent = [...new Map(runs.filter((run) => run.nativeIssueId === issueId).map((run) => [run.id, run])).values()];
const identities = parent.map((run) => ({
run: run.id,
session: run.nativeSessionId,
provider: run.runnerProfileJson?.sessionCheckpoint?.providerSessionId,
workspace: run.runnerProfileJson?.nativeExecutionInput?.binding?.executionWorkspaceId,
}));
const passed = parent.length >= 2 && ["session", "provider", "workspace"].every((key) =>
identities.every((identity) => typeof identity[key as "session"] === "string" && identity[key as "session"].length > 0) &&
new Set(identities.map((identity) => identity[key as "session"])).size === 1,
);
return { id: "native-session-continuity", passed, evidence: ["finished"],
detail: `Same-task follow-ups must retain native/provider/workspace identities: ${JSON.stringify(identities)}` };
}
export function gradeFirstTask(e: FirstTaskEvidence): FirstTaskCheck[] {
const scenario = firstTaskScenario(e.caseId, e.nonce);
const checks: FirstTaskCheck[] = [];
const add = (
id: string,
passed: boolean,
detail: string,
refs: string[],
notReached?: string,
) =>
checks.push({
id,
passed: notReached ? false : passed,
detail,
evidence: refs,
...(notReached ? { notReached } : {}),
});
const first = e.checkpoints.find((c) => c.phase === "response");
const last = e.checkpoints.at(-1);
const approval = e.checkpoints.find((c) => c.phase === "accepted");
const rejection =
e.caseId === "reject-no-execution"
? e.checkpoints.find((c) => c.phase === "rejected")
: undefined;
const extras = (c: FirstTaskCheckpoint) =>
c.tasks.filter((t) => !e.initialTaskIds.includes(t.id));
const seeded = new Set(e.checkpoints[0]?.comments.map((c) => c.id));
const replies = (c: FirstTaskCheckpoint) =>
c.comments.filter((r) => r.authorAgentId && !seeded.has(r.id));
add(
"recorded-response",
Boolean(
first &&
replies(first).length +
first.interactions.filter(
(i) =>
!e.checkpoints[0]?.interactions.some((old) => old.id === i.id),
).length >
0,
),
"An agent response or structured interaction was recorded",
first ? [first.id] : [],
);
add(
"instruction-snapshot",
e.instructions.some((i) => i.path.endsWith("first-task/SKILL.md")) &&
e.instructions.some((i) => i.path === "AGENTS.md") &&
e.instructions.every(
(i) =>
digestText(i.content) === (i.contentSha256 ?? i.sha256) &&
(i.redacted || i.sha256 === digestText(i.content)),
),
"Actual persona and skill were retained with verified hashes",
e.instructions.map((i) => i.path),
);
checks.push(gradeQuestionChoices(e.checkpoints));
if (!first || !last) return checks;
const text = replies(first)
.map((r) => r.body ?? "")
.join("\n");
const questions = first.interactions.filter(
(i) =>
i.kind === "ask_user_questions" &&
!e.checkpoints[0]?.interactions.some((old) => old.id === i.id),
);
add(
"no-reintroduction",
!/welcome to paperclip|your first agent teammate/i.test(text),
"Do not repeat the seeded welcome",
[first.id],
);
add(
"opening-not-repeated",
!questions.some((i) =>
(i.payload?.questionSet?.questions ?? i.payload?.questions ?? []).some(
(q: any) => q.id === "first-task-opening",
),
),
"Do not post another opening choice",
[first.id],
);
if (scenario.opening === "interview")
add(
"interview-questions",
questions.length === 1 &&
questionCount(questions[0]) >= 3 &&
questionCount(questions[0]) <= 4,
"One structured interview with 3–4 questions",
[first.id],
);
if (scenario.opening === "ambiguous")
add(
"focused-clarification",
questions.length > 0 || (text.match(/\?/g)?.length ?? 0) >= 2,
"Ask focused questions before proposing ambiguous work",
[first.id],
);
if (["task", "message"].includes(scenario.opening))
add(
"subtask-proposal",
[text, ...first.documents.filter(isPlanningDocument)
.map((d) => `${d.title ?? ""}\n${d.body ?? ""}`)]
.some((visible) => hasTaskProposal(visible)) ||
first.interactions.some((i) =>
["request_confirmation", "request_checkbox_confirmation"].includes(i.kind) &&
hasTaskProposal([i.title, i.summary, i.payload?.prompt, i.payload?.detailsMarkdown]
.filter(Boolean).join("\n"), true),
),
"Propose a task for the concrete request",
[first.id],
);
if (scenario.opening === "plan" || e.caseId === "interview-plan-accept") {
const proposal =
e.checkpoints.find((c) => c.phase === "clarified") ?? first;
add(
"durable-plan",
proposal.documents.some(
(d) => isPlanDocument(d) && String(d.body).trim(),
),
"Save the requested plan before acceptance",
[proposal.id],
);
}
const before = e.checkpoints.filter(
(c) =>
c.phase !== "opening" &&
(!approval || Date.parse(c.at) < Date.parse(approval.at)),
);
if (scenario.opening !== "ordinary") {
add(
"no-premature-work",
before.every(
(c) =>
extras(c).length === 0 &&
c.agents.length === e.checkpoints[0].agents.length &&
c.documents.every(isPlanningDocument) &&
(c.attachments ?? []).every(isPlanningAttachment) &&
(c.tasks.find((t) => t.id === c.issueId)?.status !== "done" ||
Boolean(rejection && Date.parse(c.at) >= Date.parse(rejection.at))),
),
"Before acceptance: no hires, execution tasks, finished output, or claimed completion; closing an unexecuted rejected task is allowed",
before.map((c) => c.id),
);
} else {
add(
"ordinary-task-control",
questions.length === 0 &&
extras(last).length === 0 &&
verifiedFirstTaskOutputs(last).some(
(d) =>
!isPlanningDocument(d) && String(d.body).includes(scenario.marker),
),
"Ordinary work produces output without onboarding questions or delegation",
[last.id],
);
}
if (e.caseId === "accept-while-running") {
const accepted = approval?.interactions.find(i => i.status === "accepted" &&
["request_confirmation", "request_checkbox_confirmation"].includes(i.kind));
const source = last.runs.find(r => r.id === accepted?.sourceRunId);
const acceptedAt = Date.parse(accepted?.resolvedAt ?? "");
const startedAt = Date.parse(source?.startedAt ?? "");
const finishedAt = Date.parse(source?.finishedAt ?? "");
const overlapped = Number.isFinite(acceptedAt) && Number.isFinite(startedAt) &&
Number.isFinite(finishedAt) && startedAt <= acceptedAt && acceptedAt < finishedAt;
add("accepted-while-running", overlapped,
"The persisted approval resolution must fall inside its source run's actual execution interval",
approval ? [approval.id, last.id] : [last.id],
overlapped ? undefined : "The recording did not prove acceptance during the source run; the concurrency regression was not exercised.");
}
if (!scenario.firstResponseOnly && e.caseId !== "reject-no-execution") {
const notReached = approval
? undefined
: "The recording did not reach the acceptance checkpoint; this part of the journey was not evaluated.";
const acceptedComment = approval?.comments.some(
(c) => !c.authorAgentId && c.body === scenario.acceptance,
);
const acceptedCard = approval?.interactions.some(
(i) =>
["request_confirmation", "request_checkbox_confirmation"].includes(
i.kind,
) &&
i.status === "accepted" &&
i.result?.outcome === "accepted",
);
add(
"acceptance-recorded",
Boolean(approval && (acceptedComment || acceptedCard)),
"Explicit acceptance is persisted as a user comment or approved confirmation card",
approval ? [approval.id] : [],
notReached,
);
if (e.caseId === "interview-plan-accept") {
add(
"accepted-plan-retained",
last.documents.some((d) => isPlanDocument(d) && String(d.body).trim()),
"The accepted plan remains available as a durable document",
[last.id],
notReached,
);
} else {
const children = extras(last);
add(
"one-scoped-subtask",
children.length === 1 &&
children[0].parentId === e.onboardingIssueId &&
children[0].assigneeAgentId === e.agentId,
"Exactly one approved subtask belongs to this onboarding issue and agent",
[last.id],
notReached,
);
add(
"creation-after-acceptance",
Boolean(approval) &&
children.every(
(t) =>
typeof t.createdAt === "string" &&
Date.parse(t.createdAt) >= Date.parse(approval!.at),
),
"Task creation must follow acceptance, including between checkpoints",
[last.id],
notReached,
);
add(
"durable-completion",
children.length === 1 &&
children[0].status === "done" &&
verifiedFirstTaskOutputs(last).some(
(d) =>
d.issueId === children[0].id &&
!isPlanningDocument(d) &&
String(d.body).includes(scenario.marker) &&
(e.caseId !== "revise-accept" ||
(!String(d.body).includes(scenario.originalMarker) &&
/sunday/i.test(d.body))),
),
"Approved output is saved on the completed child, with revised scope when applicable",
[last.id],
notReached,
);
}
}
if (e.caseId === "reject-no-execution")
add(
"rejection-respected",
extras(last).length === 0 &&
last.documents.every(isPlanningDocument) &&
(last.attachments ?? []).every(isPlanningAttachment) &&
activeRuns(last.runs).length === 0,
"Rejected work never executes",
[last.id],
rejection
? undefined
: "The recording did not reach the rejection checkpoint; the rejection response was not evaluated.",
);
add(
"provider-runs-succeeded",
last.runs.length > 0 && last.runs.every((r) => r.status === "succeeded" || isBlockedUnstartedWake(r) ||
(scenario.firstResponseOnly && r.status === "running" && answerableRuntimeRunIds(last.interactions).has(r.id))),
"Provider runs succeeded, or a first-response run is paused on its recorded answerable native question",
[last.id],
);
if (e.runtimeSettings?.adapterType === "paperclip_runner" &&
["task-reply-accept", "task-card-accept"].includes(e.caseId) && last.phase === "finished") {
checks.push({ ...gradeNativeSessionContinuity(e.checkpoints.flatMap((checkpoint) => checkpoint.runs), e.onboardingIssueId), evidence: [last.id] });
}
if (e.runtimeSettings?.completionDeliveryProbe && last) {
const worker = last.tasks.find(t => t.parentId === e.onboardingIssueId) ?? {};
checks.push(...completionDelivery({ sourceId: e.onboardingIssueId, worker, documents: last.documents,
comments: last.comments, runs: last.runs, marker: scenario.marker,
renderedLinks: e.runtimeSettings.completionRenderedLinks as Array<{ commentId: string; href: string }> | undefined,
}).checks.map(check => ({ ...check, evidence: [last.id] })));
}
return checks;
}