Files
PaperClipAI/tests/runner-e2e/first-task-scoring.ts
T
DottaandPaperclip d0b67bfe71 feat: queue approvals and answers during active runs (#13539)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users guide running agents through messages, questions, and approval
cards.
> - Messages already wait in a queue when an agent is running.
> - Card responses did not appear in that queue. Some question answers
also steered a later run without a user click.
> - A fast approval could invalidate the agent's review handoff and
cause it to stop its own run.
> - This pull request gives card responses the same queue controls and
preserves the exact response during delivery.
> - Users can wait for completion or explicitly send the response with
Interrupt or Steer.

## Linked Issues or Issue Description

Refs #13517, which is merged. This PR targets master and adds queued
interaction responses on top of the onboarding changes. Related
continuation work: #10519 and #12866.

**What happened?**

Accepting a proposal while its source run was active left a saved
response outside the message queue. The agent could then lose its review
path, reassign the task, and cancel itself. Answers to older questions
could also steer another active turn without a click.

**Expected behavior**

Save the response immediately. Queue its continuation behind the active
run. Deliver it after completion, or when the user explicitly chooses
Interrupt or Steer. Preserve approval revisions and answer choices.

**Steps to reproduce**

1. Let an agent publish a confirmation card while its run is still
active.
2. Accept the card before the agent finishes its review handoff.
3. Inspect the message queue and the task's next run.

**Paperclip version or commit**

Reproduced on da8a3876c with the onboarding changes from #13517.

**Deployment mode**

Local development from source. The fix covers legacy adapters and native
Runner turns.

## What Changed

- Project resolved cards into the existing queue as immutable responses.
Keep answers and exact approval revisions.
- Require an explicit click to steer a response into a compatible native
turn. Use Interrupt when a fresh session is required.
- Preserve typed response context through interruption, cleanup waits,
and normal queue promotion. Keep the direct answer channel for a
provider blocked on its original question request.
- Accept the source run's review handoff after its card resolves. Reject
stale agent reassignment that would orphan a queued response.
- Add deterministic regression tests and an `accept-while-running` case
to the first-task suite. Require recorded timestamp overlap before that
case can pass.
- Keep the first-task skill name out of user-facing messages.

## Verification

- Red-green: the original route failed the queue regression; the changed
route passes it.
- Focused server/UI tests: 139 passed, including 64 queue-route tests.
- Runner harness unit tests: 314 passed.
- Server, UI, and Runner E2E typechecks passed. UI token gates passed.
- Full repository typecheck and build passed. Server typecheck passed
again after review fixes.
- Review regressions: 165 queue/reopen route tests, 53 wake admission
tests, and 18 run identity tests passed. Approval acknowledgement
recovery and both message/approval arrival orders are covered.
- Full local test run: 12,401 passed; three new admission regressions
ran against a cached pre-fix module. A fresh run of that entire suite
passed (53 tests). The complete CI suite passed on the final commit.
- Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional
Storybook checks skipped. Every server/workspace/browser shard, Runner
verification, build, typecheck/release registry, canary, policy, and
security check passed. Greptile: 5/5, no unresolved threads. Earlier
interrupted CI workers were replaced by this fresh complete run.
- After integrating the updated parent: 314 harness tests, 119
queue/admission tests, 44 onboarding/question-delivery tests, and 13
native recovery tests passed locally. Full repository typecheck and
build passed.
- Clarified the skill wording preference: routine replies describe the
action without announcing the internal skill; direct questions and
permission/security/execution disclosures remain truthful.
- The paid `accept-while-running` scenario is registered for all four
local first-task profiles. It has not been run against a model in this
change.

- Rebased onto the merged parent at `11921075a`; the resulting tree
exactly matches the locally verified integration tree. Final-head CI on
`b53054807` passed: 54 successful checks, 2 optional Storybook checks
skipped, no failed checks. Every new server/browser shard, aggregate
verify/e2e gate, Runner, typecheck, build, canary, and security check
passed on the first attempt. Greptile reviewed this exact head at 5/5
with no unresolved threads.

## Risks

- Responses now wait instead of implicitly steering another active turn.
A provider blocked on the original question still receives its answer
directly.
- Approval receipts cannot be edited, discarded, or reordered as
comments. This preserves the recorded decision.
- Interruption must still prove that the prior execution stopped. The
tests cover cleanup waits and duplicate delivery.
- The new paid overlap case can be unexercised if the model finishes
before the click lands. It cannot pass without evidence of overlap.
- No database migration is required. This repairs the existing approvals
and execution controls; it does not implement the roadmap's work-stream
queues.

## Model Used

OpenAI GPT-6 through Codex. The exact deployed model ID and
context-window size were not exposed in this session. Capabilities used:
agentic reasoning, repository inspection, code editing, terminal
commands, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 14:44:36 -05:00

445 lines
16 KiB
TypeScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
import { sanitizeJson } from "./redaction.js";
import { createHash } from "node:crypto";
import { firstTaskScenario } from "./first-task-cases.js";
export type Row = { id: string; [key: string]: any };
export interface FirstTaskCheckpoint {
id: string;
at: string;
phase:
| "opening"
| "response"
| "clarified"
| "revised"
| "accepted"
| "rejected"
| "finished";
issueId: string;
tasks: Row[];
agents: Row[];
comments: Row[];
interactions: Row[];
documents: Row[];
attachments?: Row[];
runs: Row[];
}
export interface FirstTaskEvidence {
caseId: string;
nonce: string;
onboardingIssueId: string;
agentId: string;
initialTaskIds: string[];
instructions: Array<{
path: string;
content: string;
sha256: string;
contentSha256?: string;
redacted?: boolean;
}>;
source?: { sha: string; ref: string; dirty: boolean };
runtimeSettings?: Record<string, unknown>;
configuredModel: string | null;
observedModels: string[];
checkpoints: FirstTaskCheckpoint[];
checks: FirstTaskCheck[];
}
export interface FirstTaskCheck {
id: string;
passed: boolean;
/** Unexercised checks remain non-passing, but are not behavioral failures. */
notReached?: string;
evidence: string[];
detail: string;
}
export const digestText = (text: string) =>
createHash("sha256").update(text).digest("hex");
export function snapshotInstruction(
path: string,
content: string,
secrets: readonly string[] = [],
) {
const safe = sanitizeJson({ content }, secrets) as { content: string };
return {
path,
content: safe.content,
sha256: digestText(content),
contentSha256: digestText(safe.content),
redacted: safe.content !== content,
};
}
/** Only user-visible affirmative proposals count; incidental task nouns do not. */
function hasTaskProposal(text: string, confirmationCard = false) {
const affirmative = text.split(/(?<=[.!?])\s+/).filter((sentence) =>
!/\b(?:no\s+(?:(?:new|child)\s+)?(?:subtask|task)|(?:not|never|won['’]t|don['’]t|without)\b[^.!?]{0,80}\b(?:subtask|task))\b/i.test(sentence),
).join("\n");
return /\b(?:approve|accept|propose|proposed|proposing|proposal|suggest|suggested|suggesting|recommend|recommended|recommending|create|creating|set up)\b[\s\S]{0,160}\b(?:subtask|task)\b/i.test(affirmative) ||
/\b(?:subtask|task)\b[^.!?\n]{0,30}\bproposal\b/i.test(affirmative) ||
(confirmationCard && (
/\b(?:I|we)(?:['’]ll| will)\s+(?:save|attach|write|make|open|create)\b[^.!?]{0,160}\b(?:(?:new|child)\s+task|subtask)\b/i.test(affirmative) ||
/\b(?:subtask|task)\s+(?:I|we)(?:['’]ll| will)\s+(?:create|open|set up)\b/i.test(affirmative)
));
}
export const activeRuns = (runs: Row[]) =>
runs.filter((run) => ["queued", "running"].includes(run.status));
export function questionCount(interaction: Row) {
return (
interaction.payload?.questionSet?.questions ??
interaction.payload?.questions ??
[]
).length;
}
/** Inspect the presentation the user saw, including superseded cards at earlier checkpoints. */
function gradeQuestionChoices(
checkpoints: FirstTaskCheckpoint[],
): FirstTaskCheck {
const seen = new Set<string>();
const failures: string[] = [];
const refs = new Set<string>();
for (const checkpoint of checkpoints) {
for (const interaction of checkpoint.interactions) {
if (interaction.kind !== "ask_user_questions") continue;
const questions =
interaction.payload?.questionSet?.questions ??
interaction.payload?.questions ??
[];
for (const question of questions) {
const key = JSON.stringify([interaction.id, question]);
if (seen.has(key)) continue;
seen.add(key);
// A canonical text question has no visible choice controls. Its legacy
// compatibility entry may contain a single freeText option; ignore that.
if (question?.answerMode === "text") continue;
const labels = new Set<string>(
(question?.options ?? [])
.map((option: any) =>
String(option.label ?? "")
.trim()
.toLowerCase(),
)
.filter(Boolean),
);
if (labels.size >= 2) continue;
refs.add(checkpoint.id);
failures.push(
`${interaction.id} / ${question?.id ?? "unknown question"}: "${question?.prompt ?? ""}" has ${labels.size} distinct choice option(s); expected at least 2, or a text question`,
);
}
}
}
return {
id: "question-choice-options",
passed: failures.length === 0,
detail:
failures.length > 0
? failures.join("; ")
: "Every presented choice question offers at least two distinct options; open-ended text questions are allowed",
evidence: failures.length > 0 ? [...refs] : checkpoints.map((c) => c.id),
};
}
/** A terminal parent with no child is an assessable wrong outcome, not a transport timeout. */
export function firstTaskCompletionSettled(
tasks: Row[],
initialTaskIds: string[],
issueId: string,
): boolean {
const children = tasks.filter((task) => !initialTaskIds.includes(task.id));
return children.length > 0
? children.every((task) => task.status === "done")
: tasks.some((task) => task.id === issueId && task.status === "done");
}
function isPlanDocument(document: Row): boolean {
return (
document.key === "plan" ||
(/(?:^|[-_])plan(?:$|[-_])/i.test(String(document.key)) &&
(/\bplan\b/i.test(String(document.title ?? "")) ||
/^#+\s+plan\b/im.test(String(document.body ?? ""))))
);
}
/** Plan/proposal documents are planning evidence, never durable completion output. */
function isPlanningDocument(document: Row): boolean {
return (
isPlanDocument(document) ||
(/(?:^|[-_])proposal(?:$|[-_])/i.test(String(document.key)) &&
/(?:^|\n)(?:#+\s*)?(?:proposed (?:(?:child|single|first) )?task|(?:(?:first|single)[- ])?task proposal|proposal)\b/i.test(
`${document.title ?? ""}\n${document.body ?? ""}`,
))
);
}
function isVerifiedAttachment(a: Row): boolean {
return a.contentVerified === true && typeof a.body === "string" && a.contentSha256 === digestText(a.body);
}
function attachmentDocument(a: Row): Row {
return { ...a, key: String(a.originalFilename ?? a.filename ?? "").replace(/\.(?:md|txt)$/i, ""), title: a.title ?? a.originalFilename ?? a.filename };
}
function isPlanningAttachment(a: Row): boolean {
return isVerifiedAttachment(a) && isPlanningDocument(attachmentDocument(a));
}
function verifiedFirstTaskOutputs(checkpoint: FirstTaskCheckpoint): Row[] {
return [...checkpoint.documents, ...(checkpoint.attachments ?? []).filter(isVerifiedAttachment).map(attachmentDocument)];
}
export function gradeFirstTask(e: FirstTaskEvidence): FirstTaskCheck[] {
const scenario = firstTaskScenario(e.caseId, e.nonce);
const checks: FirstTaskCheck[] = [];
const add = (
id: string,
passed: boolean,
detail: string,
refs: string[],
notReached?: string,
) =>
checks.push({
id,
passed: notReached ? false : passed,
detail,
evidence: refs,
...(notReached ? { notReached } : {}),
});
const first = e.checkpoints.find((c) => c.phase === "response");
const last = e.checkpoints.at(-1);
const approval = e.checkpoints.find((c) => c.phase === "accepted");
const rejection =
e.caseId === "reject-no-execution"
? e.checkpoints.find((c) => c.phase === "rejected")
: undefined;
const extras = (c: FirstTaskCheckpoint) =>
c.tasks.filter((t) => !e.initialTaskIds.includes(t.id));
const seeded = new Set(e.checkpoints[0]?.comments.map((c) => c.id));
const replies = (c: FirstTaskCheckpoint) =>
c.comments.filter((r) => r.authorAgentId && !seeded.has(r.id));
add(
"recorded-response",
Boolean(
first &&
replies(first).length +
first.interactions.filter(
(i) =>
!e.checkpoints[0]?.interactions.some((old) => old.id === i.id),
).length >
0,
),
"An agent response or structured interaction was recorded",
first ? [first.id] : [],
);
add(
"instruction-snapshot",
e.instructions.some((i) => i.path.endsWith("first-task/SKILL.md")) &&
e.instructions.some((i) => i.path === "AGENTS.md") &&
e.instructions.every(
(i) =>
digestText(i.content) === (i.contentSha256 ?? i.sha256) &&
(i.redacted || i.sha256 === digestText(i.content)),
),
"Actual persona and skill were retained with verified hashes",
e.instructions.map((i) => i.path),
);
checks.push(gradeQuestionChoices(e.checkpoints));
if (!first || !last) return checks;
const text = replies(first)
.map((r) => r.body ?? "")
.join("\n");
const questions = first.interactions.filter(
(i) =>
i.kind === "ask_user_questions" &&
!e.checkpoints[0]?.interactions.some((old) => old.id === i.id),
);
add(
"no-reintroduction",
!/welcome to paperclip|your first agent teammate/i.test(text),
"Do not repeat the seeded welcome",
[first.id],
);
add(
"opening-not-repeated",
!questions.some((i) =>
(i.payload?.questionSet?.questions ?? i.payload?.questions ?? []).some(
(q: any) => q.id === "first-task-opening",
),
),
"Do not post another opening choice",
[first.id],
);
if (scenario.opening === "interview")
add(
"interview-questions",
questions.length === 1 &&
questionCount(questions[0]) >= 3 &&
questionCount(questions[0]) <= 4,
"One structured interview with 3–4 questions",
[first.id],
);
if (scenario.opening === "ambiguous")
add(
"focused-clarification",
questions.length > 0 || (text.match(/\?/g)?.length ?? 0) >= 2,
"Ask focused questions before proposing ambiguous work",
[first.id],
);
if (["task", "message"].includes(scenario.opening))
add(
"subtask-proposal",
[text, ...first.documents.filter(isPlanningDocument)
.map((d) => `${d.title ?? ""}\n${d.body ?? ""}`)]
.some((visible) => hasTaskProposal(visible)) ||
first.interactions.some((i) =>
["request_confirmation", "request_checkbox_confirmation"].includes(i.kind) &&
hasTaskProposal([i.title, i.summary, i.payload?.prompt, i.payload?.detailsMarkdown]
.filter(Boolean).join("\n"), true),
),
"Propose a task for the concrete request",
[first.id],
);
if (scenario.opening === "plan" || e.caseId === "interview-plan-accept") {
const proposal =
e.checkpoints.find((c) => c.phase === "clarified") ?? first;
add(
"durable-plan",
proposal.documents.some(
(d) => isPlanDocument(d) && String(d.body).trim(),
),
"Save the requested plan before acceptance",
[proposal.id],
);
}
const before = e.checkpoints.filter(
(c) =>
c.phase !== "opening" &&
(!approval || Date.parse(c.at) < Date.parse(approval.at)),
);
if (scenario.opening !== "ordinary") {
add(
"no-premature-work",
before.every(
(c) =>
extras(c).length === 0 &&
c.agents.length === e.checkpoints[0].agents.length &&
c.documents.every(isPlanningDocument) &&
(c.attachments ?? []).every(isPlanningAttachment) &&
(c.tasks.find((t) => t.id === c.issueId)?.status !== "done" ||
Boolean(rejection && Date.parse(c.at) >= Date.parse(rejection.at))),
),
"Before acceptance: no hires, execution tasks, finished output, or claimed completion; closing an unexecuted rejected task is allowed",
before.map((c) => c.id),
);
} else {
add(
"ordinary-task-control",
questions.length === 0 &&
extras(last).length === 0 &&
verifiedFirstTaskOutputs(last).some(
(d) =>
!isPlanningDocument(d) && String(d.body).includes(scenario.marker),
),
"Ordinary work produces output without onboarding questions or delegation",
[last.id],
);
}
if (e.caseId === "accept-while-running") {
const accepted = approval?.interactions.find(i => i.status === "accepted" &&
["request_confirmation", "request_checkbox_confirmation"].includes(i.kind));
const source = last.runs.find(r => r.id === accepted?.sourceRunId);
const acceptedAt = Date.parse(accepted?.resolvedAt ?? "");
const startedAt = Date.parse(source?.startedAt ?? "");
const finishedAt = Date.parse(source?.finishedAt ?? "");
const overlapped = Number.isFinite(acceptedAt) && Number.isFinite(startedAt) &&
Number.isFinite(finishedAt) && startedAt <= acceptedAt && acceptedAt < finishedAt;
add("accepted-while-running", overlapped,
"The persisted approval resolution must fall inside its source run's actual execution interval",
approval ? [approval.id, last.id] : [last.id],
overlapped ? undefined : "The recording did not prove acceptance during the source run; the concurrency regression was not exercised.");
}
if (!scenario.firstResponseOnly && e.caseId !== "reject-no-execution") {
const notReached = approval
? undefined
: "The recording did not reach the acceptance checkpoint; this part of the journey was not evaluated.";
const acceptedComment = approval?.comments.some(
(c) => !c.authorAgentId && c.body === scenario.acceptance,
);
const acceptedCard = approval?.interactions.some(
(i) =>
["request_confirmation", "request_checkbox_confirmation"].includes(
i.kind,
) &&
i.status === "accepted" &&
i.result?.outcome === "accepted",
);
add(
"acceptance-recorded",
Boolean(approval && (acceptedComment || acceptedCard)),
"Explicit acceptance is persisted as a user comment or approved confirmation card",
approval ? [approval.id] : [],
notReached,
);
if (e.caseId === "interview-plan-accept") {
add(
"accepted-plan-retained",
last.documents.some((d) => isPlanDocument(d) && String(d.body).trim()),
"The accepted plan remains available as a durable document",
[last.id],
notReached,
);
} else {
const children = extras(last);
add(
"one-scoped-subtask",
children.length === 1 &&
children[0].parentId === e.onboardingIssueId &&
children[0].assigneeAgentId === e.agentId,
"Exactly one approved subtask belongs to this onboarding issue and agent",
[last.id],
notReached,
);
add(
"creation-after-acceptance",
Boolean(approval) &&
children.every(
(t) =>
typeof t.createdAt === "string" &&
Date.parse(t.createdAt) >= Date.parse(approval!.at),
),
"Task creation must follow acceptance, including between checkpoints",
[last.id],
notReached,
);
add(
"durable-completion",
children.length === 1 &&
children[0].status === "done" &&
verifiedFirstTaskOutputs(last).some(
(d) =>
d.issueId === children[0].id &&
!isPlanningDocument(d) &&
String(d.body).includes(scenario.marker) &&
(e.caseId !== "revise-accept" ||
(!String(d.body).includes(scenario.originalMarker) &&
/sunday/i.test(d.body))),
),
"Approved output is saved on the completed child, with revised scope when applicable",
[last.id],
notReached,
);
}
}
if (e.caseId === "reject-no-execution")
add(
"rejection-respected",
extras(last).length === 0 &&
last.documents.every(isPlanningDocument) &&
(last.attachments ?? []).every(isPlanningAttachment) &&
activeRuns(last.runs).length === 0,
"Rejected work never executes",
[last.id],
rejection
? undefined
: "The recording did not reach the rejection checkpoint; the rejection response was not evaluated.",
);
add(
"provider-runs-succeeded",
last.runs.length > 0 && last.runs.every((r) => r.status === "succeeded"),
"All observed provider runs settled successfully",
[last.id],
);
return checks;
}