Files
PaperClipAI/tests/runner-e2e/native-completion-scoring.ts
DottaandPaperclip a386a59998 Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native agents receive task constraints and completion tools from
Paperclip.
> - Completion tools already define the procedure for reporting a
result.
> - Repeated procedure text adds instructions to each full task turn.
> - The final reply must still explain a blocker and link a saved
document.
> - This pull request removes repeated procedure text and keeps these
visible outcome requirements explicit.
> - A document receipt supplies the exact link, and stricter evals check
the persisted reply and browser navigation.

## Linked Issues or Issue Description

Refs: #14961. Related: #14948 and #15007.

**What happened?**

Native task envelopes repeat completion procedure text. A reduced
envelope needs explicit final-reply requirements. The `write_document`
receipt also lacks a canonical document link.

**Expected behavior**

Keep the completion tools as the source of procedure details. Require
one accepted completion result before the final reply. A blocked reply
must explain the reason, owner and unblock action. A document reply must
contain a working link to the saved document.

**Steps to reproduce**

1. Run the native assigned-skill document case and native blocker case.
2. Inspect the run-attributed provider final and its persisted comment.
3. Check the blocker explanation or open the final reply's document
link.

## What Changed

- Remove repeated completion procedure text from the native task
constraints and backend instructions.
- Keep explicit blocker and document-link requirements in full task
turns.
- Return a company/task-scoped `documentHref` from `write_document`.
Preserve the link in the idempotent mutation receipt.
- Repeat canonical links for this run's current saved revisions in
accepted completion feedback. Give blocked providers final-response
guidance for the cause, owner and unblock action.
- Keep internal document/comment anchors when Markdown issue links load
cached issue details.
- Add a manual six-cell comparison suite with strict source, build,
default-instruction and budget admission.
- Capture eighteen shared runnerd RPC projections and six direct
OpenCode HTTP projections across start, resume and continuation phases,
using scripted local transports and no provider execution.
- Apply v3 checks only to the manual instruction comparison; preserve v2
checks for the existing native completion suite. Check the actual
persisted blocker reason and exact saved-document link. Click the
rendered document link and check the original content marker in the
classic document card or the new document tab.
- Forward exact OpenCode finishing calls through the controller. Wait
for acceptance, keep accepted feedback and concrete rejection text, and
reject malformed responses. Preserve ordinary dynamic-tool response
handling.
- Settle the completion decision and tool response before mapping a
racing idle/error/abort event or handling explicit close/interruption.
Reject a concurrent finishing call before controller admission.
- Add a provider-free regression through real runnerd, the OpenCode
proxy and a fake provider. Reject the first completion, accept the
corrected report in the same turn, and propose one result.
- Keep all original verdicts unchanged. Treat replay under new checks as
separate diagnostics.

## Verification

- `pnpm -r typecheck` and `pnpm build` pass locally.
- Native document-authority tests pass, including company/run
authorization and idempotent replay.
- Native runtime-context, backend and measurement tests pass.
- Final-answer calibration, protocol scoring, source-admission and
catalog tests pass. Wrong reasons, absent links and wrong link targets
fail.
- `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six
single-attempt local cells with the declared models.
- Exported `prepareNativeInstructionPreflight` then
`verifyNativeInstructionPreflight` pass on this clean committed source.
They build locally and make zero provider calls.
- Corrective live confirmation is incomplete. Source 3a7349d passed both
Claude and both Codex cases. OpenCode saved the correct document but
omitted its final link; its blocker case was canceled before paid
execution. Preserve this failure. The e171282 confirmation was stopped
during build after fresh review found a completion-settlement race; it
executed zero providers. Source c3e0cb303 fixes that race. Two affected
OpenCode cases await fresh review and one bounded confirmation; earlier
results remain attributed to their original source.
- OpenCode proxy parsing, driver, factory and input tests: 81 pass
across retained focused runs, including six settlement races.
Evaluator/scoring/admission checks: 122 pass. The real proxy regression
passes. Fresh local prepare then verify passes with 18 shared and 6
direct scripted captures, fresh SDK/Rust builds and zero providers.
- The full local suite recorded two failures: a webhook timeout and a
Git-scan load count mismatch. Both files pass in isolation with
unchanged assertions/time budgets; preserve the original failure log.
Fresh c3e0cb303 CI and review are pending. This PR remains draft.

## Risks

- Final-answer wording can vary by provider. The checks cover the
declared release-access blocker and saved document fixture, not general
answer quality.
- A single trial does not establish general equivalence, cause, speed,
cost or live resume behavior.
- `documentHref` is an additive receipt field. It points to the current
saved document, not an immutable historical revision. Replaying an older
receipt does not fabricate a new link.
- The correction adds four production paths for document receipts,
accepted completion feedback and UI navigation, plus four OpenCode
controller/proxy paths, beyond the original three instruction paths.
Completion rejection must remain repairable; the production-boundary
regression covers it.
- Preserve the frozen comparison context for live measurement. A
merge-tree check against current master is clean. Do not relabel earlier
live results as results from a later source tree.

## Model Used

- OpenAI Codex, GPT-6 family. The exact serving model ID and context
window are unavailable in this session. Capabilities used: reasoning,
code editing, shell execution, test authoring and evidence review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 08:10:14 -05:00

151 lines
11 KiB
TypeScript

import { isValidNativePrpEnvelope } from "./native-event-envelope.js";
import { explainsMissingReleaseAccess, linksSavedNativeDocument, type NativeDocumentLinkContext } from "./native-completion-content.js";
type Row = Record<string, unknown>;
const row = (value: unknown): Row => value && typeof value === "object" && !Array.isArray(value) ? value as Row : {};
const forbiddenKinds = new Set(["commandExecution", "process", "fileChange", "file_change", "apply_patch"]);
const toolKinds = new Set(["dynamicToolCall", "mcpToolCall", "commandExecution", "tool_call", "tool_use"]);
const terminalName = (value: unknown, expected: string) => typeof value === "string" && (value === expected || value.endsWith(`.${expected}`));
export interface NativeCompletionObservation {
caseId: "assigned-skill-explicit-invocation" | "native-blocked-report";
companyId: string; agentId: string; issue: Row; runs: readonly Row[];
comments: readonly Row[]; events: readonly Row[];
state: { issueIds: string[]; agentIds: string[]; documentCount: number; interactionCount: number };
initial: { issueIds: string[]; agentIds: string[] };
workspaceChanged: boolean;
marker: string;
documentLinkContext?: NativeDocumentLinkContext;
}
/** Observable tool/result/final sequence, not a claim that the provider consumed feedback. */
export function gradeNativeCompletion(input: NativeCompletionObservation) {
return gradeObservation(input, false);
}
/** Stricter final-answer checks belong to the instruction comparison only. */
export function gradeNativeCompletionFinalAnswer(input: NativeCompletionObservation) {
return gradeObservation(input, true);
}
function gradeObservation(input: NativeCompletionObservation, finalAnswer: boolean) {
const checks: Array<{ id: string; passed: boolean; detail: string }> = [];
const check = (id: string, passed: boolean, detail: string) => checks.push({ id, passed, detail });
const blocked = input.caseId === "native-blocked-report";
const run = input.runs[0] ?? {};
const expectedDisposition = blocked ? "blocked" : "done";
const expectedTool = blocked ? "paperclip_block" : "paperclip_finish";
const nativeResult = row(row(run.resultJson).nativeResult);
check("one-native-attempt", input.runs.length === 1 && run.status === "succeeded" && run.runtimeMode === "native"
&& typeof run.id === "string" && run.id.length > 0 && typeof input.issue.id === "string" && input.issue.id.length > 0
&& run.nativeIssueId === input.issue.id
&& run.companyId === input.companyId && run.agentId === input.agentId
&& !run.retryOfRunId && !run.continuationAttempt && !input.issue.scheduledRetry,
"Exactly one succeeded native run, with no continuation/retry or uncounted extra work.");
check("durable-disposition", input.issue.companyId === input.companyId && input.issue.assigneeAgentId === input.agentId
&& input.issue.status === expectedDisposition && nativeResult.reportedWorkDisposition === expectedDisposition,
"Public durable issue status and the native semantic result agree with the requested disposition.");
if (blocked) {
const blocker = row(nativeResult.blocker);
check("exact-blocker", row(blocker.owner).name === "Release Owner" && blocker.unblockAction === "Grant deployment access"
&& blocker.scope === "task_wide", "The structured owner, exact action and whole-task scope are independently checked.");
}
const events = input.events.map(event => ({ event, envelope: row(row(event.payload).prpEvent) }));
const bound = events.length > 0 && events.every(({ event, envelope }, index) => {
const action = /^(?:tool\.execution\.|item\.)/.test(String(event.eventType));
const required = action || ["run.result.proposed", "run.result.accepted", "run.terminal"].includes(String(event.eventType));
return event.seq === index + 1
&& (Object.keys(envelope).length === 0 ? !required : isValidNativePrpEnvelope(envelope, event.protocolSchemaVersion as number | undefined)
&& (!action || envelope.sourceKind === "runner")
&& typeof envelope.sourceEventId === "string" && envelope.sourceEventId.length > 0
&& typeof envelope.sourceInstanceId === "string" && envelope.sourceInstanceId.length > 0
&& Number.isSafeInteger(envelope.sourceSeq) && Number(envelope.sourceSeq) > 0
&& envelope.runId === run.id && envelope.eventType === event.eventType
&& envelope.sourceEventId === event.sourceEventId && envelope.sourceInstanceId === event.sourceInstanceId
&& envelope.sourceSeq === event.sourceSeq);
});
check("complete-bound-events", bound, "A complete contiguous public PRP stream is bound to this run and persisted source identity.");
const selected = (type: string) => events.filter(({ event }) => event.eventType === type);
const proposed = selected("run.result.proposed");
const accepted = selected("run.result.accepted");
const terminal = selected("run.terminal");
const acceptedResult = row(row(accepted[0]?.envelope.payload).result);
const terminalResult = row(terminal[0]?.envelope.payload);
check("authoritative-native-result", proposed.length === 1 && accepted.length === 1 && terminal.length === 1
&& proposed[0]?.envelope.sourceKind === "runner" && accepted[0]?.envelope.sourceKind === "control_plane" && terminal[0]?.envelope.sourceKind === "control_plane"
&& row(proposed[0]?.envelope.payload).reportedWorkDisposition === expectedDisposition
&& acceptedResult.reportedWorkDisposition === expectedDisposition
&& terminalResult.runTerminalState === "succeeded" && terminalResult.turnTerminalState === "completed",
"One runner proposal is independently accepted and terminated by the control plane.");
const proposalSeq = Number(proposed[0]?.event.seq ?? -1);
const finishes = events.filter(({ event, envelope }) => {
const payload = row(envelope.payload), item = row(payload.item);
const completion = envelope.sourceKind === "runner" && Number(event.seq) > proposalSeq && (
event.eventType === "tool.execution.completed" && terminalName(payload.name, expectedTool) && payload.status === "completed"
|| event.eventType === "item.completed" && (payload.kind === "tool_result" || item.type === "tool_result") && item.status === "completed"
);
if (!completion) return false;
return events.some(({ event: start, envelope: startEnvelope }) => {
const started = row(startEnvelope.payload), startedItem = row(started.item);
return startEnvelope.sourceKind === "runner" && Number(start.seq) < proposalSeq && (
event.eventType === "tool.execution.completed" && start.eventType === "tool.execution.started"
&& typeof payload.executionId === "string" && payload.executionId.length > 0 && started.executionId === payload.executionId && terminalName(started.name, expectedTool)
|| event.eventType === "item.completed" && start.eventType === "item.started"
&& (started.kind === "tool_call" || startedItem.type === "tool_call")
&& typeof item.id === "string" && item.id.length > 0 && item.id === startedItem.id
&& terminalName(startedItem.name, expectedTool)
&& (item.name === undefined || item.name === null || terminalName(item.name, expectedTool))
);
});
});
const finals = events.filter(({ event, envelope }) => {
const payload = row(envelope.payload), item = row(payload.item);
return envelope.sourceKind === "runner" && event.eventType === "item.completed"
&& (payload.kind === "agentMessage" || item.type === "agentMessage")
&& (payload.channel === "final" || item.channel === "final") && item.phase === "final_answer"
&& typeof item.text === "string" && item.text.trim().length > 0;
});
const finishSeq = Number(finishes.at(-1)?.event.seq ?? -1);
const finalSeq = Number(finals[0]?.event.seq ?? -1);
const starts = events.filter(({ event, envelope }) => envelope.sourceKind === "runner" && (
event.eventType === "tool.execution.started" || event.eventType === "item.started"
&& toolKinds.has(String(row(envelope.payload).kind ?? row(row(envelope.payload).item).type))));
check("final-after-tool-result", proposed.length === 1 && finishes.length > 0 && finals.length === 1
&& proposalSeq < finishSeq && finishSeq < finalSeq
&& proposalSeq < Number(accepted[0]?.event.seq ?? -1) && Number(accepted[0]?.event.seq ?? -1) < Number(terminal[0]?.event.seq ?? -1)
&& !starts.some(({ event }) => Number(event.seq) > proposalSeq),
"The provider final follows the observed finishing result; no new invocation follows admission. Control-plane acceptance precedes its terminal receipt independently.");
const finalText = String(row(row(finals[0]?.envelope.payload).item).text ?? "").trim();
const replies = input.comments.filter(comment => comment.createdByRunId === run.id && comment.authorAgentId === input.agentId
&& typeof comment.body === "string" && comment.body.trim() === finalText);
check("persisted-provider-final", finalText.length > 0 && replies.length === 1,
"A nonempty provider final is durably projected as this run's agent reply; semantic-summary fallback is insufficient.");
if (blocked) check("visible-blocker-content", finalText.split(input.marker).length === 2
&& finalText.includes("Release Owner") && finalText.includes("Grant deployment access")
&& /\b(?:blocked|cannot proceed|can't proceed|missing|required access|not (?:yet )?granted|awaiting|waiting|unavailable)\b/i.test(finalText)
&& !/\b(?:not blocked|no longer blocked|access (?:is |has been |was )?already granted|completed Grant deployment access)\b/i.test(finalText),
"The final explains the actual unresolved blocker and its owner/action, rather than supplying only a marker.");
if (finalAnswer && blocked) check("visible-blocker-reason", explainsMissingReleaseAccess(finalText),
"The persisted provider final independently explains the missing release/deployment access; a blocked label or unblock action alone is insufficient.");
else if (finalAnswer) check("saved-document-final-link", linksSavedNativeDocument(finalText, input.documentLinkContext),
"The persisted provider final links this task's one saved, revisioned document at the canonical same-origin anchor.");
const same = (a: string[], b: string[]) => a.length === b.length && a.every(id => b.includes(id)) && new Set(a).size === a.length;
check("bounded-durable-work", same(input.state.issueIds, [...input.initial.issueIds, String(input.issue.id)])
&& same(input.state.agentIds, input.initial.agentIds) && input.state.documentCount === (blocked ? 0 : 1)
&& input.state.interactionCount === 0,
"No additional tasks, agents or interactions; completion permits its one required document and blocker permits none.");
if (blocked) {
const forbidden = events.some(({ event, envelope }) => {
const payload = row(envelope.payload), item = row(payload.item);
return envelope.sourceKind === "runner" && (forbiddenKinds.has(String(payload.kind ?? item.type))
|| String(event.eventType).startsWith("tool.execution.") && payload.transport === "process"
|| /(?:write_document|create_task|hire_agent|register_deliverable|apply_patch|exec_command)/.test(String(payload.name ?? item.name ?? "")));
});
check("no-deployment-or-file-work", !input.workspaceChanged && !forbidden,
"The fixture workspace is unchanged and no process/file/deployment or extra-work tool is observed.");
}
return { schema: finalAnswer ? "paperclip.native-completion-observation.v3" : "paperclip.native-completion-observation.v2", passed: checks.every(value => value.passed), checks,
limitations: ["Exact provider feedback identity/consumption is not measured by the public sequence."] };
}