mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native agents receive task constraints and completion tools from Paperclip. > - Completion tools already define the procedure for reporting a result. > - Repeated procedure text adds instructions to each full task turn. > - The final reply must still explain a blocker and link a saved document. > - This pull request removes repeated procedure text and keeps these visible outcome requirements explicit. > - A document receipt supplies the exact link, and stricter evals check the persisted reply and browser navigation. ## Linked Issues or Issue Description Refs: #14961. Related: #14948 and #15007. **What happened?** Native task envelopes repeat completion procedure text. A reduced envelope needs explicit final-reply requirements. The `write_document` receipt also lacks a canonical document link. **Expected behavior** Keep the completion tools as the source of procedure details. Require one accepted completion result before the final reply. A blocked reply must explain the reason, owner and unblock action. A document reply must contain a working link to the saved document. **Steps to reproduce** 1. Run the native assigned-skill document case and native blocker case. 2. Inspect the run-attributed provider final and its persisted comment. 3. Check the blocker explanation or open the final reply's document link. ## What Changed - Remove repeated completion procedure text from the native task constraints and backend instructions. - Keep explicit blocker and document-link requirements in full task turns. - Return a company/task-scoped `documentHref` from `write_document`. Preserve the link in the idempotent mutation receipt. - Repeat canonical links for this run's current saved revisions in accepted completion feedback. Give blocked providers final-response guidance for the cause, owner and unblock action. - Keep internal document/comment anchors when Markdown issue links load cached issue details. - Add a manual six-cell comparison suite with strict source, build, default-instruction and budget admission. - Capture eighteen shared runnerd RPC projections and six direct OpenCode HTTP projections across start, resume and continuation phases, using scripted local transports and no provider execution. - Apply v3 checks only to the manual instruction comparison; preserve v2 checks for the existing native completion suite. Check the actual persisted blocker reason and exact saved-document link. Click the rendered document link and check the original content marker in the classic document card or the new document tab. - Forward exact OpenCode finishing calls through the controller. Wait for acceptance, keep accepted feedback and concrete rejection text, and reject malformed responses. Preserve ordinary dynamic-tool response handling. - Settle the completion decision and tool response before mapping a racing idle/error/abort event or handling explicit close/interruption. Reject a concurrent finishing call before controller admission. - Add a provider-free regression through real runnerd, the OpenCode proxy and a fake provider. Reject the first completion, accept the corrected report in the same turn, and propose one result. - Keep all original verdicts unchanged. Treat replay under new checks as separate diagnostics. ## Verification - `pnpm -r typecheck` and `pnpm build` pass locally. - Native document-authority tests pass, including company/run authorization and idempotent replay. - Native runtime-context, backend and measurement tests pass. - Final-answer calibration, protocol scoring, source-admission and catalog tests pass. Wrong reasons, absent links and wrong link targets fail. - `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six single-attempt local cells with the declared models. - Exported `prepareNativeInstructionPreflight` then `verifyNativeInstructionPreflight` pass on this clean committed source. They build locally and make zero provider calls. - Corrective live confirmation is incomplete. Source3a7349dpassed both Claude and both Codex cases. OpenCode saved the correct document but omitted its final link; its blocker case was canceled before paid execution. Preserve this failure. Thee171282confirmation was stopped during build after fresh review found a completion-settlement race; it executed zero providers. Sourcec3e0cb303fixes that race. Two affected OpenCode cases await fresh review and one bounded confirmation; earlier results remain attributed to their original source. - OpenCode proxy parsing, driver, factory and input tests: 81 pass across retained focused runs, including six settlement races. Evaluator/scoring/admission checks: 122 pass. The real proxy regression passes. Fresh local prepare then verify passes with 18 shared and 6 direct scripted captures, fresh SDK/Rust builds and zero providers. - The full local suite recorded two failures: a webhook timeout and a Git-scan load count mismatch. Both files pass in isolation with unchanged assertions/time budgets; preserve the original failure log. Freshc3e0cb303CI and review are pending. This PR remains draft. ## Risks - Final-answer wording can vary by provider. The checks cover the declared release-access blocker and saved document fixture, not general answer quality. - A single trial does not establish general equivalence, cause, speed, cost or live resume behavior. - `documentHref` is an additive receipt field. It points to the current saved document, not an immutable historical revision. Replaying an older receipt does not fabricate a new link. - The correction adds four production paths for document receipts, accepted completion feedback and UI navigation, plus four OpenCode controller/proxy paths, beyond the original three instruction paths. Completion rejection must remain repairable; the production-boundary regression covers it. - Preserve the frozen comparison context for live measurement. A merge-tree check against current master is clean. Do not relabel earlier live results as results from a later source tree. ## Model Used - OpenAI Codex, GPT-6 family. The exact serving model ID and context window are unavailable in this session. Capabilities used: reasoning, code editing, shell execution, test authoring and evidence review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
151 lines
11 KiB
TypeScript
151 lines
11 KiB
TypeScript
import { isValidNativePrpEnvelope } from "./native-event-envelope.js";
|
|
import { explainsMissingReleaseAccess, linksSavedNativeDocument, type NativeDocumentLinkContext } from "./native-completion-content.js";
|
|
|
|
type Row = Record<string, unknown>;
|
|
const row = (value: unknown): Row => value && typeof value === "object" && !Array.isArray(value) ? value as Row : {};
|
|
const forbiddenKinds = new Set(["commandExecution", "process", "fileChange", "file_change", "apply_patch"]);
|
|
const toolKinds = new Set(["dynamicToolCall", "mcpToolCall", "commandExecution", "tool_call", "tool_use"]);
|
|
const terminalName = (value: unknown, expected: string) => typeof value === "string" && (value === expected || value.endsWith(`.${expected}`));
|
|
|
|
export interface NativeCompletionObservation {
|
|
caseId: "assigned-skill-explicit-invocation" | "native-blocked-report";
|
|
companyId: string; agentId: string; issue: Row; runs: readonly Row[];
|
|
comments: readonly Row[]; events: readonly Row[];
|
|
state: { issueIds: string[]; agentIds: string[]; documentCount: number; interactionCount: number };
|
|
initial: { issueIds: string[]; agentIds: string[] };
|
|
workspaceChanged: boolean;
|
|
marker: string;
|
|
documentLinkContext?: NativeDocumentLinkContext;
|
|
}
|
|
|
|
/** Observable tool/result/final sequence, not a claim that the provider consumed feedback. */
|
|
export function gradeNativeCompletion(input: NativeCompletionObservation) {
|
|
return gradeObservation(input, false);
|
|
}
|
|
|
|
/** Stricter final-answer checks belong to the instruction comparison only. */
|
|
export function gradeNativeCompletionFinalAnswer(input: NativeCompletionObservation) {
|
|
return gradeObservation(input, true);
|
|
}
|
|
|
|
function gradeObservation(input: NativeCompletionObservation, finalAnswer: boolean) {
|
|
const checks: Array<{ id: string; passed: boolean; detail: string }> = [];
|
|
const check = (id: string, passed: boolean, detail: string) => checks.push({ id, passed, detail });
|
|
const blocked = input.caseId === "native-blocked-report";
|
|
const run = input.runs[0] ?? {};
|
|
const expectedDisposition = blocked ? "blocked" : "done";
|
|
const expectedTool = blocked ? "paperclip_block" : "paperclip_finish";
|
|
const nativeResult = row(row(run.resultJson).nativeResult);
|
|
check("one-native-attempt", input.runs.length === 1 && run.status === "succeeded" && run.runtimeMode === "native"
|
|
&& typeof run.id === "string" && run.id.length > 0 && typeof input.issue.id === "string" && input.issue.id.length > 0
|
|
&& run.nativeIssueId === input.issue.id
|
|
&& run.companyId === input.companyId && run.agentId === input.agentId
|
|
&& !run.retryOfRunId && !run.continuationAttempt && !input.issue.scheduledRetry,
|
|
"Exactly one succeeded native run, with no continuation/retry or uncounted extra work.");
|
|
check("durable-disposition", input.issue.companyId === input.companyId && input.issue.assigneeAgentId === input.agentId
|
|
&& input.issue.status === expectedDisposition && nativeResult.reportedWorkDisposition === expectedDisposition,
|
|
"Public durable issue status and the native semantic result agree with the requested disposition.");
|
|
if (blocked) {
|
|
const blocker = row(nativeResult.blocker);
|
|
check("exact-blocker", row(blocker.owner).name === "Release Owner" && blocker.unblockAction === "Grant deployment access"
|
|
&& blocker.scope === "task_wide", "The structured owner, exact action and whole-task scope are independently checked.");
|
|
}
|
|
|
|
const events = input.events.map(event => ({ event, envelope: row(row(event.payload).prpEvent) }));
|
|
const bound = events.length > 0 && events.every(({ event, envelope }, index) => {
|
|
const action = /^(?:tool\.execution\.|item\.)/.test(String(event.eventType));
|
|
const required = action || ["run.result.proposed", "run.result.accepted", "run.terminal"].includes(String(event.eventType));
|
|
return event.seq === index + 1
|
|
&& (Object.keys(envelope).length === 0 ? !required : isValidNativePrpEnvelope(envelope, event.protocolSchemaVersion as number | undefined)
|
|
&& (!action || envelope.sourceKind === "runner")
|
|
&& typeof envelope.sourceEventId === "string" && envelope.sourceEventId.length > 0
|
|
&& typeof envelope.sourceInstanceId === "string" && envelope.sourceInstanceId.length > 0
|
|
&& Number.isSafeInteger(envelope.sourceSeq) && Number(envelope.sourceSeq) > 0
|
|
&& envelope.runId === run.id && envelope.eventType === event.eventType
|
|
&& envelope.sourceEventId === event.sourceEventId && envelope.sourceInstanceId === event.sourceInstanceId
|
|
&& envelope.sourceSeq === event.sourceSeq);
|
|
});
|
|
check("complete-bound-events", bound, "A complete contiguous public PRP stream is bound to this run and persisted source identity.");
|
|
const selected = (type: string) => events.filter(({ event }) => event.eventType === type);
|
|
const proposed = selected("run.result.proposed");
|
|
const accepted = selected("run.result.accepted");
|
|
const terminal = selected("run.terminal");
|
|
const acceptedResult = row(row(accepted[0]?.envelope.payload).result);
|
|
const terminalResult = row(terminal[0]?.envelope.payload);
|
|
check("authoritative-native-result", proposed.length === 1 && accepted.length === 1 && terminal.length === 1
|
|
&& proposed[0]?.envelope.sourceKind === "runner" && accepted[0]?.envelope.sourceKind === "control_plane" && terminal[0]?.envelope.sourceKind === "control_plane"
|
|
&& row(proposed[0]?.envelope.payload).reportedWorkDisposition === expectedDisposition
|
|
&& acceptedResult.reportedWorkDisposition === expectedDisposition
|
|
&& terminalResult.runTerminalState === "succeeded" && terminalResult.turnTerminalState === "completed",
|
|
"One runner proposal is independently accepted and terminated by the control plane.");
|
|
const proposalSeq = Number(proposed[0]?.event.seq ?? -1);
|
|
const finishes = events.filter(({ event, envelope }) => {
|
|
const payload = row(envelope.payload), item = row(payload.item);
|
|
const completion = envelope.sourceKind === "runner" && Number(event.seq) > proposalSeq && (
|
|
event.eventType === "tool.execution.completed" && terminalName(payload.name, expectedTool) && payload.status === "completed"
|
|
|| event.eventType === "item.completed" && (payload.kind === "tool_result" || item.type === "tool_result") && item.status === "completed"
|
|
);
|
|
if (!completion) return false;
|
|
return events.some(({ event: start, envelope: startEnvelope }) => {
|
|
const started = row(startEnvelope.payload), startedItem = row(started.item);
|
|
return startEnvelope.sourceKind === "runner" && Number(start.seq) < proposalSeq && (
|
|
event.eventType === "tool.execution.completed" && start.eventType === "tool.execution.started"
|
|
&& typeof payload.executionId === "string" && payload.executionId.length > 0 && started.executionId === payload.executionId && terminalName(started.name, expectedTool)
|
|
|| event.eventType === "item.completed" && start.eventType === "item.started"
|
|
&& (started.kind === "tool_call" || startedItem.type === "tool_call")
|
|
&& typeof item.id === "string" && item.id.length > 0 && item.id === startedItem.id
|
|
&& terminalName(startedItem.name, expectedTool)
|
|
&& (item.name === undefined || item.name === null || terminalName(item.name, expectedTool))
|
|
);
|
|
});
|
|
});
|
|
const finals = events.filter(({ event, envelope }) => {
|
|
const payload = row(envelope.payload), item = row(payload.item);
|
|
return envelope.sourceKind === "runner" && event.eventType === "item.completed"
|
|
&& (payload.kind === "agentMessage" || item.type === "agentMessage")
|
|
&& (payload.channel === "final" || item.channel === "final") && item.phase === "final_answer"
|
|
&& typeof item.text === "string" && item.text.trim().length > 0;
|
|
});
|
|
const finishSeq = Number(finishes.at(-1)?.event.seq ?? -1);
|
|
const finalSeq = Number(finals[0]?.event.seq ?? -1);
|
|
const starts = events.filter(({ event, envelope }) => envelope.sourceKind === "runner" && (
|
|
event.eventType === "tool.execution.started" || event.eventType === "item.started"
|
|
&& toolKinds.has(String(row(envelope.payload).kind ?? row(row(envelope.payload).item).type))));
|
|
check("final-after-tool-result", proposed.length === 1 && finishes.length > 0 && finals.length === 1
|
|
&& proposalSeq < finishSeq && finishSeq < finalSeq
|
|
&& proposalSeq < Number(accepted[0]?.event.seq ?? -1) && Number(accepted[0]?.event.seq ?? -1) < Number(terminal[0]?.event.seq ?? -1)
|
|
&& !starts.some(({ event }) => Number(event.seq) > proposalSeq),
|
|
"The provider final follows the observed finishing result; no new invocation follows admission. Control-plane acceptance precedes its terminal receipt independently.");
|
|
const finalText = String(row(row(finals[0]?.envelope.payload).item).text ?? "").trim();
|
|
const replies = input.comments.filter(comment => comment.createdByRunId === run.id && comment.authorAgentId === input.agentId
|
|
&& typeof comment.body === "string" && comment.body.trim() === finalText);
|
|
check("persisted-provider-final", finalText.length > 0 && replies.length === 1,
|
|
"A nonempty provider final is durably projected as this run's agent reply; semantic-summary fallback is insufficient.");
|
|
if (blocked) check("visible-blocker-content", finalText.split(input.marker).length === 2
|
|
&& finalText.includes("Release Owner") && finalText.includes("Grant deployment access")
|
|
&& /\b(?:blocked|cannot proceed|can't proceed|missing|required access|not (?:yet )?granted|awaiting|waiting|unavailable)\b/i.test(finalText)
|
|
&& !/\b(?:not blocked|no longer blocked|access (?:is |has been |was )?already granted|completed Grant deployment access)\b/i.test(finalText),
|
|
"The final explains the actual unresolved blocker and its owner/action, rather than supplying only a marker.");
|
|
if (finalAnswer && blocked) check("visible-blocker-reason", explainsMissingReleaseAccess(finalText),
|
|
"The persisted provider final independently explains the missing release/deployment access; a blocked label or unblock action alone is insufficient.");
|
|
else if (finalAnswer) check("saved-document-final-link", linksSavedNativeDocument(finalText, input.documentLinkContext),
|
|
"The persisted provider final links this task's one saved, revisioned document at the canonical same-origin anchor.");
|
|
const same = (a: string[], b: string[]) => a.length === b.length && a.every(id => b.includes(id)) && new Set(a).size === a.length;
|
|
check("bounded-durable-work", same(input.state.issueIds, [...input.initial.issueIds, String(input.issue.id)])
|
|
&& same(input.state.agentIds, input.initial.agentIds) && input.state.documentCount === (blocked ? 0 : 1)
|
|
&& input.state.interactionCount === 0,
|
|
"No additional tasks, agents or interactions; completion permits its one required document and blocker permits none.");
|
|
if (blocked) {
|
|
const forbidden = events.some(({ event, envelope }) => {
|
|
const payload = row(envelope.payload), item = row(payload.item);
|
|
return envelope.sourceKind === "runner" && (forbiddenKinds.has(String(payload.kind ?? item.type))
|
|
|| String(event.eventType).startsWith("tool.execution.") && payload.transport === "process"
|
|
|| /(?:write_document|create_task|hire_agent|register_deliverable|apply_patch|exec_command)/.test(String(payload.name ?? item.name ?? "")));
|
|
});
|
|
check("no-deployment-or-file-work", !input.workspaceChanged && !forbidden,
|
|
"The fixture workspace is unchanged and no process/file/deployment or extra-work tool is observed.");
|
|
}
|
|
return { schema: finalAnswer ? "paperclip.native-completion-observation.v3" : "paperclip.native-completion-observation.v2", passed: checks.every(value => value.passed), checks,
|
|
limitations: ["Exact provider feedback identity/consumption is not measured by the public sequence."] };
|
|
}
|