mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native agents receive task constraints and completion tools from Paperclip. > - Completion tools already define the procedure for reporting a result. > - Repeated procedure text adds instructions to each full task turn. > - The final reply must still explain a blocker and link a saved document. > - This pull request removes repeated procedure text and keeps these visible outcome requirements explicit. > - A document receipt supplies the exact link, and stricter evals check the persisted reply and browser navigation. ## Linked Issues or Issue Description Refs: #14961. Related: #14948 and #15007. **What happened?** Native task envelopes repeat completion procedure text. A reduced envelope needs explicit final-reply requirements. The `write_document` receipt also lacks a canonical document link. **Expected behavior** Keep the completion tools as the source of procedure details. Require one accepted completion result before the final reply. A blocked reply must explain the reason, owner and unblock action. A document reply must contain a working link to the saved document. **Steps to reproduce** 1. Run the native assigned-skill document case and native blocker case. 2. Inspect the run-attributed provider final and its persisted comment. 3. Check the blocker explanation or open the final reply's document link. ## What Changed - Remove repeated completion procedure text from the native task constraints and backend instructions. - Keep explicit blocker and document-link requirements in full task turns. - Return a company/task-scoped `documentHref` from `write_document`. Preserve the link in the idempotent mutation receipt. - Repeat canonical links for this run's current saved revisions in accepted completion feedback. Give blocked providers final-response guidance for the cause, owner and unblock action. - Keep internal document/comment anchors when Markdown issue links load cached issue details. - Add a manual six-cell comparison suite with strict source, build, default-instruction and budget admission. - Capture eighteen shared runnerd RPC projections and six direct OpenCode HTTP projections across start, resume and continuation phases, using scripted local transports and no provider execution. - Apply v3 checks only to the manual instruction comparison; preserve v2 checks for the existing native completion suite. Check the actual persisted blocker reason and exact saved-document link. Click the rendered document link and check the original content marker in the classic document card or the new document tab. - Forward exact OpenCode finishing calls through the controller. Wait for acceptance, keep accepted feedback and concrete rejection text, and reject malformed responses. Preserve ordinary dynamic-tool response handling. - Settle the completion decision and tool response before mapping a racing idle/error/abort event or handling explicit close/interruption. Reject a concurrent finishing call before controller admission. - Add a provider-free regression through real runnerd, the OpenCode proxy and a fake provider. Reject the first completion, accept the corrected report in the same turn, and propose one result. - Keep all original verdicts unchanged. Treat replay under new checks as separate diagnostics. ## Verification - `pnpm -r typecheck` and `pnpm build` pass locally. - Native document-authority tests pass, including company/run authorization and idempotent replay. - Native runtime-context, backend and measurement tests pass. - Final-answer calibration, protocol scoring, source-admission and catalog tests pass. Wrong reasons, absent links and wrong link targets fail. - `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six single-attempt local cells with the declared models. - Exported `prepareNativeInstructionPreflight` then `verifyNativeInstructionPreflight` pass on this clean committed source. They build locally and make zero provider calls. - Corrective live confirmation is incomplete. Source3a7349dpassed both Claude and both Codex cases. OpenCode saved the correct document but omitted its final link; its blocker case was canceled before paid execution. Preserve this failure. Thee171282confirmation was stopped during build after fresh review found a completion-settlement race; it executed zero providers. Sourcec3e0cb303fixes that race. Two affected OpenCode cases await fresh review and one bounded confirmation; earlier results remain attributed to their original source. - OpenCode proxy parsing, driver, factory and input tests: 81 pass across retained focused runs, including six settlement races. Evaluator/scoring/admission checks: 122 pass. The real proxy regression passes. Fresh local prepare then verify passes with 18 shared and 6 direct scripted captures, fresh SDK/Rust builds and zero providers. - The full local suite recorded two failures: a webhook timeout and a Git-scan load count mismatch. Both files pass in isolation with unchanged assertions/time budgets; preserve the original failure log. Freshc3e0cb303CI and review are pending. This PR remains draft. ## Risks - Final-answer wording can vary by provider. The checks cover the declared release-access blocker and saved document fixture, not general answer quality. - A single trial does not establish general equivalence, cause, speed, cost or live resume behavior. - `documentHref` is an additive receipt field. It points to the current saved document, not an immutable historical revision. Replaying an older receipt does not fabricate a new link. - The correction adds four production paths for document receipts, accepted completion feedback and UI navigation, plus four OpenCode controller/proxy paths, beyond the original three instruction paths. Completion rejection must remain repairable; the production-boundary regression covers it. - Preserve the frozen comparison context for live measurement. A merge-tree check against current master is clean. Do not relabel earlier live results as results from a later source tree. ## Model Used - OpenAI Codex, GPT-6 family. The exact serving model ID and context window are unavailable in this session. Capabilities used: reasoning, code editing, shell execution, test authoring and evidence review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
151 lines
12 KiB
TypeScript
151 lines
12 KiB
TypeScript
import { readFileSync } from "node:fs";
|
|
import { describe, expect, it } from "vitest";
|
|
import { rehydrateRunnerdItemNotification } from "../../packages/paperclip-runner/src/live/runnerd-codex-transport.js";
|
|
import { gradeNativeCompletion as gradeLegacyNativeCompletion, gradeNativeCompletionFinalAnswer as gradeNativeCompletion, type NativeCompletionObservation } from "./native-completion-scoring.js";
|
|
type Row = Record<string, unknown>;
|
|
function sample(blocked = true, compatibility = false): NativeCompletionObservation {
|
|
const body = blocked ? "Deployment remains blocked until access is granted. Release Owner must Grant deployment access. BLOCKED_probe" : "Saved [the requested document](/RUN/issues/RUN-1#document-output).";
|
|
const result = { reportedWorkDisposition: blocked ? "blocked" : "done", ...(blocked ? { blocker: { owner: { name: "Release Owner" }, unblockAction: "Grant deployment access", scope: "task_wide" } } : {}) };
|
|
const event = (seq: number, eventType: string, payload: Row, sourceKind = "runner"): Row => ({ seq, eventType, protocolSchemaVersion: 1,
|
|
sourceEventId: `event${seq}`, sourceInstanceId: sourceKind, sourceSeq: seq,
|
|
payload: { prpEvent: { schema: "paperclip.prp.event.v1", schemaVersion: 1, runId: "run", eventType,
|
|
sourceKind, sourceEventId: `event${seq}`, sourceInstanceId: sourceKind, sourceSeq: seq, payload } } });
|
|
const tool = blocked ? "paperclip_block" : "paperclip_finish";
|
|
return { caseId: blocked ? "native-blocked-report" : "assigned-skill-explicit-invocation", companyId: "company", agentId: "agent", marker: "BLOCKED_probe",
|
|
issue: { id: "issue", companyId: "company", assigneeAgentId: "agent", status: result.reportedWorkDisposition },
|
|
runs: [{ id: "run", nativeIssueId: "issue", companyId: "company", agentId: "agent", status: "succeeded", runtimeMode: "native", resultJson: { nativeResult: result } }],
|
|
comments: [{ createdByRunId: "run", authorAgentId: "agent", body }],
|
|
documentLinkContext: { appOrigin: "https://paperclip.example", issuePrefix: "RUN", issueIdentifier: "RUN-1",
|
|
documents: blocked ? [] : [{ key: "output", latestRevisionId: "revision", latestRevisionNumber: 1 }] },
|
|
initial: { issueIds: [], agentIds: ["agent"] }, state: { issueIds: ["issue"], agentIds: ["agent"], documentCount: blocked ? 0 : 1, interactionCount: 0 }, workspaceChanged: false,
|
|
events: [compatibility ? event(1, "item.started", { kind: "tool_call", item: { type: "tool_call", id: "tool", name: tool } }) : event(1, "tool.execution.started", { name: tool, executionId: "tool" }), event(2, "run.result.proposed", result),
|
|
compatibility ? event(3, "item.completed", { kind: "tool_result", item: { type: "tool_result", status: "completed", id: "tool" } }) : event(3, "tool.execution.completed", { name: tool, status: "completed", executionId: "tool" }),
|
|
event(4, "item.completed", { kind: "agentMessage", channel: "final", item: { type: "agentMessage", phase: "final_answer", text: body, channel: "final" } }),
|
|
event(5, "run.result.accepted", { result }, "control_plane"), event(6, "run.terminal", { runTerminalState: "succeeded", turnTerminalState: "completed" }, "control_plane")],
|
|
};
|
|
}
|
|
function payload(input: NativeCompletionObservation, index: number): Row { return ((input.events[index]!.payload as Row).prpEvent as Row).payload as Row; }
|
|
describe("native completion independent oracle", () => {
|
|
it("keeps legacy completion verdicts separate from final-answer diagnostics", () => {
|
|
const value = sample(false);
|
|
const text = "Saved the requested document.";
|
|
(payload(value, 3).item as Row).text = text; value.comments[0]!.body = text;
|
|
expect(gradeLegacyNativeCompletion(value)).toMatchObject({ schema: "paperclip.native-completion-observation.v2", passed: true });
|
|
expect(gradeNativeCompletion(value)).toMatchObject({ schema: "paperclip.native-completion-observation.v3", passed: false });
|
|
expect(gradeLegacyNativeCompletion(value).checks.some(check => check.id === "saved-document-final-link")).toBe(false);
|
|
});
|
|
|
|
it("rejects a correct structured blocker whose visible reply only labels it blocked and repeats the action", () => {
|
|
const value = sample();
|
|
const text = "The whole task is blocked. Owner: Release Owner.\n\nUnblock action: Grant deployment access\n\nBLOCKED_probe";
|
|
(payload(value, 3).item as Row).text = text; value.comments[0]!.body = text;
|
|
const grade = gradeNativeCompletion(value);
|
|
expect(grade.checks.find(check => check.id === "visible-blocker-content")?.passed).toBe(true);
|
|
expect(grade.checks.find(check => check.id === "exact-blocker")?.passed).toBe(true);
|
|
expect(grade.checks.find(check => check.id === "visible-blocker-reason")?.passed).toBe(false);
|
|
expect(grade.passed).toBe(false);
|
|
});
|
|
it.each(["missing-link", "wrong-document", "wrong-reply-run", "missing-link-context", "missing-revision"])("rejects a saved document with %s in the final", scenario => {
|
|
const value = sample(false);
|
|
if (scenario === "missing-link" || scenario === "wrong-document") {
|
|
const text = scenario === "missing-link" ? "Saved the requested document." : "Saved [document](/RUN/issues/RUN-1#document-other).";
|
|
(payload(value, 3).item as Row).text = text; value.comments[0]!.body = text;
|
|
}
|
|
if (scenario === "wrong-reply-run") value.comments[0]!.createdByRunId = "other";
|
|
if (scenario === "missing-link-context") delete value.documentLinkContext;
|
|
if (scenario === "missing-revision") value.documentLinkContext!.documents = [{ key: "output", latestRevisionNumber: 1 }];
|
|
expect(gradeNativeCompletion(value).passed).toBe(false);
|
|
});
|
|
it.each([[true, false], [true, true], [false, false], [false, true]])("accepts disposition %s compatibility %s with final before late control-plane acceptance", (blocked, compatibility) => expect(gradeNativeCompletion(sample(blocked, compatibility)).passed).toBe(true));
|
|
it.each(["paperclip_finish", "paperclip_block"])("carries normalized %s identity through rehydration into the exact-call oracle", name => {
|
|
// The Rust normalization calibration asserts these exact fixture bytes.
|
|
const fixture = JSON.parse(readFileSync(new URL("./fixtures/native-completion/terminal-tool-carrier.json", import.meta.url), "utf8")) as {
|
|
cases: Array<{ name: string; normalizedPayload: Row }>;
|
|
};
|
|
const normalized = fixture.cases.find(value => value.name === name)!.normalizedPayload;
|
|
const started = rehydrateRunnerdItemNotification(normalized, "thread", "turn");
|
|
expect(started.item).toMatchObject({ id: "terminal-call", type: "tool_call", name });
|
|
const completed = rehydrateRunnerdItemNotification({ provider: "codex", itemId: "terminal-call", kind: "tool_result", status: "completed", channel: "detail", text: null }, "thread", "turn");
|
|
const value = sample(name === "paperclip_block", true);
|
|
payload(value, 0).item = started.item;
|
|
payload(value, 2).item = completed.item;
|
|
expect(gradeNativeCompletion(value).passed).toBe(true);
|
|
(payload(value, 2).item as Row).id = "different-call";
|
|
expect(gradeNativeCompletion(value).checks.find(check => check.id === "final-after-tool-result")?.passed).toBe(false);
|
|
});
|
|
it("accepts authoritative acceptance before the provider final", () => {
|
|
const value = sample();
|
|
const events = [...value.events];
|
|
value.events = [events[0]!, events[1]!, events[2]!, events[4]!, events[3]!, events[5]!].map((event, index) => ({ ...event, seq: index + 1 }));
|
|
expect(gradeNativeCompletion(value).passed).toBe(true);
|
|
});
|
|
it("accepts complete public streams containing non-PRP lifecycle rows", () => {
|
|
const value = sample();
|
|
value.events = [{ seq: 1, eventType: "lifecycle" }, ...value.events.map(event => ({ ...event, seq: Number(event.seq) + 1 }))];
|
|
expect(gradeNativeCompletion(value).passed).toBe(true);
|
|
});
|
|
it.each(["canonical", "compatibility"])("rejects an unmatched %s terminal result", kind => {
|
|
const value = sample(true, kind === "compatibility");
|
|
if (kind === "canonical") payload(value, 2).executionId = "unmatched";
|
|
else (payload(value, 2).item as Row).id = "unmatched";
|
|
expect(gradeNativeCompletion(value).passed).toBe(false);
|
|
});
|
|
it.each(["missing-name", "unrelated-named-tool", "contradictory-result-name"])("rejects compatibility finishing identity %s despite a matching ID", name => {
|
|
const value = sample(true, true);
|
|
if (name === "missing-name") delete (payload(value, 0).item as Row).name;
|
|
if (name === "unrelated-named-tool") (payload(value, 0).item as Row).name = "write_document";
|
|
if (name === "contradictory-result-name") (payload(value, 2).item as Row).name = "write_document";
|
|
expect(gradeNativeCompletion(value).checks.find(check => check.id === "final-after-tool-result")?.passed).toBe(false);
|
|
});
|
|
it.each(["extra-run", "retry", "continuation", "wrong-account", "wrong-agent", "wrong-disposition", "punctuated-action", "wrong-scope", "missing-final", "summary-fallback", "pre-tool-final", "post-admission-call", "missing-acceptance", "runner-acceptance", "failed-terminal", "event-gap", "binding-mismatch", "extra-task", "extra-agent", "extra-document", "interaction", "changed-workspace", "process-call", "marker-only", "contradiction", "wrong-reply-run"])("rejects %s", name => {
|
|
const value = sample(); const run = value.runs[0]!;
|
|
if (name === "extra-run") value.runs = [run, { ...run, id: "extra" }];
|
|
if (name === "retry") run.retryOfRunId = "prior";
|
|
if (name === "continuation") run.continuationAttempt = 1;
|
|
if (name === "wrong-account") run.companyId = "other";
|
|
if (name === "wrong-agent") run.agentId = "other";
|
|
if (name === "wrong-disposition") value.issue.status = "done";
|
|
if (name === "punctuated-action") ((run.resultJson as Row).nativeResult as {blocker: {unblockAction: string}}).blocker.unblockAction = "Grant deployment access.";
|
|
if (name === "wrong-scope") ((run.resultJson as Row).nativeResult as {blocker: {scope: string}}).blocker.scope = "step";
|
|
if (name === "missing-final" || name === "summary-fallback") payload(value, 3).kind = "summary", (payload(value, 3).item as Row).type = "summary";
|
|
if (name === "pre-tool-final") (payload(value, 2).name = "other");
|
|
if (name === "post-admission-call") value.events[3]!.eventType = "tool.execution.started";
|
|
if (name === "missing-acceptance") value.events[4]!.eventType = "other";
|
|
if (name === "runner-acceptance") ((value.events[4]!.payload as Row).prpEvent as Row).sourceKind = "runner";
|
|
if (name === "failed-terminal") payload(value, 5).runTerminalState = "failed";
|
|
if (name === "event-gap") value.events[1]!.seq = 20;
|
|
if (name === "binding-mismatch") ((value.events[1]!.payload as Row).prpEvent as Row).runId = "other";
|
|
if (name === "extra-task") value.state.issueIds.push("extra");
|
|
if (name === "extra-agent") value.state.agentIds.push("extra");
|
|
if (name === "extra-document") value.state.documentCount = 1;
|
|
if (name === "interaction") value.state.interactionCount = 1;
|
|
if (name === "changed-workspace") value.workspaceChanged = true;
|
|
if (name === "process-call") payload(value, 0).transport = "process";
|
|
if (name === "marker-only" || name === "contradiction") {
|
|
const text = name === "marker-only" ? "BLOCKED_probe" : "Deployment is not blocked. Release Owner completed Grant deployment access. BLOCKED_probe";
|
|
(payload(value, 3).item as Row).text = text; value.comments[0]!.body = text;
|
|
}
|
|
if (name === "wrong-reply-run") value.comments[0]!.createdByRunId = "other";
|
|
expect(gradeNativeCompletion(value).passed).toBe(false);
|
|
});
|
|
it.each([undefined, "other-issue"])("rejects missing or wrong native issue attribution %s", nativeIssueId => {
|
|
const value = sample(); value.runs[0]!.nativeIssueId = nativeIssueId;
|
|
expect(gradeNativeCompletion(value).passed).toBe(false);
|
|
});
|
|
it.each(["missing-action-envelope", "wrong-action-source", "missing-source-identity", "extra-wrong-source-result"])("rejects %s", name => {
|
|
const value = sample();
|
|
if (name === "missing-action-envelope") value.events[0]!.payload = {};
|
|
if (name === "wrong-action-source") ((value.events[0]!.payload as Row).prpEvent as Row).sourceKind = "control_plane";
|
|
if (name === "missing-source-identity") {
|
|
delete value.events[0]!.sourceEventId;
|
|
delete ((value.events[0]!.payload as Row).prpEvent as Row).sourceEventId;
|
|
}
|
|
if (name === "extra-wrong-source-result") {
|
|
const extra = structuredClone(value.events[4]!);
|
|
((extra.payload as Row).prpEvent as Row).sourceKind = "runner";
|
|
value.events = [...value.events, { ...extra, seq: 7 }];
|
|
}
|
|
expect(gradeNativeCompletion(value).passed).toBe(false);
|
|
});
|
|
});
|