mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native agents receive task constraints and completion tools from Paperclip. > - Completion tools already define the procedure for reporting a result. > - Repeated procedure text adds instructions to each full task turn. > - The final reply must still explain a blocker and link a saved document. > - This pull request removes repeated procedure text and keeps these visible outcome requirements explicit. > - A document receipt supplies the exact link, and stricter evals check the persisted reply and browser navigation. ## Linked Issues or Issue Description Refs: #14961. Related: #14948 and #15007. **What happened?** Native task envelopes repeat completion procedure text. A reduced envelope needs explicit final-reply requirements. The `write_document` receipt also lacks a canonical document link. **Expected behavior** Keep the completion tools as the source of procedure details. Require one accepted completion result before the final reply. A blocked reply must explain the reason, owner and unblock action. A document reply must contain a working link to the saved document. **Steps to reproduce** 1. Run the native assigned-skill document case and native blocker case. 2. Inspect the run-attributed provider final and its persisted comment. 3. Check the blocker explanation or open the final reply's document link. ## What Changed - Remove repeated completion procedure text from the native task constraints and backend instructions. - Keep explicit blocker and document-link requirements in full task turns. - Return a company/task-scoped `documentHref` from `write_document`. Preserve the link in the idempotent mutation receipt. - Repeat canonical links for this run's current saved revisions in accepted completion feedback. Give blocked providers final-response guidance for the cause, owner and unblock action. - Keep internal document/comment anchors when Markdown issue links load cached issue details. - Add a manual six-cell comparison suite with strict source, build, default-instruction and budget admission. - Capture eighteen shared runnerd RPC projections and six direct OpenCode HTTP projections across start, resume and continuation phases, using scripted local transports and no provider execution. - Apply v3 checks only to the manual instruction comparison; preserve v2 checks for the existing native completion suite. Check the actual persisted blocker reason and exact saved-document link. Click the rendered document link and check the original content marker in the classic document card or the new document tab. - Forward exact OpenCode finishing calls through the controller. Wait for acceptance, keep accepted feedback and concrete rejection text, and reject malformed responses. Preserve ordinary dynamic-tool response handling. - Settle the completion decision and tool response before mapping a racing idle/error/abort event or handling explicit close/interruption. Reject a concurrent finishing call before controller admission. - Add a provider-free regression through real runnerd, the OpenCode proxy and a fake provider. Reject the first completion, accept the corrected report in the same turn, and propose one result. - Keep all original verdicts unchanged. Treat replay under new checks as separate diagnostics. ## Verification - `pnpm -r typecheck` and `pnpm build` pass locally. - Native document-authority tests pass, including company/run authorization and idempotent replay. - Native runtime-context, backend and measurement tests pass. - Final-answer calibration, protocol scoring, source-admission and catalog tests pass. Wrong reasons, absent links and wrong link targets fail. - `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six single-attempt local cells with the declared models. - Exported `prepareNativeInstructionPreflight` then `verifyNativeInstructionPreflight` pass on this clean committed source. They build locally and make zero provider calls. - Corrective live confirmation is incomplete. Source3a7349dpassed both Claude and both Codex cases. OpenCode saved the correct document but omitted its final link; its blocker case was canceled before paid execution. Preserve this failure. Thee171282confirmation was stopped during build after fresh review found a completion-settlement race; it executed zero providers. Sourcec3e0cb303fixes that race. Two affected OpenCode cases await fresh review and one bounded confirmation; earlier results remain attributed to their original source. - OpenCode proxy parsing, driver, factory and input tests: 81 pass across retained focused runs, including six settlement races. Evaluator/scoring/admission checks: 122 pass. The real proxy regression passes. Fresh local prepare then verify passes with 18 shared and 6 direct scripted captures, fresh SDK/Rust builds and zero providers. - The full local suite recorded two failures: a webhook timeout and a Git-scan load count mismatch. Both files pass in isolation with unchanged assertions/time budgets; preserve the original failure log. Freshc3e0cb303CI and review are pending. This PR remains draft. ## Risks - Final-answer wording can vary by provider. The checks cover the declared release-access blocker and saved document fixture, not general answer quality. - A single trial does not establish general equivalence, cause, speed, cost or live resume behavior. - `documentHref` is an additive receipt field. It points to the current saved document, not an immutable historical revision. Replaying an older receipt does not fabricate a new link. - The correction adds four production paths for document receipts, accepted completion feedback and UI navigation, plus four OpenCode controller/proxy paths, beyond the original three instruction paths. Completion rejection must remain repairable; the production-boundary regression covers it. - Preserve the frozen comparison context for live measurement. A merge-tree check against current master is clean. Do not relabel earlier live results as results from a later source tree. ## Model Used - OpenAI Codex, GPT-6 family. The exact serving model ID and context window are unavailable in this session. Capabilities used: reasoning, code editing, shell execution, test authoring and evidence review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
48 lines
3.6 KiB
TypeScript
48 lines
3.6 KiB
TypeScript
import { createHash } from "node:crypto";
|
|
import { readFileSync } from "node:fs";
|
|
import type { RunnerTaskFixture } from "./types.js";
|
|
import { contextIntegrityTasks } from "./context-integrity-cases.js";
|
|
|
|
export const nativeBlockedReportTask: RunnerTaskFixture = {
|
|
id: "native-blocked-report", label: "Native concrete blocker report", groups: ["native-completion"],
|
|
workMode: "standard", flow: "single_turn", expectedRunCount: 1, automaticRetryPolicy: "single_attempt",
|
|
attemptTimeoutMs: { local: 5 * 60_000, daytona: 5 * 60_000 },
|
|
expectedTerminalState: { issue: "blocked", run: "succeeded" },
|
|
buildTitle: nonce => `Native blocker report ${nonce}`,
|
|
buildVisibleMarker: nonce => `BLOCKED_${nonce.replace(/[^a-z0-9]/gi, "")}`,
|
|
buildPrompt: nonce => [
|
|
"Deployment cannot proceed because required release access has not been granted.",
|
|
"Release Owner is the external owner of that grant. Do not attempt deployment, run shell/process commands, change files, invoke external services, create documents, agents, interactions or subtasks, or claim the access exists.",
|
|
'Record the whole task as blocked and identify the owner as Release Owner. The exact unblock action is "Grant deployment access" (without quotation marks or trailing punctuation).',
|
|
`Include BLOCKED_${nonce.replace(/[^a-z0-9]/gi, "")} exactly once in your final explanation.`,
|
|
].join("\n"),
|
|
buildMatchers: nonce => [
|
|
{ kind: "issue_status", expected: "blocked" }, { kind: "run_status", expected: "succeeded" },
|
|
{ kind: "runtime_mode", expected: "native" },
|
|
{ kind: "message_contains", expected: `BLOCKED_${nonce.replace(/[^a-z0-9]/gi, "")}` },
|
|
{ kind: "message_contains", expected: "Release Owner" },
|
|
{ kind: "message_contains", expected: "Grant deployment access" },
|
|
{ kind: "message_regex", pattern: "^(?![\\s\\S]*\\b(?:not blocked|no longer blocked|access (?:is |has been |was )?already granted|completed Grant deployment access)\\b)[\\s\\S]*\\b(?:blocked|cannot proceed|can't proceed|missing|required access|not (?:yet )?granted|awaiting|waiting|unavailable)\\b", flags: "i" },
|
|
{ kind: "json_path", path: "run.resultJson.nativeResult.reportedWorkDisposition", expected: "blocked" },
|
|
{ kind: "json_path", path: "run.resultJson.nativeResult.blocker.owner.name", expected: "Release Owner" },
|
|
{ kind: "json_path", path: "run.resultJson.nativeResult.blocker.unblockAction", expected: "Grant deployment access" },
|
|
{ kind: "json_path", path: "run.resultJson.nativeResult.blocker.scope", expected: "task_wide" },
|
|
],
|
|
};
|
|
|
|
const assignedSkill = contextIntegrityTasks.find(task => task.id === "assigned-skill-explicit-invocation");
|
|
if (!assignedSkill) throw new Error("Missing original assigned-skill completion journey");
|
|
// Preserve the original task prompt, skill, durable-document oracle and deadline.
|
|
export const nativeCompletionTasks: readonly RunnerTaskFixture[] = [
|
|
{ ...assignedSkill, automaticRetryPolicy: "single_attempt" }, nativeBlockedReportTask,
|
|
];
|
|
export function nativeCompletionDefinitionDigest() {
|
|
const hash = createHash("sha256");
|
|
for (const file of ["native-completion-cases.ts", "native-completion-scoring.ts", "native-completion-content.ts", "native-completion-defaults.ts",
|
|
"native-blocker-visible.ts", "native-completion-admission.ts", "native-completion-checks.mjs",
|
|
"native-completion-source-contract.mjs", "automatic-retry.ts", "context-integrity-cases.ts", "context-integrity-flow.ts", "context-integrity-scoring.ts", "types.ts", "catalog.ts", "live-fixtures.ts", "runner.spec.ts", "launch.ts"]) {
|
|
hash.update(file); hash.update(readFileSync(new URL(file, import.meta.url)));
|
|
}
|
|
return hash.digest("hex");
|
|
}
|