## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native agents receive task constraints and completion tools from Paperclip. > - Completion tools already define the procedure for reporting a result. > - Repeated procedure text adds instructions to each full task turn. > - The final reply must still explain a blocker and link a saved document. > - This pull request removes repeated procedure text and keeps these visible outcome requirements explicit. > - A document receipt supplies the exact link, and stricter evals check the persisted reply and browser navigation. ## Linked Issues or Issue Description Refs: #14961. Related: #14948 and #15007. **What happened?** Native task envelopes repeat completion procedure text. A reduced envelope needs explicit final-reply requirements. The `write_document` receipt also lacks a canonical document link. **Expected behavior** Keep the completion tools as the source of procedure details. Require one accepted completion result before the final reply. A blocked reply must explain the reason, owner and unblock action. A document reply must contain a working link to the saved document. **Steps to reproduce** 1. Run the native assigned-skill document case and native blocker case. 2. Inspect the run-attributed provider final and its persisted comment. 3. Check the blocker explanation or open the final reply's document link. ## What Changed - Remove repeated completion procedure text from the native task constraints and backend instructions. - Keep explicit blocker and document-link requirements in full task turns. - Return a company/task-scoped `documentHref` from `write_document`. Preserve the link in the idempotent mutation receipt. - Repeat canonical links for this run's current saved revisions in accepted completion feedback. Give blocked providers final-response guidance for the cause, owner and unblock action. - Keep internal document/comment anchors when Markdown issue links load cached issue details. - Add a manual six-cell comparison suite with strict source, build, default-instruction and budget admission. - Capture eighteen shared runnerd RPC projections and six direct OpenCode HTTP projections across start, resume and continuation phases, using scripted local transports and no provider execution. - Apply v3 checks only to the manual instruction comparison; preserve v2 checks for the existing native completion suite. Check the actual persisted blocker reason and exact saved-document link. Click the rendered document link and check the original content marker in the classic document card or the new document tab. - Forward exact OpenCode finishing calls through the controller. Wait for acceptance, keep accepted feedback and concrete rejection text, and reject malformed responses. Preserve ordinary dynamic-tool response handling. - Settle the completion decision and tool response before mapping a racing idle/error/abort event or handling explicit close/interruption. Reject a concurrent finishing call before controller admission. - Add a provider-free regression through real runnerd, the OpenCode proxy and a fake provider. Reject the first completion, accept the corrected report in the same turn, and propose one result. - Keep all original verdicts unchanged. Treat replay under new checks as separate diagnostics. ## Verification - `pnpm -r typecheck` and `pnpm build` pass locally. - Native document-authority tests pass, including company/run authorization and idempotent replay. - Native runtime-context, backend and measurement tests pass. - Final-answer calibration, protocol scoring, source-admission and catalog tests pass. Wrong reasons, absent links and wrong link targets fail. - `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six single-attempt local cells with the declared models. - Exported `prepareNativeInstructionPreflight` then `verifyNativeInstructionPreflight` pass on this clean committed source. They build locally and make zero provider calls. - Corrective live confirmation is incomplete. Source3a7349dpassed both Claude and both Codex cases. OpenCode saved the correct document but omitted its final link; its blocker case was canceled before paid execution. Preserve this failure. Thee171282confirmation was stopped during build after fresh review found a completion-settlement race; it executed zero providers. Sourcec3e0cb303fixes that race. Two affected OpenCode cases await fresh review and one bounded confirmation; earlier results remain attributed to their original source. - OpenCode proxy parsing, driver, factory and input tests: 81 pass across retained focused runs, including six settlement races. Evaluator/scoring/admission checks: 122 pass. The real proxy regression passes. Fresh local prepare then verify passes with 18 shared and 6 direct scripted captures, fresh SDK/Rust builds and zero providers. - The full local suite recorded two failures: a webhook timeout and a Git-scan load count mismatch. Both files pass in isolation with unchanged assertions/time budgets; preserve the original failure log. Freshc3e0cb303CI and review are pending. This PR remains draft. ## Risks - Final-answer wording can vary by provider. The checks cover the declared release-access blocker and saved document fixture, not general answer quality. - A single trial does not establish general equivalence, cause, speed, cost or live resume behavior. - `documentHref` is an additive receipt field. It points to the current saved document, not an immutable historical revision. Replaying an older receipt does not fabricate a new link. - The correction adds four production paths for document receipts, accepted completion feedback and UI navigation, plus four OpenCode controller/proxy paths, beyond the original three instruction paths. Completion rejection must remain repairable; the production-boundary regression covers it. - Preserve the frozen comparison context for live measurement. A merge-tree check against current master is clean. Do not relabel earlier live results as results from a later source tree. ## Model Used - OpenAI Codex, GPT-6 family. The exact serving model ID and context window are unavailable in this session. Capabilities used: reasoning, code editing, shell execution, test authoring and evidence review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
8.9 KiB
Native completion final-answer correction
The completion constraint reduction remains held until the corrected source has live evidence for final-answer quality. Keep the original candidate, historical source and all original grades immutable.
The correction keeps the shortened completion procedure while explicitly requiring a blocker explanation, owner and unblock action in the final reply. The existing document guidance now requires a canonical clickable document link. write_document returns a company/task-scoped documentHref in its idempotent receipt, derived in the document transaction from the saved key and actual task identity. No completion tool schema/description, fixed prompt, skill, permission or budget policy changes. The accepted completion feedback repeats canonical links only for this run's still-current saved document revisions, reconstructed from company/task-scoped database state rather than supplied URLs. Blocker feedback asks for the cause, owner and unblock action instead of describing completed work. Internal Markdown links retain their query and document/comment anchor when cached issue details resolve. These receipt, feedback and UI repairs add four production paths beyond the original three-file instruction reduction.
Only the manual instruction comparison uses the v3 final-answer observation. The existing native-completion suite retains its v2 verdict checks and does not acquire the browser-navigation requirement. A calibration demonstrates the same retained input can pass v2 and fail v3; original report files are never rewritten. The harness definition digest changes truthfully for future invocations. Its additional checks read only the actual run-attributed persisted provider final. A blocker label and action do not substitute for an explicit missing-access explanation. Document links must identify the exact task and one revisioned document at the instance origin; wrong origin, task, document, query, credentials or absent evidence fail. Exact call-ID joins, acceptance/termination, ordering, side-effect and budget checks remain enforced.
After saving the strict native snapshot, each completion cell clicks the rendered final-reply link in the browser. It must open the canonical saved document and show the original content marker in exactly one visible classic document card or document-specific side-panel tab. The navigation check follows the actual configured UI instead of requiring a classic-only element. Retain the navigation receipt and screenshot independently; a valid-looking Markdown URL alone does not qualify navigation.
Calibrate with correct paraphrases and plausible wrong answers, including the action-only omission. Replay the retained observations only as labeled additional diagnostics; never overwrite or regrade the original results. Provider-free measurement v2 captures eighteen shared runnerd RPC projections plus six direct OpenCode HTTP projections across v4/v5 start, resume and continuation. The latter exercises the concrete OpenCode backend using a local fake server. Both boundaries measure Paperclip-supplied payloads, not provider stock prompts or model cognition. No synthetic receipt qualifies live model behavior.
The human's paid-run authorization and subsequent request to fix the omission support one bounded corrective confirmation: the six exact native-instruction-consolidation candidate cells, one attempt each, no automatic retries, native Codex gpt-5.6-sol, ACPX Claude claude-sonnet-5, and OpenCode openrouter/deepseek/deepseek-v4-flash-0731, original local deadlines, and 1,000-cent company/agent hard stops. Each production correction must have a new immutable source/definition admission before another bounded confirmation. Preserve stopped campaigns, failed cells and canceled-before-provider cells separately. Never reroll an unchanged frozen failure. No new baseline provider run, broader campaign or merge is part of this correction. Stop and inspect a usable behavior failure.
Report exact source, definition/grader hashes, six outcomes, retained visible replies and document targets, actual run count, cleanup and billing coverage. Passing a single trial does not prove general equivalence, cause, live resume behavior, coding quality, speed or cost trends.
The receipt-only correction at be604fcbf81ff238ee38774ff6e644bf0d84eb2d failed its first Claude document cell; five remaining jobs were canceled before paid execution. The next correction at 3a7349ddc60142283c394121882288f2db07b215 passed both Claude and both Codex cases, but OpenCode again omitted the document link; its blocker job was canceled with the paid step skipped. Preserve both genuine failures and all canceled jobs. Browser navigation was not reached in either missing-link failure.
OpenCode's inner driver admitted the finishing report locally without invoking the controller feedback handler. Its proxy also discarded the controller response envelope, including the success flag. The correction forwards the exact terminal call ID, thread, turn, tool and structured result through the existing controller request channel. It waits for the controller before committing the inner result, preserves accepted response text and concrete rejection reasons, and rejects malformed acceptance or a turn that ended during feedback. Dynamic-tool unwrapping stays unchanged. A provider-free test crosses the real runnerd, proxy and fake OpenCode boundary: the first completion is rejected, the provider corrects it in the same turn, the second receives the exact final-response feedback, and only one semantic result is proposed.
This correction changes four additional OpenCode-only production paths. The shared instruction, server receipt/feedback and UI bytes remain identical to 3a7349d. Admit the new source and measure all scripted boundaries again. Confirm only the two affected OpenCode live cases on the new immutable source, one attempt each with the original model, deadlines and budget stops. Keep the four passing Claude/Codex observations explicitly attributed to 3a7349d, backed by the exact provider-path delta and final-head CI. Do not relabel them as live runs on the new head. No new baseline or unchanged-source paid retry is included.
The e171282 confirmation campaign 37257534042 was stopped during build after fresh review identified a controller/inner-session acceptance race. No paid step started and no provider was executed. Provider-free regression reproduces accepted feedback losing its result when idle/error arrives in flight. Completion settlement now holds terminal mapping, explicit interruption and close until the bound decision and tool result finish, and refuses a concurrent terminal submission before invoking the controller. Raw SSE frames remain retained at receipt; normalized terminal settlement follows the finishing decision. Confirm the same two OpenCode cases only on the newly admitted immutable source; preserve e171282's completed failing review and zero-provider campaign.
Fresh c3e0cb303 review identified a real serial-input dependency: an interrupt waits for completion while its controller response queues behind the interrupt. No paid campaign was admitted on this source. The proxy now handles only bound response frames outside the serial command queue; commands retain their order and bootstrap failure gates. EOF rejects existing controller waiters and prevents new unanswered requests, then drains the command queue and closes. A real production-bundled proxy, verified local native launcher and fake OpenCode MCP/SSE server reproduce the deadlock before the fix and verify accepted feedback reaches the provider, interruption completes, exactly one result is proposed and EOF exits. A second EOF-without-feedback case proposes no result. This is zero-provider proxy coverage; the real runnerd boundary remains separately covered. Fixture setup and macOS launcher failures are retained, not relabeled as behavioral failures.
The exact 267aeb3 source passed fresh 5/5 review, but confirmation campaign 37259746983 stopped before provider execution. Its document job entered the harness step, then failed source ancestry hydration before result creation: the fixed eight-commit fetch did not contain the required 2a8a99 base. The remaining blocker cell was canceled. The source was not behaviorally evaluated, and no graded result or cell artifact exists; retain the original failed job log and separate zero-provider receipt.
Hosted hydration now fetches only the immutable public source at bounded depths 8, 32 and 128. It verifies exact HEAD after every fetch and still requires real base ancestry. A missing or unrelated ancestor, changed HEAD, local checkout or non-shallow checkout does not receive a substitute proof. A real twelve-commit shallow-clone regression fails before the change and passes afterward. Production, fixture prompts, model/default/budget policy and behavioral grading bytes stay unchanged from 267aeb3. Admit the new setup source and confirm the same two OpenCode cases once; no provider run or original failure is retried.