4.2 KiB
Native completion final-answer correction
The completion constraint reduction remains held until the corrected source has live evidence for final-answer quality. Keep the original candidate, historical source and all original grades immutable.
The correction keeps the shortened completion procedure while explicitly requiring a blocker explanation, owner and unblock action in the final reply. The existing document guidance now requires a canonical clickable document link. write_document returns a company/task-scoped documentHref in its idempotent receipt, derived in the document transaction from the saved key and actual task identity. No completion tool schema/description, fixed prompt, skill, permission or budget policy changes. The accepted completion feedback repeats canonical links only for this run's still-current saved document revisions, reconstructed from company/task-scoped database state rather than supplied URLs. Blocker feedback asks for the cause, owner and unblock action instead of describing completed work. Internal Markdown links retain their query and document/comment anchor when cached issue details resolve. These receipt, feedback and UI repairs add four production paths beyond the original three-file instruction reduction.
Only the manual instruction comparison uses the v3 final-answer observation. The existing native-completion suite retains its v2 verdict checks and does not acquire the browser-navigation requirement. A calibration demonstrates the same retained input can pass v2 and fail v3; original report files are never rewritten. The harness definition digest changes truthfully for future invocations. Its additional checks read only the actual run-attributed persisted provider final. A blocker label and action do not substitute for an explicit missing-access explanation. Document links must identify the exact task and one revisioned document at the instance origin; wrong origin, task, document, query, credentials or absent evidence fail. Exact call-ID joins, acceptance/termination, ordering, side-effect and budget checks remain enforced.
After saving the strict native snapshot, each completion cell clicks the rendered final-reply link in the browser. It must open the canonical saved document and show the original content marker in exactly one visible classic document card or document-specific side-panel tab. The navigation check follows the actual configured UI instead of requiring a classic-only element. Retain the navigation receipt and screenshot independently; a valid-looking Markdown URL alone does not qualify navigation.
Calibrate with correct paraphrases and plausible wrong answers, including the action-only omission. Replay the retained observations only as labeled additional diagnostics; never overwrite or regrade the original results. Provider-free measurement v2 captures eighteen shared runnerd RPC projections plus six direct OpenCode HTTP projections across v4/v5 start, resume and continuation. The latter exercises the concrete OpenCode backend using a local fake server. Both boundaries measure Paperclip-supplied payloads, not provider stock prompts or model cognition. No synthetic receipt qualifies live model behavior.
The human's paid-run authorization and subsequent request to fix the omission support one bounded corrective confirmation: the six exact native-instruction-consolidation candidate cells, one attempt each, no automatic retries, native Codex gpt-5.6-sol, ACPX Claude claude-sonnet-5, and OpenCode openrouter/deepseek/deepseek-v4-flash-0731, original local deadlines, and 1,000-cent company/agent hard stops. Each production correction must have a new immutable source/definition admission before another bounded confirmation. Preserve stopped campaigns, failed cells and canceled-before-provider cells separately. Never reroll an unchanged frozen failure. No new baseline provider run, broader campaign or merge is part of this correction. Stop and inspect a usable behavior failure.
Report exact source, definition/grader hashes, six outcomes, retained visible replies and document targets, actual run count, cleanup and billing coverage. Passing a single trial does not prove general equivalence, cause, live resume behavior, coding quality, speed or cost trends.