Files
PaperClipAI/tests/runner-e2e/native-completion-cases.ts
DottaandPaperclip a386a59998 Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native agents receive task constraints and completion tools from
Paperclip.
> - Completion tools already define the procedure for reporting a
result.
> - Repeated procedure text adds instructions to each full task turn.
> - The final reply must still explain a blocker and link a saved
document.
> - This pull request removes repeated procedure text and keeps these
visible outcome requirements explicit.
> - A document receipt supplies the exact link, and stricter evals check
the persisted reply and browser navigation.

## Linked Issues or Issue Description

Refs: #14961. Related: #14948 and #15007.

**What happened?**

Native task envelopes repeat completion procedure text. A reduced
envelope needs explicit final-reply requirements. The `write_document`
receipt also lacks a canonical document link.

**Expected behavior**

Keep the completion tools as the source of procedure details. Require
one accepted completion result before the final reply. A blocked reply
must explain the reason, owner and unblock action. A document reply must
contain a working link to the saved document.

**Steps to reproduce**

1. Run the native assigned-skill document case and native blocker case.
2. Inspect the run-attributed provider final and its persisted comment.
3. Check the blocker explanation or open the final reply's document
link.

## What Changed

- Remove repeated completion procedure text from the native task
constraints and backend instructions.
- Keep explicit blocker and document-link requirements in full task
turns.
- Return a company/task-scoped `documentHref` from `write_document`.
Preserve the link in the idempotent mutation receipt.
- Repeat canonical links for this run's current saved revisions in
accepted completion feedback. Give blocked providers final-response
guidance for the cause, owner and unblock action.
- Keep internal document/comment anchors when Markdown issue links load
cached issue details.
- Add a manual six-cell comparison suite with strict source, build,
default-instruction and budget admission.
- Capture eighteen shared runnerd RPC projections and six direct
OpenCode HTTP projections across start, resume and continuation phases,
using scripted local transports and no provider execution.
- Apply v3 checks only to the manual instruction comparison; preserve v2
checks for the existing native completion suite. Check the actual
persisted blocker reason and exact saved-document link. Click the
rendered document link and check the original content marker in the
classic document card or the new document tab.
- Forward exact OpenCode finishing calls through the controller. Wait
for acceptance, keep accepted feedback and concrete rejection text, and
reject malformed responses. Preserve ordinary dynamic-tool response
handling.
- Settle the completion decision and tool response before mapping a
racing idle/error/abort event or handling explicit close/interruption.
Reject a concurrent finishing call before controller admission.
- Add a provider-free regression through real runnerd, the OpenCode
proxy and a fake provider. Reject the first completion, accept the
corrected report in the same turn, and propose one result.
- Keep all original verdicts unchanged. Treat replay under new checks as
separate diagnostics.

## Verification

- `pnpm -r typecheck` and `pnpm build` pass locally.
- Native document-authority tests pass, including company/run
authorization and idempotent replay.
- Native runtime-context, backend and measurement tests pass.
- Final-answer calibration, protocol scoring, source-admission and
catalog tests pass. Wrong reasons, absent links and wrong link targets
fail.
- `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six
single-attempt local cells with the declared models.
- Exported `prepareNativeInstructionPreflight` then
`verifyNativeInstructionPreflight` pass on this clean committed source.
They build locally and make zero provider calls.
- Corrective live confirmation is incomplete. Source 3a7349d passed both
Claude and both Codex cases. OpenCode saved the correct document but
omitted its final link; its blocker case was canceled before paid
execution. Preserve this failure. The e171282 confirmation was stopped
during build after fresh review found a completion-settlement race; it
executed zero providers. Source c3e0cb303 fixes that race. Two affected
OpenCode cases await fresh review and one bounded confirmation; earlier
results remain attributed to their original source.
- OpenCode proxy parsing, driver, factory and input tests: 81 pass
across retained focused runs, including six settlement races.
Evaluator/scoring/admission checks: 122 pass. The real proxy regression
passes. Fresh local prepare then verify passes with 18 shared and 6
direct scripted captures, fresh SDK/Rust builds and zero providers.
- The full local suite recorded two failures: a webhook timeout and a
Git-scan load count mismatch. Both files pass in isolation with
unchanged assertions/time budgets; preserve the original failure log.
Fresh c3e0cb303 CI and review are pending. This PR remains draft.

## Risks

- Final-answer wording can vary by provider. The checks cover the
declared release-access blocker and saved document fixture, not general
answer quality.
- A single trial does not establish general equivalence, cause, speed,
cost or live resume behavior.
- `documentHref` is an additive receipt field. It points to the current
saved document, not an immutable historical revision. Replaying an older
receipt does not fabricate a new link.
- The correction adds four production paths for document receipts,
accepted completion feedback and UI navigation, plus four OpenCode
controller/proxy paths, beyond the original three instruction paths.
Completion rejection must remain repairable; the production-boundary
regression covers it.
- Preserve the frozen comparison context for live measurement. A
merge-tree check against current master is clean. Do not relabel earlier
live results as results from a later source tree.

## Model Used

- OpenAI Codex, GPT-6 family. The exact serving model ID and context
window are unavailable in this session. Capabilities used: reasoning,
code editing, shell execution, test authoring and evidence review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 08:10:14 -05:00

48 lines
3.6 KiB
TypeScript

import { createHash } from "node:crypto";
import { readFileSync } from "node:fs";
import type { RunnerTaskFixture } from "./types.js";
import { contextIntegrityTasks } from "./context-integrity-cases.js";
export const nativeBlockedReportTask: RunnerTaskFixture = {
id: "native-blocked-report", label: "Native concrete blocker report", groups: ["native-completion"],
workMode: "standard", flow: "single_turn", expectedRunCount: 1, automaticRetryPolicy: "single_attempt",
attemptTimeoutMs: { local: 5 * 60_000, daytona: 5 * 60_000 },
expectedTerminalState: { issue: "blocked", run: "succeeded" },
buildTitle: nonce => `Native blocker report ${nonce}`,
buildVisibleMarker: nonce => `BLOCKED_${nonce.replace(/[^a-z0-9]/gi, "")}`,
buildPrompt: nonce => [
"Deployment cannot proceed because required release access has not been granted.",
"Release Owner is the external owner of that grant. Do not attempt deployment, run shell/process commands, change files, invoke external services, create documents, agents, interactions or subtasks, or claim the access exists.",
'Record the whole task as blocked and identify the owner as Release Owner. The exact unblock action is "Grant deployment access" (without quotation marks or trailing punctuation).',
`Include BLOCKED_${nonce.replace(/[^a-z0-9]/gi, "")} exactly once in your final explanation.`,
].join("\n"),
buildMatchers: nonce => [
{ kind: "issue_status", expected: "blocked" }, { kind: "run_status", expected: "succeeded" },
{ kind: "runtime_mode", expected: "native" },
{ kind: "message_contains", expected: `BLOCKED_${nonce.replace(/[^a-z0-9]/gi, "")}` },
{ kind: "message_contains", expected: "Release Owner" },
{ kind: "message_contains", expected: "Grant deployment access" },
{ kind: "message_regex", pattern: "^(?![\\s\\S]*\\b(?:not blocked|no longer blocked|access (?:is |has been |was )?already granted|completed Grant deployment access)\\b)[\\s\\S]*\\b(?:blocked|cannot proceed|can't proceed|missing|required access|not (?:yet )?granted|awaiting|waiting|unavailable)\\b", flags: "i" },
{ kind: "json_path", path: "run.resultJson.nativeResult.reportedWorkDisposition", expected: "blocked" },
{ kind: "json_path", path: "run.resultJson.nativeResult.blocker.owner.name", expected: "Release Owner" },
{ kind: "json_path", path: "run.resultJson.nativeResult.blocker.unblockAction", expected: "Grant deployment access" },
{ kind: "json_path", path: "run.resultJson.nativeResult.blocker.scope", expected: "task_wide" },
],
};
const assignedSkill = contextIntegrityTasks.find(task => task.id === "assigned-skill-explicit-invocation");
if (!assignedSkill) throw new Error("Missing original assigned-skill completion journey");
// Preserve the original task prompt, skill, durable-document oracle and deadline.
export const nativeCompletionTasks: readonly RunnerTaskFixture[] = [
{ ...assignedSkill, automaticRetryPolicy: "single_attempt" }, nativeBlockedReportTask,
];
export function nativeCompletionDefinitionDigest() {
const hash = createHash("sha256");
for (const file of ["native-completion-cases.ts", "native-completion-scoring.ts", "native-completion-content.ts", "native-completion-defaults.ts",
"native-blocker-visible.ts", "native-completion-admission.ts", "native-completion-checks.mjs",
"native-completion-source-contract.mjs", "automatic-retry.ts", "context-integrity-cases.ts", "context-integrity-flow.ts", "context-integrity-scoring.ts", "types.ts", "catalog.ts", "live-fixtures.ts", "runner.spec.ts", "launch.ts"]) {
hash.update(file); hash.update(readFileSync(new URL(file, import.meta.url)));
}
return hash.digest("hex");
}