Files
PaperClipAI/tests/runner-e2e/NATIVE-COMPLETION.md

3.5 KiB

Native completion guidance qualification

This explicit-only suite independently qualifies PR #14961 on master 59c07ede72dc08b8aba149a01cc11e0b7a204621. It does not include PR #14948's reduced manuals, shared prompts or operational-skill changes. The earlier comparison on reduced defaults remains historical evidence, not admission for this context.

The six local cells pair the original assigned-skill-explicit-invocation journey and native-blocked-report across runner-codex, runner-acpx-claude and runner-opencode. Completion keeps the original task, installed output skill, explicit invocation and independently saved issue document checks. Blocking checks the durable whole-task blocker, exact external owner/action, explanatory visible reply, unchanged workspace and absence of extra tasks, agents, documents or interactions.

Both use the actual production default hire: no custom QA instructions bundle. Before execution, the public bundle receipt must hash to the unchanged master AGENTS.md, and company/agent budgets must each be 1,000 cents. Profile model, effort, credentials, skills, permissions and provider configuration are inherited.

The native oracle requires exactly one succeeded native run attributed to this issue, bound complete public events, one runner semantic proposal and authoritative control-plane acceptance/terminal. A matched finishing tool start/result precedes the actual provider final; no new invocation follows the proposal. The final must be persisted as the run's agent reply. Acceptance may precede or follow the final. This proves observable ordering and durable disposition; it does not prove the provider consumed feedback or establish general coding-quality equivalence.

single_attempt is enforced by the launcher even when hosted invocation only supplies --id. Every first failure, usage and cleanup receipt is retained; no automatic infrastructure or behavioral retry occurs. The bounded comparison is six cells per variant, twelve expected provider turns total, with the existing per-cell deadlines (skill: original deadline; blocker: five minutes).

Candidate and baseline use identical executable fixtures and public oracles. Only the five native description/fingerprint source files and their six variant-specific unit tests may differ. The historical variant restores those eleven files from master, including v13; candidate uses the archived native change and v14. Source-contract hashes reject mixed revisions and changes to master defaults, auth/runtime context or shared instructions. Derived capability metadata is independently checked; it currently requires no regeneration in either variant.

Credential-free admission runs before local environment/credential loading:

node tests/runner-e2e/native-completion-checks.mjs --output-dir=/absolute/report/path
node tests/runner-e2e/native-completion-checks.mjs --verify=/absolute/report/path/preflight.json
pnpm test:e2e:runner -- --list --suite native-completion

Admission requires a committed clean exact source, variant-appropriate schema, tool bridge/catalog and resume tests, independent oracle/default/retry calibrations, original context-integrity calibrations, typecheck, manifest checks and exact six-cell discovery. Generic stock-suite gates remain unchanged. Paid hosted runs must dispatch the trusted default-branch workflow to distinct immutable target branches using the exact six IDs, preserving authorization and secret boundaries. Live results are pending until retained cell evidence is published.