Files
PaperClipAI/tests/runner-e2e/NATIVE-COMPLETION.md

59 lines
3.5 KiB
Markdown

# Native completion guidance qualification
This explicit-only suite independently qualifies PR #14961 on master
`59c07ede72dc08b8aba149a01cc11e0b7a204621`. It does not include PR #14948's
reduced manuals, shared prompts or operational-skill changes. The earlier
comparison on reduced defaults remains historical evidence, not admission for
this context.
The six local cells pair the original `assigned-skill-explicit-invocation`
journey and `native-blocked-report` across `runner-codex`,
`runner-acpx-claude` and `runner-opencode`. Completion keeps the original task,
installed output skill, explicit invocation and independently saved issue
document checks. Blocking checks the durable whole-task blocker, exact external
owner/action, explanatory visible reply, unchanged workspace and absence of
extra tasks, agents, documents or interactions.
Both use the actual production default hire: no custom QA instructions bundle.
Before execution, the public bundle receipt must hash to the unchanged master
AGENTS.md, and company/agent budgets must each be 1,000 cents. Profile model,
effort, credentials, skills, permissions and provider configuration are inherited.
The native oracle requires exactly one succeeded native run attributed to this
issue, bound complete public events, one runner semantic proposal and authoritative
control-plane acceptance/terminal. A matched finishing tool start/result precedes
the actual provider final; no new invocation follows the proposal. The final
must be persisted as the run's agent reply. Acceptance may precede or follow the
final. This proves observable ordering and durable disposition; it does not prove
the provider consumed feedback or establish general coding-quality equivalence.
`single_attempt` is enforced by the launcher even when hosted invocation only
supplies `--id`. Every first failure, usage and cleanup receipt is retained; no
automatic infrastructure or behavioral retry occurs. The bounded comparison is
six cells per variant, twelve expected provider turns total, with the existing
per-cell deadlines (skill: original deadline; blocker: five minutes).
Candidate and baseline use identical executable fixtures and public oracles.
Only the five native description/fingerprint source files and their six
variant-specific unit tests may differ. The historical variant restores those
eleven files from master, including v13; candidate uses the archived native change
and v14. Source-contract hashes reject mixed revisions and changes to master
defaults, auth/runtime context or shared instructions. Derived capability metadata
is independently checked; it currently requires no regeneration in either variant.
Credential-free admission runs before local environment/credential loading:
```sh
node tests/runner-e2e/native-completion-checks.mjs --output-dir=/absolute/report/path
node tests/runner-e2e/native-completion-checks.mjs --verify=/absolute/report/path/preflight.json
pnpm test:e2e:runner -- --list --suite native-completion
```
Admission requires a committed clean exact source, variant-appropriate schema,
tool bridge/catalog and resume tests, independent oracle/default/retry calibrations,
original context-integrity calibrations, typecheck, manifest checks and exact
six-cell discovery. Generic stock-suite gates remain unchanged. Paid hosted runs
must dispatch the trusted default-branch workflow to distinct immutable target
branches using the exact six IDs, preserving authorization and secret boundaries.
Live results are pending until retained cell evidence is published.