Files
PaperClipAI/packages/paperclip-runner/docs/adr/0001-runner-testing-eval-package-boundaries.md
T
Dotta 5458940a6e feat(runner): add offline evaluation tooling (#12653)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner needs repeatable evaluation contracts.
> - Evaluation code must stay separate from provider launch and
production orchestration.
> - Offline fixtures need stable compatibility, scoring, traceability,
and report rules.
> - Published Runner consumers need only the supported evaluation
contract surface.
> - This pull request adds offline evaluation tooling and a
workspace-private matrix kernel.
> - The benefit is deterministic evaluation without credentials or paid
provider calls.

## Linked Issues or Issue Description

Refs #11297

This pull request extracts the offline evaluation unit from the earlier
aggregate Runner work.

## What Changed

- Add a workspace-private, provider-neutral evaluation matrix kernel.
- Add the public `@paperclipai/paperclip-runner/evals` compatibility and
native execution contracts.
- Add fail-closed runnerd artifact and protocol compatibility checks.
- Add deterministic workflow catalogs, scoring, traceability, and report
generation.
- Add sanitized Codex, OpenCode, and ACPX fixtures.
- Add package-boundary and clean-consumer checks.
- Add the eval package manifest to the Docker dependency stage.
- Add the generated protocol fixture digest without changing the
lockfile.

## Verification

GitHub Actions must run:

- Runner TypeScript and Rust type checks.
- Runner unit and protocol tests.
- Evaluation kernel tests.
- Workflow traceability checks.
- Clean-consumer and package-boundary checks.
- Repository test, type-check, build, policy, and security gates.

No local test command was run. The repository owner requested
GitHub-only verification.

## Risks

This is a large greenfield review surface with 51 files. The code does
not launch a live provider or load credentials. Package and protocol
drift fail closed. The workspace lockfile remains under the existing
CI-owned process.

## Model Used

OpenAI Codex with the GPT-5 agent model. The work used high reasoning,
repository inspection, tool use, and parallel code review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 05:19:47 -05:00

6.0 KiB

ADR 0001: Runner, Testing, and Eval Package Boundaries

  • Status: Accepted
  • Date: 2026-08-11
  • Decision owners: Paperclip App

Context

The runner package previously exported production contracts, deterministic mocks, conformance fixtures, scenario/eval helpers, and demo servers from one root. A workspace checkout hid two packaging defects: consumers could use undeclared deep surfaces, and the public semantic dispatcher loaded ajv even though ajv was only a development dependency.

External conformance consumers must consume App contracts and test helpers without making the App runtime depend on scenario corpora, provider experiments, or a separate eval package.

Decision

Ownership is:

Owner Stable responsibility
Paperclip App PRP schemas and fixtures, canonical semantic catalog and dispatcher, runnerd/client interfaces, ControlPlanePort, the production binding, deterministic mock, and mock/real parity fixtures
Eval consumers Scenario corpus, provider configuration, experiment reports, and provider-backed orchestration

The production binding remains App code at server/src/services/native-runtime/paperclip-control-plane-port.ts. It implements the App-owned ControlPlanePort; the runner package never imports the server, database, UI, or CLI.

Public runner exports are:

Export Stability and purpose
@paperclipai/paperclip-runner Runtime contracts, runner clients/backends, PRP validation/replay, canonical catalog/dispatcher, and compatibility preflight
@paperclipai/paperclip-runner/evals Versioned native-attempt/build metadata, compatibility negotiation, and explicit runnerd artifact resolution
@paperclipai/paperclip-runner/testing Deterministic mocks, PRP port conformance, and provider-neutral semantic conformance kit
./browser, ./react, ./standalone, ./styles.css Existing explicitly named UI/standalone consumers

Mock adapters and conformance constants are no longer package-root exports. Tests and external conformance consumers must use ./testing. Scenario content, provider-backed matrices, reports, and provider configuration are not public App package exports.

The testkit remains an App ./testing subpath rather than a separate package. It shares the runner protocol/catalog release cadence, and no independent consumer or version cadence currently justifies another package. Split it only after an independent release requirement exists; a directory preference is not sufficient.

Generic credential-free matrix orchestration lives in the separately versioned, workspace-private @paperclipai/paperclip-eval-kernel package. It contains no runner imports, provider configuration, scenario corpus, scorer, or report renderer. The runner may use it only as a development dependency; runtime, optional, and peer dependency sets remain free of eval packages. Paid provider campaigns remain external.

The dependency graph is acyclic:

Paperclip App production binding
            |
            v
@paperclipai/paperclip-runner (runtime contracts)
            ^
            |
Eval consumers ----> @paperclipai/paperclip-runner/evals
       |             @paperclipai/paperclip-runner/testing
       +-----------> @paperclipai/paperclip-eval-kernel

No arrow points from App runtime to an external eval repository.

Compatibility and versioning

Package semver describes distribution compatibility. Independently versioned contracts are published in PAPERCLIP_RUNNER_COMPATIBILITY:

Component Current contract Compatibility rule
Canonical catalog 1 Operation removals, renames, placement changes, or incompatible schemas require a new contract version
PRP 1 Highest overlapping required protocol version; no overlap fails closed
Runner client 1 Breaking runnerd/client interface changes require a new version
runnerd artifact 2 Binary metadata/package disagreement or digest mismatch fails before launch
Harness driver 1 Breaking descriptor/config/session/conformance behavior requires a new version
Native execution 1 Breaking App attempt-bundle semantics require a new schema version and converter
Evals integration 1 Breaking package/binary/catalog/driver join behavior requires a new version
Control-plane adapter 1 Breaking ControlPlanePort or production-binding expectations require a new version
Testkit 1 Breaking mock seed, vector, observation, or conformance behavior requires a new version
Eval corpus 1 Runner declares a supported inclusive corpus-version range; out-of-range bundles fail before execution

assertPaperclipRunnerCompatibility performs a fail-closed preflight. It emits paperclip_runner_incompatible with stable issue codes for component mismatch, unsupported corpus versions, unknown catalog operations, missing provider capability declarations, and provider-operation gaps. Provider-specific runtime errors must not stand in for this preflight.

Clean-consumer proof

Run:

pnpm --filter @paperclipai/paperclip-runner check:package-boundaries
pnpm --filter @paperclipai/paperclip-runner check:clean-consumers

The second command builds and packs the runner, installs its tarball into a clean consumer, imports only the root, ./evals, and ./testing exports, executes deterministic PRP, harness-driver, and semantic conformance, and verifies the separately staged runnerd artifact digest. The consumer uses no workspace protocol, source-relative import, or deep package path. This is the packaging gate; workspace tests alone are not proof.

Consequences

  • Existing tests importing mock/conformance values from the package root must migrate to @paperclipai/paperclip-runner/testing.
  • ajv is a runtime dependency because the public dispatcher imports it.
  • The semantic conformance kit defines normalized comparison; real App service adapters and risk-weighted vectors may evolve behind its testkit version.
  • Scenario corpus changes do not force an App runtime release unless their declared compatibility requirement changes.