Files
PaperClipAI/packages/paperclip-runner/docs/capability-eval-slice.md
Dotta 5458940a6e feat(runner): add offline evaluation tooling (#12653)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner needs repeatable evaluation contracts.
> - Evaluation code must stay separate from provider launch and
production orchestration.
> - Offline fixtures need stable compatibility, scoring, traceability,
and report rules.
> - Published Runner consumers need only the supported evaluation
contract surface.
> - This pull request adds offline evaluation tooling and a
workspace-private matrix kernel.
> - The benefit is deterministic evaluation without credentials or paid
provider calls.

## Linked Issues or Issue Description

Refs #11297

This pull request extracts the offline evaluation unit from the earlier
aggregate Runner work.

## What Changed

- Add a workspace-private, provider-neutral evaluation matrix kernel.
- Add the public `@paperclipai/paperclip-runner/evals` compatibility and
native execution contracts.
- Add fail-closed runnerd artifact and protocol compatibility checks.
- Add deterministic workflow catalogs, scoring, traceability, and report
generation.
- Add sanitized Codex, OpenCode, and ACPX fixtures.
- Add package-boundary and clean-consumer checks.
- Add the eval package manifest to the Docker dependency stage.
- Add the generated protocol fixture digest without changing the
lockfile.

## Verification

GitHub Actions must run:

- Runner TypeScript and Rust type checks.
- Runner unit and protocol tests.
- Evaluation kernel tests.
- Workflow traceability checks.
- Clean-consumer and package-boundary checks.
- Repository test, type-check, build, policy, and security gates.

No local test command was run. The repository owner requested
GitHub-only verification.

## Risks

This is a large greenfield review surface with 51 files. The code does
not launch a live provider or load credentials. Package and protocol
drift fail closed. The workspace lockfile remains under the existing
CI-owned process.

## Model Used

OpenAI Codex with the GPT-5 agent model. The work used high reasoning,
repository inspection, tool use, and parallel code review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 05:19:47 -05:00

1.5 KiB

Runner eval scoring slice

The package-local eval slice provides deterministic, credential-free scoring for runner behavior. It complements the fail-fast conformance suite by keeping each observation intact and reporting independent dimensions for safety, semantic outcome, trajectory restraint, trace completeness, and efficiency.

EvalBundle records reproducibility inputs without storing credentials. assertBundleSecretFree rejects credential-shaped keys and values before any structural error can echo them. Persisted reports contain an explicit digested bundle-evidence declaration rather than the free-form source bundle, while bundleId still derives a stable content identifier from the full canonical declaration. The exact final report serialization is scanned again before it is returned to a persistence boundary.

scoreEval is pure: the same observation and bundle always produce the same scorecard. Hard-invariant failures gate the overall score to zero. Other dimensions remain separate so a report shows whether a regression came from the semantic outcome, unnecessary calls, incomplete causal evidence, or an exceeded declared budget.

runEvalBehaviorFaultMatrix exercises deterministic green and red behavior against the package's mock authority. No provider process, network credential, or paid model invocation is required.

Run the offline slice with:

pnpm --filter @paperclipai/paperclip-runner test:eval-slice

Provider-backed campaigns and recorded evidence are intentionally outside this package boundary.