Files
PaperClipAI/packages/paperclip-eval-kernel/test/kernel.test.mjs
T
Dotta 5458940a6e feat(runner): add offline evaluation tooling (#12653)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner needs repeatable evaluation contracts.
> - Evaluation code must stay separate from provider launch and
production orchestration.
> - Offline fixtures need stable compatibility, scoring, traceability,
and report rules.
> - Published Runner consumers need only the supported evaluation
contract surface.
> - This pull request adds offline evaluation tooling and a
workspace-private matrix kernel.
> - The benefit is deterministic evaluation without credentials or paid
provider calls.

## Linked Issues or Issue Description

Refs #11297

This pull request extracts the offline evaluation unit from the earlier
aggregate Runner work.

## What Changed

- Add a workspace-private, provider-neutral evaluation matrix kernel.
- Add the public `@paperclipai/paperclip-runner/evals` compatibility and
native execution contracts.
- Add fail-closed runnerd artifact and protocol compatibility checks.
- Add deterministic workflow catalogs, scoring, traceability, and report
generation.
- Add sanitized Codex, OpenCode, and ACPX fixtures.
- Add package-boundary and clean-consumer checks.
- Add the eval package manifest to the Docker dependency stage.
- Add the generated protocol fixture digest without changing the
lockfile.

## Verification

GitHub Actions must run:

- Runner TypeScript and Rust type checks.
- Runner unit and protocol tests.
- Evaluation kernel tests.
- Workflow traceability checks.
- Clean-consumer and package-boundary checks.
- Repository test, type-check, build, policy, and security gates.

No local test command was run. The repository owner requested
GitHub-only verification.

## Risks

This is a large greenfield review surface with 51 files. The code does
not launch a live provider or load credentials. Package and protocol
drift fail closed. The workspace lockfile remains under the existing
CI-owned process.

## Model Used

OpenAI Codex with the GPT-5 agent model. The work used high reasoning,
repository inspection, tool use, and parallel code review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-01 05:19:47 -05:00

55 lines
1.7 KiB
JavaScript

import assert from "node:assert/strict";
import test from "node:test";
import {
PAPERCLIP_EVAL_KERNEL_COMPATIBILITY,
PaperclipEvalKernelConfigurationError,
runPaperclipEvalMatrix,
} from "../dist/index.js";
test("runs a caller-owned scenario/candidate matrix", async () => {
const results = await runPaperclipEvalMatrix({
scenarios: [{ id: "scenario-a", input: { value: 2 } }],
candidates: [{ id: "candidate-a", config: { multiplier: 3 } }],
execute: async ({ scenario, candidate }) => scenario.input.value * candidate.config.multiplier,
score: ({ output }) => ({ passed: output === 6 }),
});
assert.equal(PAPERCLIP_EVAL_KERNEL_COMPATIBILITY.apiVersion, 1);
assert.deepEqual(results, [{
scenarioId: "scenario-a",
candidateId: "candidate-a",
output: 6,
score: { passed: true },
}]);
});
test("fails before execution when compatibility preflight fails", async () => {
let executed = false;
await assert.rejects(
runPaperclipEvalMatrix({
scenarios: [{ id: "scenario-a", input: null }],
candidates: [{
id: "candidate-a",
config: null,
preflight: () => { throw new Error("paperclip_runner_incompatible"); },
}],
execute: async () => { executed = true; },
score: () => null,
}),
/paperclip_runner_incompatible/,
);
assert.equal(executed, false);
});
test("rejects duplicate scenario ids", async () => {
await assert.rejects(
runPaperclipEvalMatrix({
scenarios: [{ id: "duplicate", input: 1 }, { id: "duplicate", input: 2 }],
candidates: [{ id: "candidate-a", config: null }],
execute: async () => null,
score: () => null,
}),
PaperclipEvalKernelConfigurationError,
);
});