Files
PaperClipAI/packages/paperclip-runner/docs/runner-workflow-evals.md
Dotta af8439a70b feat(runner): restore direct live eval campaigns and reports (#12909)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner executes agents through native and managed provider
drivers.
> - The direct live eval layer had drifted from the current Runner
contracts.
> - The old local workflow did not provide a complete parallel campaign
or durable report history.
> - The Runner also needed current native OpenCode and OpenRouter
qualification.
> - This pull request restores the direct campaign, corrects the runtime
gaps that the campaign found, and adds safe hosted Evalbook history.
> - The benefit is repeatable model comparison against an immutable
Runner and eval source revision.

## Linked Issues or Issue Description

Refs #11297
Refs #11634

**What existing behavior does this improve?**

This improves the direct live `paperclip-runner` eval workflow, provider
execution contract, and static Evalbook reporting path.

**Current behavior**

The direct evals do not have one maintained full campaign on current
`master`. OpenCode has no qualified multi-model OpenRouter roster.
Parallel provider bursts can compact committed events before the
transport observes them. Local reports do not have a separate safe S3
history index.

**Proposed behavior**

Run one immutable roster-plus-case matrix. Use the shared paid AWS
runner fleet. Keep raw artifacts access-controlled. Publish a sanitized
canonical Evalbook report under the separate `runner-protocol-evals` S3
prefix. Keep immutable campaign directories plus root history, latest,
and latest-green pointers.

**Reason and benefit**

Maintainers can compare native Codex, native OpenCode, ACPX, Claude
Managed, and AWS AgentCore behavior over time. They can inspect failures
without mixing this direct protocol layer with browser full-stack E2E.

**Breaking changes**

None. The new workflow and S3 prefix are additive. The existing Runner
full-stack E2E workflow and report remain separate.

## What Changed

- Added a trusted two-shard direct live workflow for up to 393
roster-plus-case cells.
- Reused the numeric actor allowlist, protected paid environment, and
RunsOn fleet controls from Runner full-stack E2E.
- Added immutable Runner and eval revision resolution, exact credential
boundaries, bounded retries, and cost ceilings.
- Added a public report projection that removes sessions, transcripts,
tool payloads, state, traces, raw failures, remote profile identities,
and credential-shaped values.
- Added additive S3 history under `runner-protocol-evals`, with
immutable campaigns and mutable root index pointers.
- Added native OpenCode model injection and current OpenRouter pricing
contracts.
- Fixed direct eval completion, workflow execution, semantic discovery,
warm-attach state reset, executable binding, and event-burst handling.
- Kept Runner browser full-stack E2E behavior and publication separate.
- Documented local and hosted direct eval operation.

## Verification

- `pnpm --filter @paperclipai/paperclip-runner
test:runner-protocol-eval-publish` — 15 passed.
- `pnpm --filter @paperclipai/paperclip-runner build:typescript` —
passed.
- `actionlint .github/workflows/runner-protocol-live-evals.yml
.github/workflows/runner-full-stack-e2e.yml` — passed.
- Local current matrix at the revision in
[paperclip-evals#17](https://github.com/paperclipai/paperclip-evals/pull/17)
— 323 cells across 10 enabled configurations completed.
- Final local current matrix — 269 passed, 11 behavior failures, and 43
expected macOS-only ACPX platform failures.
- Targeted Runner checks — 13/13 eval-session tests, 15/15
publisher/security tests, and package typecheck passed; complete PR CI
is green, including all browser E2E shards.

## Risks

- Paid live campaigns can consume provider budget. Actor authorization,
exact per-cell ceilings, protected environments, and explicit schedule
enablement bound this risk.
- Public reports can leak provider data. The workflow publishes only a
separately projected report and validates every file before upload.
- The new workflow cannot publish until it is present on the default
branch. This pull request does not change the existing
`runner-full-stack-e2e` publication path.
- The campaign is large. It uses two GitHub matrices and caps combined
concurrency at the shared fleet limit.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex on GPT-5. The exact deployment ID and context-window size
are not exposed. The model used reasoning, code editing, browser
inspection, repository tools, and live provider execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-05 20:18:11 -05:00

88 lines
4.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Stress-derived Runner workflow evals
The Runner workflow eval system turns the `STRESS-001`–`STRESS-044` campaign
into complementary deterministic, live, and chaos lanes. It is additive to the
capability inventory, capability cases, and existing scoring/report readers.
The workspace-private `@paperclipai/paperclip-eval-kernel` package owns only
structural scenario-by-candidate orchestration. Runner-specific cases,
observations, scoring, and traceability remain package-local. The HTML matrix is
rendered by the canonical `paperclip-evals` report program so live results use
the same grid and drill-down pages as the direct Runner eval suite.
## Lanes
- `pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals`
runs the credential-free PR gate over sanitized Codex, OpenCode, and ACPX
normalization fixtures.
- `pnpm --filter @paperclipai/paperclip-runner report:runner-workflow-evals`
validates the deterministic fail-closed fixture matrix and writes JSON,
Markdown, JUnit, and GitHub-safe artifacts under
`.paperclip-local/evals/workflows/`. It makes no network requests.
- `pnpm --filter @paperclipai/paperclip-runner report:runner-live-evals` runs
the balanced forty-execution schedule against real provider sessions. Live
candidate failures are trend-only; missing credentials, qualification
failures, and provider outages remain unscored.
- `pnpm --filter @paperclipai/paperclip-runner report:runner-chaos-evals`
writes the eight-scenario fault schedule consumed by weekly and pre-release
restart, replay, trace, finalization, interaction, and wake-race suites.
The checked-in live manifest contains only adapter/model settings,
qualification variable names, and budgets. Credentials remain in the
environment. `PAPERCLIP_EVAL_MAX_CAMPAIGN_COST_USD` must be a positive finite
number and defaults to 12 USD for scheduled runs.
Live executions export one immutable Evalbook attempt per workflow/candidate to
`.paperclip-local/evals/workflows/evalbook-runs/`, then invoke
`evals/paperclip-runner/tools/eval_program.py report` from a `paperclip-evals`
checkout. That program writes the canonical matrix to
`.paperclip-local/evals/workflows/index.html`, plus `latest.html`, test pages,
and attempt pages. Set `PAPERCLIP_EVALBOOK_PROGRAM` to the program's absolute
path or `PAPERCLIP_EVALS_ROOT` to its repository root. Conventional sibling
worktree locations are discovered automatically. GitHub Actions checks out a
pinned `paperclip-evals` revision, so every hosted run uses the same reviewed
report implementation rather than a copied or package-local renderer.
Filtered reports with only a few candidates expand their result columns to the
available viewport, keeping the PASS, FAIL, and INFRA labels visible without a
horizontal scroll. Full matrices retain the canonical scrollable grid and
sticky test-name column.
The exported artifact contains only the safe workflow observation and
scorecard. Prompts, credentials, raw provider frames, tool arguments, and
reasoning remain excluded. `evalbook-manifest.json` records the generator path
and SHA-256 digest used for the render.
Local and manual GitHub runs can bound paid execution with comma-separated
`--candidate` and `--case` selectors plus `--limit`. For example,
`report:runner-live-evals -- --candidate codex-luna --limit 2` executes only
the first two Codex entries in that week's validated schedule. Subsets receive
a distinct bundle identity and do not contaminate full-campaign trend history.
The hosted live workflow is default-branch-only and requires an allowlisted
numeric actor plus the protected `runner-e2e-paid` environment. Scheduled runs
also remain disabled until `RUNNER_LIVE_EVALS_NIGHTLY_ENABLED` is explicitly
set to `true`. The paid job uses the full-stack workflow's literal runner
selection: `RUNNER_E2E_AWS_ENABLED=true` routes it to the RunsOn Fleet, and any
other value uses `ubuntu-latest`.
## Trace and reasoning safety
Live executions capture provider frames in a run-local mode-`0600` sidecar.
The evaluator verifies byte lengths, SHA-256 digests, order, dispositions, and
lineage, retains only redacted observations plus a digest, and destroys the
temporary trace after execution. Prompts, credentials, tool arguments, and
reasoning text never enter reports or uploaded artifacts. Evals measure visible
progress and activity; they do not inspect or grade hidden chain of thought.
## Compatibility and trends
Live bundle identity includes the Runner version/build, prompt policy, schedule
seed, adapters, resolved models, and reasoning settings. Seven-day comparisons
use only matching bundle IDs, and alerts stay disabled until seven compatible
reports exist. Safe reports are retained for 30 days; raw traces are not
uploaded.
The checked traceability manifest is
`spec/evals/stress-workflow-traceability.json`; CI fails for missing findings,
unknown workflow IDs, or missing regression-test anchors.