Files
PaperClipAI/doc/evals.md
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00

12 KiB

Paperclip evaluation guide

Paperclip has two live eval families with different questions, owners, and evidence. Choose the family before selecting a model, profile, or case.

  • Runner Evals: real Runner/provider behavior against a seeded mock control plane. Definitions live in paperclip-evals/evals/paperclip-runner; see the direct live protocol evals.
  • Product E2E Evals: real browser, Paperclip server, database, Runner, provider, and (where selected) Daytona, using an isolated instance and grading oracle. See tests/runner-e2e and Everyday Workflows.

Runner Evals answer whether a real runner/provider can perform a bounded protocol operation against the expected control-plane contract. Product E2E Evals answer whether a person can complete a product workflow through the real Paperclip surfaces and whether the resulting artifact and state are usable. The names describe the system under test; “headless” is an execution option, not an eval category.

Selecting a family

Use Runner Evals for a runner protocol, adapter, transport, native session, tool grant, or one-turn provider qualification question. The workflow checks out an exact paperclip-evals revision, builds the Runner and viewer, runs a live roster, and renders the canonical Evalbook report. The control plane is a seeded test authority, so a passing result does not prove browser UX, production server behavior, database persistence, Daytona behavior, or a real third-party mutation.

Use Product E2E Evals for browser interaction, issue/task lifecycle, approval and clarification UI, project/repository selection, persistence over a controller restart, artifact delivery, billing/evidence behavior, or runner continuity in local or Daytona environments. The harness creates a fresh Paperclip instance per cell and uses public APIs and the production browser surface. The suite's Everyday Workflows are Product E2E even when their results are imported into Evalbook.

Do not combine a partial Runner campaign and a partial Product E2E campaign into one score. A campaign is comparable when its definition/grader, model/profile, environment, and contract match. The evaluated Paperclip revision may intentionally differ for a before/after fix comparison; record it as a comparison axis.

Ownership and codepaths

Runner Evals are owned by the Runner/evals maintainers. Definitions, rosters, case prompts, and the report program live in the sibling private repository paperclipai/paperclip-evals; Runner integration, viewer, aggregation, and publication code live under packages/paperclip-runner and the runner-protocol-live-evals.yml workflow. The public-facing report uses the same Evalbook renderer and Runner Lab viewer as the trusted report after sanitization.

Product E2E Evals are owned by the runner E2E maintainers. The catalog and harness are under tests/runner-e2e; the package scripts are test:e2e:runner, test:e2e:runner:unit, test:e2e:runner:typecheck, and test:e2e:runner:report. README.md, FIXTURES.md, SECURITY.md, and EVERYDAY-WORKFLOWS.md are the detailed sources of truth. The harness starts the server and embedded database, creates the company/agent/task through the real APIs, drives Chromium, and invokes the selected local or Daytona runner.

The explicit-only agent-chat-hardening Product E2E suite covers native chat recovery, hiring, status evidence, and review handoff on local and selected warm Daytona paths. Its fixture contract distinguishes startup cancellation from active response cancellation and HTTP send replay from ambiguous provider action recovery. Select it explicitly; --all excludes it.

The explicit-only agent-chat-stories suite covers the experimental settings lifecycle for a configured native agent and follow-ups during active work. Its fixture-driven file wait and persisted-plan oracle are documented in the Product E2E guide. It does not qualify the native onboarding wizard or change the native API-tool rollout defaults.

Validation ladder

Start with credential-free checks and a catalog listing. For Product E2E:

pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
pnpm test:e2e:runner -- --list

For one explicitly selected local cell, configure only the credentials named by that cell in .env.runner-e2e.local, then run a narrow ID:

pnpm test:e2e:runner -- --id core-compatibility.runner-codex.local.message-marker

Use the selectors documented in the runner E2E README for a suite, profile, case, group, or environment. Daytona needs the immutable image digest and DAYTONA_API_KEY; follow the README and fixture security guide. --all excludes manual suites such as everyday-workflows. Select that suite explicitly; use a narrow selector while developing a fixture.

For Runner Evals, the narrowest useful local validation is the report program's help/validation path and the deterministic Runner checks documented in runner-workflow-evals.md. Hosted direct live runs must use the default-branch workflow, an exact 40 character evals_sha, an explicitly selected roster (or the maintained enabled all campaign), and the protected paid environment. The complete hosted command is intentionally kept in the workflow and direct live protocol guide. Live provider runs can spend money; use the existing workflow authorization and the user's stated scope when selecting them.

Failure taxonomy

Record the primary failure class and preserve the evidence that supports it.

  • Product failure: evidence shows Paperclip or Runner behavior violates the authored case or a hard invariant, such as wrong task state, missing approval gate, lost persistence, bad artifact, or incorrect protocol operation.
  • Model/provider behavior failure: the provider turn completed with usable evidence but the model gave the wrong answer, ignored an interaction, failed to complete the authored operation, or violated a semantic assertion. It is scored as behavior, not silently retried as infrastructure.
  • Grading/evidence failure: the case or matcher cannot establish its claim, a required recording/screenshot/result is malformed, or the report contract is invalid. Fix the harness or grader before interpreting the score.
  • Infrastructure failure: the evidence points to provider/profile unavailability, transport admission failure, service startup failure, a missing credential/image, or inability to produce usable evidence. Startup, transport, and timeout symptoms can instead be product defects when evidence implicates Paperclip or Runner; classify from the observed failure and supported cause, rather than the symptom name alone. Preserve the artifact.

Missing usage or price data means unknown, not free. Keep provider-reported costs separate from estimates, and include retry costs when available. Latency, cleanup, billing coverage, and unpriced usage are dimensions of the result and should remain visible alongside the primary class. A timeout after successful product state reads can be a product behavior failure; a failed server-health read may be infrastructure, but inspect its cause. Use the family-specific classifier and read the attempt evidence before changing an analytical label.

Evidence, provenance, and history

An Evalbook report is a presentation of immutable attempt records, not the source of truth. Keep the campaign ID, Paperclip commit, paperclip-evals commit, catalog/roster or definition fingerprint, model/profile, environment, grader version, selected cells, retries, and provider/runtime usage with the report. Public projections follow each family's reviewed allowlist and may include sanitized fixture conversation, named tool outcomes, screenshots, and structured evidence intended for public history. Credentials, secrets, private data, raw unredacted records, and hidden reasoning stay out of public projections.

Distinguish a complete campaign from a partial campaign. A narrow selector, manual diagnostic, missing cell, or infrastructure retry can be useful evidence without being a qualification run. History should retain both, with explicit coverage and completeness, while trend and latest-green views compare only compatible complete campaigns. Refreshing an existing report from retained evidence has zero provider calls and is a new presentation of the old measurement, not a new model run.

Existing public histories are available at Runner protocol history and Runner Product E2E history. The consolidated eval hub is at pages.paperclip.ing/evals.

For a repeatable workflow, use the matching skill: paperclip-evals, add-runner-eval, or add-product-e2e-eval.

Install the authoring skills

The reviewable sources live in this repository's .agents/skills. For a multi-repository workspace, install the three skills at ~/paperclipai/.agents/skills (not ~/paperclipai/skills). From the Paperclip checkout, run:

for skill in paperclip-evals add-runner-eval add-product-e2e-eval; do
  install -d "$HOME/paperclipai/.agents/skills/$skill"
  install -m 644 ".agents/skills/$skill/SKILL.md" \
    "$HOME/paperclipai/.agents/skills/$skill/SKILL.md"
done

This replaces only the three named skill entrypoints. Run it again after updating their tracked sources. Each skill locates the repository independently of its installation directory.

Maintain the public hub

The hub is a static directory with two links to the existing history systems. It displays a dated snapshot, not a live scoreboard. It does not run models, create another result archive, or change the existing campaign URLs.

Build from the public history feeds and check its summary logic:

python3 -m unittest discover -s scripts/evals-hub -p 'test_*.py'
python3 scripts/evals-hub/build.py --output .paperclip/evals-hub

The hub checks need Python 3 and do not call model providers.

For offline checks, pass --history-dir <directory> containing runner-protocol-evals-history.json and runner-e2e-history.json. For a pre-merge preview, pass --docs-ref <branch-or-sha> to link the guide at that revision. The default guide link uses master.

Publish with the Paperclip page helper and the configured page-uploader credentials. Use Bash 4 or newer; macOS's system Bash 3 cannot run this helper. On macOS with Homebrew Bash installed, put $(brew --prefix bash)/bin first in PATH before these commands:

export PAPERCLIP_PAGE_BUCKET=pages.paperclip.ing
export PAPERCLIP_PAGE_BASE_URL=https://pages.paperclip.ing
export AWS_REGION=us-east-1
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals --dry-run
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals

For later refreshes, rebuild in the same output directory and publish with --update. Keep its ignored .paperclip-page/state.json ownership record; without that record, the helper will refuse to overwrite an existing prefix. Verify the public page and its links after publication. This manual refresh does not add a scheduled workflow. Preserve the measurement date when choosing a newer rendering of the same campaign.

Remaining native chat boundaries are in the explicit-only agent-chat-qualification suite: active task reassignment, user Retry after verified worker process loss, and multi-turn answers grounded in actual task records. See the workflow and qualification limits. The 26 native first-task cells exercise onboarding before native selection becomes the UI default. Live results and semantic answer reviews must accompany any qualification claim; catalog presence alone is not a pass.