Files
PaperClipAI/doc/evals.md
T
DottaandPaperclip 9d19f98b50 fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agent chat uses native runner sessions to plan, delegate, and track
that work.
> - A user can press Stop while the native session is still starting.
> - The server can acknowledge that Stop without dispatching it, then
let the session submit a turn.
> - This leaves chat recovery waiting for an execution that the user
expected to stop.
> - This PR waits for the startup handle, dispatches cancellation, and
prevents a late startup from submitting a turn.
> - New full-stack evals check the resulting records and outputs across
Claude and Codex.
> - Those evals also exposed missing ACPX readiness fields, unbounded
polling, and an old-run identity check that rejected valid warm
handoffs.

## Linked Issues or Issue Description

**What happened?**

Stop during native startup could record an acknowledged cancellation
with `dispatched: false`. The provider could then begin work. A
subsequent `/new` stayed queued. A remote Claude follow-up also
exhausted the command journal while probing warm-session readiness: ACPX
never returned the readiness fields required by the shared transport.
Once readiness worked, attachment incorrectly compared the next run
descriptor against the old run ID. The 25 ms polling loop could issue
4,800 commands during its two-minute wait, beyond the 500-command bound.
The existing chat eval treated lifecycle logs as proof of an active
provider turn, so it did not distinguish startup cancellation from
active-turn cancellation.

**Expected behavior**

A Stop during startup must reach the pending session. A late session
must not submit a prompt after Stop. Recovery must retain control when
startup exceeds the bounded wait. Chat evals must check saved task
state, document contents, worker identity, account binding, and
duplicate effects.

**Steps to reproduce**

1. Start a native Claude or Codex chat turn.
2. Press Stop after process startup is requested but before the provider
turn starts.
3. Send `/new`, then send a fresh message.
4. On the affected base, cancellation can be acknowledged without
dispatch and the reset stays queued.

**Paperclip version or commit**

The live Claude baseline reproduced this on `29d6b3509`. The branch also
includes master commit `0f5fafe16`.

Related work: #13678, #13686, #13693, #13291, #13738. A separate runner
reliability branch also contains a startup-wait fix. Its overlap must be
reconciled before merging; this branch additionally prevents prompt
submission after a late startup.

## What Changed

- Wait for a pending native startup before acknowledging a run-scoped
Stop. Preserve the existing recovery error when that wait expires.
- Keep a Stop guard on startup. Cancel a late handle before it can
submit a provider turn.
- Add regression tests for normal handle publication and publication
after the Stop deadline.
- Back off blocked warm-attachment probes. Keep the fast two-snapshot
barrier, fail closed, and record changed blockers.
- Add red/green tests for delayed readiness, persistent blockers,
alternating readiness, and readiness near the deadline.
- Publish ACPX readiness and blockers. Preserve the old authority’s
event acknowledgement barrier; only settled sessions can proceed to
attachment.
- Bind warm ACPX descriptors to the validated next authority while
retaining old-run event correlation until activation. Preserve session
identity and provider profile checks.
- Exercise two consecutive run rotations through a qualified fake
sidecar, verifying checkpointing, provider identity, pre-activation
rejection, and new-run work admission.
- Separate startup and active-turn cancellation checkpoints in the
browser eval.
- Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells
across Claude and Codex.
- Cover hiring and reuse through managed AI accounts, source-based
review, current blocked-task status, request replay after a lost HTTP
acknowledgement, server restart continuity, and Stop/reset continuity.
- Use ordinary production agent instructions. Enable API tools only for
the two coordination cases that need them.
- Calibrate the matchers with invalid records and outputs. Require
remembered context after restart and a structured status snapshot that
distinguishes the current blocker from history and task status from
active execution. Compare the public issue mutation contract and
relationships during read-only reporting. Preserve before/after source
records in failed eval evidence.
- Fix the lost-ack browser harness and verify it against a real HTTP
server. Check the chat composer after restart instead of waiting for an
unrelated document lifecycle event.
- Document the scope and limits of each case.

## Verification

- The startup regression failed on the unfixed executor and passed after
the fix.
- `pnpm test:e2e:runner:typecheck` passed.
- `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files.
- `pnpm exec vitest run
server/src/services/native-runtime/native-session-executor.test.ts`
passed: 385 tests.
- [Baseline live
campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868):
Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed.
Codex delegation was blocked by provider capacity.
- [Eval-only startup
campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786):
both providers failed as expected. Both persisted `dispatched: false`
and left `/new` queued.
- [First fixed
campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706)
on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both
providers. Failed cases exposed eval harness defects and remote
continuity failures. All attempts remain available.
- [Original workflows and stronger memory
checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649)
on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and
startup Stop passed for both providers; Codex remote restart passed.
Claude remote restart exposed the missing readiness contract. Two Codex
planning cells hit provider capacity.
- [Unchanged-model
retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548):
Codex planning and backlog creation both passed.
- [18-cell campaign with ACPX
readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963)
on `6a98ef743`: 16/18 passed, including all local/remote Stop and
committed-send cases. Claude remote continuity exposed the
next-authority check, now fixed. Codex hiring produced its checklist,
but the runner redacted the requested marker after it appeared as
“Tracking token: …”. That content-redaction policy is unchanged and
remains an explicit limitation.
- [Structured status
grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725)
on `50448c228`: both providers passed on their first attempt, including
cleanup.
- [Complete read-only state
grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011)
on `551e13892`: both providers passed.
- [Final ACPX handoff and hiring
retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456)
on `cbd637587`: all three Claude Daytona cases passed (restart
continuity, active Stop/reset, and lost-ack replay). Codex hiring
reproduced the content-redaction failure: the saved checklist contained
`Tracking token: [REDACTED]` instead of the required business marker.
All four cases completed cleanup successfully. [Published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/).
The only subsequent commit adds the qualified-sidecar integration test;
production code is identical to this live proof.
- `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without
paid models.
- Runner TypeScript typecheck passed. All 5 warm-readiness tests pass;
two failed with the prior fixed-rate loop, and the late-readiness test
failed before the pacing correction.
- ACPX readiness and warm-identity regressions each failed before their
fixes. All 292 runner-core Rust library tests passed. The
qualified-sidecar integration test passes. Rust formatting is checked.
- Status-grader regressions for misleading historical mentions and
previously unchecked mutations each failed before tightening the oracle
and pass now.
- [Latest-head
CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307)
passed on `a4093c8f1`: full build, type checks, test partitions, browser
E2E, and native runner checks. Two unrelated tests initially failed
(Sentry fixture release attribution and local-service fixture
readiness); both passed locally together (35 passed, 5 optional SDK
tests skipped) and on the failed-job retry. No changes were made to
those tests.
- Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed
and all review threads are resolved.
- The paid live suite is not fully green: the reproducible
content-redaction case remains red. This is separate from the passing PR
merge checks. No production content-redaction, prompt, model, or
completion-policy change is included.
- Managed-account hiring and review cases explicitly enable API tools;
these do not qualify default new-user onboarding.

## Risks

- Stop can wait up to 30 seconds for startup, then use the existing
pending-recovery path. This does not prove that remote cleanup has
finished.
- Blocked warm readiness adds up to 750 ms between later probes with the
two-minute remote budget, or about 32 ms with the default five-second
budget. Ready sessions retain the short second barrier.
- Paid evals can fail because of provider capacity or agent decisions.
Each failure needs evidence-based classification.
- The HTTP request replay case checks comment idempotency and duplicate
effects. It does not prove replay safety for an ambiguous provider tool
call.
- The new suite is opt-in. It does not increase the default paid
campaign.
- No production prompts or model selection change. Review-handoff
behavior and content-redaction policy remain separate product decisions.
The latter can remove harmless business content that looks like
credential syntax; the failing attempt is retained.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 10:35:38 -05:00

11 KiB

Paperclip evaluation guide

Paperclip has two live eval families with different questions, owners, and evidence. Choose the family before selecting a model, profile, or case.

  • Runner Evals: real Runner/provider behavior against a seeded mock control plane. Definitions live in paperclip-evals/evals/paperclip-runner; see the direct live protocol evals.
  • Product E2E Evals: real browser, Paperclip server, database, Runner, provider, and (where selected) Daytona, using an isolated instance and grading oracle. See tests/runner-e2e and Everyday Workflows.

Runner Evals answer whether a real runner/provider can perform a bounded protocol operation against the expected control-plane contract. Product E2E Evals answer whether a person can complete a product workflow through the real Paperclip surfaces and whether the resulting artifact and state are usable. The names describe the system under test; “headless” is an execution option, not an eval category.

Selecting a family

Use Runner Evals for a runner protocol, adapter, transport, native session, tool grant, or one-turn provider qualification question. The workflow checks out an exact paperclip-evals revision, builds the Runner and viewer, runs a live roster, and renders the canonical Evalbook report. The control plane is a seeded test authority, so a passing result does not prove browser UX, production server behavior, database persistence, Daytona behavior, or a real third-party mutation.

Use Product E2E Evals for browser interaction, issue/task lifecycle, approval and clarification UI, project/repository selection, persistence over a controller restart, artifact delivery, billing/evidence behavior, or runner continuity in local or Daytona environments. The harness creates a fresh Paperclip instance per cell and uses public APIs and the production browser surface. The suite's Everyday Workflows are Product E2E even when their results are imported into Evalbook.

Do not combine a partial Runner campaign and a partial Product E2E campaign into one score. A campaign is comparable when its definition/grader, model/profile, environment, and contract match. The evaluated Paperclip revision may intentionally differ for a before/after fix comparison; record it as a comparison axis.

Ownership and codepaths

Runner Evals are owned by the Runner/evals maintainers. Definitions, rosters, case prompts, and the report program live in the sibling private repository paperclipai/paperclip-evals; Runner integration, viewer, aggregation, and publication code live under packages/paperclip-runner and the runner-protocol-live-evals.yml workflow. The public-facing report uses the same Evalbook renderer and Runner Lab viewer as the trusted report after sanitization.

Product E2E Evals are owned by the runner E2E maintainers. The catalog and harness are under tests/runner-e2e; the package scripts are test:e2e:runner, test:e2e:runner:unit, test:e2e:runner:typecheck, and test:e2e:runner:report. README.md, FIXTURES.md, SECURITY.md, and EVERYDAY-WORKFLOWS.md are the detailed sources of truth. The harness starts the server and embedded database, creates the company/agent/task through the real APIs, drives Chromium, and invokes the selected local or Daytona runner.

The explicit-only agent-chat-hardening Product E2E suite covers native chat recovery, hiring, status evidence, and review handoff on local and selected warm Daytona paths. Its fixture contract distinguishes startup cancellation from active response cancellation and HTTP send replay from ambiguous provider action recovery. Select it explicitly; --all excludes it.

Validation ladder

Start with credential-free checks and a catalog listing. For Product E2E:

pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
pnpm test:e2e:runner -- --list

For one explicitly selected local cell, configure only the credentials named by that cell in .env.runner-e2e.local, then run a narrow ID:

pnpm test:e2e:runner -- --id core-compatibility.runner-codex.local.message-marker

Use the selectors documented in the runner E2E README for a suite, profile, case, group, or environment. Daytona needs the immutable image digest and DAYTONA_API_KEY; follow the README and fixture security guide. --all excludes manual suites such as everyday-workflows. Select that suite explicitly; use a narrow selector while developing a fixture.

For Runner Evals, the narrowest useful local validation is the report program's help/validation path and the deterministic Runner checks documented in runner-workflow-evals.md. Hosted direct live runs must use the default-branch workflow, an exact 40 character evals_sha, an explicitly selected roster (or the maintained enabled all campaign), and the protected paid environment. The complete hosted command is intentionally kept in the workflow and direct live protocol guide. Live provider runs can spend money; use the existing workflow authorization and the user's stated scope when selecting them.

Failure taxonomy

Record the primary failure class and preserve the evidence that supports it.

  • Product failure: evidence shows Paperclip or Runner behavior violates the authored case or a hard invariant, such as wrong task state, missing approval gate, lost persistence, bad artifact, or incorrect protocol operation.
  • Model/provider behavior failure: the provider turn completed with usable evidence but the model gave the wrong answer, ignored an interaction, failed to complete the authored operation, or violated a semantic assertion. It is scored as behavior, not silently retried as infrastructure.
  • Grading/evidence failure: the case or matcher cannot establish its claim, a required recording/screenshot/result is malformed, or the report contract is invalid. Fix the harness or grader before interpreting the score.
  • Infrastructure failure: the evidence points to provider/profile unavailability, transport admission failure, service startup failure, a missing credential/image, or inability to produce usable evidence. Startup, transport, and timeout symptoms can instead be product defects when evidence implicates Paperclip or Runner; classify from the observed failure and supported cause, rather than the symptom name alone. Preserve the artifact.

Missing usage or price data means unknown, not free. Keep provider-reported costs separate from estimates, and include retry costs when available. Latency, cleanup, billing coverage, and unpriced usage are dimensions of the result and should remain visible alongside the primary class. A timeout after successful product state reads can be a product behavior failure; a failed server-health read may be infrastructure, but inspect its cause. Use the family-specific classifier and read the attempt evidence before changing an analytical label.

Evidence, provenance, and history

An Evalbook report is a presentation of immutable attempt records, not the source of truth. Keep the campaign ID, Paperclip commit, paperclip-evals commit, catalog/roster or definition fingerprint, model/profile, environment, grader version, selected cells, retries, and provider/runtime usage with the report. Public projections follow each family's reviewed allowlist and may include sanitized fixture conversation, named tool outcomes, screenshots, and structured evidence intended for public history. Credentials, secrets, private data, raw unredacted records, and hidden reasoning stay out of public projections.

Distinguish a complete campaign from a partial campaign. A narrow selector, manual diagnostic, missing cell, or infrastructure retry can be useful evidence without being a qualification run. History should retain both, with explicit coverage and completeness, while trend and latest-green views compare only compatible complete campaigns. Refreshing an existing report from retained evidence has zero provider calls and is a new presentation of the old measurement, not a new model run.

Existing public histories are available at Runner protocol history and Runner Product E2E history. The consolidated eval hub is at pages.paperclip.ing/evals.

For a repeatable workflow, use the matching skill: paperclip-evals, add-runner-eval, or add-product-e2e-eval.

Install the authoring skills

The reviewable sources live in this repository's .agents/skills. For a multi-repository workspace, install the three skills at ~/paperclipai/.agents/skills (not ~/paperclipai/skills). From the Paperclip checkout, run:

for skill in paperclip-evals add-runner-eval add-product-e2e-eval; do
  install -d "$HOME/paperclipai/.agents/skills/$skill"
  install -m 644 ".agents/skills/$skill/SKILL.md" \
    "$HOME/paperclipai/.agents/skills/$skill/SKILL.md"
done

This replaces only the three named skill entrypoints. Run it again after updating their tracked sources. Each skill locates the repository independently of its installation directory.

Maintain the public hub

The hub is a static directory with two links to the existing history systems. It displays a dated snapshot, not a live scoreboard. It does not run models, create another result archive, or change the existing campaign URLs.

Build from the public history feeds and check its summary logic:

python3 -m unittest discover -s scripts/evals-hub -p 'test_*.py'
python3 scripts/evals-hub/build.py --output .paperclip/evals-hub

The hub checks need Python 3 and do not call model providers.

For offline checks, pass --history-dir <directory> containing runner-protocol-evals-history.json and runner-e2e-history.json. For a pre-merge preview, pass --docs-ref <branch-or-sha> to link the guide at that revision. The default guide link uses master.

Publish with the Paperclip page helper and the configured page-uploader credentials. Use Bash 4 or newer; macOS's system Bash 3 cannot run this helper. On macOS with Homebrew Bash installed, put $(brew --prefix bash)/bin first in PATH before these commands:

export PAPERCLIP_PAGE_BUCKET=pages.paperclip.ing
export PAPERCLIP_PAGE_BASE_URL=https://pages.paperclip.ing
export AWS_REGION=us-east-1
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals --dry-run
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals

For later refreshes, rebuild in the same output directory and publish with --update. Keep its ignored .paperclip-page/state.json ownership record; without that record, the helper will refuse to overwrite an existing prefix. Verify the public page and its links after publication. This manual refresh does not add a scheduled workflow. Preserve the measurement date when choosing a newer rendering of the same campaign.