Files
PaperClipAI/doc/evals.md
T
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00

15 KiB

Paperclip evaluation guide

Paperclip has two live eval families with different questions, owners, and evidence. Choose the family before selecting a model, profile, or case.

  • Runner Evals: real Runner/provider behavior against a seeded mock control plane. Definitions live in paperclip-evals/evals/paperclip-runner; see the direct live protocol evals.
  • Product E2E Evals: real browser, Paperclip server, database, Runner, provider, and (where selected) Daytona, using an isolated instance and grading oracle. See tests/runner-e2e and Everyday Workflows.

Runner Evals answer whether a real runner/provider can perform a bounded protocol operation against the expected control-plane contract. Product E2E Evals answer whether a person can complete a product workflow through the real Paperclip surfaces and whether the resulting artifact and state are usable. The names describe the system under test; “headless” is an execution option, not an eval category.

The explicit Product E2E completion-updates suite compares onboarding and idle Agent Chat handoffs on native Claude/Codex. It separates mechanical completion delivery/result access from semantic review of the retained answer; see the probe contract.

Selecting a family

Use Runner Evals for a runner protocol, adapter, transport, native session, tool grant, or one-turn provider qualification question. The workflow checks out an exact paperclip-evals revision, builds the Runner and viewer, runs a live roster, and renders the canonical Evalbook report. The control plane is a seeded test authority, so a passing result does not prove browser UX, production server behavior, database persistence, Daytona behavior, or a real third-party mutation.

Use Product E2E Evals for browser interaction, issue/task lifecycle, approval and clarification UI, project/repository selection, persistence over a controller restart, artifact delivery, billing/evidence behavior, or runner continuity in local or Daytona environments. The harness creates a fresh Paperclip instance per cell and uses public APIs and the production browser surface. The suite's Everyday Workflows are Product E2E even when their results are imported into Evalbook.

Do not combine a partial Runner campaign and a partial Product E2E campaign into one score. A campaign is comparable when its definition/grader, model/profile, environment, and contract match. The evaluated Paperclip revision may intentionally differ for a before/after fix comparison; record it as a comparison axis.

Ownership and codepaths

Runner Evals are owned by the Runner/evals maintainers. Definitions, rosters, case prompts, and the report program live in the sibling private repository paperclipai/paperclip-evals; Runner integration, viewer, aggregation, and publication code live under packages/paperclip-runner and the runner-protocol-live-evals.yml workflow. The public-facing report uses the same Evalbook renderer and Runner Lab viewer as the trusted report after sanitization.

Product E2E Evals are owned by the runner E2E maintainers. The catalog and harness are under tests/runner-e2e; the package scripts are test:e2e:runner, test:e2e:runner:unit, test:e2e:runner:typecheck, and test:e2e:runner:report. README.md, FIXTURES.md, SECURITY.md, and EVERYDAY-WORKFLOWS.md are the detailed sources of truth. The harness starts the server and embedded database, creates the company/agent/task through the real APIs, drives Chromium, and invokes the selected local or Daytona runner.

The explicit-only agent-chat-hardening Product E2E suite covers native chat recovery, hiring, status evidence, and review handoff on local and selected warm Daytona paths. Its fixture contract distinguishes startup cancellation from active response cancellation and HTTP send replay from ambiguous provider action recovery. Select it explicitly; --all excludes it.

The explicit-only context-integrity Product E2E suite covers ordered public comment continuation and explicit invocation of an assigned pinned skill across the seven selected legacy/native local profiles. Select it by suite or exact execution ID because --all excludes explicit-only suites. Each cell applies a 1,000-cent company and agent budget hard stop before task creation and records both limits in its evidence.

The explicit-only agent-chat-stories suite covers the experimental settings lifecycle for a configured native agent and follow-ups during active work. Its fixture-driven file wait and persisted-plan oracle are documented in the Product E2E guide. It does not qualify the native onboarding wizard or change the native API-tool rollout defaults.

The explicit-only grok-qualification and grok-subscription-qualification Product suites exercise Grok Build with API and company subscription authentication respectively. Keep their results separate; the subscription fixture seeds an explicitly supplied login and does not qualify interactive login. See the Grok fixture contract.

Validation ladder

Start with credential-free checks and a catalog listing. For Product E2E:

pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
pnpm test:e2e:runner -- --list

For one explicitly selected local cell, configure only the credentials named by that cell in .env.runner-e2e.local, then run a narrow ID:

pnpm test:e2e:runner -- --id core-compatibility.runner-codex.local.message-marker

Use the selectors documented in the runner E2E README for a suite, profile, case, group, or environment. Daytona needs the immutable image digest and DAYTONA_API_KEY; follow the README and fixture security guide. --all excludes manual suites such as everyday-workflows. Select that suite explicitly; use a narrow selector while developing a fixture.

For Runner Evals, the narrowest useful local validation is the report program's help/validation path and the deterministic Runner checks documented in runner-workflow-evals.md. Hosted direct live runs must use the default-branch workflow, an exact 40 character evals_sha, an explicitly selected roster (or the maintained enabled all campaign), and the protected paid environment. The complete hosted command is intentionally kept in the workflow and direct live protocol guide. Live provider runs can spend money; use the existing workflow authorization and the user's stated scope when selecting them.

Failure taxonomy

Record the primary failure class and preserve the evidence that supports it.

  • Product failure: evidence shows Paperclip or Runner behavior violates the authored case or a hard invariant, such as wrong task state, missing approval gate, lost persistence, bad artifact, or incorrect protocol operation.
  • Model/provider behavior failure: the provider turn completed with usable evidence but the model gave the wrong answer, ignored an interaction, failed to complete the authored operation, or violated a semantic assertion. It is scored as behavior, not silently retried as infrastructure.
  • Grading/evidence failure: the case or matcher cannot establish its claim, a required recording/screenshot/result is malformed, or the report contract is invalid. Fix the harness or grader before interpreting the score.
  • Infrastructure failure: the evidence points to provider/profile unavailability, transport admission failure, service startup failure, a missing credential/image, or inability to produce usable evidence. Startup, transport, and timeout symptoms can instead be product defects when evidence implicates Paperclip or Runner; classify from the observed failure and supported cause, rather than the symptom name alone. Preserve the artifact.

Missing usage or price data means unknown, not free. Keep provider-reported costs separate from estimates, and include retry costs when available. Latency, cleanup, billing coverage, and unpriced usage are dimensions of the result and should remain visible alongside the primary class. A timeout after successful product state reads can be a product behavior failure; a failed server-health read may be infrastructure, but inspect its cause. Use the family-specific classifier and read the attempt evidence before changing an analytical label.

Evidence, provenance, and history

Retained result snapshots and dated measurement reports belong in paperclip-evals; application tests, Product E2E fixtures/graders, and executable scenario inventories remain in this repository. Keep a compact results index with immutable archive links and public report links, as in the lifecycle baseline. The private archive is not a dependency of app test execution. Keep large logs, traces, and videos in the existing campaign artifact storage.

An Evalbook report is a presentation of immutable attempt records, not the source of truth. Keep the campaign ID, Paperclip commit, paperclip-evals commit, catalog/roster or definition fingerprint, model/profile, environment, grader version, selected cells, retries, and provider/runtime usage with the report. Public projections follow each family's reviewed allowlist and may include sanitized fixture conversation, named tool outcomes, screenshots, and structured evidence intended for public history. Credentials, secrets, private data, raw unredacted records, and hidden reasoning stay out of public projections.

Distinguish a complete campaign from a partial campaign. A narrow selector, manual diagnostic, missing cell, or infrastructure retry can be useful evidence without being a qualification run. History should retain both, with explicit coverage and completeness, while trend and latest-green views compare only compatible complete campaigns. Refreshing an existing report from retained evidence has zero provider calls and is a new presentation of the old measurement, not a new model run.

Existing public histories are available at Runner protocol history and Runner Product E2E history. The consolidated eval hub is at pages.paperclip.ing/evals.

For a repeatable workflow, use the matching skill: paperclip-evals, add-runner-eval, or add-product-e2e-eval.

Install the authoring skills

The reviewable sources live in this repository's .agents/skills. For a multi-repository workspace, install the three skills at ~/paperclipai/.agents/skills (not ~/paperclipai/skills). From the Paperclip checkout, run:

for skill in paperclip-evals add-runner-eval add-product-e2e-eval; do
  install -d "$HOME/paperclipai/.agents/skills/$skill"
  install -m 644 ".agents/skills/$skill/SKILL.md" \
    "$HOME/paperclipai/.agents/skills/$skill/SKILL.md"
done

This replaces only the three named skill entrypoints. Run it again after updating their tracked sources. Each skill locates the repository independently of its installation directory.

Maintain the public hub

The hub is a static directory with two links to the existing history systems. It displays a dated snapshot, not a live scoreboard. It does not run models, create another result archive, or change the existing campaign URLs.

Build from the public history feeds and check its summary logic:

python3 -m unittest discover -s scripts/evals-hub -p 'test_*.py'
python3 scripts/evals-hub/build.py --output .paperclip/evals-hub

The hub checks need Python 3 and do not call model providers.

For offline checks, pass --history-dir <directory> containing runner-protocol-evals-history.json and runner-e2e-history.json. For a pre-merge preview, pass --docs-ref <branch-or-sha> to link the guide at that revision. The default guide link uses master.

Publish with the Paperclip page helper and the configured page-uploader credentials. Use Bash 4 or newer; macOS's system Bash 3 cannot run this helper. On macOS with Homebrew Bash installed, put $(brew --prefix bash)/bin first in PATH before these commands:

export PAPERCLIP_PAGE_BUCKET=pages.paperclip.ing
export PAPERCLIP_PAGE_BASE_URL=https://pages.paperclip.ing
export AWS_REGION=us-east-1
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals --dry-run
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals

For later refreshes, rebuild in the same output directory and publish with --update. Keep its ignored .paperclip-page/state.json ownership record; without that record, the helper will refuse to overwrite an existing prefix. Verify the public page and its links after publication. This manual refresh does not add a scheduled workflow. Preserve the measurement date when choosing a newer rendering of the same campaign.

Remaining native chat boundaries are in the explicit-only agent-chat-qualification suite: active task reassignment, user Retry after verified worker process loss, and multi-turn answers grounded in actual task records. See the workflow and qualification limits. The 26 native first-task cells exercise onboarding before native selection becomes the UI default. Live results and semantic answer reviews must accompany any qualification claim; catalog presence alone is not a pass.

Lifecycle behavior baseline

The credential-free lifecycle baseline joins unit, scripted-runner, and database integration assertions to a scenario inventory before changing narrative-based lifecycle policy. Run pnpm test:lifecycle-baseline to retain current passes and failures. Its Product E2E matcher calibration is separate from live execution; unrun live coverage remains explicitly unmeasured.

The separate live lifecycle baseline defines 46 real-provider Product E2E cells, including paired narrative probes and named existing controls on legacy and native Codex. Discover it with pnpm test:e2e:runner -- --list --suite lifecycle-baseline. Historical execution results and follow-up coverage are recorded in that suite's guide.

Continuation accounting has an explicit-only eight-cell Product E2E baseline suite, complementing the deterministic lifecycle inventory.

The explicit-only Product E2E api-response-reading suite verifies retrieval of large saved API responses on local and Daytona native Codex runs. See the Runner E2E guide.