mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 20:50:08 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
260 lines
14 KiB
Markdown
260 lines
14 KiB
Markdown
# Paperclip evaluation guide
|
|
|
|
Paperclip has two live eval families with different questions, owners, and
|
|
evidence. Choose the family before selecting a model, profile, or case.
|
|
|
|
- **Runner Evals:** real Runner/provider behavior against a seeded mock control
|
|
plane. Definitions live in `paperclip-evals/evals/paperclip-runner`; see the
|
|
[direct live protocol evals](../packages/paperclip-runner/docs/runner-protocol-live-evals.md).
|
|
- **Product E2E Evals:** real browser, Paperclip server, database, Runner,
|
|
provider, and (where selected) Daytona, using an isolated instance and
|
|
grading oracle. See [`tests/runner-e2e`](../tests/runner-e2e/README.md) and
|
|
[Everyday Workflows](../tests/runner-e2e/EVERYDAY-WORKFLOWS.md).
|
|
|
|
Runner Evals answer whether a real runner/provider can perform a bounded
|
|
protocol operation against the expected control-plane contract. Product E2E
|
|
Evals answer whether a person can complete a product workflow through the real
|
|
Paperclip surfaces and whether the resulting artifact and state are usable.
|
|
The names describe the system under test; “headless” is an execution option,
|
|
not an eval category.
|
|
|
|
## Selecting a family
|
|
|
|
Use **Runner Evals** for a runner protocol, adapter, transport, native session,
|
|
tool grant, or one-turn provider qualification question. The workflow checks
|
|
out an exact `paperclip-evals` revision, builds the Runner and viewer, runs a
|
|
live roster, and renders the canonical Evalbook report. The control plane is a
|
|
seeded test authority, so a passing result does not prove browser UX, production
|
|
server behavior, database persistence, Daytona behavior, or a real third-party
|
|
mutation.
|
|
|
|
Use **Product E2E Evals** for browser interaction, issue/task lifecycle,
|
|
approval and clarification UI, project/repository selection, persistence over a
|
|
controller restart, artifact delivery, billing/evidence behavior, or runner
|
|
continuity in local or Daytona environments. The harness creates a fresh
|
|
Paperclip instance per cell and uses public APIs and the production browser
|
|
surface. The suite's [Everyday Workflows](../tests/runner-e2e/EVERYDAY-WORKFLOWS.md)
|
|
are Product E2E even when their results are imported into Evalbook.
|
|
|
|
Do not combine a partial Runner campaign and a partial Product E2E campaign into
|
|
one score. A campaign is comparable when its definition/grader, model/profile,
|
|
environment, and contract match. The evaluated Paperclip revision may
|
|
intentionally differ for a before/after fix comparison; record it as a
|
|
comparison axis.
|
|
|
|
## Ownership and codepaths
|
|
|
|
Runner Evals are owned by the Runner/evals maintainers. Definitions, rosters,
|
|
case prompts, and the report program live in the sibling private repository
|
|
`paperclipai/paperclip-evals`; Runner integration, viewer, aggregation, and
|
|
publication code live under `packages/paperclip-runner` and the
|
|
`runner-protocol-live-evals.yml` workflow. The public-facing report uses the
|
|
same Evalbook renderer and Runner Lab viewer as the trusted report after
|
|
sanitization.
|
|
|
|
Product E2E Evals are owned by the runner E2E maintainers. The catalog and
|
|
harness are under `tests/runner-e2e`; the package scripts are `test:e2e:runner`,
|
|
`test:e2e:runner:unit`, `test:e2e:runner:typecheck`, and
|
|
`test:e2e:runner:report`. `README.md`, `FIXTURES.md`, `SECURITY.md`, and
|
|
`EVERYDAY-WORKFLOWS.md` are the detailed sources of truth. The harness starts
|
|
the server and embedded database, creates the company/agent/task through the
|
|
real APIs, drives Chromium, and invokes the selected local or Daytona runner.
|
|
|
|
The explicit-only `agent-chat-hardening` Product E2E suite covers native chat
|
|
recovery, hiring, status evidence, and review handoff on local and selected warm
|
|
Daytona paths. Its [fixture contract](../tests/runner-e2e/README.md) distinguishes
|
|
startup cancellation from active response cancellation and HTTP send replay
|
|
from ambiguous provider action recovery. Select it explicitly; `--all` excludes it.
|
|
|
|
The explicit-only `agent-chat-stories` suite covers the experimental settings
|
|
lifecycle for a configured native agent and follow-ups during active work. Its
|
|
fixture-driven file wait and persisted-plan oracle are documented in the
|
|
[Product E2E guide](../tests/runner-e2e/README.md). It does not qualify the native
|
|
onboarding wizard or change the native API-tool rollout defaults.
|
|
|
|
## Validation ladder
|
|
|
|
Start with credential-free checks and a catalog listing. For Product E2E:
|
|
|
|
```sh
|
|
pnpm test:e2e:runner:typecheck
|
|
pnpm test:e2e:runner:unit
|
|
pnpm test:e2e:runner -- --list
|
|
```
|
|
|
|
For one explicitly selected local cell, configure only the credentials named
|
|
by that cell in `.env.runner-e2e.local`, then run a narrow ID:
|
|
|
|
```sh
|
|
pnpm test:e2e:runner -- --id core-compatibility.runner-codex.local.message-marker
|
|
```
|
|
|
|
Use the selectors documented in the [runner E2E README](../tests/runner-e2e/README.md)
|
|
for a suite, profile, case, group, or environment. Daytona needs the immutable
|
|
image digest and `DAYTONA_API_KEY`; follow the README and fixture security guide.
|
|
`--all` excludes manual suites such as `everyday-workflows`. Select that suite
|
|
explicitly; use a narrow selector while developing a fixture.
|
|
|
|
For Runner Evals, the narrowest useful local validation is the report program's
|
|
help/validation path and the deterministic Runner checks documented in
|
|
[`runner-workflow-evals.md`](../packages/paperclip-runner/docs/runner-workflow-evals.md).
|
|
Hosted direct live runs must use the default-branch workflow, an exact 40
|
|
character `evals_sha`, an explicitly selected roster (or the maintained
|
|
enabled `all` campaign), and the protected paid environment. The complete
|
|
hosted command is intentionally kept in the workflow and
|
|
[direct live protocol guide](../packages/paperclip-runner/docs/runner-protocol-live-evals.md).
|
|
Live provider runs can spend money; use the existing workflow authorization and
|
|
the user's stated scope when selecting them.
|
|
|
|
## Failure taxonomy
|
|
|
|
Record the primary failure class and preserve the evidence that supports it.
|
|
|
|
- **Product failure:** evidence shows Paperclip or Runner behavior violates the
|
|
authored case or a hard invariant, such as wrong task state, missing approval
|
|
gate, lost persistence, bad artifact, or incorrect protocol operation.
|
|
- **Model/provider behavior failure:** the provider turn completed with usable
|
|
evidence but the model gave the wrong answer, ignored an interaction, failed
|
|
to complete the authored operation, or violated a semantic assertion. It is
|
|
scored as behavior, not silently retried as infrastructure.
|
|
- **Grading/evidence failure:** the case or matcher cannot establish its claim,
|
|
a required recording/screenshot/result is malformed, or the report contract
|
|
is invalid. Fix the harness or grader before interpreting the score.
|
|
- **Infrastructure failure:** the evidence points to provider/profile
|
|
unavailability, transport admission failure, service startup failure, a
|
|
missing credential/image, or inability to produce usable evidence. Startup,
|
|
transport, and timeout symptoms can instead be product defects when evidence
|
|
implicates Paperclip or Runner; classify from the observed failure and
|
|
supported cause, rather than the symptom name alone. Preserve the artifact.
|
|
|
|
Missing usage or price data means unknown, not free. Keep provider-reported
|
|
costs separate from estimates, and include retry costs when available.
|
|
Latency, cleanup, billing coverage, and unpriced usage are dimensions of the
|
|
result and should remain visible alongside the primary class. A timeout after
|
|
successful product state reads can be a product behavior failure; a failed
|
|
server-health read may be infrastructure, but inspect its cause. Use the
|
|
family-specific classifier and read the attempt evidence before changing an
|
|
analytical label.
|
|
|
|
## Evidence, provenance, and history
|
|
|
|
Retained result snapshots and dated measurement reports belong in
|
|
`paperclip-evals`; application tests, Product E2E fixtures/graders, and executable
|
|
scenario inventories remain in this repository. Keep a compact results index
|
|
with immutable archive links and public report links, as in the
|
|
[lifecycle baseline](../tests/lifecycle-baseline/README.md#recorded-results-moved-to-paperclip-evals).
|
|
The private archive is not a dependency of app test execution. Keep large logs,
|
|
traces, and videos in the existing campaign artifact storage.
|
|
|
|
An Evalbook report is a presentation of immutable attempt records, not the
|
|
source of truth. Keep the campaign ID, Paperclip commit, `paperclip-evals`
|
|
commit, catalog/roster or definition fingerprint, model/profile, environment,
|
|
grader version, selected cells, retries, and provider/runtime usage with the
|
|
report. Public projections follow each family's reviewed allowlist and may
|
|
include sanitized fixture conversation, named tool outcomes, screenshots, and
|
|
structured evidence intended for public history. Credentials, secrets, private
|
|
data, raw unredacted records, and hidden reasoning stay out of public
|
|
projections.
|
|
|
|
Distinguish a complete campaign from a partial campaign. A narrow selector,
|
|
manual diagnostic, missing cell, or infrastructure retry can be useful evidence
|
|
without being a qualification run. History should retain both, with explicit
|
|
coverage and completeness, while trend and latest-green views compare only
|
|
compatible complete campaigns. Refreshing an existing report from retained
|
|
evidence has zero provider calls and is a new presentation of the old
|
|
measurement, not a new model run.
|
|
|
|
Existing public histories are available at
|
|
[Runner protocol history](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/index.html)
|
|
and [Runner Product E2E history](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/).
|
|
The consolidated eval hub is at
|
|
[pages.paperclip.ing/evals](https://pages.paperclip.ing/evals/).
|
|
|
|
For a repeatable workflow, use the matching skill: [paperclip-evals](../.agents/skills/paperclip-evals/SKILL.md),
|
|
[add-runner-eval](../.agents/skills/add-runner-eval/SKILL.md), or
|
|
[add-product-e2e-eval](../.agents/skills/add-product-e2e-eval/SKILL.md).
|
|
|
|
## Install the authoring skills
|
|
|
|
The reviewable sources live in this repository's `.agents/skills`. For a
|
|
multi-repository workspace, install the three skills at
|
|
`~/paperclipai/.agents/skills` (not `~/paperclipai/skills`). From the Paperclip
|
|
checkout, run:
|
|
|
|
```sh
|
|
for skill in paperclip-evals add-runner-eval add-product-e2e-eval; do
|
|
install -d "$HOME/paperclipai/.agents/skills/$skill"
|
|
install -m 644 ".agents/skills/$skill/SKILL.md" \
|
|
"$HOME/paperclipai/.agents/skills/$skill/SKILL.md"
|
|
done
|
|
```
|
|
|
|
This replaces only the three named skill entrypoints. Run it again after
|
|
updating their tracked sources. Each skill locates the repository independently
|
|
of its installation directory.
|
|
|
|
## Maintain the public hub
|
|
|
|
The hub is a static directory with two links to the existing history systems.
|
|
It displays a dated snapshot, not a live scoreboard. It does not run models,
|
|
create another result archive, or change the existing campaign URLs.
|
|
|
|
Build from the public history feeds and check its summary logic:
|
|
|
|
```sh
|
|
python3 -m unittest discover -s scripts/evals-hub -p 'test_*.py'
|
|
python3 scripts/evals-hub/build.py --output .paperclip/evals-hub
|
|
```
|
|
|
|
The hub checks need Python 3 and do not call model providers.
|
|
|
|
For offline checks, pass `--history-dir <directory>` containing
|
|
`runner-protocol-evals-history.json` and `runner-e2e-history.json`.
|
|
For a pre-merge preview, pass `--docs-ref <branch-or-sha>` to link the guide
|
|
at that revision. The default guide link uses `master`.
|
|
|
|
Publish with the [Paperclip page helper](../.agents/skills/paperclip-page/SKILL.md)
|
|
and the configured page-uploader credentials. Use Bash 4 or newer; macOS's
|
|
system Bash 3 cannot run this helper. On macOS with Homebrew Bash installed,
|
|
put `$(brew --prefix bash)/bin` first in `PATH` before these commands:
|
|
|
|
```sh
|
|
export PAPERCLIP_PAGE_BUCKET=pages.paperclip.ing
|
|
export PAPERCLIP_PAGE_BASE_URL=https://pages.paperclip.ing
|
|
export AWS_REGION=us-east-1
|
|
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals --dry-run
|
|
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals
|
|
```
|
|
|
|
For later refreshes, rebuild in the same output directory and publish with
|
|
`--update`. Keep its ignored `.paperclip-page/state.json` ownership record;
|
|
without that record, the helper will refuse to overwrite an existing prefix.
|
|
Verify the public page and its links after publication. This manual refresh
|
|
does not add a scheduled workflow. Preserve the measurement date when choosing
|
|
a newer rendering of the same campaign.
|
|
|
|
Remaining native chat boundaries are in the explicit-only
|
|
`agent-chat-qualification` suite: active task reassignment, user Retry after
|
|
verified worker process loss, and multi-turn answers grounded in actual task
|
|
records. See the [workflow and qualification limits](../tests/runner-e2e/README.md#remaining-native-agent-chat-qualification).
|
|
The 26 native `first-task` cells exercise onboarding before native selection
|
|
becomes the UI default. Live results and semantic answer reviews must accompany
|
|
any qualification claim; catalog presence alone is not a pass.
|
|
|
|
## Lifecycle behavior baseline
|
|
|
|
The credential-free [lifecycle baseline](../tests/lifecycle-baseline/README.md)
|
|
joins unit, scripted-runner, and database integration assertions to a scenario
|
|
inventory before changing narrative-based lifecycle policy. Run
|
|
`pnpm test:lifecycle-baseline` to retain current passes and failures. Its Product
|
|
E2E matcher calibration is separate from live execution; unrun live coverage
|
|
remains explicitly unmeasured.
|
|
|
|
The separate [live lifecycle baseline](../tests/runner-e2e/LIFECYCLE-BASELINE.md)
|
|
defines 46 real-provider Product E2E cells, including paired narrative probes and
|
|
named existing controls on legacy and native Codex. Discover it with
|
|
`pnpm test:e2e:runner -- --list --suite lifecycle-baseline`. Historical execution
|
|
results and follow-up coverage are recorded in that suite's guide.
|
|
|
|
Continuation accounting has an explicit-only eight-cell Product E2E [baseline suite](../tests/runner-e2e/CONTINUATION-ACCOUNTING.md), complementing the deterministic lifecycle inventory.
|