<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task descriptions, comments, continuation data, skills, and execution rules enter several agent adapters. > - The same source can be rendered by more than one automatic input carrier. > - Failed resumes can also rebuild input from stale or compact context. > - This pull request gives each Paperclip-owned source one delivery owner and preserves the required transport boundaries. > - It adds deterministic adapter, interaction, runner, and browser tests for these boundaries. > - The benefit is more predictable context delivery with explicit evidence for later live qualification. ## Linked Issues or Issue Description Related: #13144 removes a duplicate environment payload and bounds wake lists. Related: #11360 addresses Hermes resume behavior. This pull request preserves compatible active-session formats while repairing context ownership and stale question creation. **What happened?** Task descriptions and comments could enter more than one automatic context block. Native transports could wrap a complete model input in a second task envelope. Some legacy and gateway adapters could omit the owned assignment on ordinary tasks or rebuild a failed resume with stale compact context. A continuation could also request a question after newer human comments had arrived. **Expected behavior** Each task or comment source has one automatic model-facing owner. Distinct comment IDs and repeated wording remain distinct. Fresh fallback attempts rebuild the required full context. A question request is rejected when newer queued human direction makes it stale. Harness access policy remains owned by execution configuration. **Steps to reproduce** 1. Build a task with a description and current comments. 2. Capture the actual adapter or runner input. 3. Compare source ownership and task-envelope nesting. 4. Queue a human comment before a continuation requests a question. 5. Trigger a failed resume and inspect the fresh retry input. 6. Run the focused adapter, interaction, runner, and browser checks. ## What Changed - Add shared prompt-section selection at the provider-attempt boundary. - Deliver owned assignment context through native, legacy CLI, ACP, gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and Hermes paths. - Rebuild full or compact context after resume recovery changes the attempt. Add native and Claude ACP tests of actual recovery requests. - Preserve custom templates, loaded instruction files, execution policies, and older active-session formats. - Record continuation source metadata and reject stale question creation under the issue-row lock. - Add explicit Product E2E context-integrity profiles, prerequisite gates, credential-isolation checks, and report fixtures. - Bypass service-worker forwarding for same-origin Vite development modules. A real Chromium test fails with resource exhaustion before the repair and passes after it. Production asset caching keeps its existing policy. - Add browser diagnostics and service-worker module-loading regressions. - Add an explicit zero-retry eval option. The default retry behavior remains unchanged. Each campaign records its effective policy. - Remove the model-facing working-directory sentence from four prompt builders. Existing workspace, sandbox, permission, and custom-template configuration remains unchanged. - Align the everyday workflow assertion with the current 47-entry catalog. Compared with current upstream master, the branch carries the context-ownership implementation and its tests, the explicit context-integrity catalog and evidence harness, and the focused browser regression checks. ## Verification **Merge assessment:** focused regression evidence supports merge. This is not full completion of the original broad qualification matrix. The maintainer has authorized merge after fresh verification of the master integration. - Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`. All 14 conflicts are resolved. Cancellation checks, workspace finalization, native Grok support, and both sets of tests are retained. - Current-head Greptile: **5/5**, with no blocking findings. The review names this exact commit. All **59 reported checks are terminal: 55 successful, 4 skipped, zero pending or failing**. This includes the full root general and serialized suites, separate runner checks, typecheck, build, canary, browser E2E, Docker, and security checks. The successful legacy security status is included in that total. - After integration: workspace typecheck and full build passed. Separate runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust tests, and 39 preparation checks**. Other passing checks include 621 Product E2E harness units, 376 focused shared/adapter tests, 160 real-database/API tests, 86 Hermes tests, 18 browser-support checks, and Product E2E typechecking. The complete root suite passed in CI. The duplicate local monolithic root run was stopped after that CI result; it is not counted as a completed local pass. - New native recovery coverage retains full assignment, completion contract, and explicit skill selection after safe replacement, for old and prepared input formats. Full native session test file: **136/136 passed**. - New Claude ACP coverage captures actual fresh, resumed, and missing-session fallback requests. It verifies one assignment copy, comment order, identical text under distinct comment IDs, and full fallback context. Full file: **33/33 passed**. Both affected TypeScript checks passed. - Existing deterministic tests cover source revisions, approval and trust boundaries, completion validation, custom templates, compatible sessions, standalone driver wrapping, and maintained adapter transport requests. - Provider-free browser support: **17/17 passed** after the master merge. Service-worker unit tests: **33/33 passed**. The module-overload regression failed before the repair and passed after it in real Chromium. ### Fresh live comparisons The new batch ran exactly four Product E2E attempts. **All four passed on the first attempt; no retries.** Each has six terminal matchers plus the existing browser lifecycle and invariant checks. | Exact case ID | Control | Candidate | |---|---|---| | `core-compatibility.runner-codex.local.plan-revise-accept` | Passed | Passed | | `local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume` | Passed | Passed | The plan case checks a revised canonical plan and revision-bound approval before completion. The question case restarts the server before submitting the answer, then verifies the continuation completes. Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical frozen definitions and provider versions: Codex `0.156.0` with `gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with `claude-sonnet-5`. The September 24 head added master browser recovery and test-only changes. The September 28 head also integrates newer master changes, including cancellation, workspace finalization, and native Grok. These are frozen-source live results, not exact-head live runs. The candidate received one description copy where the control initially received three. The submitted initial plan envelopes were 7,969 versus 19,097 characters. Question envelopes were 7,592 versus 18,919. These are structural measurements, not whole-provider token or dollar savings. ### Earlier evidence and failed attempts - The preceding fresh batch has four effective passing pairs: OpenCode comment continuation and assigned skill, native Codex comment continuation, and native Claude comment continuation. It retains **11 attempts: eight passed and three failed**. - Original failures remain recorded: missing local PostgreSQL library links before task creation; host-sleep cleanup after task/page checks passed; and a Claude **control** session-open rejection before a model turn. Setup was repaired identically on both worktrees. The permitted unchanged infrastructure retries passed. The underlying Claude provider startup error was not retained and remains unknown. - Older R2 retains **17 passes and one failure** across 18 attempts, including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode blank-page failure led to the service-worker repair. R2 is historical evidence: master changed the native fixed prompt and removed duplicate wake environment data afterward. - The September 24 CI run initially failed one unrelated preview readiness test (`ECONNREFUSED` on its local fixture). Its test and production code match master. Isolated local verification passed **28 tests, 3 skipped**. One unchanged CI retry passed the full shard: **831 passed, 1 skipped**, including all **31 preview-exposure tests**. The aggregate CI gate passed afterward. The precise startup cause remains unknown; a port race is a hypothesis, not a proved cause. ### Limits The original wider profile/workflow matrix, repeated trials, and remote Daytona qualification are incomplete. These results support a focused merge recommendation, not statistical equivalence or universal harness qualification. Some usage receipts are missing in both variants, so no token or dollar savings are claimed. The $500 ceiling was preserved using conservative allowances; failed attempts and unknown charges remain in the ledger. Reproduce the focused additions with `pnpm exec vitest run packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/native-session-runtime.test.ts`. Full checks use `pnpm -r typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner checks. Paid evals require the frozen definitions, profiles, and credentials; do not use `--all` as a substitute for the selected cases. ## Risks - Context placement changes can affect model behavior. Deterministic checks cover the selected paths, but live qualification remains incomplete. - The stale-question guard can reject a request when queued human comments arrived during the run. This is intended. - New stored inputs and model envelopes retain compatibility readers for older active sessions. - Custom templates may intentionally repeat content. - Removing a model-facing working-directory sentence does not change filesystem, command, sandbox, or permission configuration. - The worker bypass applies only to same-origin development module paths. Cache-policy tests preserve private-response handling and production asset caching. Mounted HTTP fixture changes remain test-only. - This PR does not claim measured token savings or statistical equivalence across every harness. ## Model Used OpenAI Codex, exact model gpt-6-astra, with repository tools and code execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The serving context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR using the required issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused local checks and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect these changes - [x] I have considered and documented risks above - [x] All current-head Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups for the current head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
15 KiB
Paperclip evaluation guide
Paperclip has two live eval families with different questions, owners, and evidence. Choose the family before selecting a model, profile, or case.
- Runner Evals: real Runner/provider behavior against a seeded mock control
plane. Definitions live in
paperclip-evals/evals/paperclip-runner; see the direct live protocol evals. - Product E2E Evals: real browser, Paperclip server, database, Runner,
provider, and (where selected) Daytona, using an isolated instance and
grading oracle. See
tests/runner-e2eand Everyday Workflows.
Runner Evals answer whether a real runner/provider can perform a bounded protocol operation against the expected control-plane contract. Product E2E Evals answer whether a person can complete a product workflow through the real Paperclip surfaces and whether the resulting artifact and state are usable. The names describe the system under test; “headless” is an execution option, not an eval category.
The explicit Product E2E completion-updates suite compares onboarding and
idle Agent Chat handoffs on native Claude/Codex. It separates mechanical
completion delivery/result access from semantic review of the retained answer;
see the probe contract.
Selecting a family
Use Runner Evals for a runner protocol, adapter, transport, native session,
tool grant, or one-turn provider qualification question. The workflow checks
out an exact paperclip-evals revision, builds the Runner and viewer, runs a
live roster, and renders the canonical Evalbook report. The control plane is a
seeded test authority, so a passing result does not prove browser UX, production
server behavior, database persistence, Daytona behavior, or a real third-party
mutation.
Use Product E2E Evals for browser interaction, issue/task lifecycle, approval and clarification UI, project/repository selection, persistence over a controller restart, artifact delivery, billing/evidence behavior, or runner continuity in local or Daytona environments. The harness creates a fresh Paperclip instance per cell and uses public APIs and the production browser surface. The suite's Everyday Workflows are Product E2E even when their results are imported into Evalbook.
Do not combine a partial Runner campaign and a partial Product E2E campaign into one score. A campaign is comparable when its definition/grader, model/profile, environment, and contract match. The evaluated Paperclip revision may intentionally differ for a before/after fix comparison; record it as a comparison axis.
Ownership and codepaths
Runner Evals are owned by the Runner/evals maintainers. Definitions, rosters,
case prompts, and the report program live in the sibling private repository
paperclipai/paperclip-evals; Runner integration, viewer, aggregation, and
publication code live under packages/paperclip-runner and the
runner-protocol-live-evals.yml workflow. The public-facing report uses the
same Evalbook renderer and Runner Lab viewer as the trusted report after
sanitization.
Product E2E Evals are owned by the runner E2E maintainers. The catalog and
harness are under tests/runner-e2e; the package scripts are test:e2e:runner,
test:e2e:runner:unit, test:e2e:runner:typecheck, and
test:e2e:runner:report. README.md, FIXTURES.md, SECURITY.md, and
EVERYDAY-WORKFLOWS.md are the detailed sources of truth. The harness starts
the server and embedded database, creates the company/agent/task through the
real APIs, drives Chromium, and invokes the selected local or Daytona runner.
The explicit-only agent-chat-hardening Product E2E suite covers native chat
recovery, hiring, status evidence, and review handoff on local and selected warm
Daytona paths. Its fixture contract distinguishes
startup cancellation from active response cancellation and HTTP send replay
from ambiguous provider action recovery. Select it explicitly; --all excludes it.
The explicit-only context-integrity Product E2E suite covers ordered public
comment continuation and explicit invocation of an assigned pinned skill across
the seven selected legacy/native local profiles. Select it by suite or exact
execution ID because --all excludes explicit-only suites. Each cell applies a
1,000-cent company and agent budget hard stop before task creation and records
both limits in its evidence.
The explicit-only agent-chat-stories suite covers the experimental settings
lifecycle for a configured native agent and follow-ups during active work. Its
fixture-driven file wait and persisted-plan oracle are documented in the
Product E2E guide. It does not qualify the native
onboarding wizard or change the native API-tool rollout defaults.
The explicit-only grok-qualification and grok-subscription-qualification
Product suites exercise Grok Build with API and company subscription
authentication respectively. Keep their results separate; the subscription
fixture seeds an explicitly supplied login and does not qualify interactive
login. See the Grok fixture contract.
Validation ladder
Start with credential-free checks and a catalog listing. For Product E2E:
pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
pnpm test:e2e:runner -- --list
For one explicitly selected local cell, configure only the credentials named
by that cell in .env.runner-e2e.local, then run a narrow ID:
pnpm test:e2e:runner -- --id core-compatibility.runner-codex.local.message-marker
Use the selectors documented in the runner E2E README
for a suite, profile, case, group, or environment. Daytona needs the immutable
image digest and DAYTONA_API_KEY; follow the README and fixture security guide.
--all excludes manual suites such as everyday-workflows. Select that suite
explicitly; use a narrow selector while developing a fixture.
For Runner Evals, the narrowest useful local validation is the report program's
help/validation path and the deterministic Runner checks documented in
runner-workflow-evals.md.
Hosted direct live runs must use the default-branch workflow, an exact 40
character evals_sha, an explicitly selected roster (or the maintained
enabled all campaign), and the protected paid environment. The complete
hosted command is intentionally kept in the workflow and
direct live protocol guide.
Live provider runs can spend money; use the existing workflow authorization and
the user's stated scope when selecting them.
Failure taxonomy
Record the primary failure class and preserve the evidence that supports it.
- Product failure: evidence shows Paperclip or Runner behavior violates the authored case or a hard invariant, such as wrong task state, missing approval gate, lost persistence, bad artifact, or incorrect protocol operation.
- Model/provider behavior failure: the provider turn completed with usable evidence but the model gave the wrong answer, ignored an interaction, failed to complete the authored operation, or violated a semantic assertion. It is scored as behavior, not silently retried as infrastructure.
- Grading/evidence failure: the case or matcher cannot establish its claim, a required recording/screenshot/result is malformed, or the report contract is invalid. Fix the harness or grader before interpreting the score.
- Infrastructure failure: the evidence points to provider/profile unavailability, transport admission failure, service startup failure, a missing credential/image, or inability to produce usable evidence. Startup, transport, and timeout symptoms can instead be product defects when evidence implicates Paperclip or Runner; classify from the observed failure and supported cause, rather than the symptom name alone. Preserve the artifact.
Missing usage or price data means unknown, not free. Keep provider-reported costs separate from estimates, and include retry costs when available. Latency, cleanup, billing coverage, and unpriced usage are dimensions of the result and should remain visible alongside the primary class. A timeout after successful product state reads can be a product behavior failure; a failed server-health read may be infrastructure, but inspect its cause. Use the family-specific classifier and read the attempt evidence before changing an analytical label.
Evidence, provenance, and history
Retained result snapshots and dated measurement reports belong in
paperclip-evals; application tests, Product E2E fixtures/graders, and executable
scenario inventories remain in this repository. Keep a compact results index
with immutable archive links and public report links, as in the
lifecycle baseline.
The private archive is not a dependency of app test execution. Keep large logs,
traces, and videos in the existing campaign artifact storage.
An Evalbook report is a presentation of immutable attempt records, not the
source of truth. Keep the campaign ID, Paperclip commit, paperclip-evals
commit, catalog/roster or definition fingerprint, model/profile, environment,
grader version, selected cells, retries, and provider/runtime usage with the
report. Public projections follow each family's reviewed allowlist and may
include sanitized fixture conversation, named tool outcomes, screenshots, and
structured evidence intended for public history. Credentials, secrets, private
data, raw unredacted records, and hidden reasoning stay out of public
projections.
Distinguish a complete campaign from a partial campaign. A narrow selector, manual diagnostic, missing cell, or infrastructure retry can be useful evidence without being a qualification run. History should retain both, with explicit coverage and completeness, while trend and latest-green views compare only compatible complete campaigns. Refreshing an existing report from retained evidence has zero provider calls and is a new presentation of the old measurement, not a new model run.
Existing public histories are available at Runner protocol history and Runner Product E2E history. The consolidated eval hub is at pages.paperclip.ing/evals.
For a repeatable workflow, use the matching skill: paperclip-evals, add-runner-eval, or add-product-e2e-eval.
Install the authoring skills
The reviewable sources live in this repository's .agents/skills. For a
multi-repository workspace, install the three skills at
~/paperclipai/.agents/skills (not ~/paperclipai/skills). From the Paperclip
checkout, run:
for skill in paperclip-evals add-runner-eval add-product-e2e-eval; do
install -d "$HOME/paperclipai/.agents/skills/$skill"
install -m 644 ".agents/skills/$skill/SKILL.md" \
"$HOME/paperclipai/.agents/skills/$skill/SKILL.md"
done
This replaces only the three named skill entrypoints. Run it again after updating their tracked sources. Each skill locates the repository independently of its installation directory.
Maintain the public hub
The hub is a static directory with two links to the existing history systems. It displays a dated snapshot, not a live scoreboard. It does not run models, create another result archive, or change the existing campaign URLs.
Build from the public history feeds and check its summary logic:
python3 -m unittest discover -s scripts/evals-hub -p 'test_*.py'
python3 scripts/evals-hub/build.py --output .paperclip/evals-hub
The hub checks need Python 3 and do not call model providers.
For offline checks, pass --history-dir <directory> containing
runner-protocol-evals-history.json and runner-e2e-history.json.
For a pre-merge preview, pass --docs-ref <branch-or-sha> to link the guide
at that revision. The default guide link uses master.
Publish with the Paperclip page helper
and the configured page-uploader credentials. Use Bash 4 or newer; macOS's
system Bash 3 cannot run this helper. On macOS with Homebrew Bash installed,
put $(brew --prefix bash)/bin first in PATH before these commands:
export PAPERCLIP_PAGE_BUCKET=pages.paperclip.ing
export PAPERCLIP_PAGE_BASE_URL=https://pages.paperclip.ing
export AWS_REGION=us-east-1
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals --dry-run
bash .agents/skills/paperclip-page/scripts/publish.sh .paperclip/evals-hub --slug evals
For later refreshes, rebuild in the same output directory and publish with
--update. Keep its ignored .paperclip-page/state.json ownership record;
without that record, the helper will refuse to overwrite an existing prefix.
Verify the public page and its links after publication. This manual refresh
does not add a scheduled workflow. Preserve the measurement date when choosing
a newer rendering of the same campaign.
Remaining native chat boundaries are in the explicit-only
agent-chat-qualification suite: active task reassignment, user Retry after
verified worker process loss, and multi-turn answers grounded in actual task
records. See the workflow and qualification limits.
The 26 native first-task cells exercise onboarding before native selection
becomes the UI default. Live results and semantic answer reviews must accompany
any qualification claim; catalog presence alone is not a pass.
Lifecycle behavior baseline
The credential-free lifecycle baseline
joins unit, scripted-runner, and database integration assertions to a scenario
inventory before changing narrative-based lifecycle policy. Run
pnpm test:lifecycle-baseline to retain current passes and failures. Its Product
E2E matcher calibration is separate from live execution; unrun live coverage
remains explicitly unmeasured.
The separate live lifecycle baseline
defines 46 real-provider Product E2E cells, including paired narrative probes and
named existing controls on legacy and native Codex. Discover it with
pnpm test:e2e:runner -- --list --suite lifecycle-baseline. Historical execution
results and follow-up coverage are recorded in that suite's guide.
Continuation accounting has an explicit-only eight-cell Product E2E baseline suite, complementing the deterministic lifecycle inventory.
The explicit-only Product E2E api-response-reading suite verifies retrieval of
large saved API responses on local and Daytona native Codex runs. See the
Runner E2E guide.