Files
PaperClipAI/server
DottaandPaperclip 428b617fc9 test(evals): qualify native question resume paths (#15616)
## Thinking Path

> - Paperclip lets agents pause tasks for human answers and continue the
work.
> - Native questions can yield a run or pause inside the provider's
running turn.
> - Those paths have different run identities and different forms.
> - The documentation-placement test requires a semantic three-turn
journey.
> - A valid provider question therefore needs separate behavioral
qualification.
> - This pull request adds an opt-in two-answer journey with exact
source and resume checks.
> - Qualification exposed a Vite startup crash in idle-handler
traversal, so the branch also distinguishes Connect mount paths from
Express Route objects.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Product E2E evaluation of native task questions and answer delivery.

**Current behavior**

The semantic documentation case rejects the provider adapter's optional
Other field before answering. Its three-run expectation also cannot
prove a provider question that resumes within the same run.

**Proposed behavior**

Keep that semantic test intact. Add a separate opt-in case that
classifies each recorded question by authoritative identities. Accept
the optional Other field only for a verified provider question. Require
both user answers, the correct continuation, and one saved final
document.

Related: #15564 added strict lifecycle checks. #15522 introduced
idle-request tracking; this PR corrects its Connect/Vite compatibility
after reproducing the startup crash. #14591 changes ACP question drafts
and provider support; this PR has no overlapping files with that change.

## What Changed

- Add the explicit-only `question-resume` suite for native Codex and
ACPX Claude. Reuse the existing neutral choice-then-text task request.
- Bind each question to its company, task, agent and run. Require an
applied semantic tool receipt or the exact provider request identity.
- Require a provider answer to continue the same run. Require a semantic
answer to start one new run with the matching interaction and source-run
wake fields.
- Preserve question forms and both real board answers through the final
checkpoint. Reject missing inputs, duplicate answers, extra runs and
substituted identities.
- Keep the existing lifecycle, output, task ownership and budget checks.
Limit the new case to one attempt and one to three runs.
- Add mutation calibrations and document the separate qualification
claims. Keep production guidance, scheduling, provider code and the
historical semantic case unchanged.
- Fix startup with Vite middleware: treat Connect string mount paths
separately from Express Route objects when traversing idle-request
handlers. A real-Vite regression verifies startup and retained async
work through an idle hold.

## Verification

- 81 focused continuation, question and calibration tests passed before
review. After the reference-preservation correction, all 50
question/resume tests pass, including two new negative cases that failed
against the previous grader.
- Final eval support suite passes: 1,921 Vitest tests and 128 Node
tests, with one intentional skip.
- Eval and workspace typechecks pass. Workspace build passes.
- Catalog discovery finds only the two declared opt-in cells. Existing
continuation remains 23 cells; default paid scope is unchanged.
- Original [campaign
37831726391](https://github.com/paperclipai/paperclip/actions/runs/37831726391),
source `1ac4b2873466896ff2892f49be82904927b6c57d`, failed during server
startup before any test or provider run. Its original
`transient_infrastructure` result and `cleanup: not_started` are
preserved. Local reproduction identifies the Connect string route
traversal crash; the question grader was never reached.
- Setup-corrected [campaign
37833872898](https://github.com/paperclipai/paperclip/actions/runs/37833872898)
measures source `9781a772389d41421df55759004676e1b905e512` with trusted
workflow `8e59efc50b162126f33816841e06509634eba65c`. The workflow bytes
are unchanged from the first dispatch. Same single selected native
Claude cell, one attempt, at most three runs, ten-minute cell deadline,
and 1,000-cent company/agent hard stops. Original result **PASS**, 24/24
checks and cleanup pass. Exactly three successful native Claude runs:
two applied Paperclip `request_human_input` questions and one finishing
run. Both real board answers persist, exact interaction/source-run
response wakes match, one revision-1 task document contains Afternoon
and the supplied reference, and final task state is done with no active
lock, retry, recovery, monitor, pending interaction or child task. No
model retry.
- Current review head `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e` differs
from the measured source only in the question/resume grader and its
tests (11 insertions, 3 deletions). The saved document must contain the
complete submitted reference, including its prefix and punctuation.
Separate provider-free replay of the retained successful campaign passes
all 24 continuation/lifecycle checks under this stricter grader. The
original result and its measured source remain unchanged; no new
provider run was made.
- Observed paths were **semantic → semantic**. The provider built-in
question/optional Other path was not exercised by this new live attempt.
Its form acceptance has retained-original calibration; same-run
answer/resume has deterministic positive/negative calibration. This
neutral-task pass alone does not qualify the provider path; the
subsequent explicit bridge trial below is separate evidence. Neither
trial regrades the old Claude failure or establishes a causal
behavior/performance comparison.
- Subsequent explicit provider-path [campaign
37841107684](https://github.com/paperclipai/paperclip/actions/runs/37841107684)
measures the final PR source `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`,
using trusted workflow `2a5f65c9ae501b49b2a38ce7edd04209b806c9d0`. Exact
cell: `continuation.runner-acpx-claude.local.provider-question-bridge`,
Claude `claude-sonnet-5`, suite fingerprint
`fc77203bf85fd99b06ebc53cd52982186c0ae0dc369fcc38a9ec3c737e9a31e5`.
**Original PASS, 15/15 checks, cleanup passed, one actual provider run
and zero eval rerolls.** The built-in question produced one choice plus
its optional Other companion. The browser selected the requested
reference, left Other blank, and submitted. Exact runtime
request/response IDs and the selected option match the saved
interaction; the same paused native run created one revision-1 document
and finished Done with no pending interaction, lock, retry, recovery,
monitor or child task. Independent evidence and runtime-receipt audits
pass; marked initial/final screenshots were visually checked. [Original
public
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37841107684-1/index.html).
- This separate provider-directed trial qualifies one choice-answer
bridge and traversal of an empty optional Other field. It does not
establish natural tool selection, a typed Other answer, two consecutive
built-in questions, free-text-only input, restart recovery or general
reliability. No production, prompt or grader change was made for this
dispatch. The old failure and the separate neutral semantic-path result
retain their original grades.
- Preserve within-run friction: one `write_task_document` input-schema
rejection and two HTTP 409 “Document does not exist yet” responses
preceded a successful `write_document` call. These are three
unsuccessful tool calls inside the same run, not three extra provider
runs; no flawless document-tool behavior is claimed. Billing reports
token usage but remains unpriced/incomplete, so actual charges are
unknown and local/hosted runtime is unmetered. Both company and agent
hard stops were 1,000 cents, with one attempt, one selected cell,
concurrency one and a ten-minute cell deadline.
- Billing records all three model runs with token usage, but no priced
dollar receipt (`unpriced`, `complete: false`). Actual charges are
unknown; local/hosted runtime is unmetered. Numeric zero reported cost
does not mean free.
- Startup repair: real-Vite reproduction passes; 35 targeted
idle-tracking tests pass, 34 database-dependent checks skip because
embedded PostgreSQL is unavailable locally. Workspace typecheck and
build pass again after the repair.
- Full local repository database tests were not repeated because
embedded PostgreSQL was unavailable in the preceding workspace
verification. Full Linux CI now passes on the final review head.

- Final-head [CI run
37837804284](https://github.com/paperclipai/paperclip/actions/runs/37837804284)
passes. Complete current-head audit: 53 successful check-runs, two
intentional Storybook skips, and separate Snyk success. The ready
transition’s contributor and security checks also pass (security scan:
no findings). [Fresh
review](https://github.com/paperclipai/paperclip/pull/15616#issuecomment-6068092601)
is 5/5 on `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`; zero unresolved
threads and no merge conflicts.

## Risks

This is a new behavioral definition. The neutral two-question trial
records semantic-tool use; the separate provider-directed trial records
one built-in question. They do not establish consistent tool selection,
documentation placement, default hiring, remote environments, crash
recovery or general reliability. The separate explicit bridge trial
covers one provider choice, an empty optional Other companion and
same-run completion. Typed Other, two consecutive built-in questions and
free-text-only provider input remain unqualified. The observed
document-tool schema rejection and two 409 responses are retained as a
separate follow-up, without attributing their cause. The historical case
and its original grades remain intact. The reference question must still
be text-only; the oracle does not turn an arbitrary extra question into
an optional companion. Missing or unfamiliar identity evidence fails
closed. The only production change distinguishes Connect string mount
paths from Express route objects during idle-handler traversal. Unknown
route objects still fail closed. No API, schema, migration or workflow
change is included.

## Model Used

OpenAI Codex, GPT-6-based, with tool use and code execution. The exact
serving model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 16:00:01 -05:00
..