mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 20:50:08 +02:00
## Thinking Path > - Paperclip lets agents pause tasks for human answers and continue the work. > - Native questions can yield a run or pause inside the provider's running turn. > - Those paths have different run identities and different forms. > - The documentation-placement test requires a semantic three-turn journey. > - A valid provider question therefore needs separate behavioral qualification. > - This pull request adds an opt-in two-answer journey with exact source and resume checks. > - Qualification exposed a Vite startup crash in idle-handler traversal, so the branch also distinguishes Connect mount paths from Express Route objects. ## Linked Issues or Issue Description **What existing behavior does this improve?** Product E2E evaluation of native task questions and answer delivery. **Current behavior** The semantic documentation case rejects the provider adapter's optional Other field before answering. Its three-run expectation also cannot prove a provider question that resumes within the same run. **Proposed behavior** Keep that semantic test intact. Add a separate opt-in case that classifies each recorded question by authoritative identities. Accept the optional Other field only for a verified provider question. Require both user answers, the correct continuation, and one saved final document. Related: #15564 added strict lifecycle checks. #15522 introduced idle-request tracking; this PR corrects its Connect/Vite compatibility after reproducing the startup crash. #14591 changes ACP question drafts and provider support; this PR has no overlapping files with that change. ## What Changed - Add the explicit-only `question-resume` suite for native Codex and ACPX Claude. Reuse the existing neutral choice-then-text task request. - Bind each question to its company, task, agent and run. Require an applied semantic tool receipt or the exact provider request identity. - Require a provider answer to continue the same run. Require a semantic answer to start one new run with the matching interaction and source-run wake fields. - Preserve question forms and both real board answers through the final checkpoint. Reject missing inputs, duplicate answers, extra runs and substituted identities. - Keep the existing lifecycle, output, task ownership and budget checks. Limit the new case to one attempt and one to three runs. - Add mutation calibrations and document the separate qualification claims. Keep production guidance, scheduling, provider code and the historical semantic case unchanged. - Fix startup with Vite middleware: treat Connect string mount paths separately from Express Route objects when traversing idle-request handlers. A real-Vite regression verifies startup and retained async work through an idle hold. ## Verification - 81 focused continuation, question and calibration tests passed before review. After the reference-preservation correction, all 50 question/resume tests pass, including two new negative cases that failed against the previous grader. - Final eval support suite passes: 1,921 Vitest tests and 128 Node tests, with one intentional skip. - Eval and workspace typechecks pass. Workspace build passes. - Catalog discovery finds only the two declared opt-in cells. Existing continuation remains 23 cells; default paid scope is unchanged. - Original [campaign 37831726391](https://github.com/paperclipai/paperclip/actions/runs/37831726391), source `1ac4b2873466896ff2892f49be82904927b6c57d`, failed during server startup before any test or provider run. Its original `transient_infrastructure` result and `cleanup: not_started` are preserved. Local reproduction identifies the Connect string route traversal crash; the question grader was never reached. - Setup-corrected [campaign 37833872898](https://github.com/paperclipai/paperclip/actions/runs/37833872898) measures source `9781a772389d41421df55759004676e1b905e512` with trusted workflow `8e59efc50b162126f33816841e06509634eba65c`. The workflow bytes are unchanged from the first dispatch. Same single selected native Claude cell, one attempt, at most three runs, ten-minute cell deadline, and 1,000-cent company/agent hard stops. Original result **PASS**, 24/24 checks and cleanup pass. Exactly three successful native Claude runs: two applied Paperclip `request_human_input` questions and one finishing run. Both real board answers persist, exact interaction/source-run response wakes match, one revision-1 task document contains Afternoon and the supplied reference, and final task state is done with no active lock, retry, recovery, monitor, pending interaction or child task. No model retry. - Current review head `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e` differs from the measured source only in the question/resume grader and its tests (11 insertions, 3 deletions). The saved document must contain the complete submitted reference, including its prefix and punctuation. Separate provider-free replay of the retained successful campaign passes all 24 continuation/lifecycle checks under this stricter grader. The original result and its measured source remain unchanged; no new provider run was made. - Observed paths were **semantic → semantic**. The provider built-in question/optional Other path was not exercised by this new live attempt. Its form acceptance has retained-original calibration; same-run answer/resume has deterministic positive/negative calibration. This neutral-task pass alone does not qualify the provider path; the subsequent explicit bridge trial below is separate evidence. Neither trial regrades the old Claude failure or establishes a causal behavior/performance comparison. - Subsequent explicit provider-path [campaign 37841107684](https://github.com/paperclipai/paperclip/actions/runs/37841107684) measures the final PR source `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`, using trusted workflow `2a5f65c9ae501b49b2a38ce7edd04209b806c9d0`. Exact cell: `continuation.runner-acpx-claude.local.provider-question-bridge`, Claude `claude-sonnet-5`, suite fingerprint `fc77203bf85fd99b06ebc53cd52982186c0ae0dc369fcc38a9ec3c737e9a31e5`. **Original PASS, 15/15 checks, cleanup passed, one actual provider run and zero eval rerolls.** The built-in question produced one choice plus its optional Other companion. The browser selected the requested reference, left Other blank, and submitted. Exact runtime request/response IDs and the selected option match the saved interaction; the same paused native run created one revision-1 document and finished Done with no pending interaction, lock, retry, recovery, monitor or child task. Independent evidence and runtime-receipt audits pass; marked initial/final screenshots were visually checked. [Original public report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37841107684-1/index.html). - This separate provider-directed trial qualifies one choice-answer bridge and traversal of an empty optional Other field. It does not establish natural tool selection, a typed Other answer, two consecutive built-in questions, free-text-only input, restart recovery or general reliability. No production, prompt or grader change was made for this dispatch. The old failure and the separate neutral semantic-path result retain their original grades. - Preserve within-run friction: one `write_task_document` input-schema rejection and two HTTP 409 “Document does not exist yet” responses preceded a successful `write_document` call. These are three unsuccessful tool calls inside the same run, not three extra provider runs; no flawless document-tool behavior is claimed. Billing reports token usage but remains unpriced/incomplete, so actual charges are unknown and local/hosted runtime is unmetered. Both company and agent hard stops were 1,000 cents, with one attempt, one selected cell, concurrency one and a ten-minute cell deadline. - Billing records all three model runs with token usage, but no priced dollar receipt (`unpriced`, `complete: false`). Actual charges are unknown; local/hosted runtime is unmetered. Numeric zero reported cost does not mean free. - Startup repair: real-Vite reproduction passes; 35 targeted idle-tracking tests pass, 34 database-dependent checks skip because embedded PostgreSQL is unavailable locally. Workspace typecheck and build pass again after the repair. - Full local repository database tests were not repeated because embedded PostgreSQL was unavailable in the preceding workspace verification. Full Linux CI now passes on the final review head. - Final-head [CI run 37837804284](https://github.com/paperclipai/paperclip/actions/runs/37837804284) passes. Complete current-head audit: 53 successful check-runs, two intentional Storybook skips, and separate Snyk success. The ready transition’s contributor and security checks also pass (security scan: no findings). [Fresh review](https://github.com/paperclipai/paperclip/pull/15616#issuecomment-6068092601) is 5/5 on `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`; zero unresolved threads and no merge conflicts. ## Risks This is a new behavioral definition. The neutral two-question trial records semantic-tool use; the separate provider-directed trial records one built-in question. They do not establish consistent tool selection, documentation placement, default hiring, remote environments, crash recovery or general reliability. The separate explicit bridge trial covers one provider choice, an empty optional Other companion and same-run completion. Typed Other, two consecutive built-in questions and free-text-only provider input remain unqualified. The observed document-tool schema rejection and two 409 responses are retained as a separate follow-up, without attributing their cause. The historical case and its original grades remain intact. The reference question must still be text-only; the oracle does not turn an arbitrary extra question into an optional companion. Missing or unfamiliar identity evidence fails closed. The only production change distinguishes Connect string mount paths from Express route objects during idle-handler traversal. Unknown route objects still fail closed. No API, schema, migration or workflow change is included. ## Model Used OpenAI Codex, GPT-6-based, with tool use and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>