mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
codex/opencode-stock-routing-qualified
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3c561642b4 |
fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask for decisions and optional details through cards in chat. > - A clear approval in a message can leave the matching card pending. > - An unanswered question can also block an unrelated later reply. > - Decisions need a saved source message, while optional questions need to remain answerable in history. > - This pull request records conversational decisions and lets users move on from questions and answer them later. ## Linked Issues or Issue Description **What happened?** Native Claude and Codex could act on approval in chat while the original approval card stayed pending. Pending question forms stayed above the composer, were absent from history, and could suppress later chat replies. A late native question answer could wait for a finished run to reconnect. **Expected behavior** The active agent records a clear approval or refusal against the exact card and user message. Ambiguous replies do not grant consent. Users can send another message without answering a question. The question remains pending in history and can be reopened and answered later. The saved answer reaches the agent. **Steps to reproduce** 1. Ask an agent to propose work with a confirmation card, then approve it in chat. 2. Check that the original card records that approval before work starts. 3. Ask an interactive question, send an unrelated message, and reload. 4. Open the unanswered question from history and submit an answer. Related work: #14408 added completion delivery. #14607 tests completion reporting turns. Neither records conversational answers on approval cards. ## What Changed - Add a confirmation endpoint backed by a user comment, with schema validation, OpenAPI discovery, and native Plan-mode access. Ask mode remains read-only. - Check company, active run, actor, current session, message provenance, revision, and resolver policy. Save the decision and audit in one transaction. Retries do not repeat effects. Emit resolution telemetry after commit. - Give fresh and resumed chat turns the actual pending confirmation identities. Teach agents to save clear conversational decisions before acting and to clarify ambiguity. - Keep unanswered Agent Chat questions as compact history entries. A newer user message closes the old form. Question cards never contribute to composer pending counts or navigation, including after dismissing a fresh form. The history card is the sole reminder; clicking it restores that exact form and draft. - Preserve Agent Chat questions when later messages or questions arrive. Historical ordinary inputs no longer gate later chat replies. Current-run requests, task execution, and governed approvals keep their gates. Remove the special acknowledgement-publication proof helpers that this rule replaces. - Route answers to finished native runs through durable fresh-wake delivery, with existing idempotency and source-question context. Settle late replies against contiguous completed conversation turns and freeze their history replay; failed, unhandled, and newly arriving messages remain actionable. - Add real-component Storybook scenarios, database and UI regressions, and a three-turn native Claude/Codex E2E case. Capture distinct, UI-ready screenshots and report the individual assertions. ## Verification - Focused decision/publication/UI regressions after merging master: 288 passed; subsequent UI draft, failed-send, and conversation checks: 199 passed. - Native question and durable delivery regressions: 106 passed, including all four terminal run states and exactly-once late delivery. Seven targeted regressions fail against the original implementation and pass with the fix. - Latest conversation/decision/native-delivery regressions after the master merge: 121 passed. Covers completed progress, missing or failed intervening turns, new messages during a late reply, stale sessions, and frozen retry/replay boundaries. Four new assertions fail before the ordering fix. - E2E support suite after the master merge: 792 passed. Negative controls reject expired cards, wrong questions/answers, stale or missing replies, unrelated clarification forms, and unexpected tasks. - The embedded-browser walkthrough caught one additional defect: dismissing a fresh question still showed a composer badge. Both Cancel and close-button regressions failed before the fix. The fix at `65f2ade12` passes 170 chat-thread tests and 792 E2E support tests. After merging master, 232 chat-thread/confirmation tests, server/UI typechecks, and token gates pass. The preview and two-provider live E2E pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at that commit. All 55 checks are now successful at `e5512a206` (four conditional checks skipped), including the aggregate verification gate and clean-install canary test. The first attempt was interrupted by simultaneous CI worker shutdowns; one failed-job rerun passed without code changes. - [Published Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on): nine real-component scenarios. Manually exercised move on, reopen, preserve draft, answer later, answer one of multiple questions, and a custom mobile answer in the embedded browser. Retested fresh Cancel and close-button dismissal in the updated build, then reopened and submitted the preserved Green selection and inspected its answered receipt. Static preview has no live model/backend; its callbacks are fixture responses. - [First live campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/) reproduced the late-answer completion-state defect on both providers despite correct saved answers and acknowledgements. It also exposed a valid imperative clarification rejected by the old oracle. Both issues are fixed with regression controls; this failing run is retained as evidence. - [Four-cell qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/) passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous confirmation, each on native Claude and Codex. Inspected saved state, source-message decisions, visible cards, and agent replies. Both late-answer chats settled to waiting; no unrequested tasks were created. [Final branch rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/) passed 2/2 at `142630720`: the same unanswered-question journey after merging master, plus an additional screenshot and browser assertion for the actual late-answer acknowledgement. - [Composer-reminder E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/) passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each, with explicit no-badge assertions before and after reload. Inspected saved pending/answered state, both screenshots with a clear composer, and actual Blue acknowledgements; all five behavioral matchers passed per provider and neither created tasks. Cost coverage is partial; this is bounded workflow qualification. - [Fresh-dismissal E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/) passed 2/2 at `e5512a206`: native Claude and Codex, including fresh Cancel, clear composer, reopen, unrelated message, reload, late Blue answer, and actual agent acknowledgement. All five behavioral matchers pass per provider. Inspected the fresh-dismissal screenshots and saved pending/answered identity; neither created tasks. Cost coverage is partial (4/6 runs). - Prior evidence remains available in [the earlier campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/). Its early loading screenshot and overwritten final capture prompted the UI-ready, distinct screenshot fixes. ## Risks - The model interprets intent. The server verifies permission and provenance; it does not infer consent from text. Ambiguous and unrelated replies are not approvals. - Historical questions can accumulate. They remain visible, pending, and answerable; no automatic answer or expiry is invented. - The change to completion gates is scoped to Agent Chat and ordinary historical inputs. Current-turn and governed approvals retain their existing controls. - Live qualification is limited to the selected stories. Broader native onboarding finalization remains separate work. - No database migration. Telemetry adds no fields or values; the contract and README document the commit boundary. Privacy review was requested on the PR. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser-test orchestration. The exact model ID and context-window size are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
83076d7e7c |
feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage. Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3f4f8b37ba |
fix: grade Codex clarification and refusal outcomes from evidence (#14570)
Grade clarification lists, obsolete unstarted wakes, and refusal cancellation from persisted evidence. Preserve execution and ownership assertions, add boundary regressions, and version the affected eval definitions. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
cea8dda472 |
test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users can delegate work through onboarding and Agent Chat. > - A completed task does not prove that its result reached the original conversation. > - Existing tests do not isolate completion after the source chat becomes idle. > - This pull request adds four explicit native-runner probes across Claude and Codex. > - The probes preserve the result and reply so we can separate delivery failures from inaccurate answers. ## Linked Issues or Issue Description Refs #13775. Refs #13813. These evals extend native-runner qualification. They measure completion updates before we choose a product change. ## What Changed - Add the opt-in `completion-updates` suite with two stories for each native provider. - Test completion in the existing onboarding task flow and after an Agent Chat handoff becomes idle. - Gate the chat worker on a brief inside its managed project workspace. Prove the source is idle before releasing the worker. - Check durable task completion, saved output, a subsequent source reply, and rendered access to the result. - Preserve replies, task state, screenshots, run events, and a separate semantic review rubric. - Add grader regression tests and update the documented eval contract. - Preserve the suites added on master and include four completion cases in the 306-cell catalog. Production behavior and prompts are unchanged. ## Verification - Passed all 565 eval support tests across 45 files after merging current master: `node node_modules/vitest/vitest.mjs run --config tests/runner-e2e/vitest.config.ts`. - Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p tests/runner-e2e/tsconfig.json`. - Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs tests/runner-e2e/launch.ts --list --suite completion-updates`. - Four-cell behavior campaign on source `ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`: https://github.com/paperclipai/paperclip/actions/runs/36072337485. - A screenshot-only follow-up waits for the restored source reply to render after result-link navigation. Its one-cell Claude onboarding verification passed on final head: https://github.com/paperclipai/paperclip/actions/runs/36075716141. Corrected report: https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/. The original four-cell onboarding screenshots caught navigation loading; its saved reply evidence remains valid. The follow-up again found stale wording: "That work will run next" was posted 38 seconds after the child was Done. The four-cell campaign keeps its original source and measurements. - Published evidence: https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/. - Suite definition: `afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`, version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local execution, one attempt per cell. All four cleanup checks passed. Onboarding billing coverage is partial; reported zero cost must not be read as a free run. | Story | Automated delivery/access | Separate semantic review | | --- | --- | --- | | Codex onboarding | Pass | Pass: accurate completion reply with an accessible result | | Claude onboarding | Pass | Fail: reply says it will save the note once the task runs, after the note is already saved and the task is Done | | Codex idle chat handoff | Fail | Worker completed and saved the note; no completion reply during the full observation window | | Claude idle chat handoff | Fail | Worker completed and saved the note; no completion reply during the full observation window | Both chat cases positively recorded the source waiting and the worker at the brief gate before release. Both saved outputs include the brief-only start time. The opt-in campaign is red because it exposes current behavior. It is not a required merge gate. The PR does not fix that product behavior. Semantic review is a recorded human/agent assessment of retained evidence; it is not an automated prose-quality judge. - Second campaign: https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex chat reached the idle boundary and completed its task, then received no completion reply during the full window. Claude onboarding again returned a stale handoff answer. Claude chat exceeded the prior 110-second handoff setup budget; this revision raises that bounded setup window to 180 seconds. - Retained baseline: https://github.com/paperclipai/paperclip/actions/runs/36069427676. Onboarding passed delivery/access for both providers, but Claude gave a stale handoff answer. Chat cases stopped at fixture problems; they do not establish a completion-delivery failure. This revision fixes the workspace path and competing reference requirements. - On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR checks passed and two were skipped, including typecheck, tests, and build. Broad checks ran in CI, not locally. That head received Greptile 5/5 with no unresolved findings. The unchanged mobile repository-settings browser test passed on one targeted retry after a detached/disabled Save-button timeout. - Merged current master in `9b4491e1f` and resolved the catalog-count conflict. Eval support tests and eval TypeScript pass locally. All individual CI jobs passed on this merge commit, including build, typecheck, server tests, runner checks, and browser shards. The final aggregate check also passed: 54 checks passed and two were skipped. Greptile reviewed this exact commit at 5/5 with no unresolved findings. ## Risks - These explicit probes can expose current product failures. They do not change the default paid test selection. - Mechanical delivery and result access do not establish answer accuracy. The preserved reply still requires semantic review. - A fixture failure before the idle boundary or worker completion cannot establish a completion-update failure. - The handoff setup window lasts three minutes. The worker brief wait is bounded at four minutes. The observation window lasts two minutes after worker completion. It retains later replies without erasing earlier accessible delivery. ## Model Used OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository inspection, code execution, and GitHub tool use. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
43acbcc398 |
fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task state to provider sessions. > - Follow-up turns must retain provider memory and carry new user direction. > - Lost session IDs caused repeated context and extra input tokens. > - Native question answers and approval races could leave valid work blocked. > - This pull request repairs those paths and adds regression coverage. > - Agents can continue accepted work without repeating the conversation or losing the user's answer. ## Linked Issues or Issue Description Refs #13574. That merged PR shortened continuation prompts and moved question instructions into tool documentation. This change preserves sessions and fixes failures exposed by broader testing. Related runtime work: #13408 and #13410. **What happened?** Native follow-up turns could lose the provider session ID. Completion guidance could replace the original task with its latest comment. Claude native questions could remain pending after the user answered. Approval during a running tool call could suspend the run before the tool response arrived. Onboarding and chat handoff instructions also caused repeated planning or missing plan documents. **Expected behavior** Reuse a valid provider session. Send only new events when that session already has the history. Preserve the task requirements and apply later user direction. Store the question answer and deliver it to the waiting run. Finish governed tool responses before suspending. Execute the accepted plan without asking for the same approval again. **Steps to reproduce** Run the continuation, local-session-integrity, first-task, and agent-chat suites with native Codex and Claude. Include provider-question-bridge, accept-while-running, and plan-handoff. **Paperclip version or commit** This branch is based on master |
||
|
|
d0b67bfe71 |
feat: queue approvals and answers during active runs (#13539)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users guide running agents through messages, questions, and approval cards. > - Messages already wait in a queue when an agent is running. > - Card responses did not appear in that queue. Some question answers also steered a later run without a user click. > - A fast approval could invalidate the agent's review handoff and cause it to stop its own run. > - This pull request gives card responses the same queue controls and preserves the exact response during delivery. > - Users can wait for completion or explicitly send the response with Interrupt or Steer. ## Linked Issues or Issue Description Refs #13517, which is merged. This PR targets master and adds queued interaction responses on top of the onboarding changes. Related continuation work: #10519 and #12866. **What happened?** Accepting a proposal while its source run was active left a saved response outside the message queue. The agent could then lose its review path, reassign the task, and cancel itself. Answers to older questions could also steer another active turn without a click. **Expected behavior** Save the response immediately. Queue its continuation behind the active run. Deliver it after completion, or when the user explicitly chooses Interrupt or Steer. Preserve approval revisions and answer choices. **Steps to reproduce** 1. Let an agent publish a confirmation card while its run is still active. 2. Accept the card before the agent finishes its review handoff. 3. Inspect the message queue and the task's next run. **Paperclip version or commit** Reproduced on da8a3876c with the onboarding changes from #13517. **Deployment mode** Local development from source. The fix covers legacy adapters and native Runner turns. ## What Changed - Project resolved cards into the existing queue as immutable responses. Keep answers and exact approval revisions. - Require an explicit click to steer a response into a compatible native turn. Use Interrupt when a fresh session is required. - Preserve typed response context through interruption, cleanup waits, and normal queue promotion. Keep the direct answer channel for a provider blocked on its original question request. - Accept the source run's review handoff after its card resolves. Reject stale agent reassignment that would orphan a queued response. - Add deterministic regression tests and an `accept-while-running` case to the first-task suite. Require recorded timestamp overlap before that case can pass. - Keep the first-task skill name out of user-facing messages. ## Verification - Red-green: the original route failed the queue regression; the changed route passes it. - Focused server/UI tests: 139 passed, including 64 queue-route tests. - Runner harness unit tests: 314 passed. - Server, UI, and Runner E2E typechecks passed. UI token gates passed. - Full repository typecheck and build passed. Server typecheck passed again after review fixes. - Review regressions: 165 queue/reopen route tests, 53 wake admission tests, and 18 run identity tests passed. Approval acknowledgement recovery and both message/approval arrival orders are covered. - Full local test run: 12,401 passed; three new admission regressions ran against a cached pre-fix module. A fresh run of that entire suite passed (53 tests). The complete CI suite passed on the final commit. - Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional Storybook checks skipped. Every server/workspace/browser shard, Runner verification, build, typecheck/release registry, canary, policy, and security check passed. Greptile: 5/5, no unresolved threads. Earlier interrupted CI workers were replaced by this fresh complete run. - After integrating the updated parent: 314 harness tests, 119 queue/admission tests, 44 onboarding/question-delivery tests, and 13 native recovery tests passed locally. Full repository typecheck and build passed. - Clarified the skill wording preference: routine replies describe the action without announcing the internal skill; direct questions and permission/security/execution disclosures remain truthful. - The paid `accept-while-running` scenario is registered for all four local first-task profiles. It has not been run against a model in this change. - Rebased onto the merged parent at `11921075a`; the resulting tree exactly matches the locally verified integration tree. Final-head CI on `b53054807` passed: 54 successful checks, 2 optional Storybook checks skipped, no failed checks. Every new server/browser shard, aggregate verify/e2e gate, Runner, typecheck, build, canary, and security check passed on the first attempt. Greptile reviewed this exact head at 5/5 with no unresolved threads. ## Risks - Responses now wait instead of implicitly steering another active turn. A provider blocked on the original question still receives its answer directly. - Approval receipts cannot be edited, discarded, or reordered as comments. This preserves the recorded decision. - Interruption must still prove that the prior execution stopped. The tests cover cleanup waits and duplicate delivery. - The new paid overlap case can be unexercised if the model finishes before the click lands. It cannot pass without evidence of overlap. - No database migration is required. This repairs the existing approvals and execution controls; it does not implement the roadmap's work-stream queues. ## Model Used OpenAI GPT-6 through Codex. The exact deployed model ID and context-window size were not exposed in this session. Capabilities used: agentic reasoning, repository inspection, code editing, terminal commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
11921075a4 |
Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first task helps a new user define and approve useful work. > - That workflow needs reusable instructions and tests against the production experience. > - Native Codex and Claude must load the assigned skill, including after resume. > - Maintainers need recorded conversations and precise failed checks to judge regressions. > - This pull request adds the first-task skill and a suite in the shared Runner E2E harness. > - It keeps behavior results separate from informational quality scores and incomplete recordings. ## Linked Issues or Issue Description **What existing behavior does this improve?** The first onboarding task and the Runner E2E report used to review it. **Current behavior** Onboarding embeds its policy in a hidden brief. Native Codex drops the skill-instructions setting at the Rust boundary. The shared E2E harness has no onboarding suite or full conversation view. **Proposed behavior** Assign and invoke `/first-task` for the onboarding task. Send selected Codex skills as structured protocol inputs. Run twelve scenarios across legacy Codex, legacy Claude, native Codex, and native ACPX Claude. Include all 48 cells in full campaigns. Show recorded chat, question and approval cards, exact checks, instructions, and billing in the shared dashboard. **Reason and benefit** Measure the real onboarding experience before changing prompts. Distinguish infrastructure failures, behavior failures, and unexercised journey steps. **Breaking changes** No database migration or production API change. First-task instructions now live in an assigned skill. The user-edited persona is preserved; the skill includes the maintainer-approved proposal-mode mapping and saved-plan requirement. Related: #11043 is earlier onboarding work. #13422 already fixes native Claude model pinning, context delivery, and read permissions on master; this branch includes those fixes through its base. The new Claude recovery test supplements them. ## What Changed - Extract and assign the first-task skill while retaining the production greeting and opening question. - Carry the Codex skill-instructions flag through thread start and resume. Resolve explicit task skill references only against assigned skills and send native skill inputs. - Invoke an unambiguously selected assigned skill through Claude ACPX’s native slash-command parser on initial and resumed turns, retaining the entire task/wake envelope as its argument. Do not carry that invocation into ordinary tasks. - Restore the saved single-task proposal modes: confirmation card, or saved plan with revision-targeted checkbox approval. Explicit plan requests also require a saved plan. - Add first-response and complete-journey cases with fixed user facts, acceptance checkpoints, durable outcome checks, and accounting for child runs. - Fail the eval when choice questions have fewer than two real options. Recognize planning documents without treating them as completed work. - Add optional, bounded quality judging as explicit post-processing. - Render full conversations and static interaction cards in the shared report. Conversations start folded. Show original and regraded results and incomplete journeys distinctly. - Keep credential-persistence scanning outside the first-task behavioral suite; retain public evidence redaction. - Refresh generated capability references after the API-reference edits. - Correct shared native question guidance and tool schemas: choices need at least two meaningful options; open-ended questions use canonical text fields with the required compatibility payload. Verify both formats through real tool-authority persistence. - Disable announcements automatically for every isolated Runner E2E process and label the gallery environment/provider/target explicitly. - Remove CI races in the GitHub connection browser test and native session recovery test by waiting for the actual async work before asserting its results. ## Verification - `pnpm exec vitest run server/src/services/onboarding-first-task-assets.test.ts server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19 passed. - `pnpm --dir packages/paperclip-runner exec vitest run src/drivers/acpx/runtime-host.test.ts src/drivers/acpx/native-skill-prompt.test.ts src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command forwarding and the 1 MiB input boundary both failed before their fixes and passed afterward. Coverage includes changed skills on reopen, approval context, and an ordinary subsequent task. - Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64 first-task fixture and grader tests also pass. - Full repository typecheck and build passed locally. Server typecheck and Runner build passed again after the native-command change. - Full GitHub Actions CI passed on `23e56447b`: all server/workspace/browser shards, Runner verification, typecheck/release registry, build, canary, policy, and Docker checks. Greptile reviewed this exact head at 5/5 with no unresolved threads. The earlier broad local run had database startup/timing failures that passed isolated retries; the complete remote suite is green. - Merge verification against current master: 312 harness tests and 13 native recovery tests passed. Regenerated semantic contracts and fixture hashes pass their consistency check. Full local typecheck and build also passed on the stacked queue branch. After merging the latest master and preserving the GitHub setup timing regression in the split browser suite, both focused GitHub browser tests passed. Three CI timing/startup flakes passed local verification and one remote retry; all latest-head checks are green. - Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local mock API confirmed that `/skill-name` expands the assigned skill body before the model request and retains the task arguments. A prose mention does not. The probes made no paid model calls. The ACP probe used the current first-task skill body and retained the wake arguments. - [Full 48-case campaign and report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/): 44 passed after three interrupted Codex cases completed in targeted reruns. Original results, regrades, and all 51 executions remain in the report provenance. - [Claude campaign after the shared-question fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/): 10/12 passed with zero single-option failures. All 12 recorded the current assigned skill and corrected guidance. The failures exposed skipped skill invocation and a missing saved plan. This PR adds native command invocation and explicit saved-plan instructions; the subsequent report below still shows behavior failures. - [Fresh 12-case Claude report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/) at `78452129e`: 10/12 pass after correcting two false proposal-matcher failures. The recordings said “Here is the task I will create and run/complete” in approval cards; the old matcher missed that word order. Regression tests failed before the fix and pass after it. Original results and offline regrade provenance remain linked. No agent rerun was needed. Zero single-option-question failures; two behavior failures remain: direct work before acceptance on a plain first message, and an explicit plan request without a saved plan. Neither check was relaxed. The follow-up `82087ac7e` fixes command-prefix size accounting; `94aefb1f3` fixes only that proposal matcher. - Report browser checks confirm folded conversations, rendered cards, explicit Local/Daytona labels, and no page errors. The published-object audit scanned 1,306 text files across 2,154 objects with no credential-format findings or prohibited files. Image pixels and unknown token formats are outside that scan. ## Risks - Model behavior is nondeterministic. One campaign is evidence, not a guarantee. The two remaining Claude behavior failures are visible in the report and require further product work; this PR does not claim all onboarding scenarios pass. - The suite checks persisted Paperclip effects. It cannot prove the absence of arbitrary external effects. - Historical recordings can miss later journey steps. These remain incomplete, never passes. - Native profiles switch runtime after the production onboarding wizard because it does not yet expose a native option. - Quality scores are informational and cannot override behavioral failures. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, and code execution. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |