mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 14:10:50 +02:00
5ee9e751fbb38787c055ee793c3412ecfc89da2e
24
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
96bf004a79 |
fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f56aee5423 |
fix(evals): capture complete durable run event streams (#13977)
## Thinking Path > - Paperclip manages AI agents and records their durable work outcomes. > - Product E2E checks those outcomes through the browser and public API. > - A successful long run can emit more than 1,000 durable events. > - The harness read one page and missed the later completion evidence. > - This pull request reads every page before it checks runtime invariants. > - Invalid or incomplete capture still fails. A missing page cannot produce a pass. ## Linked Issues or Issue Description Related: #13882 (Grok qualification). A duplicate search found no existing event-pagination fix. **What happened?** The local structured-question case in [campaign 36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537) completed the task and passed its six outcome matchers. It failed native runtime invariants because the capture contained exactly 1,000 events. The last captured event preceded the run's completion by more than a minute. The API caps each response at 1,000 rows. The harness did not request the next page. **Expected behavior** Read the complete durable event stream through the public API before checking semantic-result and terminal-event counts. Reject incomplete or malformed evidence. **Steps to reproduce** 1. Complete a native task that emits more than 1,000 durable events. 2. Place the semantic-result and terminal events after row 1,000. 3. Capture the run with the Product E2E harness. 4. Before this fix, the invariant checker sees only the first page. **Paperclip version or commit** Observed at `4196a4cd76db434854b679035e4146c7f69689ce`. The same single-page capture exists on master. The original failed result remains unchanged; missing historical tail evidence is not reconstructed or graded as a pass. ## What Changed - Add a bounded event collector that advances through the public `afterSeq` cursor. - Use it for task success/failure evidence and shared chat run evidence. - Reject invalid pages, missing or non-increasing sequence numbers, repeated cursors, failed later requests, and an exhausted page limit. - Test completion events beyond the first page, exact page boundaries, and malformed evidence. - Correct the existing Everyday catalog test from 38 to the maintained 47 cells. The suite stays explicit-only. - Document the complete-capture requirement and its bound. ## Verification - Product harness typecheck passes. - The full credential-free harness suite passed 477 tests. After adding the chat integration regression, all 34 chat evidence tests pass. - Seventeen pagination tests cover the valid tail and malformed-evidence cases. - `git diff --check` passes. - [Full repository CI](https://github.com/paperclipai/paperclip/actions/runs/36080683422) passes at `aed6f79c089226be79e75dcf390969561ec4f787`: 52 successful checks and two intentional skips, including Rust, typecheck, build, server tests, and browser shards. Greptile gives this exact head 5/5 with no findings. No local Docker, Rust build, browser suite, or model invocation was used for this change. - [Live Grok requalification](https://github.com/paperclipai/paperclip/actions/runs/36080870743) is running on combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, with the runner, scheduler, and evidence fixes. Its full credential-free harness passes all 493 tests and typecheck. Live results are pending; the original campaign remains a failed measurement. ## Risks Long runs need more read-only API requests and larger private evidence files. Capture stops with an explicit error after 100 full pages. This changes neither production APIs nor provider behavior. It does not relax an invariant or change a historical grade. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9a76390cfa |
test(runner): improve blank-page diagnostics and infrastructure coverage (#13824)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner E2E tests verify real tasks and retain evidence for failures.
> - Some exposure tests assumed that port 42000 was free.
> - A blank task page could also fail without enough browser startup
evidence.
> - This pull request tests occupied ports and task reloads, and records
private startup diagnostics.
> - These changes make test failures easier to reproduce and explain.
## Linked Issues or Issue Description
Refs #13815. Report publication already received its production fix in
#13750; this PR adds a regression check for that workflow.
**What happened?**
Three exposure tests assumed that the allocator would select port 42000.
The synthetic failure disappeared when another process occupied that
port. A prior Daytona run also retained an empty task page after
navigation, but its evidence did not record pending modules or
service-worker control.
**Expected behavior**
Exposure tests must exercise the intended failure on the actual assigned
port. Browser failure evidence must distinguish an empty root from
loaded content. A saved task must remain usable after navigation and
reload.
**Steps to reproduce**
Run the exposure regression with the base port pair marked unavailable.
Run the browser-support tests with an unresolved entry module. Open a
saved task under the service worker, navigate to the same URL, and
reload it.
**Paperclip version or commit**
Based on master at
|
||
|
|
846336e5a0 |
test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path > - Paperclip lets people manage agents through ongoing conversations. > - Chat users can change instructions while a provider is already working. > - Existing chat evals wait for each turn to settle before the next message. > - They cannot prove delivery during active work or the saved effect of a correction. > - Existing fixtures also enable Agent Chat through the API rather than the settings UI. > - This PR adds bounded browser workflows and checks their persisted outcomes. ## Linked Issues or Issue Description Refs #13741, #13752, #13750. **What happened?** The chat suites cover planning, delegation, status, and recovery. They lack active-turn follow-ups and the experimental settings lifecycle. A sequential conversation can pass even if messages sent during work are lost. **Expected behavior** A follow-up submitted during a provider turn survives and affects the final reply. A changed launch day appears in the saved plan. Disabling Agent Chat rejects new messages while preserving history; re-enabling resumes the same conversation. **Steps to reproduce** Run the explicit `agent-chat-stories` suite. It selects three local cases for each native Claude and Codex profile. An ordinary provider command waits for a fixture brief file so the browser can send the follow-up at an observed active-run boundary. ## What Changed - Add six opt-in Product E2E cells for settings, active follow-ups, and plan corrections. - Drive experimental settings through the UI and verify disabled sends are rejected by the public API. - Use a bounded file wait in the actual isolated agent workspace, with provider-written readiness and an undisclosed brief reference. - Grade persisted user messages, final replies, native run outcomes, and exact saved plan fields. - Accept active-turn steering or one queued successor; reject lost input, duplicate input, and stale outputs. - Allow one steered run or two sequential runs throughout the shared harness, while preserving exact counts for other cases. - Require a single marker-bearing response attributed to the final provider run. - Unload the development browser client before restarting the server, avoiding reconnect/navigation races without weakening the post-restart memory check. - Add browser regressions for restart isolation and asynchronously saved settings switches. - Document prepared-agent setup, native onboarding limits, and the separate API-tool rollout gate. ## Verification - Eval TypeScript check passed. - Eval support suite: 436 tests passed in 39 files. - New oracle calibration: six tests passed, including plausible invalid outcomes. - Browser support regressions: seven tests passed; the restart regression was observed failing before the fix. - Catalog discovery selects exactly six local native cases and leaves default paid selection unchanged. - [Consolidated existing native chat report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/): master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a browser navigation timeout across restart; the page request returned 200 and the chat rendered. - [Nine targeted restart/replay cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/) passed on `1fe2fe275`, including the original failure, across native Claude/Codex and local/Daytona; all cleanup passed. - [Initial six-story campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817) retained all six failures: asynchronous switch assertions, unavailable fixture paths, and rich-text escaping in raw command comparisons. The corrected fixtures preserve the same behavioral assertions. - [Six-story campaign v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/) on `8232773a0`: 4/6 passed (both settings cases and both Claude interruptions). Codex could not see the host-temp fixture outside its workspace; this failed before follow-up delivery was exercised. - [Four affected interruption cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/) all passed, including cleanup, on definition v3 / `ad6ac0545`. Files live inside the actual agent workspace and the observed run workspace is verified. Both providers saved Friday in the real plan with the undisclosed brief reference; follow-ups persisted while the original run was active. Together with both unchanged settings cases from v2, all six new scenario variants have passing live evidence. - Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful checks, two intentional skips, zero pending/failing checks; mergeable and clean. Fresh Greptile 5/5, zero unresolved findings. - Full typecheck, tests, build, and browser CI passed remotely. One earlier head encountered a signoff-policy browser timing failure; the final head passed that shard. - Local pnpm wrapper could not fetch its version/signature metadata in the restricted environment; local eval checks used the installed Node executables. Repo-wide validation was completed by GitHub Actions. ## Risks These are eval-only changes. The file wait is a timing fixture in the isolated agent workspace, not a production runner hook. Native Codex host-filesystem isolation stays unchanged. It has a two-minute limit and is released in `finally`. The prepared-agent settings case is not full native onboarding: the wizard currently offers legacy adapters. The disabled-entry assertion uses full document navigation, which clears the prior React Query cache; preserved history is checked through the public API and re-enabled chat. No production prompt, rollout default, adapter behavior, or credential policy changes. Active-task reassignment and worker-crash recovery remain outside these new cases. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3790ca2f13 |
fix(runner): repair approval and Stop races and eval infrastructure (#13750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Runner tasks must continue after approval and stop when the user presses Stop. > - Live evals found races at approval delivery and provider startup. > - Browser readiness and CI setup errors also hid the actual task results. > - This pull request fixes those races and the related test infrastructure. > - Regression tests and saved live reports show which cases now pass. ## Linked Issues or Issue Description Companion eval definitions PR: https://github.com/paperclipai/paperclip-evals/pull/25 (AgentCore paused and provider/environment infrastructure). Related: #13741 now supplies the late-startup Stop fence and warm-attachment recovery; this PR retains that fence and extends startup tracking and regression coverage to both native backend paths. #13539 introduced queued approvals during active runs. #13738 fixes child assignment, task replies, and warm process continuity and is already in the base. #13291 concerns automatic continuation of interrupted legacy sandbox runs; this PR fixes native startup cancellation and does not change that recovery policy. **What happened?** An accepted service approval could wait after its source run stopped. Stop could return success before the provider handle existed. Work could then start after Stop, or a cancelled run could be recorded as failed. Some E2E tests also failed on unloaded browser content or irrelevant reply wording. Runner CI could fail before model work because of dependency or sandbox setup. **Expected behavior** Deliver each settled approval once after its source run stops. Do not start work after an acknowledged Stop. Preserve the audited cancellation. Test the intended product behavior with a ready browser and verified runtime dependencies. **Steps to reproduce** 1. Approve a service request while its source run is active. Let the run finish. Check that its result starts one continuation. 2. Delay provider startup. Press Stop before its handle is available. Check cancellation, then submit `/new`. 3. Run the browser, warm-workspace, and Stop-and-redirect cases from the linked report. **Paperclip version or commit** The branch includes master at `9d19f98b5`. The report records the original source for each focused attempt. **Deployment mode** Isolated local development instances and disposable Daytona sandboxes. ## What Changed - Deliver settled tool-action results for the exact company and source run during final cleanup. Keep the existing idempotent receipt and periodic recovery sweep. - Wait for startup to hand off its provider handle before acknowledging Stop. Reject first-turn admission after cancellation. Preserve a matching audited pending or acknowledged cancellation. - Wait for mounted task history and connector controls in browser tests. Record failure evidence. Grade workspace contents and process continuity separately from exact reply wording. Require each warm-turn marker once and in order, allowing surrounding prose. - Stop-and-redirect now checks that the source file exists and work is active before Stop. - Resolve target dependency locks in an uncredentialed CI job. Verify the lock artifact hash. Keep orchestration and publication on the trusted workflow revision. - Materialize the pinned OpenCode executable and configure the exact Codex executable's user-namespace profile before provider credentials are available. - Compress Daytona directory uploads with gzip. Preserve files, executable modes, symlinks, empty directories, and confinement checks. - Classify file-transfer RPC deadlines as infrastructure. Keep unrelated runner RPC failures visible. ## Verification - [Focused live report with screenshots and original attempts](https://pages.paperclip.ing/runner-reliability-20260921/): 14 of 15 selected Product E2E cases pass across the recorded revisions. Claude and Codex Stop → `/new`, Claude service approval, delegation, both hiring/reuse cases, and native Daytona warm continuity pass. - Two credentialed Runner smoke cases pass. These are not full protocol coverage. - E2E harness after the master merge: 429 tests pass. E2E and server TypeScript checks pass. - Daytona plugin: 239 tests pass, 6 skipped. Plugin TypeScript build passes. The compression test fails against the old code and passes with the change. - Runner backend/runtime regression group: 161 tests pass. Cancellation/startup selection: 26 tests pass. Approval delivery: 34 real-database tests pass. - Workflow security: 7 tests pass. Both edited workflows pass actionlint. Runner TypeScript and Rust builds pass. - After merging master, all 389 native executor tests pass, including both native backend paths and late startup after the Stop deadline. - Post-merge `pnpm -r typecheck` and `pnpm build` pass. The monolithic local `pnpm test:run` was interrupted to integrate master and is inconclusive. The [hosted CI test partitions](https://github.com/paperclipai/paperclip/actions/runs/35620461738) pass on `50a3e43822bcba1e0d07b1b45b0be91cbf9312da`. An unchanged sandbox callback schema test initially received HTTP 503. It passed five isolated local runs, its full local test file, and one failed-job CI retry. No assertion was weakened. ## Risks - Stop can wait for the bounded startup handoff. If it cannot settle, the existing pending-recovery state remains instead of a false acknowledgement. - Immediate approval delivery must remain idempotent across cleanup and recovery sweeps. Tests cover duplicate delivery and company/run boundaries. - The workflow changes still need hosted Linux verification. They retain the trusted workflow and credential boundaries. - Gzip reduces the observed provider upload from about 1.8 GB to 663 MB. It does not yet fix the remaining Claude Daytona transfer timeout. That recovery test never reached Claude, so recovery remains unverified. Use a matching image with the verified provider package preinstalled for the next recovery test; retain cold-upload coverage separately. - The report preserves diagnostic runs with missing source metadata and marks them as such. It does not claim a new full-suite pass. - This PR adds no new prompt policy or historical status reconciliation. ## Model Used OpenAI GPT-6 through Codex performed the primary implementation and review. The exact primary backend model ID is not exposed in this session. OpenAI `gpt-5.6-luna` assisted with bounded infrastructure work and verification. The agents used repository tools, code execution, and browser tests. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 <noreply@openai.com> |
||
|
|
0f5fafe16b |
fix(runner): preserve task replies and warm process continuity (#13738)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - The Runner connects provider sessions to task state, replies, and delegated work. > - Full-stack tests found lost final replies, rejected helper calls that stopped the parent, and unnecessary process restarts. > - A completed child could also receive a new assignment wake that the scheduler then cancelled. > - This change fixes those boundaries and gives agents clearer teammate instructions. > - The tests retain strict completion and process-continuity requirements. ## Linked Issues or Issue Description **What happened?** A generated attachment comment could suppress an agent's final reply. A known Codex helper could stop its parent when it requested a Paperclip tool. Native Daytona processes restarted between turns because Paperclip minted an unused GitHub broker token. Reassigning a completed child queued a run that immediately cancelled. Revision instructions also allowed agents to do work assigned to a named teammate themselves. **Expected behavior** Keep the final reply. Reject helper tool requests without borrowing parent authority or stopping the parent. Keep an unconfigured sandbox process alive between turns. Treat assignment-only changes to completed tasks as metadata changes. Preserve explicit teammate assignments during revisions. **Steps to reproduce** Run the retained Runner E2E cases for file handoff, teammate reuse, Daytona warm continuity, and Legacy Claude interview/plan acceptance. The focused regression tests reproduce the reply, helper, process-lifetime, and assignment-wake defects without provider calls. **Paperclip version or commit** The live lifetime and completion campaign used `db3857807`. This PR replays the changes on master `c65fc9e3c`. See Verification for the limits of that evidence. **Deployment mode** Isolated local instances and native Runner sessions in Daytona sandboxes. Related work: #13546 handles a different queued-run issue after an issue-lock compare-and-set failure. This PR prevents the unnecessary assignment wake earlier. #13410 covers retained user services; this PR covers the provider process. No duplicate fix was found. ## What Changed - Exclude generated deliverable-binding comments from final-reply deduplication. Preserve the attachment and explicit user-facing replies. - Reject Paperclip tool and input requests from known Codex helper threads without terminating the parent. Keep unknown-thread rejection intact. - Explain how to hire or reuse a persistent teammate and preserve named delegation on revisions. Update generated protocol fixtures. - Use stable, token-free GitHub wrappers for unconfigured native sandboxes. Preserve credential isolation, configured-account rotation, and cleanup after partial staging failures. - Do not queue assignment-only wakes for done or cancelled tasks. Keep explicit reopening behavior. - Make warm-continuity fixtures create real review cards. Read the persisted final response selected by production presentation logic. Missing selected evidence still fails. ## Verification - Before rebase: 560 focused route, native-executor, and launcher tests passed. The new regressions were reproduced before their fixes. - Live E2E: Legacy Claude interview/plan acceptance passed 3/3 repetitions. Daytona warm continuity passed 2/3 full repetitions. Each successful run retained one process and provider session for all three turns. - The remaining Daytona repetition stopped after a same-URL browser reload left the page blank. Both completed turns retained the same process. Its failed verdict remains unchanged; this PR does not claim the blank-page cause is fixed. - Reports: https://pages.paperclip.ing/runner-e2e-lifetime-race-20260920/investigation.html and https://pages.paperclip.ing/runner-e2e-behavior-followups-20260919-results/investigation.html - Post-rebase `pnpm build` and `pnpm -r typecheck` passed. All 414 Runner E2E harness unit tests and its typecheck passed. Codex protocol tests: 88 passed, 2 ignored. - Latest-head CI: 55 successful checks and 2 intentional skips. Greptile: 5/5 with no review threads. The unchanged workspace exposure tests hit a fixed-port collision on the first CI attempt; their local suite passed (25 tests, 3 platform skips), and the CI shard passed on one retry. - The duplicate local `pnpm test:run` was stopped after the full hosted general and serialized test shards passed. It did not finish locally and is not counted as a local full-suite pass. ## Risks Configured GitHub accounts retain run-scoped credential rotation and can still restart warm processes. That limitation requires a separate design. Known provider helpers cannot use Paperclip coordination tools directly; they must return findings to the parent. The delegation prompt is an instruction, not an enforced guarantee; Codex Mini hiring/reuse failures remain open. No schema or workflow changes are included. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and parallel coding agents. The exact deployment suffix and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub issues and shared reports) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; full hosted test shards also pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
43acbcc398 |
fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task state to provider sessions. > - Follow-up turns must retain provider memory and carry new user direction. > - Lost session IDs caused repeated context and extra input tokens. > - Native question answers and approval races could leave valid work blocked. > - This pull request repairs those paths and adds regression coverage. > - Agents can continue accepted work without repeating the conversation or losing the user's answer. ## Linked Issues or Issue Description Refs #13574. That merged PR shortened continuation prompts and moved question instructions into tool documentation. This change preserves sessions and fixes failures exposed by broader testing. Related runtime work: #13408 and #13410. **What happened?** Native follow-up turns could lose the provider session ID. Completion guidance could replace the original task with its latest comment. Claude native questions could remain pending after the user answered. Approval during a running tool call could suspend the run before the tool response arrived. Onboarding and chat handoff instructions also caused repeated planning or missing plan documents. **Expected behavior** Reuse a valid provider session. Send only new events when that session already has the history. Preserve the task requirements and apply later user direction. Store the question answer and deliver it to the waiting run. Finish governed tool responses before suspending. Execute the accepted plan without asking for the same approval again. **Steps to reproduce** Run the continuation, local-session-integrity, first-task, and agent-chat suites with native Codex and Claude. Include provider-question-bridge, accept-while-running, and plan-handoff. **Paperclip version or commit** This branch is based on master |
||
|
|
84fe89906d |
fix: complete native agent review handoffs (#13581)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Native execution uses durable runs, issue locks, wake requests, and typed tool authority > - A child can finish with a native agent review request while its original assignee stays responsible for the work > - The reviewer then needs a bounded execution path that can inspect the child, record one decision, and finish safely > - Before this change, assignee-only gates rejected the reviewer or left the parent waiting after the child review ended > - This pull request adds typed reviewer admission, scoped reviewer tools, durable wake and recovery handling, and parent continuation evidence > - The benefit is that native review handoffs complete without changing child ownership or granting broad mutation access ## Linked Issues or Issue Description Refs: #13314 Refs: #13574 **What happened?** A native child run could report `needs_review` for an agent reviewer. The reviewer wake then failed assignee and execution-lock checks. The child remained in review and the parent remained waiting. **Expected behavior** The named reviewer should receive one durable wake. The reviewer should inspect the child and resolve the exact review card. The child assignee should stay unchanged. The parent should receive the recorded review outcome after the child reaches its terminal state. **Steps to reproduce** 1. Run a native task with a different named agent reviewer. 2. Keep the child assigned to its original worker. 3. Let the worker finish with a native completion review request. 4. Start the durable reviewer wake. 5. Resolve the review and finish the reviewer run. 6. Observe the child and parent state. **Paperclip version or commit** Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification source: `eea171aae`. **Deployment mode** Built from source. **Installation method** Built from source (pnpm build). **Agent adapter(s) involved** Not adapter-specific (core bug). **Access context** Both. **Database mode** Embedded PostgreSQL in the isolated live test fixtures. ## What Changed - Add server-validated native review assignment facts. - Admit only the exact company, issue, source run, decision, revision, addressee, and resolver policy. - Give reviewer runs a narrow set of Paperclip read and resolve tools. File and shell access follow the configured agent and environment policy, so reviewers can run tests. - Separate server-owned reviewer instructions from untrusted persisted review data. Escape the data boundary; retain server-enforced authorization. - Keep the child assignee unchanged. Atomically claim the reviewer run, wake request, and issue execution lock. A competing lock prevents provider startup. - Require the exact running reviewer session and current issue lock to resolve its assigned card. Reject missing, unrelated, or terminal reviewer runs. - Add durable reviewer wake, lock, stale-card, and abandoned-run recovery handling. - Prevent duplicate native wake dispatches during deferred admission and recovery. - Carry accepted or rejected child review outcomes into parent task context and continuation evidence. - Add focused server, runner, and native protocol coverage. - Preserve upstream continuation rules. Add child review decisions as separate evidence, while keeping real human answers in their own field. - Return actionable completion validation feedback to both providers. Permit a corrected completion after rejection. Keep strict terminal acknowledgment validation. - Apply exclusive shared-workspace locks to sandbox environments. Local and SSH folders can run concurrently, including when old settings request serialization. - Repair test timing, native event parsing, and the review artifact assertion. Allow a valid reject, correct, and accept review sequence. Check the accepted card against its reviewer run and decision. Keep polling within the existing deadline when review acceptance precedes the parent wake projection; report a specific missing-continuation error at timeout. - Apply the ACPX pending-call limit to reserved finish/block calls, with capacity-release and cancellation tests. ## Verification - `pnpm build`: passed on `eea171aae`. - `pnpm -r typecheck`: passed on `eea171aae`. - `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on `b31ad9ab8`; runner E2E typecheck also passed. - `pnpm check:token-gates`: passed. - Focused DB review, reviewer authority, and prompt-boundary checks: 31 tests passed. They cover invalid reviewer runs, competing locks, atomic admission, duplicate claims, and valid resolution. - Heartbeat, workspace, and recovery checks: 30 tests passed. - ACPX sidecar suite: 27 tests passed. Moving the capacity guard back below reserved handling makes both new regression cases fail. - Four focused live continuation checks passed on their first attempt at `f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and question tool guidance (12/12 each). These cases do not use the reviewer prompt path changed afterward. - Fresh Codex and Claude review-handoff checks passed all 29 native checks each on their first attempt at `eea171aae`. Both runs received the expected fixed prompt and completed cleanup. Only the six selected live flows were tested; no full paid provider catalog run. - The final commit only extracts the existing test-harness timeout diagnostic into a shared helper and adds positive and negative coverage. Removing the accepted-review guard makes two regression assertions fail; restoring it passes all six timeout tests. Production runtime code, prompts, deadlines, and grading criteria are unchanged by this final commit. - Deadline regressions: a valid continuation delayed 20 seconds succeeds within its 30-second unit-test deadline; an absent wake returns a specific candidate-failure diagnostic at that same deadline. Both assertions failed before the fix. Production E2E deadlines remain unchanged. - Historical native failures remain recorded: Docker availability failures; a valid reject/correct/accept sequence that the first-card grader misread; and a test that rejected the gap between accepted child review and parent wake projection. No failed result was regraded. The latest tests use a protected reference to the pinned Docker image and the unchanged artifact oracle and time limits. - Full repository verification runs in GitHub CI. Local verification uses the focused suites above, full build, and full typecheck. An unchanged Codex shutdown timing test failed once in CI, passed in isolation, and its full shard passed on the final commit without changes to that test or its causal code path. The original failure is retained in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no outstanding actionable findings. All review threads are resolved. All current-head CI gates passed, including the isolated native runner Docker build (55 successful checks; two skipped by the workflow). ## Risks - Reviewer admission depends on exact persisted decision and interaction bindings. A stale or changed card is rejected. - Paperclip control-plane tools are limited to inspection and review resolution. This is not a filesystem permission boundary; provider file and shell access retain the configured policy. - Deferred wake recovery changes dispatch receipt coalescing. A scheduler regression could delay a continuation if the receipt state is wrong. - Parent review outcomes are evidence for the model. They do not grant tool authority or change issue ownership. - This change does not address legacy lease-hold handoff behavior. > Roadmap review: native execution, review gates, and durable recovery are existing roadmap capabilities. This PR completes a narrow reliability path for those capabilities. ## Model Used OpenAI `gpt-6-astra` with reasoning, tool use, and code execution. OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and journal work. Context window size is not exposed by this session. Live test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the PR authors. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e26d787928 |
Shorten continuation prompts and verify question tool guidance (#13574)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents must continue tasks using user answers without losing earlier requirements or approval gates. > - The wake prompt mixed human decisions with prior tool evidence and repeated detailed question instructions. > - Those instructions belong with the question tool, with a short routing hint in the wake. > - The Runner evals need to prove that answers, approvals, and completed work survive later turns. > - This PR shortens the prompts, separates authenticated answers, and adds continuation tests with useful screenshots. ## Linked Issues or Issue Description Refs #13517. This is a follow-up to the merged onboarding skill and Runner E2E work. Related #13539 covers responses received while a run is active; this PR preserves its cases and adds continuation coverage. Existing continuation/recovery and question PRs were searched; none covers this prompt/documentation and eval change. **What existing behavior does this improve?** The instructions sent when an agent continues a task, the native human-input tool documentation, and the evidence captured by Runner full-stack E2E. **Current behavior** The wake repeats a long question-tool guide. Human answers appear alongside untrusted prior results. Screenshot capture can finish at DOM load while the task still shows a spinner, even when backend behavior checks pass. **Proposed behavior** Keep earlier requirements unless the user changes them. Treat clarification as distinct from approval. Give authenticated human responses a scoped field. Keep tool and agent results as evidence. Put detailed question behavior in the tool descriptor and retain one routing sentence in the native wake. Wait for the correct task and loaded conversation before taking screenshots. **Reason and benefit** Reduce repeated prompt text and make authority boundaries clear. Test that real question cards, later answers, approval gates, and completed child tasks still work. Make screenshots useful for human review. ## What Changed - Shorten shared continuation instructions for legacy and native runners. Separate authenticated user responses from tool results and agent summaries. - Remove the detailed question guide from native wake prompts. Keep its behavior in the canonical `request_human_input` descriptor and existing payload schema. Regenerate semantic contracts and fixture hashes. - Add five continuation cases across four local profiles. Add a dedicated choice-then-text case for native Codex and native Claude. All 22 cells join the shared full E2E campaign. - Cover revised scope, clarification without approval, hostile instructions in a handoff file, and reuse of a completed child after restart. Keep production instructions and fixed user facts. - Capture continuation screenshots only when the intended task and conversation have rendered. Add provider-free browser regressions for loaders and wrong-task capture. - Preserve current master’s extra tool and onboarding cases. The default campaign now contains 166 cells; 35 manual everyday cells remain separate. ## Verification - `pnpm -r typecheck`: passed after replay on current master. - `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed. - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm test:e2e:runner:browser-support`: 4 passed. These tests failed against immediate screenshot capture and passed after the fix. - Focused continuation and native-input tests: 36 passed locally. The tool-authority suite could not initialize embedded PostgreSQL locally, including one isolated retry; its 17 assertions did not run locally. The full remote server shards passed on this PR commit. - `pnpm build`: passed after replay on current master. `pnpm test:run` was attempted locally but hit the same embedded PostgreSQL initialization failure; the remaining local run was stopped after complete remote CI passed. This is not claimed as a full local test pass. - [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685): passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All server/chat/workspace/serialized shards, browser shards, Runner checks, typecheck, build, canary and policy checks passed. The isolated native Runner build and security checks also passed: 57 successful checks, with two expected Storybook skips. - Greptile reviewed the exact PR head at 5/5, with no findings or unresolved review threads. The PR has no merge conflicts. - [Live question-docs report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/): 3/3 passed at source `83dd132f2` before replay on master. Native Codex and Claude each asked a choice, waited, asked a text question, and saved both answers. Claude also passed a completed-child restart case. All three native turns are checked for absence of the old question block. - [Earlier continuation report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/): all five continuation cases passed on native Claude. The report retains campaign and revision provenance and separately shows two unresolved onboarding behavior failures. - [Before/after prompt report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/): full text, current recorded Claude inputs, and reproducible reference-token counts. The controlled wake comparison removes 401 reference tokens; the net counted input reduction is 339 after charging the larger tool description. These are text-size estimates, not measured billing savings. ## Risks - Prompt wording affects model behavior. Live results cover the stated cases, not every provider or conversation. Legacy profiles are registered but were not rerun for this change. - The optional continuation field changes prompt data only; there is no database migration or new production API. - Authenticated answer projection excludes generated summaries and agent-resolved interactions. It preserves the answer’s question or approval scope. - The screenshot guard can expose UI loading failures that earlier runs hid. Backend grading alone no longer makes those captures valid. - The two prior onboarding failures remain separate product issues: work before acceptance and a missing saved plan. This PR does not claim the entire onboarding suite passes. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, code execution, and browser verification. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted tests above; the full local database-startup limit is documented - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
11921075a4 |
Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first task helps a new user define and approve useful work. > - That workflow needs reusable instructions and tests against the production experience. > - Native Codex and Claude must load the assigned skill, including after resume. > - Maintainers need recorded conversations and precise failed checks to judge regressions. > - This pull request adds the first-task skill and a suite in the shared Runner E2E harness. > - It keeps behavior results separate from informational quality scores and incomplete recordings. ## Linked Issues or Issue Description **What existing behavior does this improve?** The first onboarding task and the Runner E2E report used to review it. **Current behavior** Onboarding embeds its policy in a hidden brief. Native Codex drops the skill-instructions setting at the Rust boundary. The shared E2E harness has no onboarding suite or full conversation view. **Proposed behavior** Assign and invoke `/first-task` for the onboarding task. Send selected Codex skills as structured protocol inputs. Run twelve scenarios across legacy Codex, legacy Claude, native Codex, and native ACPX Claude. Include all 48 cells in full campaigns. Show recorded chat, question and approval cards, exact checks, instructions, and billing in the shared dashboard. **Reason and benefit** Measure the real onboarding experience before changing prompts. Distinguish infrastructure failures, behavior failures, and unexercised journey steps. **Breaking changes** No database migration or production API change. First-task instructions now live in an assigned skill. The user-edited persona is preserved; the skill includes the maintainer-approved proposal-mode mapping and saved-plan requirement. Related: #11043 is earlier onboarding work. #13422 already fixes native Claude model pinning, context delivery, and read permissions on master; this branch includes those fixes through its base. The new Claude recovery test supplements them. ## What Changed - Extract and assign the first-task skill while retaining the production greeting and opening question. - Carry the Codex skill-instructions flag through thread start and resume. Resolve explicit task skill references only against assigned skills and send native skill inputs. - Invoke an unambiguously selected assigned skill through Claude ACPX’s native slash-command parser on initial and resumed turns, retaining the entire task/wake envelope as its argument. Do not carry that invocation into ordinary tasks. - Restore the saved single-task proposal modes: confirmation card, or saved plan with revision-targeted checkbox approval. Explicit plan requests also require a saved plan. - Add first-response and complete-journey cases with fixed user facts, acceptance checkpoints, durable outcome checks, and accounting for child runs. - Fail the eval when choice questions have fewer than two real options. Recognize planning documents without treating them as completed work. - Add optional, bounded quality judging as explicit post-processing. - Render full conversations and static interaction cards in the shared report. Conversations start folded. Show original and regraded results and incomplete journeys distinctly. - Keep credential-persistence scanning outside the first-task behavioral suite; retain public evidence redaction. - Refresh generated capability references after the API-reference edits. - Correct shared native question guidance and tool schemas: choices need at least two meaningful options; open-ended questions use canonical text fields with the required compatibility payload. Verify both formats through real tool-authority persistence. - Disable announcements automatically for every isolated Runner E2E process and label the gallery environment/provider/target explicitly. - Remove CI races in the GitHub connection browser test and native session recovery test by waiting for the actual async work before asserting its results. ## Verification - `pnpm exec vitest run server/src/services/onboarding-first-task-assets.test.ts server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19 passed. - `pnpm --dir packages/paperclip-runner exec vitest run src/drivers/acpx/runtime-host.test.ts src/drivers/acpx/native-skill-prompt.test.ts src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command forwarding and the 1 MiB input boundary both failed before their fixes and passed afterward. Coverage includes changed skills on reopen, approval context, and an ordinary subsequent task. - Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64 first-task fixture and grader tests also pass. - Full repository typecheck and build passed locally. Server typecheck and Runner build passed again after the native-command change. - Full GitHub Actions CI passed on `23e56447b`: all server/workspace/browser shards, Runner verification, typecheck/release registry, build, canary, policy, and Docker checks. Greptile reviewed this exact head at 5/5 with no unresolved threads. The earlier broad local run had database startup/timing failures that passed isolated retries; the complete remote suite is green. - Merge verification against current master: 312 harness tests and 13 native recovery tests passed. Regenerated semantic contracts and fixture hashes pass their consistency check. Full local typecheck and build also passed on the stacked queue branch. After merging the latest master and preserving the GitHub setup timing regression in the split browser suite, both focused GitHub browser tests passed. Three CI timing/startup flakes passed local verification and one remote retry; all latest-head checks are green. - Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local mock API confirmed that `/skill-name` expands the assigned skill body before the model request and retains the task arguments. A prose mention does not. The probes made no paid model calls. The ACP probe used the current first-task skill body and retained the wake arguments. - [Full 48-case campaign and report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/): 44 passed after three interrupted Codex cases completed in targeted reruns. Original results, regrades, and all 51 executions remain in the report provenance. - [Claude campaign after the shared-question fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/): 10/12 passed with zero single-option failures. All 12 recorded the current assigned skill and corrected guidance. The failures exposed skipped skill invocation and a missing saved plan. This PR adds native command invocation and explicit saved-plan instructions; the subsequent report below still shows behavior failures. - [Fresh 12-case Claude report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/) at `78452129e`: 10/12 pass after correcting two false proposal-matcher failures. The recordings said “Here is the task I will create and run/complete” in approval cards; the old matcher missed that word order. Regression tests failed before the fix and pass after it. Original results and offline regrade provenance remain linked. No agent rerun was needed. Zero single-option-question failures; two behavior failures remain: direct work before acceptance on a plain first message, and an explicit plan request without a saved plan. Neither check was relaxed. The follow-up `82087ac7e` fixes command-prefix size accounting; `94aefb1f3` fixes only that proposal matcher. - Report browser checks confirm folded conversations, rendered cards, explicit Local/Daytona labels, and no page errors. The published-object audit scanned 1,306 text files across 2,154 objects with no credential-format findings or prohibited files. Image pixels and unknown token formats are outside that scan. ## Risks - Model behavior is nondeterministic. One campaign is evidence, not a guarantee. The two remaining Claude behavior failures are visible in the report and require further product work; this PR does not claim all onboarding scenarios pass. - The suite checks persisted Paperclip effects. It cannot prove the absence of arbitrary external effects. - Historical recordings can miss later journey steps. These remain incomplete, never passes. - Native profiles switch runtime after the production onboarding wizard because it does not yet expose a native option. - Quality scores are informational and cannot override behavioral failures. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, and code execution. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cceeb0aa66 |
test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e200104727 |
feat: review connection actions from tasks (#13063)
Bring governed connection reviews into task history and composer approvals. Share resolution with Connections, add scoped remembered permissions, and resume agents through durable outcome receipts. Keep cards compact, collapse raw results, isolate untrusted provider output, bound continuation payloads, and reconcile missed live events. Add Storybook coverage, browser journeys, and service regression tests. Verification: all PR CI gates passed, Greptile 5/5, security scans passed, five connection-review browser journeys passed, and real native Codex approval/continuation was verified against the local MCP fixture. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
05735b3d87 |
fix(ui): show continuation actions in confirmation receipts (#12939)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task interactions let an operator approve completion or ask an agent to continue. > - The server stores the continue choice as a rejected completion request so it can resume the same task. > - The task feed ignored the configured action label and showed the generic text `Declined request`. > - The generic text made a successful three-turn continuation look like a failed request. > - This pull request keeps the server state and shows the action that the operator selected. > - The benefit is an accurate task feed for legacy Codex and Runner Codex. ## Linked Issues or Issue Description **What happened?** A rejected confirmation always appeared as `Declined request`. The feed did not use a custom rejection action such as `Continue work`. **Expected behavior** The resolved receipt and success toast must show the selected custom action. Confirmations without a custom action must keep the current fallback text. **Steps to reproduce** 1. Create a completion confirmation with `rejectLabel` set to `Continue work`. 2. Select `Continue work` and enter a continuation note. 3. Open the completed task feed. 4. Observe that the old UI says `Declined request` instead of the selected action. **Paperclip version or commit** `539c9212f4b98e37643f5a8e3b603f1b5845b5d7` **Deployment mode** GitHub Actions Runner E2E with warm Daytona sandboxes. ## What Changed - Show `Selected “Continue work”` when a rejected confirmation has that custom action label. - Use the same action-aware text in the success toast. - Keep `Declined request` as the fallback for confirmations without a custom rejection label. - Make the legacy warm-turn prompt ask if the task is ready to complete. - Require both warm Daytona matrix cells to show two continuation receipts and no generic decline receipt. ## Verification - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx ui/src/lib/issue-thread-interactions.test.ts ui/src/components/task-chat/TaskChatInteractionCard.test.tsx` - `pnpm test:e2e:runner:unit` - `pnpm test:e2e:runner:typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/ui build` - `pnpm check:token-gates` - Paid `daytona-warm-continuity` campaign: both three-turn cells passed on attempt 1 ([workflow run](https://github.com/paperclipai/paperclip/actions/runs/34052047946)); the downloaded evidence aggregates locally as 2/2. The trusted report job did not publish because default branch `pnpm-lock.yaml` was transiently behind its manifest. - `pnpm typecheck` reached an unrelated `plugin-workspace-diff` dependency type error after the current master manifest and lockfile resolved different versions. The focused UI and runner checks pass. ## Risks - Low risk. The database state and continuation behavior do not change. - A custom rejection label now appears in resolved receipts and success toasts. - The paid warm Daytona suite has a stricter browser assertion. ## Model Used - OpenAI Codex, GPT-5, with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in this PR with the bug template fields - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal task identifier - [x] I have run focused tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated the relevant E2E fixture and assertions - [x] I have considered and documented the risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0c1e7504c0 |
fix(runner): persist warm Daytona workspaces across turns (#12904)
## Thinking Path > - Daytona preserves a stopped sandbox filesystem, but deleting or replacing a sandbox removes its only remote copy. > - Warm reuse therefore improves latency but cannot be Paperclip's durability boundary. > - The host execution workspace must remain authoritative after every successful turn, while same-run recovery must avoid overwriting unexported remote work. > - Result proposal, workspace export/merge, and terminal completion need a durable, replayable ordering so a crash never starts a duplicate provider turn. > - A paid browser acceptance suite must exercise both legacy Codex and Runner Codex for three real turns on one continuously warm Daytona sandbox. ## Linked Issues or Issue Description Refs #12901. Runner Codex did not previously export successful Daytona workspace changes back to the authoritative host workspace. That made warm reuse depend on Daytona's remote filesystem and left deleted/replacement sandboxes without a reliable reconstruction path. The existing paid fixture also lacked a focused three-turn continuity case for both Codex adapters. ## What Changed - Persist versioned, atomic native workspace-sync descriptors and durable seeds in `PAPERCLIP_HOME`, without credentials or a database migration. - Classify fresh, warm, replacement, and same-run-recovery workspace preparation explicitly; ambiguous lease/root/digest evidence fails closed. - Finalize native workspace export/merge after semantic result proposal and before run completion, with idempotent replay that never submits a second provider turn. - Surface legacy Codex workspace restoration failures instead of masking them, while preserving an earlier provider error when both fail. - Keep healthy reusable Daytona leases warm for legacy and native adapters, stamp finalized workspace generations, and retain existing cleanup behavior for per-turn or unhealthy leases. - Preserve Runner Codex's provider process/session across warm turns, including bounded post-terminal tail draining and exact authority rotation. - Add the exact paid `daytona-warm-continuity` matrix: - `legacy-codex × daytona × warm-three-turn` - `runner-codex × daytona × warm-three-turn` - Drive all three turns through the browser, verify ordered file continuity and stable lease/workspace/runtime identities, capture per-turn timings, and delete the sandbox immediately after assertions. - Document `pnpm test:e2e:runner -- --suite daytona-warm-continuity`; no package script was added. ## Verification - `pnpm typecheck` — passed, including migration safety (no migration added) - Focused server/runner Vitest coverage — 144 passed - `pnpm test:e2e:runner:unit` — 114 passed - `pnpm test:e2e:runner:typecheck` — passed - `pnpm --filter @paperclipai/paperclip-runner test:codex` — 66 passed, 1 helper ignored - `native-session-executor.test.ts` — 139 passed, including safe fail-closed cleanup after remote runner identity capture failure - Paid local browser acceptance, exact post-rebase Linux/amd64 runner binary: - Runner Codex — passed in 1.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - Legacy Codex — passed in 2.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - [Protected paid GitHub Actions campaign](https://github.com/paperclipai/paperclip/actions/runs/34026735033) against `7da42a91b95fa7fb2df126668ef7e37afb3b2b9d` — passed 2/2: - Runner Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Legacy Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Merge/enforcement, S3 history, and Pages publication jobs passed - Paid result artifacts were scanned for both provider credentials; neither secret was present. - Current PR checks — 31 passed, 1 expected Storybook skip; Greptile 5/5; Superagent security scan passed - `git diff --check origin/master...HEAD` — passed - Confirmed no `package.json`, lockfile, migration, or SQL changes. ## Risks - Workspace synchronization now sits on the terminal-success path, so a remote export failure deliberately prevents false success. Retryable state retains its lease/seed; loss of the only unexported remote copy fails closed. - Warm provider reuse has strict identity and quiescence checks. Mismatched or ambiguous evidence blocks reuse rather than risking concurrent provider work. - The paid suite incurs Daytona and Codex cost only in the existing protected scheduled/manual workflow and explicitly destroys its sandbox after each cell. ## Model Used OpenAI Codex with GPT-5 agentic reasoning, repository inspection, real browser E2E execution, Rust/TypeScript test execution, and GitHub Actions diagnostics. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing issue or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the dedicated suite invocation without adding a package script - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green on the current revision - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the current revision - [x] I will address all reviewer comments before requesting merge |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
87832c48fd |
feat(runner-e2e): publish declared screenshots (#12895)
## Thinking Path > - Paperclip uses runner end-to-end reports to compare agent profiles and execution environments > - The report dashboard shows each reviewed final-state screenshot as a thumbnail and gallery item > - The public history publisher removed all per-attempt images before it regenerated the dashboard > - Therefore the public dashboard had the new layout but could not show the screenshots from the run > - The publisher needs a narrow rule that keeps only screenshots from the exact live fixture issue route > - This pull request keeps those trusted PNG files in every future public S3 and GitHub Pages report > - The benefit is that each future report can show its screenshot gallery without exposing logs, traces, videos, archives, arbitrary images, or generated report trees ## Linked Issues or Issue Description **What happened?** The runner E2E job captured final-state screenshots in its private artifact. The public S3 and GitHub Pages publication step removed those screenshots before it regenerated the dashboard. As a result, the public report showed the new dashboard controls but no screenshot thumbnails or gallery items. **Expected behavior** Each future public runner E2E report must include reviewed PNG screenshots from the live fixture issue. Other captures and active or unsafe evidence must stay private. **Steps to reproduce** 1. Run the runner full-stack E2E workflow on `master` before this change. 2. Open the private `runner-e2e-report-*` artifact and confirm that it contains per-attempt PNG screenshots. 3. Open the public campaign URL and confirm that the dashboard has no screenshot gallery items. **Paperclip version or commit** The issue was reproduced on commit `64d8929`, after the report design change in PR #12889. **Deployment mode** GitHub Actions with the public S3 and GitHub Pages report publishers. Related design work: Refs #12889. ## What Changed - Mark screenshots from the exact server-created live fixture issue route with `public-runner-fixture`. - Keep marked PNG files in both the S3 history bundle and the GitHub Pages bundle. - Keep captures from other issue routes, sensitive routes, and external origins private. - Bind public files to the normalized execution ID, attempt, and safe PNG base name. - Validate every retained image with the existing PNG signature and 12 MiB size checks. - Skip missing-artifact sentinel results with attempt `0` when they have no public screenshots. - Continue to remove unmarked images, videos, traces, archives, generated HTML reports, and other private evidence. - Update publisher tests, workflow checks, report copy, and the public-evidence security documentation. ## Verification - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts tests/runner-e2e/history.test.ts tests/runner-e2e/report.test.ts` - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts tests/runner-e2e/workflow-security.test.ts -t "uses environment-scoped OIDC"` - `pnpm test:e2e:runner:typecheck` - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/runner-e2e-dashboard.spec.ts` - `pnpm -r typecheck` - `pnpm build` - Regenerated the dashboard from retained evidence for Actions run `33968240659` without a paid matrix rerun. The public-stage proof contained 121 screenshot gallery items and thumbnail frames, with zero generated HTML report files. The trusted-fixture marker and route gate have separate focused tests. - All pull request CI checks pass on commit `ccd2b49e1`. ## Risks - This change intentionally makes marked fixture screenshots public at the campaign URL. A screenshot can show data that a raw-byte secret scan cannot detect. - The capture helper marks a screenshot only on the exact loopback issue route for the fixture that the harness created. A different issue, sensitive page, or external origin stays private. - The publisher also requires the marker, a safe normalized path, a valid PNG signature, and the size limit. - The change does not publish videos, logs, traces, archives, arbitrary images, or generated browser report trees. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5.6-sol`, with high reasoning, repository tool use, shell execution, browser inspection, and GitHub CLI access. The working context was the Codex desktop task context. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bcc6fe7a44 |
fix(runner): restore multi-turn remote sessions (#12840)
## Thinking Path > - Paperclip manages AI agents and their work. > - The runner executes agent turns on local and remote providers. > - A remote per-turn session must save its state before Paperclip releases its sandbox. > - The session runtime returned after 100 milliseconds while the remote checkpoint still ran. > - The next turn also checked the local state path instead of the verified remote backup. > - This pull request waits for the bounded remote close and accepts only a verified suspended backup. > - The benefit is reliable multi-turn execution without weaker identity checks. ## Linked Issues or Issue Description **What happened?** A successful remote agent turn released its sandbox before the runner saved the verified continuation backup. The next turn failed with `runner_state_identity_mismatch`. **Expected behavior** Paperclip must finish the bounded remote checkpoint before it releases the sandbox. A later turn must validate and restore the digest-matched suspended backup. **Steps to reproduce** 1. Run a native ACPX Claude Plan test in a non-reusable Daytona sandbox. 2. Reject the first plan to start a second turn. 3. Observe that the second turn fails before provider execution. **Paperclip version or commit** The failure reproduced at `13775a90b078ff64872f50961ea1b83d575e7bc6`. **Deployment mode** GitHub Actions with a Daytona sandbox. ## What Changed - Wait for the internally bounded remote runner close and checkpoint before the host returns. - Preserve the existing short cleanup bound for other providers. - Validate remote continuation lifecycle from a complete digest-verified backup when local runner state is absent. - Keep corrupt, non-suspended, mismatched, and unverified state fail-closed. - Make native Plan completion and accepted-Plan wake prompts deterministic. ## Verification - A prior 45-cell local campaign passed 44 cells. The only failure was the OpenCode Plan prompt variance fixed here. - A focused OpenCode local Plan rerun passed. - ACPX Claude Daytona message and question cells passed. - Focused regressions cover delayed checkpoint close and verified remote backup lifecycle. - GitHub Build and the focused ACPX Claude Daytona Plan cell will validate this exact head. ## Risks Remote runnerd sessions now wait for their internally bounded close/checkpoint path before returning; generic provider cleanup retains the existing 100 millisecond bound. Durable run success still cannot be reversed. The environment release guard still blocks sandbox destruction when no verified backup stamp exists. ## Model Used OpenAI Codex, GPT-5.6, extended reasoning, with code execution and GitHub Actions inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal task id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open findings - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b84964e5a2 |
fix(runner): stabilize local paid E2E recovery (#12836)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paid runner E2E tests verify the complete runner, control-plane, and UI path. > - A server restart could load a fresh task page while Playwright still waited on an unsettled Vite navigation lifecycle. > - The current server also ignored the isolated Vite cache path and skipped Vite's per-request HTML transform from the known-green runner snapshot. > - A one-cell paid run then exposed that download-artifact v8 removes the artifact-name directory for one pattern match. > - This pull request restores the Vite contract, proves a fresh document after restart, and accepts only the exact singleton artifact layout. > - The benefit is reliable local runner qualification without weaker UI, source, or artifact checks. ## Linked Issues or Issue Description Refs #12769 Refs #12828 Refs #12829 Refs #12833 **What happened?** The structured-question restart test could time out after the replacement server returned the task route and rendered the durable pending interaction. A focused one-cell rerun passed the paid test but failed aggregation because download-artifact v8 flattened its single artifact. **Expected behavior** The test must prove that a new document loaded after the server restart and that the same pending interaction survived. The aggregate must accept the exact documented singleton download layout while it continues to reject ambiguous or foreign artifacts. **Steps to reproduce** 1. Run the local ACPX-Codex structured-question restart-resume cell. 2. Restart the isolated server while the question waits for an answer. 3. Observe that the route and task UI can reload before Playwright settles the navigation promise. 4. Run a paid campaign with one selected cell. 5. Observe download-artifact v8 extract the sole campaign directory directly into the requested path. **Paperclip version or commit** The local campaign reproduced the navigation failure at `3586956a1b794b3cb4a9c5f57ffb7355e2b0c46d`. The one-cell aggregate reproduced the singleton layout at `f487660c0a06ba06ca140b57386f21ed39f13120`. This fix is `de4ccceff453a4b39436bf9a2eb8f03924151af7`. **Deployment mode** Local development and paid GitHub Actions. **Installation method** Built from source. **Agent adapter(s) involved** ACPX-Codex. The Vite and aggregate fixes are provider-neutral. ## What Changed - Prove a new post-restart browser document with an in-memory sentinel. - Tolerate only Playwright's navigation timeout before the exact UI and API checks run. - Honor `PAPERCLIP_VITE_CACHE_DIR` in the embedded Vite server. - Limit dependency optimization to the real UI entry. - Run `vite.transformIndexHtml` for each request while caching only the branded source template. - Accept download-artifact v8's flattened layout only for one expected cell with one unique recognized campaign. - Keep source SHA, source ref, workflow URL, execution ID, attempt, and unexpected-entry validation. - Add focused positive and negative regressions for Vite rendering and singleton artifact selection. ## Verification - Exact 45-cell local campaign https://github.com/paperclipai/paperclip/actions/runs/33888939013 passed 44/45. Its only failure was the post-restart navigation false negative fixed here. - Exact focused rerun https://github.com/paperclipai/paperclip/actions/runs/33891207957 passed the ACPX-Codex restart cell first attempt with the same session, two durable runs, the terminal marker once, and cleanup complete. - The focused Vite renderer suite passed 2/2 tests. - The focused rerun-artifact selector suite passed 12/12 tests. - Prettier and `git diff --check` passed. - An exact-head 45-cell confirmation is pending. ## Risks Low to medium risk. The Vite change restores known-green per-request transforms and isolated cache behavior. It can affect all development UI loads. The paid matrix and ordinary CI will verify that behavior. The singleton selector remains fail-closed for ambiguous layouts and validates every result source. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, and parallel focused agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
da1e403022 |
test(runner): harden native OpenCode paid fixtures (#12833)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paid runner E2E tests verify the full control-plane path for supported providers. > - Native OpenCode could write the reserved terminal marker through progress and final output. > - The restart fixture also waited for all development assets after the recovered UI was already usable. > - These behaviors made two valid local runner paths fail qualification. > - This pull request makes the OpenCode write contract explicit and uses the visible UI as the restart readiness gate. > - The benefit is reliable local OpenCode qualification without weaker duplicate detection. ## Linked Issues or Issue Description Refs #12769 Refs #12829 Refs #12828 **What happened?** The native OpenCode ask fixture allowed a progress tool call before the final response. OpenCode could write the reserved terminal marker in both places. The structured restart fixture could also time out while it waited for `DOMContentLoaded` after the recovered UI was visible and usable. **Expected behavior** The ask fixture must write the reserved marker once. The restart fixture must continue when the recovered UI and interaction API prove that the application is ready. **Steps to reproduce** 1. Run the local native OpenCode ask-question paid cell. 2. Observe a run that calls `report_progress`, calls `paperclip_finish`, and then emits the exact marker. 3. Run the local native OpenCode structured-question restart-resume cell with a fresh Vite graph. 4. Observe that the page is usable before the navigation lifecycle event completes. **Paperclip version or commit** The failures reproduced at `06cdf88bd9ac0fad82588025d23a68e810b20fd0`. The fixes are at `f7e044e71df11a0582eafe28d2fd52ea7cd07948`. **Deployment mode** Local dev. **Installation method** Built from source. **Agent adapter(s) involved** OpenCode through the native runner. ## What Changed - Require `paperclip_finish` to be the only tool call in the native ask fixture. - Forbid `report_progress` and other tool calls in that fixture. - Wait for navigation commit after a server restart. - Keep the explicit recovered UI and interaction API readiness checks. - Add prompt contract assertions. ## Verification - The exact two-cell paid run passed both affected cells on the AWS runner fleet: https://github.com/paperclipai/paperclip/actions/runs/33883334853 - Native OpenCode ask-question passed in job https://github.com/paperclipai/paperclip/actions/runs/33883334853/job/101058952662 - Native OpenCode structured restart-resume passed in job https://github.com/paperclipai/paperclip/actions/runs/33883334853/job/101058952823 - Prettier passed for all three changed files. - `git diff --check` passed. - The run-level aggregate failed only because the workflow source still used the pre-repair lockfile on `master`. PR #12828 repairs that lockfile. ## Risks Low risk. The prompt change affects native ask fixtures across provider profiles. The navigation change remains guarded by explicit UI and API assertions. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
af3023f1e3 |
fix(runner): repair paid provider startup paths (#12769)
## Thinking Path > - Paperclip manages AI agents that perform work. > - Paperclip Runner connects durable task runs to local provider processes. > - The full-stack paid matrix exposed failures after the runner integrity repair. > - Verified JavaScript entrypoints lost their relative module graph when Linux executed them through descriptor paths. > - Returned provider startup errors also remained pending and became indeterminate after recovery. > - Sparse Codex tool lifecycle events lost the `write_document` identity before task transcript projection. > - This pull request repairs those three boundaries and makes the structured-question fixture deterministic. > - The benefit is repeatable provider startup, exact failure replay, and correct inline Plan placement. ## Linked Issues or Issue Description Refs #12721 and #12700. **What happened?** The paid runner matrix failed ACPX and OpenCode startup before provider session creation. The runner journal then replaced the original startup error with an indeterminate recovery result. Native Codex saved a Plan but rendered it only as a fallback card. A legacy Claude waiting reply could also echo the reserved terminal marker before the answer arrived. **Expected behavior** Verified JavaScript providers must start from immutable descriptor-backed artifacts. Returned startup failures must persist as terminal failed command results. Native tool lifecycle updates must preserve the `write_document` boundary. Pre-answer fixture output must not contain the reserved terminal marker. **Steps to reproduce** 1. Run the local provider cells in the Runner Full-Stack E2E workflow. 2. Observe ACPX and OpenCode fail during `session.open` before provider execution. 3. Observe recovery report `execution_indeterminate` instead of the original startup error. 4. Run the native Codex Plan cell and observe the fallback Plan card after the tool activity row. 5. Run the legacy Claude structured-question resume cell and observe an early marker echo in waiting prose. **Paperclip version or commit** `0f9452101740835ce0b1488a204bf48acd5bafc3` **Deployment mode** Local development with the paid GitHub Actions acceptance workflow. ## What Changed - Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM entrypoints before hashing and verified descriptor launch. - Anchor ACPX dynamic provider package resolution at a controller-derived provider-pack root and keep that root out of the provider child environment. - Persist executor-returned startup errors as redacted durable failed command results while retaining indeterminate recovery for true process death. - Coalesce sparse native tool items by stable ID so a late `write_document` name, input, and result reach the transcript boundary once. - Forbid the structured-question fixture from spelling or announcing its reserved terminal marker before the user answers. ## Verification - Rust and TypeScript regression tests cover durable failed replay, true crash ambiguity, bundle closure, package-root derivation, environment filtering, exact Codex tool lifecycle coalescing, and prompt determinism. - Local execution is intentionally limited to formatters and static diff checks. GitHub Actions will run tests, type checks, builds, and security checks. - After ordinary CI is green, scoped paid cells will validate one ACPX launch, one OpenCode launch, native Codex Plan projection, and legacy Claude structured resume before a complete matrix rerun. - Prior failing matrix: https://github.com/paperclipai/paperclip/actions/runs/33682434315 ## Risks - Bundling changes the bytes covered by provider launch hashes. Provider-pack generation already hashes the final built files. - ACPX still loads qualified provider packages dynamically. The controller supplies a normalized package root, while existing version, digest, path, and descriptor checks remain active. - Durable `failed` is terminal. Replays return the same redacted result and do not execute the provider effect twice. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5 with agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions coordination. The exact deployed snapshot and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked related public work or described the bug in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] No documentation change is required for this runtime repair - [x] I have considered and documented the risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0f94521017 |
fix(runner): restore local session and task integrity (#12721)
## Thinking Path > - Paperclip is the control plane for agents that perform work. > - Paperclip Runner connects durable provider sessions to individual task runs through PRP. > - Provider continuity and per-run authority are different lifetimes. > - The existing implementation mixed those lifetimes and lost event metadata between provider frames, runnerd, persistence, API sanitization, and the task thread. > - That caused failed continuation, missing progress and Plans, duplicate replies, hidden failures, and unsafe recovery. > - This repair gives every heartbeat fresh authority, preserves qualified provider-session continuity, and restores one lossless presentation path without changing direct adapters. ## Linked Issues or Issue Description **What happened?** A second native heartbeat could reuse tickets, leases, command receipts, sequence state, and run identity from the first heartbeat. Provider phase and item identity could be lost before the UI read them. Redaction could corrupt protocol discriminators while still missing malformed credential tails. The task thread could fold progress into the final response, hide failures, or show more than one final answer. Native Codex also exposed approval modes that do not yet have a durable approval bridge. **Expected behavior** Each heartbeat uses a new PRP authority epoch. Codex and OpenCode preserve exact qualified provider sessions; ACPX emits an explicit continuity event when its qualified process-replacement policy is used. Every accepted provider event is presented, classified as internal, or surfaced as unsupported. The task page shows chronological progress, reasoning summaries, activity, Plans, interactions, terminal failures, and exactly one final reply. Direct adapters retain their existing path. **Steps to reproduce** 1. Enable the unified experimental Paperclip Runner setting. 2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex agent. 3. Run response, Plan, structured-question/resume, restart, cancellation, and failure scenarios. 4. Reload the task while active, waiting, failed, and settled. 5. On the old implementation, observe stale run authority, missing classifications, incomplete output, or duplicated/folded replies. **Paperclip version or commit** The repair is based directly on `master` at `87d05e194b643810d16d20612115acd01d735d43`. **Deployment mode** Local development with the embedded database. Related work: Refs #12616, #12646, #12666, #12685, and #12700. ## What Changed - Rotates PRP control-plane, outbox, ticket, lease, command, receipt, and sequence authority for each heartbeat while carrying forward only a validated provider-session identity. - Reads `control-plane-state.json`, validates both durable schemas and lifecycle values, resumes coherent current runs, archives qualified settled authority, and quarantines malformed or mismatched scoped state without moving ambiguous live legacy state. - Preserves Codex provider phase and stable item identities so commentary remains progress and only `final_answer` becomes final. - Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning lifecycle mapping. - Makes ACPX normalization lossless for visible reasoning, tool lifecycle metadata, stable bounded identities, Plan revisions, structured requests, failures, and qualified process replacement. Only the compatible terminal assistant message is promoted as final. - Applies schema-aware redaction before generic JWT-shaped detection and scans every diagnostic string leaf. Malformed raw/escaped quoted credential tails are redacted in both server and durable Rust state. - Restores snapshot-style chronological task presentation, expandable tool activity, inline Plan cards, visible waiting/resume/cancel/failure states, and exactly one final answer. - Makes `never` the only qualified native Codex permission mode and rejects unsupported persisted native modes with remediation. OpenCode and ACPX policies remain intact. - Keeps the unified experimental Runner setting as the only enablement flag. Onboarding and direct Codex, Claude, and OpenCode stay on their legacy execution/finalization paths. - Adds cross-language goldens, authority/recovery/fault coverage, exact response/count assertions, and native plus legacy acceptance scenarios. ## Verification - Pull-request GitHub Actions run Rust formatting/tests, TypeScript checks, server/UI tests, builds, protocol drift checks, browser E2E, and security scans. - A separate workflow-only validation ref is pinned directly on this PR head and runs the 35-cell paid local matrix: three core scenarios plus structured-question resume and restart/resume for native Codex, native OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode. Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315 - Acceptance requires exact single visible replies, monotonic sequences, matching envelope discriminators, one semantic terminal, one run terminal, no unresolved interaction, no duplicate mutation, no secret leakage, provider continuity, and zero native rows for direct adapters. - Per maintainer direction, tests are running in GitHub Actions rather than on the slower local host. Only formatters and static diff checks were run locally. ## Risks - Recovery from old or partial filesystem state is sensitive. The repair fails closed, preserves active or unverifiable authority, and quarantines only state whose scoped ownership is safe to move. - Provider event formats can change. Closed validators and boundary goldens turn new or malformed events into visible diagnostics instead of silent drops. - Shared task presentation could affect direct adapters. Runtime-fact gating plus the direct-adapter matrix protect the existing path. - Managed and remote providers are not qualified here. Shared code continues to compile and fail safely, but live qualification is deferred. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact deployed snapshot and context-window size are not exposed to this task. It used agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] The paid local-provider matrix is green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5716fe907e |
test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner subsystem executes agent work across local and managed provider backends. > - The lower pull requests restore the task runtime, provider backends, and managed-provider control plane. > - The restored system needs repeatable full-stack checks before it can ship safely. > - Paid live checks also need clear access, cost, and secret controls. > - This pull request adds acceptance, live evaluation, chaos, and release gates for the restored runner stack. > - The benefit is measurable runner parity with safer release decisions. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change covers runner tests, release workflows, server contracts, and evaluation tools. **Problem or motivation** The runner stack did not have one complete acceptance surface for native Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could miss provider drift, task-view regressions, cost-policy errors, and destructive cleanup errors. **Proposed solution** Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid workflows. Add live evaluation, chaos, cost-limit, redaction, and release contract checks. Add AWS AgentCore infrastructure and guarded provisioning tools. Keep the native runner experimental flag off by default. **Alternatives considered** We considered manual smoke tests only. They do not give repeatable evidence and they do not protect release branches. We also considered one large pull request. The stacked pull requests keep each review below the Greptile file limit. **Roadmap alignment** This work supports the shipped Cloud / Sandbox agents milestone and the shipped Agent evals & feedback milestone in `ROADMAP.md`. Related stack: - #12699 adds managed provider backends and lifecycle support. - #12691 adds qualified OpenCode and ACPX provider backends. - #12685 restores task runtime rendering and steering. ## What Changed - Add the runner full-stack harness with 57 catalog cells and 60 unit tests. - Add a Daytona runner image with digest-pinned base images and base-aware image-content checks. - Add guarded live evaluation and chaos workflows with a fixed 40-execution matrix; live and full-stack paid schedules now run only on Sundays or by manual dispatch. - Add in-flight reported-usage cost stops, post-turn cost caps, exact-threshold failure classification, secret redaction, retry classification, and actor authorization. - Reattach stream and hard-budget listeners before restart-recovery continuations so restored paid sessions cannot bypass in-flight interruption. - Preserve OpenCode usage and cost across tool-loop messages and turns while exposing an explicit current-run delta to durable accounting. - Keep PNG/WebM evidence in access-controlled artifacts only, reject SVG, and publish only pruned inert structured per-attempt evidence. - Add AWS AgentCore infrastructure, provisioning checks, and smoke tools; reject unsafe model identifiers, require exact stack ownership markers, and make failed-stack replacement explicit. - Add evaluation-session contracts and capability reports. - Add release workflow checks for immutable action pins, frozen dependency installs, exact weekly cron shape, paid-run guards, provider-secret isolation, and chaos test paths. - Reauthorize the original and triggering numeric actor IDs as the first step of every provider-secret job, including partial reruns, before checkout or provider access. - Give each full-stack matrix cell only its matching provider credential, expose Daytona only to Daytona cells, and disable shared dependency caches anywhere paid credentials or OIDC write access are present. - Protect the legacy manual E2E workflow with the same default-branch, allowlist, environment, and per-job authorization boundary. - Rotate live-eval candidates by week and retain 120 days of compatible history so the seven-week trend window remains viable. - Restore the root runner-acceptance commands and reconcile reported snapshots, raw receipts, and terminal usage without double counting or losing late usage. - Mark ACPX token deltas exact only when every budget field is present, keep cumulative cost/request authority separate, reject non-USD cost labeling, and include thought tokens in output-token budgets. - Keep `enableNativeRunner` off by default. The acceptance harness enables it only in its isolated test instance. ## Verification Passed locally: - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm test:runner-acceptance:typecheck` - `pnpm test:runner-acceptance` (19 tests) - focused OpenCode proxy, driver, runnerd transport, live-session, and turn-stream tests (106 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/clean-room-server.test.ts` (22 tests) - `pnpm test:e2e:runner:typecheck` - `pnpm test:e2e:runner:unit` (62 tests) - `node --test scripts/__tests__/release-verify-workflow.test.mjs` - `pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals` (22 tests) - `pnpm -r typecheck` - `pnpm build` - `node --test packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs` (6 tests) - `git diff --check` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core --lib --locked` (161 tests) - focused ACPX provider-event tests (10 tests) - The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged. I did not run paid live provider jobs or provision AWS resources. Those checks need credentials and can create cost. ## Risks The paid workflows can create provider cost. They require an allowlisted original and triggering actor, the protected `runner-e2e-paid` environment, explicit opt-in variables, and cost limits. The four provider credentials exist only in that master-only environment, which requires allowlisted reviewer approval and disables administrator bypass; repository and organization Actions scopes contain no copies. Provider usage arrives after a billable request, so the live guard cannot prevent one request from crossing a threshold. It interrupts immediately on the first reported threshold hit and permits no continuation. Visual evidence can contain secrets rendered as pixels. PNG/WebM remain only in access-controlled workflow artifacts; SVG and per-attempt XML are excluded, and S3/Pages receive a pruned structured dashboard. The AWS scripts can create cloud resources. They use explicit commands, least-privilege roles, KMS encryption, saved nonsecret metadata, and explicit teardown. This pull request does not enable the experimental native runner for existing instances. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The model used extended reasoning, tool use, code execution, and parallel subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |