mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
docs/connector-launch-kit-inputs
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5842185e4f |
fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat needs reliable native execution before native runners become the onboarding default. > - Existing stories covered idle reassignment and controller restart, but not an executing worker handoff or worker process loss. > - Status answer tests also need to reject stale claims and invented facts. > - This pull request adds six opt-in full-stack cells with independent state assertions and retained evidence. > - The probes exposed a misleading Retry across server projection and recovery-banner paths; the fix reports the blocked recovery honestly. > - The tests preserve failures without changing recovery policy, production prompts, or onboarding defaults. ## Linked Issues or Issue Description Refs: #13762. Related: #13765 (Retry targets the latest failed attempt), #13753 (task context ownership), #13746 (native recovery work). ## What Changed - Add active reassignment with saved draft and plan preservation, old-worker cancellation, and successor completion checks. - Preserve recovery-needed projection when native cleanup fails before its coordinator exists, refuse a generic retry that would immediately fail again, and replace the recovery banner's misleading Retry with Inspect run. - Add verified local worker process loss with a required successful continuation; retain a failing qualification result when recovery is unavailable, while independently verifying the UI/API refuse doomed retries. - Add two-turn factual answer checks for current blockers, stale claims, inactive backlog work, and unknown facts. Retain prose for separate semantic review. - Add positive and negative oracle calibration and document fault isolation, cleanup, billing, and qualification limits. ## Verification - Eval TypeScript check passes. - All 442 eval support tests pass locally. The 89 focused server tests and server typecheck pass. Six recovery-banner UI tests and token gates pass. - Initial new-cell campaign: https://github.com/paperclipai/paperclip/actions/runs/35657128077. All six results are retained; four failed on fixture-contract issues and two exposed real worker cleanup quarantine. - All 26 existing native onboarding cells: https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26 passed on master |
||
|
|
846336e5a0 |
test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path > - Paperclip lets people manage agents through ongoing conversations. > - Chat users can change instructions while a provider is already working. > - Existing chat evals wait for each turn to settle before the next message. > - They cannot prove delivery during active work or the saved effect of a correction. > - Existing fixtures also enable Agent Chat through the API rather than the settings UI. > - This PR adds bounded browser workflows and checks their persisted outcomes. ## Linked Issues or Issue Description Refs #13741, #13752, #13750. **What happened?** The chat suites cover planning, delegation, status, and recovery. They lack active-turn follow-ups and the experimental settings lifecycle. A sequential conversation can pass even if messages sent during work are lost. **Expected behavior** A follow-up submitted during a provider turn survives and affects the final reply. A changed launch day appears in the saved plan. Disabling Agent Chat rejects new messages while preserving history; re-enabling resumes the same conversation. **Steps to reproduce** Run the explicit `agent-chat-stories` suite. It selects three local cases for each native Claude and Codex profile. An ordinary provider command waits for a fixture brief file so the browser can send the follow-up at an observed active-run boundary. ## What Changed - Add six opt-in Product E2E cells for settings, active follow-ups, and plan corrections. - Drive experimental settings through the UI and verify disabled sends are rejected by the public API. - Use a bounded file wait in the actual isolated agent workspace, with provider-written readiness and an undisclosed brief reference. - Grade persisted user messages, final replies, native run outcomes, and exact saved plan fields. - Accept active-turn steering or one queued successor; reject lost input, duplicate input, and stale outputs. - Allow one steered run or two sequential runs throughout the shared harness, while preserving exact counts for other cases. - Require a single marker-bearing response attributed to the final provider run. - Unload the development browser client before restarting the server, avoiding reconnect/navigation races without weakening the post-restart memory check. - Add browser regressions for restart isolation and asynchronously saved settings switches. - Document prepared-agent setup, native onboarding limits, and the separate API-tool rollout gate. ## Verification - Eval TypeScript check passed. - Eval support suite: 436 tests passed in 39 files. - New oracle calibration: six tests passed, including plausible invalid outcomes. - Browser support regressions: seven tests passed; the restart regression was observed failing before the fix. - Catalog discovery selects exactly six local native cases and leaves default paid selection unchanged. - [Consolidated existing native chat report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/): master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a browser navigation timeout across restart; the page request returned 200 and the chat rendered. - [Nine targeted restart/replay cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/) passed on `1fe2fe275`, including the original failure, across native Claude/Codex and local/Daytona; all cleanup passed. - [Initial six-story campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817) retained all six failures: asynchronous switch assertions, unavailable fixture paths, and rich-text escaping in raw command comparisons. The corrected fixtures preserve the same behavioral assertions. - [Six-story campaign v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/) on `8232773a0`: 4/6 passed (both settings cases and both Claude interruptions). Codex could not see the host-temp fixture outside its workspace; this failed before follow-up delivery was exercised. - [Four affected interruption cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/) all passed, including cleanup, on definition v3 / `ad6ac0545`. Files live inside the actual agent workspace and the observed run workspace is verified. Both providers saved Friday in the real plan with the undisclosed brief reference; follow-ups persisted while the original run was active. Together with both unchanged settings cases from v2, all six new scenario variants have passing live evidence. - Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful checks, two intentional skips, zero pending/failing checks; mergeable and clean. Fresh Greptile 5/5, zero unresolved findings. - Full typecheck, tests, build, and browser CI passed remotely. One earlier head encountered a signoff-policy browser timing failure; the final head passed that shard. - Local pnpm wrapper could not fetch its version/signature metadata in the restricted environment; local eval checks used the installed Node executables. Repo-wide validation was completed by GitHub Actions. ## Risks These are eval-only changes. The file wait is a timing fixture in the isolated agent workspace, not a production runner hook. Native Codex host-filesystem isolation stays unchanged. It has a two-minute limit and is released in `finally`. The prepared-agent settings case is not full native onboarding: the wizard currently offers legacy adapters. The disabled-entry assertion uses full document navigation, which clears the prior React Query cache; preserved history is checked through the public API and re-enabled chat. No production prompt, rollout default, adapter behavior, or credential policy changes. Active-task reassignment and worker-crash recovery remain outside these new cases. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3790ca2f13 |
fix(runner): repair approval and Stop races and eval infrastructure (#13750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Runner tasks must continue after approval and stop when the user presses Stop. > - Live evals found races at approval delivery and provider startup. > - Browser readiness and CI setup errors also hid the actual task results. > - This pull request fixes those races and the related test infrastructure. > - Regression tests and saved live reports show which cases now pass. ## Linked Issues or Issue Description Companion eval definitions PR: https://github.com/paperclipai/paperclip-evals/pull/25 (AgentCore paused and provider/environment infrastructure). Related: #13741 now supplies the late-startup Stop fence and warm-attachment recovery; this PR retains that fence and extends startup tracking and regression coverage to both native backend paths. #13539 introduced queued approvals during active runs. #13738 fixes child assignment, task replies, and warm process continuity and is already in the base. #13291 concerns automatic continuation of interrupted legacy sandbox runs; this PR fixes native startup cancellation and does not change that recovery policy. **What happened?** An accepted service approval could wait after its source run stopped. Stop could return success before the provider handle existed. Work could then start after Stop, or a cancelled run could be recorded as failed. Some E2E tests also failed on unloaded browser content or irrelevant reply wording. Runner CI could fail before model work because of dependency or sandbox setup. **Expected behavior** Deliver each settled approval once after its source run stops. Do not start work after an acknowledged Stop. Preserve the audited cancellation. Test the intended product behavior with a ready browser and verified runtime dependencies. **Steps to reproduce** 1. Approve a service request while its source run is active. Let the run finish. Check that its result starts one continuation. 2. Delay provider startup. Press Stop before its handle is available. Check cancellation, then submit `/new`. 3. Run the browser, warm-workspace, and Stop-and-redirect cases from the linked report. **Paperclip version or commit** The branch includes master at `9d19f98b5`. The report records the original source for each focused attempt. **Deployment mode** Isolated local development instances and disposable Daytona sandboxes. ## What Changed - Deliver settled tool-action results for the exact company and source run during final cleanup. Keep the existing idempotent receipt and periodic recovery sweep. - Wait for startup to hand off its provider handle before acknowledging Stop. Reject first-turn admission after cancellation. Preserve a matching audited pending or acknowledged cancellation. - Wait for mounted task history and connector controls in browser tests. Record failure evidence. Grade workspace contents and process continuity separately from exact reply wording. Require each warm-turn marker once and in order, allowing surrounding prose. - Stop-and-redirect now checks that the source file exists and work is active before Stop. - Resolve target dependency locks in an uncredentialed CI job. Verify the lock artifact hash. Keep orchestration and publication on the trusted workflow revision. - Materialize the pinned OpenCode executable and configure the exact Codex executable's user-namespace profile before provider credentials are available. - Compress Daytona directory uploads with gzip. Preserve files, executable modes, symlinks, empty directories, and confinement checks. - Classify file-transfer RPC deadlines as infrastructure. Keep unrelated runner RPC failures visible. ## Verification - [Focused live report with screenshots and original attempts](https://pages.paperclip.ing/runner-reliability-20260921/): 14 of 15 selected Product E2E cases pass across the recorded revisions. Claude and Codex Stop → `/new`, Claude service approval, delegation, both hiring/reuse cases, and native Daytona warm continuity pass. - Two credentialed Runner smoke cases pass. These are not full protocol coverage. - E2E harness after the master merge: 429 tests pass. E2E and server TypeScript checks pass. - Daytona plugin: 239 tests pass, 6 skipped. Plugin TypeScript build passes. The compression test fails against the old code and passes with the change. - Runner backend/runtime regression group: 161 tests pass. Cancellation/startup selection: 26 tests pass. Approval delivery: 34 real-database tests pass. - Workflow security: 7 tests pass. Both edited workflows pass actionlint. Runner TypeScript and Rust builds pass. - After merging master, all 389 native executor tests pass, including both native backend paths and late startup after the Stop deadline. - Post-merge `pnpm -r typecheck` and `pnpm build` pass. The monolithic local `pnpm test:run` was interrupted to integrate master and is inconclusive. The [hosted CI test partitions](https://github.com/paperclipai/paperclip/actions/runs/35620461738) pass on `50a3e43822bcba1e0d07b1b45b0be91cbf9312da`. An unchanged sandbox callback schema test initially received HTTP 503. It passed five isolated local runs, its full local test file, and one failed-job CI retry. No assertion was weakened. ## Risks - Stop can wait for the bounded startup handoff. If it cannot settle, the existing pending-recovery state remains instead of a false acknowledgement. - Immediate approval delivery must remain idempotent across cleanup and recovery sweeps. Tests cover duplicate delivery and company/run boundaries. - The workflow changes still need hosted Linux verification. They retain the trusted workflow and credential boundaries. - Gzip reduces the observed provider upload from about 1.8 GB to 663 MB. It does not yet fix the remaining Claude Daytona transfer timeout. That recovery test never reached Claude, so recovery remains unverified. Use a matching image with the verified provider package preinstalled for the next recovery test; retain cold-upload coverage separately. - The report preserves diagnostic runs with missing source metadata and marks them as such. It does not claim a new full-suite pass. - This PR adds no new prompt policy or historical status reconciliation. ## Model Used OpenAI GPT-6 through Codex performed the primary implementation and review. The exact primary backend model ID is not exposed in this session. OpenAI `gpt-5.6-luna` assisted with bounded infrastructure work and verification. The agents used repository tools, code execution, and browser tests. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 <noreply@openai.com> |
||
|
|
9d19f98b50 |
fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent chat uses native runner sessions to plan, delegate, and track that work. > - A user can press Stop while the native session is still starting. > - The server can acknowledge that Stop without dispatching it, then let the session submit a turn. > - This leaves chat recovery waiting for an execution that the user expected to stop. > - This PR waits for the startup handle, dispatches cancellation, and prevents a late startup from submitting a turn. > - New full-stack evals check the resulting records and outputs across Claude and Codex. > - Those evals also exposed missing ACPX readiness fields, unbounded polling, and an old-run identity check that rejected valid warm handoffs. ## Linked Issues or Issue Description **What happened?** Stop during native startup could record an acknowledged cancellation with `dispatched: false`. The provider could then begin work. A subsequent `/new` stayed queued. A remote Claude follow-up also exhausted the command journal while probing warm-session readiness: ACPX never returned the readiness fields required by the shared transport. Once readiness worked, attachment incorrectly compared the next run descriptor against the old run ID. The 25 ms polling loop could issue 4,800 commands during its two-minute wait, beyond the 500-command bound. The existing chat eval treated lifecycle logs as proof of an active provider turn, so it did not distinguish startup cancellation from active-turn cancellation. **Expected behavior** A Stop during startup must reach the pending session. A late session must not submit a prompt after Stop. Recovery must retain control when startup exceeds the bounded wait. Chat evals must check saved task state, document contents, worker identity, account binding, and duplicate effects. **Steps to reproduce** 1. Start a native Claude or Codex chat turn. 2. Press Stop after process startup is requested but before the provider turn starts. 3. Send `/new`, then send a fresh message. 4. On the affected base, cancellation can be acknowledged without dispatch and the reset stays queued. **Paperclip version or commit** The live Claude baseline reproduced this on `29d6b3509`. The branch also includes master commit `0f5fafe16`. Related work: #13678, #13686, #13693, #13291, #13738. A separate runner reliability branch also contains a startup-wait fix. Its overlap must be reconciled before merging; this branch additionally prevents prompt submission after a late startup. ## What Changed - Wait for a pending native startup before acknowledging a run-scoped Stop. Preserve the existing recovery error when that wait expires. - Keep a Stop guard on startup. Cancel a late handle before it can submit a provider turn. - Add regression tests for normal handle publication and publication after the Stop deadline. - Back off blocked warm-attachment probes. Keep the fast two-snapshot barrier, fail closed, and record changed blockers. - Add red/green tests for delayed readiness, persistent blockers, alternating readiness, and readiness near the deadline. - Publish ACPX readiness and blockers. Preserve the old authority’s event acknowledgement barrier; only settled sessions can proceed to attachment. - Bind warm ACPX descriptors to the validated next authority while retaining old-run event correlation until activation. Preserve session identity and provider profile checks. - Exercise two consecutive run rotations through a qualified fake sidecar, verifying checkpointing, provider identity, pre-activation rejection, and new-run work admission. - Separate startup and active-turn cancellation checkpoints in the browser eval. - Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells across Claude and Codex. - Cover hiring and reuse through managed AI accounts, source-based review, current blocked-task status, request replay after a lost HTTP acknowledgement, server restart continuity, and Stop/reset continuity. - Use ordinary production agent instructions. Enable API tools only for the two coordination cases that need them. - Calibrate the matchers with invalid records and outputs. Require remembered context after restart and a structured status snapshot that distinguishes the current blocker from history and task status from active execution. Compare the public issue mutation contract and relationships during read-only reporting. Preserve before/after source records in failed eval evidence. - Fix the lost-ack browser harness and verify it against a real HTTP server. Check the chat composer after restart instead of waiting for an unrelated document lifecycle event. - Document the scope and limits of each case. ## Verification - The startup regression failed on the unfixed executor and passed after the fix. - `pnpm test:e2e:runner:typecheck` passed. - `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files. - `pnpm exec vitest run server/src/services/native-runtime/native-session-executor.test.ts` passed: 385 tests. - [Baseline live campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868): Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed. Codex delegation was blocked by provider capacity. - [Eval-only startup campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786): both providers failed as expected. Both persisted `dispatched: false` and left `/new` queued. - [First fixed campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706) on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both providers. Failed cases exposed eval harness defects and remote continuity failures. All attempts remain available. - [Original workflows and stronger memory checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649) on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and startup Stop passed for both providers; Codex remote restart passed. Claude remote restart exposed the missing readiness contract. Two Codex planning cells hit provider capacity. - [Unchanged-model retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548): Codex planning and backlog creation both passed. - [18-cell campaign with ACPX readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963) on `6a98ef743`: 16/18 passed, including all local/remote Stop and committed-send cases. Claude remote continuity exposed the next-authority check, now fixed. Codex hiring produced its checklist, but the runner redacted the requested marker after it appeared as “Tracking token: …”. That content-redaction policy is unchanged and remains an explicit limitation. - [Structured status grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725) on `50448c228`: both providers passed on their first attempt, including cleanup. - [Complete read-only state grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011) on `551e13892`: both providers passed. - [Final ACPX handoff and hiring retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456) on `cbd637587`: all three Claude Daytona cases passed (restart continuity, active Stop/reset, and lost-ack replay). Codex hiring reproduced the content-redaction failure: the saved checklist contained `Tracking token: [REDACTED]` instead of the required business marker. All four cases completed cleanup successfully. [Published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/). The only subsequent commit adds the qualified-sidecar integration test; production code is identical to this live proof. - `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without paid models. - Runner TypeScript typecheck passed. All 5 warm-readiness tests pass; two failed with the prior fixed-rate loop, and the late-readiness test failed before the pacing correction. - ACPX readiness and warm-identity regressions each failed before their fixes. All 292 runner-core Rust library tests passed. The qualified-sidecar integration test passes. Rust formatting is checked. - Status-grader regressions for misleading historical mentions and previously unchecked mutations each failed before tightening the oracle and pass now. - [Latest-head CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307) passed on `a4093c8f1`: full build, type checks, test partitions, browser E2E, and native runner checks. Two unrelated tests initially failed (Sentry fixture release attribution and local-service fixture readiness); both passed locally together (35 passed, 5 optional SDK tests skipped) and on the failed-job retry. No changes were made to those tests. - Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed and all review threads are resolved. - The paid live suite is not fully green: the reproducible content-redaction case remains red. This is separate from the passing PR merge checks. No production content-redaction, prompt, model, or completion-policy change is included. - Managed-account hiring and review cases explicitly enable API tools; these do not qualify default new-user onboarding. ## Risks - Stop can wait up to 30 seconds for startup, then use the existing pending-recovery path. This does not prove that remote cleanup has finished. - Blocked warm readiness adds up to 750 ms between later probes with the two-minute remote budget, or about 32 ms with the default five-second budget. Ready sessions retain the short second barrier. - Paid evals can fail because of provider capacity or agent decisions. Each failure needs evidence-based classification. - The HTTP request replay case checks comment idempotency and duplicate effects. It does not prove replay safety for an ambiguous provider tool call. - The new suite is opt-in. It does not increase the default paid campaign. - No production prompts or model selection change. Review-handoff behavior and content-redaction policy remain separate product decisions. The latter can remove harmless business content that looks like credential syntax; the failing attempt is retained. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f5fafe16b |
fix(runner): preserve task replies and warm process continuity (#13738)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - The Runner connects provider sessions to task state, replies, and delegated work. > - Full-stack tests found lost final replies, rejected helper calls that stopped the parent, and unnecessary process restarts. > - A completed child could also receive a new assignment wake that the scheduler then cancelled. > - This change fixes those boundaries and gives agents clearer teammate instructions. > - The tests retain strict completion and process-continuity requirements. ## Linked Issues or Issue Description **What happened?** A generated attachment comment could suppress an agent's final reply. A known Codex helper could stop its parent when it requested a Paperclip tool. Native Daytona processes restarted between turns because Paperclip minted an unused GitHub broker token. Reassigning a completed child queued a run that immediately cancelled. Revision instructions also allowed agents to do work assigned to a named teammate themselves. **Expected behavior** Keep the final reply. Reject helper tool requests without borrowing parent authority or stopping the parent. Keep an unconfigured sandbox process alive between turns. Treat assignment-only changes to completed tasks as metadata changes. Preserve explicit teammate assignments during revisions. **Steps to reproduce** Run the retained Runner E2E cases for file handoff, teammate reuse, Daytona warm continuity, and Legacy Claude interview/plan acceptance. The focused regression tests reproduce the reply, helper, process-lifetime, and assignment-wake defects without provider calls. **Paperclip version or commit** The live lifetime and completion campaign used `db3857807`. This PR replays the changes on master `c65fc9e3c`. See Verification for the limits of that evidence. **Deployment mode** Isolated local instances and native Runner sessions in Daytona sandboxes. Related work: #13546 handles a different queued-run issue after an issue-lock compare-and-set failure. This PR prevents the unnecessary assignment wake earlier. #13410 covers retained user services; this PR covers the provider process. No duplicate fix was found. ## What Changed - Exclude generated deliverable-binding comments from final-reply deduplication. Preserve the attachment and explicit user-facing replies. - Reject Paperclip tool and input requests from known Codex helper threads without terminating the parent. Keep unknown-thread rejection intact. - Explain how to hire or reuse a persistent teammate and preserve named delegation on revisions. Update generated protocol fixtures. - Use stable, token-free GitHub wrappers for unconfigured native sandboxes. Preserve credential isolation, configured-account rotation, and cleanup after partial staging failures. - Do not queue assignment-only wakes for done or cancelled tasks. Keep explicit reopening behavior. - Make warm-continuity fixtures create real review cards. Read the persisted final response selected by production presentation logic. Missing selected evidence still fails. ## Verification - Before rebase: 560 focused route, native-executor, and launcher tests passed. The new regressions were reproduced before their fixes. - Live E2E: Legacy Claude interview/plan acceptance passed 3/3 repetitions. Daytona warm continuity passed 2/3 full repetitions. Each successful run retained one process and provider session for all three turns. - The remaining Daytona repetition stopped after a same-URL browser reload left the page blank. Both completed turns retained the same process. Its failed verdict remains unchanged; this PR does not claim the blank-page cause is fixed. - Reports: https://pages.paperclip.ing/runner-e2e-lifetime-race-20260920/investigation.html and https://pages.paperclip.ing/runner-e2e-behavior-followups-20260919-results/investigation.html - Post-rebase `pnpm build` and `pnpm -r typecheck` passed. All 414 Runner E2E harness unit tests and its typecheck passed. Codex protocol tests: 88 passed, 2 ignored. - Latest-head CI: 55 successful checks and 2 intentional skips. Greptile: 5/5 with no review threads. The unchanged workspace exposure tests hit a fixed-port collision on the first CI attempt; their local suite passed (25 tests, 3 platform skips), and the CI shard passed on one retry. - The duplicate local `pnpm test:run` was stopped after the full hosted general and serialized test shards passed. It did not finish locally and is not counted as a local full-suite pass. ## Risks Configured GitHub accounts retain run-scoped credential rotation and can still restart warm processes. That limitation requires a separate design. Known provider helpers cannot use Paperclip coordination tools directly; they must return findings to the parent. The delegation prompt is an instruction, not an enforced guarantee; Codex Mini hiring/reuse failures remain open. No schema or workflow changes are included. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and parallel coding agents. The exact deployment suffix and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub issues and shared reports) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; full hosted test shards also pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5716fe907e |
test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner subsystem executes agent work across local and managed provider backends. > - The lower pull requests restore the task runtime, provider backends, and managed-provider control plane. > - The restored system needs repeatable full-stack checks before it can ship safely. > - Paid live checks also need clear access, cost, and secret controls. > - This pull request adds acceptance, live evaluation, chaos, and release gates for the restored runner stack. > - The benefit is measurable runner parity with safer release decisions. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change covers runner tests, release workflows, server contracts, and evaluation tools. **Problem or motivation** The runner stack did not have one complete acceptance surface for native Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could miss provider drift, task-view regressions, cost-policy errors, and destructive cleanup errors. **Proposed solution** Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid workflows. Add live evaluation, chaos, cost-limit, redaction, and release contract checks. Add AWS AgentCore infrastructure and guarded provisioning tools. Keep the native runner experimental flag off by default. **Alternatives considered** We considered manual smoke tests only. They do not give repeatable evidence and they do not protect release branches. We also considered one large pull request. The stacked pull requests keep each review below the Greptile file limit. **Roadmap alignment** This work supports the shipped Cloud / Sandbox agents milestone and the shipped Agent evals & feedback milestone in `ROADMAP.md`. Related stack: - #12699 adds managed provider backends and lifecycle support. - #12691 adds qualified OpenCode and ACPX provider backends. - #12685 restores task runtime rendering and steering. ## What Changed - Add the runner full-stack harness with 57 catalog cells and 60 unit tests. - Add a Daytona runner image with digest-pinned base images and base-aware image-content checks. - Add guarded live evaluation and chaos workflows with a fixed 40-execution matrix; live and full-stack paid schedules now run only on Sundays or by manual dispatch. - Add in-flight reported-usage cost stops, post-turn cost caps, exact-threshold failure classification, secret redaction, retry classification, and actor authorization. - Reattach stream and hard-budget listeners before restart-recovery continuations so restored paid sessions cannot bypass in-flight interruption. - Preserve OpenCode usage and cost across tool-loop messages and turns while exposing an explicit current-run delta to durable accounting. - Keep PNG/WebM evidence in access-controlled artifacts only, reject SVG, and publish only pruned inert structured per-attempt evidence. - Add AWS AgentCore infrastructure, provisioning checks, and smoke tools; reject unsafe model identifiers, require exact stack ownership markers, and make failed-stack replacement explicit. - Add evaluation-session contracts and capability reports. - Add release workflow checks for immutable action pins, frozen dependency installs, exact weekly cron shape, paid-run guards, provider-secret isolation, and chaos test paths. - Reauthorize the original and triggering numeric actor IDs as the first step of every provider-secret job, including partial reruns, before checkout or provider access. - Give each full-stack matrix cell only its matching provider credential, expose Daytona only to Daytona cells, and disable shared dependency caches anywhere paid credentials or OIDC write access are present. - Protect the legacy manual E2E workflow with the same default-branch, allowlist, environment, and per-job authorization boundary. - Rotate live-eval candidates by week and retain 120 days of compatible history so the seven-week trend window remains viable. - Restore the root runner-acceptance commands and reconcile reported snapshots, raw receipts, and terminal usage without double counting or losing late usage. - Mark ACPX token deltas exact only when every budget field is present, keep cumulative cost/request authority separate, reject non-USD cost labeling, and include thought tokens in output-token budgets. - Keep `enableNativeRunner` off by default. The acceptance harness enables it only in its isolated test instance. ## Verification Passed locally: - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm test:runner-acceptance:typecheck` - `pnpm test:runner-acceptance` (19 tests) - focused OpenCode proxy, driver, runnerd transport, live-session, and turn-stream tests (106 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/clean-room-server.test.ts` (22 tests) - `pnpm test:e2e:runner:typecheck` - `pnpm test:e2e:runner:unit` (62 tests) - `node --test scripts/__tests__/release-verify-workflow.test.mjs` - `pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals` (22 tests) - `pnpm -r typecheck` - `pnpm build` - `node --test packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs` (6 tests) - `git diff --check` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core --lib --locked` (161 tests) - focused ACPX provider-event tests (10 tests) - The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged. I did not run paid live provider jobs or provision AWS resources. Those checks need credentials and can create cost. ## Risks The paid workflows can create provider cost. They require an allowlisted original and triggering actor, the protected `runner-e2e-paid` environment, explicit opt-in variables, and cost limits. The four provider credentials exist only in that master-only environment, which requires allowlisted reviewer approval and disables administrator bypass; repository and organization Actions scopes contain no copies. Provider usage arrives after a billable request, so the live guard cannot prevent one request from crossing a threshold. It interrupts immediately on the first reported threshold hit and permits no continuation. Visual evidence can contain secrets rendered as pixels. PNG/WebM remain only in access-controlled workflow artifacts; SVG and per-attempt XML are excluded, and S3/Pages receive a pruned structured dashboard. The AWS scripts can create cloud resources. They use explicit commands, least-privilege roles, KMS encryption, saved nonsecret metadata, and explicit teardown. This pull request does not enable the experimental native runner for existing instances. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The model used extended reasoning, tool use, code execution, and parallel subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |