mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip lets people manage agents through ongoing conversations. > - Chat users can change instructions while a provider is already working. > - Existing chat evals wait for each turn to settle before the next message. > - They cannot prove delivery during active work or the saved effect of a correction. > - Existing fixtures also enable Agent Chat through the API rather than the settings UI. > - This PR adds bounded browser workflows and checks their persisted outcomes. ## Linked Issues or Issue Description Refs #13741, #13752, #13750. **What happened?** The chat suites cover planning, delegation, status, and recovery. They lack active-turn follow-ups and the experimental settings lifecycle. A sequential conversation can pass even if messages sent during work are lost. **Expected behavior** A follow-up submitted during a provider turn survives and affects the final reply. A changed launch day appears in the saved plan. Disabling Agent Chat rejects new messages while preserving history; re-enabling resumes the same conversation. **Steps to reproduce** Run the explicit `agent-chat-stories` suite. It selects three local cases for each native Claude and Codex profile. An ordinary provider command waits for a fixture brief file so the browser can send the follow-up at an observed active-run boundary. ## What Changed - Add six opt-in Product E2E cells for settings, active follow-ups, and plan corrections. - Drive experimental settings through the UI and verify disabled sends are rejected by the public API. - Use a bounded file wait in the actual isolated agent workspace, with provider-written readiness and an undisclosed brief reference. - Grade persisted user messages, final replies, native run outcomes, and exact saved plan fields. - Accept active-turn steering or one queued successor; reject lost input, duplicate input, and stale outputs. - Allow one steered run or two sequential runs throughout the shared harness, while preserving exact counts for other cases. - Require a single marker-bearing response attributed to the final provider run. - Unload the development browser client before restarting the server, avoiding reconnect/navigation races without weakening the post-restart memory check. - Add browser regressions for restart isolation and asynchronously saved settings switches. - Document prepared-agent setup, native onboarding limits, and the separate API-tool rollout gate. ## Verification - Eval TypeScript check passed. - Eval support suite: 436 tests passed in 39 files. - New oracle calibration: six tests passed, including plausible invalid outcomes. - Browser support regressions: seven tests passed; the restart regression was observed failing before the fix. - Catalog discovery selects exactly six local native cases and leaves default paid selection unchanged. - [Consolidated existing native chat report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/): master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a browser navigation timeout across restart; the page request returned 200 and the chat rendered. - [Nine targeted restart/replay cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/) passed on `1fe2fe275`, including the original failure, across native Claude/Codex and local/Daytona; all cleanup passed. - [Initial six-story campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817) retained all six failures: asynchronous switch assertions, unavailable fixture paths, and rich-text escaping in raw command comparisons. The corrected fixtures preserve the same behavioral assertions. - [Six-story campaign v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/) on `8232773a0`: 4/6 passed (both settings cases and both Claude interruptions). Codex could not see the host-temp fixture outside its workspace; this failed before follow-up delivery was exercised. - [Four affected interruption cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/) all passed, including cleanup, on definition v3 / `ad6ac0545`. Files live inside the actual agent workspace and the observed run workspace is verified. Both providers saved Friday in the real plan with the undisclosed brief reference; follow-ups persisted while the original run was active. Together with both unchanged settings cases from v2, all six new scenario variants have passing live evidence. - Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful checks, two intentional skips, zero pending/failing checks; mergeable and clean. Fresh Greptile 5/5, zero unresolved findings. - Full typecheck, tests, build, and browser CI passed remotely. One earlier head encountered a signoff-policy browser timing failure; the final head passed that shard. - Local pnpm wrapper could not fetch its version/signature metadata in the restricted environment; local eval checks used the installed Node executables. Repo-wide validation was completed by GitHub Actions. ## Risks These are eval-only changes. The file wait is a timing fixture in the isolated agent workspace, not a production runner hook. Native Codex host-filesystem isolation stays unchanged. It has a two-minute limit and is released in `finally`. The prepared-agent settings case is not full native onboarding: the wizard currently offers legacy adapters. The disabled-entry assertion uses full document navigation, which clears the prior React Query cache; preserved history is checked through the public API and re-enabled chat. No production prompt, rollout default, adapter behavior, or credential policy changes. Active-task reassignment and worker-crash recovery remain outside these new cases. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
9 lines
421 B
TypeScript
9 lines
421 B
TypeScript
import { expect, type Locator } from "@playwright/test";
|
|
|
|
/** Controlled settings switches update after their asynchronous save handler. */
|
|
export async function setSettingsToggle(toggle: Locator, enabled: boolean) {
|
|
await expect(toggle).toBeVisible();
|
|
if (await toggle.getAttribute("aria-checked") !== String(enabled)) await toggle.click();
|
|
await expect(toggle).toHaveAttribute("aria-checked", String(enabled));
|
|
}
|