mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 00:54:38 +02:00
codex/private-task-storage
2076
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6d6c7331cb |
Update private task storage with master migrations
Preserve the reviewed privacy rules and existing PR history. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
799e4d556f |
fix: make accounting durable and synchronize cost reporting (#14997)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
99a9de9940 |
fix(mcp): personalize assistant connections and hide revoked grants (#15411)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Assistant connections let people use their organization from another assistant. > - Setup begins with an invitation and returns to a list of connected assistants. > - Client-only labels obscure the authorizing person, while revoked rows clutter that list. > - This pull request shows the person’s avatar, names connections by owner and client, and hides revoked rows. > - It also shortens the copied invitation while keeping the approval instructions. ## Linked Issues or Issue Description Related: #14933 and #15380. **What existing behavior does this improve?** Invitation copy and assistant connection management, in the organization’s Connections screen and the account-wide management page. **Subsystem affected** Cross-cutting: shared MCP connection types, server profile projection, UI and Storybook. **Current behavior** Connections are labeled only with a client name such as “Codex.” Revoked connections remain visible. The invitation includes an extra sentence about agent identity. **Proposed behavior** Show the authorizing person’s avatar and use names such as “Dotta’s Codex connection.” Hide revoked rows after successful revocation and when loading retained revoked grants. Failed revocation leaves the connection visible. Remove the extra identity sentence from invitation copy. **Reason and benefit** Make connection identity clear and keep the list focused on usable connections. **Breaking changes** The connection response adds optional `user` metadata with name and image. Older servers remain usable. Names and revoked-row visibility change in the UI; OAuth client identity, authorization and audit retention remain unchanged. ## What Changed - Shorten the shared invitation text. - Project the authorizing person’s name and avatar through the user-scoped connection endpoint, without returning email or credentials. - Reuse the existing Identity component and owner naming conventions across the connection page, catalog card and account-wide list. - Hide revoked grants and remove a successfully revoked row from the shared cache, even if the subsequent refresh fails. - Update documentation, regression tests and production-page Storybook fixtures and revocation journeys. ## Verification - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook` and `pnpm check:token-gates` pass. Final UI type checks also pass. - Focused consent and connection UI tests: 31 pass, including company filtering, legacy metadata, custom client names, user-initial fallbacks, failed revocation, retained revoked rows and refresh failure after successful revocation. - Existing MCP regression suite: 6 pass; 71 database checks are skipped locally because embedded PostgreSQL cannot start on this machine. The added database check verifies user-profile isolation and retained revocation history; CI runs these checks. - Browser verification with Storybook fixtures: owner avatar and name render; revocation removes the selected row in both production pages, leaves other connections visible, and restores the empty state after the last revocation. - All 54 current-head CI checks pass, with two optional Storybook jobs skipped. CI includes database, browser, runner, typecheck, build and clean-install canary coverage. - Greptile reviewed commit `1368d79e1066b418712224378d89d64c2b11cb86`: 5/5, no actionable findings or unresolved threads. - The full local `pnpm test:run` was stopped after complete CI passed. Local database coverage remains unavailable because embedded PostgreSQL cannot start; no full local-suite pass is claimed. ## Risks Low risk. The additive profile field is optional for compatibility. Revoked grants are filtered only from management UI and retained for audit. Revocation failure does not hide an active connection. No authorization scopes, token handling or schema changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, command execution and browser verification. The exact deployment ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a6306ba606 |
feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dd05c14e87 |
Update private task storage with master migrations
Preserve the reviewed privacy rules and existing PR history. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2d0c138122 |
Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also use assistants in Codex, Claude, and other MCP clients. > - The existing assistant connection can read work and create tasks or comments. > - It cannot edit tasks, exchange files, or manage normal agent and project settings. > - These operations must retain the person's permissions and Paperclip's execution rules. > - This pull request adds an explicit operation registry and separately consented configuration access. > - Assistants can manage work without receiving credentials or runner authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: assistant MCP, domain routes, consent UI, storage, and Product E2E. **Problem or motivation** A connected assistant cannot update tasks, maintain documents, attach files, or configure existing agents, projects, and skills. Users must leave the assistant for these routine actions. **Proposed solution** Add named tools and a restricted API registry to direct connections. Require separate configuration consent. Reuse domain routes and retry receipts. File uploads save the attachment when the byte transfer succeeds. **Alternatives considered** Arbitrary REST forwarding would expose administration and credential operations. Runner impersonation would bypass execution ownership. Separate upload completion calls add unnecessary client state. **Roadmap alignment** Checked ROADMAP.md and related MCP pull requests. This extends the human-authorized connection from #14933. It does not replace the runner or introduce agent impersonation. Companion Cloud routing and directory isolation: https://github.com/paperclipai/paperclip-cloud/pull/678. ## What Changed - Add task editing, finish/block, documents/revisions, deliverables, agent settings/instructions, projects/repositories, and skills/files. - Add an allowlisted API search/call registry with identical field restrictions, scopes, and retry identities. - Add unchecked configuration consent. Existing write grants retain their current authority. - Add hashed, expiring file transfer tickets and atomic upload receipts. No completion call is required. - Preserve company boundaries, human attribution, active native execution ownership, and execution review gates. - Add protocol/domain tests, consent stories, and eight paid Product E2E workflows. - Repair two CI fixture races: await cold route setup before assertions, and wait for asynchronously loaded connection copy. Both fixture suites pass (24 + 48 tests). ## Verification - Consent revision: one write-access checkbox controls requested work and configuration permissions in browser and device flows. All 16 consent tests, UI typecheck/build and token gates pass. Updated interactive stories cover default approval, opt-out and viewer restrictions. The paid browser helper uses the new exact label. Real GPT-5.4 Mini Product E2E passes 2/2 at `64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic retries, cleanup passed; $0.04149375 estimated assistant cost plus unpriced worker usage. Raw results, usage and source fingerprints are retained in the worktree. UI and Product E2E typechecks pass. - Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks pass; two optional Storybook checks skip. Greptile 5/5 on that head, no unresolved review threads. Final consent head `64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing and Greptile 5/5 with no unresolved threads. The unchanged Cursor sandbox test had one 10-second timeout, passed in local isolation, and passed its single CI rerun; the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37554106934). The existing chat retry-denial browser test had one visibility failure; its single rerun passes, and the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37542735691). - Full workspace `pnpm -r typecheck` and `pnpm build` pass at final runtime source `b2196fae1`. UI token gates pass. - 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests pass, including one-connection PostgreSQL OAuth and concurrent upload retries. - Paid Product E2E: all eight expanded cases qualified across Mini, Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell timed out before application startup. Final affected-case qualification passes 9/9 on all three models with grader v16, including the failed cell. Automatic retries disabled; failures, costs, source hashes and independent durable-state/file assertions are retained in [the verification record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md). - Actual Codex CLI, Claude Code and OpenCode clients completed local reads/mutations. Codex wrote a report, Claude updated it in a later conversation, and OpenCode uploaded/downloaded a file with matching SHA-256 and registered the attachment. Revoking the CLI grant rejects subsequent bridge initialization. - Butter staging is verified on final runtime `b2196fae1` ([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)). A fresh OpenCode workspace fetched the copied invitation, configured remote MCP, started OAuth and reached real consent with configuration unchecked. Invalid transfer tickets return 403 through Cloud. Human approval for the new persistent staging grant is pending; hosted task/file success is not yet claimed. The final transaction fix is deployed. - Full local `pnpm test:run` passed 15,614 general-server tests but stopped on two macOS timeouts. The heartbeat test passed in isolation; the existing 40,000-file Git stress fixture timed out again. Its Linux CI lane passes. Later local full-suite phases did not run after the timeout; this is not an all-green local full-suite claim. - Instructions and security limits are in `doc/public-mcp.md`; the saved plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`. ## Risks - This expands the experimental direct MCP surface. Explicit schemas and domain permissions must stay synchronized. - Migration 0311 adds transfer tickets and upload receipts. Expired orphan cleanup must not remove committed attachments. - Configuration requires a new consent request containing that scope; the single write-access choice controls it alongside work mutations. Refreshing an old grant does not add it. - The public directory keeps its original ten tools through the companion Cloud change. - Hosted consent/work proof remains the final delivery gate. The PR stays draft while approval of the new staging grant is pending; code checks and review are green. Merging is a separate action. ## Model Used OpenAI Codex (GPT-6, tool use and code execution). The exact serving model ID and context window are not exposed in this session. Paid evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001; claude-sonnet-4-6. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite macOS limitation disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
892b0b3606 |
fix: scope quota reports to authorized subscription accounts (#14994)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d05ad82b8d |
Repair master CI regressions for private task stack
Preserve existing PR history and verified downstream privacy rules. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3ebd7bc0c9 |
test: isolate accounting fixtures and include all adapter suites (#14989)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6daa031600 |
Update private task storage with master migrations
Preserve the reviewed privacy rules and existing PR history. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ac8f3eb143 |
feat(ui): remove star and leave buttons from agent index (#15400)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Agents index page (`/agents/all`) shows one row per agent in both the streamlined and production sides of the UI > - Each row had a star toggle and a leave/join button that write per-user state through the resource-memberships API (`agent_memberships` table) > - The default streamlined sidebar no longer lists agents, so the star and leave actions on the index no longer feed a visible pinned list > - This pull request removes the star toggle and the leave/join button from agent index rows and org-tree nodes > - The backend resource-memberships endpoints and the `agent_memberships` table stay live because the agent detail page and the legacy sidebar still use them > - The benefit is a cleaner index page with no redundant actions for state that is not surfaced on the page itself ## Linked Issues or Issue Description No public GitHub issue exists for this change; the description follows the feature request issue template inline. **Subsystem affected** ui/ — React + Vite board UI. The change is presentational only; no backend or schema code is touched. **Problem or motivation** The agent index rows show a star toggle and a leave/join button. Those two buttons write per-user membership state that the default streamlined UI no longer surfaces anywhere on the page itself. The default sidebar shows a plain "Agents" nav item and does not list agents. The buttons are now redundant on the index. **Proposed solution** Remove both buttons from the agent index list rows and org-tree nodes in both UI variants. Keep the data layer. The agent detail page still renders the header star toggle and the join/leave banner, and the legacy sidebar still renders star and leave rows for instances that opt out of the streamlined UI. **Alternatives considered** - Remove the backend endpoints too. Rejected: the agent detail page and the legacy sidebar still read and write the same state. - Keep the buttons hidden behind a setting. Rejected: no product need, and the plan approved the straightforward removal. **Roadmap alignment** This is a small presentational change. It does not overlap any section in `ROADMAP.md`. **Additional context** Rows whose membership state is `left` keep their dimmed styling on the index. ## What Changed - Removed the `StarToggle` component from agent index list rows in `ui/src/pages/Agents.tsx` and `ui/src/pages/Agents.production.tsx` - Removed the `MembershipAction` (Leave/Join) component from agent index list rows in both files - Removed both actions from the org-tree nodes in both files - Dropped now-unused imports and derived values (`StarToggle`, `MembershipAction`, `isStarred`, `useResourceMembershipMutation`, and the per-row pending/starred variables) - Kept `useResourceMemberships` and the dim-on-left row styling so rows whose membership is `left` remain visually identified - Updated `ui/src/pages/Agents.test.tsx`: a new test asserts both modes omit star and leave/join actions from list and org-chart views; the dim-on-left tests are kept ## Verification - `pnpm --filter @paperclipai/ui typecheck` passes - `pnpm --filter @paperclipai/ui build` passes - UI Vitest suite for `Agents.test.tsx` passes (21 tests) - `pnpm check:token-gates` is clean - Manual check on `/agents/all`: no star and no Leave/Join buttons in the list view or the org-chart view; star/leave still work on the agent detail page ## Risks Low risk. This is a presentational UI change with no backend or schema changes. Star/leave remain available on the agent detail page and the legacy sidebar. No migration is involved. ## Model Used - Provider: DeepSeek (via the Paperclip opencode_local adapter, model `openrouter/~deepseek/deepseek-v4-flash-latest`) - Model: deepseek-v4-flash (OpenRouter `openrouter/~deepseek/deepseek-v4-flash-latest`) - Context window: 128K; used with tool use in the repository - Assisted with the code change and this PR body; the plan and scope came from the tracking work item ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f77fcbf4bf |
feat(apps): add Telem.AI web search connection (#15379)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Research agents need current web information. > - The Apps catalog connects agents to remote MCP tools through the normal access rules. > - Telem.AI supplies web search and page reading through one API key. > - This PR adds its catalog entry, optional search settings, artwork, and setup guide. > - The connection keeps an operator's saved header policy when they reconnect. ## Linked Issues or Issue Description Refs #15302. Related catalog work: #13881. This PR continues #15302 by Yifei Ai (@aiwen324). Thank you for the connector and its review fixes. All seven original commits are preserved. GitHub denied the attempt to push to the contributor's fork, so this branch retains the repair commit and merges current master. Master now includes the same test fix. For a squash merge, keep the original author in the final commit message: ```text Co-Authored-By: Yifei Ai <aiwen324@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> ``` A search found no separate public Telem issue or competing Telem PR. The existing Apps path matches the roadmap. **Agent or provider** Telem.AI provides web search and page reading through a hosted MCP server. **Why this adapter is useful** Agents can use multiple search providers through one governed connection. Operators can set the search tier, auto routing, and provider lists. **How the agent is invoked** The remote MCP server uses Streamable HTTP at `https://mcp.telem.ai/mcp`. The API key uses an `Authorization: Bearer` header. See the [official MCP guide](https://docs.telem.ai/integrations/mcp/). ## What Changed - Add the Telem.AI definition, research entry, permission review, and generated registry entry. - Add four optional settings. Unset settings send no request header. - Add official light and dark artwork, source records, and a setup guide. - Forward company, issue, agent, run, project, and correlation IDs by default. Preserve a saved policy, including disabled forwarding, on reconnect. - Add catalog and connection tests. ## Verification Current head: `b4164477fb1b312a504789bf17b51c963244c0fd`. Merged master: `228f0e2807c5b59d2aa129cf2d80b9777ebabf07`. - Resolved five shared catalog conflicts after the Superagent connection merged. - Keep both providers in the research ledger, generated registry, generator, branding manifest, and connection guide index. - Correct the combined catalog totals: 52 self-serve candidates, 55 research entries, and 68 Apps entries. - The published Git tree exactly matches the tested local resolution. - Catalog and Apps UI suites: **295 tests pass** after the catalog count fixes. - Connection service suite: **387 tests pass** in the full run. Its only failure was the old catalog count. That test passes on a focused rerun after the fix. This gives **388 passing service tests** across the two runs. - Total focused coverage: **683 passing tests**. The first runs exposed four fixed-count assertions that needed the combined totals. - Shared package build and plugin SDK compile pass. Token gates and whitespace checks pass. - Generation with `--definitions-only` reproduces the Telem definition and registry. The unrelated AgentMail and Linear drift remains excluded. - Local UI typecheck ended with exit 137 at the container memory limit. Full local typecheck, test, and build are not claimed. Earlier runs also recorded missing Cargo and Node development headers. - GitHub reports a clean merge state against master `228f0e280`. - All 54 checks are complete: **52 passed and two Storybook checks skipped**. No check failed or remains pending. - [CI](https://github.com/paperclipai/paperclip/actions/runs/37545412406) passes on this head. This includes typecheck, build, tests, browser shards, Runner checks, and Canary Dry Run. - [Greptile](https://github.com/paperclipai/paperclip/pull/15379#issuecomment-6024519537) is **5/5 on this head**. There are no review threads, open P2s, recommendations, or follow-ups. - [Superagent](https://github.com/paperclipai/paperclip/runs/112548142930) passes. - Final recovery checks confirm all seven original commits and current master remain in history. Token gates and whitespace checks pass. - The final recovery run makes no source change. It verifies the published repair and retains the local check limits below. - [Commitperclip](https://github.com/paperclipai/paperclip/actions/runs/37545407943) passes with no failures. Its only informational note asks the merger to keep the author trailer above. - All seven original contribution commits remain in history. The diff against master contains the same 14 Telem files. It adds no dependency, lockfile, schema, or workflow change. - The managed GitHub CLI capability was missing in this run. The installed GitHub connection applied the base files, merged master, then restored the tested combined catalog. No history was rewritten. The previous head `2d32a0094` passed all remote gates and had Greptile 5/5. Those results do not verify this new head. The original PR reports live setup, discovery, settings headers, gateway calls, and context-header forwarding. This repair does not repeat those account-bound checks. The permission record still marks maintainer live qualification as outstanding. ## Risks - Telem.AI receives the six context IDs by default. The saved header policy controls forwarding. Search use is billed to the account that owns the key. - All agents on a connection share its search settings. - The merge uses master's route-test setup unchanged. The company-boundary assertions remain intact. - No schema, dependency, or workflow change is included. - Live provider evidence is attributed to the original contributor. Maintainer live qualification remains outside this CI repair. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Original contribution: Anthropic Claude Opus 5.5 (`claude-opus-5-5`), 1M-token context, through Claude Code with shell, editing, and test tools, as disclosed in #15302. - CI repair and review: OpenAI `gpt-6-astra`, through Codex with reasoning, shell, editing, and GitHub tools. The runtime does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Yifei Ai <aiwen324@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
228f0e2807 |
fix(ui): make pull request review waits visible (#15375)
## Thinking Path > - Paperclip lets people supervise agent work and review its outputs. > - A task can wait for a human to review a pull request while an agent schedules checks. > - The PR work product and detected external links use separate display paths. > - A private PR can disappear from the useful sidebar view when its provider lookup fails. > - The monitor countdown says when the next check runs but omits the saved review request. > - This change shows the saved PR and requested review in the task's existing surfaces. ## Linked Issues or Issue Description **What happened?** A task kept checking whether a private GitHub PR had merged. Its work product requested board review, but the Properties sidebar used detected external objects instead. The PR URL was stored only in metadata, which the PR refresh path did not read. The composer showed a generic monitor countdown with no PR link or review action. **Expected behavior** Show saved PR links even if provider access fails. During a GitHub monitor wait, make outstanding review requests visible with the next check time and a status-check action. Stop requesting review when a PR merges or closes. **Steps to reproduce** 1. Register a PR work product with `reviewState: needs_board_review` and its URL in `metadata.url`. 2. Schedule an external-service monitor for GitHub. 3. Open the task with external-object lookup unavailable or unable to access the private repository. 4. Inspect Properties and the composer wait strip. Related work: #14469 added rich artifact cards. #8759 addresses attention on monitored task blockers. This change uses the existing work products and monitor action; it adds no blocker or approval mechanism. ## What Changed - Show saved PRs in Properties, including metadata-only links, independent of external-object availability. - Deduplicate equivalent GitHub PR URLs while preserving the saved navigation link, and put explicit review requests first. - Keep provider status and freshness on the combined PR row; keep private PR links usable when lookup fails. - Show the saved PR review request above the composer and in the existing monitor banner during a GitHub monitor wait. - Label the existing monitor action `Check status` and show check failures inline. - Suppress review prompts for merged, closed, or archived PRs even if their review flag is stale. - Refresh PR metadata using `metadata.url` and the existing `repository` alias. - Share saved work-product reads across the thread, Properties, and Artifacts. Refresh GitHub in a separate query and enrich only matching PR versions. - Start a fresh saved-row request on live invalidation so late provider responses cannot hide new artifacts or changed review requests. - Cover stalled GitHub lookups in both panels so refresh latency cannot hide saved work. - Scope monitor-check mutation state to the task so failures and late responses do not leak across navigation. - Refresh GitHub status when the displayed run finishes, including when saved PR rows have not changed. - Clarify that PR review and external release handoffs need a saved human-input interaction with an agent assigned for continuation; a flag, monitor, or handoff comment alone does not create that card. ## Verification - 386 tests pass across eleven affected UI and server suites on `184c7d8746`. Regressions cover saved links, provider status, cold-cache loading, panel reopening, monitor errors during navigation, and live updates during provider refresh. - Three live-update regressions and two run-completion regressions fail before their fixes and pass afterward. These tests use the real API client's GET coalescing and abort handling. - Full repository typecheck and build passed on the merged parent `1a49112dd2`. The latest UI changes pass UI typecheck and token gates; CI also verifies the latest build. Capability contract/inventory checks pass. - Storybook build and browser checks passed before the query race fix. Browser checks cover desktop and 390px mobile review waits, plus video, mixed-file, and empty artifact galleries after background refresh. - The earlier full local `pnpm test:run` was stopped under disk pressure. It reported failures outside the changed suites; an isolated skill-cache run reproduced three existing macOS permission failures. The full local suite was not rerun for this follow-up. - Latest-head CI passes on `184c7d8746`, including the browser shards and canary dry run. Apex is 5/5 with no actionable findings and no unresolved threads. All eight reported findings are addressed. ## Risks - This uses saved PR review state. When GitHub access fails, the saved state can remain stale until the agent updates it. The link remains visible and the status check remains available. - `Check status` wakes the existing monitor owner. It does not merge a PR, accept an approval, or mark the task done. - No schema, permissions, or scheduler behavior changes. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository analysis, code editing, and test execution. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — 277 affected tests - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cbd21b3e7 |
feat(apps): add Superagent connection (#15394)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents reach external services through the Apps catalog. Each catalog entry is a reviewed `AppDefinition` that connects a provider's hosted MCP server to Paperclip's shared vault, grants, policies, gateway, and audit trail. > - Superagent (`superagent.sh`) is a security platform. Its hosted MCP server lets agents read and triage security findings, start red-team reports, run Contributor Trust and dependency update jobs, and score web pages, email, files, skills, MCP repositories, and packages before an agent trusts them. > - Superagent is not in the catalog. Operators must use the generic "Connect your own MCP server" flow. That flow has no branding, no key guidance, and no warning about billable or destructive tools. > - The connector playbook supports this provider with existing definition fields. The server accepts an organization API key as a bearer header. It publishes no OAuth authorization-server metadata, so browser sign-in is not possible. > - This pull request adds the Superagent definition, its artwork, the research and permission-review ledger rows, documentation, deterministic tests, and one reviewed risk-classification rule. > - The benefit is a branded, governed Superagent connection with clear key guidance and a warning about tools that cost credits or delete data. ## Linked Issues or Issue Description **Problem or motivation** Teams that use Superagent for PR security, red teaming, and agent guardrails want their Paperclip agents to read findings, return structured reports, and score content before they use it. Superagent is not in the Apps catalog. Operators must paste the MCP URL and an `Authorization` header into the generic remote-MCP flow. That flow gives no branding and no provider guidance. It also does not tell the operator that the key reaches the whole organization, or that some tools consume credits or permanently delete findings. **Proposed solution** Add a catalog-only Superagent connection that follows the connector playbook. It has one method: a customer organization API key (`sk_live_...`), sent as an `Authorization: Bearer` header to `https://www.superagent.sh/mcp`. The field helper text explains that Superagent keys are not scoped. The method warning tells operators to set billable and destructive actions to Ask first before agents run unattended. All discovered tools stay governed by the normal per-action policies. **Alternatives considered** A browser sign-in method was not added. The server's protected-resource metadata names `https://superagent.sh` as its authorization server, but that origin publishes no `oauth-authorization-server` or `openid-configuration` document, so Paperclip cannot discover OAuth endpoints. A plugin was not needed because the connection needs no custom UI, tables, workers, or webhooks. Relying on the generic risk classifier was not enough. Several Superagent mutations (`triage_finding`, `scan_*`, `restore_agent_builtin_rule`) use names that it reads as reads, so a narrow reviewed Superagent rule was added instead. **Roadmap alignment** This extends the existing self-serve remote-MCP connection catalog. It does not overlap planned core work. ## What Changed - Added the `superagent` row to `packages/shared/src/self-serve-mcp-research.json` (API-key auth, risk tier S4). - Added the `superagent` provider to `scripts/ingest-app-definitions.mjs` (category, key placement and placeholder, console links, guidance, description). Regenerated `packages/shared/src/app-definitions/superagent.json` and the generated registry. - Added the `superagent/mcp-api-key` permission review to `doc/connections/tool-method-permission-reviews.json`, with key-permission text and evidence links. - Added Superagent's official mark (`ui/public/brands/apps/superagent.png`, the 460×460 avatar of the official `superagent-ai` GitHub organization) and the brand manifest entry. - Added gallery copy for the Superagent card. - Added a reviewed Superagent rule to `classifyRisk` in `server/src/services/tool-access.ts`. Only `list_*` and `get_*` tools, and tools that Superagent marks read-only, are reads. `delete_*` and `revoke_agent_client` are destructive. All other tools are writes, so billable and rule-changing tools can be set to Ask first. - Put the Ask-first advice in the API-key helper text, because the key form shows helper text and not method warnings. - Added `doc/connections/SUPERAGENT.md` (transport and auth, why there is no OAuth, administrator setup, capabilities and policy, manifest, brand provenance, validation hook). Linked it from the connections README and the permission audit. - Tests: definition shape, store visibility and artwork, URL recognition, the bearer header on discovery with the key kept out of connection config, read/write/destructive classification of fixture tools, the Superagent risk rule (including `triage_finding` and `restore_agent_builtin_rule`), the visible Ask-first advice and API-key gating of the connect form, and the pinned catalog counts. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts packages/shared/src/app-definitions-url.test.ts ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx`: 317 passed. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts`: 385 passed. - `node scripts/check-app-brand-assets.mjs` and `node --test scripts/app-brand-validation.test.mjs`: passed. - `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter @paperclipai/server typecheck`, and `pnpm --filter @paperclipai/ui typecheck`: clean. - Manual: in a local instance, open Apps → Browse and confirm the Superagent card and icon. Open `/apps/connect?source=superagent`. Confirm the single API-key method, and confirm that Connect enables only after a key is entered. - Live metadata probe on 2026-10-06: an unauthenticated `initialize` on `https://www.superagent.sh/mcp` returns 401 with `resource_metadata="https://www.superagent.sh/.well-known/oauth-protected-resource"`. That document returns 200. No authorization-server metadata exists at the named issuer. ## Risks - Low risk to existing providers. The change is additive catalog data plus tests. The generated registry only gains one import. The new risk rule runs only for Superagent connections. - A Superagent key reaches its whole organization. Some tools consume credits (`create_*_report`, `triage_finding`) or delete data permanently (`delete_finding`). Every action starts Allowed under the current product default. The key helper text tells operators to set these actions to Ask first. - The permission-review ledger records live proof as not run. No Superagent account was used. The lifecycle checklist in `doc/connections/SUPERAGENT.md` needs a documented pass before the entry is fully qualified. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with extended thinking and tool use (shell, file editing, web fetch, browser checks). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ee3c728b95 |
fix(ui): keep No project available in new composer (#15382)
Keep the no-project action first and available during project search, while preserving matching-project keyboard selection. Add focused regression coverage for ordering, projectless submission, and filtered Enter behavior. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
582911ba74 |
fix(access): give Operators default company editing permissions (#15377)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Human company members receive permission grants from their role preset. > - Human invitations use the Operator role by default. > - The old Operator preset only granted task assignment, so ordinary members could not edit agents or manage connections. > - This pull request gives Operators company editing and audit access while keeping join approval and member-permission management separate. > - Members can configure their agents and accounts without an Owner role. ## Linked Issues or Issue Description **What existing behavior does this improve?** The default permission grants for human Operators and the invitation role description. **Subsystem affected** server/ permission presets and invitation authorization, plus ui/ invitation and tool access gates. **Current behavior** Operators receive only `tasks:assign` as an explicit default grant. Agent configuration and managed account setup fail their permission checks. **Proposed behavior** Operators receive agent creation/configuration, skills, environments, invitations, task assignment, pipelines, connections, tool management/use, and audit grants. The preset excludes `joins:approve` and `users:manage_permissions`. Operators can invite Operators and Viewers. Inviting a role with either excluded power requires that power, so invitations cannot bypass the restriction. **Reason and benefit** Ordinary company members can edit company work and connect accounts through the existing routes. **Breaking changes** The Operator preset grants more permissions. The existing startup and Cloud sign-in seeding paths can insert these missing grants for existing Operators. They retain custom grant scopes. Explicit invitation grants still take precedence. Inviting Admins now requires join approval; inviting Owners also requires member-permission management. This PR adds no migration or new backfill path. Related: #5945 describes missing role grants in another membership entry point. This PR changes the preset and does not change that endpoint. ## What Changed - Expanded the Operator preset to 14 company editing, invitation, tool, and audit grants. - Kept join approval and member-permission management out of the preset. - Required those powers when an invitation's selected human role includes them, preventing Operators from delegating the excluded powers through Owner/Admin invitations. - Updated the invitation role copy and the product specifications. - Allowed active Operators (including legacy Member roles) through the invite shortcut and advanced tool/profile UI gates. Kept inactive memberships, Viewers, and other companies excluded; API grants remain authoritative. - Added tests for the default Operator grants, both excluded actions, agent editing, company boundaries, and invitation role authorization. - Simplified adapter-route test imports so authorization errors and the error handler use one module graph; retained the upstream mock-initialization fix. - Kept Owner/Admin/Viewer presets and explicit invitation grants unchanged. Added no EE code or database migration. ## Verification - Permission and invitation tests: 38 passed across access service, invitation defaults, and invitation creation routes. - UI access tests: 61 passed across invitation shortcuts, Operator tool/profile access, and invitation UI. - Adapter route tests: 15 passed after resolving the upstream test setup conflict. - `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`: passed on the final branch. - The initial local `pnpm test:run` general-server batch reported three tool-access failures and eight upload-helper timeouts. Both complete affected suites passed in isolation on the final branch: 403 tests. The initial full local run was not green. - Both complete tool-access and adapter route suites also passed on an untouched snapshot of base `88ff98b83`: 387 tests. The earlier tool failures did not reproduce there. - [All CI gates passed](https://github.com/paperclipai/paperclip/actions/runs/37527422346) for `068be153b9c2a064f2aaccc0627fdf458e73747e`, including the previously failing serialized adapter lane, general tests, E2E, typecheck, build, Runner checks, and canary dry run. - Greptile scored that exact commit 5/5. Superagent passed. Both invitation review threads are resolved. ## Risks - Operators gain broad company editing and invitation access by design. - Existing default seeding can add the new missing grants during startup or Cloud sign-in. It does not replace existing scopes. - Company boundaries and the two excluded permission checks still apply. - An Admin without member-permission management can no longer invite an Owner. Operators retain invitation access for Operator/Viewer roles and agent-only invites. - This PR changes company permissions. Cloud workspace invitation rules remain separate. ## Model Used OpenAI Codex, GPT-6. The exact deployment model ID and context window were not exposed in this session. Used repository inspection, code editing, shell execution, and test tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting a merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ca0d26a99 |
fix(connections): recover missing personal AI credentials in chat (#15376)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e38d6d16b6 |
feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect accounts and choose an agent harness and model. > - The runtime change in #14970 supports custom providers on those connections. > - Normal setup must stay simple while advanced users can choose a compatible gateway. > - Shared connector rows and access controls keep these choices consistent. > - This pull request refines the agent setup UI and adds review stories and repeatable browser qualification. > - The qualification checks real tools and downloaded outputs, not only a successful run status. ## Linked Issues or Issue Description Refs #14970, #37, #13083, #14104, #14565, #12692. The core implementation in #14970 is merged. This branch incorporates its squash commit and targets `master`. Both PRs contain our implementation. #14016 is a reference only and is not a dependency. This PR has 96 changed files. ## What Changed - Complete model-provider connector presentation beside other connectors. Each row uses the existing Connect action and connection list. Tags are stored without category UI. The base PR includes the provider forms and routes. - Show persistent Subscription, API Key, and Advanced choices. Label Advanced as Custom Gateway. Reuse provider logos, connection lists, and permissions controls. Default access to the organization and all agents when permitted; keep narrowing controls under Advanced. - Keep Configure reachable before subscription sign-in, so users can select a supported environment when the default cannot sign in. Testing and saving still require a connection. Show the execution environment in Configure. Preserve the confirmed Connect choice. Editing a method, credential, saved account, or advanced choice requires that current choice to connect before testing or saving. Use matching model and thinking-effort dropdowns and retain connection icons in selected values. - Preserve the new harness model default when switching an existing OpenCode agent to Codex or Claude, and resolve user-selected model names with the effective harness. - Load popular OpenRouter models through the shared connection-model discovery path. Keep explicit model lists and manual model entry available. - Group onboarding, connection setup, agent runtime, management, recovery, and production-component stories under AI Connections / Provider routing. - Add an explicit-only provider-connections browser suite for managed local or existing local/staging targets. Use private browser profiles and credential handoffs. Support human-assisted subscription sign-in without sharing passwords or tokens in reports. - Verify persisted connection identity, runtime probes, tool execution, exact artifact bytes, completion, and context-dependent follow-up. Retain source/model provenance, cost bounds, closed error diagnostics, original failures, and cleanup evidence. - Add Gemini startup-model and skill-root fixes, Grok private-history detection, ACP filesystem regression fixtures, selected-workspace handling for local Hermes, and artifact-helper workspace fallback. - Keep managed Grok runtime homes disposable. Remove host-side transcript retention/restoration because private file modes do not isolate same-user agent processes. Ignore earlier development archives and use a fresh task handoff when history is unavailable. Verify the absence of restored transcripts with a separate same-user process. - Capture stopped-run diagnostics before deleting an attached-company fixture agent. Track creation and owned sign-in receipts; revoke only this attempt's accounts and never adopt a concurrent campaign's newly created account. Preserve failure signals and final status through cleanup. - Require the requested environment in the saved agent and every run, including follow-ups. Reject a forced incompatible target. Keep one cancellation state through startup, every cell, reporting, and teardown for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after interruption. Document qualification limits. ## Verification - Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes master `d9f600043`. The security fix in `a758fde31` passes full workspace typecheck, production build, and 119 connection/Grok regressions. The unchanged UI passes all 126 configuration/model-discovery tests and token gates. The final published-guide correction passes Grok adapter typecheck. Earlier head `eebd8225c` passed the complete deterministic runner suite (1,404 Vitest tests and 128 Node tests) and all CI jobs. Current-head CI run `37520147514` passed all 47 jobs, including the full sharded Vitest and browser matrix, production build, and canary dry run. All 55 checks completed: 53 successes and two expected skips. The current-head security scan passed, Greptile is 5/5, and no review threads remain open. - A separate same-user process reproduced reading a restored Grok transcript before the security fix. The regression now finds no transcript. Existing fresh-session fallback and ordinary session metadata behavior pass. - The final account-choice and cleanup fixes pass 85 setup tests and 26 qualification-harness tests. Regressions verify that editing a connection invalidates confirmation, Configure remains reachable before sign-in, diagnostics are captured before fixture deletion, and concurrent campaigns cannot adopt or revoke each other's accounts. UI and E2E typechecks pass. - The Storybook build and actual Chromium production-component stories passed during this change. Review the neighboring AI Connections / Provider routing stories, regular connector rows, three connection modes, model discovery, and the single execution-environment control in Configure. - Cancellation smoke verified authenticated cleanup before browser close for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during startup and reporting, missing-file ACP resource errors, and preserved permission denials. Both ACP runtime versions and 54 ACPX/Grok regressions passed. The deterministic connection-intent browser suite passed two tests. - Historical local qualification retained 43 passing API/gateway cells out of 46, with downloaded outputs and follow-up receipts. These attempts span earlier builds; they do not qualify this exact commit or staging. Subscription combinations, Gemini overloads, and the unresolved follow-up failure remain recorded rather than counted as passing. - Use `pnpm test:e2e:runner -- --list --suite provider-connections` to inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md` for credentials, target URL, sign-in assistance, budget, evidence, and cleanup. Paid live tests remain opt-in. ## Risks - The core implementation in #14970 is merged. This PR adds no database migration of its own. - Subscription login needs an interactive provider session. Dedicated accounts and staging qualification remain follow-up work; this PR does not certify every login combination for production. - Managed Grok transcript resume is deferred until provider history has an OS isolation or authorized broker solution. Follow-ups start fresh with Paperclip task context; earlier live Grok results do not qualify this behavior. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Live overloads and one unresolved follow-up timeout remain recorded. The stock CLI is unchanged, and those cases are not marked as passing. - Real-provider tests spend credits and use private credential/evidence directories. The launcher requires explicit selection and checks target ownership. It must not attach to a developer's database by accident. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local remain outside custom provider setup. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a7a2ed63a0 |
fix(ui): restore wheel and touch scrolling in task selectors (#15374)
Restore wheel and touch scrolling in nested task selectors and keep mobile picker controls reachable above the keyboard. Stabilize the affected server route test harnesses for Vitest 5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
88ff98b83d |
feat(apps): add read-only Enterpret MCP with OAuth and org tokens (#13906)
## Thinking Path > - Paperclip agents need governed access to Enterpret's hosted MCP tools for customer feedback. > - The official Enterpret MCP is read-only. Enterpret Agent’s beta write MCP is a separate service and is outside this PR. > - This PR supports both organization auth tokens and personal browser OAuth for the official MCP. > - Enterpret committed to fixing OAuth scope behavior and reducing its token-revocation cache window. The provider fixes will follow after connector deployment and do not gate this release. Paperclip functional validation remains required. New tools stay quarantined on refresh. > - This PR combines the original connector and Vivek's token-path work on current master, with live self-hosted QA recorded in the validation guide. ## Linked Issues or Issue Description No public issue covers this connector. This PR incorporates the work from #14737. The independent OAuth scope-provenance fix is #14059; it is not required for the organization-token path. #13890 is a merged connector precedent. #13692 provides the connector-authoring skill. **Problem or motivation** The generic MCP setup can reach Enterpret, but it has no reviewed Enterpret entry, method-specific setup, artwork, or policy defaults. The original connector branch had also fallen behind master. **Proposed solution** Publish the Enterpret catalog entry with an organization auth token as the selectable method. Support browser OAuth through instance-local DCR on the same read-only endpoint. Default `run_graph_query` to Ask first and quarantine tools newly found on refresh. ## What Changed - Added the Enterpret app definition, generated registry, official icon, Browse card, setup copy, tests, and Storybook states. - Combined #14737 and resolved the current-master conflicts, including the connector permission audit. - Preserved catalog quarantine on token reconnect and updated expiry guidance to follow the token's dashboard expiry. The QA token displayed three months, contradicting the former six-month copy. - Enabled personal OAuth for the official read-only MCP. Removed the write-capability inference and OAuth draft-only setup restriction. The beta Agent endpoint is excluded. - Preserved tool quarantine during token reconnect, with a regression test for existing quarantined and newly discovered tools. - Added OAuth scope and revocation-cache disclosures and an interactive Storybook retry check. - Merged master `d9f600043` into the existing PR branch without rewriting history. Resolved four conflicts, regenerated the combined registry, and retained the model-provider assertions with a catalog count of 66. - Recorded live QA and its limits in [doc/connections/ENTERPRET.md](https://github.com/paperclipai/paperclip/blob/950bacc52/doc/connections/ENTERPRET.md). ## Verification Current head: `6c73d3084202ccbb25080e7a4e9c3467bc1341dd`. All 513 focused connector, discovery, OAuth, policy, shared-definition, and UI-handoff tests passed across 11 suites. Full repository typecheck, build, Storybook build, token gates, module boundaries, Node policy, no-git-push policy, and typecheck-build-gaps passed. Release-registry tests passed 129/129 with modern Bash. The four failures with macOS Bash 3.2 also reproduced on untouched master. The full test suite passed in CI on this exact head, including every general-test and serialized-server shard and all eight E2E shards. The duplicate local `pnpm test:run` was stopped after these remote test lanes passed; it did not complete locally. The combined registry was regenerated. Regeneration also reproduces existing Linear and AgentMail definition drift on untouched master; those unrelated changes were excluded. On an isolated self-hosted Paperclip instance, an organization token connected and discovered eight actions. `get_organization_details` returned the intended organization before and after reconnect. Catalog refresh, invalid-token failure and valid-token recovery, activity logging, local disable, and local removal were exercised. No feedback quotes or graph query were requested. Paperclip revoked both QA connection secrets and archived the QA connections. The token-path agent access preview correctly showed an action Off, but the board Test endpoint still invoked it as the board user. This shared Test-path mismatch is documented; it is not evidence that a real agent can bypass policy. An actual token-path agent-session denial, a real agent process, expiry recovery, server/VPS, and Cloud remain unverified. Enterpret's dashboard confirmed the QA token was revoked, but the same token immediately reconnected and completed a safe read. Vivek explained this as a 24-hour validity cache and committed to reducing the window. The new maximum delay and deployment remain unverified. The earlier OAuth check reported `mcp:write` and `email` for an `mcp:read` request. This is a scope mismatch; it did not prove access to write tools. On 2026-10-01 the account holder relayed Vivek’s clarification that the official MCP is read-only and the beta Agent write MCP is separate. ## Risks The organization token can read the organization's customer feedback, and provider-side revocation did not immediately stop access. Operators should check each token's displayed expiry and remove Paperclip's local connection when retiring access. The account holder confirmed the provider fixes are fast follows after deployment. Release can proceed with the current scope reporting and up-to-24-hour revocation delay documented, once fresh OAuth authorization/refresh, real-agent allow/deny, CI, and review pass. Retest the reduced revocation window after Enterpret deploys it. Scope labels alone must not be used to claim write access. Greptile is 5/5 on head `6c73d3084` with no actionable findings. All current-head CI checks pass, including the canary dry run. Live agent-session denial and fresh OAuth refresh remain explicit deployment QA items. The board Test panel does not prove agent authorization. Provider scope-reporting and revocation-cache fixes are accepted fast follows after deployment. ## Model Used Original connector and validation record: Claude Opus 5 through Claude Code. Integration, current-master reconciliation, and token-path QA: OpenAI Codex, GPT-6 family (exact host variant not exposed), using repository tools and an isolated runtime. Contributor commits retain their authors. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal issue ID - [x] I have run relevant local tests and checks - [x] I have updated relevant documentation - [x] I have considered and documented the risks above - [x] All current-head CI gates are green - [x] Greptile is 5/5 with no actionable follow-ups If squash-merging, preserve contributor credit in the squash message: Co-Authored-By: Vivek Kaushal <kaushalvivek@users.noreply.github.com> --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Vivek Kaushal <vivek@enterpret.com> |
||
|
|
854af7df19 |
build(deps-dev): bump vitest from 4.1.11 to 5.0.3 (#12969)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.11 to 5.0.3. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v5.0.3</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Isolate <code>result.status</code> between <code>repeats</code> runs - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11218">vitest-dev/vitest#11218</a> <a href="https://github.com/vitest-dev/vitest/commit/5dbebe9e3"><!-- raw HTML omitted -->(5dbeb)<!-- raw HTML omitted --></a></li> <li>Don't print an interceptor warning in browser mode - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11377">vitest-dev/vitest#11377</a> <a href="https://github.com/vitest-dev/vitest/commit/15cc006aa"><!-- raw HTML omitted -->(15cc0)<!-- raw HTML omitted --></a></li> <li>Don't retry when <code>test.fails</code> expectedly failed - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11219">vitest-dev/vitest#11219</a> <a href="https://github.com/vitest-dev/vitest/commit/b24585f08"><!-- raw HTML omitted -->(b2458)<!-- raw HTML omitted --></a></li> <li>Scope cache key generators to projects - by <a href="https://github.com/ecoyoung"><code>@ecoyoung</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11281">vitest-dev/vitest#11281</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11301">vitest-dev/vitest#11301</a> <a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1d"><!-- raw HTML omitted -->(92ba7)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Delay server <code>listen</code> until tests start running - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11366">vitest-dev/vitest#11366</a> <a href="https://github.com/vitest-dev/vitest/commit/7d8ed3e9b"><!-- raw HTML omitted -->(7d8ed)<!-- raw HTML omitted --></a></li> <li>Check mock path boundaries - by <a href="https://github.com/saryn17"><code>@saryn17</code></a>, <strong>Ryosei Sato</strong> and <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11361">vitest-dev/vitest#11361</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11362">vitest-dev/vitest#11362</a> <a href="https://github.com/vitest-dev/vitest/commit/1c3888bce"><!-- raw HTML omitted -->(1c388)<!-- raw HTML omitted --></a></li> <li>Keep config of browser-consumed environments - by <a href="https://github.com/kasperpeulen"><code>@kasperpeulen</code></a> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11378">vitest-dev/vitest#11378</a> <a href="https://github.com/vitest-dev/vitest/commit/aafc0996f"><!-- raw HTML omitted -->(aafc0)<!-- raw HTML omitted --></a></li> <li>Ignore page crash while cancelling - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11386">vitest-dev/vitest#11386</a> <a href="https://github.com/vitest-dev/vitest/commit/7c36748fa"><!-- raw HTML omitted -->(7c367)<!-- raw HTML omitted --></a></li> <li><code>toMatchScreenshot</code> uses wrong reference on retried tests - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11393">vitest-dev/vitest#11393</a> <a href="https://github.com/vitest-dev/vitest/commit/c22aba992"><!-- raw HTML omitted -->(c22ab)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>cache</strong>: <ul> <li>Revalidate imports of cached modules - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11381">vitest-dev/vitest#11381</a> <a href="https://github.com/vitest-dev/vitest/commit/38f98855f"><!-- raw HTML omitted -->(38f98)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>deps</strong>: <ul> <li>Pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into <code>ERR_PNPM_TRUST_DOWNGRADE</code> - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11403">vitest-dev/vitest#11403</a> <a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977"><!-- raw HTML omitted -->(f6c9a)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Pass current equality testers to <code>expect.extend</code> asymmetric matchers - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11401">vitest-dev/vitest#11401</a> <a href="https://github.com/vitest-dev/vitest/commit/3e794a96b"><!-- raw HTML omitted -->(3e794)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Support Blob on jsdom 30.1 - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11379">vitest-dev/vitest#11379</a> <a href="https://github.com/vitest-dev/vitest/commit/6c49b7197"><!-- raw HTML omitted -->(6c49b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>pool</strong>: <ul> <li>Preserve unique pool ids when <code>groupOrder</code> is set - by <a href="https://github.com/mtorp"><code>@mtorp</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11392">vitest-dev/vitest#11392</a> <a href="https://github.com/vitest-dev/vitest/commit/50312ebb4"><!-- raw HTML omitted -->(50312)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>ui</strong>: <ul> <li>Split-pane handle overlapping iframe - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11221">vitest-dev/vitest#11221</a> <a href="https://github.com/vitest-dev/vitest/commit/f91db0dfd"><!-- raw HTML omitted -->(f91db)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vitest</strong>: <ul> <li>Remove root temp dir on close - by <a href="https://github.com/abhinav-phi"><code>@abhinav-phi</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11248">vitest-dev/vitest#11248</a> <a href="https://github.com/vitest-dev/vitest/commit/7c7119cf7"><!-- raw HTML omitted -->(7c711)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vm</strong>: <ul> <li>Do not optimize deps from index.html - by <a href="https://github.com/ezefernandezyf"><code>@ezefernandezyf</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11329">vitest-dev/vitest#11329</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11360">vitest-dev/vitest#11360</a> <a href="https://github.com/vitest-dev/vitest/commit/caf2887de"><!-- raw HTML omitted -->(caf28)<!-- raw HTML omitted --></a></li> <li>Don't reuse scripts across vite environments - by <a href="https://github.com/MO2k4"><code>@MO2k4</code></a>, <strong>Martin Oehlert</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11395">vitest-dev/vitest#11395</a> <a href="https://github.com/vitest-dev/vitest/commit/346d3896b"><!-- raw HTML omitted -->(346d3)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v5.0.2...v5.0.3">View changes on GitHub</a></h5> <h2>v5.0.2</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Bind <code>process</code> in case global is overwritten - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11343">vitest-dev/vitest#11343</a> <a href="https://github.com/vitest-dev/vitest/commit/0b79231ad"><!-- raw HTML omitted -->(0b792)<!-- raw HTML omitted --></a></li> <li><strong>detect-async-leaks</strong>: <ul> <li>Ignore <code>process.stdio</code> handles - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11333">vitest-dev/vitest#11333</a> <a href="https://github.com/vitest-dev/vitest/commit/0fd6b9790"><!-- raw HTML omitted -->(0fd6b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Fix <code>toMatchObject</code> with asymmetric matchers - by <a href="https://github.com/ShreeBohara"><code>@ShreeBohara</code></a>, <strong>Claude Opus 5</strong>, <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-5)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11100">vitest-dev/vitest#11100</a> <a href="https://github.com/vitest-dev/vitest/commit/42523289e"><!-- raw HTML omitted -->(42523)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Fix <code>Request</code> with <code>Blob</code> body on jsdom 28+ - by <a href="https://github.com/harshit-d3v"><code>@harshit-d3v</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11295">vitest-dev/vitest#11295</a> <a href="https://github.com/vitest-dev/vitest/commit/d1c3ecc93"><!-- raw HTML omitted -->(d1c3e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporter</strong>: <ul> <li><code>agent</code> to respect <code>--silent</code> - by <a href="https://github.com/Raj4478"><code>@Raj4478</code></a> and <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11271">vitest-dev/vitest#11271</a> <a href="https://github.com/vitest-dev/vitest/commit/5b95efb6d"><!-- raw HTML omitted -->(5b95e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporters</strong>: <ul> <li>Handle concurrent <code>createReport</code> calls - by <a href="https://github.com/7rulnik"><code>@7rulnik</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11278">vitest-dev/vitest#11278</a> <a href="https://github.com/vitest-dev/vitest/commit/e8e556ff7"><!-- raw HTML omitted -->(e8e55)<!-- raw HTML omitted --></a></li> <li><code>hanging-process</code> to use ESM entrypoint - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11316">vitest-dev/vitest#11316</a> <a href="https://github.com/vitest-dev/vitest/commit/4e91e5668"><!-- raw HTML omitted -->(4e91e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>spy</strong>: <ul> <li>Fix stack overflow when spying <code>Set.prototype.add</code> - by <a href="https://github.com/fengmk2"><code>@fengmk2</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11299">vitest-dev/vitest#11299</a> <a href="https://github.com/vitest-dev/vitest/commit/a0a939653"><!-- raw HTML omitted -->(a0a93)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/33cadea62e8763c455c7fca38d9ab1dda87c5f75"><code>33cadea</code></a> chore: release v5.0.3 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11409">#11409</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/346d3896b65c3c907174447a035807342799f346"><code>346d389</code></a> fix(vm): don't reuse scripts across vite environments (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11395">#11395</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977ad3363f796a737834572e54c6ad5c18"><code>f6c9a49</code></a> fix(deps): pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into `...</li> <li><a href="https://github.com/vitest-dev/vitest/commit/062c75d8b63519211d951d8293ea81b5a9e3c124"><code>062c75d</code></a> chore: fix standalone docs build, update exports maps (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11394">#11394</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/caf2887dee8987a60118d53933f6e9cabd6b3e2a"><code>caf2887</code></a> fix(vm): do not optimize deps from index.html (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11329">#11329</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11360">#11360</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/50312ebb4eca6a98f6d0b2b61d5d9d38cbbabcef"><code>50312eb</code></a> fix(pool): preserve unique pool ids when <code>groupOrder</code> is set (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11392">#11392</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/7c36748fad1eae9687312f2f7ceadce6ec88b5df"><code>7c36748</code></a> fix(browser): ignore page crash while cancelling (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11386">#11386</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1df16a4fa5bbee3f198c582fbd56689d8"><code>92ba7fc</code></a> fix: scope cache key generators to projects (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11281">#11281</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11301">#11301</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/38f98855fa9cd7fd376afb84094eba0fda256a74"><code>38f9885</code></a> fix(cache): revalidate imports of cached modules (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11381">#11381</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/b24585f08f2ea267746a2d6ca0e43edcbb29726f"><code>b24585f</code></a> fix: don't retry when <code>test.fails</code> expectedly failed (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11219">#11219</a>)</li> <li>Additional commits viewable in <a href="https://github.com/vitest-dev/vitest/commits/v5.0.3/packages/vitest">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a590ab769d |
Give managed agents persistent cryptographic identities (#15352)
Give agents persistent Ed25519 identities encrypted with the existing instance master key. Create keys transactionally for new agents and lazily before supported managed runs, expose public identities in the API and agent UI, and protect private material during runtime delivery and output persistence. Preserve identities in recovery backups while giving imported and development-cloned agents fresh keys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9f7057e122 |
feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use a harness, a model, and a credential to run tasks. > - Connections already store credentials and control who can use them. > - Custom providers also need an endpoint and a supported API format. > - A per-agent endpoint would duplicate credentials and access rules. > - This pull request stores routing on the connection and projects it into the harness. > - Isolated credentials and protocol checks keep the selected connection authoritative. ## Linked Issues or Issue Description Refs #37, #13083, #14104, #14565, #12692. #14016 is a reference only. This PR has its own schema, vault persistence, routing validation, runtime projection, and tests. None of the five commits in #14016 is an ancestor of this branch. We do not depend on or plan to merge it. #14967 addresses task-pinned account pools. #14422 addresses another provider integration. This is the first of two linked PRs. Merge this connection change before #15341, which refines agent setup and adds the qualification harness. The split keeps each review below 100 changed files. Provider catalog entries have usable setup forms in this PR. Local browser subscription sign-in is included. ## What Changed - Store non-secret routing metadata on AI connections. Vault provider API keys, including Bedrock bearer API keys. Reject general AWS access keys. - Enforce company, owner, human audience, agent access, connection status, and protocol checks before resolving credentials. Keep reconnect destinations immutable and retain connection identity during key rotation. - Project OpenRouter and compatible custom endpoints into Codex, Claude, OpenCode, and local Hermes. Carry these settings through both legacy and native runner transports. Clear conflicting host credentials and redact keys from diagnostics. - Preserve older OpenRouter accounts and native personal defaults. Add Google API-key accounts and migration `0306` for the two provider-default constraints. - Run local Claude and Codex subscription sign-in behind the existing browser sign-in card. Use private attempt homes and owner-bound completion instead of a copied terminal command. - Seed isolated Gemini authentication and preserve OpenCode workspace permissions. Keep the selected connection authoritative. The independent Gemini and Grok workflow fixes are in #15341. - Keep native OpenCode custom gateway keys in a runner-owned selected-model proxy; the harness config contains only a session-scoped capability. Honor runtime outgoing proxy and certificate settings. Preserve streamed responses and revoke the proxy on close or startup failure. - Allow ordinary members to connect native personal accounts before an agent exists. - Repair routed accounts from task cards using the saved provider destination, protocol, model aliases, and connection identity. - Add provider catalog definitions, model discovery, pinned logos, and complete native and routed setup forms. Allow a personal routed connection before a new agent exists. Keep endpoint authentication keys out of Hermes terminal children. - Recover cancelled or restarted browser sign-in with a clear restart action. Support no-auth endpoints without a vault credential. Add isolation and recovery regressions and runtime documentation. ## Verification - Updated with `origin/master` at `22a3ea341`. Migration `0306` follows the new master migration and passes migration and snapshot checks. - The integrated connection regressions passed 152 tests and 50 native OpenCode driver tests, including key-free child-shell configuration reads, authenticated/no-auth forwarding, streaming, model/path restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass. Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34 executable also completed a turn through the proxy against a local synthetic provider; the reusable key was absent from its config. A second real-executable smoke passed with an HTTPS CONNECT proxy and runtime-specific synthetic certificate trust. Certificate-file and certificate-directory regressions pass. - Task-card repair passed 48 tests, including OpenRouter, Bedrock, and custom gateway reconnect cases. UI typecheck and token gates passed. - The prior core regression set passed 133 tests across new-agent setup, provider forms, browser sign-in, routing projection, and connection authorization. Token gates and UI typecheck passed. - Full workspace typecheck and production build passed again after the latest integration and credential-proxy fix. The merged deterministic runner E2E suite passed 1,400 Vitest tests and 128 Node tests. - Full workspace typecheck passed on the prior linked combined implementation. Production build, Storybook build, 1,316 browser-harness Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the latest master integration and regenerated migration. This exact head passed 54 remote checks with four expected skips and Greptile 5/5; no review threads remain open. An unchanged server fixture had a random six-character issue-prefix collision on its first attempt. All 245 tests passed locally and the single CI retry passed. - A provider-free terminal check used the cited supported Hermes source and dummy keys. Gateway and OpenRouter terminal children could not read the selected key. - The broad local Vitest attempt passed 15,442 tests but was not green. It had an embedded-Postgres startup failure, an HTTP logger timeout, an origin socket error, and a browser cancellation wait timeout. The cancellation wait was corrected. The relevant connection tests and the full origin test file passed separately. Latest-head CI must pass before merge. - Prior credential-backed acceptance exercised task creation, tool use, artifact delivery, completion, and context-dependent follow-up. Claude legacy and native runners passed Bedrock with `us-east-1` and `us.anthropic.claude-sonnet-4-6`. - Historical local qualification retained 43 passing API/gateway cells out of 46. Those attempts span earlier builds. They do not qualify this exact commit or staging. All subscription combinations and staging remain unqualified. - Verify native subscription and API-key setup. Connect a regular provider catalog row. Verify an incompatible harness and a changed reconnect URL are rejected. Use #15341 for the complete browser campaign. ## Risks - Migration `0306` changes two check constraints. It preserves rows and is safe to reapply. It takes normal constraint-change locks. - Credential projection touches several harnesses. CLI upgrades can change provider configuration and session behavior. - The native OpenCode proxy adds a loopback hop, pins requests to the selected model, limits request bodies to 16 MiB, rejects redirects, and expires at session close. It prevents reusable keys in the child configuration; it is not an OS isolation boundary against a process debugger running as the same user. - Custom endpoints must be reachable from the agent environment. Saving a connection does not prove connectivity. Bedrock keys require rotation before expiry. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Provider overloads and an unresolved follow-up timeout also affect live Gemini qualification. We have not patched the installed CLI or marked those cases as passing. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS identity, arbitrary auth headers, and custom routing for other harnesses are excluded. - These PRs do not establish production or staging qualification for every provider and login method. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
22a3ea3414 |
Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Public MCP lets people use their organization from an external assistant. > - Operators need a visible control for this experimental access. > - Hosted users should select an organization once and then approve its permissions. > - This pull request adds the setting, invitation-first setup, and browser or device consent. > - Connections provides a copyable invitation with public instructions that grant no access. > - Users reach browser consent from their assistant and return to inspect or revoke access. ## Linked Issues or Issue Description Builds on merged foundation #14846. This PR now targets master. Related settings convention: #13905. **Current behavior** The preview uses an environment variable to enable MCP. Hosted consent repeats organization selection. Assistant access has no entry in Connections, so users must already know the endpoint and how to reach consent. **Proposed behavior** An administrator enables Settings → Experimental → Assistant connections (MCP). A hosted connection shows the selected organization and its icon, then asks for permissions. Requested write access starts checked when the user’s role permits it; the user can opt out before connecting. Direct instance connections show an organization picker with the first available organization selected. The selection stays fixed across refetches and still requires an explicit Connect action. Connections includes Assistant Connection (MCP). Its setup page explains the canonical endpoint, client configuration, browser authentication, and connected access. It connects as the current person and does not select or impersonate an agent. **Reason and benefit** Operators manage access with the other experiments. Users select one organization, and both the UI and server enforce that choice. **Breaking changes** The old enable variable has no effect. Preview operators must enable the setting once. Apply the additive consent-request migration before deploying the tenant, then deploy the compatible Cloud broker. Existing direct requests and grants keep their behavior. ## What Changed - Simplify OAuth and device consent: show the Paperclip logo beside “Connect {client} to Paperclip”, fall back to “your assistant”, and show the identifying origin plus its favicon below, with the callback URL also visible when different. Remove the hosted-organization creation action. Default to the first available organization without silently changing it on refetch; preserve company restrictions and write opt-outs. Keep the button row contained on narrow screens. - Make Copy invitation the primary action, using the shared animated AgentSetupPrompt and a collapsed manual setup section with icon-labeled line tabs. Remove redundant link actions, copy-status text, the extra first-prompt well and revocation explanation from the setup page. Serve shared version-aware HTML and Markdown instructions without private organization data. - Support guarded Client ID Metadata Documents alongside dynamic registration, and include authorization response issuer identification. - Add RFC 8628 device authorization with separately hashed codes, expiry, shared request quotas, persistent polling backoff and atomic redemption. Reuse human consent, role checks, scoped grants, audit and revocation. - Add CLI device login and a local stdio bridge. Store credentials separately with private permissions and serialize rotating refreshes. - Add device consent stories and five cold-start paid Product E2E cases with independent grant, configuration and durable-work assertions. - Add `enablePublicMcp` to the settings validator, normalizer, feature catalog, and toggle UI. Check it live for OAuth, tools, subscriptions, and event delivery. Keep connection management and revocation available while disabled. - Default the MCP origin to the existing auth public URL, with strict validation and an explicit override. - Persist the optional OAuth `company_id` restriction. Describe only that company and reject approval for any other company, even if the person belongs to both. Keep active-membership and role checks. - Show the Paperclip icon and a large organization icon during consent. Return the saved company logo through the company-scoped request response and reuse the standard fallback icon. Use the requested concise permission labels: “Read all of your Paperclip data” and “Allow write access and creating tasks as me”. Use concise permission copy, retain a compact client and callback-origin disclosure, and remove the footer link. - Default requested write access on for eligible roles. Preserve opt-out across organization changes and refetch, reset defaults for a new request, and submit read-only access when the request or role does not allow writes. Align the shared checkbox with its label. - Use organization wording in consent, management, settings, and walkthroughs. Keep the organization fixed for hosted requests and retain direct-instance choice. - Add an Assistant Connection (MCP) card to the Connectors catalog, a setup page in the app shell, and a return link from Experimental settings. Include Codex, Claude Code, OpenCode, and generic remote MCP instructions. - Read the live gate and canonical server URL through authenticated setup metadata. Show only the current person’s grants for the selected organization, refresh after consent, and support revocation. Surface catalog status failures with an explicit retry action; do not present them as an empty connection list. Opening setup grants no authority. - Start the eight guided chapters in Connections. Keep presenter notes and chapter controls around real product pages in the app shell. Explain the terminal, consent, delegation, retrieval, and revocation handoffs. Mark conversation examples as illustrative. Cover first use, client setup, connected, loading, and error states. Keep the existing consent and management stories. - Keep the paid-eval setup and browser helper aligned with the setting and consent button. ## Verification - Warm-standby integration fix `ec64ea05e`: public MCP ingress now follows the Cloud claim guard; MCP and discovery paths return 503 instead of SPA HTML while unclaimed. Event polling checks the in-memory claim before reading the persisted experimental setting. All 97 focused OAuth/Cloud tests and server typecheck pass, including new request and timer regressions for idle-before-claim and resume-after-claim behavior. Fresh review is 5/5 with no unresolved threads, and all security scans pass on this final head. All browser shards, typecheck, build, canary installation and other test groups passed on the first attempt. The unchanged Cursor sandbox default-command test timed out at 10 seconds; the exact test passed locally without edits in 587 ms. The single failed-job retry passed, with the original failure retained in workflow 37500711895. All 54 final-head checks pass on `ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs are intentionally skipped). - Final master integration `8457828fc`: merged foundation #14846 and current master, preserving the invitation changes and all 33 files from the two newer upstream changes. No migration renumbering was required. All 95 focused OAuth/Cloud integration tests, full recursive typecheck and token gates pass. All CI gates passed on that integration head; review identified the warm-standby issue fixed above. - Security-review fix `8c1d0b696`: commit shared global/per-source admission before outbound CIMD work, preserve failed-attempt receipts, and validate resource/scope before fetching. Added migration `0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive typecheck and production build pass. The security scanner passed that commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18 authorization requests sharing one proxy across two service instances use just two fetches. All 70 OAuth/metadata tests and server typecheck pass after that refinement. Final follow-up `8ebeae84c` reports admission-storage failures as retryable HTTP 503 instead of invalid client metadata. Its regression proves no outbound request before admission and successful retry after storage recovers. All 71 OAuth/metadata tests and server typecheck pass. Final-head security scanning passes; Greptile is 5/5 with no unresolved findings. CI passed all browser shards, typecheck, build, token gates and canary installation. One unchanged adapter-utils bridge test raced a response-file write (expected a JSON error, received the safe file-changed error). The exact test passed locally without edits. The single failed-job retry passed; the original failure is retained in workflow 37490609192. All 54 checks now pass on final head `8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final integration above now targets master. - Integration with current master: preserved the new Connections source filters and pagination, kept all eval suites, and regenerated the consent/device snapshots as migrations 0303/0304. All four MCP migration SQL hashes are unchanged from the staging versions. Full recursive typecheck and production build, 132 focused UI tests (including catalog filtering), 89 server authorization/settings tests, 26 migration tests, 120 eval calibration tests and token gates pass. Review follow-up `4820ce74c` also keeps active assistant grants in Installed, with pending/error recovery and revocation/company-isolation coverage. All 76 setup/catalog tests, UI typecheck and token gates pass after that fix. The unchanged signoff browser test timed out waiting for a heartbeat in CI at `4820ce74c`; the exact test passed locally without code changes, and the preceding CI head passed that shard. That same unchanged test failed at the reviewer stage in the next CI run. All five signoff tests passed three times locally (15/15), without test changes. All eight browser shards pass at final head `8ebeae84c`; no browser-test edits or failed-browser-job retries were needed. - Setup-page refinement at `9ab009178`: all 17 focused setup/consent tests pass, along with UI typecheck, production build, Storybook build and token gates. Browser exercised the shared prompt preview and client tab switching, and the updated InvitationCopied Storybook interaction checks its clipboard fixture. All final-head CI checks pass at `9ab009178`, with no unresolved review findings. Deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524. Verified the actual page, tab switching and line styling, removed actions/copy, and successful native copy/paste of the complete Butter invitation into a local-only test field. The existing Claude grant was left intact. - Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates pass. UI typecheck and production build passed again at `4e4d5e4d9`; Storybook build and eval-helper typecheck passed for `28101cf91`. Follow-ups let the primary button wrap on narrow screens, preserve a distinct callback URL, and use only bundled icons to avoid pre-consent requests to client-selected sites. Browser-verified the real consent component in desktop and 320px mobile stories, including default selection, write access and preserved opt-out. Updated E2E heading/default-selection helpers. All CI checks passed at `dc8e9fd11`, with review 5/5 and no unresolved threads. The Butter preview publication needed a retry because npm initially accepted the DB package before making it visible; the retry succeeded and `dc8e9fd11` deployed. Verified a fresh, unapproved native Codex CIMD request on Butter: default organization/write selection, known-client heading and icon, distinct callback origin, and removed creation action. No grant was approved for this UI check. Prior paid runs below retain their exact source provenance; this UI-only follow-up did not rerun paid qualification. - Source-pinned paid matrix at `2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign `local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config, unavailable host, denied consent and reconnect/later retrieval, with independent configuration/grant/task/run/document assertions. Original failures, transcripts, source fingerprints and billing remain retained. - Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`. Latest `0637b9f1c` shares that same guidance across HTML, Markdown and manual UI after review; generated Markdown is verified byte-identical to the paid-evaluated version. Shared build, server/UI typechecks, token gates and 63 auth/metadata tests passed again. Every CI gate passed at prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads. - Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval calibration tests, server/UI/eval typechecks, token gates and Storybook build passed. Full recursive typecheck and production build passed during implementation; CI also passed them at `2992ef271`. - Local full-suite limitations: a large-file Git streaming test times out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh MCP reruns passed, and the corresponding CI groups passed. No claim that the local full suite is green. - Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD consent; device CLI reaches verification/consent. New grants await human approval. Existing local OpenCode retrieved a saved result in a fresh conversation through its previously approved grant. - Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received the exact copied invitation, read public setup, configured its server and started PKCE consent. Its shell command timed out; background retry reached the client's own callback deadline while approval remained pending. Latest instructions cover that handoff. **No completed Butter read/delegation/result retrieval is claimed.** - Cloud companion https://github.com/paperclipai/paperclip-cloud/pull/672 passes checks/review and deployed. Anonymous setup and device-protocol routing verified. Core `2992ef271` deployed successfully and the actual Claude web flow now reaches consent. Its extra JWT-bearer metadata is filtered to implemented grants; unsupported token grants remain rejected. Final `0637b9f1c` deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300; live HTML and Markdown both contain the final guidance. The superseded instruction-only build was canceled before deployment. This is a core-only staging preview; private Cloud plugins are omitted. ChatGPT web is signed out, so browser connector use is unverified. - Screenshot gallery begins at Butter's dashboard and distinguishes real setup/pending consent from local reuse and fixtures. It records the timeout finding. New persistent access needs human confirmation before the remaining actual-client acceptance work. - Manual path: Connectors → Assistant Connection (MCP) → Copy invitation → paste into assistant → configure and start authorization → sign in and approve → verify `paperclip_connection` → delegate → retrieve the saved report later. - Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md` and `doc/public-mcp.md`. ## Risks - Apply additive, replay-safe migration `0304_curvy_shadow_king.sql` before using device authorization. The public setup link carries no credential. Device codes and tokens stay private; neither sharing instructions nor installing a plugin authorizes access. - Apply additive migration `0305_chubby_vin_gonzales.sql` before deploying the shared metadata admission gate. It retains at most 60 short-lived, hashed-source receipts per instance and rejects excess attempts with 429. - CIMD metadata fetching is a new external-input boundary. It requires HTTPS, exact client ID and redirect validation, bounded responses and guarded DNS/network access. Client names remain self-reported. - Device support is per-instance. The central Cloud broker retains its existing grant support. Host installation and tool reload capabilities vary by client; instructions describe manual settings and restart requirements. - Consent names the registered client in its heading and displays its identifying origin below. Known-origin icons are bundled; all other origins show a neutral site icon without contacting client-selected sites. Client names are self-reported; the callback origin is the recipient check. The Cloud chooser also displays the original client and receiving origin before tenant handoff. - A user who accepts the preselected write permission can create tasks and comments. Task creation and comments can start or wake agents and use execution budget; the consent label uses the concise wording explicitly requested by the maintainer. Scope requests, role checks, and the final Connect action still apply. - Migration `0303_supreme_garia.sql` adds one nullable UUID column with `IF NOT EXISTS`. Requests without a company restriction keep the direct-instance picker. The binding stays recorded if its company is deleted; consent then fails closed. - Deploy tenant support before the Cloud broker sends `company_id`. Unknown or inaccessible organizations must never fall back to a different company. - The setting defaults off. Disabling access does not cancel work already delegated. Existing tokens and unexpired subscriptions can resume when enabled again; revocation remains separate. - The catalog entry is visible for discovery while the feature is off. Setup instructions, OAuth, and tool execution remain gated. No access is granted by viewing the entry. - Assistant sign-in starts in the external client so it owns PKCE and callback state. Client command syntax can change and links to official setup documentation are included. - An authenticated instance and valid public URL are required. Hosting, paid execution, and store publication remain separate rollout steps. ## Model Used OpenAI GPT-6 in Codex, with tool use and code execution. The exact serving model version and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks pass; unrelated local full-suite timeouts are explicitly recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e34abee670 |
feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path > - Paperclip gives teams durable tasks, agent execution, budgets, and approvals. > - People also use assistants in Codex, Claude, and other MCP clients. > - Those assistants need a scoped connection that preserves the person’s permissions and attribution. > - Delegating a task must not turn the assistant into the assigned agent. > - This PR adds opt-in user OAuth, ten first-party tools, browser consent, and workflow packages. > - Paid product evals verify the resulting tasks, documents, attribution, retries, and access boundaries. > - The team keeps working after the assistant conversation ends. ## Linked Issues or Issue Description **Problem or motivation** A person cannot connect an external assistant to an existing team through browser consent and safely delegate durable work as themselves. **Proposed solution** Expose an opt-in `/mcp/paperclip` endpoint with individually described first-party operations. Bind every connection to a person, client, company, resource, and scopes. Reuse domain authorization and scheduling. Package shared team-review, delegation, and follow-up workflows for OpenAI/Codex and Claude. **Alternatives considered** Related PRs #9393 and #12549 cover earlier remote MCP and board-operator approaches. This change uses user OAuth and a bounded public catalog. It does not expose a generic executor, operator administration, static shared board credentials, or external agent execution. Registry listing work in #9851 is a separate distribution step. **Roadmap alignment** This maintainer-requested implementation extends the governed MCP gateway, activity attribution, durable work products, and hosted deployment direction in `ROADMAP.md`. It implements the first release of the saved design plan; external agent participation and granted third-party tools remain later releases. ## What Changed - Add MCP 2.0 discovery and task status/comment/document Events on the same authenticated endpoint. Persist subscriptions and delivery receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt callback material, recheck permissions/Cloud membership, and bound retries/expiry. Older MCP clients keep their existing tools. - Add discovery, dynamic client registration, S256 PKCE, resource validation, rotating refresh tokens, revocation, and company consent. Store credentials as hashes and recheck membership at execution. - Add tools for connection identity, agents/projects, task search/read/create, human comments, documents/deliverables, and pending-approval links. Preserve current domain permissions and scheduling. - Add durable mutation receipts across reconnects. Matching retries replay results; uncertain outcomes keep the same request ID and require inspection. - Add consent and connection-management pages, OAuth log redaction, shared plugin workflows, and separate OpenAI/Codex and Claude package outputs. - Add eight paid Product E2E cases across three models, independent durable-state grading, usage evidence, cleanup, and report integration. Add task-document guidance and regenerate the runner capability inventories. - Add migrations 0301 and 0302, the dated implementation plan, result notes, and direct-client setup instructions in `doc/public-mcp.md`. ## Verification - Merge integration `e180b1948`: resolved conflicts with current master, preserved both eval registries, regenerated capability catalogs, and regenerated migrations as 0301/0302 while keeping the original replay-safe SQL byte-identical. Local migration safety/snapshot tests (26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval catalog/grading tests (198) pass. Token and capability gates pass. Full recursive typecheck passed. Fresh Greptile review is 5/5 with no unresolved findings. CI is green on this exact head (55 successes, two intentional skips, one neutral result): one unchanged Cursor sandbox test timed out at 10 seconds, then passed locally in 856 ms. A single retry of that failed shard and the aggregate workflow passed. Merge remains blocked on the repository code-owner approval rule. Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55 successes, two intentional skips and one neutral result. [The earlier CI run](https://github.com/paperclipai/paperclip/actions/runs/36901592350) includes all test shards, browser tests, typecheck, build and canary dry run. Greptile was 5/5 on that commit with no unresolved review threads. GitHub still requires code-owner review under the repository merge rules; passing checks do not bypass that approval. Paid source fingerprints remain separate below and in the dated result note. - Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku 4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback, signature verification and report retrieval in a fresh conversation. A final Mini regression passes after the quota/status fixes. All evidence validates. Bounded tunnel startup retries occur before provider calls and remain visible; failed earlier attempts retain their original grades. - The earlier complete seven-case matrix passes **21/21**, with a separate **3/3** delegation regression. Two preceding matrices also passed 21/21 each. A complete 24-cell matrix including Events has not been run. [The dated results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain exact source fingerprints, failures, model IDs and partial costs. - Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass after merging master. Server typecheck passes after the final quota/status changes. Eval typecheck and all 892 eval-support tests pass. - All 33 real MCP/OAuth tests pass. The preceding combined MCP, redaction, private-address and DNS-rebinding run passed 129 tests; two later MCP regressions cover quota reuse and unchanged-status suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI then found a null checkout result in the existing concurrent-workspace path; logging now uses optional status access. All 12 closed-workspace tests and all 33 MCP tests pass after that correction. The exact-start event calibration exposed a timestamp gap; scanning now includes the subscription start, with all 33 MCP tests and server typecheck passing. These two narrow corrections follow the paid regression. - A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0 subscription/delivery, current membership loss, unsubscribe, legacy SDK tools, refresh and revocation. Its callback transport is a fixture with independent HMAC verification. The paid Events campaigns separately prove public HTTPS delivery. - Earlier component qualification passed UI 7,117 tests, CLI 502, shared 832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module boundaries, migration order and plugin regeneration passed. CI covers general/serialized suites, eight browser shards, runner checks, typecheck, build and canary dry run. - **Local full-suite limitation:** the earlier monolithic run was not clean. It encountered overlapping schema rebuilding, Mac database shared-memory limits and isolated CLI/fixture failures. Targeted reruns passed. The existing >32 MiB Git filename stress test still hit its 300-second Mac timeout. The additional serialized sweep stopped after 62 passing suites once CI passed. Original failures and partial logs remain; this PR does not claim a wholly green local monolithic run. - Local Codex CLI and Claude Code OAuth login and MCP SDK interoperability were verified. Public-store installation, actual ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted newcomer provisioning remain release gates. Enablement is moving to **Settings → Experimental → Assistant connections (MCP)** in the stacked follow-up [#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge both for the intended setup experience. This foundation branch alone still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set `PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and connect to `/mcp/paperclip`. Select a team and allow writes in browser consent. Configure an available agent and budget, then delegate and retrieve results later. For Events, rescan the deployed plugin catalog in ChatGPT Work Cloud; the host supplies its webhook credentials when the user asks to watch a task. See [the setup runbook](doc/public-mcp.md). ## Risks - Events are at-least-once and may arrive out of order. No replay cursor is advertised. Clients must refresh finite subscriptions, read current state and avoid comment feedback loops. Callback material uses the instance secrets master key; hosted subscriptions require the updated Cloud broker and are bounded to five minutes/the access proof expiry. - ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging subscription remain deployment gates. Local signed-webhook and paid model evidence does not claim those surfaces have been exercised. - Disabled by default. Merging adds schema and opt-in code; it does not deploy a public endpoint, publish a store listing, create a team, or start paid agents. - Migrations 0301 and 0302 are additive and idempotent. Their SQL is unchanged from the earlier preview numbers, so hash-aware upgrade reconciliation preserves prior staging applications. Normal instance upgrades must apply it before enabling MCP. - Task creation and comments can schedule paid agent work. Consent and tool descriptions disclose that effect. Revocation blocks future calls but does not undo delegated work. - Public deployments need edge rate limits and credential-safe logging. Internal dispatch is restricted to the closed catalog and carries a request-local verified actor. - Hosted onboarding requires the companion Cloud broker, encryption-key configuration, and tenant rollout. Self-hosted direct connections can use this PR alone. - Store acceptance and agent-mode participation are not claimed. Checked-in plugin endpoints are development defaults; rebuild packages for a real deployment before installation. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A more specific serving version and context-window size were not exposed by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`, `claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted/component checks; full local-run limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
202c2d307e |
fix(server): leave unclaimed warm Cloud databases idle (#15314)
## Thinking Path > - Paperclip manages work performed by AI agents. > - Managed deployments prepare empty applications before an owner claims them. > - Those applications start database pollers even though no company can have work. > - Health probes also query SQL, so idle databases cannot remain suspended. > - This pull request adds an explicit standby marker for empty, unclaimed Cloud apps. > - The existing signed, durable claim resumes normal processing without restarting the app. ## Linked Issues or Issue Description **What happened?** An empty, unclaimed warm application runs recurring chat, email, plugin, heartbeat, cleanup, and reconciliation queries. Its health route also opens the database. This prevents idle database compute from suspending. **Expected behavior** An explicitly marked unclaimed application should keep its HTTP process and sandbox provider plugins ready while leaving the database idle. A successful signed claim should resume normal API behavior and background processing. Claimed and self-hosted instances should keep their current behavior. **Steps to reproduce** Start an empty Cloud-managed application and leave it unclaimed. Observe database activity while repeatedly requesting `/api/health`. Before this change, periodic queries continue without company data. Related: #15153 reduces allocation during chat polling. This change suppresses polling only for explicitly marked, empty, unclaimed Cloud apps. ## What Changed - Add `PAPERCLIP_CLOUD_WARM_STANDBY=1`. Check company emptiness once after restoring the persisted Cloud runtime identity. Missing Cloud configuration, existing data, or a persisted claim leaves normal processing active. - Gate recurring database pollers with an in-memory predicate. Keep startup preparation and sandbox provider plugin loading intact. - Serve unclaimed health probes without session or database reads and report `warmStandby: true`. Serve standby pages/assets directly from the UI router, bypassing session, bearer, tenant, and dynamic handlers. Refuse API requests and all WebSocket upgrades before authentication can query SQL or seed company data. - Exit standby after the existing signed identity assertion commits. Normal timers resume at their next tick; a restart restores the claim even with stale provider variables. - Document the marker, readiness semantics, rollout checks, and rollback. ## Verification - `pnpm -r typecheck` passed. A final server typecheck also passed after adding tests. - `pnpm build` passed. - Focused standby, signed claim, restart, health, static/Vite routing, hostname, HMR, and live-events suites: 72 passed after the review fixes. Includes real HTTP upgrade admission before/after claim. - `pnpm test:run` was attempted locally; both superseded runs were stopped after encountering checkout/platform failures. A clean-checkout rerun eliminated ancestor skill-directory lookup failures. The company-skills/runtime-cache families encounter macOS read-only-directory rename failures (`EACCES`); all three company-skills failures reproduce on unmodified base `bf14f803d5`. The initial full run also reported one native runner API test failure; an isolated comparison on both revisions was blocked by local embedded PostgreSQL startup failures. The full [Linux CI run](https://github.com/paperclipai/paperclip/actions/runs/37423263993) passed on final commit `5016c415ea`, including all server and workspace test shards, browser suites, typecheck, build, and release canary. This is not a claim that the full local suite passed. - Isolated full server with local PostgreSQL: after startup and connection expiry, 70 health probes, 70 page requests carrying valid synthetic tenant credentials, and 70 rejected WebSocket upgrades over 70 seconds observed zero app database connections. The signed claim completed in 62 ms and normal polling resumed (475 database transactions over 12 seconds). Restart with stale provider variables restored the durable claim. The latency is local-only, not a provider wake measurement. - Apex review: **5/5** on `5016c415ea`, both earlier threads resolved, no open recommendations. - No live-provider test or production deployment was performed. An actual database suspension/resume canary remains required before enabling the control-plane switch. ## Risks - Standby health reports HTTP readiness rather than current database connectivity. The signed claim still requires a durable database write; claimed health checks retain the SQL probe and 503 failure behavior. - Pollers resume at their usual intervals. A suspended database may add claim latency. Validate the real provider before enabling the marker. - Startup preparation and sandbox plugins remain loaded. New plugins or background loops must respect the same standby contract. - The marker is off by default. Remove it or set it to `0` and restart to roll back. No schema migration or claimed-workspace inactivity policy changes. ## Model Used OpenAI Codex, based on GPT-6. The exact serving snapshot and configured context-window size are not exposed in this session. Assistance included source review, TypeScript changes, command execution, and PostgreSQL tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted tests pass; full local suite limitations are documented above, and full Linux CI is green - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f2715e02bb |
fix(native): surface model capacity errors and retry automatically (#15347)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs preserve provider events and results, then decide task state. > - Codex can fail a turn because the selected model is temporarily at capacity. > - Paperclip displayed a generic native-session failure and left the task in recovery. > - A known capacity failure needs a clear message and a durable delayed retry. > - This pull request uses committed terminal evidence, the status-effect ledger, and existing dispatch gates. > - The task can continue automatically without an unbounded retry loop or a model switch. ## Linked Issues or Issue Description Related: #13993 adds launcher capacity deferrals before provider startup. This change handles a committed native Codex `serverOverloaded` terminal after provider work has started. **What happened?** A native Codex turn ended with `codexErrorInfo: serverOverloaded` and “Selected model is at capacity. Please try a different model.” Paperclip saved the result, showed a generic failure, and required recovery instead of waiting and retrying. **Expected behavior** Show the capacity error directly. Schedule bounded retries after a short delay. Preserve saved work and honor current execution gates. **Steps to reproduce** 1. Run a native Codex task. 2. Commit a `turn.failed` event with `serverOverloaded`, then accept a failed result for the same turn. 3. Inspect the task error and its recovery state. The regression test reproduces this without a paid provider call. **Paperclip version or commit** Reproduced against master `16b7db35ffa0f9a95913c8cbdeea3d595435691f`. ## What Changed - Classify capacity failures from committed runner events and the pinned execution identity. Preserve accepted results and display the specific capacity error. - Atomically persist one scheduled successor with the status decision. Retry after one minute, then two minutes. Share the existing failure budget and stop after two automatic retries. - Reuse dispatch gates for ownership, task holds, dependencies, budget, and locks. Wait for predecessor execution, finalization, and cleanup before claiming a retry. - Preserve pending reviewer authority and suppress retries after reassignment or a successor claim. - Consume the failed run's resume receipt and delivered wake input. Rebuild ordinary continuation from the failed run so explicitly resumed tasks can retry without borrowing one-run authorization. - Label scheduled retries “Model at capacity.” Add a Storybook example and avoid duplicate punctuation in the existing retry card. - Add regression coverage and document the runtime contract. ## Verification - Targeted recovery and UI suites: 125 tests passed. Additional final cleanup and lock regressions passed. - Latest-head continuation and authorization regressions: 57 tests passed, including all four capacity integration tests against an isolated PostgreSQL database. - Rechecked all four capacity integration tests with the full runner's isolated `PAPERCLIP_HOME`, config, temporary directory, and serial fork settings: passed. - `pnpm -r typecheck`: passed. Final server typecheck passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Browser: checked the real Storybook card. It shows the capacity message, automatic retry time, and existing Retry now action. - `pnpm test:run`: attempted and restarted after an interruption. The resumed run started before the final review correction and was stopped after recorded workspace/native test failures and the five-minute Git streaming timeout already documented on master. Exit 130; no complete local full-suite pass is claimed. Current-head isolated capacity tests and the complete CI suite pass. - A combined local recovery-suite attempt also hit PostgreSQL initialization failures at this macOS host's global shared-memory limit (32 slots). The focused recovery suites passed separately; final-head capacity tests also pass with the full runner's isolated environment settings. - Latest-head GitHub checks: all 56 checks green, including general and serialized test shards, browser E2E, typecheck, build, Runner verification, and canary dry run. No merge conflicts. - Greptile: 5/5 on `813a470b8c18c05aeb7e31e63e573c3a3a2f5cac`, with zero unresolved findings after fixing the resumed-task continuation issue. ## Risks - Capacity retries can repeat a task turn after partial work. They start a fresh provider session with task history and wait for predecessor cleanup. They do not resume the failed turn. - Retries retain the configured model unless an operator changes configuration. Persistent overload consumes the existing failure budget and then needs an explicit retry or model change. - Only run-bound native Codex `serverOverloaded` failures qualify. Usage-limit exhaustion, model/account incompatibility, unbound text, and unknown provider failures retain their current recovery behavior. - No schema migration is required. ## Model Used - OpenAI GPT-6 through Codex. The exact model ID and context-window size are not exposed to this session. Used reasoning, repository tools, code execution, and browser verification. No subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9b3fe260ba |
fix(tasks): surface Codex ChatGPT model rejection (#15299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native Runner runs report provider errors and can also save a structured failure result. > - Codex rejects a model that the user's ChatGPT account cannot use. > - Master now diagnoses this rejection, but the task's compact run data omits its message. The thread can still call it a generic run failure. > - Users need the account restriction and a clear step to repair the model selection. > - This pull request shows the existing diagnosis as “Model unavailable” on the task. ## Linked Issues or Issue Description **What happened?** Codex returns HTTP 400 with `invalid_request_error` and the message `The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account.` Master now stores an actionable diagnosis for this rejection. The task thread still labels it “Run failed” and does not receive its error text in compact run data. **Expected behavior** Show the account restriction on the task and in the run error. Tell the user to choose a supported model or clear the task's model override before retrying. **Steps to reproduce** 1. Use a Native Runner Codex agent signed in with a ChatGPT account. 2. Select a model that produces the rejection above and start a task. 3. Let the runner save its generic failed result. Inspect the task's failure marker and recovery notice. **Additional context** Refs: #15304. That merged PR diagnoses the provider failure and preserves worker and review recovery rules. This PR adds its task-facing message and guidance without changing that diagnosis or those rules. Refs: #13134. That PR improves model discovery for ChatGPT accounts. This PR exposes the rejection when a configured model still fails at execution time. ## What Changed - Return a bounded model rejection message in compact issue-run data. - Show “Model unavailable” and model-change guidance in the task thread and recovery notice. - Add database and UI regression tests, including both the original rejection text and master's fixed diagnosis. Document the new failure message. ## Verification - Focused merged-branch validation: 279 tests passed across activity service, native provider failure observation and PostgreSQL integration, TaskChatThread, and ExecutionBlockerNotice. - `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` passed on the merged source tree. - Greptile reviewed conflict-resolution commit `1db39c4e636a92e2e76ba224b71c6f1e2f556c81` at 5/5 with no findings or inline comments. - All 55 check runs completed without failure on that commit. The two optional Storybook jobs were skipped. The legacy Snyk status passed. The branch has no merge conflicts. - UI regression assertion: the task's failure marker says “Model unavailable” and retains the account restriction. It no longer says that this failure happened after a final response. ## Risks - Provider recognition and recovery are owned by the existing master implementation. This PR exposes only bounded error text for failed runs with `native_provider_model_rejected`. - Historical runs with the generic `adapter_failed` code are not reclassified. No stored run is rewritten. - Existing retry and reconciliation gates remain in place. No schema migration or model configuration change is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser inspection. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
858094ba81 |
Fix browser polling for unsaved agent chats (#15298)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat uses the shared task surface and browser panel. > - An unsaved chat has a temporary ID instead of a stored task UUID. > - The browser poll sent that ID to a PostgreSQL UUID query and returned a server error. > - This pull request waits for a saved task and validates task IDs before the query. > - Saved chats keep their browser tools and access checks. ## Linked Issues or Issue Description Refs #14331. That report describes the same draft-ID problem in email requests. This change covers the browser routes; it does not change email behavior. Opening an unsaved Agent Chat sends `chat:<agent-id>` to `/issues/:issueId/browsers` every three seconds. PostgreSQL rejects this as a task UUID. The draft is only a UI view model. The first send or upload creates its stored task. ## What Changed - Disable browser polling for unsaved chat IDs. Resume polling when the chat has a task UUID. - Return the existing task-not-found response for malformed task IDs before database access. Keep board authentication first and company and credential checks unchanged. - Test all six task browser routes, draft-to-saved polling, normal tasks, saved chats, and denied access. Keep canonical UUID v4, v7, and nil values queryable. - Document the draft and saved-chat browser contract. ## Verification - `pnpm install --frozen-lockfile` passed with pnpm 9.15.4 and Node 24.21.0. - `pnpm exec vitest run server/src/__tests__/browser-use-route-scope.test.ts server/src/__tests__/browser-use-connection.test.ts` passed all 35 tests. - `pnpm exec vitest run ui/src/hooks/useTaskBrowsers.test.ts` passed all 5 tests. - Independent review passed 14 hook and route tests plus both real task and saved-chat authorization cases. - `pnpm check:token-gates` and `git diff --check` passed. - Full workspace `pnpm -r typecheck` and `pnpm build` passed. - Full Linux CI passed on `dafff44a32`: 53 successful checks and two intentional Storybook skips. This includes all general and serialized test groups, UI and browser tests, typecheck, build, and the canary dry run. No CI retries were needed. - The full local `TMPDIR=/private/tmp pnpm test:run` started but did not complete. It was stopped after full Linux CI passed. Before the stop, three company-skills cache cases failed on macOS. A focused rerun reproduced all three as `EACCES` while renaming immutable cache staging directories. The test, company-skills service, and runtime cache files are byte-identical to the baseline previously reproduced on `e99854249c` for #15291. Other local groups were not reached; the full Linux CI run provides aggregate coverage. - Greptile scored the exact head `dafff44a32` 5/5 with no findings or unresolved review threads. ## Risks - Malformed task IDs now return 404 instead of reaching PostgreSQL and returning 500. The route does not map draft IDs to agents or tasks. - No schema, browser provider, task creation, or authorization policy changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and independent agent review. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full local run limitations are reported above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
81cd03b3b5 |
test(ui): keep initial reasoning with the saved reply (#15107)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The issue thread shows saved replies and agent run history. > - History that appears after the reply can move the page. > - Master now waits for initial run history before it shows the thread in #15228. > - This pull request checks that reasoning and the saved reply appear together in the first visible frame. > - The test protects that behavior from a later change. ## Linked Issues or Issue Description No public issue tracks this report. This is a bug regression test. **What happened?** The issue thread showed a saved answer before its reasoning and tool history loaded. The later history moved the page. **Expected behavior** The first visible thread contains the initial reasoning and the saved reply in order. **Steps to reproduce** 1. Open an issue with an agent run and a saved reply. 2. Delay initial run-history hydration. 3. Observe the saved reply before its reasoning appears. Related work: #15228 now contains the reveal gate. This PR adds a direct regression assertion for that gate. #14667 takes a different loading approach and remains open. ## What Changed - Add a saved agent reply to the initial-history test. - Assert that the native run's reasoning appears before that reply when the thread becomes visible. - Assert that the thread stays inert until history is ready. ## Verification - `vitest run ui/src/components/TaskChatThread.test.tsx ui/src/pages/IssueDetail.test.tsx` passed: 353 tests. - `git diff --check origin/master...HEAD` passed. - GitHub CI checks are in progress for the rebased head. ## Risks - Low risk. This PR changes only a test. - The production fix is already in #15228. This PR must not restore the older code from the previous branch head. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex CLI assisted with tool use and code execution. The exact model ID and context window were not exposed to this agent session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My remote branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing documentation changed) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e0b63e5a5 |
feat(connections): add experimental task-pinned AI routing (#14967)
## Thinking Path > - Paperclip manages AI agents and their work. > - AI Connections separate account access from models and harnesses. > - A pool must act as one connection while retaining each task’s account. > - Core must enforce member access and preserve session and recovery rules. > - A plugin supplies rotation policy without receiving credentials. > - This change adds durable routing and native connector setup and management. ## Linked Issues or Issue Description **Subsystem affected** AI Connections, Connectors, plugins, run dispatch, and session compatibility. **Problem or motivation** Operators need to rotate new tasks across saved accounts while each task keeps its account and session. Pool setup must fit the existing connector catalog and account workflow. **Proposed solution** Add an experimental router binding, a capability-gated plugin hook, and transactional task pins. Plugins declare native pooled connectors through `aiConnectionRouter`. Core hosts the existing-account picker, ordering step, and account settings. Related usage contract: #14936. Companion private plugin: https://github.com/paperclipai/paperclip-cloud/pull/643. **Roadmap alignment** This extends Apps and AI Connections. Core supplies generic enforcement and native connector UI; the private plugin owns rotation and quota policy. The prior duplicate search found no matching router implementation. ## What Changed - Add a router binding without changing existing concrete bindings. Keep the instance flag and new pools disabled by default. Require manual operator configuration. Show no routing toggle in Experimental settings on either open-source or Cloud installs, even after routing is enabled. - Persist company-scoped pools, one shared cursor per pool, and pins keyed by company, pool, agent, and task. Commit pins and cursor advances together with revision checks and bounded retries. Persist run-ID affinity before allocation. - Pass only authorized metadata and normalized usage to plugins. Core retains credential handling, member access checks, runtime qualification, and recovery evidence. Probe outside locks with a shared 15-second budget and freshness cache. - Resolve routing before credential preparation and backend selection. Preserve pins through turns, session resets, removed members, and quota waits. Retain admitted recovery after disable or uninstall. - Separate credential session epochs from token generations. Verified refresh preserves the epoch; reconnect and manual replacement change it. Include the credential slot ID in session and usage-cache identity, so reconnecting an indexed legacy account invalidates its old session even when both epochs are zero. - Validate pool member installations before accepting saved-agent bindings and recheck compatibility when the harness changes. Install only authorized members in the new-agent transaction and record their IDs in local activity. Pool membership cannot install a restricted shared connection. - Preserve pool bindings when agents hire teammates through either creation API or native caller runtime inheritance. Block stale manager credential references; retain explicit child authentication precedence and reject incompatible inherited pools. - Add native connector registration through plugin metadata. Reuse the Connectors catalog, setup header, account header, sidebar, dialogs, and usage display. Setup selects and orders saved connections. Advanced settings hold usage rules and member runtime defaults. New-account setup opens in another tab. - Use revision-checked pool archival from the Connectors catalog and account page. Keep task pins, cursors, recovery evidence, and underlying connections. Reject ordinary connection updates or removals that bypass pool revisions. - Add pool selectors, composer models, override notes, quota status, run details, activity records, and local run-log records. Keep session-adoption copy minimal. - Show **Used by** below the pool connections. List current company agents with shared avatars and profile links. Include paused agents; exclude terminated agents and agents using another pool. - Add Core stories for the generic connector workflow and runtime surfaces. Cloud stories reuse these production routes and tokens through a preview-only alias. ## Verification - Final head `73cb953bca30ed83e4505dd820edd9b5edffd28b`: full workspace `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass locally. - All 422 focused connector/settings/shared-contract/migration tests and all 156 database-backed AI connection, hiring, reconnect, and durable-routing cases pass (69 hiring cases rerun after the final auth-precedence fix). The merged shared contract retains connection instructions and pool metadata. The pool migration is generated at sequence 0299 after the latest upstream migrations; this PR makes no lockfile changes. - All four full-app Playwright tests pass on the final head after a cold restart and migration, against the installed private plugin and isolated database, with no route or pool-API mocks. They cover hidden routing controls after manual opt-in, native pool creation, ordering, membership edits, rename, paused defaults, enabling/save/refresh persistence, stale edits, cancellation/removal, preserved underlying accounts, unavailable routers, and Used by avatars and profile links. Exact command: `PAPERCLIP_CONNECTION_POOL_E2E=1 AI_CONNECTIONS_TEST_COMPANY_ID=a37b9625-5ecf-4e29-8081-04df3d6e7d6f AI_CONNECTIONS_TEST_URL=http://127.0.0.1:3108 pnpm exec playwright test --config tests/ai-connections-app/playwright.config.ts connection-pools.spec.ts`. - [Native setup, ordering, and management screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-6006976278) address the review follow-up. [Earlier selector, quota, and run-detail screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-5971537316) show the runtime surfaces. Core previews: `pnpm --filter @paperclipai/ui storybook`, then **Connectors / Pool host** or **AI Connections / Connection pools**. Cloud owns its host-backed plugin stories; both repositories’ Operator Setup Required story assertions pass. - Live acceptance used OpenAI/Codex and Anthropic/Claude ACPX, resumed both exact sessions after restart, preserved pinned accounts through explicit reset and controlled quota deferral/recovery, and committed only two allocations across fourteen runs. A later UI-created task test again rotated OpenAI then Anthropic and resumed OpenAI through follow-up/restart/quota recovery. That later Anthropic execution was blocked by its saved OAuth token expiring (provider 401). No live usage probes ran. - The full local `pnpm test:run` was attempted earlier and did not complete because of macOS embedded PostgreSQL bootstrap/shared-memory failures and the 40,000-file Git fixture timeout. The focused database suites above now pass; full-suite verification is provided by the split CI lanes. The preceding CI run had one runtime readiness timeout; it passes locally both alone and inside the larger runtime suite. That larger local suite also encountered an embedded PostgreSQL setup failure and two macOS temporary-path alias assertions; those two assertions pass with canonical TMPDIR=/private/tmp. All final-head CI checks are terminal green, including full general/serialized server suites, Runner checks, browser E2E shards, canary verification, build, and typecheck. Greptile is 5/5 on that exact head with no unresolved threads. ## Risks - The migration adds routing tables and a credential epoch column. Install the private plugin only with the compatible Core contract. - Routing and each pool require opt-in. Production distribution and fleet defaults remain unchanged. - Unknown usage stays eligible. Known pinned exhaustion waits; revoked access requires operator repair. - Legacy adapters require compatible members. Runner model and effort overrides remain limited by qualified backend support. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and browser testing. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes:` / `Refs:` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted suites; full-suite limitations are reported above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4857799a88 |
feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes. Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2c43b39167 |
refactor(ui): share task composer and project/worktree controls (#15091)
Use the shared composer for new tasks, including project and worktree selection, rich mentions and slash commands, and remembered task settings. Save project preferences after successful task updates. Show known agent models by name and use Default for unknown runtime defaults. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ad0c4f0767 |
fix(ui): keep mobile task output above the composer (#15282)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People read agent output in the task conversation. > - Mobile tasks scroll the document and keep the composer above the footer. > - The document body keeps a fixed height while the thread grows. > - The old observer misses late thread growth, and a layout scroll event can cancel follow mode. > - This pull request observes the thread and preserves follow mode until the reader scrolls up. > - The benefit is visible new output without repeated manual scrolling. ## Linked Issues or Issue Description **What happened?** New task output can move behind the mobile composer while the reader is already at the bottom. Late content layout does not resize the fixed-height body. A scroll event after layout growth can also clear follow mode before the resize observer runs. **Expected behavior** Keep new output visible above the composer while the reader follows the bottom. Hold the reading position after an upward scroll. Resume following when the reader returns to the bottom. **Steps to reproduce** 1. Open a long task conversation on mobile and scroll to the bottom. 2. Let a streamed row or delayed content grow after the parent renders. 3. Check whether the latest output stays above the composer. 4. Scroll up to read earlier output, then return to the bottom. **Paperclip version or commit** Reproduced by regression tests against master at |
||
|
|
6c36c07a4f |
feat(adapters): add GPT-6.1 Sol and refresh shared coding harness pins (#14942)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents run through coding-agent adapters and the native runner. Both use the same installed provider CLIs, model catalogs, and reasoning controls. > - OpenAI released GPT-6.1 Sol (`gpt-6.1-sol`) in Codex. Anthropic released Claude Sonnet 5.5. The static Codex, Bedrock, and OpenCode catalogs do not list these IDs. > - The shared provider pack pins Codex 0.156.0 and OpenCode 1.18.32. The evaluation image pins older Grok, Gemini, Kimi, Cursor, and GitHub CLI releases. Codex 0.156.0 has no bundled metadata for GPT-6.1 Sol. > - A model entry without a current harness, or a harness pin without its runner integrity checks, fails at run time. > - This pull request adds the verified model IDs and moves the harness pins, executable digests, controller checks, and image pins together. > - The benefit is that operators can select the current models, and the native and local adapters share one current CLI installation. ## Linked Issues or Issue Description Refs #13829 and #13838 (the September 22, 2026 model and harness refresh). Related pull requests: #14993 (merged October 5, 2026, superseding #14816) added the direct Claude Sonnet 5.5 entry and refreshed the Claude runtime to Agent SDK 0.3.286 / Claude Code 2.1.286. This pull request does not change the Claude runtime or the direct Claude model list; it keeps the #14993 pins and adds only the Bedrock Sonnet 5.5 ID. After #14993 merged, this branch was rebased onto `master` (October 5, 2026). The six overlapping pin regions (`docker/daytona-runner/Dockerfile`, `docker/daytona-runner/README.md`, `package.json`, `pnpm-workspace.yaml`, `packages/adapters/claude-local/src/index.test.ts`, `packages/paperclip-runner/src/backends/native-backend-factory.test.ts`) were resolved by keeping this pull request's Codex 0.160.0 and OpenCode 1.18.34 pins next to #14993's Claude 0.3.286 / 2.1.286 pins, taking the union of the Sonnet 5.5 model IDs in the Claude test, and merging both README paragraphs. The Sonnet 5.5 effort and CLI-gate lines in the Claude adapter were identical in both pull requests and merged without a diff. #14917 and #14918 reordered the Claude and Codex model lists earlier; the new entries sit where those ordering rules put them. Sources checked on 2026-10-02: - [OpenAI Codex models](https://learn.chatgpt.com/docs/models): GPT-6.1 Sol uses `gpt-6.1-sol`, supports reasoning efforts from Light to Ultra, and has Standard and Fast modes at launch. The page also records that `gpt-5.4` and `gpt-5.4-mini` retired from Codex with ChatGPT sign-in on August 31, 2026, and that `gpt-5.5` retires on October 14, 2026. Neither retirement applies to the OpenAI API. - [Codex CLI releases](https://github.com/openai/codex/releases) 0.157.0 through 0.160.0. The bundled model metadata in the 0.160.0 Linux binary contains `gpt-6.1-sol`. - [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview): Bedrock ID `anthropic.claude-sonnet-5-5`, released September 28, 2026. - [OpenCode releases](https://github.com/anomalyco/opencode/releases) 1.18.33 and 1.18.34 (fixes only). The OpenCode model registry lists both added provider-qualified IDs. - npm `latest` tags for `@xai-official/grok` 1.0.46, `@google/gemini-cli` 0.62.0, and `@moonshot-ai/kimi-code` 2.1.1. [xAI](https://docs.x.ai/docs/models), [Google](https://ai.google.dev/gemini-api/docs/models), and [Kimi](https://www.kimi.com/code/docs/en/kimi-code/models.html) list no newer coding models. - Cursor CLI 2026.10.01-e373342 is the version the official installer resolves. The pinned digest is the SHA-256 of the versioned Linux x64 archive. - [GitHub CLI 2.102.0](https://github.com/cli/cli/releases/tag/v2.102.0) (security fixes). The pinned digest matches the release `checksums.txt`. ## What Changed - Codex adapter: add `gpt-6.1-sol` to the model list, the Fast mode list, and the Ultra effort set. It is the first entry: #14918 orders the list newest version first, and its description notes the ChatGPT app lists GPT-6.1 Sol first. Update the adapter documentation text. - Claude adapter: add `us.anthropic.claude-sonnet-5-5` (Bedrock Sonnet 5.5) to the Bedrock catalog in the newest-Sonnet slot after Opus 5.5; `us.anthropic.claude-sonnet-5` moves into the older-Sonnet group, matching what `sortClaudeModels` from #14917 produces at runtime. Any Sonnet 5.5 ID (direct or Bedrock-qualified) now gets the documented `xhigh` and `max` efforts and requires Claude Code 2.1.284 or later on the CLI lane (the Claude Code changelog entry for 2.1.284 adds `claude-sonnet-5-5`). These two lines are identical to the ones #14993 merged, so the branch carries no diff for them. - OpenCode adapter: add `openai/gpt-6.1-sol` and `anthropic/claude-sonnet-5-5` to the static fallback catalog. - Codex runtime pin 0.156.0 → 0.160.0 in the root and workspace overrides, the runner package, the Codex ACP package patch, the qualified ACPX profiles, the Linux x64 executable digest, the Rust provider backend and its tests, the provider-pack manifest pins, the remote controller pins, the sandbox npm install spec, and the opt-in qualification scripts. - Remote Codex compatibility window: upper bound 0.157.0 → 0.161.0. The minimum stays at 0.149.0. - OpenCode runtime pin 1.18.32 → 1.18.34 in the runner package, the materialization script, the server and Rust qualified versions, the eval and live-session labels, fixtures, and the configuration label. - Evaluation image (`docker/daytona-runner/Dockerfile`): Grok CLI 1.0.46, Gemini CLI 0.62.0, Kimi Code 2.1.1, Cursor CLI 2026.10.01-e373342 with its digest, GitHub CLI 2.102.0 with its digest, Codex and OpenCode version probes, and the refreshed lockfile digest. The Claude Code 2.1.286 probe comes from #14993 and is unchanged here. - `pnpm-lock.yaml` is not part of this pull request. The repository's pull request gate rejects lockfile edits, and the refresh bot regenerates the lockfile on master (the same flow #13838 used). The Dockerfile `PAPERCLIP_RUNNER_LOCK_SHA256` default is the digest of the lockfile that `pnpm install --resolution-only --ignore-scripts --no-frozen-lockfile` (the refresh workflow's command) produces for the combined pins on the rebased branch (`e1856797…`); that lockfile differs from master only in the `@openai/codex` 0.160.0 platform packages, the `@anthropic-ai/claude-agent-sdk` 0.3.286 override that #14993 introduced (the open refresh-bot pull request #14872 carries that part), `opencode-ai` 1.18.34 with its Linux x64 baseline, and the `codex-acp` patch hash. - Documentation: runner README, runner compatibility doc, environment variable example, and a new `doc/adapter-model-audit-2026-10-02.md` with sources and deferred items. - Tests: Codex adapter catalog, server adapter models, Codex compatibility window, native session executor pins, runner package contract, OpenCode materialization, and UI effort options. Unchanged on purpose: Claude Agent SDK 0.3.286 / Claude Code 2.1.286 (already on `master` from #14993), ACP bridges (`acpx` 0.13.1, `claude-agent-acp` 0.73.0, `codex-acp` 1.6.2; newer upstream releases need a separate qualification), the native Grok runtime 1.0.13, Pi 0.84.2 / 0.87.1 (the Pi 1.0 runner stack covers it), and Hermes 0.19.0 (current). `gpt-5.4` and `gpt-5.4-mini` stay in the picker because the OpenAI API still serves them. ## Verification Run on Linux x64 with Node 25.9.0 and pnpm 9.15.4 after `pnpm install --no-frozen-lockfile` (the refreshed lockfile stays local; see above). The results below were re-run on the rebased head (October 5, 2026) for the suites the conflict resolution touches; the other rows are from the original run and are covered by CI on every push: - Rebased head: `packages/adapters/codex-local` 482 passed; `packages/adapters/claude-local` 340 passed, 4 failed (`execute.remote`, `test.probe`, `execute.acp-fallback`, `acp` spawn/env-hardening cases that fail identically on unchanged `master` in this host environment); `server` adapter-models + codex-runtime-compatibility + native-session-executor + adapter-registry 607 passed, 1 failed (the same adapter-registry override-pause case as before, also failing on `master` here); `packages/paperclip-runner` native-backend-factory + qualified-profiles 36 passed; `ui` codex-reasoning-effort + config-fields + model-utils 19 passed. Rust, full typecheck, build, and the Docker image are left to CI as before. - `vitest run` in `packages/adapters/codex-local`: 13 passed. `vitest run` in `packages/adapters/claude-local` (whole package, including the new Sonnet 5.5 gate and effort tests): see the latest CI run and the comment below. `vitest run` in `packages/adapters/opencode-local`: 48 passed, 1 failed (`runtime-config.test.ts` reads the host `PAPERCLIP_OPENCODE_PROVIDERS` variable; it fails the same way on the unchanged base). - `vitest run src/__tests__/adapter-models.test.ts src/services/native-runtime/codex-runtime-compatibility.test.ts src/__tests__/adapter-registry.test.ts` in `server`: 84 passed, 1 failed (`adapter-registry.test.ts` override pause test; it fails the same way on the unchanged base). - `vitest run` in `ui` for `codex-reasoning-effort`, `agent-setup-fields`, `config-fields`, and `ComposerRunSettingsPicker`: 25 passed. - `node --test test/acpx-codex-package-contract.test.mjs scripts/materialize-opencode-binary.test.mjs scripts/runner-protocol-eval-campaign.test.mjs` in `packages/paperclip-runner`: 23 passed. The package contract test verifies the installed Codex ACP executable digest and the 0.160.0 patch pin. - `vitest run src/drivers/acpx src/backends src/drivers/opencode src/live/live-session.test.ts` in `packages/paperclip-runner`: 626 passed, 5 failed, 1 skipped. The 5 failures (`installation-integrity.test.ts` `/proc/self/fd` module loading and one OpenCode answer-selection test) also fail on the unchanged base under Node 25; Linux CI runs Node 24. - `pnpm run test:opencode:qualification` in `packages/paperclip-runner` against the installed OpenCode 1.18.34 executable: passed. - `codex --version` from the installed pack prints `codex-cli 0.160.0`. The Linux x64 executable digest `12eb3e81…652aad` was computed from the `@openai/codex@0.160.0-linux-x64` archive after checking its registry `dist.integrity`. - `pnpm check:token-gates`: all gates clean. - `pnpm run typecheck:typescript` in `packages/paperclip-runner`: passed. Package typechecks ran one at a time; see the comment below for the server and UI results. Not run here, and needed from CI: - Rust tests and `pnpm -r typecheck` / `pnpm build` for the server (no `cargo` in this environment; the server typecheck prepares the runner vendor build). - The Docker evaluation image build and the real-binary Codex startup and session-resume probes (no Docker; the probes need the compiled `paperclip-runnerd`). The trusted CI runner workflow covers them. - Authenticated inference with any new model. This change is metadata and startup validation only. ## Risks - Codex 0.160.0 changes the bundled model catalog and app-server behaviour (authoritative provider catalogs, incremental running-turn tracking). The patched `codex-acp` 1.6.2 bridge is unchanged and declares `^0.148.0`; it worked with 0.156.0 under the same override. If CI probes show a protocol change, the pin can return to 0.156.0 by reverting this pull request. - The compatibility window upper bound moves to `<0.161.0`. Remote images with Codex 0.157 to 0.160 become accepted. Older images stay accepted down to 0.149.0. - Until the refresh bot lands the regenerated lockfile on master, the Dockerfile lockfile digest default does not match the committed lockfile. The trusted CI workflow computes the digest from its own resolution at build time, so this affects only a local build that passes no digest. - Existing saved model selections and effort settings are not changed. Agents on `gpt-5.4` or `gpt-5.5` with ChatGPT sign-in need a model change before the OpenAI retirement dates; that is documented, not enforced. - Rollout order: deploy the controller and runner from this change before promoting a sandbox image that carries these pins. Older controllers reject the new provider-pack pins. ## Model Used - Claude Fable 5.1 (Anthropic, model ID `claude-fable-5-1`), 1M context window, adaptive thinking, tool use. The model ran as a Paperclip agent through the Claude Code harness, performed the web research, edited the code, and ran the tests listed above. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Bender (Fable) <noreply@paperclip.ing> Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
b43073d11f |
feat(connections): sync and group accounts managed by aggregators (#15254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external tools. > - Aggregator gateways can expose accounts that users already connected upstream. > - The Apps catalog did not show those accounts or their current provider status. > - Separate cards and setup tasks also made account ownership unclear. > - This pull request discovers upstream accounts and groups them under one app card. > - Users can find connected apps while each provider keeps control of its accounts. ## Linked Issues or Issue Description **Subsystem affected** Connections across the database, shared contracts, server, and board UI. **Problem or motivation** Users cannot see which apps are connected through a saved aggregator gateway. Native and upstream accounts need one app card. Discovery must preserve company, user, gateway, and credential boundaries. **Proposed solution** Sync account metadata from Composio, Arcade, and supported Executor gateways. Keep upstream account management in each provider. Use source chips and search to browse the catalog. Preserve native setup and the gateway's existing access policy. **Alternatives considered** Creating a local executable connection for each upstream account would duplicate authorization state. Using an agent task for routine Composio setup would add an unnecessary step. The board now calls the saved gateway directly for that setup. **Roadmap alignment** This extends the shipped Connected Apps and MCP Tool Gateway features in ROADMAP.md. The duplicate search found no open PR for managed account discovery. Related work: Refs #13755, Refs #13941, Refs #14725, Refs #13855. Open PR #12906 covers adjacent toolkit routing work. ## What Changed - Add provider-neutral discovery, sync, and refresh APIs. Preserve the Composio API paths. - Cache observations by company, saved gateway, viewing user, and credential version. Retain stale observations after failed or incomplete scans. - Add optional Arcade account sync credentials in the vault. Discover Executor accounts through its supported inventory interface. - Group native and upstream accounts in one app card. Imported account menus open their provider. Gateway menus own refresh and sync setup. - Add Paperclip, Composio, Arcade, Installed, and All chips. Show 50 catalog entries per page. Keep connected accounts above discovery. Keep explicit provider searches scoped. - Simplify Composio app setup and refresh its connected app list on the gateway Permissions page. - Add a compact agent access card and task creation defaults for connection setup. Preserve explicit blocks and approval policies. - Add two replay-safe migrations, service and UI tests, Storybook journeys, and acceptance stories. ## Verification - Passed the repository typecheck, full build, token gates, and migration ordering check. - Passed the focused provider adapter, connection interaction, and catalog tests after rebasing onto master. - Passed all nine database sync and migration replay tests using a disposable database on the test-drive PostgreSQL cluster. Removed that database after the run. - Verified Arcade cursor pagination against its official Go SDK and passed all eight adapter tests, including short and incomplete pages. - Passed all 45 interaction tests after making the exact requested tools and their Allowed/Ask first permissions visible before granting access. Verified the compact card in Storybook. - Passed the complete UI suite on the final code: 683 files and 7,432 tests, including the corrected Composio destination assertions. Passed 130 focused tests for the UUID, management-link, and health-status corrections. - Passed 22 Composio setup/sync tests, 23 connection-intent service tests, and the connection migration test in separate disposable databases. Database startup alone was substituted; the suites exercised their real SQL and services. - Passed all 10 OpenAPI route checks and the full-stack connection-intent browser test, including scoped consent, agent continuation, and task completion. - The local full runner encountered embedded PostgreSQL startup failures on this loaded macOS host. The earlier in-flight run also held the pre-fix Arcade transform; a fresh run of the final provider suite passes. The final-head CI is queued during GitHub’s active Actions incident: https://www.githubstatus.com/. The previous run also lost several runners simultaneously; its real catalog assertion failures are fixed and the fresh complete UI suite passes. - Tested the real test-drive server in the embedded browser with a live Composio gateway. Detected Airtable and Circleback. Verified refresh progress, account rows, source chips, search scope, and 50-entry pagination. - Arcade and Executor coverage uses provider fixtures. Live credentials were unavailable. - Storybook builds successfully and includes grouped native/provider accounts, stale and unavailable discovery, optional Arcade setup, and mobile states. The acceptance document records the simulated and live coverage separately. - Greptile reviewed final commit `217b024c27b5933e773ce9419c4e92b1032042c6` at 5/5. All six review threads are resolved, security scans pass, and the PR has no merge conflicts. The outstanding remote checks are `ci / Select trusted runner` and `review`, queued by GitHub. They need to complete before merge. ## Risks - Provider response changes can break inventory discovery. Failed scans retain observations and show stale status. - Composio scans only the supported catalog and can take time. Large inventories run in the background with progress and a bounded lease. - Arcade requires a project API key and user ID when the gateway cannot supply them. This key is used only for discovery. - Executor discovery depends on the server's exposed inventory tools. Unsupported servers report unavailable discovery. - Cached account rows do not grant access or create executable connections. Gateway policies still govern tool use. Account deletion and per-app authorization remain upstream. - The migrations add tables and one nullable column. Replay preserves existing rows and company-scoped foreign keys. ## Model Used OpenAI Codex, based on GPT-6. The session does not expose a more specific serving model ID or context limit. Used reasoning, repository tools, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
55e0c895ef |
test(ui): finish sidebar focus cleanup before JSDOM teardown (#15255)
Finish deferred Radix focus cleanup before the sidebar test environment closes. Use React act and controlled timers, check the real unmount focus events for both menus, and assert no timer work remains. Validation: all 7,368 UI tests, repository typecheck and build passed. Hosted CI passed, including server tests, browser tests and canary packaging. Greptile 5/5 on the exact reviewed head; no unresolved comments. Local full-suite limitations and clean-base comparisons are recorded in the PR. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
cab4263dc9 |
feat(claude-local): add Sonnet 5.5 and refresh the qualified Claude runtime (#14993)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Claude local adapter lists models for users with a Claude subscription. > - The list needs Claude Sonnet 5.5 and its supported effort levels. > - Sonnet 5.5 needs Claude Code 2.1.284 or later on both execution paths. > - The qualified ACP runtime previously used Claude Code 2.1.280. > - This change adds Sonnet 5.5 and pins Agent SDK 0.3.286, which includes Claude Code 2.1.286. > - Users can select the model and run it with a qualified runtime. ## Linked Issues or Issue Description Refs #3936. This replaces #14816 because the maintainer integration cannot write to the contributor fork. Thank you to @SkilLab-Tech for the model support, runtime refresh, tests, and platform digest verification. This branch preserves both original commits: `8d2f3261af61a2ac1120e51e8a8618732ace543b` and `e68d1d002a3ed745f016fb11c50ac5a3c5a9ff8d`. Related work: - #14917 added Claude model ordering. This branch includes that merged change and resolves its conflicts with #14816. - #14942 updates the other models and harnesses. It remains separate. Its matching Sonnet effort and CLI-gate changes are identical. Both PRs merge with master. The second PR will need a rebase after the first merges because adjacent runtime-pin and test edits conflict. - #14954 is another Sonnet 5.5 change. It overlaps with the model additions but does not include the qualified runtime refresh. - #14039 makes the per-task effort picker model-aware. #3937 is also related to effort selection. The original author checked the [Claude Code changelog](https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md) and [effort documentation](https://platform.claude.com/docs/en/build-with-claude/effort) on 2026-10-01. ## What Changed - Add the direct `claude-sonnet-5-5` model and Low, Medium, High, X-High, and Max effort levels. - Require Claude Code 2.1.284 or later for that model on the CLI path. - Put Sonnet 5.5 after Opus 5.5 in the current-model group. Keep Sonnet 5 in the older-model group. - Retain the Sonnet 5.5 assertions and the model-order assertions in the server tests. - Pin Agent SDK 0.3.286 and Claude Code 2.1.286 across overrides, integrity digests, qualified profiles, Rust provider pins, and the Daytona version check. - Update the related adapter and runtime documentation. ## Verification Local verification uses the resolved source tree and pnpm 9.15.4. Model tests passed on Node 25.9.0. Runtime integrity tests use CI's Node 24.21.0. - Five focused Claude test files pass: 62 tests. They cover model defaults, model ordering, CLI gates, and remote execution probes. - Server model-list and UI setup tests pass: 30 tests. - The Claude adapter typecheck passes. - Runner integrity and qualification tests pass on Node 24.21.0: 78 tests. Four descriptor-loader tests fail on Node 25.9.0; all four pass on the CI version. - The runner package contract passes: 10 tests. - Full local typecheck stopped with exit 137 in the database package under the container's 4 GB memory limit. The production build reached the runner Rust build, then stopped because `cargo` is absent. - The full stable local Vitest run was stopped after all current-head CI test shards passed. It did not complete locally. The 180 focused tests listed above passed. - `git diff --check` passes. The branch changes 20 files against master. It has no lockfile or workflow changes. - The original author verified all three platform digests against registry integrity and ran the Linux executable. Its version was `2.1.286 (Claude Code)`. See #14816 for that evidence. - Greptile reviewed head `6db3d3f1` and gave 5/5 with zero comments. Both Superagent scans and Commitperclip pass. All current-head CI jobs pass, including build, typecheck, Rust, test shards, browser tests, and the canary dry run. ## Risks - The controller and provider pack must use matching runtime pins. Deploy them together. - CI owns `pnpm-lock.yaml`. The master lockfile refresh must resolve the SDK override. Refresh the Daytona lock digest with that lockfile. - Images built with Claude Code older than 2.1.284 need a rebuild before the CLI path can use Sonnet 5.5. - The runtime remains at SDK 0.3.286. This PR does not take the later 0.3.287 patch. - A live Sonnet 5.5 session and a Daytona image build are not part of the local verification. ## Model Used - Original work: Anthropic Claude Code, `claude-sonnet-5-5`. Review: `claude-opus-5-5`. The author reported `xhigh` effort, tool use, and code execution. The original context window was not reported. - Merge repair and PR preparation: OpenAI Codex, based on GPT-6, with tool use and code execution. The runtime does not expose the exact model identifier or context window in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Squash Attribution Keep these trailers in the squash commit to preserve the original author and AI attribution: ```text Co-Authored-By: Claude Code (Ivan) <SkilLab-Tech@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> ``` --------- Co-authored-by: Claude Code (Ivan) <ivan@skillab.com.br> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7498705642 |
fix(ui): stabilize mobile task reading and document navigation (#15228)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People read tasks and send instructions from phones as well as desktop browsers. > - The mobile footer let page text show through its labels, and small text fields made Safari zoom on focus. > - Small scroll changes made the footer switch direction, and its changing page padding moved the conversation. > - Task pages also showed comments before question cards and run history arrived, so the composer and reading position moved again. > - Document links also used native navigation, which reset the reading position or reloaded a task through its UUID URL. Desktop tabs were crowded in the mobile drawer. > - This pull request keeps navigation steady, opens documents in the mounted task, and gives mobile readers a full-height panel with a vertical tab selector. > - The benefit is a stable task view and smoother scrolling on mobile. ## Linked Issues or Issue Description **What happened?** The mobile footer was translucent. Safari zoomed when a person focused a small text field. The footer switched abruptly while scrolling. A large task could show saved comments, then move the page again when a question card or run history arrived. In a local test with delayed responses, a late question moved the mobile composer by about 374 pixels. Opening a plan from the feed could reset the view or reload the task through a UUID link. The mobile document drawer left part of the feed exposed above small desktop tab controls. **Expected behavior** The footer has an opaque surface and moves smoothly after deliberate scrolling. Text fields do not cause automatic focus zoom. A task shows its initial conversation and composer together at the final scroll position. Background refreshes keep the existing conversation visible. Document links open in the mounted task with its existing cache and reading position. Mobile documents fill the viewport, show a clear close button, and offer a vertical list of open tabs. **Steps to reproduce** 1. Open a task with many long comments and a pending question in iOS Safari. 2. Delay its interactions, activity, and runs responses by different amounts. 3. Reload the page and watch the conversation and composer move as each response arrives. 4. Scroll down and back up, including small direction changes and edge bounce. 5. Focus the task composer, search field, and new-task title and description. 6. Open a plan or another task document from the feed, including a link that uses the task UUID. Close the panel and check the reading position. 7. Open several documents on a phone. Switch tabs and close both active and inactive tabs. **Paperclip version or commit** Developed from `1c07b5903` and rebased onto `59015846a`. **Deployment mode** Built from source. Tested in an isolated local test drive with iOS 26.5 Simulator Safari and Chrome. A temporary local proxy delayed independent responses for the layout test. Related work found in the duplicate search: - Refs #14727. It made saved replies appear before supporting history. This PR keeps its parallel requests and narrows the tradeoff in favor of a stable first layout. - Refs #14667. This open PR takes a different approach with per-run placeholders and retries. This PR fixes the observed question/composer movement and mobile navigation behavior. - Refs #13095 and #13597. These earlier fixes added task scroll anchors and skipped transcript waits for scheduled retries. - Refs #6550. Earlier mobile board polish. - Refs #9467. This related open PR changes list and generic tab reflow. The task-pane selector uses a separate component. ## What Changed - Give the mobile footer an opaque semantic surface. - Set a base-size floor for editable text on touch devices to prevent Safari focus zoom. Preserve larger title text. - Share mobile scroll tracking between both layouts. Accumulate scroll distance, ignore edge bounce and changed document bounds, and update once per frame. - Use shared motion tokens for the footer and composer. Keep page padding stable and honor reduced motion. - Wait for the initial question cards, attachments, work products, activity, runtime selection, plan, and relevant transcript history before the first reveal. Skip scheduled retries and older runs outside the initial comment window. - Bound the first reveal to 15 seconds. A stalled supporting request leaves saved conversation and the composer accessible with an explicit loading notice. - Keep concealed mobile history from stretching the document. Keep the composer mounted but concealed until the same reveal. Keep both visible during later refreshes. - Route first and repeated same-task document clicks in place. Recognize UUID and identifier links. Preserve the thread history entry and feed position. Keep modifier clicks, downloads, external links, and classic document behavior. - Give the mobile task panel the full viewport and safe-area padding. Use a visible X and 44-pixel touch controls. Replace the horizontal tab strip with a vertical selector that wraps titles and supports keyboard focus. - Add four interactive Storybook states for a few tabs, long names, many tabs, and the last tab. Reuse the production selector and tab controller. - Add navigation and tab regressions, update first-reveal regressions, and document the behavior in `DESIGN.md`. ## Verification - 392 tests passed across the seven focused task-loading, scroll, mobile-navigation, layout, and composer suites. After review fixes, all 339 tests across the four affected suites passed, including stalled-loading fallback on mobile and desktop and the motion-token catalog. - All 442 focused document, tab, task-thread, and scroll tests pass. The final click-propagation cleanup also passes all 136 task-detail tests. UI typecheck, UI production build, and `pnpm check:token-gates` passed. - All four cases in `artifact-tab-arrival.spec.ts` and `text-attachment-tabs.spec.ts` pass locally, covering desktop and mobile selection, composer focus, document rendering, and downloads of the original bytes. - `pnpm --filter @paperclipai/ui build-storybook` passed. Open the mobile tab stories under `Prototypes/Task detail/Mobile tabs`. - A local diagnostic proxy measured cached plan content at about 250 ms after the first click. The HTML load count and task request count did not change. Feed scroll stayed at the same position. First, repeated, and UUID document links were tested at phone and desktop widths. Task-reference links close their preview before the document reader opens. - In Chrome at desktop and phone widths, the delayed-response task showed one complete reveal. The late question no longer moved an already visible composer. - In iOS Simulator Safari, verified the large-task reload, opaque footer, navigation hide/reveal, and search/new-task/composer focus without automatic zoom. - Full workspace typecheck and build passed. The updated UI also passes typecheck, production build, and token gates. - All 54 checks pass on the final commit `bf5c9914e61833e7cc8794a69d72cf8c7057b952` (two additional checks are intentionally skipped). Greptile reviewed that commit at 5/5, and all review threads are resolved. - Full local `pnpm test:run` was attempted but stopped after server-fixture failures. Embedded PostgreSQL startup failure reproduced in an isolated native-interaction fixture after five startup attempts. The broad run also reported a rapid Slack callback ordering test failure. These server paths are unchanged by this PR, and their CI shards pass on the latest head. The full local suite is not claimed as passing. - A localhost proxy stalled the activity response for 30 seconds. Chrome revealed the available conversation after the 15-second deadline at both desktop and phone widths, kept the composer accessible, and cleared the loading notice when the response arrived. ## Risks Slow initial history requests can delay the first conversation reveal by up to 15 seconds. If that deadline expires, late data can change the available conversation while a loading notice remains visible. The reveal waits only for runs in the initial comment window, and later refreshes do not conceal an existing conversation. The larger editable text can change line wrapping on phones. Mobile navigation and composer motion use shared tokens and respect reduced-motion settings. Mobile tab selection changes the control layout. Document links retain URL history while sharing the task reading position; other tasks and external links keep their normal navigation behavior. I checked `ROADMAP.md`. This is a fix for existing UI behavior. ## Model Used OpenAI GPT-6 in Codex. The runtime does not expose a more specific model ID or context-window size. The agent used reasoning, code editing, terminal tools, and Chrome and iOS Simulator testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
59015846ae |
fix(chat): keep dismissed task questions in the feed (#15229)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask users questions in task chats and Agent Chat. > - Agent Chat keeps unanswered questions as compact entries in the feed. > - Regular task chats still show a pending composer badge after dismissal. > - A page reload can also open a dismissed question form again. > - This pull request applies the same feed behavior to both chat types and saves dismissal per person and task. > - Users can continue the chat and return to the original question later. ## Linked Issues or Issue Description **What happened?** Dismissing a question in a regular task chat leaves a pending composer badge. The question can return after a page reload. The earlier feed behavior only applied to Agent Chat. **Expected behavior** A dismissed question stays in the feed. Its form and pending badge leave the composer. Reload preserves dismissal. Opening the feed entry restores the original question and draft answer. **Steps to reproduce** 1. Open a regular task chat with a pending question. 2. Select an option, then dismiss the form. 3. Reload the page. Check that the composer stays clear. 4. Open the question in the feed. Check that the draft is restored. 5. Submit the answer. Check that the answered receipt appears. **Paperclip version or commit** Reproduced on master at `a386a599983519eb1d399f8b770bfccdb2a74762`. **Deployment mode** Browser UI in local and authenticated instances. This change does not depend on the agent adapter. Related work: Refs #14613, which added the Agent Chat feed behavior. Refs #9141, which validates real answers on the server. Refs #11434, which tracks comment-driven changes to interaction state. This PR changes question presentation only. ## What Changed - Show compact unanswered question entries in regular task chats. - Exclude durable questions from composer pending counts and navigation. - Save dismissal in local storage per person and task. Merge the latest saved IDs so dismissals from another tab survive reload. A new question can still open its form. - Keep the original question pending and answerable. Keep approval and permission controls. - Run the question-history regressions in both chat modes. Add reload, new-question, user/task scope, and stale-tab coverage. - Share the real-component Storybook fixture. Add a regular task test drive and an interactive dismissal/answer scenario. - Update the planning-mode browser test to check a dismissed question in the feed after reload and on mobile. - Update the design rules, implementation spec, and preview instructions. ## Verification - 329 focused thread, composer, and interaction-card tests pass. - `pnpm -r typecheck` passes. The UI typecheck also passes after the review fix. - `pnpm build` passes for the full repository. - UI build, Storybook build, and token gates pass. The UI build and token gates were rerun after the review fix. - Greptile gives final commit `98f71f221` a 5/5 score. The current-head check passes, and there are no unresolved review threads. - All 56 current-head checks are terminal: 54 pass and two conditional Storybook jobs skip. The CI run includes the full test shards, build, typecheck, browser tests, aggregate verification gate, and package canary. - The stale-tab regression fails in both chat modes before the review fix and passes after it. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/planning-mode-visual-verification.spec.ts` passes against a throwaway local server. It checks dismissal, reload, task navigation, and desktop/mobile planning controls. - The duplicate local `pnpm test:run` attempt was stopped after the full CI test suites passed. It did not complete locally. - Manual browser test: select Green, dismiss, reload, reopen, submit the saved answer, and inspect the answered receipt. Also send a new message while the unanswered question remains in the feed. The test drive uses real UI components with fixture response callbacks. - To repeat the browser test, run `pnpm storybook`. Open **Chat & Comments → Task Chat Unanswered Questions → Test Drive**. ## Risks - Dismissal is a browser-local preference. It does not sync to another browser or device. Clearing local storage removes it. - When local storage is unavailable, dismissal lasts for the mounted thread only. - Unanswered questions can accumulate in the feed. They stay pending until answered or resolved through the existing API. - No database migration, API change, or change to approval permissions. ## Model Used OpenAI Codex, `gpt-6.1-sol`, with xhigh reasoning, repository editing, code execution, and browser control. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a386a59998 |
Reduce repeated native completion guidance and preserve final replies (#15151)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native agents receive task constraints and completion tools from Paperclip. > - Completion tools already define the procedure for reporting a result. > - Repeated procedure text adds instructions to each full task turn. > - The final reply must still explain a blocker and link a saved document. > - This pull request removes repeated procedure text and keeps these visible outcome requirements explicit. > - A document receipt supplies the exact link, and stricter evals check the persisted reply and browser navigation. ## Linked Issues or Issue Description Refs: #14961. Related: #14948 and #15007. **What happened?** Native task envelopes repeat completion procedure text. A reduced envelope needs explicit final-reply requirements. The `write_document` receipt also lacks a canonical document link. **Expected behavior** Keep the completion tools as the source of procedure details. Require one accepted completion result before the final reply. A blocked reply must explain the reason, owner and unblock action. A document reply must contain a working link to the saved document. **Steps to reproduce** 1. Run the native assigned-skill document case and native blocker case. 2. Inspect the run-attributed provider final and its persisted comment. 3. Check the blocker explanation or open the final reply's document link. ## What Changed - Remove repeated completion procedure text from the native task constraints and backend instructions. - Keep explicit blocker and document-link requirements in full task turns. - Return a company/task-scoped `documentHref` from `write_document`. Preserve the link in the idempotent mutation receipt. - Repeat canonical links for this run's current saved revisions in accepted completion feedback. Give blocked providers final-response guidance for the cause, owner and unblock action. - Keep internal document/comment anchors when Markdown issue links load cached issue details. - Add a manual six-cell comparison suite with strict source, build, default-instruction and budget admission. - Capture eighteen shared runnerd RPC projections and six direct OpenCode HTTP projections across start, resume and continuation phases, using scripted local transports and no provider execution. - Apply v3 checks only to the manual instruction comparison; preserve v2 checks for the existing native completion suite. Check the actual persisted blocker reason and exact saved-document link. Click the rendered document link and check the original content marker in the classic document card or the new document tab. - Forward exact OpenCode finishing calls through the controller. Wait for acceptance, keep accepted feedback and concrete rejection text, and reject malformed responses. Preserve ordinary dynamic-tool response handling. - Settle the completion decision and tool response before mapping a racing idle/error/abort event or handling explicit close/interruption. Reject a concurrent finishing call before controller admission. - Add a provider-free regression through real runnerd, the OpenCode proxy and a fake provider. Reject the first completion, accept the corrected report in the same turn, and propose one result. - Keep all original verdicts unchanged. Treat replay under new checks as separate diagnostics. ## Verification - `pnpm -r typecheck` and `pnpm build` pass locally. - Native document-authority tests pass, including company/run authorization and idempotent replay. - Native runtime-context, backend and measurement tests pass. - Final-answer calibration, protocol scoring, source-admission and catalog tests pass. Wrong reasons, absent links and wrong link targets fail. - `pnpm test:e2e:runner:typecheck` passes. Discovery lists exactly six single-attempt local cells with the declared models. - Exported `prepareNativeInstructionPreflight` then `verifyNativeInstructionPreflight` pass on this clean committed source. They build locally and make zero provider calls. - Corrective live confirmation is incomplete. Source |
||
|
|
1c07b5903b |
feat: Chat leads the left nav, agent work beside chats, and a Combined Inbox + Task List flag (#15100)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The left nav is the main way people move between tasks, the inbox, and Agent Chat > - The nav has separate Inbox and Tasks rows that show overlapping work, and Chat is one row among many > - The side panel beside a chat shows the conversation's own artifacts, not the work the agent did > - People want Chat to be easy to find, and they want one place for their task views > - This pull request moves Chat to the top of Work, shows the agent's tasks and artifacts beside each chat, and adds an experimental flag that folds Inbox into Tasks > - The benefit is a shorter nav and a chat view that shows what the agent is working on. Both changes stay off until an operator enables them ## Linked Issues or Issue Description Refs #14706 (the secondary Agent Chat navigation this change builds on) Refs #14848 (reopen the last visited agent chat) **Subsystem affected** UI navigation (left nav, mobile tab bar), Agent Chat side panel, task list and inbox, and the company artifacts API. **Problem or motivation** Inbox and Tasks are two nav rows for overlapping work. Chat sits in the top group with no clear home. The chat rail lists only agents you already talked to, so you cannot see your other teammates there. The side panel beside a chat shows only the conversation's own artifacts. It does not show the tasks and files the agent made. **Proposed solution** With Agent Chat on, Chat leads the Work section and the rail lists every eligible agent. The chat side panel opens on the agent's tasks as cards, and the agent's artifacts are available from +. A new experimental flag, Combined Inbox + Task List, makes Inbox a set of views inside Tasks. **Alternatives considered** Rebuilding the inbox inside the task list. Instead, `/issues` hosts the existing Inbox component for inbox views and the existing task list for status views, so all inbox behaviour stays the same. **Roadmap alignment** Agent Chat (ROADMAP.md, "Agent Chat (including CEO Chat)"). All changes are behind experimental flags that are off by default. ## What Changed - **Agent Chat nav (streamlined shell):** Chat is the first row of Work, not a top-group row. Workspaces leaves the nav while Agent Chat is on. The mobile tab bar is Home · Chat · + · Tasks · Agents. The legacy shell keeps master's top-group Chat row. - **Chat rail:** `AgentConversationsSidebar` lists every eligible agent. The open chat is first, then conversations by recent activity, then the rest of the roster alphabetically. Terminated agents and agents you left are omitted unless you have history with them. The picker still marks only real conversations as "Open chat". - **Chat side panel:** a new default Tasks tab shows one card per task the agent created, was assigned, commented on, or acted on, newest first. It has the task list's filter popover and a sort control. **+ → Artifacts** shows the agent's artifacts as cards. Cards open in a new tab. Agent Chat off keeps the old Artifacts tab. - **Artifacts API:** `GET /api/companies/:companyId/artifacts` accepts `agentId`. The filter applies to documents, work products, and attachments by the agent each result is attributed to. The shared validator and the UI client carry the new parameter, and the OpenAPI entry picks it up from the shared schema. - **Combined Inbox + Task List flag (`enableCombinedInboxTasks`, off by default):** new card in Settings > Experimental. The Inbox row goes away and its badge moves to Tasks. A Views menu on `/issues` covers Mine, Unread, Blocked, Recent, Everything, All, Active, Backlog, and Done. Bare `/issues` opens the last-used view (default Mine). Links that carry `assignee`, `workspace`, `participantAgentId`, or `q` open All so the filter is kept. `/inbox/*` and `/issues/{all,active,backlog,done,recent}` redirect to the matching view. `/inbox/requests` stays its own page. - **Task detail breadcrumb:** the view key now decides the source, so quick-archive still works after a reload from an inbox view. - **Docs:** `doc/PRODUCT.md` and `doc/SPEC.md` describe the chat rail, the chat side panel, and the new flag. ## Verification - `cd ui && npx vitest run --no-file-parallelism src/components/chat src/components/task-side-panel/TaskSidePanel.test.tsx src/components/AgentConversationsSidebar.test.tsx src/components/Sidebar.test.tsx src/components/SidebarCompanyMenu.test.tsx src/components/Layout.test.tsx src/pages/AgentChats.test.tsx src/pages/InstanceExperimentalSettings.test.tsx src/lib/task-views.test.ts src/lib/issueDetailBreadcrumb.test.ts src/pages/Inbox.test.tsx src/pages/Issues.test.tsx src/App.test.tsx src/App.activity-routing.test.tsx src/components/MobileBottomNav.test.tsx src/components/CommandPalette.test.tsx`: 20 files, 356 tests pass. - `cd server && npx vitest run src/__tests__/company-artifacts-service.test.ts`: 13/13 pass, including the new agent-filter test across all three artifact sources. - The new rail test fails against the unmodified rail. - `pnpm check:token-gates`: all gates clean. - Manual: enable Agent Chat in Settings > Experimental. Open Chat. The rail lists all agents. Open a chat. The side panel shows the agent's tasks. Use **+ → Artifacts** to see the agent's artifacts. Then enable Combined Inbox + Task List. The Inbox row goes away, and Tasks shows a Views menu. - Snapshot baselines are intentionally not updated. See `doc/design/DECISION-SHEET.md`, "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks - With both flags off, the app behaves like master. The only exception is the API: it accepts a new optional query parameter. - With Agent Chat on, the rail can list many agents in a large company. It uses the agent list the app already loads, and search filters it. - The Tasks panel reads at most 200 recently updated tasks per agent and says so when it reaches the limit. The Artifacts panel reads at most 500 of the agent's artifacts. - Combined Inbox + Task List changes what bare `/issues` opens for people who enable it. Deep links with a task filter still open All. ## Model Used - Claude (Anthropic), model ID `claude-opus-5-5`, through Claude Code with tool use (shell, file edit, test runs). Extended thinking was enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: scotttong <squadbot000@gmail.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
db72ad4c73 |
fix(ui): explain branch artifacts without remote links (#15051)
## Thinking Path > - Paperclip helps people manage agent work and inspect its results. > - The issue artifact list presents those results as work-product cards. > - A branch can have no remote URL or structured metadata. > - The current card hides its summary and has no action in that case. > - This pull request shows the summary and provides an accessible details action. > - The card labels the absent link without making a false link. ## Linked Issues or Issue Description **What happened?** A branch work product without a URL showed a short title and an icon. It did not show the saved summary or provide an action. **Expected behavior** The card must show what the branch contains. People must be able to open the saved details. It must not imply that a remote URL exists. **Steps to reproduce** 1. Open the artifact list of an issue with a branch work product that has `url: null` and `metadata: null`. 2. Find the branch card. Its summary and details action are missing before this change. Related work: #12717 introduced rich work-product cards. ## What Changed - Keep linkless branch cards concise: show the title and no-link state. Keep the full summary behind an optional Saved description toggle. - Label a branch with no URL as having no remote link. Let a person open its details with a keyboard-accessible button. - Add a focused interaction test and two Storybook states for the no-link branch. ## Verification - `pnpm --filter @paperclipai/ui typecheck` passed. - Focused Vitest tests passed: 45 tests in two files. - `pnpm --filter @paperclipai/ui build` passed. - `pnpm --filter @paperclipai/ui build-storybook` passed. - `pnpm check:token-gates` passed. - The Storybook collapsed and expanded states rendered in Chromium. Both screenshots were captured. - Full workspace typecheck could not finish: the runner Rust package requires `cargo`, which is not installed in this environment. ## Risks - Low risk. The change affects card text and the action for records without a URL. It does not change the saved data or routes. - Long saved summaries on other linkless work products expand the card height after a person selects Details. ## Model Used - OpenAI Codex coding agent. The execution environment does not expose the exact model ID or context window to this agent. It used code execution and a browser to validate the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details (the managed execution workspace fixes this branch name) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (Storybook examples) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2a8a99e4a5 |
fix(ui): reveal completed thinking caret beside label (#15048)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task thread shows completed agent work and its reasoning. > - The completed work row places the disclosure caret at the far right. > - The caret stays visible even when the reader does not inspect the row. > - The row needs a quiet control that stays close to the work label. > - This pull request places the caret after that label and shows it on hover or keyboard focus. > - The change keeps the timestamp on the right and does not move content on hover. ## Linked Issues or Issue Description **What happened?** The completed agent work row shows a disclosure caret at the far right of the task thread. The caret stays visible when the pointer is elsewhere. **Expected behavior** The caret stays near the work label. It appears on hover or keyboard focus. The timestamp stays on the right. **Steps to reproduce** 1. Open a task with a completed agent run. 2. Look at the completed work row without hovering over it. 3. Hover over the row and then use the keyboard to focus its button. Related: #11772 addresses a different caret and composer layout. ## What Changed - Move the completed work caret next to the work label. - Keep its space reserved. Show it on hover or keyboard focus. - Add a test for caret position, visibility classes, and the open state. ## Verification - `node node_modules/vitest/vitest.mjs run ui/src/components/IssueChatThread.test.tsx` passes all 100 tests. - `node scripts/check-token-gates.mjs` passes all token gates. - `git diff origin/master...HEAD --check` passes. - CI passes on the latest head, including typecheck, build, tests, and verification. The local isolated worktree could not run typecheck because dependency installation failed while applying the existing `codex-acp` patch. ## Risks - Low risk. The same button still opens the work detail. Only the caret position and visibility change. - On a touch device, the caret is not visible without hover. The work row stays a button that users can tap. > This is a focused UI fix. It does not add a planned core feature from `ROADMAP.md`. ## Model Used - OpenAI Codex coding agent. The runner does not expose the exact model ID or context window. The agent used reasoning, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with the details available in this runner - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked public issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket ID - [x] I have run focused tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have checked documentation impact; no documentation change is needed for this visual fix - [x] I have considered and documented the risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1c3abf5075 |
fix(ui): use available space for composer labels (#15050)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task composer shows the next assignee and model before a message starts work. > - Fixed width limits shortened both labels when the composer had unused space. > - People could not read the selected agent and model even when the full text could fit. > - This pull request makes the capsule use the available composer width. > - It keeps the ellipsis when Plan mode or a narrow layout causes real space pressure. > - The benefit is clearer run settings without damage to the compact composer layout. ## Linked Issues or Issue Description **What happened?** The task composer truncated the assignee name at 6rem and the complete assignee and model capsule at 16rem. It did this even when the composer had more available space. **Expected behavior** The composer must show the complete assignee and model labels when they fit. It must use an ellipsis only when another control or a narrow viewport limits the available width. **Steps to reproduce** 1. Open a task composer with a long assignee name and a long model name. 2. Use a wide desktop layout. 3. Observe that the old capsule shortened both labels while unused space remained. **Paperclip version or commit** Current `master` before this change. **Deployment mode** Local dev (`pnpm dev`). **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific. This is a core UI bug. **Database mode** Not database-related. **Access context** Board. **Additional context** The Storybook cases cover a wide composer and a narrow composer with Plan mode. ## What Changed - Removed the fixed maximum width from the assignee and model capsule. - Removed the fixed maximum width from the assignee label. - Kept overflow ellipsis behavior when the parent row has insufficient space. - Added stable label selectors and focused component coverage. - Added Storybook cases for complete labels and Plan-mode truncation. ## Verification - `pnpm --filter @paperclipai/plugin-sdk build` - `pnpm --filter @paperclipai/ui typecheck` - `vitest run ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` - `node scripts/check-token-gates.mjs` - Storybook production build under Node.js 24.20.0 - Captured and inspected the two new Storybook cases. - The complete repository typecheck and build reach the Rust Runner step. This local environment does not have `cargo`. Hosted CI supplies the Rust toolchain. - The repository test suite reaches workspace-runtime tests. This local runner does not allow their required temporary home directories. Hosted CI supplies a writable test home. ## Risks - Low risk. The change only removes fixed width limits from one flex item. - A very narrow composer can still shorten both labels. This is the intended fallback. - The Storybook constrained case verifies that Plan mode and Send remain usable. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 through the Paperclip Codex runner. The run used reasoning, tool use, code execution, browser automation, and image inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
569c7203aa |
fix(ui): use latest issue runtime callbacks (#14863)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - The issue page uses an external-store adapter for comments. > - The adapter must keep one identity during an unrelated render. > - The adapter must also call the newest send and cancel functions. > - Passive effects update those functions too late for a synchronous runtime call. > - This pull request updates the function refs during render and adds a regression test. > - The benefit is a stable comment thread with no stale comment action. ## Linked Issues or Issue Description Refs #3678 **What happened?** The issue comment runtime can call the prior send function after a render. The callback refs do not update until passive effects run. **Expected behavior** The stable runtime adapter must call the newest function as soon as the render supplies it. **Steps to reproduce** 1. Render the issue runtime with one send function. 2. Render it again with a new send function and the same thread data. 3. Call the stable adapter before passive effects run. 4. Observe that the prior send function runs. **Paperclip version or commit** `6395cae072` **Deployment mode** Local development from source. **Agent adapter(s) involved** This is a core UI bug. It is not adapter-specific. ## What Changed - Update the latest send and cancel refs during render. - Keep the external-store adapter stable across callback-only renders. - Add a test for a runtime callback before passive effects run. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/hooks/usePaperclipIssueRuntime.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` ## Risks - Low risk. The change only updates two refs earlier in the same render. - The adapter identity and its data dependencies do not change. > This bug fix does not add or overlap with a roadmap feature. ## Model Used - OpenAI Codex with GPT-5. The exact deployed model ID and context window are not exposed. Reasoning, tool use, and local code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cc67d4e1d8 |
fix: preserve steering and recover stopped task conversations (#15015)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task conversation must let a user guide a running agent and resume stopped work. > - The active run owns its input protocol, even when the user changes the next model or effort. > - Queue delivery waits for a provider receipt, which must be able to persist during the request. > - A stopped startup also needs a clear user action that passes normal task admission. > - This pull request fixes steering delivery, makes queue actions immediate, and restores explicit continuation. > - The benefit is a responsive conversation that can recover without losing saved input. ## Linked Issues or Issue Description **What happened?** A queued message could change from Steer to Interrupt while a native run prepared. A steer request could wait on its own database lock and fail to deliver. A stopped startup could then leave the conversation without a working Retry or message continuation. Interrupt also waited for the server and showed a toast. **Expected behavior** The active run keeps its input protocol. Steer delivers input to that run. Steer and Interrupt clear the submitted queue rows and show the input in the conversation immediately. Failed delivery restores the latest queue with an inline error. An eligible stopped run offers Retry, and authenticated user input can start a fresh turn through normal task admission. **Steps to reproduce** 1. Start a task with a native Paperclip Runner. 2. Change the selected model or effort while that run prepares. 3. Queue a message and press Steer. 4. Observe the provider receipt and queue state during the request. 5. Stop a startup before its provider process begins, then try Retry or send a new message. 6. Repeat queued delivery with a legacy runner and press Interrupt. **Paperclip version or commit** Reproduced on the parent of this branch, `59c07ede7`. **Deployment mode** Authenticated private deployment. The fixes also cover local task conversations. Related work: Refs #12834, Refs #13354, Refs #13275. The open refactor in #13160 moves the same queue route; it does not fix the receipt lock or stopped-run continuation addressed here. ## What Changed - Select queue behavior from the active run's immutable dispatch and runtime resolution. - Leave the run row unlocked during provider acknowledgement, then lock and read it before merging the receipt. - Retain queued input if the target run stops during that wait. Keep inline delivery errors visible after empty queue updates. - Permit exact Retry and authenticated continuation after verified native startup cancellation. Preserve pause, approval, budget, ownership, and process-stop gates. - Carry undelivered native queue input into a fresh turn once the old execution is confirmed stopped. - Show Steer and Interrupt input in the conversation and clear submitted composer rows immediately. Restore the latest queue inline on failure. Remove delivery toasts. - Keep optimistic delivery stable across stale polls, empty queues, and paginated history. Preserve classic Interrupt error handling. - Document recovery and optimistic delivery behavior. Add regression tests across server, shared queue projection, and UI boundaries. ## Verification - Red-green regression tests reproduced the queue protocol, receipt lock, stopped-startup continuation, and optimistic delivery failures. - The focused server route, continuation, queue, and runner boundary suites passed during implementation. - The queue-route suite passes with 78 tests. The three complete conversation UI suites pass with 347 tests. - UI typecheck, production build, and `pnpm check:token-gates` pass. - Workspace `pnpm -r typecheck` and `pnpm build` pass. The local monolithic `pnpm test:run` is still running; remote CI verifies the complete suite on the latest commit. - All CI gates pass on `bd9031ad56abfcde13d13a13488c1b9217c2fd3a`, including the full test shards, runner verification, browser E2E, typecheck, release registry, and canary dry run. - Greptile reports 5/5 for that commit. Both review threads are resolved. ## Risks This changes queue display and explicit continuation admission. The UI must restore rejected delivery without losing other-session edits. The server must preserve concurrent provider result updates and must not resume a process whose stop is uncertain. Focused tests cover these boundaries. This change has no database migration. ## Model Used OpenAI Codex, an agent based on GPT-6. The exact runtime model ID and context window are not exposed in this session. Used reasoning, repository tools, code execution, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
59c07ede72 |
fix(ui): let a selected saved subscription be used without a tile click (#14996)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A user creates an agent in the New Agent flow and connects a model provider in the model step. > - When a saved subscription exists, the step shows it as the selected default and labels the primary button "Use saved subscription". > - The button stays disabled until the user clicks the provider tile, which adds no auth step on this path. > - This extra click is a papercut: the visible choice looks ready but does not work. > - This pull request lets a selected saved subscription enable the button without the tile click. > - The benefit is one less confusing step, with no change to new sign-ins or API keys. ## Linked Issues or Issue Description No public issue exists. The description follows the bug template. **What happened** In New Agent > Connect model, choose a provider that has a saved subscription (for example OpenAI with a saved default). The saved subscription shows as selected and the primary button reads "Use saved subscription". The button stays disabled until you click the provider tile. **Expected behavior** The button is enabled when a saved subscription is selected. A click on it reuses that subscription. **Steps to reproduce** 1. Have a saved OpenAI subscription. 2. Open New Agent and go to the model connection step for a Codex agent. 3. Do not click the OpenAI Subscription tile. 4. See that "Use saved subscription" is disabled. **Paperclip version or commit** master at `4abff286c`. ## What Changed - `AgentProviderConnection.tsx`: the `!opened` gate on the footer primary button no longer applies when `method === "subscription"` and a saved subscription is selected. All other disabled conditions stay (pending auth, loading saved keys, login readiness, API key input). - `connect()` never reads `opened`, so no auth step is skipped. New sign-ins and API keys still require the user to open the provider tile. - `AgentProviderConnection.test.tsx`: a regression test that renders the initial state (tile not opened, saved subscription selected) and expects the button to be enabled and to connect with the saved subscription. A guard test checks that a new sign-in still requires opening the tile. ## Verification - `cd ui && npx vitest run src/components/new-agent src/components/ai-connections --no-file-parallelism`: 6 files, 64/64 tests pass. - The new regression test fails on master and passes with this change. - `pnpm --filter @paperclipai/ui typecheck`: 0 errors. - `node scripts/check-token-gates.mjs`: all gates clean. - Component QA in Storybook (real component, provider verification simulated): before, the button is disabled and skipped by Tab until the tile click. After, a direct click and Tab+Enter both connect, with exactly one verification callback. Selecting a new subscription still requires opening the tile. - Full-app before/after QA with a disposable empty backend and a simulated provider response: same result. - Not tested: a live OpenAI sign-in. Repo-wide `pnpm -r typecheck`, `pnpm test:run` and `pnpm build` are left to CI. Manual check: open New Agent with a saved subscription, go to the model step, and click "Use saved subscription" without clicking the tile. ## Risks - Low risk. The change is one condition in one component. - If a future change makes `connect()` depend on the tile being opened for saved subscriptions, this path would skip that work. The comment at the condition records why the gate is skipped. - Open PR #14698 edits nearby lines in the same component. A small merge conflict is possible. ## Model Used - Claude Opus 5.5 (`claude-opus-5-5`), Anthropic, via Claude Code with tool use (shell, file edit, browser automation). Extended reasoning on. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: scotttong <squadbot000@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d7bdfc422c |
fix(ui): show agent avatar and align chat header controls (#14726)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent chat shows which agent receives each message. > - The chat header used initials instead of the agent avatar. > - The name and controls did not use the same center alignment. > - This change uses the shared avatar and centers the header content. > - Users can identify the agent and open its settings from the same row. ## Linked Issues or Issue Description **What happened?** The agent chat header showed initials instead of the agent avatar. The settings control did not align with the name and avatar. **Expected behavior** Show the agent avatar. Put the avatar, name, and settings control on the same horizontal center line. **Steps to reproduce** 1. Enable Agent Chat in Experimental settings. 2. Open a conversation with an agent that has an appearance set. 3. Check the avatar and settings control in the top bar. Related navigation work: #14706. ## What Changed - Use `AgentAvatar` in the conversation breadcrumb. - Update the breadcrumb key when the agent appearance changes. - Center breadcrumb labels that have an adjacent action. Keep task identifier baseline alignment unchanged. - Add regression checks for avatar props, appearance changes, and center alignment. ## Verification - PASS: affected Vitest suites, 129 tests. - PASS: `pnpm check:token-gates`. - PASS: `pnpm --filter @paperclipai/ui typecheck` on an isolated retry. - PASS: `pnpm --filter @paperclipai/ui build`. - PASS: browser checks at 1000 and 390 CSS pixels. Avatar, name, and settings control all have center Y = 29.5 CSS pixels. - Screenshots use the real header and avatar renderer with isolated context fixtures. No screenshot or fixture is part of this diff. - Full typecheck and build cannot complete because Cargo is not installed. - Full tests stopped with SIGKILL. The first UI typecheck also stopped with exit 137. These runs do not prove full-suite success. ## Risks - Low risk. The change only affects UI rendering. It changes no API or database contract. - Breadcrumbs with adjacent actions now use center alignment. Ordinary task identifiers keep baseline alignment. ## Model Used OpenAI Codex agent, with code editing, shell tools, and browser checks. The runtime did not expose the exact model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge No behavior or command documentation needs an update for this rendering fix. The execution environment requires the assigned branch name to stay unchanged. Pending review gates are not marked complete. Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5b8b2b38ca |
feat(apps): add Neon connection (#14980)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents reach external services through the Apps catalog. Each catalog entry is a reviewed `AppDefinition` that wires a provider's hosted MCP server into Paperclip's shared vault, grants, policies, gateway, and audit trail. > - Neon is a widely used serverless Postgres provider with an official hosted MCP server, but it is not in the catalog. Teams that run their databases on Neon must use the generic "connect your own MCP server" path, which has no branding, no guidance, and no project or read-only controls. > - The connector playbook requires a catalog entry for a provider like this: the hosted server supports dynamic client registration and bearer API keys, and the common definition fields can express every option Paperclip can serialize. > - This pull request adds the Neon definition, its official artwork, the research and permission-review ledger rows, documentation, and deterministic tests, without any provider-specific runtime code. > - The benefit is a one-click, governed Neon connection with optional project pinning and read-only mode, and a documented path to live qualification. ## Linked Issues or Issue Description **Problem or motivation** Neon is a common Postgres host for the applications agents work on, but Paperclip's Apps catalog has no Neon entry. Operators who want agents to inspect schemas, run SQL, or manage branches must paste the MCP URL into the generic remote-MCP flow, which gives no branding, no provider guidance, no project boundary, and no read-only switch. **Proposed solution** Add a catalog-only Neon connection built from the connector playbook: browser sign-in through Neon's dynamic client registration with the reviewed `read` and `write` scopes, or a customer API key sent as an Authorization bearer header. Both methods expose Neon's documented `projectId` pin and `readonly` switch as optional Advanced fields. Every discovered tool stays governed by the normal per-action policies. **Alternatives considered** A plugin was not needed because no custom UI, tables, workers, or webhooks are involved. A separate read-only method was not added because the playbook treats read-only switches as advanced fields rather than methods. Neon's repeatable `category` query filter was left out because tenant fields serialize lists as one comma-joined value, so it cannot be sent correctly without new runtime code; per-action policies cover catalog narrowing instead. **Roadmap alignment** This extends the existing self-serve remote-MCP connection catalog and does not overlap planned core work. ## What Changed - Added the `neon` provider to `scripts/ingest-app-definitions.mjs` (category, API-key placement, methods, tenant fields, guidance, warnings) and regenerated `packages/shared/src/app-definitions/neon.json` plus the generated registry. - Added the Neon row to the self-serve MCP research ledger with `dcr_or_api_key` auth and risk tier S4. - Added permission reviews for `neon/mcp-oauth` (explicit scopes `read`, `write`, taken from Neon's live authorization-server metadata) and `neon/mcp-api-key` (provider key), with evidence links. - Added Neon's official tile icon (`ui/public/brands/apps/neon.png`, copied byte-for-byte from the icon linked by neon.com) and the brand manifest entry. - Added prosumer gallery copy for the Neon card. - Added `doc/connections/NEON.md` (service involvement, endpoints, administrator setup, capabilities and policy, manifest, brand provenance, validation hook) and linked it from the connections README and the permission audit. - Tests: Neon definition shape, store visibility and artwork, URL recognition, reviewed scopes with scope-widening rejection, URL projection of the project pin and read-only flag, invalid project ID rejection, the connect form's API-key gating, and the pinned catalog counts. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts packages/shared/src/app-definitions-url.test.ts` — 34 passed. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts` — 367 passed. - `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx` — all passed. - `node scripts/check-app-brand-assets.mjs` and `node --test scripts/app-brand-validation.test.mjs` — passed. - `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter @paperclipai/server typecheck`, `pnpm --filter @paperclipai/ui typecheck`, `pnpm check:token-gates` — clean. - Manual: in a local instance, open Apps → Browse, confirm the Neon card and icon, open `/apps/connect?source=neon`, confirm both methods, the Advanced project pin and read-only toggle, and that Connect enables after an API key is entered. The operator completed a live connection against a Neon account on this build. - Live metadata probed on 2026-10-02: both `.well-known` documents at `mcp.neon.tech` return the recorded endpoints and scopes; an unauthenticated `initialize` returns 401 with `resource_metadata`. ## Risks - Low risk to existing providers: the change is additive catalog data plus tests. The generated registry only gains one import. - Neon's hosted server grants broad project and database management. The definition carries two warnings, recommends a development project, and keeps every write under the normal action policies; the read-only switch is enforced by Neon's server, not locally. - The permission-review ledger records live proof for both methods as not run; the full lifecycle checklist in `doc/connections/NEON.md` still needs a documented pass before the entry is considered fully qualified. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with extended thinking and tool use (shell, file editing, browser verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |