mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
master
2061
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
99a9de9940 |
fix(mcp): personalize assistant connections and hide revoked grants (#15411)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Assistant connections let people use their organization from another assistant. > - Setup begins with an invitation and returns to a list of connected assistants. > - Client-only labels obscure the authorizing person, while revoked rows clutter that list. > - This pull request shows the person’s avatar, names connections by owner and client, and hides revoked rows. > - It also shortens the copied invitation while keeping the approval instructions. ## Linked Issues or Issue Description Related: #14933 and #15380. **What existing behavior does this improve?** Invitation copy and assistant connection management, in the organization’s Connections screen and the account-wide management page. **Subsystem affected** Cross-cutting: shared MCP connection types, server profile projection, UI and Storybook. **Current behavior** Connections are labeled only with a client name such as “Codex.” Revoked connections remain visible. The invitation includes an extra sentence about agent identity. **Proposed behavior** Show the authorizing person’s avatar and use names such as “Dotta’s Codex connection.” Hide revoked rows after successful revocation and when loading retained revoked grants. Failed revocation leaves the connection visible. Remove the extra identity sentence from invitation copy. **Reason and benefit** Make connection identity clear and keep the list focused on usable connections. **Breaking changes** The connection response adds optional `user` metadata with name and image. Older servers remain usable. Names and revoked-row visibility change in the UI; OAuth client identity, authorization and audit retention remain unchanged. ## What Changed - Shorten the shared invitation text. - Project the authorizing person’s name and avatar through the user-scoped connection endpoint, without returning email or credentials. - Reuse the existing Identity component and owner naming conventions across the connection page, catalog card and account-wide list. - Hide revoked grants and remove a successfully revoked row from the shared cache, even if the subsequent refresh fails. - Update documentation, regression tests and production-page Storybook fixtures and revocation journeys. ## Verification - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook` and `pnpm check:token-gates` pass. Final UI type checks also pass. - Focused consent and connection UI tests: 31 pass, including company filtering, legacy metadata, custom client names, user-initial fallbacks, failed revocation, retained revoked rows and refresh failure after successful revocation. - Existing MCP regression suite: 6 pass; 71 database checks are skipped locally because embedded PostgreSQL cannot start on this machine. The added database check verifies user-profile isolation and retained revocation history; CI runs these checks. - Browser verification with Storybook fixtures: owner avatar and name render; revocation removes the selected row in both production pages, leaves other connections visible, and restores the empty state after the last revocation. - All 54 current-head CI checks pass, with two optional Storybook jobs skipped. CI includes database, browser, runner, typecheck, build and clean-install canary coverage. - Greptile reviewed commit `1368d79e1066b418712224378d89d64c2b11cb86`: 5/5, no actionable findings or unresolved threads. - The full local `pnpm test:run` was stopped after complete CI passed. Local database coverage remains unavailable because embedded PostgreSQL cannot start; no full local-suite pass is claimed. ## Risks Low risk. The additive profile field is optional for compatibility. Revoked grants are filtered only from management UI and retained for audit. Revocation failure does not hide an active connection. No authorization scopes, token handling or schema changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, command execution and browser verification. The exact deployment ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a9a20fb5c6 |
feat(security): add read-only customer-success inspection APIs (#15405)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a persistent identity and a verified active run for governed access. > - Customer-success inspection needs broad reads without tenant writes or secret access. > - Ordinary board login and database credentials give more authority than this task needs. > - This pull request adds a dedicated inspection API and strict managed-run authority. > - Cloud owns short grants, human approval, replay protection, and audit records. > - The benefit is inspectable access that an operator can disable immediately. ## Linked Issues or Issue Description Refs: #15352. This change reuses the persistent Ed25519 identity from that PR. **Subsystem affected** Server authentication, pure resource readers, and the shared wire contract. **Problem or motivation** One internal Paperclip agent must inspect customer onboarding work. It must not receive owner login or database credentials. Reads must not create customer sessions, memberships, activity, or read receipts. **Proposed solution** Add disabled-by-default run authority and versioned tenant inspection endpoints. Require strict instance-bound managed-run JWTs on the home instance. Require exact-operation, single-use Cloud permits on tenants. Execute a reviewed company-scoped catalog in read-only transactions. Cloud applies seven-day stack-age eligibility and human exceptions. **Roadmap alignment** This is access support for Cloud deployments and governed agent identities. Bot creation, scheduling, scoring, and reports are separate work. The maintainer requested this implementation. ## What Changed - Reuse existing public identity reads and managed private-key injection. Reject unprovisioned keys, paused agents, ended runs, legacy signatures, and wrong instances. - Mount `/api/customer-success/v1` before actor/session synchronization. Verify Cloud permits and consume them centrally before reading. - Add explicit company-scoped database readers and bounded instruction, skill snapshot, run log, workspace, and asset reads. Preserve existing redactions and file protections. - Add protocol, security, database immutability, and managed-agent qualification tests. Add deployment and rollback documentation. ## Verification - Full `pnpm -r typecheck` and `pnpm build` passed. Server typecheck passed after review fixes. - The broad local `pnpm test:run` recorded 14,277 passes and four failures in unchanged suites: two timeouts and two PR-metadata mock assertions. All three affected suites passed on isolated reruns (36 tests). The complete CI matrix passes at the final head, including every test lane, typecheck, build, runner checks, canary dry run, and the security scan. - Focused inspection, JWT, and existing identity tests pass. The catalog test compares every public database table before and after reads. - Inspection and route-contract tests: 22 passed. The coordinated test runs a real managed process agent against separate home/customer PostgreSQL databases and a PostgreSQL broker over HTTP. It proves wake through the existing controller, bounded binary file reads, single challenge consumption across replicas, concurrent grants with a two-connection pool, scoped SQL audits, append-only runtime auditing, one-year retention, and unchanged tenant data/files. - Run the coordinated test with `PAPERCLIP_INSPECTION_CLOUD_DIST` pointing at the sibling Cloud build. Normal unit runs skip that optional private integration. - Final-head Greptile is 5/5 with no unresolved findings. - No production deployment or customer inspection occurred. ## Risks - This adds an authentication boundary. Keep both feature flags disabled until coordinated staging and canary qualification. - Cloud support must deploy after this API. Unsupported tenants fail closed. There is no owner-login or database fallback. - Existing redactions remain the content boundary. Arbitrary pasted secrets in readable prose or files may remain. - Remote files and suppressed provider traces remain unavailable. Wake can cause normal startup/background writes; test those separately. - Disable Cloud policy first during rollback. Preserve existing identity material and Cloud audit history. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, and browser testing. The session does not expose a more specific deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; broad-run flakes are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
caf120105c |
test: prepare neutral native connection guidance evals (#15407)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need to discover connections, obtain consent, and continue from saved decisions. > - We want to reduce repeated instructions only when measured behavior supports the change. > - The existing decline tasks tell the model not to retry. One provider-decline check can pass without an explanation or an observed service counter. > - This PR adds neutral tasks and stricter saved-evidence checks before any connection instruction reduction. > - Production instructions remain unchanged. The new cells are configured, not live-qualified. ## Linked Issues or Issue Description Refs #15218. Refs #15389. **What existing behavior does this improve?** The Product E2E connection workflow evaluation and its instruction measurement provenance. **Current behavior** Some decline prompts supply the policy they intend to test. The provider-decline workflow does not require a saved post-decision explanation. Its old no-call check can use a missing fixture counter as zero. OpenCode has no connection cases in the original Everyday matrix. **Proposed behavior** Add an explicit-only suite with five connection stories on native Codex, ACPX Claude, and OpenCode. Require an explanation attributed by exact run ID after a saved decline. Observe the provider fixture counter. Preserve the original cases and grades. ## What Changed - Add fifteen configured cells with one attempt, twelve-minute deadlines, and verified 1,000-cent company and agent budget stops. - Remove procedure hints from the three new decline prompts. Keep a user-permitted explanation fallback and the existing positive controls. - Require saved decline state, one decision, unchanged connections, observed zero service calls where applicable, and a post-decision explanation from a successful run on the same task. - Add negative grader calibration and test the actual fixture budget payloads. Exclude the suite from default and generic selection. - Extend the existing full-catalog measurement source manifest with connection descriptions and schemas. Add an audit of fixed text, tool descriptions, returned instructions, and unqualified behavior. - Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and preserve its new Cursor suites. No production, credential, workflow, or lockfile change. ## Verification - Before rebase: Product E2E support passed 1,424 TypeScript tests and 128 Node checks. Six catalog measurement tests, repository typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery passed. - The full pre-rebase repository test run was stopped when master advanced. Its partial result is not a pass. - After rebase and the review correction: repository build/typecheck, Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128 Node checks, six measurement tests, and exact fifteen-cell discovery pass. The duplicate local full-suite run was stopped incomplete after about 20 minutes once complete CI passed; no local full-suite pass is claimed. - Review found that the initial grader read `runId` instead of public `createdByRunId`. A regression calibration reproduced both rejection of valid public comments and acceptance of the wrong alias. The fix uses the actual field and binds the evidence type to the shared `IssueComment` contract. A subsequent type-only import path correction passes Product E2E typecheck. - Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes [complete CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545): 51 successful checks and two intentional Storybook skips, plus separate Snyk success. Fresh Greptile review is 5/5 with the single review thread resolved and no new findings. The PR is clean and mergeable. - Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm test:e2e:runner -- --list --suite native-connection-guidance`. The measurement uses `PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json pnpm exec vitest run --project @paperclipai/server server/src/__tests__/native-procedure-measurement.test.ts`. Validation used pinned pnpm 9.15.4. - No paid provider campaign was started. There is no baseline/candidate behavior result for these new cells. - The audit records 654 UTF-8 bytes of fixed connection guidance. A clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41 supplied tools, 53,341 normalized bytes at start/resume, 50,949 at compact continuation, and a 48,195-byte authenticated OpenCode MCP catalog. These are byte counts, not tokens, bills, vendor-private prompt sizes, or savings from this PR. ## Risks - This is eval preparation. Passing support tests do not establish live model behavior or qualify an instruction reduction. - The explanation oracle checks attributed saved output. It does not prove cognition or arbitrary prose truthfulness. One saved interaction also does not prove the absence of repeated idempotent tool calls. - Successful new authentication and tool refresh, existing-connection agent grants, independent work while waiting, explicit retry after decline, and blocking when mandatory work remains still need separate coverage. - Notion setup decline does not execute a real Notion service. Positive service approval uses an already installed deterministic service; it does not qualify new connection creation. - The original historical failures remain unchanged. Future comparisons must freeze source, fixture, model, input, and grading controls and retain every actual attempt. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code execution. The exact serving model ID and context-window size were not exposed in this session; they are not inferred. No model provider was invoked by the eval suite in this PR preparation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a6306ba606 |
feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
faa8e452c7 |
fix(tasks): stop repeated reminders for historical questions (#15392)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task questions remain saved so a person can answer them later. > - The composer moves an old question into history after a newer human message. > - Native completion still treated every pending question as a required response. > - This caused agents to demand an old answer after the person moved work forward. > - This pull request shares the historical-question rule across context and task execution. > - Agents can finish verified work while the original question remains answerable. ## Linked Issues or Issue Description Refs #15229, Refs #14613, Refs #13130. **What happened?** An agent repeatedly asked a person to answer a question that had moved into feed history. Completion feedback explicitly told the agent to request a response. The saved pending row also blocked task completion. **Expected behavior** A question before newer human direction remains answerable in the feed. Its pending state alone must not require another reminder or stop completed work. A new input blocker, an approval, or a configured review stage must keep its gate. **Steps to reproduce** 1. Let a task agent create an ordinary question. 2. Dismiss the question and send a newer task message. 3. Let the agent finish the requested work and submit its completion report. 4. Observe a demand to answer the old question and a retained completion gate. ## What Changed - Add one company-scoped predicate for historical questions. Only later human comments count. Exclude agent attribution, run attribution, system notices, and untrusted source data. - Apply the predicate to completion feedback, native waits, finalization, commit validation, retry validation, blocked routing, and successful-run handoff. - Include question classification and guidance in heartbeat context and both native task-context tools. Add the guidance to fresh and resumed task prompts. - Replace automatic reminders for current ordinary questions with instructions to assess the real blocker, continue independent work, and withdraw obsolete questions through the existing API. - Preserve historical question rows during completion while cancelling their live native source runs through the existing post-commit and recovery paths. Let an authorized human answer them after completion without reopening work or creating a response wake. Preserve cancellation, current-input, approval, permission, credential, connection, and review gates. Add no dismissal storage or migration. - Document the rule in the execution contract and agent skill. Refresh generated capability source anchors. Add database-backed status, context, attribution, and governance regression tests. ## Verification - `pnpm build` passed. The server rebuild also passed after the lifecycle fix. Generated capability contract and inventory checks passed after the agent documentation update. - `pnpm -r typecheck` passed. Final `pnpm --filter @paperclipai/server exec tsc --noEmit` also passed after the last test additions. - Lifecycle and interaction regressions passed: 224 tests in 3 suites. Context and prompt tests also passed. Final historical-question cases passed (33 tests), native cancellation/recovery cases passed (6 tests), and the existing interaction/confirmation suites passed (72 tests). The full local `pnpm test:run` was attempted and stopped after more than two hours with unrelated fixture/hook timeout failures; it did not pass. All 52 successful GitHub checks are green on the latest commit, including the complete test matrix; no checks are pending or failing. Greptile is 5/5 and both review threads are resolved. - Regression cases cover the old-question/new-human-message sequence, final task status, answering after completion with no wake, live-run cancellation and crash recovery, current input blockers, both task-context tools, API context, timestamp precision, attribution boundaries, and protected gates. ## Risks - A later human task message makes an earlier ordinary question historical even if its input is still missing. The agent must identify the current blocker and ask only for information that still prevents work. - Browser dismissal remains a local preference. Dismissal without a later human message is not recorded by this change. - No schema change or data migration. Completion retains ordinary historical questions; cancellation still expires them. A completed task accepts historical answers only from an authorized human and creates no response-delivery outbox row. Governed requests retain their gates. ## Model Used OpenAI Codex, GPT-6, with repository editing, code execution, and browser diagnostics. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2d0c138122 |
Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also use assistants in Codex, Claude, and other MCP clients. > - The existing assistant connection can read work and create tasks or comments. > - It cannot edit tasks, exchange files, or manage normal agent and project settings. > - These operations must retain the person's permissions and Paperclip's execution rules. > - This pull request adds an explicit operation registry and separately consented configuration access. > - Assistants can manage work without receiving credentials or runner authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: assistant MCP, domain routes, consent UI, storage, and Product E2E. **Problem or motivation** A connected assistant cannot update tasks, maintain documents, attach files, or configure existing agents, projects, and skills. Users must leave the assistant for these routine actions. **Proposed solution** Add named tools and a restricted API registry to direct connections. Require separate configuration consent. Reuse domain routes and retry receipts. File uploads save the attachment when the byte transfer succeeds. **Alternatives considered** Arbitrary REST forwarding would expose administration and credential operations. Runner impersonation would bypass execution ownership. Separate upload completion calls add unnecessary client state. **Roadmap alignment** Checked ROADMAP.md and related MCP pull requests. This extends the human-authorized connection from #14933. It does not replace the runner or introduce agent impersonation. Companion Cloud routing and directory isolation: https://github.com/paperclipai/paperclip-cloud/pull/678. ## What Changed - Add task editing, finish/block, documents/revisions, deliverables, agent settings/instructions, projects/repositories, and skills/files. - Add an allowlisted API search/call registry with identical field restrictions, scopes, and retry identities. - Add unchecked configuration consent. Existing write grants retain their current authority. - Add hashed, expiring file transfer tickets and atomic upload receipts. No completion call is required. - Preserve company boundaries, human attribution, active native execution ownership, and execution review gates. - Add protocol/domain tests, consent stories, and eight paid Product E2E workflows. - Repair two CI fixture races: await cold route setup before assertions, and wait for asynchronously loaded connection copy. Both fixture suites pass (24 + 48 tests). ## Verification - Consent revision: one write-access checkbox controls requested work and configuration permissions in browser and device flows. All 16 consent tests, UI typecheck/build and token gates pass. Updated interactive stories cover default approval, opt-out and viewer restrictions. The paid browser helper uses the new exact label. Real GPT-5.4 Mini Product E2E passes 2/2 at `64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic retries, cleanup passed; $0.04149375 estimated assistant cost plus unpriced worker usage. Raw results, usage and source fingerprints are retained in the worktree. UI and Product E2E typechecks pass. - Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks pass; two optional Storybook checks skip. Greptile 5/5 on that head, no unresolved review threads. Final consent head `64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing and Greptile 5/5 with no unresolved threads. The unchanged Cursor sandbox test had one 10-second timeout, passed in local isolation, and passed its single CI rerun; the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37554106934). The existing chat retry-denial browser test had one visibility failure; its single rerun passes, and the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37542735691). - Full workspace `pnpm -r typecheck` and `pnpm build` pass at final runtime source `b2196fae1`. UI token gates pass. - 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests pass, including one-connection PostgreSQL OAuth and concurrent upload retries. - Paid Product E2E: all eight expanded cases qualified across Mini, Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell timed out before application startup. Final affected-case qualification passes 9/9 on all three models with grader v16, including the failed cell. Automatic retries disabled; failures, costs, source hashes and independent durable-state/file assertions are retained in [the verification record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md). - Actual Codex CLI, Claude Code and OpenCode clients completed local reads/mutations. Codex wrote a report, Claude updated it in a later conversation, and OpenCode uploaded/downloaded a file with matching SHA-256 and registered the attachment. Revoking the CLI grant rejects subsequent bridge initialization. - Butter staging is verified on final runtime `b2196fae1` ([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)). A fresh OpenCode workspace fetched the copied invitation, configured remote MCP, started OAuth and reached real consent with configuration unchecked. Invalid transfer tickets return 403 through Cloud. Human approval for the new persistent staging grant is pending; hosted task/file success is not yet claimed. The final transaction fix is deployed. - Full local `pnpm test:run` passed 15,614 general-server tests but stopped on two macOS timeouts. The heartbeat test passed in isolation; the existing 40,000-file Git stress fixture timed out again. Its Linux CI lane passes. Later local full-suite phases did not run after the timeout; this is not an all-green local full-suite claim. - Instructions and security limits are in `doc/public-mcp.md`; the saved plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`. ## Risks - This expands the experimental direct MCP surface. Explicit schemas and domain permissions must stay synchronized. - Migration 0311 adds transfer tickets and upload receipts. Expired orphan cleanup must not remove committed attachments. - Configuration requires a new consent request containing that scope; the single write-access choice controls it alongside work mutations. Refreshing an old grant does not add it. - The public directory keeps its original ten tools through the companion Cloud change. - Hosted consent/work proof remains the final delivery gate. The PR stays draft while approval of the new staging grant is pending; code checks and review are green. Merging is a separate action. ## Model Used OpenAI Codex (GPT-6, tool use and code execution). The exact serving model ID and context window are not exposed in this session. Paid evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001; claude-sonnet-4-6. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite macOS limitation disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
892b0b3606 |
fix: scope quota reports to authorized subscription accounts (#14994)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
eab93fd4a0 |
fix: checkpoint adapter usage and preserve unknown prices (#14991)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3ebd7bc0c9 |
test: isolate accounting fixtures and include all adapter suites (#14989)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f77fcbf4bf |
feat(apps): add Telem.AI web search connection (#15379)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Research agents need current web information. > - The Apps catalog connects agents to remote MCP tools through the normal access rules. > - Telem.AI supplies web search and page reading through one API key. > - This PR adds its catalog entry, optional search settings, artwork, and setup guide. > - The connection keeps an operator's saved header policy when they reconnect. ## Linked Issues or Issue Description Refs #15302. Related catalog work: #13881. This PR continues #15302 by Yifei Ai (@aiwen324). Thank you for the connector and its review fixes. All seven original commits are preserved. GitHub denied the attempt to push to the contributor's fork, so this branch retains the repair commit and merges current master. Master now includes the same test fix. For a squash merge, keep the original author in the final commit message: ```text Co-Authored-By: Yifei Ai <aiwen324@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> ``` A search found no separate public Telem issue or competing Telem PR. The existing Apps path matches the roadmap. **Agent or provider** Telem.AI provides web search and page reading through a hosted MCP server. **Why this adapter is useful** Agents can use multiple search providers through one governed connection. Operators can set the search tier, auto routing, and provider lists. **How the agent is invoked** The remote MCP server uses Streamable HTTP at `https://mcp.telem.ai/mcp`. The API key uses an `Authorization: Bearer` header. See the [official MCP guide](https://docs.telem.ai/integrations/mcp/). ## What Changed - Add the Telem.AI definition, research entry, permission review, and generated registry entry. - Add four optional settings. Unset settings send no request header. - Add official light and dark artwork, source records, and a setup guide. - Forward company, issue, agent, run, project, and correlation IDs by default. Preserve a saved policy, including disabled forwarding, on reconnect. - Add catalog and connection tests. ## Verification Current head: `b4164477fb1b312a504789bf17b51c963244c0fd`. Merged master: `228f0e2807c5b59d2aa129cf2d80b9777ebabf07`. - Resolved five shared catalog conflicts after the Superagent connection merged. - Keep both providers in the research ledger, generated registry, generator, branding manifest, and connection guide index. - Correct the combined catalog totals: 52 self-serve candidates, 55 research entries, and 68 Apps entries. - The published Git tree exactly matches the tested local resolution. - Catalog and Apps UI suites: **295 tests pass** after the catalog count fixes. - Connection service suite: **387 tests pass** in the full run. Its only failure was the old catalog count. That test passes on a focused rerun after the fix. This gives **388 passing service tests** across the two runs. - Total focused coverage: **683 passing tests**. The first runs exposed four fixed-count assertions that needed the combined totals. - Shared package build and plugin SDK compile pass. Token gates and whitespace checks pass. - Generation with `--definitions-only` reproduces the Telem definition and registry. The unrelated AgentMail and Linear drift remains excluded. - Local UI typecheck ended with exit 137 at the container memory limit. Full local typecheck, test, and build are not claimed. Earlier runs also recorded missing Cargo and Node development headers. - GitHub reports a clean merge state against master `228f0e280`. - All 54 checks are complete: **52 passed and two Storybook checks skipped**. No check failed or remains pending. - [CI](https://github.com/paperclipai/paperclip/actions/runs/37545412406) passes on this head. This includes typecheck, build, tests, browser shards, Runner checks, and Canary Dry Run. - [Greptile](https://github.com/paperclipai/paperclip/pull/15379#issuecomment-6024519537) is **5/5 on this head**. There are no review threads, open P2s, recommendations, or follow-ups. - [Superagent](https://github.com/paperclipai/paperclip/runs/112548142930) passes. - Final recovery checks confirm all seven original commits and current master remain in history. Token gates and whitespace checks pass. - The final recovery run makes no source change. It verifies the published repair and retains the local check limits below. - [Commitperclip](https://github.com/paperclipai/paperclip/actions/runs/37545407943) passes with no failures. Its only informational note asks the merger to keep the author trailer above. - All seven original contribution commits remain in history. The diff against master contains the same 14 Telem files. It adds no dependency, lockfile, schema, or workflow change. - The managed GitHub CLI capability was missing in this run. The installed GitHub connection applied the base files, merged master, then restored the tested combined catalog. No history was rewritten. The previous head `2d32a0094` passed all remote gates and had Greptile 5/5. Those results do not verify this new head. The original PR reports live setup, discovery, settings headers, gateway calls, and context-header forwarding. This repair does not repeat those account-bound checks. The permission record still marks maintainer live qualification as outstanding. ## Risks - Telem.AI receives the six context IDs by default. The saved header policy controls forwarding. Search use is billed to the account that owns the key. - All agents on a connection share its search settings. - The merge uses master's route-test setup unchanged. The company-boundary assertions remain intact. - No schema, dependency, or workflow change is included. - Live provider evidence is attributed to the original contributor. Maintainer live qualification remains outside this CI repair. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Original contribution: Anthropic Claude Opus 5.5 (`claude-opus-5-5`), 1M-token context, through Claude Code with shell, editing, and test tools, as disclosed in #15302. - CI repair and review: OpenAI `gpt-6-astra`, through Codex with reasoning, shell, editing, and GitHub tools. The runtime does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Yifei Ai <aiwen324@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
228f0e2807 |
fix(ui): make pull request review waits visible (#15375)
## Thinking Path > - Paperclip lets people supervise agent work and review its outputs. > - A task can wait for a human to review a pull request while an agent schedules checks. > - The PR work product and detected external links use separate display paths. > - A private PR can disappear from the useful sidebar view when its provider lookup fails. > - The monitor countdown says when the next check runs but omits the saved review request. > - This change shows the saved PR and requested review in the task's existing surfaces. ## Linked Issues or Issue Description **What happened?** A task kept checking whether a private GitHub PR had merged. Its work product requested board review, but the Properties sidebar used detected external objects instead. The PR URL was stored only in metadata, which the PR refresh path did not read. The composer showed a generic monitor countdown with no PR link or review action. **Expected behavior** Show saved PR links even if provider access fails. During a GitHub monitor wait, make outstanding review requests visible with the next check time and a status-check action. Stop requesting review when a PR merges or closes. **Steps to reproduce** 1. Register a PR work product with `reviewState: needs_board_review` and its URL in `metadata.url`. 2. Schedule an external-service monitor for GitHub. 3. Open the task with external-object lookup unavailable or unable to access the private repository. 4. Inspect Properties and the composer wait strip. Related work: #14469 added rich artifact cards. #8759 addresses attention on monitored task blockers. This change uses the existing work products and monitor action; it adds no blocker or approval mechanism. ## What Changed - Show saved PRs in Properties, including metadata-only links, independent of external-object availability. - Deduplicate equivalent GitHub PR URLs while preserving the saved navigation link, and put explicit review requests first. - Keep provider status and freshness on the combined PR row; keep private PR links usable when lookup fails. - Show the saved PR review request above the composer and in the existing monitor banner during a GitHub monitor wait. - Label the existing monitor action `Check status` and show check failures inline. - Suppress review prompts for merged, closed, or archived PRs even if their review flag is stale. - Refresh PR metadata using `metadata.url` and the existing `repository` alias. - Share saved work-product reads across the thread, Properties, and Artifacts. Refresh GitHub in a separate query and enrich only matching PR versions. - Start a fresh saved-row request on live invalidation so late provider responses cannot hide new artifacts or changed review requests. - Cover stalled GitHub lookups in both panels so refresh latency cannot hide saved work. - Scope monitor-check mutation state to the task so failures and late responses do not leak across navigation. - Refresh GitHub status when the displayed run finishes, including when saved PR rows have not changed. - Clarify that PR review and external release handoffs need a saved human-input interaction with an agent assigned for continuation; a flag, monitor, or handoff comment alone does not create that card. ## Verification - 386 tests pass across eleven affected UI and server suites on `184c7d8746`. Regressions cover saved links, provider status, cold-cache loading, panel reopening, monitor errors during navigation, and live updates during provider refresh. - Three live-update regressions and two run-completion regressions fail before their fixes and pass afterward. These tests use the real API client's GET coalescing and abort handling. - Full repository typecheck and build passed on the merged parent `1a49112dd2`. The latest UI changes pass UI typecheck and token gates; CI also verifies the latest build. Capability contract/inventory checks pass. - Storybook build and browser checks passed before the query race fix. Browser checks cover desktop and 390px mobile review waits, plus video, mixed-file, and empty artifact galleries after background refresh. - The earlier full local `pnpm test:run` was stopped under disk pressure. It reported failures outside the changed suites; an isolated skill-cache run reproduced three existing macOS permission failures. The full local suite was not rerun for this follow-up. - Latest-head CI passes on `184c7d8746`, including the browser shards and canary dry run. Apex is 5/5 with no actionable findings and no unresolved threads. All eight reported findings are addressed. ## Risks - This uses saved PR review state. When GitHub access fails, the saved state can remain stale until the agent updates it. The link remains visible and the status check remains available. - `Check status` wakes the existing monitor owner. It does not merge a PR, accept an approval, or mark the task done. - No schema, permissions, or scheduler behavior changes. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository analysis, code editing, and test execution. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — 277 affected tests - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cbd21b3e7 |
feat(apps): add Superagent connection (#15394)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents reach external services through the Apps catalog. Each catalog entry is a reviewed `AppDefinition` that connects a provider's hosted MCP server to Paperclip's shared vault, grants, policies, gateway, and audit trail. > - Superagent (`superagent.sh`) is a security platform. Its hosted MCP server lets agents read and triage security findings, start red-team reports, run Contributor Trust and dependency update jobs, and score web pages, email, files, skills, MCP repositories, and packages before an agent trusts them. > - Superagent is not in the catalog. Operators must use the generic "Connect your own MCP server" flow. That flow has no branding, no key guidance, and no warning about billable or destructive tools. > - The connector playbook supports this provider with existing definition fields. The server accepts an organization API key as a bearer header. It publishes no OAuth authorization-server metadata, so browser sign-in is not possible. > - This pull request adds the Superagent definition, its artwork, the research and permission-review ledger rows, documentation, deterministic tests, and one reviewed risk-classification rule. > - The benefit is a branded, governed Superagent connection with clear key guidance and a warning about tools that cost credits or delete data. ## Linked Issues or Issue Description **Problem or motivation** Teams that use Superagent for PR security, red teaming, and agent guardrails want their Paperclip agents to read findings, return structured reports, and score content before they use it. Superagent is not in the Apps catalog. Operators must paste the MCP URL and an `Authorization` header into the generic remote-MCP flow. That flow gives no branding and no provider guidance. It also does not tell the operator that the key reaches the whole organization, or that some tools consume credits or permanently delete findings. **Proposed solution** Add a catalog-only Superagent connection that follows the connector playbook. It has one method: a customer organization API key (`sk_live_...`), sent as an `Authorization: Bearer` header to `https://www.superagent.sh/mcp`. The field helper text explains that Superagent keys are not scoped. The method warning tells operators to set billable and destructive actions to Ask first before agents run unattended. All discovered tools stay governed by the normal per-action policies. **Alternatives considered** A browser sign-in method was not added. The server's protected-resource metadata names `https://superagent.sh` as its authorization server, but that origin publishes no `oauth-authorization-server` or `openid-configuration` document, so Paperclip cannot discover OAuth endpoints. A plugin was not needed because the connection needs no custom UI, tables, workers, or webhooks. Relying on the generic risk classifier was not enough. Several Superagent mutations (`triage_finding`, `scan_*`, `restore_agent_builtin_rule`) use names that it reads as reads, so a narrow reviewed Superagent rule was added instead. **Roadmap alignment** This extends the existing self-serve remote-MCP connection catalog. It does not overlap planned core work. ## What Changed - Added the `superagent` row to `packages/shared/src/self-serve-mcp-research.json` (API-key auth, risk tier S4). - Added the `superagent` provider to `scripts/ingest-app-definitions.mjs` (category, key placement and placeholder, console links, guidance, description). Regenerated `packages/shared/src/app-definitions/superagent.json` and the generated registry. - Added the `superagent/mcp-api-key` permission review to `doc/connections/tool-method-permission-reviews.json`, with key-permission text and evidence links. - Added Superagent's official mark (`ui/public/brands/apps/superagent.png`, the 460×460 avatar of the official `superagent-ai` GitHub organization) and the brand manifest entry. - Added gallery copy for the Superagent card. - Added a reviewed Superagent rule to `classifyRisk` in `server/src/services/tool-access.ts`. Only `list_*` and `get_*` tools, and tools that Superagent marks read-only, are reads. `delete_*` and `revoke_agent_client` are destructive. All other tools are writes, so billable and rule-changing tools can be set to Ask first. - Put the Ask-first advice in the API-key helper text, because the key form shows helper text and not method warnings. - Added `doc/connections/SUPERAGENT.md` (transport and auth, why there is no OAuth, administrator setup, capabilities and policy, manifest, brand provenance, validation hook). Linked it from the connections README and the permission audit. - Tests: definition shape, store visibility and artwork, URL recognition, the bearer header on discovery with the key kept out of connection config, read/write/destructive classification of fixture tools, the Superagent risk rule (including `triage_finding` and `restore_agent_builtin_rule`), the visible Ask-first advice and API-key gating of the connect form, and the pinned catalog counts. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts packages/shared/src/app-definitions-url.test.ts ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx`: 317 passed. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts`: 385 passed. - `node scripts/check-app-brand-assets.mjs` and `node --test scripts/app-brand-validation.test.mjs`: passed. - `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter @paperclipai/server typecheck`, and `pnpm --filter @paperclipai/ui typecheck`: clean. - Manual: in a local instance, open Apps → Browse and confirm the Superagent card and icon. Open `/apps/connect?source=superagent`. Confirm the single API-key method, and confirm that Connect enables only after a key is entered. - Live metadata probe on 2026-10-06: an unauthenticated `initialize` on `https://www.superagent.sh/mcp` returns 401 with `resource_metadata="https://www.superagent.sh/.well-known/oauth-protected-resource"`. That document returns 200. No authorization-server metadata exists at the named issuer. ## Risks - Low risk to existing providers. The change is additive catalog data plus tests. The generated registry only gains one import. The new risk rule runs only for Superagent connections. - A Superagent key reaches its whole organization. Some tools consume credits (`create_*_report`, `triage_finding`) or delete data permanently (`delete_finding`). Every action starts Allowed under the current product default. The key helper text tells operators to set these actions to Ask first. - The permission-review ledger records live proof as not run. No Superagent account was used. The lifecycle checklist in `doc/connections/SUPERAGENT.md` needs a documented pass before the entry is fully qualified. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with extended thinking and tool use (shell, file editing, web fetch, browser checks). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b508a05c43 |
feat: add internal agent complaints and suggestions (#15367)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use legacy skills or native runner tools to work on tasks. > - Those agents can encounter friction that does not belong in the task thread. > - A complaint should preserve the raw reaction. A suggestion should describe an improvement. > - This pull request adds attributed local storage and both submission paths. > - Agents can submit feedback once and continue their primary work. ## Linked Issues or Issue Description **Subsystem affected** Server, database, shared contracts, runtime skills, and native runner tools. **Problem or motivation** Agents have no default internal channel for incidental complaints and suggestions. Sending this feedback through task comments adds noise and can alter task workflows. **Proposed solution** Store free-form feedback in the current instance database. Derive agent, run, company, and task attribution from active authority. Provide default legacy skills and provider-neutral native actions. Keep the instructions close to Warp's MIT-licensed originals. **Alternatives considered** Task comments and external Slack delivery add unwanted side effects. Mandatory suggestion fields and short editorial limits would discard useful feedback. This release has no listing API, UI, read tool, automatic triage, or external forwarding. **Roadmap alignment** This is a maintainer-requested addition to the existing runtime skills and runner tool paths. It does not duplicate a listed roadmap milestone. Searches for complaint tooling, suggestion-box, and agent commentary found no overlapping public PR or issue. ## What Changed - Add the company-scoped `agent_commentary` table, shared validation, and idempotent migration `0310`. - Add one transactional service and the agent-only POST route. Validate active authority before writes or replay. Redact known credentials. Commit a content-free audit with each new record. - Add `submit_complaint` and `submit_suggestion` to standard, ask, and planning modes. Keep review, revocation, and completion restrictions. Store replay identity on the commentary row. - Mount `complain` and `suggestion-box` by default for legacy agents. Bundle a dependency-free Node.js stdin helper in the operational skill and allow its POST through the sandbox bridge. - Preserve Warp's complaint voice and suggestion guidance, with attribution and local transport adaptations. Keep source attribution and MIT notices in each skill's LICENSE, outside runtime instructions. - Document custom-runtime HTTP use and database inspection. Add real-database tests and a repeatable live Codex smoke for local and Daytona execution. - Pin the lagging-source migration fixture before the identity-repair migration so later migrations preserve its regression coverage. ## Verification - Personally ran real Codex submissions in all four environments on 2026-10-06. Local runs passed at 20:35 UTC. Daytona native passed at 20:31 UTC; Daytona legacy passed at 20:33 UTC. Each stored exactly two rows with company, agent, run, and task attribution, wrote the continuation marker, exited zero, created no task comments, and left task status unchanged. Each recorded two content-free activity entries. - Daytona used production provider hooks, real remote execution and file transfer, the legacy queue callback bridge, and native private WebSocket ingress. The current Linux runner was built from `abf47b595`, staged, and verified against controller contracts. Both sandboxes were confirmed deleted. This is a focused feedback transport smoke; it does not claim full Runner E2E catalog or browser qualification. - The immutable base image and Linux binary digest are recorded in [the verification documentation](https://github.com/paperclipai/paperclip/blob/codex/agent-commentary/doc/agent-commentary.md#verification). The smoke script can save content-free JSON evidence. No credentials or feedback bodies are in these reports. | Environment | Runner | Complaint row | Suggestion row | | --- | --- | --- | --- | | local | legacy Codex | `59413a00-1de2-4bb1-bcc6-9c4b54c64aa6` | `3db2364d-3e15-4f47-846f-875d3902999d` | | local | native Codex | `5da22b5f-41df-4de5-8ba0-d9345ab01267` | `2d5abe17-dd41-403c-a5ee-4729f2d58921` | | daytona | legacy Codex | `27c9d0aa-8477-409f-9da0-e8ffa48dee50` | `209681c9-d1e9-4ce1-999e-48fa07692389` | | daytona | native Codex | `6eb001bb-4bcf-43f7-8717-f662f53dc7c3` | `77c383d8-a997-49e5-a33e-25c70e15c0b2` | - Run the local check with `node cli/node_modules/tsx/dist/cli.mjs server/scripts/verify-agent-commentary-live.ts`. The documentation gives the Daytona invocation. Both use disposable instance databases and normal Codex provider usage. - Repository `pnpm -r typecheck` and `pnpm build` passed after the test extension. The build includes runner generation, contracts, and replay checks. The smoke scripts also passed a separate TypeScript check. The lagging-source migration regression passed. All equivalent current-head Vitest CI shards passed. The local monolithic `pnpm test:run` invocation was stopped after CI supplied that coverage; it did not complete locally. - Focused tests cover company isolation, spoofing, revoked credentials, stale ownership, post-finish rejection, concurrent replay, conflicting keys, atomic rollback, and deletion through existing services. Boundary tests cover empty text, Unicode, text beyond 8,000 characters, and the 524,288-character ceiling without truncation. Mounting tests cover Codex, Claude, and sandbox staging. Helper tests cover standalone Node execution, stdin, invalid UTF-8, redirects, HTTP failure, and its deadline. Privacy and bridge tests cover successful and rejected requests. - Instructions were compared with Warp's originals. MIT notices and source credits live only in LICENSE files. Native tools preserve truthful disclosure when asked, without routine announcements. - [Full CI](https://github.com/paperclipai/paperclip/actions/runs/37508559190) and Greptile 5/5 passed on the earlier feature commit `5209c3501`. The later head found the migration-fixture assumption fixed in this update. On `8a4965164`, all 55 check contexts passed after one browser shard rerun. Its initial reviewer signoff failure also passed an isolated local browser run (1 test). Greptile scored that head 5/5 and identified one smoke cleanup gap. `6ecbafb0b` fixes failed-acquisition cleanup with four passing tests and a passing smoke-script typecheck. Fresh CI is pending for this final test-only fix. No commentary production code changed during verification. ## Risks - Feedback is internally attributed. It is not anonymous. Existing redaction removes known credentials, but agents must still omit sensitive content. Normal provider transcripts can include their submitted arguments. - Default skill availability changes for existing legacy agents. Runtime policy filtering still applies. The helper uses the existing Node.js runtime with no extra dependencies; custom runtimes can call the HTTP endpoint. - Feedback is removed with its run, agent, or company. Task deletion clears only the issue pointer. Normal database backups include the table. - The migration is additive and has no backfill. Writes serialize on the active run for replay consistency. No server suggestion quota is imposed. ## Model Used OpenAI `gpt-6-astra` through Codex, with `xhigh` reasoning effort and a reported 258,400-token context window. Capabilities used: repository inspection, code execution, and live runtime verification. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fe9d865250 |
fix(native): wake parents after revised child completions (#15372)
## Thinking Path > - Paperclip manages agents and the tasks they perform. > - A parent needs a new notification when a child finishes requested revisions. > - Native completion currently identifies that wake by the parent and child only. > - An already consumed notification can suppress the child's next completion. > - This PR binds ordinary child notifications to the committed status decision. > - A replay keeps the same identity, while a new completion can wake the parent. > - Live evaluation then exposed new child results being swallowed by an active parent. > - The repair preserves that input for a fresh turn and prevents premature Done from discarding it. ## Linked Issues or Issue Description Refs: #15218. The procedure experiments remain unshipped. This fixes native completion wake identity and durable delivery to a busy parent. Related open work: #10559 and #4507 address legacy wake deduplication, #11179 addresses delivery during an active parent run, and #13044 addresses watchdog signals. This PR changes native committed-decision keys and their delivery to an active parent, while preserving exact watchdog behavior. **What happened?** A native child completed, its parent consumed the notification, and feedback reopened the child. The second completion found the old completed wake and created no new notification. **Expected behavior** Each new ordinary child completion can notify the parent. Replaying the same committed completion must not add a notification. If the parent is already running, the new result must remain available for a fresh turn; an unread result must survive a parent Done claim. **Steps to reproduce** Complete a native child, consume its parent wake, reopen and complete that same child, then inspect the parent wakes. The regression fails on unchanged master because only one wake exists after two completions. **Paperclip version or commit** Baseline: `0fe47882cfcb12082035113c59ca96091c46ebfc`. **Deployment mode** Native runner with the standard server and PostgreSQL control plane. ## What Changed - Include the durable status-decision ID in ordinary child completion wake keys. - Defer new ordinary native child completions behind an active parent, carrying the committed decision identity and revised summary. - Keep a parent in progress while that result is still queued, claimed or deferred; recheck before the status commit. Use the existing continuation without reviving cancelled tasks. - Lock the parent before ordinary child status writes, making the notification/Done ordering explicit without upgrading an implicit foreign-key lock. - Cover exact sequential delivery, dispatcher replay, and queued/claimed/deferred versus consumed/current-run completion identities in database tests. - Apply it when the child is also a dependency and when it is only a child. - Preserve the stable key for exact `task_watchdog` origins to avoid repeated watchdog loops. - Extend real database conformance coverage for both relationships, ordinary and near-match origins, watchdogs, revised summaries, and replay. - Reconcile the working checklist with the merged guidance PRs and record the bounded next step. - Reuse the current composer helper for Everyday task creation: capture the returned task ID, preserve the exact prompt and chosen assignee/project, and cover the setup with paused-agent browser tests. The same setup correction is present in both comparison variants. ## Verification **Ready for review and merge at `e2fc0c8e3ddb84dd9bc045704c3d1ecb23ee3447`: both original live cases pass, zero new failures against the frozen baseline, all checks green, CLEAN/MERGEABLE and out of draft. Not merged.** | Profile | Original baseline | Initial candidate | Fixed candidate | | --- | --- | --- | --- | | native Codex / `gpt-5.6-sol` | PASS | FAIL | PASS | | ACPX Claude / `claude-sonnet-5` | FAIL | PASS | PASS | One new pass, one unchanged pass, zero new failures and no pending pairs against the original baseline. The earlier failed candidate is preserved; it was fixed and measured at a new source, not regraded or rerolled unchanged. ### Current source and live evidence - [Completed campaign 37523025407](https://github.com/paperclipai/paperclip/actions/runs/37523025407) and [public report with original screenshots](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37523025407-1/index.html). Report and screenshots return HTTP 200; screenshot bytes match original retained evidence. - Measured candidate `e2fc0c8e3ddb84dd9bc045704c3d1ecb23ee3447`; reused frozen baseline `7d700f43e93e89e4c196f8799b0aa3cef41d20da`. Trusted workflow `88ff98b83d15eaa6640ee1848df4f0ad9bc82af3` is distinct from measured sources; workflow blob `0600886144d3e22ea2e4a38329a79177882f3948`. Prompts, fixture, oracle, model/profile and input controls are unchanged. Both variants include the same composer setup repair. - Both original candidate grades PASS: all 31 checks per cell, parent and child Done, independently tested revised ZIPs, successful cleanup. Exactly one first attempt per profile on this source: 10 actual runs, no retries. Reusing nine baseline runs gives 19 matched runs; the baseline was not rerun. - Claude directly exercised the repair: its revised child completed while the parent was active, the parent's Done claim became InProgress with `native_child_completion_pending`, and a fresh parent turn then completed with the revised artifact. - Both persisted final provider comments refer to the same attachment whose parent-registration hash matches the revised ZIP tested by the original oracle; native final evidence also references that attachment. Codex provides a clickable download link. Claude describes the new `--max-length` behavior but uses a backticked attachment ID, without a clickable URL in the final prose. The artifact is registered on the parent and independently downloaded/tested; prose-link usability remains a presentation limit outside the original oracle. Artifact identity and observed execution do not prove cognitive review. - 68 focused scheduling/conformance/arbiter tests pass. Four regressions fail against the original production files while four controls pass. The database scheduler test proves one sequential continuation with the revised summary and replay deduplication. It uses a mock adapter and is separate from the live proof. - Full local `pnpm -r typecheck` and `pnpm build` pass. [Exact-head full CI 37522126627](https://github.com/paperclipai/paperclip/actions/runs/37522126627): 51 successful checks, two intentional skips, separate Snyk success. Fresh exact-head Greptile [5/5](https://github.com/paperclipai/paperclip/pull/15372#issuecomment-6022179172), no new actionable findings; all review threads resolved. - Ready-transition Contributor trust and Superagent Security Scan both pass. The security scan completed at 2026-10-06T20:26:53Z with zero annotations. Final total: 53 successful check runs, two intentional skips and separate Snyk success; aggregate SUCCESS, source unchanged, zero unresolved threads. - The explicit parent lock precedes child writes. Its controlled PostgreSQL ordering also passes on previous production through an implicit foreign-key lock; this is hardening, not a reproduced additional live failure. The four original failing regressions remain the before/after proof of the active-parent repair. ### Preserved failures and accounting - Initial candidate `a2ae2324ce7692e704fc43e99904b076d6246fee`: Codex PASS → FAIL, Claude FAIL → PASS. Equal totals concealed a new failure and did not qualify that source. Its Codex parent ended Blocked after the revised child result coalesced into the active parent; there was no retained later parent execution and the independent ZIP oracle was never reached. A later server backstop log did not prove recovery. - In the initial Claude cells, both parent finals referenced an earlier parent ZIP while the oracle tested the revised child's ZIP. The original candidate PASS did not establish latest-artifact delivery. Earlier parent ZIP bytes are absent, so different hashes alone do not prove missing functionality. Those original grades and content findings are unchanged. - Original baseline [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37509893406-1/index.html); original candidate Claude [artifact](https://github.com/paperclipai/paperclip/actions/runs/37509916344/artifacts/11434353106); failed candidate Codex [campaign](https://github.com/paperclipai/paperclip/actions/runs/37513070706). The original cohort contains 17 actual runs, no retries, successful cleanup. - Campaign 37509916344 was cancelled after its Codex job waited over 16 minutes without a runner or steps. Completed Claude evidence was preserved; only the missing Codex first trial was dispatched in 37513070706. The trusted workflow revision differs but workflow bytes are identical. The cancelled campaign has no published HTML. Earlier campaigns 37506546098 / 37506550087 retain four composer setup failures, zero agent runs; the identical fixture-only repair restored setup without changing the prompt or outcome oracle. - Intermediate `3c1cf6830ce9db4123b834ffb99c594c383a8986` [campaign 37518652522](https://github.com/paperclipai/paperclip/actions/runs/37518652522) was cancelled after a concurrency review finding. Both paid-cell steps started and both tasks were created. Only invocation policies survived; no grade, run inventory or cleanup receipt. Provider activity and charges are unknown: two incomplete attempts, not passes or zero-provider setup failures. - Cumulative accounting: **27 known actual runs** (17 original + 10 fixed-candidate), plus unknown activity in those two cancelled intermediate attempts. The 19-run matched comparison reuses nine baseline runs and is not additional execution. Reported LLM amounts are zero with original billing `complete=true`; actual charges are unknown and local/hosted runtime is unmetered. No free-run, speed or cost claim. - Original and new results pass the canonical result validator. Retained source, input, result/API/story/final-ledger/usage run identities and artifact hashes were audited. Each downloaded package omits the pre-upload-declared `playwright-output/.last-run.json`; primary result, API, final ledger, story, screenshots and reached ZIP oracles are retained. No full-package completeness claim. - The initial redundant local full test invocation was stopped after 2,492.5 seconds once that head's CI passed; completed groups recorded 23,045 passes and 87 skips. That local invocation remains incomplete. Current full CI is the repository-wide test evidence. ## Risks A revised completion can schedule another sequential parent run and its normal budget use. A parent Done claim remains non-terminal while a newer native child result awaits delivery. Company scope, governance, workspace-finalization, terminal cancellation and exact watchdog behavior remain enforced. This changes native completion authority and scheduling, so the durable identity, replay and concurrent-commit controls matter. The two live trials qualify the observed revised-child handoff, not broad task quality or causal/general equivalence. Claude's final prose still gives an attachment ID without a clickable URL, although the registered revised artifact passes the original download oracle. Completion before any new child result exists, removal of unfinished dependencies and broader instruction reduction remain separate questions. There is no schema or prompt change. ## Model Used OpenAI Codex, GPT-6 family, with repository inspection, shell execution and code editing. The exact serving model ID, reasoning setting and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
582911ba74 |
fix(access): give Operators default company editing permissions (#15377)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Human company members receive permission grants from their role preset. > - Human invitations use the Operator role by default. > - The old Operator preset only granted task assignment, so ordinary members could not edit agents or manage connections. > - This pull request gives Operators company editing and audit access while keeping join approval and member-permission management separate. > - Members can configure their agents and accounts without an Owner role. ## Linked Issues or Issue Description **What existing behavior does this improve?** The default permission grants for human Operators and the invitation role description. **Subsystem affected** server/ permission presets and invitation authorization, plus ui/ invitation and tool access gates. **Current behavior** Operators receive only `tasks:assign` as an explicit default grant. Agent configuration and managed account setup fail their permission checks. **Proposed behavior** Operators receive agent creation/configuration, skills, environments, invitations, task assignment, pipelines, connections, tool management/use, and audit grants. The preset excludes `joins:approve` and `users:manage_permissions`. Operators can invite Operators and Viewers. Inviting a role with either excluded power requires that power, so invitations cannot bypass the restriction. **Reason and benefit** Ordinary company members can edit company work and connect accounts through the existing routes. **Breaking changes** The Operator preset grants more permissions. The existing startup and Cloud sign-in seeding paths can insert these missing grants for existing Operators. They retain custom grant scopes. Explicit invitation grants still take precedence. Inviting Admins now requires join approval; inviting Owners also requires member-permission management. This PR adds no migration or new backfill path. Related: #5945 describes missing role grants in another membership entry point. This PR changes the preset and does not change that endpoint. ## What Changed - Expanded the Operator preset to 14 company editing, invitation, tool, and audit grants. - Kept join approval and member-permission management out of the preset. - Required those powers when an invitation's selected human role includes them, preventing Operators from delegating the excluded powers through Owner/Admin invitations. - Updated the invitation role copy and the product specifications. - Allowed active Operators (including legacy Member roles) through the invite shortcut and advanced tool/profile UI gates. Kept inactive memberships, Viewers, and other companies excluded; API grants remain authoritative. - Added tests for the default Operator grants, both excluded actions, agent editing, company boundaries, and invitation role authorization. - Simplified adapter-route test imports so authorization errors and the error handler use one module graph; retained the upstream mock-initialization fix. - Kept Owner/Admin/Viewer presets and explicit invitation grants unchanged. Added no EE code or database migration. ## Verification - Permission and invitation tests: 38 passed across access service, invitation defaults, and invitation creation routes. - UI access tests: 61 passed across invitation shortcuts, Operator tool/profile access, and invitation UI. - Adapter route tests: 15 passed after resolving the upstream test setup conflict. - `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`: passed on the final branch. - The initial local `pnpm test:run` general-server batch reported three tool-access failures and eight upload-helper timeouts. Both complete affected suites passed in isolation on the final branch: 403 tests. The initial full local run was not green. - Both complete tool-access and adapter route suites also passed on an untouched snapshot of base `88ff98b83`: 387 tests. The earlier tool failures did not reproduce there. - [All CI gates passed](https://github.com/paperclipai/paperclip/actions/runs/37527422346) for `068be153b9c2a064f2aaccc0627fdf458e73747e`, including the previously failing serialized adapter lane, general tests, E2E, typecheck, build, Runner checks, and canary dry run. - Greptile scored that exact commit 5/5. Superagent passed. Both invitation review threads are resolved. ## Risks - Operators gain broad company editing and invitation access by design. - Existing default seeding can add the new missing grants during startup or Cloud sign-in. It does not replace existing scopes. - Company boundaries and the two excluded permission checks still apply. - An Admin without member-permission management can no longer invite an Owner. Operators retain invitation access for Operator/Viewer roles and agent-only invites. - This PR changes company permissions. Cloud workspace invitation rules remain separate. ## Model Used OpenAI Codex, GPT-6. The exact deployment model ID and context window were not exposed in this session. Used repository inspection, code editing, shell execution, and test tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting a merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ca0d26a99 |
fix(connections): recover missing personal AI credentials in chat (#15376)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ea8e686197 |
fix(tools): distinguish requested from granted OAuth scopes (#14059)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps use OAuth connection grants to let agents reach external services. > - Paperclip stored a requested scope as a granted scope when a provider omitted `scope` from its token response. > - A live MCP connection showed why that matters: the provider reported scopes beyond the requested scope. Scope names alone do not establish write capability. > - This PR records whether a scope came from the provider or from Paperclip's request, and flags asserted extra scopes. > - The result is an honest scope record across authorization and refresh, including for self-hosted and Cloud instances. **Review order:** shared scope resolver and OAuth write paths in `server/src/services/tool-access.ts` → optional schema/types/validator fields → focused tests in `server/src/__tests__/tool-access-service.test.ts`. This PR changes the shared OAuth record. The Enterpret catalog definition is separate in **#13906**. ## Linked Issues or Issue Description No public issue covers this bug. Related: **#13906** adds the official read-only Enterpret connector with OAuth and organization tokens. **What happened?** The OAuth callback stored `normalizeOauthScopes(token.scope ?? requestedScopes)`. When the provider omitted `scope`, Paperclip recorded the request as though it were a verified grant. In the live case: ```text Paperclip requested mcp:read Provider token reply no scope field Paperclip recorded ["mcp:read"] Token introspection openid email profile mcp:read mcp:write ``` The access token is opaque. A negative introspection control returned `active: false` with no scope, confirming the wider scope belonged to the live token. The grants API could therefore present requested scopes as though the provider had asserted them. This observation did not establish access to write tools. **Expected behavior** Store the provider's asserted scope when present. When absent, identify the value as an inference from the request. Preserve that provenance through refresh and report extra scopes only when the provider actually asserts them. **Steps to reproduce** 1. Use an OAuth provider that omits `scope` from its token response. 2. Complete consent after requesting `mcp:read`. 3. Read the connection or grant: before this fix, the stored `scopes` looked like an asserted read-only grant. 4. The focused fixture tests reproduce the callback and refresh record without needing a live provider. **Deployment mode** The shared OAuth path affects self-hosted and Cloud. The live provider evidence came from an isolated self-hosted runtime. ## What Changed - Added `resolveGrantedOauthScopes`. It records `scopeSource: provider` when the token response asserts scope, or `requested_fallback` when it does not. `unrequestedScopes` contains only provider-asserted scopes outside the request. - Applied the resolver at initial authorization and refresh. A refresh without `scope` retains a previous provider assertion and warning; a fresh assertion can replace them. - Stored the actual authorization request per grant as `providerTenant.oauth.requestedScopes`. This keeps a multi-user connection's refresh baseline tied to the right grant. Generic MCP OAuth also records discovered scopes when it sends them; curated apps retain their reviewed request. - Carried provenance onto the connection and default organization grant. The organization grant does not receive token expiry, so its existing refresh/reconnect behavior remains intact. - Added optional fields to database JSONB schema types, shared types, and validation. Existing grants need no migration or backfill. This PR **does not change authorization decisions**, reject tokens, introspect providers, or show a new UI warning. If a provider hides an over-grant by omitting `scope`, Paperclip still cannot discover it. `unrequestedScopes: []` with `requested_fallback` means **unknown**, not least privilege. ## Verification **Current head:** `7fde922a8`. Reconciled with master `88ff98b83`; the final diff is six OAuth implementation, contract, and test files. Preserved master's generic `offline_access` consent and legacy callback baseline, and verified requested versus provider-asserted scopes on each grant. Removed four duplicated GitHub token-method fixture properties after fresh review. All 504 focused tests in seven suites pass on the final head. Local workspace build, recursive typecheck, build-gap typecheck, module boundaries, node-version, token gates, and runtime push-policy checks pass. The full local Vitest run completed with 15,610 passing tests, 91 skipped, and six failures in heartbeat/workspace suites; all six failures reproduced on unchanged master `88ff98b83` in an isolated baseline worktree. These suites and their runtime code are unchanged by this PR. Greptile is 5/5 on this exact head with no actionable findings. All current-head GitHub checks are green: 53 successful check runs, two intentional Storybook skips, and the successful Snyk status. The initially failed adapter-access and signoff-heartbeat shards both passed their single rerun. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37521588330). The regression coverage includes omitted and explicit scopes, asserted extra scopes, refresh preservation, per-user request baselines, generic MCP discovery, and organization grants. No live provider token is needed to reproduce these recordkeeping cases. ## Risks - Existing records keep their scope list but have no source marker. Historical provenance cannot be recovered from them. - A refresh that omits `scope` preserves the last known assertion. A provider that silently narrows a grant will not update the record until it asserts a new scope. - A provider that omits `scope` leaves the actual scope unasserted. This PR cannot discover scopes omitted from the response or establish endpoint capabilities. Enterpret’s provider fixes are separate follow-ups. - Scope provenance is additive metadata. No policy path uses these new fields to allow or deny an action. ## Model Used Implementation and tests: Claude Opus 5 (`claude-opus-5`) through Claude Code with extended thinking, tools, and code execution. PR text cleanup: Codex GPT-6 (exact host model ID and context window were not exposed; tool use). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] Focused tests pass locally; full-suite baseline failures are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All current-head CI gates are green - [x] Greptile is 5/5 on the current head with no actionable follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge This shared scope-recording fix can be reviewed and merged independently of the Enterpret connector. --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e38d6d16b6 |
feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect accounts and choose an agent harness and model. > - The runtime change in #14970 supports custom providers on those connections. > - Normal setup must stay simple while advanced users can choose a compatible gateway. > - Shared connector rows and access controls keep these choices consistent. > - This pull request refines the agent setup UI and adds review stories and repeatable browser qualification. > - The qualification checks real tools and downloaded outputs, not only a successful run status. ## Linked Issues or Issue Description Refs #14970, #37, #13083, #14104, #14565, #12692. The core implementation in #14970 is merged. This branch incorporates its squash commit and targets `master`. Both PRs contain our implementation. #14016 is a reference only and is not a dependency. This PR has 96 changed files. ## What Changed - Complete model-provider connector presentation beside other connectors. Each row uses the existing Connect action and connection list. Tags are stored without category UI. The base PR includes the provider forms and routes. - Show persistent Subscription, API Key, and Advanced choices. Label Advanced as Custom Gateway. Reuse provider logos, connection lists, and permissions controls. Default access to the organization and all agents when permitted; keep narrowing controls under Advanced. - Keep Configure reachable before subscription sign-in, so users can select a supported environment when the default cannot sign in. Testing and saving still require a connection. Show the execution environment in Configure. Preserve the confirmed Connect choice. Editing a method, credential, saved account, or advanced choice requires that current choice to connect before testing or saving. Use matching model and thinking-effort dropdowns and retain connection icons in selected values. - Preserve the new harness model default when switching an existing OpenCode agent to Codex or Claude, and resolve user-selected model names with the effective harness. - Load popular OpenRouter models through the shared connection-model discovery path. Keep explicit model lists and manual model entry available. - Group onboarding, connection setup, agent runtime, management, recovery, and production-component stories under AI Connections / Provider routing. - Add an explicit-only provider-connections browser suite for managed local or existing local/staging targets. Use private browser profiles and credential handoffs. Support human-assisted subscription sign-in without sharing passwords or tokens in reports. - Verify persisted connection identity, runtime probes, tool execution, exact artifact bytes, completion, and context-dependent follow-up. Retain source/model provenance, cost bounds, closed error diagnostics, original failures, and cleanup evidence. - Add Gemini startup-model and skill-root fixes, Grok private-history detection, ACP filesystem regression fixtures, selected-workspace handling for local Hermes, and artifact-helper workspace fallback. - Keep managed Grok runtime homes disposable. Remove host-side transcript retention/restoration because private file modes do not isolate same-user agent processes. Ignore earlier development archives and use a fresh task handoff when history is unavailable. Verify the absence of restored transcripts with a separate same-user process. - Capture stopped-run diagnostics before deleting an attached-company fixture agent. Track creation and owned sign-in receipts; revoke only this attempt's accounts and never adopt a concurrent campaign's newly created account. Preserve failure signals and final status through cleanup. - Require the requested environment in the saved agent and every run, including follow-ups. Reject a forced incompatible target. Keep one cancellation state through startup, every cell, reporting, and teardown for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after interruption. Document qualification limits. ## Verification - Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes master `d9f600043`. The security fix in `a758fde31` passes full workspace typecheck, production build, and 119 connection/Grok regressions. The unchanged UI passes all 126 configuration/model-discovery tests and token gates. The final published-guide correction passes Grok adapter typecheck. Earlier head `eebd8225c` passed the complete deterministic runner suite (1,404 Vitest tests and 128 Node tests) and all CI jobs. Current-head CI run `37520147514` passed all 47 jobs, including the full sharded Vitest and browser matrix, production build, and canary dry run. All 55 checks completed: 53 successes and two expected skips. The current-head security scan passed, Greptile is 5/5, and no review threads remain open. - A separate same-user process reproduced reading a restored Grok transcript before the security fix. The regression now finds no transcript. Existing fresh-session fallback and ordinary session metadata behavior pass. - The final account-choice and cleanup fixes pass 85 setup tests and 26 qualification-harness tests. Regressions verify that editing a connection invalidates confirmation, Configure remains reachable before sign-in, diagnostics are captured before fixture deletion, and concurrent campaigns cannot adopt or revoke each other's accounts. UI and E2E typechecks pass. - The Storybook build and actual Chromium production-component stories passed during this change. Review the neighboring AI Connections / Provider routing stories, regular connector rows, three connection modes, model discovery, and the single execution-environment control in Configure. - Cancellation smoke verified authenticated cleanup before browser close for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during startup and reporting, missing-file ACP resource errors, and preserved permission denials. Both ACP runtime versions and 54 ACPX/Grok regressions passed. The deterministic connection-intent browser suite passed two tests. - Historical local qualification retained 43 passing API/gateway cells out of 46, with downloaded outputs and follow-up receipts. These attempts span earlier builds; they do not qualify this exact commit or staging. Subscription combinations, Gemini overloads, and the unresolved follow-up failure remain recorded rather than counted as passing. - Use `pnpm test:e2e:runner -- --list --suite provider-connections` to inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md` for credentials, target URL, sign-in assistance, budget, evidence, and cleanup. Paid live tests remain opt-in. ## Risks - The core implementation in #14970 is merged. This PR adds no database migration of its own. - Subscription login needs an interactive provider session. Dedicated accounts and staging qualification remain follow-up work; this PR does not certify every login combination for production. - Managed Grok transcript resume is deferred until provider history has an OS isolation or authorized broker solution. Follow-ups start fresh with Paperclip task context; earlier live Grok results do not qualify this behavior. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Live overloads and one unresolved follow-up timeout remain recorded. The stock CLI is unchanged, and those cases are not marked as passing. - Real-provider tests spend credits and use private credential/evidence directories. The launcher requires explicit selection and checks target ownership. It must not attach to a developer's database by accident. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local remain outside custom provider setup. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a7a2ed63a0 |
fix(ui): restore wheel and touch scrolling in task selectors (#15374)
Restore wheel and touch scrolling in nested task selectors and keep mobile picker controls reachable above the keyboard. Stabilize the affected server route test harnesses for Vitest 5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
88ff98b83d |
feat(apps): add read-only Enterpret MCP with OAuth and org tokens (#13906)
## Thinking Path > - Paperclip agents need governed access to Enterpret's hosted MCP tools for customer feedback. > - The official Enterpret MCP is read-only. Enterpret Agent’s beta write MCP is a separate service and is outside this PR. > - This PR supports both organization auth tokens and personal browser OAuth for the official MCP. > - Enterpret committed to fixing OAuth scope behavior and reducing its token-revocation cache window. The provider fixes will follow after connector deployment and do not gate this release. Paperclip functional validation remains required. New tools stay quarantined on refresh. > - This PR combines the original connector and Vivek's token-path work on current master, with live self-hosted QA recorded in the validation guide. ## Linked Issues or Issue Description No public issue covers this connector. This PR incorporates the work from #14737. The independent OAuth scope-provenance fix is #14059; it is not required for the organization-token path. #13890 is a merged connector precedent. #13692 provides the connector-authoring skill. **Problem or motivation** The generic MCP setup can reach Enterpret, but it has no reviewed Enterpret entry, method-specific setup, artwork, or policy defaults. The original connector branch had also fallen behind master. **Proposed solution** Publish the Enterpret catalog entry with an organization auth token as the selectable method. Support browser OAuth through instance-local DCR on the same read-only endpoint. Default `run_graph_query` to Ask first and quarantine tools newly found on refresh. ## What Changed - Added the Enterpret app definition, generated registry, official icon, Browse card, setup copy, tests, and Storybook states. - Combined #14737 and resolved the current-master conflicts, including the connector permission audit. - Preserved catalog quarantine on token reconnect and updated expiry guidance to follow the token's dashboard expiry. The QA token displayed three months, contradicting the former six-month copy. - Enabled personal OAuth for the official read-only MCP. Removed the write-capability inference and OAuth draft-only setup restriction. The beta Agent endpoint is excluded. - Preserved tool quarantine during token reconnect, with a regression test for existing quarantined and newly discovered tools. - Added OAuth scope and revocation-cache disclosures and an interactive Storybook retry check. - Merged master `d9f600043` into the existing PR branch without rewriting history. Resolved four conflicts, regenerated the combined registry, and retained the model-provider assertions with a catalog count of 66. - Recorded live QA and its limits in [doc/connections/ENTERPRET.md](https://github.com/paperclipai/paperclip/blob/950bacc52/doc/connections/ENTERPRET.md). ## Verification Current head: `6c73d3084202ccbb25080e7a4e9c3467bc1341dd`. All 513 focused connector, discovery, OAuth, policy, shared-definition, and UI-handoff tests passed across 11 suites. Full repository typecheck, build, Storybook build, token gates, module boundaries, Node policy, no-git-push policy, and typecheck-build-gaps passed. Release-registry tests passed 129/129 with modern Bash. The four failures with macOS Bash 3.2 also reproduced on untouched master. The full test suite passed in CI on this exact head, including every general-test and serialized-server shard and all eight E2E shards. The duplicate local `pnpm test:run` was stopped after these remote test lanes passed; it did not complete locally. The combined registry was regenerated. Regeneration also reproduces existing Linear and AgentMail definition drift on untouched master; those unrelated changes were excluded. On an isolated self-hosted Paperclip instance, an organization token connected and discovered eight actions. `get_organization_details` returned the intended organization before and after reconnect. Catalog refresh, invalid-token failure and valid-token recovery, activity logging, local disable, and local removal were exercised. No feedback quotes or graph query were requested. Paperclip revoked both QA connection secrets and archived the QA connections. The token-path agent access preview correctly showed an action Off, but the board Test endpoint still invoked it as the board user. This shared Test-path mismatch is documented; it is not evidence that a real agent can bypass policy. An actual token-path agent-session denial, a real agent process, expiry recovery, server/VPS, and Cloud remain unverified. Enterpret's dashboard confirmed the QA token was revoked, but the same token immediately reconnected and completed a safe read. Vivek explained this as a 24-hour validity cache and committed to reducing the window. The new maximum delay and deployment remain unverified. The earlier OAuth check reported `mcp:write` and `email` for an `mcp:read` request. This is a scope mismatch; it did not prove access to write tools. On 2026-10-01 the account holder relayed Vivek’s clarification that the official MCP is read-only and the beta Agent write MCP is separate. ## Risks The organization token can read the organization's customer feedback, and provider-side revocation did not immediately stop access. Operators should check each token's displayed expiry and remove Paperclip's local connection when retiring access. The account holder confirmed the provider fixes are fast follows after deployment. Release can proceed with the current scope reporting and up-to-24-hour revocation delay documented, once fresh OAuth authorization/refresh, real-agent allow/deny, CI, and review pass. Retest the reduced revocation window after Enterpret deploys it. Scope labels alone must not be used to claim write access. Greptile is 5/5 on head `6c73d3084` with no actionable findings. All current-head CI checks pass, including the canary dry run. Live agent-session denial and fresh OAuth refresh remain explicit deployment QA items. The board Test panel does not prove agent authorization. Provider scope-reporting and revocation-cache fixes are accepted fast follows after deployment. ## Model Used Original connector and validation record: Claude Opus 5 through Claude Code. Integration, current-master reconciliation, and token-path QA: OpenAI Codex, GPT-6 family (exact host variant not exposed), using repository tools and an isolated runtime. Contributor commits retain their authors. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal issue ID - [x] I have run relevant local tests and checks - [x] I have updated relevant documentation - [x] I have considered and documented the risks above - [x] All current-head CI gates are green - [x] Greptile is 5/5 with no actionable follow-ups If squash-merging, preserve contributor credit in the squash message: Co-Authored-By: Vivek Kaushal <kaushalvivek@users.noreply.github.com> --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Vivek Kaushal <vivek@enterpret.com> |
||
|
|
1cebdd4c4f |
Add plugin lifecycle delivery and an initial resource baseline (#15343)
## Thinking Path > - Paperclip manages agents and projects for work. > - Their lifecycle changes already commit to a durable journal. > - Plugins need to read these records and track completed work. > - Existing resources need a one-time baseline before delivery is enabled. > - Worker crashes must leave unfinished records available for retry. > - Each plugin needs its own progress and company access checks. > - This PR adds a pull inbox through the existing plugin SDK and job system. ## Linked Issues or Issue Description **Problem or motivation** Plugins cannot consume the durable resource lifecycle journal. In-process notifications can disappear during a restart and cannot record successful completion. Resource plugins need retries, company boundaries, and ordered transitions. **Proposed solution** Add `ctx.events.listLifecycle(companyId, limit?, afterId?)` and `ctx.events.acknowledgeLifecycle(companyId, eventId)` under `events.subscribe`. Seed a one-time baseline from current resource state. Deliver creation first, then the remaining transitions in ID order. Store acknowledgments for each plugin. Use existing plugin jobs to poll configured companies. **Alternatives considered** A global sequence cursor can skip lower IDs that commit later. Fire-and-forget subscriptions cannot record completion. A new dispatcher is unnecessary because plugin jobs support polling. Consumers serialize polling and use provider idempotency keys. **Roadmap alignment** This extends the existing plugin system and builds on #15280 and #15306. Searches found no duplicate resource inbox work. Related #13306 exposes decision events on the in-process bus. This PR includes the initial journal baseline. Provider provisioning remains separate work. ## What Changed - Add company-scoped lifecycle reads and acknowledgments to the SDK and worker RPC host. - Gate both methods by capability, invocation or proactive company scope, plugin readiness, and company enablement. - Seed hired agents and all projects once. Preserve paused/terminated agent state and archived project state. Keep pending hires behind approval. - Preserve existing history and deliver backfilled creation before partial transition histories. - Deliver project archive events from #15371 through the SDK. Seed archive intents for archived projects and update intents for active projects with a partial archive history. - Store acknowledgments per plugin and reject acknowledgments that skip earlier resource events. - Page past failed resources while retaining their pending records. Reset the page cursor each sweep to include late commits. - Add acknowledgment storage, an index for resource ordering, tests, and authoring guidance. - Preserve native identity definitions and sequence progress in JavaScript backups. Repair journal ID generators lost by older backups before seeding the baseline, without changing existing IDs. ## Verification - `pnpm -r typecheck` passed after the final origin/master rebase. - `pnpm build` passed before the final metadata rebase. The final CI build also passed. - All 46 focused lifecycle, SDK harness/RPC, migration snapshot, and legacy restore checks passed again with migration 0309. - Checks cover retries, per-plugin progress, resource order, capability/company boundaries, baseline idempotency, approval gates, archive delivery, restored projects with partial histories, and atomic migration rollback. - Full CI passed on final commit `0e674b87fc`. One unrelated Cursor fixture hit a 10-second timeout; it passed locally in under one second, and the failed server shard passed on retry. - Greptile scored the final commit 5/5 with no unresolved review threads. - The branch is rebased onto origin/master and is conflict-free. - Earlier full local runs were stopped as scope changed. CI ran the complete repository suite. - `git diff --check` and a local secret/PII scan passed. ## Risks - Apply migration `0309_loving_the_hood.sql` before starting the new server. It creates acknowledgment storage and an index, then seeds the baseline in the same transaction. Agent/project writes wait for the migration to commit. Keep these writes quiesced through the migration and activation of the new capture-capable server; do not resume an older runtime that lacks capture after the baseline. - The baseline records current desired state, not historical transitions. It includes archived projects and terminated-agent cleanup intents. It runs once with the delivery migration; no later or runtime journal backfill is planned. - Delivery is at least once. Concurrent reads can repeat an event. Consumers must serialize polling, use stable company/event idempotency keys, and acknowledge successful operations only. - Reset `afterId` at the start of every polling sweep. It is a page cursor, not a persisted high-water mark. - Deleting a plugin removes its acknowledgment records. A new installation may replay existing journal events. - Records contain identity and action. Consumers must load current authorized data before acting. Cleanup and retention policy belong to the provider plugin. - Provider calls, VM/volume provisioning, and journal retention are outside this PR. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code execution, and tool use. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
854af7df19 |
build(deps-dev): bump vitest from 4.1.11 to 5.0.3 (#12969)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.11 to 5.0.3. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v5.0.3</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Isolate <code>result.status</code> between <code>repeats</code> runs - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11218">vitest-dev/vitest#11218</a> <a href="https://github.com/vitest-dev/vitest/commit/5dbebe9e3"><!-- raw HTML omitted -->(5dbeb)<!-- raw HTML omitted --></a></li> <li>Don't print an interceptor warning in browser mode - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11377">vitest-dev/vitest#11377</a> <a href="https://github.com/vitest-dev/vitest/commit/15cc006aa"><!-- raw HTML omitted -->(15cc0)<!-- raw HTML omitted --></a></li> <li>Don't retry when <code>test.fails</code> expectedly failed - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11219">vitest-dev/vitest#11219</a> <a href="https://github.com/vitest-dev/vitest/commit/b24585f08"><!-- raw HTML omitted -->(b2458)<!-- raw HTML omitted --></a></li> <li>Scope cache key generators to projects - by <a href="https://github.com/ecoyoung"><code>@ecoyoung</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11281">vitest-dev/vitest#11281</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11301">vitest-dev/vitest#11301</a> <a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1d"><!-- raw HTML omitted -->(92ba7)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Delay server <code>listen</code> until tests start running - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11366">vitest-dev/vitest#11366</a> <a href="https://github.com/vitest-dev/vitest/commit/7d8ed3e9b"><!-- raw HTML omitted -->(7d8ed)<!-- raw HTML omitted --></a></li> <li>Check mock path boundaries - by <a href="https://github.com/saryn17"><code>@saryn17</code></a>, <strong>Ryosei Sato</strong> and <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11361">vitest-dev/vitest#11361</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11362">vitest-dev/vitest#11362</a> <a href="https://github.com/vitest-dev/vitest/commit/1c3888bce"><!-- raw HTML omitted -->(1c388)<!-- raw HTML omitted --></a></li> <li>Keep config of browser-consumed environments - by <a href="https://github.com/kasperpeulen"><code>@kasperpeulen</code></a> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11378">vitest-dev/vitest#11378</a> <a href="https://github.com/vitest-dev/vitest/commit/aafc0996f"><!-- raw HTML omitted -->(aafc0)<!-- raw HTML omitted --></a></li> <li>Ignore page crash while cancelling - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11386">vitest-dev/vitest#11386</a> <a href="https://github.com/vitest-dev/vitest/commit/7c36748fa"><!-- raw HTML omitted -->(7c367)<!-- raw HTML omitted --></a></li> <li><code>toMatchScreenshot</code> uses wrong reference on retried tests - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11393">vitest-dev/vitest#11393</a> <a href="https://github.com/vitest-dev/vitest/commit/c22aba992"><!-- raw HTML omitted -->(c22ab)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>cache</strong>: <ul> <li>Revalidate imports of cached modules - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11381">vitest-dev/vitest#11381</a> <a href="https://github.com/vitest-dev/vitest/commit/38f98855f"><!-- raw HTML omitted -->(38f98)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>deps</strong>: <ul> <li>Pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into <code>ERR_PNPM_TRUST_DOWNGRADE</code> - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11403">vitest-dev/vitest#11403</a> <a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977"><!-- raw HTML omitted -->(f6c9a)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Pass current equality testers to <code>expect.extend</code> asymmetric matchers - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11401">vitest-dev/vitest#11401</a> <a href="https://github.com/vitest-dev/vitest/commit/3e794a96b"><!-- raw HTML omitted -->(3e794)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Support Blob on jsdom 30.1 - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11379">vitest-dev/vitest#11379</a> <a href="https://github.com/vitest-dev/vitest/commit/6c49b7197"><!-- raw HTML omitted -->(6c49b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>pool</strong>: <ul> <li>Preserve unique pool ids when <code>groupOrder</code> is set - by <a href="https://github.com/mtorp"><code>@mtorp</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11392">vitest-dev/vitest#11392</a> <a href="https://github.com/vitest-dev/vitest/commit/50312ebb4"><!-- raw HTML omitted -->(50312)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>ui</strong>: <ul> <li>Split-pane handle overlapping iframe - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11221">vitest-dev/vitest#11221</a> <a href="https://github.com/vitest-dev/vitest/commit/f91db0dfd"><!-- raw HTML omitted -->(f91db)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vitest</strong>: <ul> <li>Remove root temp dir on close - by <a href="https://github.com/abhinav-phi"><code>@abhinav-phi</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11248">vitest-dev/vitest#11248</a> <a href="https://github.com/vitest-dev/vitest/commit/7c7119cf7"><!-- raw HTML omitted -->(7c711)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vm</strong>: <ul> <li>Do not optimize deps from index.html - by <a href="https://github.com/ezefernandezyf"><code>@ezefernandezyf</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11329">vitest-dev/vitest#11329</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11360">vitest-dev/vitest#11360</a> <a href="https://github.com/vitest-dev/vitest/commit/caf2887de"><!-- raw HTML omitted -->(caf28)<!-- raw HTML omitted --></a></li> <li>Don't reuse scripts across vite environments - by <a href="https://github.com/MO2k4"><code>@MO2k4</code></a>, <strong>Martin Oehlert</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11395">vitest-dev/vitest#11395</a> <a href="https://github.com/vitest-dev/vitest/commit/346d3896b"><!-- raw HTML omitted -->(346d3)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v5.0.2...v5.0.3">View changes on GitHub</a></h5> <h2>v5.0.2</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Bind <code>process</code> in case global is overwritten - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11343">vitest-dev/vitest#11343</a> <a href="https://github.com/vitest-dev/vitest/commit/0b79231ad"><!-- raw HTML omitted -->(0b792)<!-- raw HTML omitted --></a></li> <li><strong>detect-async-leaks</strong>: <ul> <li>Ignore <code>process.stdio</code> handles - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11333">vitest-dev/vitest#11333</a> <a href="https://github.com/vitest-dev/vitest/commit/0fd6b9790"><!-- raw HTML omitted -->(0fd6b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Fix <code>toMatchObject</code> with asymmetric matchers - by <a href="https://github.com/ShreeBohara"><code>@ShreeBohara</code></a>, <strong>Claude Opus 5</strong>, <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-5)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11100">vitest-dev/vitest#11100</a> <a href="https://github.com/vitest-dev/vitest/commit/42523289e"><!-- raw HTML omitted -->(42523)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Fix <code>Request</code> with <code>Blob</code> body on jsdom 28+ - by <a href="https://github.com/harshit-d3v"><code>@harshit-d3v</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11295">vitest-dev/vitest#11295</a> <a href="https://github.com/vitest-dev/vitest/commit/d1c3ecc93"><!-- raw HTML omitted -->(d1c3e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporter</strong>: <ul> <li><code>agent</code> to respect <code>--silent</code> - by <a href="https://github.com/Raj4478"><code>@Raj4478</code></a> and <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11271">vitest-dev/vitest#11271</a> <a href="https://github.com/vitest-dev/vitest/commit/5b95efb6d"><!-- raw HTML omitted -->(5b95e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporters</strong>: <ul> <li>Handle concurrent <code>createReport</code> calls - by <a href="https://github.com/7rulnik"><code>@7rulnik</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11278">vitest-dev/vitest#11278</a> <a href="https://github.com/vitest-dev/vitest/commit/e8e556ff7"><!-- raw HTML omitted -->(e8e55)<!-- raw HTML omitted --></a></li> <li><code>hanging-process</code> to use ESM entrypoint - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11316">vitest-dev/vitest#11316</a> <a href="https://github.com/vitest-dev/vitest/commit/4e91e5668"><!-- raw HTML omitted -->(4e91e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>spy</strong>: <ul> <li>Fix stack overflow when spying <code>Set.prototype.add</code> - by <a href="https://github.com/fengmk2"><code>@fengmk2</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11299">vitest-dev/vitest#11299</a> <a href="https://github.com/vitest-dev/vitest/commit/a0a939653"><!-- raw HTML omitted -->(a0a93)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/33cadea62e8763c455c7fca38d9ab1dda87c5f75"><code>33cadea</code></a> chore: release v5.0.3 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11409">#11409</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/346d3896b65c3c907174447a035807342799f346"><code>346d389</code></a> fix(vm): don't reuse scripts across vite environments (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11395">#11395</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977ad3363f796a737834572e54c6ad5c18"><code>f6c9a49</code></a> fix(deps): pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into `...</li> <li><a href="https://github.com/vitest-dev/vitest/commit/062c75d8b63519211d951d8293ea81b5a9e3c124"><code>062c75d</code></a> chore: fix standalone docs build, update exports maps (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11394">#11394</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/caf2887dee8987a60118d53933f6e9cabd6b3e2a"><code>caf2887</code></a> fix(vm): do not optimize deps from index.html (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11329">#11329</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11360">#11360</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/50312ebb4eca6a98f6d0b2b61d5d9d38cbbabcef"><code>50312eb</code></a> fix(pool): preserve unique pool ids when <code>groupOrder</code> is set (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11392">#11392</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/7c36748fad1eae9687312f2f7ceadce6ec88b5df"><code>7c36748</code></a> fix(browser): ignore page crash while cancelling (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11386">#11386</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1df16a4fa5bbee3f198c582fbd56689d8"><code>92ba7fc</code></a> fix: scope cache key generators to projects (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11281">#11281</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11301">#11301</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/38f98855fa9cd7fd376afb84094eba0fda256a74"><code>38f9885</code></a> fix(cache): revalidate imports of cached modules (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11381">#11381</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/b24585f08f2ea267746a2d6ca0e43edcbb29726f"><code>b24585f</code></a> fix: don't retry when <code>test.fails</code> expectedly failed (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11219">#11219</a>)</li> <li>Additional commits viewable in <a href="https://github.com/vitest-dev/vitest/commits/v5.0.3/packages/vitest">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a590ab769d |
Give managed agents persistent cryptographic identities (#15352)
Give agents persistent Ed25519 identities encrypted with the existing instance master key. Create keys transactionally for new agents and lazily before supported managed runs, expose public identities in the API and agent UI, and protect private material during runtime delivery and output persistence. Preserve identities in recovery backups while giving imported and development-cloned agents fresh keys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d9f600043f |
Capture project archive and restore lifecycle events (#15371)
## Thinking Path > - Paperclip manages projects and their repositories. > - Resource changes commit to a durable lifecycle journal. > - Project archive changes currently leave no journal record. > - Consumers need to observe archive and restore transitions. > - The record must commit with the project change. > - This PR adds project archive capture before plugin delivery in #15343. ## Linked Issues or Issue Description **Problem or motivation** Archiving a project does not record a lifecycle event. Restoring an archived project also emits no event when it is the only change. Consumers cannot observe these transitions through the journal. **Proposed solution** Record `archive` when an active project becomes archived. Record `update` when an archived project is restored. Keep both writes in the existing project transaction and row lock. Repeated status requests emit no new status record. **Alternatives considered** An in-process notification can disappear on restart. A separate archive service would duplicate the existing mutation path. Use the existing journal and project transaction. **Roadmap alignment** This extends lifecycle capture from #15280 and #15306. It should merge before the delivery PR #15343. The roadmap and related PR search showed no duplicate archive lifecycle capture. ## What Changed - Allow project `archive` records in the journal constraint and TypeScript type. - Capture archive and restore transitions under the existing project row lock. - Emit `update` then `archive` for an edit combined with archive. - Preserve repository records and roll back the project change if capture fails. - Add migration `0307_cool_naoko.sql`, tests, and database documentation. ## Verification - `pnpm -r typecheck` passed after the final master rebase. - `pnpm build` passed before the final master rebase; the final CI build also passed. - 20 targeted lifecycle, migration snapshot, and legacy restore tests passed on the current migration. - Tests cover concurrent archives, repeated archive/restore requests, repository retention, combined edits, atomic rollback, and upgrades from older JavaScript backups without the action constraint. - All 53 applicable CI checks passed on final commit `1021953035`; the two Storybook checks were skipped as expected. - Greptile scored the final commit 5/5 with no unresolved review threads. - `git diff --check` and a local secret/PII scan passed. - Prior CI found a missing-constraint restore failure; the migration now handles it. A separate runtime readiness timeout passed locally. The current full CI run passes both paths. ## Risks - Older JavaScript restores may omit the action check. The migration tolerates its absence and installs the complete check. - Apply the constraint migration before running the new capture code. It takes a short table lock and validates existing journal rows; it changes no existing records. - This PR captures future transitions only. The one-time baseline and plugin consumption remain in #15343, which must be rebased after this PR merges. - Archive records authorize no provider cleanup. Provider behavior and volume retention remain separate work. - An edit combined with archive emits two ordered records in one transaction. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code execution, and tool use. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9f7057e122 |
feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use a harness, a model, and a credential to run tasks. > - Connections already store credentials and control who can use them. > - Custom providers also need an endpoint and a supported API format. > - A per-agent endpoint would duplicate credentials and access rules. > - This pull request stores routing on the connection and projects it into the harness. > - Isolated credentials and protocol checks keep the selected connection authoritative. ## Linked Issues or Issue Description Refs #37, #13083, #14104, #14565, #12692. #14016 is a reference only. This PR has its own schema, vault persistence, routing validation, runtime projection, and tests. None of the five commits in #14016 is an ancestor of this branch. We do not depend on or plan to merge it. #14967 addresses task-pinned account pools. #14422 addresses another provider integration. This is the first of two linked PRs. Merge this connection change before #15341, which refines agent setup and adds the qualification harness. The split keeps each review below 100 changed files. Provider catalog entries have usable setup forms in this PR. Local browser subscription sign-in is included. ## What Changed - Store non-secret routing metadata on AI connections. Vault provider API keys, including Bedrock bearer API keys. Reject general AWS access keys. - Enforce company, owner, human audience, agent access, connection status, and protocol checks before resolving credentials. Keep reconnect destinations immutable and retain connection identity during key rotation. - Project OpenRouter and compatible custom endpoints into Codex, Claude, OpenCode, and local Hermes. Carry these settings through both legacy and native runner transports. Clear conflicting host credentials and redact keys from diagnostics. - Preserve older OpenRouter accounts and native personal defaults. Add Google API-key accounts and migration `0306` for the two provider-default constraints. - Run local Claude and Codex subscription sign-in behind the existing browser sign-in card. Use private attempt homes and owner-bound completion instead of a copied terminal command. - Seed isolated Gemini authentication and preserve OpenCode workspace permissions. Keep the selected connection authoritative. The independent Gemini and Grok workflow fixes are in #15341. - Keep native OpenCode custom gateway keys in a runner-owned selected-model proxy; the harness config contains only a session-scoped capability. Honor runtime outgoing proxy and certificate settings. Preserve streamed responses and revoke the proxy on close or startup failure. - Allow ordinary members to connect native personal accounts before an agent exists. - Repair routed accounts from task cards using the saved provider destination, protocol, model aliases, and connection identity. - Add provider catalog definitions, model discovery, pinned logos, and complete native and routed setup forms. Allow a personal routed connection before a new agent exists. Keep endpoint authentication keys out of Hermes terminal children. - Recover cancelled or restarted browser sign-in with a clear restart action. Support no-auth endpoints without a vault credential. Add isolation and recovery regressions and runtime documentation. ## Verification - Updated with `origin/master` at `22a3ea341`. Migration `0306` follows the new master migration and passes migration and snapshot checks. - The integrated connection regressions passed 152 tests and 50 native OpenCode driver tests, including key-free child-shell configuration reads, authenticated/no-auth forwarding, streaming, model/path restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass. Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34 executable also completed a turn through the proxy against a local synthetic provider; the reusable key was absent from its config. A second real-executable smoke passed with an HTTPS CONNECT proxy and runtime-specific synthetic certificate trust. Certificate-file and certificate-directory regressions pass. - Task-card repair passed 48 tests, including OpenRouter, Bedrock, and custom gateway reconnect cases. UI typecheck and token gates passed. - The prior core regression set passed 133 tests across new-agent setup, provider forms, browser sign-in, routing projection, and connection authorization. Token gates and UI typecheck passed. - Full workspace typecheck and production build passed again after the latest integration and credential-proxy fix. The merged deterministic runner E2E suite passed 1,400 Vitest tests and 128 Node tests. - Full workspace typecheck passed on the prior linked combined implementation. Production build, Storybook build, 1,316 browser-harness Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the latest master integration and regenerated migration. This exact head passed 54 remote checks with four expected skips and Greptile 5/5; no review threads remain open. An unchanged server fixture had a random six-character issue-prefix collision on its first attempt. All 245 tests passed locally and the single CI retry passed. - A provider-free terminal check used the cited supported Hermes source and dummy keys. Gateway and OpenRouter terminal children could not read the selected key. - The broad local Vitest attempt passed 15,442 tests but was not green. It had an embedded-Postgres startup failure, an HTTP logger timeout, an origin socket error, and a browser cancellation wait timeout. The cancellation wait was corrected. The relevant connection tests and the full origin test file passed separately. Latest-head CI must pass before merge. - Prior credential-backed acceptance exercised task creation, tool use, artifact delivery, completion, and context-dependent follow-up. Claude legacy and native runners passed Bedrock with `us-east-1` and `us.anthropic.claude-sonnet-4-6`. - Historical local qualification retained 43 passing API/gateway cells out of 46. Those attempts span earlier builds. They do not qualify this exact commit or staging. All subscription combinations and staging remain unqualified. - Verify native subscription and API-key setup. Connect a regular provider catalog row. Verify an incompatible harness and a changed reconnect URL are rejected. Use #15341 for the complete browser campaign. ## Risks - Migration `0306` changes two check constraints. It preserves rows and is safe to reapply. It takes normal constraint-change locks. - Credential projection touches several harnesses. CLI upgrades can change provider configuration and session behavior. - The native OpenCode proxy adds a loopback hop, pins requests to the selected model, limits request bodies to 16 MiB, rejects redirects, and expires at session close. It prevents reusable keys in the child configuration; it is not an OS isolation boundary against a process debugger running as the same user. - Custom endpoints must be reachable from the agent environment. Saving a connection does not prove connectivity. Bedrock keys require rotation before expiry. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Provider overloads and an unresolved follow-up timeout also affect live Gemini qualification. We have not patched the installed CLI or marked those cases as passing. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS identity, arbitrary auth headers, and custom routing for other harnesses are excluded. - These PRs do not establish production or staging qualification for every provider and login method. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
22a3ea3414 |
Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Public MCP lets people use their organization from an external assistant. > - Operators need a visible control for this experimental access. > - Hosted users should select an organization once and then approve its permissions. > - This pull request adds the setting, invitation-first setup, and browser or device consent. > - Connections provides a copyable invitation with public instructions that grant no access. > - Users reach browser consent from their assistant and return to inspect or revoke access. ## Linked Issues or Issue Description Builds on merged foundation #14846. This PR now targets master. Related settings convention: #13905. **Current behavior** The preview uses an environment variable to enable MCP. Hosted consent repeats organization selection. Assistant access has no entry in Connections, so users must already know the endpoint and how to reach consent. **Proposed behavior** An administrator enables Settings → Experimental → Assistant connections (MCP). A hosted connection shows the selected organization and its icon, then asks for permissions. Requested write access starts checked when the user’s role permits it; the user can opt out before connecting. Direct instance connections show an organization picker with the first available organization selected. The selection stays fixed across refetches and still requires an explicit Connect action. Connections includes Assistant Connection (MCP). Its setup page explains the canonical endpoint, client configuration, browser authentication, and connected access. It connects as the current person and does not select or impersonate an agent. **Reason and benefit** Operators manage access with the other experiments. Users select one organization, and both the UI and server enforce that choice. **Breaking changes** The old enable variable has no effect. Preview operators must enable the setting once. Apply the additive consent-request migration before deploying the tenant, then deploy the compatible Cloud broker. Existing direct requests and grants keep their behavior. ## What Changed - Simplify OAuth and device consent: show the Paperclip logo beside “Connect {client} to Paperclip”, fall back to “your assistant”, and show the identifying origin plus its favicon below, with the callback URL also visible when different. Remove the hosted-organization creation action. Default to the first available organization without silently changing it on refetch; preserve company restrictions and write opt-outs. Keep the button row contained on narrow screens. - Make Copy invitation the primary action, using the shared animated AgentSetupPrompt and a collapsed manual setup section with icon-labeled line tabs. Remove redundant link actions, copy-status text, the extra first-prompt well and revocation explanation from the setup page. Serve shared version-aware HTML and Markdown instructions without private organization data. - Support guarded Client ID Metadata Documents alongside dynamic registration, and include authorization response issuer identification. - Add RFC 8628 device authorization with separately hashed codes, expiry, shared request quotas, persistent polling backoff and atomic redemption. Reuse human consent, role checks, scoped grants, audit and revocation. - Add CLI device login and a local stdio bridge. Store credentials separately with private permissions and serialize rotating refreshes. - Add device consent stories and five cold-start paid Product E2E cases with independent grant, configuration and durable-work assertions. - Add `enablePublicMcp` to the settings validator, normalizer, feature catalog, and toggle UI. Check it live for OAuth, tools, subscriptions, and event delivery. Keep connection management and revocation available while disabled. - Default the MCP origin to the existing auth public URL, with strict validation and an explicit override. - Persist the optional OAuth `company_id` restriction. Describe only that company and reject approval for any other company, even if the person belongs to both. Keep active-membership and role checks. - Show the Paperclip icon and a large organization icon during consent. Return the saved company logo through the company-scoped request response and reuse the standard fallback icon. Use the requested concise permission labels: “Read all of your Paperclip data” and “Allow write access and creating tasks as me”. Use concise permission copy, retain a compact client and callback-origin disclosure, and remove the footer link. - Default requested write access on for eligible roles. Preserve opt-out across organization changes and refetch, reset defaults for a new request, and submit read-only access when the request or role does not allow writes. Align the shared checkbox with its label. - Use organization wording in consent, management, settings, and walkthroughs. Keep the organization fixed for hosted requests and retain direct-instance choice. - Add an Assistant Connection (MCP) card to the Connectors catalog, a setup page in the app shell, and a return link from Experimental settings. Include Codex, Claude Code, OpenCode, and generic remote MCP instructions. - Read the live gate and canonical server URL through authenticated setup metadata. Show only the current person’s grants for the selected organization, refresh after consent, and support revocation. Surface catalog status failures with an explicit retry action; do not present them as an empty connection list. Opening setup grants no authority. - Start the eight guided chapters in Connections. Keep presenter notes and chapter controls around real product pages in the app shell. Explain the terminal, consent, delegation, retrieval, and revocation handoffs. Mark conversation examples as illustrative. Cover first use, client setup, connected, loading, and error states. Keep the existing consent and management stories. - Keep the paid-eval setup and browser helper aligned with the setting and consent button. ## Verification - Warm-standby integration fix `ec64ea05e`: public MCP ingress now follows the Cloud claim guard; MCP and discovery paths return 503 instead of SPA HTML while unclaimed. Event polling checks the in-memory claim before reading the persisted experimental setting. All 97 focused OAuth/Cloud tests and server typecheck pass, including new request and timer regressions for idle-before-claim and resume-after-claim behavior. Fresh review is 5/5 with no unresolved threads, and all security scans pass on this final head. All browser shards, typecheck, build, canary installation and other test groups passed on the first attempt. The unchanged Cursor sandbox default-command test timed out at 10 seconds; the exact test passed locally without edits in 587 ms. The single failed-job retry passed, with the original failure retained in workflow 37500711895. All 54 final-head checks pass on `ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs are intentionally skipped). - Final master integration `8457828fc`: merged foundation #14846 and current master, preserving the invitation changes and all 33 files from the two newer upstream changes. No migration renumbering was required. All 95 focused OAuth/Cloud integration tests, full recursive typecheck and token gates pass. All CI gates passed on that integration head; review identified the warm-standby issue fixed above. - Security-review fix `8c1d0b696`: commit shared global/per-source admission before outbound CIMD work, preserve failed-attempt receipts, and validate resource/scope before fetching. Added migration `0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive typecheck and production build pass. The security scanner passed that commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18 authorization requests sharing one proxy across two service instances use just two fetches. All 70 OAuth/metadata tests and server typecheck pass after that refinement. Final follow-up `8ebeae84c` reports admission-storage failures as retryable HTTP 503 instead of invalid client metadata. Its regression proves no outbound request before admission and successful retry after storage recovers. All 71 OAuth/metadata tests and server typecheck pass. Final-head security scanning passes; Greptile is 5/5 with no unresolved findings. CI passed all browser shards, typecheck, build, token gates and canary installation. One unchanged adapter-utils bridge test raced a response-file write (expected a JSON error, received the safe file-changed error). The exact test passed locally without edits. The single failed-job retry passed; the original failure is retained in workflow 37490609192. All 54 checks now pass on final head `8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final integration above now targets master. - Integration with current master: preserved the new Connections source filters and pagination, kept all eval suites, and regenerated the consent/device snapshots as migrations 0303/0304. All four MCP migration SQL hashes are unchanged from the staging versions. Full recursive typecheck and production build, 132 focused UI tests (including catalog filtering), 89 server authorization/settings tests, 26 migration tests, 120 eval calibration tests and token gates pass. Review follow-up `4820ce74c` also keeps active assistant grants in Installed, with pending/error recovery and revocation/company-isolation coverage. All 76 setup/catalog tests, UI typecheck and token gates pass after that fix. The unchanged signoff browser test timed out waiting for a heartbeat in CI at `4820ce74c`; the exact test passed locally without code changes, and the preceding CI head passed that shard. That same unchanged test failed at the reviewer stage in the next CI run. All five signoff tests passed three times locally (15/15), without test changes. All eight browser shards pass at final head `8ebeae84c`; no browser-test edits or failed-browser-job retries were needed. - Setup-page refinement at `9ab009178`: all 17 focused setup/consent tests pass, along with UI typecheck, production build, Storybook build and token gates. Browser exercised the shared prompt preview and client tab switching, and the updated InvitationCopied Storybook interaction checks its clipboard fixture. All final-head CI checks pass at `9ab009178`, with no unresolved review findings. Deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524. Verified the actual page, tab switching and line styling, removed actions/copy, and successful native copy/paste of the complete Butter invitation into a local-only test field. The existing Claude grant was left intact. - Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates pass. UI typecheck and production build passed again at `4e4d5e4d9`; Storybook build and eval-helper typecheck passed for `28101cf91`. Follow-ups let the primary button wrap on narrow screens, preserve a distinct callback URL, and use only bundled icons to avoid pre-consent requests to client-selected sites. Browser-verified the real consent component in desktop and 320px mobile stories, including default selection, write access and preserved opt-out. Updated E2E heading/default-selection helpers. All CI checks passed at `dc8e9fd11`, with review 5/5 and no unresolved threads. The Butter preview publication needed a retry because npm initially accepted the DB package before making it visible; the retry succeeded and `dc8e9fd11` deployed. Verified a fresh, unapproved native Codex CIMD request on Butter: default organization/write selection, known-client heading and icon, distinct callback origin, and removed creation action. No grant was approved for this UI check. Prior paid runs below retain their exact source provenance; this UI-only follow-up did not rerun paid qualification. - Source-pinned paid matrix at `2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign `local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config, unavailable host, denied consent and reconnect/later retrieval, with independent configuration/grant/task/run/document assertions. Original failures, transcripts, source fingerprints and billing remain retained. - Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`. Latest `0637b9f1c` shares that same guidance across HTML, Markdown and manual UI after review; generated Markdown is verified byte-identical to the paid-evaluated version. Shared build, server/UI typechecks, token gates and 63 auth/metadata tests passed again. Every CI gate passed at prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads. - Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval calibration tests, server/UI/eval typechecks, token gates and Storybook build passed. Full recursive typecheck and production build passed during implementation; CI also passed them at `2992ef271`. - Local full-suite limitations: a large-file Git streaming test times out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh MCP reruns passed, and the corresponding CI groups passed. No claim that the local full suite is green. - Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD consent; device CLI reaches verification/consent. New grants await human approval. Existing local OpenCode retrieved a saved result in a fresh conversation through its previously approved grant. - Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received the exact copied invitation, read public setup, configured its server and started PKCE consent. Its shell command timed out; background retry reached the client's own callback deadline while approval remained pending. Latest instructions cover that handoff. **No completed Butter read/delegation/result retrieval is claimed.** - Cloud companion https://github.com/paperclipai/paperclip-cloud/pull/672 passes checks/review and deployed. Anonymous setup and device-protocol routing verified. Core `2992ef271` deployed successfully and the actual Claude web flow now reaches consent. Its extra JWT-bearer metadata is filtered to implemented grants; unsupported token grants remain rejected. Final `0637b9f1c` deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300; live HTML and Markdown both contain the final guidance. The superseded instruction-only build was canceled before deployment. This is a core-only staging preview; private Cloud plugins are omitted. ChatGPT web is signed out, so browser connector use is unverified. - Screenshot gallery begins at Butter's dashboard and distinguishes real setup/pending consent from local reuse and fixtures. It records the timeout finding. New persistent access needs human confirmation before the remaining actual-client acceptance work. - Manual path: Connectors → Assistant Connection (MCP) → Copy invitation → paste into assistant → configure and start authorization → sign in and approve → verify `paperclip_connection` → delegate → retrieve the saved report later. - Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md` and `doc/public-mcp.md`. ## Risks - Apply additive, replay-safe migration `0304_curvy_shadow_king.sql` before using device authorization. The public setup link carries no credential. Device codes and tokens stay private; neither sharing instructions nor installing a plugin authorizes access. - Apply additive migration `0305_chubby_vin_gonzales.sql` before deploying the shared metadata admission gate. It retains at most 60 short-lived, hashed-source receipts per instance and rejects excess attempts with 429. - CIMD metadata fetching is a new external-input boundary. It requires HTTPS, exact client ID and redirect validation, bounded responses and guarded DNS/network access. Client names remain self-reported. - Device support is per-instance. The central Cloud broker retains its existing grant support. Host installation and tool reload capabilities vary by client; instructions describe manual settings and restart requirements. - Consent names the registered client in its heading and displays its identifying origin below. Known-origin icons are bundled; all other origins show a neutral site icon without contacting client-selected sites. Client names are self-reported; the callback origin is the recipient check. The Cloud chooser also displays the original client and receiving origin before tenant handoff. - A user who accepts the preselected write permission can create tasks and comments. Task creation and comments can start or wake agents and use execution budget; the consent label uses the concise wording explicitly requested by the maintainer. Scope requests, role checks, and the final Connect action still apply. - Migration `0303_supreme_garia.sql` adds one nullable UUID column with `IF NOT EXISTS`. Requests without a company restriction keep the direct-instance picker. The binding stays recorded if its company is deleted; consent then fails closed. - Deploy tenant support before the Cloud broker sends `company_id`. Unknown or inaccessible organizations must never fall back to a different company. - The setting defaults off. Disabling access does not cancel work already delegated. Existing tokens and unexpired subscriptions can resume when enabled again; revocation remains separate. - The catalog entry is visible for discovery while the feature is off. Setup instructions, OAuth, and tool execution remain gated. No access is granted by viewing the entry. - Assistant sign-in starts in the external client so it owns PKCE and callback state. Client command syntax can change and links to official setup documentation are included. - An authenticated instance and valid public URL are required. Hosting, paid execution, and store publication remain separate rollout steps. ## Model Used OpenAI GPT-6 in Codex, with tool use and code execution. The exact serving model version and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks pass; unrelated local full-suite timeouts are explicitly recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e34abee670 |
feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path > - Paperclip gives teams durable tasks, agent execution, budgets, and approvals. > - People also use assistants in Codex, Claude, and other MCP clients. > - Those assistants need a scoped connection that preserves the person’s permissions and attribution. > - Delegating a task must not turn the assistant into the assigned agent. > - This PR adds opt-in user OAuth, ten first-party tools, browser consent, and workflow packages. > - Paid product evals verify the resulting tasks, documents, attribution, retries, and access boundaries. > - The team keeps working after the assistant conversation ends. ## Linked Issues or Issue Description **Problem or motivation** A person cannot connect an external assistant to an existing team through browser consent and safely delegate durable work as themselves. **Proposed solution** Expose an opt-in `/mcp/paperclip` endpoint with individually described first-party operations. Bind every connection to a person, client, company, resource, and scopes. Reuse domain authorization and scheduling. Package shared team-review, delegation, and follow-up workflows for OpenAI/Codex and Claude. **Alternatives considered** Related PRs #9393 and #12549 cover earlier remote MCP and board-operator approaches. This change uses user OAuth and a bounded public catalog. It does not expose a generic executor, operator administration, static shared board credentials, or external agent execution. Registry listing work in #9851 is a separate distribution step. **Roadmap alignment** This maintainer-requested implementation extends the governed MCP gateway, activity attribution, durable work products, and hosted deployment direction in `ROADMAP.md`. It implements the first release of the saved design plan; external agent participation and granted third-party tools remain later releases. ## What Changed - Add MCP 2.0 discovery and task status/comment/document Events on the same authenticated endpoint. Persist subscriptions and delivery receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt callback material, recheck permissions/Cloud membership, and bound retries/expiry. Older MCP clients keep their existing tools. - Add discovery, dynamic client registration, S256 PKCE, resource validation, rotating refresh tokens, revocation, and company consent. Store credentials as hashes and recheck membership at execution. - Add tools for connection identity, agents/projects, task search/read/create, human comments, documents/deliverables, and pending-approval links. Preserve current domain permissions and scheduling. - Add durable mutation receipts across reconnects. Matching retries replay results; uncertain outcomes keep the same request ID and require inspection. - Add consent and connection-management pages, OAuth log redaction, shared plugin workflows, and separate OpenAI/Codex and Claude package outputs. - Add eight paid Product E2E cases across three models, independent durable-state grading, usage evidence, cleanup, and report integration. Add task-document guidance and regenerate the runner capability inventories. - Add migrations 0301 and 0302, the dated implementation plan, result notes, and direct-client setup instructions in `doc/public-mcp.md`. ## Verification - Merge integration `e180b1948`: resolved conflicts with current master, preserved both eval registries, regenerated capability catalogs, and regenerated migrations as 0301/0302 while keeping the original replay-safe SQL byte-identical. Local migration safety/snapshot tests (26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval catalog/grading tests (198) pass. Token and capability gates pass. Full recursive typecheck passed. Fresh Greptile review is 5/5 with no unresolved findings. CI is green on this exact head (55 successes, two intentional skips, one neutral result): one unchanged Cursor sandbox test timed out at 10 seconds, then passed locally in 856 ms. A single retry of that failed shard and the aggregate workflow passed. Merge remains blocked on the repository code-owner approval rule. Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55 successes, two intentional skips and one neutral result. [The earlier CI run](https://github.com/paperclipai/paperclip/actions/runs/36901592350) includes all test shards, browser tests, typecheck, build and canary dry run. Greptile was 5/5 on that commit with no unresolved review threads. GitHub still requires code-owner review under the repository merge rules; passing checks do not bypass that approval. Paid source fingerprints remain separate below and in the dated result note. - Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku 4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback, signature verification and report retrieval in a fresh conversation. A final Mini regression passes after the quota/status fixes. All evidence validates. Bounded tunnel startup retries occur before provider calls and remain visible; failed earlier attempts retain their original grades. - The earlier complete seven-case matrix passes **21/21**, with a separate **3/3** delegation regression. Two preceding matrices also passed 21/21 each. A complete 24-cell matrix including Events has not been run. [The dated results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain exact source fingerprints, failures, model IDs and partial costs. - Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass after merging master. Server typecheck passes after the final quota/status changes. Eval typecheck and all 892 eval-support tests pass. - All 33 real MCP/OAuth tests pass. The preceding combined MCP, redaction, private-address and DNS-rebinding run passed 129 tests; two later MCP regressions cover quota reuse and unchanged-status suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI then found a null checkout result in the existing concurrent-workspace path; logging now uses optional status access. All 12 closed-workspace tests and all 33 MCP tests pass after that correction. The exact-start event calibration exposed a timestamp gap; scanning now includes the subscription start, with all 33 MCP tests and server typecheck passing. These two narrow corrections follow the paid regression. - A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0 subscription/delivery, current membership loss, unsubscribe, legacy SDK tools, refresh and revocation. Its callback transport is a fixture with independent HMAC verification. The paid Events campaigns separately prove public HTTPS delivery. - Earlier component qualification passed UI 7,117 tests, CLI 502, shared 832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module boundaries, migration order and plugin regeneration passed. CI covers general/serialized suites, eight browser shards, runner checks, typecheck, build and canary dry run. - **Local full-suite limitation:** the earlier monolithic run was not clean. It encountered overlapping schema rebuilding, Mac database shared-memory limits and isolated CLI/fixture failures. Targeted reruns passed. The existing >32 MiB Git filename stress test still hit its 300-second Mac timeout. The additional serialized sweep stopped after 62 passing suites once CI passed. Original failures and partial logs remain; this PR does not claim a wholly green local monolithic run. - Local Codex CLI and Claude Code OAuth login and MCP SDK interoperability were verified. Public-store installation, actual ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted newcomer provisioning remain release gates. Enablement is moving to **Settings → Experimental → Assistant connections (MCP)** in the stacked follow-up [#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge both for the intended setup experience. This foundation branch alone still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set `PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and connect to `/mcp/paperclip`. Select a team and allow writes in browser consent. Configure an available agent and budget, then delegate and retrieve results later. For Events, rescan the deployed plugin catalog in ChatGPT Work Cloud; the host supplies its webhook credentials when the user asks to watch a task. See [the setup runbook](doc/public-mcp.md). ## Risks - Events are at-least-once and may arrive out of order. No replay cursor is advertised. Clients must refresh finite subscriptions, read current state and avoid comment feedback loops. Callback material uses the instance secrets master key; hosted subscriptions require the updated Cloud broker and are bounded to five minutes/the access proof expiry. - ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging subscription remain deployment gates. Local signed-webhook and paid model evidence does not claim those surfaces have been exercised. - Disabled by default. Merging adds schema and opt-in code; it does not deploy a public endpoint, publish a store listing, create a team, or start paid agents. - Migrations 0301 and 0302 are additive and idempotent. Their SQL is unchanged from the earlier preview numbers, so hash-aware upgrade reconciliation preserves prior staging applications. Normal instance upgrades must apply it before enabling MCP. - Task creation and comments can schedule paid agent work. Consent and tool descriptions disclose that effect. Revocation blocks future calls but does not undo delegated work. - Public deployments need edge rate limits and credential-safe logging. Internal dispatch is restricted to the closed catalog and carries a request-local verified actor. - Hosted onboarding requires the companion Cloud broker, encryption-key configuration, and tenant rollout. Self-hosted direct connections can use this PR alone. - Store acceptance and agent-mode participation are not claimed. Checked-in plugin endpoints are development defaults; rebuild packages for a real deployment before installation. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A more specific serving version and context-window size were not exposed by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`, `claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted/component checks; full local-run limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
202c2d307e |
fix(server): leave unclaimed warm Cloud databases idle (#15314)
## Thinking Path > - Paperclip manages work performed by AI agents. > - Managed deployments prepare empty applications before an owner claims them. > - Those applications start database pollers even though no company can have work. > - Health probes also query SQL, so idle databases cannot remain suspended. > - This pull request adds an explicit standby marker for empty, unclaimed Cloud apps. > - The existing signed, durable claim resumes normal processing without restarting the app. ## Linked Issues or Issue Description **What happened?** An empty, unclaimed warm application runs recurring chat, email, plugin, heartbeat, cleanup, and reconciliation queries. Its health route also opens the database. This prevents idle database compute from suspending. **Expected behavior** An explicitly marked unclaimed application should keep its HTTP process and sandbox provider plugins ready while leaving the database idle. A successful signed claim should resume normal API behavior and background processing. Claimed and self-hosted instances should keep their current behavior. **Steps to reproduce** Start an empty Cloud-managed application and leave it unclaimed. Observe database activity while repeatedly requesting `/api/health`. Before this change, periodic queries continue without company data. Related: #15153 reduces allocation during chat polling. This change suppresses polling only for explicitly marked, empty, unclaimed Cloud apps. ## What Changed - Add `PAPERCLIP_CLOUD_WARM_STANDBY=1`. Check company emptiness once after restoring the persisted Cloud runtime identity. Missing Cloud configuration, existing data, or a persisted claim leaves normal processing active. - Gate recurring database pollers with an in-memory predicate. Keep startup preparation and sandbox provider plugin loading intact. - Serve unclaimed health probes without session or database reads and report `warmStandby: true`. Serve standby pages/assets directly from the UI router, bypassing session, bearer, tenant, and dynamic handlers. Refuse API requests and all WebSocket upgrades before authentication can query SQL or seed company data. - Exit standby after the existing signed identity assertion commits. Normal timers resume at their next tick; a restart restores the claim even with stale provider variables. - Document the marker, readiness semantics, rollout checks, and rollback. ## Verification - `pnpm -r typecheck` passed. A final server typecheck also passed after adding tests. - `pnpm build` passed. - Focused standby, signed claim, restart, health, static/Vite routing, hostname, HMR, and live-events suites: 72 passed after the review fixes. Includes real HTTP upgrade admission before/after claim. - `pnpm test:run` was attempted locally; both superseded runs were stopped after encountering checkout/platform failures. A clean-checkout rerun eliminated ancestor skill-directory lookup failures. The company-skills/runtime-cache families encounter macOS read-only-directory rename failures (`EACCES`); all three company-skills failures reproduce on unmodified base `bf14f803d5`. The initial full run also reported one native runner API test failure; an isolated comparison on both revisions was blocked by local embedded PostgreSQL startup failures. The full [Linux CI run](https://github.com/paperclipai/paperclip/actions/runs/37423263993) passed on final commit `5016c415ea`, including all server and workspace test shards, browser suites, typecheck, build, and release canary. This is not a claim that the full local suite passed. - Isolated full server with local PostgreSQL: after startup and connection expiry, 70 health probes, 70 page requests carrying valid synthetic tenant credentials, and 70 rejected WebSocket upgrades over 70 seconds observed zero app database connections. The signed claim completed in 62 ms and normal polling resumed (475 database transactions over 12 seconds). Restart with stale provider variables restored the durable claim. The latency is local-only, not a provider wake measurement. - Apex review: **5/5** on `5016c415ea`, both earlier threads resolved, no open recommendations. - No live-provider test or production deployment was performed. An actual database suspension/resume canary remains required before enabling the control-plane switch. ## Risks - Standby health reports HTTP readiness rather than current database connectivity. The signed claim still requires a durable database write; claimed health checks retain the SQL probe and 503 failure behavior. - Pollers resume at their usual intervals. A suspended database may add claim latency. Validate the real provider before enabling the marker. - Startup preparation and sandbox plugins remain loaded. New plugins or background loops must respect the same standby contract. - The marker is off by default. Remove it or set it to `0` and restart to roll back. No schema migration or claimed-workspace inactivity policy changes. ## Model Used OpenAI Codex, based on GPT-6. The exact serving snapshot and configured context-window size are not exposed in this session. Assistance included source review, TypeScript changes, command execution, and PostgreSQL tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted tests pass; full local suite limitations are documented above, and full Linux CI is green - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f2715e02bb |
fix(native): surface model capacity errors and retry automatically (#15347)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs preserve provider events and results, then decide task state. > - Codex can fail a turn because the selected model is temporarily at capacity. > - Paperclip displayed a generic native-session failure and left the task in recovery. > - A known capacity failure needs a clear message and a durable delayed retry. > - This pull request uses committed terminal evidence, the status-effect ledger, and existing dispatch gates. > - The task can continue automatically without an unbounded retry loop or a model switch. ## Linked Issues or Issue Description Related: #13993 adds launcher capacity deferrals before provider startup. This change handles a committed native Codex `serverOverloaded` terminal after provider work has started. **What happened?** A native Codex turn ended with `codexErrorInfo: serverOverloaded` and “Selected model is at capacity. Please try a different model.” Paperclip saved the result, showed a generic failure, and required recovery instead of waiting and retrying. **Expected behavior** Show the capacity error directly. Schedule bounded retries after a short delay. Preserve saved work and honor current execution gates. **Steps to reproduce** 1. Run a native Codex task. 2. Commit a `turn.failed` event with `serverOverloaded`, then accept a failed result for the same turn. 3. Inspect the task error and its recovery state. The regression test reproduces this without a paid provider call. **Paperclip version or commit** Reproduced against master `16b7db35ffa0f9a95913c8cbdeea3d595435691f`. ## What Changed - Classify capacity failures from committed runner events and the pinned execution identity. Preserve accepted results and display the specific capacity error. - Atomically persist one scheduled successor with the status decision. Retry after one minute, then two minutes. Share the existing failure budget and stop after two automatic retries. - Reuse dispatch gates for ownership, task holds, dependencies, budget, and locks. Wait for predecessor execution, finalization, and cleanup before claiming a retry. - Preserve pending reviewer authority and suppress retries after reassignment or a successor claim. - Consume the failed run's resume receipt and delivered wake input. Rebuild ordinary continuation from the failed run so explicitly resumed tasks can retry without borrowing one-run authorization. - Label scheduled retries “Model at capacity.” Add a Storybook example and avoid duplicate punctuation in the existing retry card. - Add regression coverage and document the runtime contract. ## Verification - Targeted recovery and UI suites: 125 tests passed. Additional final cleanup and lock regressions passed. - Latest-head continuation and authorization regressions: 57 tests passed, including all four capacity integration tests against an isolated PostgreSQL database. - Rechecked all four capacity integration tests with the full runner's isolated `PAPERCLIP_HOME`, config, temporary directory, and serial fork settings: passed. - `pnpm -r typecheck`: passed. Final server typecheck passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Browser: checked the real Storybook card. It shows the capacity message, automatic retry time, and existing Retry now action. - `pnpm test:run`: attempted and restarted after an interruption. The resumed run started before the final review correction and was stopped after recorded workspace/native test failures and the five-minute Git streaming timeout already documented on master. Exit 130; no complete local full-suite pass is claimed. Current-head isolated capacity tests and the complete CI suite pass. - A combined local recovery-suite attempt also hit PostgreSQL initialization failures at this macOS host's global shared-memory limit (32 slots). The focused recovery suites passed separately; final-head capacity tests also pass with the full runner's isolated environment settings. - Latest-head GitHub checks: all 56 checks green, including general and serialized test shards, browser E2E, typecheck, build, Runner verification, and canary dry run. No merge conflicts. - Greptile: 5/5 on `813a470b8c18c05aeb7e31e63e573c3a3a2f5cac`, with zero unresolved findings after fixing the resumed-task continuation issue. ## Risks - Capacity retries can repeat a task turn after partial work. They start a fresh provider session with task history and wait for predecessor cleanup. They do not resume the failed turn. - Retries retain the configured model unless an operator changes configuration. Persistent overload consumes the existing failure budget and then needs an explicit retry or model change. - Only run-bound native Codex `serverOverloaded` failures qualify. Usage-limit exhaustion, model/account incompatibility, unbound text, and unknown provider failures retain their current recovery behavior. - No schema migration is required. ## Model Used - OpenAI GPT-6 through Codex. The exact model ID and context-window size are not exposed to this session. Used reasoning, repository tools, code execution, and browser verification. No subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
16b7db35ff |
Shorten planning skills and measure task decomposition (#15296)
## Thinking Path > - Paperclip manages work for AI agents. > - Planning guidance helps agents choose owners and dependencies. > - The runtime skill favors few tasks, but the catalog skill requires a child-task breakdown. > - Both add repeated process instructions that can distract from the requested outcome. > - This change keeps the ownership and dependency rules and removes the required matrix and repeated checklist. > - A bounded Product E2E comparison measures saved outcomes and task handoffs before qualification. ## Linked Issues or Issue Description Refs #11057. Related measurement work: #15218. **What existing behavior does this improve?** Planning and delegation through the runtime plan-to-tasks and bundled task-planning skills. **Current behavior** The two skills contain about 1,900 words and conflicting guidance on whether plans require child tasks. **Proposed behavior** Keep cohesive work with one owner. Split only for a real owner, parallel output, dependency, independent review, or follow-up lifecycle. Preserve existing authorization and planning mechanics. ## What Changed - Shorten both skills to about 400 words combined. Preserve their keys and installed-version behavior. - Remove the duplicate operational-skill pointer and regenerate affected source metadata. - Add twelve explicit Product E2E cells: four scenarios with current, short and disabled planning skills. - Use the current task composer and actual create-response ID; calibrate public skill APIs and browser creation without providers. - Eliminate an observed collision in chat-test company prefixes with a per-suite sequence. - Grade saved documents, exact author/run attribution, child count, prerequisite execution order, review boundaries and completion handoffs. - Retain current skill bytes and report source, selections, run accounting and failures. ## Verification - `pnpm test:e2e:runner:typecheck`: pass. - `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks pass. - `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve local Codex cells. - Archived current skills match master `72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly. - Corrected fixture: three real public-API/database calibrations pass with zero provider runs; all 35 evaluator checks and Product E2E typecheck pass. - Setup campaign [37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253) was canceled after source review found unsupported bundled edits and automatic core reinstallation. Its paid-cell step was skipped: zero provider runs, no behavioral grade. - The next setup [37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799) failed before task creation on the old title-field selector: zero actual runs, original FAIL retained, cleanup passed. A real browser/API calibration of the new helper passes with paused non-provider agents and zero runs. - Full local typecheck/build pass. Full local tests retain one unchanged five-minute Git streaming timeout (also fails isolated), 9,591 passes and 5,796 skips. CI's chat failure was a proven random fixture-prefix collision; five affected cases pass after the test-only repair. - Paid behavior comparison and new-head CI/review remain pending. This PR remains a draft. ## Risks - The shorter text may change delegation decisions. Live outcomes are not yet qualified. - The initial comparison uses one profile and one attempt per cell. It cannot establish cross-model reliability or cost trends. - Disabled means unassigned company-owned copies; the company library remains discoverable. This does not qualify global removal, automatic accepted-plan wiring changes, or installed-copy migration. - Skill availability does not prove a model read or cognitively used it. - No provider/tool protocol, permission, timeout or runtime lifecycle behavior changes in production. ## Model Used OpenAI Codex (GPT-6), with repository inspection, code editing and tool use. The exact backend model ID and context-window size are not exposed in this session. The declared eval model is native Codex `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
44e4979d23 |
Capture project and repository update lifecycle events (#15306)
## Thinking Path > - Existing lifecycle capture records agent transitions and project creation. > - Project and repository edits need matching project update records. > - The record must commit with the mutation so failed writes cannot lose a hook. > - Creation with repositories is one creation, and repository replacement is one update. > - Archive-only changes preserve project state without creating a hook. > - Plugin consumption and provider behavior are separate work. ## Linked Issues or Issue Description **Problem or motivation** Project edits and repository/workspace changes lack durable lifecycle records. A future resource plugin needs those changes captured alongside the existing project creation hook. Archive-only changes must produce no hook. **Proposed solution** Allow project `update` records in the existing lifecycle journal. Record project and workspace mutations in their database transaction while holding the project row lock. Suppress intermediate workspace hooks during project creation and aggregate repository replacement. **Alternatives considered** Route-only hooks miss shared service callers. Recording after commit can lose an event. Emitting a hook for each child mutation exposes intermediate repository state. **Roadmap alignment** This completes project lifecycle capture begun in #15280. Plugin delivery, VM/volume provisioning, and backfill remain separate. Searches found no duplicate project lifecycle work; related #13306 concerns decision events on the in-process plugin bus. ## What Changed - Record project edits and workspace additions, updates, and removals as project `update` events. - Commit each event atomically with its mutation under the project row lock. - Keep project creation with repositories to one creation event and repository replacement to one aggregate update. - Ignore archive-only changes and retain workspace records. - Extend the journal action constraint and document project update capture. ## Verification - `pnpm -r typecheck` passed on the narrowed scope. - 57 tests passed across six lifecycle, project, repository, and chat-project suites. - Seven managed-sandbox workspace route tests and the CLI lagging-worktree migration regression passed (65 targeted tests total). - Full GitHub CI passed on `8ba6f97f22`; all required gates are green. - Greptile scored the final project-only commit 5/5 with zero unresolved review threads. - The branch is current with `master` and has no merge conflicts. - `git diff --check` and a local secret/PII scan passed. ## Risks - Apply migration `0300_chunky_chamber.sql` before running the new server. It permits project update actions and tolerates older JavaScript worktree backups that omitted the prior CHECK constraint. - Event-write failure intentionally rolls back the project or repository mutation. - Records contain identity and action; future consumers must load current authorized project/workspace data. - Plugin consumption, provider calls, volume cleanup, and backfill are outside this PR. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code execution, and tool use. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1477d1ecea |
test: remove the no-op sequential describe modifier (#15286)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs the server and runner test suites with Vitest. > - Vitest 5 removes the deprecated `describe.sequential` property, so the pending Vitest 5 upgrade fails the type-check and test jobs. > - `describe.sequential` only changes behaviour inside a `describe.concurrent` suite, or when `sequence.concurrent` is on. > - This repository has neither, so the modifier changed nothing at run time. > - The benefit is that the Vitest 5 upgrade can land, and the test files lose a modifier that did no work. ## Linked Issues or Issue Description Refs: #12969 ## What Changed - Replace every `describe.sequential` use with a plain `describe` call. - Drop the `{ concurrent: false }` suite option from the two runner test files. - Add a comment to `server/vitest.config.ts` that records why these suites must run one test at a time. - Leave the package manifests and the lockfile unchanged. ## Why the modifier did nothing The Vitest documentation states that `describe.sequential` is useful to run tests in sequence inside a `describe.concurrent` suite, or with the `--sequence.concurrent` option. `sequence.concurrent` defaults to `false`. This repository satisfies neither condition: - No test file uses `describe.concurrent`, `it.concurrent`, or `test.concurrent`. - `server/vitest.config.ts` sets `sequence.concurrent: false`, with `maxWorkers: 1`, `maxConcurrency: 1`, and `isolate: true`. - `packages/paperclip-runner/vitest.config.ts` sets no `sequence` block, so the `false` default applies. `packages/db` and `cli` already run the same embedded-Postgres suites with a plain `describe`, and those jobs are green. The server package was the only outlier. The modifier did carry one real piece of knowledge: these suites need their tests to run one at a time. The new comment in `server/vitest.config.ts` records that reason next to the setting that enforces it. ## Verification - `git grep` for `describe.sequential` returns nothing outside `node_modules`. - The author ran the changed server test files under the installed Vitest 4, and the results match the results without this change. - Two very large embedded-Postgres test files exceeded the author's local memory limit, so the CI test jobs cover those two. - The two changed runner test files have pre-existing local failures caused by a missing Rust toolchain and a missing global `pnpm` binary. The failures are identical with and without this change. - The author type-checked the changed files and found no new error. - CI must pass the typecheck, build, server test, and runner verify jobs. ## Risks - Low risk. Suite execution stays serial, because the Vitest config enforces it. - The change adds no dependency and changes no package manifest or lockfile. - A future change that turns `sequence.concurrent` on would break these suites. The new config comment warns against it. ## Model Used - Claude Sonnet 5 — code edits and local verification. - OpenAI Codex, GPT-5 — the earlier revision of this branch. ## Test plan - [x] Every CI check reaches a terminal green state. A pending or queued check is not a pass. - [x] The `Typecheck + Release Registry` job passes. This change must not introduce a type error. - [x] The `Build` job passes. - [x] The server test jobs and the runner verify jobs pass. - [x] Greptile re-reviews this commit set and posts a passing verdict. The dependabot waiver does not apply to this pull request. - [x] `mergeable` reads `MERGEABLE` as a terminal value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the related public issue with `Refs: #12969` - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Priya Raman <priya.raman@paperclip.ing> --------- Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com> |
||
|
|
90182b4f8b |
Report bounded Cloud portfolio failure diagnostics (#15340)
## Thinking Path > - Paperclip manages work for AI agents and their companies. > - Cloud instances proxy a trusted user's portfolio through the control plane. > - A failed request returns a generic error to the client. > - That replacement error loses the failure phase and network code in Sentry. > - Operators need bounded evidence without upstream messages or credentials. > - This PR adds safe diagnostics while preserving the existing request behavior. ## Linked Issues or Issue Description Refs #10850, which added the portfolio proxy. No open PR for this diagnostic gap was found. **What happened?** A rejected portfolio fetch becomes a generic 502 in Sentry. The event cannot distinguish a connection reset, deadline, HTTP response failure, or body failure. The route's previous warning also included the original error and a stack identifier. **Steps to reproduce** Make the portfolio proxy's fetch reject with a TypeError whose cause has `code: ECONNRESET`. The client correctly receives the generic 502, but the captured replacement error loses that code. **Expected behavior** Keep the existing client response. Attach only bounded server-side diagnostic fields to the failure event. Do not retry the request or expose the original error. ## What Changed - Add a typed portfolio error with a private frozen diagnostic record: phase, upstream HTTP status, elapsed milliseconds, and an allowlisted network code. - Read at most four error/cause objects through own data properties. Unknown codes, messages, getters, and out-of-range values do not enter the record. - Replace the route's raw-error warnings with safe fields. Send a plain error plus event-local context through the existing optional Sentry gate. Keep the route callsite and default fingerprint policy. - Preserve authentication, trusted headers, cookies, exact HTTP error bodies, cache behavior, the ten-second deadline, and one fetch per request. Public responses receive no diagnostic fields. - Document the fields and test HTTP behavior, privacy, and event isolation with the real Sentry SDK. ## Verification - `PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1 pnpm exec vitest run server/src/__tests__/cloud-portfolio-error.test.ts server/src/__tests__/cloud-routes.test.ts server/src/__tests__/sentry.test.ts server/src/__tests__/run-failure-sentry-real-sdk.test.ts`: 65 passed, with the audited optional SDK installed. - Full local `pnpm -r typecheck` and `pnpm build` passed on Node 24.21.0 and pnpm 9.15.4. - Independent review found no blockers and independently passed all 65 tests, including the real SDK checks, on this exact commit. - Full Linux CI passed on this exact commit and provides aggregate suite coverage (54 successful checks, 2 intentional skips). A duplicate full local aggregate was not run. - The first SDK contract job failed before tests when npm could not resolve an OpenTelemetry transitive package. A subsequent empty-cache install first encountered a missing tarball, then succeeded after the registry artifact became available. All 6 real-SDK tests passed against that fresh install; the single unchanged-head CI retry passed. The SDK pin, workflow, and dependency files are unchanged. - The first browser shard 8 run timed out waiting for the inbox retry reply after 45 seconds. The unchanged isolated case passed (1/1), and the test, UI handler, fixture, and recovery files match the base commit. The failed log contains no wakeup POST before the test's immediate navigation; a navigation/request timing race is suspected but unproven without a trace. The single unchanged-head shard retry passed (20 passed, 1 skipped); no timeout or source change was made. - Greptile reviewed this exact commit at 5/5 with no unresolved review threads. - No live portfolio request was replayed. Route tests use controlled local upstream responses. - The added diff passed the secret and PII scan and `git diff --check`. ## Risks This is a diagnostic change. It does not identify or repair the origin of a connection reset. Unknown transport failures remain `unknown`. Elapsed values outside 0–60,000 ms become null. Default Sentry fingerprinting remains enabled; exact historical group membership is not guaranteed. No retry, migration, deployment, or configuration change is included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and independent agent review. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9b3fe260ba |
fix(tasks): surface Codex ChatGPT model rejection (#15299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native Runner runs report provider errors and can also save a structured failure result. > - Codex rejects a model that the user's ChatGPT account cannot use. > - Master now diagnoses this rejection, but the task's compact run data omits its message. The thread can still call it a generic run failure. > - Users need the account restriction and a clear step to repair the model selection. > - This pull request shows the existing diagnosis as “Model unavailable” on the task. ## Linked Issues or Issue Description **What happened?** Codex returns HTTP 400 with `invalid_request_error` and the message `The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account.` Master now stores an actionable diagnosis for this rejection. The task thread still labels it “Run failed” and does not receive its error text in compact run data. **Expected behavior** Show the account restriction on the task and in the run error. Tell the user to choose a supported model or clear the task's model override before retrying. **Steps to reproduce** 1. Use a Native Runner Codex agent signed in with a ChatGPT account. 2. Select a model that produces the rejection above and start a task. 3. Let the runner save its generic failed result. Inspect the task's failure marker and recovery notice. **Additional context** Refs: #15304. That merged PR diagnoses the provider failure and preserves worker and review recovery rules. This PR adds its task-facing message and guidance without changing that diagnosis or those rules. Refs: #13134. That PR improves model discovery for ChatGPT accounts. This PR exposes the rejection when a configured model still fails at execution time. ## What Changed - Return a bounded model rejection message in compact issue-run data. - Show “Model unavailable” and model-change guidance in the task thread and recovery notice. - Add database and UI regression tests, including both the original rejection text and master's fixed diagnosis. Document the new failure message. ## Verification - Focused merged-branch validation: 279 tests passed across activity service, native provider failure observation and PostgreSQL integration, TaskChatThread, and ExecutionBlockerNotice. - `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` passed on the merged source tree. - Greptile reviewed conflict-resolution commit `1db39c4e636a92e2e76ba224b71c6f1e2f556c81` at 5/5 with no findings or inline comments. - All 55 check runs completed without failure on that commit. The two optional Storybook jobs were skipped. The legacy Snyk status passed. The branch has no merge conflicts. - UI regression assertion: the task's failure marker says “Model unavailable” and retains the account restriction. It no longer says that this failure happened after a final response. ## Risks - Provider recognition and recovery are owned by the existing master implementation. This PR exposes only bounded error text for failed runs with `native_provider_model_rejected`. - Historical runs with the generic `adapter_failed` code are not reclassified. No stored run is rewritten. - Existing retry and reconciliation gates remain in place. No schema migration or model configuration change is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser inspection. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
63f3aa2dbf |
fix(runner): continue restart-interrupted Codex turns (#15297)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs retain their provider conversation across server restarts. > - A dead runner can restore a Codex conversation after its active turn is lost. > - The old recovery path synthesized a failed task result from that interruption. > - The task then required operator action even though its conversation and workspace were available. > - This pull request preserves the interruption cause and uses the admitted restart attempt for one continuation in the same conversation. > - The agent can reconcile unfinished actions and complete the current request without resending the original task. ## Linked Issues or Issue Description Related: #12845 added native restart recovery. #15042 covers admission during shutdown. #14796 covers legacy shutdown recovery. This change covers a lost native Codex turn after successful conversation restoration. **What happened?** After a server restart killed the local runner, Paperclip restored the saved Codex thread. Runnerd found that the old turn was no longer active. It synthesized a failed terminal and a needs-review result from the last progress message. The transport also discarded the terminal error when it reconstructed thread history. The task failed instead of continuing. **Expected behavior** After proving that the old process stopped and admitting a bounded recovery attempt, resume the current request in the same conversation. Preserve the workspace. Inspect unfinished actions before proceeding. Keep real provider failures, accepted results, intentional stops, unknown unreconciled effects, and exhausted attempts subject to their existing rules. **Steps to reproduce** 1. Start a local native Codex run and leave its turn active. 2. Kill the isolated runner and provider processes, as can happen during a server restart. 3. Restore the same provider thread with no active turn. 4. Observe the synthetic task failure. The new real-process regression reproduces this boundary with a scripted provider. ## What Changed - Record an explicit recoverable process-loss cause without inventing a task result. - Preserve terminal errors and prior turns in reconstructed provider history. Recover the authoritative saved result when adopting an accepted continuation. - Send one continuation in the same conversation for an admitted dead-runner recovery. Require reconciliation of unfinished commands and external actions. - Persist the interrupted terminal before submission and retain the existing recovery marker across another controller loss. - Keep provider attempt limits, terminal failures, and intentional cancellation behavior. - Add red/green regressions, real process-kill coverage, restart checkpoint coverage, and retry-budget coverage. Document the behavior and run-log evidence. ## Verification - Red: the new native runtime regression rejected with `NativeProviderTerminalFailure` on the original code; the Rust restore regression found a missing recovery cause. - Red/green: if restoring the conversation fails and replacement is allowed, the replacement receives the full task and fresh-session handoff. Both prepared and legacy execution inputs are covered. - Red: a second controller crash after the provider accepted the continuation caused an extra `turn/start`. The regression now proves there are exactly two submissions total: the original and its continuation. - Green: focused runtime, backend, driver recovery, and real-process restart suites (207 tests). After the final history/result changes, driver recovery and real-process restart suites passed again (36 tests). - Green: complete Codex transport suite (186 tests), server restart classification/database integration suites (34 tests), and Rust Codex provider suite (92 passed, 2 ignored). - Full `pnpm -r typecheck` and `pnpm build` passed on `60141e649`. The subsequent replacement-prompt guard passed the Runner TypeScript check and the complete runtime plus process-restart suites (148 tests). - Local full-suite attempt: `pnpm test:run` reported two failures in the untouched chat integration suite. Both passed individually, and the complete chat suite passed on rerun (1,063 tests). After all remote test shards passed, the duplicate serial local run was stopped with SIGINT; it is not claimed as a full local-suite pass. - Latest-head CI (`b1297dcd4`): 55 successful checks and 4 intentionally skipped checks, including all test shards, typecheck, build, native Runner verification, end-to-end tests, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37405346303). - Greptile: 5/5 on the latest head, with no unresolved review threads. The PR is mergeable. - The process tests use the real runner binary and a scripted Codex provider. They do not call a live model service. ## Risks - This changes local Codex recovery after process loss. A continuation can execute more work in the retained conversation. Its prompt requires state inspection before repeating an uncertain action; the system does not replay tool calls. - Recovery shares the existing three-attempt budget and one-shot continuation marker. Real failures and older unmarked failed checkpoints are not reopened. - No database migration or API change is required. ## Model Used - OpenAI Codex, based on GPT-6. The exact model ID and context-window size are not exposed in this session. - Capabilities: reasoning, source inspection, tool use, code editing, code execution, and test analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a1ab55a56d |
fix(agents): grant configuration access by default (#15283)
## Thinking Path > - Paperclip manages agent work and permissions for a company. > - Agent setup can require one agent to configure another agent. > - New standard agents could create agents, but they had no direct configuration grant. > - Existing agents must keep their current permissions after an upgrade. > - The requested new-agent defaults include 13 more direct permissions for suggestions, skills, tools, audit, inbox, and task assignment. > - This pull request adds the 14-grant set at creation and approval activation, retains it through invitations, and keeps existing agents unchanged. > - The change keeps low trust and built-in agents on their narrower permissions. ## Linked Issues or Issue Description **What existing behavior does this improve?** The agent creation and approval flows, and the agent configuration authorization path. **Current behavior** A new standard agent can create an agent. It cannot make protected changes to a peer agent unless an operator adds an `agents:configure` grant. **Proposed behavior** Add direct grants for `agents:configure`, `agents:suggest-changes`, `skills:create`, `skills:suggest-changes`, `tools:manage_connections`, `tools:manage_profiles`, `tools:view_audit`, `audit:view_agent_actions`, `tools:use`, `tools:manage_runtime`, `inbox:manage`, `tasks:assign`, `tasks:assign_scope`, and `tasks:manage_active_checkouts` to new standard agents. Scope `tasks:assign_scope` to the agent’s reporting subtree. Keep existing agents and their grants unchanged. **Reason and benefit** New standard agents can complete agent setup and the requested tool, skill, inbox, and task workflows under existing route, scope, and approval checks. **Breaking changes** Existing agents keep their current permissions. There is no permission migration. Related PR: #12212 adds scoped grant routes. This PR changes the default grant. ## What Changed - Add the 14 requested direct grants in agent creation and approval activation. Keep them when invitation approval replaces grants, preserving explicit scopes. Remove them when an agent is deleted. - Apply defaults only to new agents. Remove the existing-agent permission migration, its snapshot, and its journal entry. Preserve existing scoped grants. - Add permission, scope, invitation, and existing-agent regression tests. Document all default and excluded permissions. - Prevent agent keys from creating, changing, or restoring host-executed process or local adapter command settings, and from restoring workspace commands through rollback. ## Verification - Run `node_modules/.bin/vitest run server/src/__tests__/agent-default-configure-grants.test.ts server/src/__tests__/invite-join-grants.test.ts packages/db/src/migration-snapshot-drift.test.ts`. All 15 tests pass. These cover new-agent defaults, unchanged existing grants, pending approval, invitation grants, and migration history. - Run `node_modules/.bin/tsc -p server/tsconfig.json --noEmit`. It passes. - Run `packages/db/node_modules/.bin/tsx packages/db/src/check-migration-numbering.ts` and `packages/db/node_modules/.bin/tsx packages/db/src/check-migration-safety.ts`. Both pass. - Confirm that this PR has no files under `packages/db/src/migrations/` in its final diff. - GitHub CI at `e8173b2c41` passes build, typecheck, server suites, browser shards, and canary verification. One unchanged OpenCode transport test reached its five-second timeout in the first Runner shard run. That test passes locally in 2.55 seconds. The shard passed on one retry. All 54 checks pass, with two expected skips. - Greptile gives this exact head 5/5. Security review passes. The PR has no unresolved review threads or merge conflicts. ## Risks - The 14 default permissions apply only to new standard agents. Existing permissions stay unchanged. Low trust and managed built-in agents are excluded. - New grants include connection, runtime, and active-checkout management. Existing company, responsible-user, scope, and approval checks remain in force. The default inbox grant carries a responsible-user-only scope so it cannot override another user's inbox settings. - Protected changes still need the responsible user's authority when that check applies. Company boundaries and approval gates still apply. - Agent-authenticated requests cannot configure process adapters or host command settings on local adapters. Known provider credential references remain allowed, as do narrow plain authentication overrides on new peers when the caller uses an AI connection pool. Arbitrary environment settings remain blocked. Board operators retain the host configuration paths. > I checked `ROADMAP.md`. This is a narrow fix to existing agent configuration behavior. ## Model Used - OpenAI Codex, GPT-6 series. This runtime did not expose its exact hosted model ID or context window. This revision uses OpenAI GPT-6 through Codex with tool use, code execution, and tests. The runtime does not expose the exact hosted model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f858207161 |
Guard routine UUID lookups without narrowing valid inputs (#15313)
## Thinking Path > - Paperclip manages work for AI agents and their companies. > - Routines expose details, triggers, and management actions through resource IDs. > - These resource IDs are UUIDs. The UI and CLI do not resolve short prefixes. > - A malformed ID reaches a PostgreSQL UUID comparison and returns a server error. > - A lookup guard can return the existing not-found response before that query. > - This PR preserves valid PostgreSQL UUID input forms and the existing access checks. ## Linked Issues or Issue Description Refs #11471. This is a credited continuation of the lookup-guard approach from @mv2woods. That older PR remains open and unchanged. Its review requested the required PR description sections. This continuation uses current master and adds UUID compatibility and authorization coverage. The earlier `isUuidLike` guard restricts UUID versions to 1–5 and trims input. The database currently accepts other UUID values and input forms. This guard preserves the existing [PostgreSQL UUID input contract](https://www.postgresql.org/docs/current/datatype-uuid.html), including uppercase, paired braces, omitted hyphens, and hyphens after groups of four digits. It rejects whitespace and short prefixes without rewriting the value sent to the database. ## What Changed - Guard the shared routine and private-trigger UUID lookups. Malformed resource IDs return null, so existing routes return their normal 404 response. - Keep company access, assignee permissions, body validation order, and database error propagation unchanged. - Add real PostgreSQL tests for stored UUIDv4, UUIDv7, nil, and max values in six input forms. Add no-query checks for malformed input and HTTP coverage across all 15 root-resource route handlers. - Document complete resource IDs and distinguish them from opaque public webhook IDs. ## Verification - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts server/src/__tests__/routines-e2e.test.ts`: 91 passed, including the new PostgreSQL and HTTP regressions. - Independent review found no blockers and passed all 9 focused PostgreSQL and HTTP regression cases on this commit. - Full local `pnpm -r typecheck` and `pnpm build` passed on Node 24.21.0 with pnpm 9.15.4. - Full Linux CI passed on `0a2990eca2f904335821027a322e4db868151eb3`: 53 successful checks and 2 inapplicable Storybook skips. No retries were needed. This provides aggregate suite coverage; a duplicate full local aggregate was not run. - [Greptile reviewed this exact commit](https://github.com/paperclipai/paperclip/pull/15313#issuecomment-6010222312) at 5/5 with no findings or unresolved threads. The PR is ready for review and has no merge conflicts. - `git diff --check` and the added-diff secret and PII scan passed. ## Risks Malformed routine and private-trigger IDs now return 404 instead of a database error. Full UUIDs retain the existing company and assignee checks. This change does not add prefix lookup, alter public webhook IDs, or validate unrelated nested revision/thread IDs or query parameters. No migration is required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and independent agent review. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1d5a439e2b |
Diagnose native model rejections and preserve repairable reviews (#15304)
## Thinking Path > - Paperclip manages agent execution and recovery. > - A native provider can return a failed turn with an accepted result. > - This path replaced a structured model rejection with a generic failure. > - Finalization could also queue the same failed request for recovery. > - This change retains a bounded model and authentication diagnosis. > - Workers wait for a configuration fix; reviews retain their assignment. ## Linked Issues or Issue Description **Describe the bug** A native Codex turn can report HTTP 400 because the selected model is not supported by its ChatGPT connection. When the turn also returns a failed semantic result, the run shows `Native session failed` and `adapter_failed`. The native failed-result finalizer can create a retry intent before the adapter error is saved. **To Reproduce** Use a native Codex session that returns a structured `invalid_request_error` model/ChatGPT rejection and an accepted failed result. The regression tests reproduce this with a stub provider and a real isolated PostgreSQL database. No provider account or live request is needed. **Expected behavior** Show an actionable error and stop automatic retries of the rejected request. Block the current worker for a configuration fix. Preserve a pending review so its reviewer can explicitly retry after repair and resolve the original assignment. A stale run must not block a new task owner or override a completed review. **Version and deployment mode** The affected source path is present at `858094ba81`. This change applies to native Codex execution in local and remote environments. Related work: #13853 covers CLI protocol compatibility, #14942 refreshes runtime/model pins, and #12767 proposes legacy adapter model validation. This change concerns accepted native failed results and does not replace those changes. ## What Changed - Recognize only a bounded structured HTTP 400 provider error for the exact selected model and failed turn. - Save a fixed actionable error and four fixed diagnostic fields. Exclude raw provider text and model names from the new diagnostic. - Re-derive finalization evidence from the pinned execution and committed runner event, including its canonical hash and identity bindings. - Block only the current worker. For a current pending reviewer, preserve the original review decision and status version, release execution, and create no automatic recovery action. Recheck execution ownership under the issue lock. Preserve superseded and completed reviews. - Keep unknown failures, accepted results, cleanup, and ownership fences intact. Advance the status policy version to `phase6-v8`. - Route this exact failure code through the existing configuration recovery path. Permit an authenticated Board retry of this native failure only for its saved, still-pending review; discard caller review markers and revalidate at dispatch. ## Verification - Node 24: `pnpm -r typecheck` and `pnpm build` passed. - Independent focused validation: 222 tests passed across the new observer and PostgreSQL integration suites, status arbiter, diagnostic redaction, recovery classification, wake policy, and wake use cases. - The integration tests use actual result persistence, finalization, and the wake dispatcher. They cover zero retry dispatch for an authorized rejection, repeated reconciliation, pending/superseded/resolved reviews, reassignment, and a same-agent successor at an unchanged status version. - Final-commit verification: full typecheck and build passed; independent route/integration/arbiter/review validation passed 136 tests with no remaining source findings. This includes authenticated retry refusals, actual heartbeat retry admission, atomic queued-review claim, and acceptance of the original review. - The local aggregate test run was stopped to avoid duplicating the canonical Linux CI aggregate on this development host. Its partial log showed no failures; no full local-suite pass is claimed. The full hosted Linux CI aggregate passed on the final commit. - No live provider calls, runtime changes, or schema changes were made. ## Risks The classifier is intentionally narrow. A changed provider response format remains an unknown failure. No global model restriction is introduced. The worker blocker requires current task authority and does not replace newer assignments or completed review decisions. A pending review remains in review and requires an explicit retry after configuration repair; it is not rebound to a new status decision. The retry decision is durable. There remains a small existing gap between status/effect commit and diagnostic metadata persistence; a crash there can leave a generic diagnostic while preserving the blocked decision. No schema migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The exact serving model suffix and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #123` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal ticket id or instance-derived details - [x] I have run focused tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d21fe31243 |
Retry pooled execution projection reads after disconnects (#15300)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task run lists show current execution and recovery state. > - Their execution projection reads runs, finalizations, pending interactions, and recovery records. > - A temporary database disconnect in one lookup makes the run list fail. > - This pull request uses the existing bounded retry policy for these read-only lookups. > - It keeps transaction reads, authorization, and liveness writes outside the retry scope. ## Linked Issues or Issue Description Refs #15248. This extends the same bounded SELECT retry policy to the execution projection used by task run lists. A wrapped `CONNECTION_CLOSED` error from the pending-interaction SELECT makes `GET /issues/:id/runs` fail. These lookups are safe to repeat on the pooled database connection. The shared projection helper also serves transaction callers, so retries must remain an explicit opt-in. ## What Changed - Opt the activity run list into per-query projection retries. Keep the default path single-attempt for transaction callers. - Rebuild only the failed SELECT. Use the existing three-attempt limit and 50/100 ms delays. - Keep completed reads, route authorization, and liveness backfill writes outside the retry callbacks. - Test each query, exhaustion, exact error preservation, non-transient errors, empty results, transaction defaults, run-list backfill behavior, and denied route access. - Document the pooled-read boundary. ## Verification - Frozen install passed with Node 24.21.0 and pnpm 9.15.4. - 92 focused tests passed across execution projection, activity service/routes, prior service-read retries, and the shared retry policy. - Independent review passed 56 projection, retry-policy, and route authorization tests with no blocking findings. - `git diff --check` and the staged diff and PR-text secret/PII scan passed. - Full workspace `pnpm -r typecheck` and `pnpm build` passed. - Greptile scored exact head `78783e84a5` 5/5 with no findings or open review threads. - Initial CI had one Runner test timeout: `stops and closes a timed-out consumer even when the caller requested a warm session` exceeded 5 seconds. Its fixture uses a 1 ms session timeout. The Runner package is unchanged from the base. The focused case passed locally (1 passed, 138 skipped). One unchanged-head rerun of that job passed all 1,219 Runner tests and the additional native-runtime integration test. All other CI lanes passed on the initial run. - Full Linux CI passed on exact head `78783e84a5`: 53 successful checks, two intentional Storybook skips, and no pending checks. It provides aggregate test coverage. A full local aggregate was not launched. The same host's prior aggregate exposed three known macOS immutable-cache rename failures in unchanged company-skills code, documented in #15298. ## Risks - An affected lookup can take up to two additional connection attempts plus 150 ms of delay. Persistent failures still return the final error. - The opt-in is only used by the pooled activity read path. No transaction, mutation, full request, or provider action is replayed. - No schema or response-shape changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and independent agent review. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; aggregate coverage is described above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
858094ba81 |
Fix browser polling for unsaved agent chats (#15298)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat uses the shared task surface and browser panel. > - An unsaved chat has a temporary ID instead of a stored task UUID. > - The browser poll sent that ID to a PostgreSQL UUID query and returned a server error. > - This pull request waits for a saved task and validates task IDs before the query. > - Saved chats keep their browser tools and access checks. ## Linked Issues or Issue Description Refs #14331. That report describes the same draft-ID problem in email requests. This change covers the browser routes; it does not change email behavior. Opening an unsaved Agent Chat sends `chat:<agent-id>` to `/issues/:issueId/browsers` every three seconds. PostgreSQL rejects this as a task UUID. The draft is only a UI view model. The first send or upload creates its stored task. ## What Changed - Disable browser polling for unsaved chat IDs. Resume polling when the chat has a task UUID. - Return the existing task-not-found response for malformed task IDs before database access. Keep board authentication first and company and credential checks unchanged. - Test all six task browser routes, draft-to-saved polling, normal tasks, saved chats, and denied access. Keep canonical UUID v4, v7, and nil values queryable. - Document the draft and saved-chat browser contract. ## Verification - `pnpm install --frozen-lockfile` passed with pnpm 9.15.4 and Node 24.21.0. - `pnpm exec vitest run server/src/__tests__/browser-use-route-scope.test.ts server/src/__tests__/browser-use-connection.test.ts` passed all 35 tests. - `pnpm exec vitest run ui/src/hooks/useTaskBrowsers.test.ts` passed all 5 tests. - Independent review passed 14 hook and route tests plus both real task and saved-chat authorization cases. - `pnpm check:token-gates` and `git diff --check` passed. - Full workspace `pnpm -r typecheck` and `pnpm build` passed. - Full Linux CI passed on `dafff44a32`: 53 successful checks and two intentional Storybook skips. This includes all general and serialized test groups, UI and browser tests, typecheck, build, and the canary dry run. No CI retries were needed. - The full local `TMPDIR=/private/tmp pnpm test:run` started but did not complete. It was stopped after full Linux CI passed. Before the stop, three company-skills cache cases failed on macOS. A focused rerun reproduced all three as `EACCES` while renaming immutable cache staging directories. The test, company-skills service, and runtime cache files are byte-identical to the baseline previously reproduced on `e99854249c` for #15291. Other local groups were not reached; the full Linux CI run provides aggregate coverage. - Greptile scored the exact head `dafff44a32` 5/5 with no findings or unresolved review threads. ## Risks - Malformed task IDs now return 404 instead of reaching PostgreSQL and returning 500. The route does not map draft IDs to agents or tasks. - No schema, browser provider, task creation, or authorization policy changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and independent agent review. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full local run limitations are reported above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e0b63e5a5 |
feat(connections): add experimental task-pinned AI routing (#14967)
## Thinking Path > - Paperclip manages AI agents and their work. > - AI Connections separate account access from models and harnesses. > - A pool must act as one connection while retaining each task’s account. > - Core must enforce member access and preserve session and recovery rules. > - A plugin supplies rotation policy without receiving credentials. > - This change adds durable routing and native connector setup and management. ## Linked Issues or Issue Description **Subsystem affected** AI Connections, Connectors, plugins, run dispatch, and session compatibility. **Problem or motivation** Operators need to rotate new tasks across saved accounts while each task keeps its account and session. Pool setup must fit the existing connector catalog and account workflow. **Proposed solution** Add an experimental router binding, a capability-gated plugin hook, and transactional task pins. Plugins declare native pooled connectors through `aiConnectionRouter`. Core hosts the existing-account picker, ordering step, and account settings. Related usage contract: #14936. Companion private plugin: https://github.com/paperclipai/paperclip-cloud/pull/643. **Roadmap alignment** This extends Apps and AI Connections. Core supplies generic enforcement and native connector UI; the private plugin owns rotation and quota policy. The prior duplicate search found no matching router implementation. ## What Changed - Add a router binding without changing existing concrete bindings. Keep the instance flag and new pools disabled by default. Require manual operator configuration. Show no routing toggle in Experimental settings on either open-source or Cloud installs, even after routing is enabled. - Persist company-scoped pools, one shared cursor per pool, and pins keyed by company, pool, agent, and task. Commit pins and cursor advances together with revision checks and bounded retries. Persist run-ID affinity before allocation. - Pass only authorized metadata and normalized usage to plugins. Core retains credential handling, member access checks, runtime qualification, and recovery evidence. Probe outside locks with a shared 15-second budget and freshness cache. - Resolve routing before credential preparation and backend selection. Preserve pins through turns, session resets, removed members, and quota waits. Retain admitted recovery after disable or uninstall. - Separate credential session epochs from token generations. Verified refresh preserves the epoch; reconnect and manual replacement change it. Include the credential slot ID in session and usage-cache identity, so reconnecting an indexed legacy account invalidates its old session even when both epochs are zero. - Validate pool member installations before accepting saved-agent bindings and recheck compatibility when the harness changes. Install only authorized members in the new-agent transaction and record their IDs in local activity. Pool membership cannot install a restricted shared connection. - Preserve pool bindings when agents hire teammates through either creation API or native caller runtime inheritance. Block stale manager credential references; retain explicit child authentication precedence and reject incompatible inherited pools. - Add native connector registration through plugin metadata. Reuse the Connectors catalog, setup header, account header, sidebar, dialogs, and usage display. Setup selects and orders saved connections. Advanced settings hold usage rules and member runtime defaults. New-account setup opens in another tab. - Use revision-checked pool archival from the Connectors catalog and account page. Keep task pins, cursors, recovery evidence, and underlying connections. Reject ordinary connection updates or removals that bypass pool revisions. - Add pool selectors, composer models, override notes, quota status, run details, activity records, and local run-log records. Keep session-adoption copy minimal. - Show **Used by** below the pool connections. List current company agents with shared avatars and profile links. Include paused agents; exclude terminated agents and agents using another pool. - Add Core stories for the generic connector workflow and runtime surfaces. Cloud stories reuse these production routes and tokens through a preview-only alias. ## Verification - Final head `73cb953bca30ed83e4505dd820edd9b5edffd28b`: full workspace `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass locally. - All 422 focused connector/settings/shared-contract/migration tests and all 156 database-backed AI connection, hiring, reconnect, and durable-routing cases pass (69 hiring cases rerun after the final auth-precedence fix). The merged shared contract retains connection instructions and pool metadata. The pool migration is generated at sequence 0299 after the latest upstream migrations; this PR makes no lockfile changes. - All four full-app Playwright tests pass on the final head after a cold restart and migration, against the installed private plugin and isolated database, with no route or pool-API mocks. They cover hidden routing controls after manual opt-in, native pool creation, ordering, membership edits, rename, paused defaults, enabling/save/refresh persistence, stale edits, cancellation/removal, preserved underlying accounts, unavailable routers, and Used by avatars and profile links. Exact command: `PAPERCLIP_CONNECTION_POOL_E2E=1 AI_CONNECTIONS_TEST_COMPANY_ID=a37b9625-5ecf-4e29-8081-04df3d6e7d6f AI_CONNECTIONS_TEST_URL=http://127.0.0.1:3108 pnpm exec playwright test --config tests/ai-connections-app/playwright.config.ts connection-pools.spec.ts`. - [Native setup, ordering, and management screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-6006976278) address the review follow-up. [Earlier selector, quota, and run-detail screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-5971537316) show the runtime surfaces. Core previews: `pnpm --filter @paperclipai/ui storybook`, then **Connectors / Pool host** or **AI Connections / Connection pools**. Cloud owns its host-backed plugin stories; both repositories’ Operator Setup Required story assertions pass. - Live acceptance used OpenAI/Codex and Anthropic/Claude ACPX, resumed both exact sessions after restart, preserved pinned accounts through explicit reset and controlled quota deferral/recovery, and committed only two allocations across fourteen runs. A later UI-created task test again rotated OpenAI then Anthropic and resumed OpenAI through follow-up/restart/quota recovery. That later Anthropic execution was blocked by its saved OAuth token expiring (provider 401). No live usage probes ran. - The full local `pnpm test:run` was attempted earlier and did not complete because of macOS embedded PostgreSQL bootstrap/shared-memory failures and the 40,000-file Git fixture timeout. The focused database suites above now pass; full-suite verification is provided by the split CI lanes. The preceding CI run had one runtime readiness timeout; it passes locally both alone and inside the larger runtime suite. That larger local suite also encountered an embedded PostgreSQL setup failure and two macOS temporary-path alias assertions; those two assertions pass with canonical TMPDIR=/private/tmp. All final-head CI checks are terminal green, including full general/serialized server suites, Runner checks, browser E2E shards, canary verification, build, and typecheck. Greptile is 5/5 on that exact head with no unresolved threads. ## Risks - The migration adds routing tables and a credential epoch column. Install the private plugin only with the compatible Core contract. - Routing and each pool require opt-in. Production distribution and fleet defaults remain unchanged. - Unknown usage stays eligible. Known pinned exhaustion waits; revoked access requires operator repair. - Legacy adapters require compatible members. Runner model and effort overrides remain limited by qualified backend support. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and browser testing. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes:` / `Refs:` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted suites; full-suite limitations are reported above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
72ff3a9f27 |
Measure native tool context and expand bounded workflow evals (#15218)
## Thinking Path > - Paperclip manages persistent agents and their assigned work. > - Agents receive both fixed instructions and tool definitions. > - Moving a procedure into a tool description still adds model context. > - We need to measure the complete delivered catalog and test real outcomes. > - The tested reductions saved little space and introduced failing outcomes. > - This PR keeps measurement, bounded eval coverage and the original evidence. > - Production prompts, tools and runtime behavior stay unchanged. ## Linked Issues or Issue Description Refs: #15151, #14961, #14948, #14985. This adds the measurement and eval coverage needed to assess further native instruction changes. The attempted hiring and dependency reduction failed qualification and is excluded from the final diff. ## What Changed - Measure the actual standard-mode tool authority, including all 39 tools and their input schemas. Capture scripted native start, resume and continuation payloads and the OpenCode MCP declaration list. - Add OpenCode to the two explicit-only local hiring/reuse and delegation/feedback stories. Preserve the original task requests and independent oracles. - Apply one attempt per selected story and explicit company and lead-agent budget stops. - Select managed hiring credentials from the requested profile, including OpenRouter. - Run the existing Node test files under Node instead of collecting them as Vitest suites. - Retain sanitized comparison reports, original failed grades, source hashes and evidence gaps. ## Verification Final source: `2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8`. Local verification passes: - Full repository `pnpm -r typecheck` and `pnpm build`. - Eval-support unit tests: 1,252 Vitest checks and 128 Node checks. - Eval TypeScript check and six final-source measurement tests. All 36 normalized components across nine scripted deliveries and the OpenCode MCP catalog match the baseline exactly; the complete standing projection is 49,200 bytes. Fresh review of this exact head is [5/5 with no remaining findings](https://github.com/paperclipai/paperclip/pull/15218#issuecomment-5996723170); all three review threads are resolved. [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/37354539126) passes all 47 jobs, including repository typecheck, build, tests, runner checks, browser shards and the canary check. One initial annotation-test timeout is retained in attempt 1; its seven-test suite passed locally, and the affected CI lane passed on one targeted retry. No source or paid eval rerun was needed. Both ready-transition security scans passed with zero annotations. The PR is out of draft and conflict-free. GitHub still requires code-owner approval for the `package.json` test-script change; its requested reviewers are already set. A byte-for-byte comparison against master context `a65ca0950834a85bb93bcc4b4042ecacdebfef53` confirms no production changes under `packages/` or production `server/` paths. The only server addition is a measurement test. No new paid rerun is needed to compare unchanged production bytes. This does not claim that existing product defects have been fixed. The rejected corrected experiment had baseline **5 PASS / 1 FAIL** and candidate **3 PASS / 3 FAIL**, including **two newly failing pairs**. The later readiness experiment had candidate **3 PASS / 3 FAIL** and baseline **3 PASS / 2 FAIL / one setup cell without a behavioral grade**. Its five comparable pairs had two new failures, two new passes and one unchanged pass. The missing baseline Codex cell never reached its provider step because Docker setup timed out. The observed missing behaviors include parent continuation, waiting for the latest child revision, revised ZIP delivery and OpenCode credential persistence. Those failures remain failures. Source review and passing CI do not regrade them. The original reduction saved 460 bytes; its first repair saved only 125 bytes, and the larger unqualified runtime repair increased the full projection. None of those production changes is shipped here. Read the [report](https://github.com/paperclipai/paperclip/blob/2e7cef78eef7cdfe02265e0dcb03e855b8e50bd8/doc/plans/2026-10-05-native-procedure-guidance.md) and its linked sanitized receipts for exact sources, original campaign links, pair-level results and evidence limitations. ## Risks - Full-catalog bytes are not model tokens, invoices, private vendor prompts, lazy-loading behavior or truncation proof. - The two stories are explicit-only and do not prove general coding quality or arbitrary resume behavior. Single trials do not establish causation or performance trends. - The retained OpenCode candidate credential guard failed. Cleanup removed the original provider database, so the precise persistence mechanism remains unknown. The guard is unchanged. - ACPX provider-execution IDs lack proven host-call mapping. No extra-work or feedback-consumption claim is inferred by matching names, order or counts. - PostgreSQL cannot start locally while the host's shared-memory slots are exhausted. Hosted CI must supply the full database checks; the full local database suite is not claimed green. ## Model Used OpenAI Codex, based on GPT-6, with repository tools and code execution. The exact deployment model ID and context-window limit are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
efac8ff2f4 |
Add bounded Git integration restore diagnostics (#15291)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox runs must restore their workspace before finalization can succeed. > - Restore diagnostics identify the failed phase and operation. > - A Git integration exit code can still describe several different failures. > - This pull request adds fixed command labels and supported failure classes. > - Operators can distinguish these failures without collecting private Git output. ## Linked Issues or Issue Description Refs #15005. This branch includes merged #15268. It labels that change's locked ref transaction and nested branch probe without changing their behavior. **What happened?** A failed Git integration can report only `git_integration`, `unknown`, and an exit code. That evidence does not identify the failed command. Some Git versions also return exit 1 for both a merge conflict and an invalid object. **Expected behavior** Record a fixed command family and a supported failure class. Keep unknown cases as `unknown`. Exclude command arguments, process output, paths, repository URLs, filenames, and ref names. **Steps to reproduce** The tests create local repositories with conflicting commits, a missing object, and an expected-old ref mismatch. They call the real Git operations and inspect the resulting diagnostic. No hosted workspace or external provider is used. **Paperclip version or commit** Base: `e99854249c`. **Deployment mode** Built from source. The diagnostic applies to sandbox workspace restore. ## What Changed - Label Git integration calls with a closed command enum. Preserve arguments, options, errors, and retry behavior. - Recognize supported object and ref errors and OS permission codes. Require both exit 1 and completed tree output for a merge conflict. - Carry the closed fields through the existing restore receipt, saved adapter result, and Sentry projection. Revalidate saved metadata before projection. - Test nested wrappers, parallel failures, reused errors, handled probes, result settlement, and privacy with the real Sentry SDK. - Document the field contract and its limits. ## Verification - Eight focused suites pass: 295 tests. These cover Git sync, restore diagnostics, result settlement, teardown, sandbox runtime, failure projection, the real Sentry SDK, and native warm-workspace Git history. - `PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1 pnpm exec vitest run packages/adapter-utils/src/workspace-restore-diagnostics.test.ts packages/adapter-utils/src/workspace-restore-result.test.ts packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/workspace-restore-teardown.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts server/src/services/__tests__/run-failure-diagnostics.test.ts server/src/__tests__/run-failure-sentry-real-sdk.test.ts server/src/__tests__/native-workspace-sync-history.test.ts` - Local checks use Node 24.21.0 and the pinned pnpm 9.15.4. The optional Sentry SDK is pinned to the declared 10.71.0 and uses an in-memory transport. - `pnpm -r typecheck` and `pnpm build`: pass on the rebased head. - The full local `pnpm test:run` was stopped before source changes when #15268 merged and required a rebase. It reported three pre-existing company-skill cache test failures on macOS. The same three cases fail on clean bases `2c43b39167` and `e99854249c` and the earlier head with `EACCES` when publishing a read-only cache directory. The relevant test, service, and cache source blobs are identical. No cache changes are included. Later local test groups were not reached. - All 54 checks pass on exact head `8448896bc6`, including the post-ready security scan, with two intentional Storybook skips. Greptile scores this head 5/5 with no review threads. Full Linux CI covers the local groups that were not reached. - Independent review found no blocker. Its diagnostics, Git workspace, and native history suites pass 115/115 on this head. - A source comparison confirms that removing only diagnostic wrappers yields the merged upstream Git integration code exactly, including ref locks, transaction protocol, arguments, options, and retry decisions. - The clean base has stale dependency overrides in its lockfile. Local installation resolved them as the existing PR CI fallback does. No manifest or lockfile change is included. - `git diff --check` and a local secrets/PII review pass. ## Risks A Git version or localized message may not match a known failure form. Such cases remain `unknown`. Up to 16 KiB of stderr and the bounded tree-ID prefix of stdout are inspected only in memory. The saved fields contain enum values only. These diagnostics do not establish workspace recovery or authorize retries. Restore decisions and Git mutations are unchanged. No schema migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The service does not expose the exact model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e99854249c |
fix(workspaces): restore rebased sandbox history against its starting snapshot (#15268)
## Thinking Path > - Paperclip manages agents and preserves their work across runs. > - Sandbox execution restores Git history and files to the host workspace. > - An agent can rebase or amend its branch before it finishes. > - Restore currently treats the original and rewritten tips as concurrent work. > - That merge can conflict even when the host has not changed. > - This change uses the starting Git snapshot to accept rewritten history safely. ## Linked Issues or Issue Description **What happened?** A successful sandbox turn can end with `workspace_restore_failed` after the agent rebases and pushes its branch. The restore step merges the original host tip with the rewritten sandbox tip. This can recreate conflicts that the agent already resolved. **Expected behavior** Accept the rewritten history when the host still has the recorded starting branch and commit. Preserve concurrent host work through the existing merge and recovery paths. **Steps to reproduce** 1. Start a sandbox from a feature branch. 2. Rebase that branch onto an upstream commit that changes the same file. Resolve the conflict in the sandbox. 3. Restore the sandbox while the host remains on the original commit. 4. The old implementation attempts a conflicting Git merge and fails the run after the agent finishes. **Paperclip version or commit** Reproduced on `cab4263dc9` with a real Git rebase fixture. **Deployment mode** Self-hosted server with sandbox execution. Related work: #15005 records restore failure stages. #11638 preserves unrelated imported history with a graft. #10601 handles bundle prerequisites. Those changes do not distinguish a rebase from a concurrent host edit. ## What Changed - Pass the run's starting Git branch and commit into sandbox history integration. - Adopt a related rewritten tip when the host still matches that snapshot. Keep the expected-old-value ref update and bounded retry. - Verify the branch attachment inside a prepared Git transaction while Git holds its ref locks. Abort if a checkout changed the branch. - Check host and sandbox Git identity before warm reuse, including nested repositories. Restage when their tips or branches differ. - Reject host branch changes and unrelated imports after a concurrent host commit. - Retain the existing unrelated-history graft for an unchanged host and the conservative behavior for callers without a snapshot. - Export a full bundle for an intentional reset to an ancestor so restore receives the actual sandbox tip. - Add real Git and sandbox restore regressions. Document the restore contract. ## Verification - Eight Git sync, sandbox restore, and native workspace suites pass: 255 tests. - The new checkout-race and warm-reuse regressions failed before the fixes. Real Git hooks verify that prepared transactions prevent a concurrent HEAD change. - `pnpm -r typecheck` passed after rebasing onto `984f092ddf` and applying the Apex findings. - `pnpm build` passed on `f5132603d6`. - The earlier `pnpm test:run` attempt reported three `company-skills-service.test.ts` failures on macOS (`EACCES` renaming a read-only staging directory). The same failures reproduced on unchanged master. The broader run was stopped after confirming that baseline failure; it was not a full-suite pass. Those source and test files are unchanged in the current base. - [Apex review](https://github.com/paperclipai/paperclip/pull/15268#issuecomment-6002549219): 5/5 on `f5132603d6`, requested with `@greptileai apex review`. All three historical findings are addressed and all review threads are resolved. - All CI gates passed on `f5132603d6` (53 successful checks, two skipped). The signoff-policy browser fixture initially timed out waiting for a local process-agent run; its one retry passed without code changes. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37387511152) ## Risks - A recorded starting snapshot now authorizes replacement of related rewritten history. A stale or changed host tip retains the concurrent-history path. Ref writes still compare the expected old commit. - Branch changes and unrelated rewrites after host advancement require recovery instead of replacing host work. - An intentional reset to an ancestor uses a full bundle. Large histories can increase transfer time in that case. - The existing directory merge rules remain in force. The fix does not resolve an earlier failed restore or replay its external actions. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository analysis, code editing, and local test execution. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — 255 affected tests; the broader-suite baseline failure is documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4857799a88 |
feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes. Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
984f092ddf |
Add project and agent resource lifecycle hooks (#15280)
## Thinking Path > - Paperclip manages agents and projects for work. > - Plugins need reliable lifecycle hooks for these resources. > - In-process notifications can disappear during a restart. > - A lifecycle record must commit with the resource change. > - This pull request adds a generic lifecycle journal without provider calls. > - A later plugin delivery layer can use these records without losing lifecycle changes. ## Linked Issues or Issue Description **Problem or motivation** Resource provisioning needs durable hooks for agent hiring, pause, resume, termination, and project creation. Hooks must cover shared service callers, budget actions, and approval paths. Capture should work for all plugins and deployment modes. **Proposed solution** Record content-free events in the resource transaction. Deduplicate creation and unchanged status. Preserve each pause/resume cycle. Commit approval, activation, and creation together. Commit termination and API-key revocation together. **Alternatives considered** Event subscriptions alone cannot survive process failure. Cloud-only capture would exclude other plugins. Provider calls inside tenant transactions would couple resource creation to external services. **Roadmap alignment** This is lifecycle infrastructure for the existing plugin system. It adds no provider, adapter, UI, or plugin read API. Repository mutations and backfill remain separate work. Searches found no duplicate lifecycle-journal PR. ## What Changed - Add a company-scoped lifecycle journal and an additive migration. - Record hired-agent and project creation from shared services. - Record agent pause, resume, and termination, including generic updates and budget actions. - Lock agent state changes to suppress concurrent duplicate hooks. - Make hire approval, rejection, and termination transactions atomic with their lifecycle records. - Document capture, ordering, future per-plugin acknowledgments, and migration scope. ## Verification - Full workspace typecheck: `pnpm -r typecheck` passed. - Production build: `pnpm build` passed. - Focused agent, project, approval, budget, and built-in regressions: 102 tests passed across 9 suites. - Final lifecycle journal check: 14 tests passed. It covers repeated/concurrent transitions, rollback, key revocation, self-hosted capture, and company boundaries. - Review regression: 32 lifecycle and approval tests passed, including rejected-hire rollback and retry; server typecheck passed after the fix. - The initial local `pnpm test:run` overlapped the rejection fix and reported the new rollback regression against the earlier service code. A fresh lifecycle run passed all 14 tests. Fresh full local shards were stopped once the complete CI test matrix passed on the final commit. - `git diff --check` and a local secret/PII scan passed. - Greptile: 5/5 on `3ee3903f8e1a171ea3ba2bea9caa2b5766a3f807`, with the review thread resolved. - [Complete CI passed](https://github.com/paperclipai/paperclip/actions/runs/37374227222) on `3ee3903f8e1a171ea3ba2bea9caa2b5766a3f807`: general server/chat/workspace tests, serialized server suites, runner checks, build, typecheck, canary, and all end-to-end shards. The requester waived CI during the Actions outage, but the workflow subsequently completed successfully. ## Risks - The migration creates an empty table. It does not scan or backfill existing resources. - Apply the normal database migration before running this server version. Event-write failure intentionally rolls back the resource change. - Pending hires cannot bypass approval through pause or resume. - Capture works on all deployments. Plugin delivery, retention, retries, and provider actions remain separate work. No plugin can read this journal through a new API in this PR. - Future delivery must enforce company scope, track acknowledgments per plugin, and preserve resource order. A global sequence cursor can skip uncommitted transactions. - A termination hook does not authorize deleting persistent volumes. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code execution, and tool use. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6c36c07a4f |
feat(adapters): add GPT-6.1 Sol and refresh shared coding harness pins (#14942)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents run through coding-agent adapters and the native runner. Both use the same installed provider CLIs, model catalogs, and reasoning controls. > - OpenAI released GPT-6.1 Sol (`gpt-6.1-sol`) in Codex. Anthropic released Claude Sonnet 5.5. The static Codex, Bedrock, and OpenCode catalogs do not list these IDs. > - The shared provider pack pins Codex 0.156.0 and OpenCode 1.18.32. The evaluation image pins older Grok, Gemini, Kimi, Cursor, and GitHub CLI releases. Codex 0.156.0 has no bundled metadata for GPT-6.1 Sol. > - A model entry without a current harness, or a harness pin without its runner integrity checks, fails at run time. > - This pull request adds the verified model IDs and moves the harness pins, executable digests, controller checks, and image pins together. > - The benefit is that operators can select the current models, and the native and local adapters share one current CLI installation. ## Linked Issues or Issue Description Refs #13829 and #13838 (the September 22, 2026 model and harness refresh). Related pull requests: #14993 (merged October 5, 2026, superseding #14816) added the direct Claude Sonnet 5.5 entry and refreshed the Claude runtime to Agent SDK 0.3.286 / Claude Code 2.1.286. This pull request does not change the Claude runtime or the direct Claude model list; it keeps the #14993 pins and adds only the Bedrock Sonnet 5.5 ID. After #14993 merged, this branch was rebased onto `master` (October 5, 2026). The six overlapping pin regions (`docker/daytona-runner/Dockerfile`, `docker/daytona-runner/README.md`, `package.json`, `pnpm-workspace.yaml`, `packages/adapters/claude-local/src/index.test.ts`, `packages/paperclip-runner/src/backends/native-backend-factory.test.ts`) were resolved by keeping this pull request's Codex 0.160.0 and OpenCode 1.18.34 pins next to #14993's Claude 0.3.286 / 2.1.286 pins, taking the union of the Sonnet 5.5 model IDs in the Claude test, and merging both README paragraphs. The Sonnet 5.5 effort and CLI-gate lines in the Claude adapter were identical in both pull requests and merged without a diff. #14917 and #14918 reordered the Claude and Codex model lists earlier; the new entries sit where those ordering rules put them. Sources checked on 2026-10-02: - [OpenAI Codex models](https://learn.chatgpt.com/docs/models): GPT-6.1 Sol uses `gpt-6.1-sol`, supports reasoning efforts from Light to Ultra, and has Standard and Fast modes at launch. The page also records that `gpt-5.4` and `gpt-5.4-mini` retired from Codex with ChatGPT sign-in on August 31, 2026, and that `gpt-5.5` retires on October 14, 2026. Neither retirement applies to the OpenAI API. - [Codex CLI releases](https://github.com/openai/codex/releases) 0.157.0 through 0.160.0. The bundled model metadata in the 0.160.0 Linux binary contains `gpt-6.1-sol`. - [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview): Bedrock ID `anthropic.claude-sonnet-5-5`, released September 28, 2026. - [OpenCode releases](https://github.com/anomalyco/opencode/releases) 1.18.33 and 1.18.34 (fixes only). The OpenCode model registry lists both added provider-qualified IDs. - npm `latest` tags for `@xai-official/grok` 1.0.46, `@google/gemini-cli` 0.62.0, and `@moonshot-ai/kimi-code` 2.1.1. [xAI](https://docs.x.ai/docs/models), [Google](https://ai.google.dev/gemini-api/docs/models), and [Kimi](https://www.kimi.com/code/docs/en/kimi-code/models.html) list no newer coding models. - Cursor CLI 2026.10.01-e373342 is the version the official installer resolves. The pinned digest is the SHA-256 of the versioned Linux x64 archive. - [GitHub CLI 2.102.0](https://github.com/cli/cli/releases/tag/v2.102.0) (security fixes). The pinned digest matches the release `checksums.txt`. ## What Changed - Codex adapter: add `gpt-6.1-sol` to the model list, the Fast mode list, and the Ultra effort set. It is the first entry: #14918 orders the list newest version first, and its description notes the ChatGPT app lists GPT-6.1 Sol first. Update the adapter documentation text. - Claude adapter: add `us.anthropic.claude-sonnet-5-5` (Bedrock Sonnet 5.5) to the Bedrock catalog in the newest-Sonnet slot after Opus 5.5; `us.anthropic.claude-sonnet-5` moves into the older-Sonnet group, matching what `sortClaudeModels` from #14917 produces at runtime. Any Sonnet 5.5 ID (direct or Bedrock-qualified) now gets the documented `xhigh` and `max` efforts and requires Claude Code 2.1.284 or later on the CLI lane (the Claude Code changelog entry for 2.1.284 adds `claude-sonnet-5-5`). These two lines are identical to the ones #14993 merged, so the branch carries no diff for them. - OpenCode adapter: add `openai/gpt-6.1-sol` and `anthropic/claude-sonnet-5-5` to the static fallback catalog. - Codex runtime pin 0.156.0 → 0.160.0 in the root and workspace overrides, the runner package, the Codex ACP package patch, the qualified ACPX profiles, the Linux x64 executable digest, the Rust provider backend and its tests, the provider-pack manifest pins, the remote controller pins, the sandbox npm install spec, and the opt-in qualification scripts. - Remote Codex compatibility window: upper bound 0.157.0 → 0.161.0. The minimum stays at 0.149.0. - OpenCode runtime pin 1.18.32 → 1.18.34 in the runner package, the materialization script, the server and Rust qualified versions, the eval and live-session labels, fixtures, and the configuration label. - Evaluation image (`docker/daytona-runner/Dockerfile`): Grok CLI 1.0.46, Gemini CLI 0.62.0, Kimi Code 2.1.1, Cursor CLI 2026.10.01-e373342 with its digest, GitHub CLI 2.102.0 with its digest, Codex and OpenCode version probes, and the refreshed lockfile digest. The Claude Code 2.1.286 probe comes from #14993 and is unchanged here. - `pnpm-lock.yaml` is not part of this pull request. The repository's pull request gate rejects lockfile edits, and the refresh bot regenerates the lockfile on master (the same flow #13838 used). The Dockerfile `PAPERCLIP_RUNNER_LOCK_SHA256` default is the digest of the lockfile that `pnpm install --resolution-only --ignore-scripts --no-frozen-lockfile` (the refresh workflow's command) produces for the combined pins on the rebased branch (`e1856797…`); that lockfile differs from master only in the `@openai/codex` 0.160.0 platform packages, the `@anthropic-ai/claude-agent-sdk` 0.3.286 override that #14993 introduced (the open refresh-bot pull request #14872 carries that part), `opencode-ai` 1.18.34 with its Linux x64 baseline, and the `codex-acp` patch hash. - Documentation: runner README, runner compatibility doc, environment variable example, and a new `doc/adapter-model-audit-2026-10-02.md` with sources and deferred items. - Tests: Codex adapter catalog, server adapter models, Codex compatibility window, native session executor pins, runner package contract, OpenCode materialization, and UI effort options. Unchanged on purpose: Claude Agent SDK 0.3.286 / Claude Code 2.1.286 (already on `master` from #14993), ACP bridges (`acpx` 0.13.1, `claude-agent-acp` 0.73.0, `codex-acp` 1.6.2; newer upstream releases need a separate qualification), the native Grok runtime 1.0.13, Pi 0.84.2 / 0.87.1 (the Pi 1.0 runner stack covers it), and Hermes 0.19.0 (current). `gpt-5.4` and `gpt-5.4-mini` stay in the picker because the OpenAI API still serves them. ## Verification Run on Linux x64 with Node 25.9.0 and pnpm 9.15.4 after `pnpm install --no-frozen-lockfile` (the refreshed lockfile stays local; see above). The results below were re-run on the rebased head (October 5, 2026) for the suites the conflict resolution touches; the other rows are from the original run and are covered by CI on every push: - Rebased head: `packages/adapters/codex-local` 482 passed; `packages/adapters/claude-local` 340 passed, 4 failed (`execute.remote`, `test.probe`, `execute.acp-fallback`, `acp` spawn/env-hardening cases that fail identically on unchanged `master` in this host environment); `server` adapter-models + codex-runtime-compatibility + native-session-executor + adapter-registry 607 passed, 1 failed (the same adapter-registry override-pause case as before, also failing on `master` here); `packages/paperclip-runner` native-backend-factory + qualified-profiles 36 passed; `ui` codex-reasoning-effort + config-fields + model-utils 19 passed. Rust, full typecheck, build, and the Docker image are left to CI as before. - `vitest run` in `packages/adapters/codex-local`: 13 passed. `vitest run` in `packages/adapters/claude-local` (whole package, including the new Sonnet 5.5 gate and effort tests): see the latest CI run and the comment below. `vitest run` in `packages/adapters/opencode-local`: 48 passed, 1 failed (`runtime-config.test.ts` reads the host `PAPERCLIP_OPENCODE_PROVIDERS` variable; it fails the same way on the unchanged base). - `vitest run src/__tests__/adapter-models.test.ts src/services/native-runtime/codex-runtime-compatibility.test.ts src/__tests__/adapter-registry.test.ts` in `server`: 84 passed, 1 failed (`adapter-registry.test.ts` override pause test; it fails the same way on the unchanged base). - `vitest run` in `ui` for `codex-reasoning-effort`, `agent-setup-fields`, `config-fields`, and `ComposerRunSettingsPicker`: 25 passed. - `node --test test/acpx-codex-package-contract.test.mjs scripts/materialize-opencode-binary.test.mjs scripts/runner-protocol-eval-campaign.test.mjs` in `packages/paperclip-runner`: 23 passed. The package contract test verifies the installed Codex ACP executable digest and the 0.160.0 patch pin. - `vitest run src/drivers/acpx src/backends src/drivers/opencode src/live/live-session.test.ts` in `packages/paperclip-runner`: 626 passed, 5 failed, 1 skipped. The 5 failures (`installation-integrity.test.ts` `/proc/self/fd` module loading and one OpenCode answer-selection test) also fail on the unchanged base under Node 25; Linux CI runs Node 24. - `pnpm run test:opencode:qualification` in `packages/paperclip-runner` against the installed OpenCode 1.18.34 executable: passed. - `codex --version` from the installed pack prints `codex-cli 0.160.0`. The Linux x64 executable digest `12eb3e81…652aad` was computed from the `@openai/codex@0.160.0-linux-x64` archive after checking its registry `dist.integrity`. - `pnpm check:token-gates`: all gates clean. - `pnpm run typecheck:typescript` in `packages/paperclip-runner`: passed. Package typechecks ran one at a time; see the comment below for the server and UI results. Not run here, and needed from CI: - Rust tests and `pnpm -r typecheck` / `pnpm build` for the server (no `cargo` in this environment; the server typecheck prepares the runner vendor build). - The Docker evaluation image build and the real-binary Codex startup and session-resume probes (no Docker; the probes need the compiled `paperclip-runnerd`). The trusted CI runner workflow covers them. - Authenticated inference with any new model. This change is metadata and startup validation only. ## Risks - Codex 0.160.0 changes the bundled model catalog and app-server behaviour (authoritative provider catalogs, incremental running-turn tracking). The patched `codex-acp` 1.6.2 bridge is unchanged and declares `^0.148.0`; it worked with 0.156.0 under the same override. If CI probes show a protocol change, the pin can return to 0.156.0 by reverting this pull request. - The compatibility window upper bound moves to `<0.161.0`. Remote images with Codex 0.157 to 0.160 become accepted. Older images stay accepted down to 0.149.0. - Until the refresh bot lands the regenerated lockfile on master, the Dockerfile lockfile digest default does not match the committed lockfile. The trusted CI workflow computes the digest from its own resolution at build time, so this affects only a local build that passes no digest. - Existing saved model selections and effort settings are not changed. Agents on `gpt-5.4` or `gpt-5.5` with ChatGPT sign-in need a model change before the OpenAI retirement dates; that is documented, not enforced. - Rollout order: deploy the controller and runner from this change before promoting a sandbox image that carries these pins. Older controllers reject the new provider-pack pins. ## Model Used - Claude Fable 5.1 (Anthropic, model ID `claude-fable-5-1`), 1M context window, adaptive thinking, tool use. The model ran as a Paperclip agent through the Claude Code harness, performed the web research, edited the code, and ran the tests listed above. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Bender (Fable) <noreply@paperclip.ing> Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
41c1431d66 |
fix(server): remove the fixture-cleanup deadlock in the stale queued-run test (#14968)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server heartbeat tests use shared database fixtures and concurrent run execution. > - The stale queued-run test checked the database before all active executions finished. > - A late run writer could then hold a row lock while fixture cleanup held another table lock. > - PostgreSQL could report a deadlock, and the test cleanup could fail. > - This pull request waits for active executions and shares the deadlock retry helper. > - The benefit is reliable test cleanup without a production code change. ## Linked Issues or Issue Description Refs: #14198 **What happened?** The stale queued-run test could fail during fixture cleanup. Its cleanup poll could miss a run that had not written its database row. A later writer could then deadlock with the cleanup `TRUNCATE` operation. **Expected behavior** Fixture cleanup must wait for active run executions. It must retry transient PostgreSQL deadlocks and foreign-key races. **Steps to reproduce** 1. Run the server heartbeat tests with concurrent test activity. 2. Let an active run write its database row after the cleanup poll starts. 3. Observe a PostgreSQL `40P01` deadlock during fixture cleanup. **Paperclip version or commit** `master` at the base commit for this pull request. **Deployment mode** Test-only change against the server test database. ## What Changed - Replace the local cleanup poll with `drainHeartbeatRunsToQuiescence`. - Add the shared `truncateTablesWithDeadlockRetry` test helper. - Retry PostgreSQL `40P01` deadlocks and the existing foreign-key race. - Use the shared helper in both heartbeat test files. - Remove the private duplicate helper. - Change no production code. ## Verification - `vitest run server/src/__tests__/helpers/truncate-with-deadlock-retry.test.ts` passes. - `vitest run server/src/__tests__/heartbeat-active-run-output-watchdog.test.ts` passes. - The target stale queued-run test passes in the engineer validation run. - CI must confirm the full stale queued-run test file and the required server checks. ## Risks Low risk. This change affects test cleanup only. The helper keeps a bounded retry count and rethrows persistent errors. The full stale queued-run test file needs CI because the local sandbox lacks the managed Codex home required by an existing pre-dispatch credential gate. ## Model Used OpenAI Codex, GPT-5, with tool use and code review assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no documentation change is needed for this test-only change) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Test plan - [x] CI `General tests` passes for the `server` shards. - [x] CI `verify` passes. - [x] `vitest run server/src/__tests__/helpers/truncate-with-deadlock-retry.test.ts` passes. - [x] `vitest run server/src/__tests__/heartbeat-active-run-output-watchdog.test.ts` passes. - [x] `vitest run server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts` passes. ## Note on the sandbox test gap The engineer could not run the full `heartbeat-stale-queue-invalidation.test.ts` file green in the development sandbox. A pre-dispatch credential gate in `heartbeat.ts` needs a usable Codex home for `codex_local` agents, and that sandbox has no managed home. The engineer proved the failure is pre-existing: they stashed both test-file changes, reran the unmodified file, and got the identical 25 failures with the identical error. CI supplies the full-file run. Treat the CI result for this file as the authoritative check. Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
b43073d11f |
feat(connections): sync and group accounts managed by aggregators (#15254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external tools. > - Aggregator gateways can expose accounts that users already connected upstream. > - The Apps catalog did not show those accounts or their current provider status. > - Separate cards and setup tasks also made account ownership unclear. > - This pull request discovers upstream accounts and groups them under one app card. > - Users can find connected apps while each provider keeps control of its accounts. ## Linked Issues or Issue Description **Subsystem affected** Connections across the database, shared contracts, server, and board UI. **Problem or motivation** Users cannot see which apps are connected through a saved aggregator gateway. Native and upstream accounts need one app card. Discovery must preserve company, user, gateway, and credential boundaries. **Proposed solution** Sync account metadata from Composio, Arcade, and supported Executor gateways. Keep upstream account management in each provider. Use source chips and search to browse the catalog. Preserve native setup and the gateway's existing access policy. **Alternatives considered** Creating a local executable connection for each upstream account would duplicate authorization state. Using an agent task for routine Composio setup would add an unnecessary step. The board now calls the saved gateway directly for that setup. **Roadmap alignment** This extends the shipped Connected Apps and MCP Tool Gateway features in ROADMAP.md. The duplicate search found no open PR for managed account discovery. Related work: Refs #13755, Refs #13941, Refs #14725, Refs #13855. Open PR #12906 covers adjacent toolkit routing work. ## What Changed - Add provider-neutral discovery, sync, and refresh APIs. Preserve the Composio API paths. - Cache observations by company, saved gateway, viewing user, and credential version. Retain stale observations after failed or incomplete scans. - Add optional Arcade account sync credentials in the vault. Discover Executor accounts through its supported inventory interface. - Group native and upstream accounts in one app card. Imported account menus open their provider. Gateway menus own refresh and sync setup. - Add Paperclip, Composio, Arcade, Installed, and All chips. Show 50 catalog entries per page. Keep connected accounts above discovery. Keep explicit provider searches scoped. - Simplify Composio app setup and refresh its connected app list on the gateway Permissions page. - Add a compact agent access card and task creation defaults for connection setup. Preserve explicit blocks and approval policies. - Add two replay-safe migrations, service and UI tests, Storybook journeys, and acceptance stories. ## Verification - Passed the repository typecheck, full build, token gates, and migration ordering check. - Passed the focused provider adapter, connection interaction, and catalog tests after rebasing onto master. - Passed all nine database sync and migration replay tests using a disposable database on the test-drive PostgreSQL cluster. Removed that database after the run. - Verified Arcade cursor pagination against its official Go SDK and passed all eight adapter tests, including short and incomplete pages. - Passed all 45 interaction tests after making the exact requested tools and their Allowed/Ask first permissions visible before granting access. Verified the compact card in Storybook. - Passed the complete UI suite on the final code: 683 files and 7,432 tests, including the corrected Composio destination assertions. Passed 130 focused tests for the UUID, management-link, and health-status corrections. - Passed 22 Composio setup/sync tests, 23 connection-intent service tests, and the connection migration test in separate disposable databases. Database startup alone was substituted; the suites exercised their real SQL and services. - Passed all 10 OpenAPI route checks and the full-stack connection-intent browser test, including scoped consent, agent continuation, and task completion. - The local full runner encountered embedded PostgreSQL startup failures on this loaded macOS host. The earlier in-flight run also held the pre-fix Arcade transform; a fresh run of the final provider suite passes. The final-head CI is queued during GitHub’s active Actions incident: https://www.githubstatus.com/. The previous run also lost several runners simultaneously; its real catalog assertion failures are fixed and the fresh complete UI suite passes. - Tested the real test-drive server in the embedded browser with a live Composio gateway. Detected Airtable and Circleback. Verified refresh progress, account rows, source chips, search scope, and 50-entry pagination. - Arcade and Executor coverage uses provider fixtures. Live credentials were unavailable. - Storybook builds successfully and includes grouped native/provider accounts, stale and unavailable discovery, optional Arcade setup, and mobile states. The acceptance document records the simulated and live coverage separately. - Greptile reviewed final commit `217b024c27b5933e773ce9419c4e92b1032042c6` at 5/5. All six review threads are resolved, security scans pass, and the PR has no merge conflicts. The outstanding remote checks are `ci / Select trusted runner` and `review`, queued by GitHub. They need to complete before merge. ## Risks - Provider response changes can break inventory discovery. Failed scans retain observations and show stale status. - Composio scans only the supported catalog and can take time. Large inventories run in the background with progress and a bounded lease. - Arcade requires a project API key and user ID when the gateway cannot supply them. This key is used only for discovery. - Executor discovery depends on the server's exposed inventory tools. Unsupported servers report unavailable discovery. - Cached account rows do not grant access or create executable connections. Gateway policies still govern tool use. Account deletion and per-app authorization remain upstream. - The migrations add tables and one nullable column. Replay preserves existing rows and company-scoped foreign keys. ## Model Used OpenAI Codex, based on GPT-6. The session does not expose a more specific serving model ID or context limit. Used reasoning, repository tools, code execution, and browser verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cab4263dc9 |
feat(claude-local): add Sonnet 5.5 and refresh the qualified Claude runtime (#14993)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Claude local adapter lists models for users with a Claude subscription. > - The list needs Claude Sonnet 5.5 and its supported effort levels. > - Sonnet 5.5 needs Claude Code 2.1.284 or later on both execution paths. > - The qualified ACP runtime previously used Claude Code 2.1.280. > - This change adds Sonnet 5.5 and pins Agent SDK 0.3.286, which includes Claude Code 2.1.286. > - Users can select the model and run it with a qualified runtime. ## Linked Issues or Issue Description Refs #3936. This replaces #14816 because the maintainer integration cannot write to the contributor fork. Thank you to @SkilLab-Tech for the model support, runtime refresh, tests, and platform digest verification. This branch preserves both original commits: `8d2f3261af61a2ac1120e51e8a8618732ace543b` and `e68d1d002a3ed745f016fb11c50ac5a3c5a9ff8d`. Related work: - #14917 added Claude model ordering. This branch includes that merged change and resolves its conflicts with #14816. - #14942 updates the other models and harnesses. It remains separate. Its matching Sonnet effort and CLI-gate changes are identical. Both PRs merge with master. The second PR will need a rebase after the first merges because adjacent runtime-pin and test edits conflict. - #14954 is another Sonnet 5.5 change. It overlaps with the model additions but does not include the qualified runtime refresh. - #14039 makes the per-task effort picker model-aware. #3937 is also related to effort selection. The original author checked the [Claude Code changelog](https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md) and [effort documentation](https://platform.claude.com/docs/en/build-with-claude/effort) on 2026-10-01. ## What Changed - Add the direct `claude-sonnet-5-5` model and Low, Medium, High, X-High, and Max effort levels. - Require Claude Code 2.1.284 or later for that model on the CLI path. - Put Sonnet 5.5 after Opus 5.5 in the current-model group. Keep Sonnet 5 in the older-model group. - Retain the Sonnet 5.5 assertions and the model-order assertions in the server tests. - Pin Agent SDK 0.3.286 and Claude Code 2.1.286 across overrides, integrity digests, qualified profiles, Rust provider pins, and the Daytona version check. - Update the related adapter and runtime documentation. ## Verification Local verification uses the resolved source tree and pnpm 9.15.4. Model tests passed on Node 25.9.0. Runtime integrity tests use CI's Node 24.21.0. - Five focused Claude test files pass: 62 tests. They cover model defaults, model ordering, CLI gates, and remote execution probes. - Server model-list and UI setup tests pass: 30 tests. - The Claude adapter typecheck passes. - Runner integrity and qualification tests pass on Node 24.21.0: 78 tests. Four descriptor-loader tests fail on Node 25.9.0; all four pass on the CI version. - The runner package contract passes: 10 tests. - Full local typecheck stopped with exit 137 in the database package under the container's 4 GB memory limit. The production build reached the runner Rust build, then stopped because `cargo` is absent. - The full stable local Vitest run was stopped after all current-head CI test shards passed. It did not complete locally. The 180 focused tests listed above passed. - `git diff --check` passes. The branch changes 20 files against master. It has no lockfile or workflow changes. - The original author verified all three platform digests against registry integrity and ran the Linux executable. Its version was `2.1.286 (Claude Code)`. See #14816 for that evidence. - Greptile reviewed head `6db3d3f1` and gave 5/5 with zero comments. Both Superagent scans and Commitperclip pass. All current-head CI jobs pass, including build, typecheck, Rust, test shards, browser tests, and the canary dry run. ## Risks - The controller and provider pack must use matching runtime pins. Deploy them together. - CI owns `pnpm-lock.yaml`. The master lockfile refresh must resolve the SDK override. Refresh the Daytona lock digest with that lockfile. - Images built with Claude Code older than 2.1.284 need a rebuild before the CLI path can use Sonnet 5.5. - The runtime remains at SDK 0.3.286. This PR does not take the later 0.3.287 patch. - A live Sonnet 5.5 session and a Daytona image build are not part of the local verification. ## Model Used - Original work: Anthropic Claude Code, `claude-sonnet-5-5`. Review: `claude-opus-5-5`. The author reported `xhigh` effort, tool use, and code execution. The original context window was not reported. - Merge repair and PR preparation: OpenAI Codex, based on GPT-6, with tool use and code execution. The runtime does not expose the exact model identifier or context window in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Squash Attribution Keep these trailers in the squash commit to preserve the original author and AI attribution: ```text Co-Authored-By: Claude Code (Ivan) <SkilLab-Tech@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> ``` --------- Co-authored-by: Claude Code (Ivan) <ivan@skillab.com.br> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |