mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 14:10:50 +02:00
codex/slack-managed-setup
1633
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
483c466008 |
fix(slack): make manager refresh and installation recovery resumable
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
785ae16699 |
feat(slack): add gated managed setup alongside customer-owned apps
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
835a022936 |
perf(server): wake delivery queues from transaction-aware work signals (#15625)
## Thinking Path > - Paperclip manages AI-agent work and its delivery to users and agents. > - Five delivery queues scan the database even when empty. > - Each queue already has durable rows; empty polling wastes idle hosting capacity. > - Producers can register intent inside their existing database transaction. > - Central transaction tracking can wake consumers after the outer commit without extra caller wrappers. > - A lost response needs reconciliation against PostgreSQL before an empty queue is safe to leave idle. > - This change uses one scheduler for outstanding work and lets settled, empty queues go quiet. ## Linked Issues or Issue Description Supersedes the closed #15614. Related idle-safety work: #15522 and #15599. **What existing behavior does this improve?** Delivery scheduling for feedback exports, chat completions, connection continuations, question answers, and tool-action receipts. **Current behavior** Feedback scans every five seconds. Four delivery queues scan on the heartbeat interval, normally 30 seconds. Startup and some event paths also scan them. **Proposed behavior** Each producer awaits one intent registration before its queue write. The existing `createDb` transaction boundary tracks nested savepoints and wakes the consumer after outer settlement. Startup scans restore durable work. Pending rows, failures, and unresolved transactions retain retry deadlines. Empty, settled queues have no timer or recurring scan. **Reason and benefit** Normal delivery starts after commit. Central tracking removes caller-specific post-commit plumbing. PostgreSQL transaction status resolves lost responses without new tables or a permanent uncertainty latch. **Breaking changes** No API or schema changes. The notification scope is one DB owner and its dedicated child pools. Direct SQL, separate roots, and other processes require an explicit wake/recovery integration. This path requires PostgreSQL 14+ transaction-information functions; embedded PostgreSQL uses 18. ## What Changed - Add one named-deadline scheduler and app-owned coordinator for five existing queues. - Instrument `createDb().transaction()` centrally, including nested savepoints. Dedicated child connections share their owner's signal scope. - Register work before the five queue insert paths. The first registration obtains the outer XID; unrelated transactions issue no extra SQL. Tool receipt insertion gains a small transaction around its existing insert. - Query `pg_xact_status` only for rejected, tracked transactions. Probe before scanning the queue. Retain the idle hold while PostgreSQL still reports an in-progress transaction or the probe fails. Release it after settlement and reconciliation. - Remove unconditional delivery recovery scans and caller-specific completion post-commit actions. Retain the existing task-scoped completion activity fast path and durable claims. - Coalesce wakes without postponing an earlier deadline. Retry outstanding rows and failures; suppress dispatch during idle drain while allowing transaction and queue reconciliation; suppress both during warm standby. Cancel deadlines and await active delivery sweeps on shutdown. - Serialize feedback flushes. Vote routes return after saving. Uploads have a 30-second deadline and shutdown cancellation; unfinished exports remain recoverable. - Document the writer contract and update the dated sleep inventory. ## Verification - Final head: `a4609956841107ad60c4cb50fe2acb05636b1363`. - Passed: all final-head hosted checks, with 54 successes and two intentional skips. The full test matrix, typecheck, build, canary, and aggregate verification passed in [run 37940948440](https://github.com/paperclipai/paperclip/actions/runs/37940948440). - Passed: Greptile 5/5 on the same head, with a successful check run and zero unresolved review threads. The branch is mergeable. - Passed locally after the final rebase: 50 coordinator, scheduler, and startup tests; server typecheck; server build. - Passed locally before the final import-only rebase: repo-wide typecheck and build; 237 PostgreSQL producer/delivery and database signal tests; feedback, native-question, and idle-safety regression tests. - Real PostgreSQL tests hold transactions open after an injected client response loss, then commit or abort. They verify that an empty scan cannot clear an in-progress transaction. Another test terminates a real backend and verifies recovery without replay. - Coordinator tests cover an hour with no empty-queue timers, pending/error retries, worker replacement, concurrent writes, earliest-deadline preservation, rollback during idle drain, committed work during drain, standby, and shutdown. - Local `pnpm test:run` was started and stopped after embedded PostgreSQL startup failures appeared. It did not finish and is not reported as a pass. Later local database reruns had skipped suites; those skips are not claimed as verification. The earlier PostgreSQL tests did execute and pass, and the final hosted full matrix passed. ## Risks - One XID query is added per transaction that registers delivery work. Nested transactions share that ID. Registration must be awaited before writing; new enqueue paths must follow this contract. - Work signals are local to a DB owner and its dedicated child pools. Independent roots, direct SQL, or other processes are not observed. Startup scans recover already committed rows, but do not fence late transactions from a previous process. Cross-process ownership and database failover require separate work. - A database outage or genuinely unresolved transaction keeps an idle hold until status can be reconciled. No timer clears uncertainty by assumption. - These changes remove five empty polling loops. They do not implement whole-instance sleep, database teardown, or an external waker. The earliest deadline covers only the registered queues. - Waiting tool reviews retain their existing retry cadence until their receipts can be delivered. - Feedback vote responses no longer wait for the remote upload. Sharing consent and the saved vote response are unchanged. - Real hosting-provider sleep and cost savings have not been measured. ## Model Used OpenAI Codex, GPT-6. Used reasoning, repository inspection, code editing, and command execution. The runtime did not expose the exact model revision or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
57e977be72 |
feat: integrate Pi 1.0 into the experimental Runner (#14921)
## Thinking Path > - Paperclip manages AI agents and their work. > - The experimental Runner owns provider processes and durable sessions. > - Pi needs working task execution and human controls. > - The five-PR stack must preserve changes already on master. > - Each layer now carries the complete integrated source for a safe sequential fallback. > - This PR belongs to native GitHub stack #15602, ending at #14956. ## Linked Issues or Issue Description Refs #14436, #14631, #14743 and #14956. Ship Pi 1.0 through the experimental Paperclip Runner. The five PRs are #14921, #14922, #14923, #14924 and #14956. The user authorized the complete merge after checks pass. Existing `pi_local` execution is unchanged. Accounting and wider provider/platform qualification remain deferred. ## What Changed - Recover missing final replies after workspace finalization changes owners, using accepted-turn evidence without rerunning work or granting external-chat publication. - Preserve the admitted Pi instruction root across warm runs, while retaining changed-root rejection. - Give Pi a bounded 15-second default shutdown grace so stop, drain acknowledgement and durable suspension can complete. Explicit deadlines and other providers retain their existing behavior. - Integrate the Pi 1.0 runtime and master contracts. - Use Pi profile 22. Preserve explicit caller-selected models and exact native thinking levels. Keep Pi's wrapper, helper, extension and question/control behavior unchanged from the qualified profile-19 runtime. - Preserve master's Dot lifecycle and consent fields, configured task environment, status guards and current Codex/Claude dependency versions. Cursor stays qualified. Copilot stays pending; profile 17 binds the changed shared protocol validation sources. - Exclude general AWS IAM credentials from Pi static/custom provider bindings and selected task projections; preserve the provider-scoped Bedrock bearer key. Profile 21 is retained as historical provenance. Rust and cloud install probes use the current declaration. - Patch bundled brace-expansion 5.0.9 to the exact official 5.0.12 payload. Pin the patch and complete runtime closures. Include the patch in normal installed setup tooling. Keep the upstream Pi shrinkwrap as provenance and permit only this exact security correction. - Include current attestation files in the Docker build context. Keep the repository lockfile unchanged from master. CI and private image builds resolve manifest changes before their frozen installation. ## Verification - Full local `pnpm -r typecheck` passes, including Runner Rust, server and UI. Focused integration checks pass: 194 Runner admission/environment tests, 63 profile/credential tests with one expected skip, 152 Dot/UI configuration tests, and Pi transcript/notice tests. - Full local `pnpm build` passes on the final source. - Fresh final-source checks pass: all 698 Rust workspace tests (32 binaries), 156 credential/profile/controller tests with one expected skip, Runner TypeScript typecheck, and 20 package/setup/sandbox tests. - The profile-21 Pi materializer passes on the native host with the official pinned Node 24.21.0 and its npm. It verifies all 150 locked packages, the patched dependency and the exact closure. Setup/package bundle tests and UI token gates pass. - The old hashes were reproduced for all three supported targets before calculating the patched graph. New closure hashes are darwin-arm64 `282022db10150c6632b3444df421342e7d534bdf5d5fb1097a2e79d0625a2bcf`, darwin-x64 `64e251e19009f755c0b04f73ce2138246faab71a961b0f13d75ebfcc34bef12e`, and linux-x64 `713b1fdff42fb56a1518bdc084f181d70bee8ebadc3e4b1d76321ed9108c8410`. Independent native platform execution is separate from graph identity reproduction. - Historical cloud qualification remains unchanged: all seven core cases pass on shipping source `10dc43c9ec65d88c2f782d62afb296d09494f215`, harness `1a4408a48cfb5a1f094a311141c257c92cd7a893`, image `sha256:5b3a775b383591bda1b0c1889e509acc70ce7f37c53f09733c81d59037f02280`, and accepted Sonnet 4.6/low fixture. All 215 canonical files and all seven cleanup checks pass independent verification. These are profile-19 results and are not relabeled as fresh profile-22 runs. - Current Pi digest: `sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`. [The readiness plan](https://github.com/paperclipai/paperclip/blob/codex/pi-production-readiness/doc/plans/2026-10-02-pi-production-readiness.md) preserves campaign and failed-attempt provenance. - Merge only after every PR's current-head CI and fresh review pass. Linux CI covers the full suites, build and browser tests. The local embedded Postgres API-authority suite cannot start on this macOS/Node 26 host, so Linux CI must confirm that suite. ### Fresh profile-22 core qualification — 2026-10-08 All seven accepted core cases pass canonically on Pi profile 22, with `openrouter/anthropic/claude-sonnet-4.6` and native-confirmed low thinking. This model is a fixture; production accepts the caller's explicit Pi provider/model. Runtime/install source: `3241a992f2a7703e59e97ed0fd3e5d6405de4401`. Frozen accepted harness: `1a4408a48cfb5a1f094a311141c257c92cd7a893`. Immutable cloud image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:506f22db7edd78f37c0c40bec1cc084af1850455026dbf467194bfbb8fcef141`. Pi digest: `sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`. [Hosted Linux image and clean-install verification](https://github.com/paperclipai/paperclip/actions/runs/37868328023) passes, including all 20 source-bound archives, normal CLI/Pi setup, companion import and the production pack reader. This exact installation source includes the latest master integration and the corrected Pi warm instruction-root fence. Full local typecheck/build and current-head hosted CI verify the final stack. All 13 focused real-root regressions pass. The full local executor suite passed 662 tests; 15 database tests could not start the Mac embedded PostgreSQL service. Hosted Linux CI passes the full required verification and E2E checks. These fresh results keep their own source identity; profile-19 results remain historical. | Core path | Canonical campaign | Retained archive SHA-256 | | --- | --- | --- | | File edit, validation, download and Done | `pi-core22-replyfix-0-1791511228` | 23 files; `a473e8603a3dd4737863291f8d3d1e392391f0b16d433c3e0e0e9d8baf7a97b0` | | Pending question and controller restart | `pi-core22-replyfix-1-1791511376` | 33 files; `6b829c4eb74e1f32a89c692a4ae7130dbfc1c6d3cf13915effe2103d9e242c8e` | | Three-turn session/process/workspace continuity | `pi-core22-replyfix-2-1791511587` | 23 files; `7a87021f8f9a3fdd3c58bb4467f8d82c635e3ea4795d6e75f144d9aa14818df8` | | Four typed questions and browser reconnects | `pi-core22-replyfix-3-1791511881` | 42 files; `9e31755252be1f4f9cb0626c984c142d4d1ae5f5bee3a7af08444db8d12c280a` | | Plan approval and completion | `pi-core22-replyfix-4-1791512031` | 22 files; `a0383ce1aab38e7b5a25ce0e9dd3bebea5c037ebd96ae6b29dae19015da2ae2c` | | Same-turn steering and permission denial | `pi-core22-replyfix-5-1791512261` | 39 files; `c929b8c7070f0b66aedc17e65ca46e6beab1e363926ac9f7e2a75fb250f05949` | | Stop during pending permission | `pi-core22-replyfix-6-1791512390` | 33 files; `7f0a58ae0f4d5bfc76149435f4e322537089c5bd16e7ffe9b5ad71f10a621a07` | All 215 canonical files (28714587 bytes) are independently hash-verified. All seven cleanup grades pass, with no owned runtime process or temporary root after each case. Automatic retries are zero. The owned cloud host stopped normally after retention. The prior profile-22 warm attempt remains failed and separately retained: archive SHA-256 `1e54eba5ec72b50cee1534b23d1d1d4f21a090006b8a64501ba70db972abfde5`. Its original canonical classification is preserved. Diagnosis reproduced a product bug comparing an agent-files root against an unset checkpoint-only field. The fix stores the admitted physical root separately from the adopted per-run collection capability. The real-root regression fails before the fix and passes afterward, including rejection of a changed physical root. Fixture, grader, model and all seven accepted case IDs are unchanged; this fresh campaign tests final-reply publication after file registration first. The intermediate restart attempt also remains failed and retained: archive SHA-256 `5dcaefdf1d17cf4cd54fd4cf810f45e736667392339b8ce7caf08bb4e225277f`. Its original canonical classification is preserved. Pi resumed, wrote the verified answer and completed its task; exact runner suspension was proven, but idle stop consumed about 5.2s and left under 3s for the drain acknowledgement. The Pi-only default shutdown grace is now 15s, preserving a full 5s drain round trip and a finite suspension reserve. Explicit caller deadlines, other provider defaults, literal drain receipts and exact suspension identity checks remain unchanged. The timing regression fails before this correction and passes afterward; all 18 focused settlement tests and Runner typecheck pass. The final-source file attempt is also preserved as failed (`candidate_failure`), archive SHA-256 `db6767b6773ea618997927ac77bdb005a5ac81492c7b9c0ffbc900449f829bc9`. Native edit, validation, exact downloadable artifact and Done/succeeded all passed, and the exact final reply was durably recorded. A workspace recovery owner completed before the live heartbeat reached presentation, leaving that reply absent from task chat. Recovery now materializes only a completed final reply from the accepted turn of an ordinary internal Done task, preserving issue/run/contract binding, suppression, external-chat authorization and same-run deduplication. The database regression covers the generated file-preparation receipt, suppression, unapproved external continuation and replay. Server typecheck and all 49 response-selection tests pass; hosted Linux verifies the database regression because embedded PostgreSQL cannot start on this Mac. The delayed-final-answer database regression passes on [the final root-source Linux server shard](https://github.com/paperclipai/paperclip/actions/runs/37868262553/job/113628594152), alongside 1,108 passing tests. The first root Runner shard had one unchanged durable-resume test exceed its 5-second timeout; the identical top-source shard and the isolated exact test passed. One rerun of that failed job and its required aggregate passed without source or test changes. The original failed job log and the single-rerun receipt remain retained. ### October 9 merge verification Current merge head: `5a8fe63512a7166aaef5cf50065a25008aa8b44b`. All current-head checks pass, including `ci / verify` and `ci / e2e`; exact-head Greptile review is 5/5 with no unresolved threads. Current master conflicts are resolved. The user authorized the maintainer override of the code-owner review gate after these checks. The seven retained live core cases remain bound to source `3241a992f2a7703e59e97ed0fd3e5d6405de4401` and its recorded cloud image. ## Risks - The security correction changes the dependency closure and profile identity. Old sessions must reopen on the new profile. Exact identities and credential bindings fail closed. - The runner remains experimental and requires explicit selection. Legacy Pi Local is unchanged. Caller model IDs pass through; the E2E model is a fixture. - Accounting and the broad platform/provider matrix remain deferred. This merge does not publish a release or deploy a service. ## Model Used OpenAI GPT-6 through Codex assisted with reasoning, repository inspection, editing and tool use. The exact serving ID and context window are not exposed in this session. Final live qualification uses Pi 1.0.0 with `openrouter/anthropic/claude-sonnet-4.6` and native-confirmed low thinking. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6f9d0a56ba |
fix: resolve installed Codex and preserve npm host dependencies (#15555)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Installed agents need their own runtime dependencies. > - Native Codex startup and browser login could depend on a global CLI. > - npm bundles do not inherit workspace dependency overrides or patches. > - Codex native packages must come from npm for the consumer host. > - This PR fixes executable resolution and the required npm dependency layout. > - Usable installed CLI versions can run without a numeric version gate. ## Linked Issues or Issue Description Refs: #15422. This is a prerequisite for the later Codex-default change. **What happened?** Packaged native Codex startup and browser login could require a global Codex CLI. The server's vendored runner lacked its own declared Codex bridge. A bundled wrapper could retain producer-host binaries or select the legacy adapter's separate platform version. **Expected behavior** Prefer the installed Codex executable. Accept usable older or newer versions. Use the selected host's PATH when the dependency is absent. Keep the patched JavaScript graph. Let npm install official native packages for the consumer platform. Grant sandbox reads only to the selected vendor resources. **Steps to reproduce** 1. Prepare a clean npm installation without a global Codex command. 2. Start native Codex or its browser login. 3. Inspect the selected executable, installed host package, and sandbox resource paths. **Paperclip version or commit** Frozen master base: `1881894973a2b25838d8abed9bd8aeebc3af4441`. **Deployment mode** Source and packaged self-hosted installation. ## What Changed - Share installed Codex command resolution across native startup, direct evals, and browser login. Accept usable version differences. Preserve explicit commands, recorded sessions, remote host boundaries, and the Linux ARM64 legacy login path. Return safe errors for missing executables and failed login terminals. - Declare the server's Codex bridge. Retain the patched JavaScript dependency graph in npm packages. Strip Codex native payloads. Declare official optional host packages on the published server manifest so native and legacy versions remain separate. Preserve consumer esbuild platform dependencies. - Adjust resource lookup for npm's separate platform packages. Retain package identity, path containment, resolver ownership, and narrow sandbox reads. Keep exact version and digest checks for explicit ACPX artifact qualification. - Extend existing packaging, installed-consumer, login, selection, recovery, and integrity tests. Document the retained behavior. The diff is now 23 files, 1,709 additions, and 53 deletions. The prior diff had 35 files and about 3,700 changed lines. Removed work is preserved on `codex/runner-packaging-full-snapshot` at `6f3060beaa842ffcee21af058370d3cab5f571e9`. Removed from this PR: release workflows and assembly, provider-pack changes, Docker materialization, Git installer changes, extra login HOME/working-directory isolation, and unrelated CI fixture repairs. This PR does not change agent defaults, stored runner choices, UI, schema, or provider qualification. ## Verification - Final candidate: `dde37d7ed0121c10b60b8801eda80f8dc17909ad`. [All 47 ordinary CI jobs passed](https://github.com/paperclipai/paperclip/actions/runs/37842288835) on attempt 1, including typecheck, tests, build, E2E, runner checks, and the installed-consumer canary. [Fresh Greptile review](https://github.com/paperclipai/paperclip/pull/15555#issuecomment-6059581546) is 5/5 on this head; no unresolved threads remain. Human approval is still outstanding. - Reused focused controls passed with Node 24: 71 initial narrowed checks, then 56 affected Codex/selection/eval checks after the relative PATH correction. Runner TypeScript no-emit, syntax, and diff checks passed. The existing fixture reproduces the original relative PATH failure and verifies working-directory selection, empty entries, ordering, and absolute launch. One CLI entrypoint test was initially blocked by sandbox IPC and passed with its existing local socket allowed. Login HOME, config-directory assignments, and working directory match master. - [The actual clean Linux npm consumer](https://github.com/paperclipai/paperclip/actions/runs/37842288835/job/113535523044) passed on the final candidate's CI integration. All 17 Paperclip tarballs omit native Codex payloads. Official npm host packages retain their own integrity and `inBundle=false`. Native Codex selects 0.160.0; the legacy closure retains 0.156.1. Package admission, command leases, narrow native sandbox resources, consumer hooks, preserved lock, and offline lifecycle controls passed. Provider calls were zero. No new workflow is added. - Actual official Codex 0.156.1 passed on the final committed source on macOS ARM64: installed dependency preference, absolute and relative PATH fallback, narrow resource lookup, and app-server initialize/initialized. Executed module hashes match the candidate. One scripts-disabled install, three version probes, and one handshake completed in 26 seconds; owned files and process group were removed. No login, account/model request, or provider task ran. - CI checked out `dfda8e708e87306d22ed735d78bdfcbe770e99b7` on base `65b558180533039a891ee0cd1ccab9988aa79adc`. Its 13 upstream paths do not overlap this PR's 23 paths or alter packaging/Codex inputs. The consumer report records producer `7b8e94c08657b5ddc265946459762340f125726c`, after the existing canary staged a generated-lock-only commit (one file, three insertions). These identities are kept separate; raw generated-lock bytes were not retained. - No local Docker or Rust build, paid provider turn, merge, or deployment. Later PRs must prove live onboarding and production cloud packaging before changing defaults. ## Risks - npm must install optional host dependencies. Missing dependencies still return errors. Paperclip tarballs do not pre-bundle Codex executables. - The repository requires CI-owned lockfile updates. PR CI resolves the changed manifest and stages its own producer lockfile. A raw source Docker build with `--frozen-lockfile` must wait for the existing master lockfile bot to merge its refresh, or use a disposable resolved checkout. No Docker build or deployment is qualified by this PR. - Ordinary Codex startup accepts version differences. Actual protocol or login failures remain errors. Explicit ACPX artifact checks retain their release pins. - The published server delegates platform installation outside its bundled JavaScript graph. Focused negative controls reject unsafe package metadata and paths. The hosted consumer test passed on this candidate’s CI integration. - Existing agents retain their stored runner choices. There is no data migration or automatic upgrade. Release pipeline and platform qualification work remain separate prerequisites for later defaults. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and parallel agents. The exact serving snapshot and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2b688a8fe3 |
feat(connections): add verified MCP providers and setup fixes (#15621)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents governed access to external tools. > - Several provider MCP servers lacked a supported catalog entry or failed during setup. > - Real browser tests identified specific registration, session, and form defects. > - This pull request adds seven catalog entries and fixes the shared paths used by eleven verified providers. > - Users can connect these providers through the existing Apps flow and control each action. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Apps catalog and remote MCP setup, discovery, and action tester. **Subsystem affected** Shared app definitions, the server tool services, and the Apps UI. **Current behavior** Seven providers lack a catalog entry. Airtable can select an advertised client-metadata flow that fails. Calendly rejects the registration name. Tavily needs an initialized session before a call. Firecrawl exposes invalid defaults for optional nested form objects. HTTPS setup can display a callback that differs from the server callback. **Proposed behavior** Add Calendly, Exa, Firecrawl, GSC Wizard, Parallel Search, Tavily, and Windsor.ai. Preserve Airtable, Linear, Make, and PostHog. Use reviewed provider options through the shared MCP and OAuth paths. Display the callback supplied by the server. Omit untouched optional object inputs. **Reason and benefit** Eleven providers passed bounded reads through the real Paperclip browser action tester. This PR includes that verified set and its required shared fixes. Unfinished providers remain outside this change. AgentMail keeps its existing integration. **Breaking changes** No database migration. Existing OAuth credentials and action policies keep their ownership and access rules. Airtable's reviewed method re-registers a retained client-metadata binding through DCR. Tavily's reviewed method initializes sessions before dispatch. The existing customer, managed, and Vercel OAuth gateway paths now refresh on upstream 401 and return `oauth_refreshed_retry_required` (409), requiring an explicit caller retry. They do not replay the rejected call automatically; the next invocation initializes a fresh credential-scoped session when required. Related catalog work: #15545. This branch preserves current master catalog entries and does not add another Google Workspace integration. Open and closed provider PRs and public MCP issues were searched. No duplicate for this verified set was found. The change extends the shipped Connected Apps roadmap item. ## What Changed - Add seven catalog entries with official provider branding and reviewed permissions. - Update Airtable, Linear, and Make metadata and show Make in Apps. - Add authoritative provider source overrides that use the existing generation and review checks. - Add the reviewed Airtable DCR option and compatible Calendly registration names. - Initialize Tavily sessions before discovery and calls, including retained connections. Keep credential-scoped session caching and single call dispatch. - Use the server's callback URL in the OAuth setup UI. - Omit untouched optional object defaults; keep supplied empty strings, false, zero, and object values intact in the action tester, with strict required-child validation. - Require an explicit retry after OAuth refresh instead of replaying a tools/call with missing session headers. - Record all eleven successful reads, write policies, auth-method limits, and qualification caveats. Omit credentials and private account data. - Find chat connector cards through catalog search in browser tests, so pagination does not hide Slack or other later entries. ## Verification - Real browser qualification: eleven bounded reads passed. Catalog refresh and reload proof are recorded in `doc/connections/verified-mcp-qualification-2026-10-08.md`. - Provider writes and actual agent-adapter sessions were not run. Alternative documented auth methods and expiry refresh remain unqualified. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - Focused shared/catalog tests: 41 passed. Source override tests: 3 passed. Catalog regeneration fixture suite: 9 passed. Latest form regression suites: 31 passed. Gateway suite: 38 passed. Other focused UI suites passed before these last fixes. - UI token gates and both token sync checks: passed. - The catalog regeneration fixture includes the authoritative provider overrides; its full suite passes. The focused OAuth socket case passed on isolated rerun. - Complete local UI suite: 7,866 passed. CLI suite: 511 passed, 6 skipped. Complete shared suite: 892 passed. - The unchanged AgentMail/ClickUp discovery fallback suite passes all 45 cases. An OpenAI login test passed on isolated rerun. - Local broad database coverage is limited by macOS PostgreSQL shared-memory exhaustion (`shmget: No space left on device`, not disk space). The broad run was interrupted after diagnosis; no host settings or other running services were changed. - CI found that the Slack browser test assumed its catalog card was on the first page. The test now uses catalog search; the Slack case passes locally. Adjacent GitHub and iMessage cases could not start locally because embedded PostgreSQL initialization failed before browser assertions. - Final commit [`febe4520e`](https://github.com/paperclipai/paperclip/commit/febe4520e13d4b3a4121eafc3d22540c2b11379d): all 54 checks passed, with none pending or failed. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37847182065) passed typecheck, build, all general and serialized test suites, all eight browser shards, runner checks, canary dry run, and the aggregate verification gate. - Greptile reviewed that exact final commit at 2026-10-08 21:33 UTC and returned 5/5 with no outstanding findings or unresolved threads. - PR UI preview against the existing local test server: Calendly catalog/setup observed; Firecrawl read passed after expanding More options with nested optional inputs untouched. This retest proves frontend behavior, not the revised OAuth backend against a live provider. <details> <summary>Browser evidence</summary>    </details> ## Risks - Provider registration and consent behavior can change. Airtable's DCR option is explicit and keeps the existing issuer, resource, redirect, and PKCE checks. - Tavily uses initialized sessions without automatic call retries. After a successful OAuth refresh following 401, callers now receive a retry-required result. A later explicit retry can still require reauthorization if the provider rejects the refreshed token. Provider writes are not live-qualified. - Optional object cleanup affects the shared action tester. Regression tests cover absent objects, required children, defaults, and populated values. - GSC Wizard's underlying Google data scopes and Google flow completion were not independently verified. Its account reported paid/trial metadata of unknown origin. No purchase was performed. - No database migration, new AgentMail integration, or change to existing action grants is included. ## Model Used OpenAI GPT-6 through Codex assisted with implementation, research, tool use, and code execution. The root backend model ID and context window are not exposed in this session. The cheaper subagents used OpenAI `gpt-6-luna`. No model context size is inferred. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted checks and complete UI/CLI/shared suites passed; the broad local database run was blocked by the macOS PostgreSQL startup limitation documented above. The full database/workspace suite passed in CI. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b9750b152f |
build(deps): bump rustls from 0.23.43 to 0.23.45 in /packages/paperclip-runner/runner (#15295)
Bumps [rustls](https://github.com/rustls/rustls) from 0.23.43 to 0.23.45. <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/rustls/rustls/commit/2976d90fd1c2db6b518700dd101b714069cfcb17"><code>2976d90</code></a> Prepare 0.23.45</li> <li><a href="https://github.com/rustls/rustls/commit/f9d4ee92b20da535ef5982ffcaf6f62385f79641"><code>f9d4ee9</code></a> Handshake "alignment" check covers previously-received messages</li> <li><a href="https://github.com/rustls/rustls/commit/c4e92e8c9ee29ca3f552d4938efa73680c179179"><code>c4e92e8</code></a> Keep data needed for HRR processing together</li> <li><a href="https://github.com/rustls/rustls/commit/e553a7a70f99cad4722f06b32cc37857008a7d3f"><code>e553a7a</code></a> server: TLS1.2 is not available after a HRR</li> <li><a href="https://github.com/rustls/rustls/commit/5dff9da73798253d8d35909dd8548db522cd973c"><code>5dff9da</code></a> Test whether server negotiates TLS1.2 after HRR</li> <li><a href="https://github.com/rustls/rustls/commit/185a063b5fe1e9698c589d8be6a83d67224b95a1"><code>185a063</code></a> reject a second ClientHello that changes the cipher suite</li> <li><a href="https://github.com/rustls/rustls/commit/9fafe6dd0c15ea6e47399e94b2c6b647861f874b"><code>9fafe6d</code></a> reject a second ClientHello that drops pre_shared_key</li> <li><a href="https://github.com/rustls/rustls/commit/1b42c5f78e8e826d6afa72ce0ead3e3a4e624d79"><code>1b42c5f</code></a> providers: zeroize private key DER</li> <li><a href="https://github.com/rustls/rustls/commit/64ad386785c718c74262b79e6893813db381adfe"><code>64ad386</code></a> Bump version to 0.23.44</li> <li><a href="https://github.com/rustls/rustls/commit/1efbf662ed79d1392c542c0822805a818d232ac0"><code>1efbf66</code></a> bogo: remove PostQuantum setup</li> <li>Additional commits viewable in <a href="https://github.com/rustls/rustls/compare/v/0.23.43...v/0.23.45">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) You can disable automated security fix PRs for this repo from the [Security Alerts page](https://github.com/paperclipai/paperclip/network/alerts). </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
1c4ce44e13 |
fix: retain bounded workspace transfer failure evidence (#15637)
## Thinking Path > - Paperclip manages AI agents and the work they produce. > - Sandbox work must return to the host when an agent run ends. > - A failed transfer must keep its source available for recovery. > - Some transfer errors lose their useful fields before restore diagnostics reach Sentry. > - This change records fixed transfer stages, failure kinds, and numeric RPC codes. > - The next failure can identify the operation that failed without exposing workspace contents. ## Linked Issues or Issue Description **What happened?** A workspace transfer can fail after model execution succeeds. Several command, archive validation, and download errors become plain errors. The existing RPC envelope then reports `unknown`, with no transfer stage or command exit code. A host RPC timeout also loses its code from restore diagnostics. **Expected behavior** Keep bounded evidence through the provider, worker RPC, restore result, and Sentry context. Keep the same exception, error message, retry rules, and source retention policy. A transfer timeout must remain separate from the model execution timeout flag. **Steps to reproduce** Return a nonzero exit code from the sandbox archive command, return a per-file download error, exceed a tar listing limit, or time out the sync-out RPC. Observe the missing fields in the saved restore diagnostic. The added tests use local fixtures for these cases. Related public work: #15479 preserves the source after restore failure. #15481 carries bounded sync-out diagnostics through RPC. This change supplies missing producer evidence and extends that same envelope. I checked the roadmap and searched open PRs for duplicate transfer diagnostic work. ## What Changed - Add optional transfer stage and failure kind fields to the existing diagnostic envelope. Capture command exits and listing deadlines at their producers. - Record only known codes from typed RPC errors at the sync-out boundary. Revalidate every field before persistence and Sentry projection. - Keep annotations private to each outbound request and restore settlement. Preserve frozen error identity and prevent evidence from leaking across concurrent or later calls. - Document the fields. Cover producer failures, quota limits, old workers, concurrent error reuse, and redaction through the real Sentry SDK. ## Verification - Focused SDK, provider, host, persistence, and real Sentry tests: 449 passed, no skips. Independent review also ran the focused contracts and all provider tests. The optional Sentry SDK is required for this check, so the contract tests cannot skip. - A compiled Daytona transfer and compiled SDK worker pass a local RPC round trip. This uses the development TypeScript loader for workspace dependency exports. The transfer and capture modules are compiled JavaScript. - `pnpm -r typecheck` and `pnpm build` passed. The excluded Daytona package also passed `tsc --noEmit`. - The initial local `pnpm test:run` stopped in the general-server group: 60 suites could not initialize embedded PostgreSQL because the isolated install skipped its Darwin library-link setup. The existing package postinstall repairs this; `initdb --version` now passes. All 60 affected suites then passed: 1,342 tests, no skips. No local full aggregate pass is claimed. - Independent source review covers privacy, concurrency, packaging, and recovery behavior. ## Risks - These diagnostics help identify future failures. They do not establish or repair the cause of a past transfer failure. - New fields are optional. Older workers remain compatible. Unknown values are omitted. - The change adds a small request-local diagnostic scope. It keeps the original thrown errors, archive confinement, cleanup order, timeout values, retry count, and source retention policy. - No schema or deployment changes are required. No live provider actions are part of validation. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and independent agent review. The exact deployed model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9b624a110e |
fix: checkpoint database backups before idle sleep (#15629)
## Thinking Path > - Paperclip coordinates agent work and preserves company state. > - Hosted instances can stop compute when admission and all work sources are quiet. > - The periodic database backup timer currently blocks every automatic sleep. > - Disabling backups would remove recovery points. > - This change offers an explicit final-backup checkpoint before sleep. > - A verified archive and restart marker preserve recovery while compute is stopped. ## Linked Issues or Issue Description **What happened?** An otherwise idle instance with automatic backups enabled always reports background work. Operators cannot reclaim idle compute without disabling those backups. The work inventory also queries the database before checking known local blockers. **Expected behavior** An operator can enable checkpoint mode. The server may authorize sleep only after all other work is quiet and a fresh backup is verified under the same owned hold. Backups resume on restart with the existing retention policy. **Steps to reproduce** Acquire a bounded idle drain on an instance with no application work and automatic backups enabled. Read its owned idle safety report. It reports present even when all other work is clear. With this change and `PAPERCLIP_DB_BACKUP_IDLE_CHECKPOINT_ENABLED=1`, a successful final backup can permit sleep; backup failures still refuse it. **Paperclip version or commit** Reproduced from master at |
||
|
|
4265cb3a2b |
fix(ci): cache native server integration builds in release verification (#15619)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Release verification must run the same server integration coverage as PR verification. > - Three server suites use real Rust Runner binaries. > - PR checks already run them with the shared Rust dependency cache, but release checks cold-build them inside ordinary server test shards. > - A cold Dot Runner build can consume its entire setup deadline before any test runs. > - This pull request gives release verification a required cached native integration lane and preserves every test. ## Linked Issues or Issue Description **What happened?** On master commit `1881894973a2b25838d8abed9bd8aeebc3af4441`, [Cloud readiness run 37837810707](https://github.com/paperclipai/paperclip/actions/runs/37837810707) failed in server shard 7. Dot Runner's `cargo build --release` exceeded its 300-second child-process deadline. The other 1,640 tests in that shard passed. The failed setup then tried to remove an undefined temporary path and emitted a second error. Cargo output was captured as an opaque buffer, which obscured build progress. **What did you expect to happen?** Build fixture binaries before tests in a lane with the existing Rust dependency cache. Run all native integration tests and make their result required for release readiness. A setup failure should keep its original diagnostic. **Steps to reproduce** Run release verification on a clean runner. The existing `general-server-without-chat` group retains the three Cargo-backed suites outside the PR workflow. The Dot suite can time out during a cold release build. Run the new workflow and partition tests against the previous source to reproduce five routing and coverage failures deterministically. **Version / commit** Observed at `1881894973a2b25838d8abed9bd8aeebc3af4441`; this change is based on `ed6abbf158b`. **Deployment mode** GitHub Actions Release and Cloud readiness verification. No application runtime or deployment action changes. Related work: #15581 also edits Dot integration tests for onboarding behavior. It does not repair release test routing. Searches found no open PR for this failure. ## What Changed - Add an explicit server test group that excludes the dedicated chat and native suites. Preserve the existing PR and local test groups. - Run all three native server suites in one required matrix lane with the existing trusted Rust dependency cache. - Build debug and release fixture binaries in a visible step with a 10-minute limit before tests start. Keep Cargo freshness checks and the existing test deadlines. - Stream Cargo diagnostics and clean up safely when Dot setup stops before creating its temporary directory. - Test complete, non-overlapping partitions for Release, Cloud readiness, other callers, and existing PR/local groups. Check cache restrictions, build order, and the source verification dependency. ## Verification - Passed 44 workflow and partition tests: `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - The same tests fail in five relevant cases against the previous workflow and selector. They pass after this correction. - An independent review repeated all 44 tests successfully. - Passed `actionlint .github/workflows/release-verify.yml`, `git diff --check`, and a secret scan. - Passed all 381 workflow policy tests: `node --test '.github/scripts/tests/*.test.mjs'`. - Passed full `pnpm build` in a clean worktree. - Built debug and release fixture binaries, then passed all 32 tests in the three real native integration suites: `pnpm test:run:general -- --group general-server-native-runner` (49.5 seconds, no skips). - Passed full `pnpm -r typecheck`. - Injected a synthetic Cargo setup failure. The suite reports that failure without the secondary undefined-path cleanup error. - Exact-head CI completed: 53 successful checks, including all server shards, both native Runner lanes, build, typecheck, browser tests, and the canary dry run. Two optional Storybook checks were skipped. - Started the duplicate local `pnpm test:run` aggregate and stopped it after exact-head CI passed. No completed local aggregate result is claimed. All test processes owned by that run exited. - Greptile scored the final head `ffe5ab5a28af2dfea9db1c1952c652f070bf52f8` at 5/5 with no review threads. - `Superagent Supply Chain Scan` is neutral, not passed: it only supports exact dependency pin replacements and cannot verify structural edits to `.github/workflows/release-verify.yml`. It reported no annotations. Independent source review, Greptile, actionlint, and the workflow policy tests cover this structural change. - GitHub still requires a code-owner review for the workflow files. There are no merge conflicts. ## Risks - Adds one CI matrix job, which increases concurrent runner demand. The existing cache writer and trust restrictions stay in place. - Missing dependencies still compile from the lockfile. A build that exceeds its step limit fails visibly; failed tests still block source verification. - Workflow execution uses the new lane after merge. PR tests validate the workflow contract and run the existing native test lane. - The supply chain scanner cannot analyze this workflow structure. Its neutral result is an explicit coverage limit; it is not counted as a successful scan. - No runtime, schema, deployment, publication, credential, or test-timeout changes. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and independent agent review. The exact deployed model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2b4a1d7073 |
feat(dot): add guided invitations and persistent agent connections (#15581)
Add guided external-agent invitations with live Dot connection checks and ongoing access until revoked. Keep expired credentials invalid and enforce permissions for private avatar uploads. Preserve Runner recovery and remote launch boundaries. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3d1d5294d9 |
Drain idle plugin workers before automatic sleep (#15599)
## Thinking Path > - Paperclip runs autonomous work through agents, schedules, and plugins. > - Automatic idle sleep must preserve accepted work and unfinished cleanup. > - Every enabled plugin currently blocks sleep, including unused providers. > - An enabled flag or package version cannot prove that a worker is idle. > - This change adds a live worker drain under the existing owned hold. > - Idle workers can allow sleep while unknown work stays protected. ## Linked Issues or Issue Description Refs #15522. Related: #15391 adds plugin readiness for agent admission; this change concerns instance sleep and does not replace that contract. **Subsystem affected** Plugin worker lifecycle and automatic idle sleep. **Problem or motivation** A workspace with no pending work cannot sleep when any plugin is enabled. Removing that check alone would lose accepted RPCs, background tasks, or cleanup after a caller timeout. **Proposed solution** Require a live `onIdleDrain` handshake from each worker. Close admission in both processes for the exact owner and expiry. Count accepted work until completion. Continue checking durable work separately. ## What Changed - Add bounded worker holds, exact-owner release, automatic expiry, and an abort signal for plugin-owned background work. - Count host and worker requests through their real completion receipts. Keep timed-out work counted. Check active notifications and terminal routes. - Accept enabled plugins only when their current workers provide matching runtime receipts. Missing workers, old SDKs, crashes, invalid replies, and unknown cleanup still prevent sleep. - Let unused Daytona workers opt in. Once a worker contacts the provider, it remains a blocker for that process lifetime. This restriction avoids treating its existing timeout and terminal-close behavior as a cleanup receipt. - Document the plugin author contract. No schema or user-facing API is added. ## Verification - All hosted CI checks passed on |
||
|
|
53105d5830 |
feat(apps): add Gauge connection (#15596)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents reach external services through the Apps catalog. Each catalog entry is a reviewed `AppDefinition` that connects a provider's hosted MCP server to Paperclip's shared vault, grants, policies, gateway, and audit trail. > - Gauge (`withgauge.com`) measures how AI answer engines mention and cite a brand. Its hosted MCP server also gives SEO, analytics and ads reports, and runs a content pipeline that can publish to a connected CMS. > - Gauge is not in the catalog. Operators must use the generic "Connect your own MCP server" flow. That flow has no branding, no organization guidance, and no warning about live CMS publishing. > - The connector playbook supports this provider with existing definition fields. The server supports dynamic client registration with a reviewed `mcp` scope, and it accepts an organization API key as a bearer header. > - This pull request adds the Gauge definition, its artwork, the research and permission-review ledger rows, documentation, and deterministic tests. > - The benefit is a branded, governed Gauge connection with browser sign-in, an API-key option, and a clear publish warning. ## Linked Issues or Issue Description **Problem or motivation** Marketing and growth teams use Gauge to track their brand in AI answers and to run content workflows. They want their Paperclip agents to read visibility, keyword and traffic data, and to prepare content. Gauge is not in the Apps catalog. Operators must paste the MCP URL into the generic remote-MCP flow, which gives no branding, no method guidance, and no warning that content tools can publish live. **Proposed solution** Add a catalog-only Gauge connection built from the connector playbook. Browser sign-in uses Gauge's dynamic client registration and requests only the reviewed `mcp` scope. The user selects one Gauge organization on the consent screen. An optional method sends a customer-created organization API key as an `Authorization: Bearer` header. Both methods warn the operator to set publish actions to Ask first. Every discovered tool stays governed by the normal per-action policies. **Alternatives considered** A plugin was not needed because no custom UI, tables, workers, or webhooks are involved. The identity scopes that Gauge also advertises (`openid`, `profile`, `email`, `organizations`) are not requested, because Gauge selects the organization on its consent screen. A Gauge-specific `classifyRisk` rule was not added, because the tool names cannot be seen without an account. The generic rule already classifies `publish`, `create` and `update` tools as writes. **Roadmap alignment** This extends the existing self-serve remote-MCP connection catalog and does not overlap planned core work. ## What Changed - Added the `gauge` provider to `scripts/ingest-app-definitions.mjs` (category `analytics`, API-key placement, guidance, warnings, description) and regenerated `packages/shared/src/app-definitions/gauge.json` and the generated registry. - Added the Gauge row to the self-serve MCP research ledger with `dcr_or_api_key` auth and risk tier S3. - Added permission reviews for `gauge/mcp-oauth` (explicit scope `mcp`, from Gauge's live authorization-server metadata) and `gauge/mcp-api-key` (provider key), with evidence links. - Added Gauge's mark (`ui/public/brands/apps/gauge.png`, the avatar of Gauge's official GitHub organization) and the brand manifest entry. - Added gallery copy for the Gauge card. - Added `doc/connections/GAUGE.md` (endpoints, scopes, administrator setup, capabilities and policy, manifest, brand provenance, validation hook) and linked it from the connections README and the permission audit. - Tests: Gauge definition shape, store visibility and artwork, URL recognition, reviewed scope, bearer-header placement, the connect form's sign-in default and API-key gating, and the pinned catalog counts. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts packages/shared/src/app-definitions-url.test.ts ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx server/src/__tests__/tool-access-service.test.ts` — 738 passed. - `node scripts/check-app-brand-assets.mjs` and `node --test scripts/app-brand-validation.test.mjs` — passed. - `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter @paperclipai/server typecheck`, `pnpm --filter @paperclipai/ui typecheck` — clean. - `node scripts/ingest-app-definitions.mjs --definitions-only` produces no Gauge drift. - Manual: on a local instance, open Apps → Browse and confirm the Gauge card and icon. Open `/apps/connect?source=gauge` and confirm that sign-in is the default and that the API-key method is under Advanced. - Live metadata probe on 2026-10-08: an unauthenticated `initialize` on `https://app.withgauge.com/mcp` returns 401 with `resource_metadata`. Both `.well-known` documents return the recorded endpoints and scopes. Dynamic client registration succeeds. The authorize endpoint accepts `scope=mcp` and rejects an unknown scope with HTTP 400. ## Risks - Low risk to existing providers: the change is additive catalog data plus tests. The generated registry only gains one import. - Gauge content tools can publish to a connected CMS, including live, and a bulk keyword update replaces each prompt's keyword list. Both methods warn the operator, and the API-key helper text tells operators to set publish actions to Ask first. - Gauge API keys are not scoped and reach the whole organization. Paperclip cannot narrow an issued key. - The permission-review ledger records live proof for both methods as not run. The lifecycle checklist in `doc/connections/GAUGE.md` still needs a documented pass with a Gauge organization. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with tool use (shell, file editing, web research). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9f4acc8f14 |
build(deps): bump @modelcontextprotocol/sdk from 1.30.0 to 1.31.0 (#15361)
Bumps [@modelcontextprotocol/sdk](https://github.com/modelcontextprotocol/typescript-sdk) from 1.30.0 to 1.31.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/modelcontextprotocol/typescript-sdk/releases">@modelcontextprotocol/sdk's releases</a>.</em></p> <blockquote> <h2>1.31.0</h2> <h2>Upgrade notes</h2> <ul> <li>Stored OAuth tokens and client information now include an <code>issuer</code> field. Storage that rejects unknown fields needs to allow it.</li> <li>Pass <code>expectedIssuer</code> when constructing <code>ClientCredentialsProvider</code>, <code>PrivateKeyJwtProvider</code> or <code>StaticPrivateKeyJwtProvider</code>. Constructing them without it is deprecated.</li> </ul> <h2>What's Changed</h2> <ul> <li>[v1.x] Bind stored OAuth credentials to the authorization server that issued them by <a href="https://github.com/maxisbey"><code>@maxisbey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2888">modelcontextprotocol/typescript-sdk#2888</a></li> <li>chore: bump version to 1.31.0 by <a href="https://github.com/claude"><code>@claude</code></a>[bot] in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2890">modelcontextprotocol/typescript-sdk#2890</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.1...1.31.0">https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.1...1.31.0</a></p> <h2>1.30.1</h2> <h2>What's Changed</h2> <ul> <li>[v1.x] fix(server): read HTTP request bodies with a size limit and bound JSON-RPC batch length by <a href="https://github.com/maxisbey"><code>@maxisbey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2717">modelcontextprotocol/typescript-sdk#2717</a></li> <li>fix(auth): preserve resource URI without trailing slash (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1968">#1968</a>) by <a href="https://github.com/MukundaKatta"><code>@MukundaKatta</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1972">modelcontextprotocol/typescript-sdk#1972</a></li> <li>chore: bump version to 1.30.1 by <a href="https://github.com/claude"><code>@claude</code></a>[bot] in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2848">modelcontextprotocol/typescript-sdk#2848</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/MukundaKatta"><code>@MukundaKatta</code></a> made their first contribution in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1972">modelcontextprotocol/typescript-sdk#1972</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.30.1">https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.30.1</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/4b0051f400219f8d8855f9a5433c6df35f15a639"><code>4b0051f</code></a> chore: bump version to 1.31.0 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2890">#2890</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/51ad4f03190dc5c84ab8b1a25f9b78b277be0dc7"><code>51ad4f0</code></a> [v1.x] Bind stored OAuth credentials to the authorization server that issued ...</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/289ac2c3af7e1536160e80414296b175171a1a87"><code>289ac2c</code></a> chore: bump version to 1.30.1 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2848">#2848</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/12b425678a76cd54b0452a2ccf1e5dc7740f73ef"><code>12b4256</code></a> fix(auth): preserve resource URI without trailing slash (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1968">#1968</a>) (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1972">#1972</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/a9f6eb709b85459d01e8f2a9e881fef2621756c1"><code>a9f6eb7</code></a> [v1.x] fix(server): read HTTP request bodies with a size limit and bound JSON...</li> <li>See full diff in <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.31.0">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
3d0e74e743 |
Keep proven external connection failures out of Sentry (#15590)
Apply the reviewed change for Keep proven external connection failures out of Sentry. Validation: local typecheck/build and focused regression tests, passing exact-head CI, Greptile 5/5, and independent source review. The PR records full-suite evidence and any local environment limitations. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0ac194450a |
fix: make Copilot provider-pack wrappers portable (#15586)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Container images include a movable provider pack for agent execution. > - pnpm executable wrappers can contain the temporary build directory. > - New optional Copilot packages add wrappers that the build does not replace. > - This pull request gives those installed wrappers relative executable paths. > - Image publication can finish while the existing path check remains enforced. ## Linked Issues or Issue Description Refs #15572 and #15560. The small portability loop comes from cryppadotta's larger Copilot runtime PR #15560. This separate fix repairs image publication without waiting for that feature's qualification and runtime changes. That PR can remove its duplicate loop after this lands. **What happened?** The standard Docker build stops with `Provider pack shim copilot-linux-x64 retains its temporary build path`. The failure occurred before and after #15522. See [the failed master build](https://github.com/paperclipai/paperclip/actions/runs/37802065316). **Expected behavior** Installed native Copilot wrappers resolve their pinned executable after the provider pack moves. Optional packages that are absent do not gain a command. The builder still rejects wrappers with temporary paths. **Steps to reproduce** 1. Run the provider-pack build from the affected master revision on Linux x64. 2. Let `pnpm deploy --prod` install the optional Copilot package. 3. The wrapper scan rejects its temporary `NODE_PATH`. **Paperclip version or commit** Master `3367b75ccce34d02f355cda1f1ed3fe0b34cf93d`. **Deployment mode** Docker and provider-pack builds. ## What Changed - Replace installed Copilot platform wrappers with relative native executable launchers. - Move the existing executable wrapper writer into an importable helper. Retain the same launch behavior for Node, Claude and OpenCode. - Test relocation, argument handling, exit status, absent optional packages and missing executable packages. Register the tests in the existing Runner test preparation command. - Hash the helper in Daytona image identity and prove that helper changes invalidate the image cache. - Document the packaging rule. Keep dependencies, provider qualification and the final temporary-path check unchanged. ## Verification - 17 Node packaging tests passed across the new wrapper tests, provider-pack release tests and candidate selection tests. - Six existing bundled remote-provider-pack tests passed using the Runner Vitest configuration. - Nine Daytona image identity tests and Runner E2E typecheck passed after the cache-input correction. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - Real `pnpm deploy --prod` reproduction on macOS ARM64: the original Copilot wrapper contained the temporary path. After the rewrite and directory relocation, the actual executable returned Copilot CLI 1.0.88 with exit code zero using only `/usr/bin:/bin` in `PATH`. - Syntax checks, `git diff --check` and the pre-push secret scan passed. - The full local `pnpm test:run` reproduced the same five skill/connector fixture failures observed earlier in this workspace. It was stopped after current-head clean-checkout CI passed; later local phases were not run. This local run is not claimed as passing. Focused packaging tests, workspace typecheck/build and all hosted CI passed. - [Hosted Docker verification passed](https://github.com/paperclipai/paperclip/actions/runs/37804902678): Linux AMD64 and ARM64 image builds, multi-architecture publication and the process-reaping smoke check. This run tested `779d94989c54d9abbeba0838194c951186af66a6`; the only later changes are Daytona cache identity and its regression test. Provider-pack build code is identical. Local Docker did not respond within the bounded probe. - Current head `db2c19e270b8d5a7bab5db39d3de0e4761cf57c3`: 54 successful checks, two conditional skips, no failures and no merge conflicts. Apex is 5/5 with no unresolved comments. ## Risks The helper uses each installed package's exported executable. An installed wrapper with a missing package still fails the build. This change does not execute Copilot during image construction, alter dependency pins, change runtime admission or weaken the temporary-path check. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use and code execution. The exact serving model identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; the full local-suite limitation is documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3367b75ccc |
fix: fence accepted work and cleanup before idle sleep (#15522)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A host can stop an idle instance to reduce unused compute. > - Zero active runs do not prove that requests, saved work, or cleanup are complete. > - A client can disconnect while its request still writes, and failed cleanup can remain only in memory or on disk. > - This pull request holds new admission and checks accepted work, cleanup, and persisted work under an owned drain. > - The host gets an empty report only while those checks remain valid. ## Linked Issues or Issue Description Refs #13413. That companion PR uses the same owned-hold vocabulary for runtime services. This PR covers HTTP requests, scheduler work, accounting, cleanup, and the instance work inventory. It does not include the preview gateway changes. **What existing behavior does this improve?** The instance task-drain API reports process counters. It does not prove that a host can safely stop the instance for idle sleep. **Current behavior** A quiet run set can coexist with an unfinished request, cleanup after a completed run, future work, or a failed accounting write. **Proposed behavior** Provide a bounded owned idle hold. Block new ingress, track accepted handler promises, inspect durable and local work, and return `none` only when the same hold stays quiet through the checks. Keep normal deployment drains compatible. **Reason and benefit** Hosts can identify eligible idle instances without treating a disconnect or failed cleanup write as completed work. ## What Changed - Add `purpose: "idle"`, a bounded TTL, unique owners, and owner-checked release to task drain. Existing holds cannot be replaced by another API request. - Gate HTTP ingress before parsers, auth, webhooks, and MCP. Gate new WebSocket upgrades. Track async handlers in nested Express routers and error middleware until they settle, even after the response or client disconnect. - Count accepted live-event WebSocket authentication through settlement, even after disconnects. Count detached built-in agent, managed-home and runtime-service startup reconciliation after readiness. - Keep health and control mutations tracked. Count control-request authentication separately from the read-only report, including concurrent user/company/membership writes. - Pause new scheduler admissions during idle holds. Count work already in flight, including database backups, and reject scans whose work generation changes. Periodic backups block sleep without a host wake schedule. - Inspect accounting and orphan-cleanup spool directories without skipping temporary or malformed entries. Retain orphan tokens through queue splices, flush failures, and buffer overflow. Keep failed usage capture counted until its database failure fence is written. - Check persisted work across companies in a bounded read-only transaction. Include deferred agent-file cleanup and saved watchdogs whose watched issues are complete. Enabled plugins and unsupported retained work remain blockers. - Require the exact idle owner and completed startup tracking before returning an empty report. Keep reports free of tenant details. - Document the hosting protocol, retry behavior, conservative blockers, and the remaining external provider-stop race. ## Verification - Current head: `8d9a599b732c2047b8c671d4799065d7c90a3567`, rebased on master `941a3fa991aeb97eb1ac390c65b7973b5f6de1ad`. Both heartbeat helper extractions are preserved. GitHub confirms no merge conflicts. - Focused heartbeat renderer/run-log, drain, control-auth, admission, route and PostgreSQL inventory checks: 234 tests passed in 10 suites after the rebase. - The POST task-drain contract includes the expected `409` conflict response. The 84 OpenAPI and instance-settings route tests and server typecheck passed after that final documentation fix. - Final accepted-upgrade/startup regression run: 104 tests passed in five suites, including success and failure after readiness or disconnect. Final server typecheck also passed. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm install --frozen-lockfile` and the PR diff secret scan passed. No dependency or lockfile changes are added by this PR. - The prior verification covered HTTP disconnects and early responses, async error handlers, concurrent authentication writes, signed bootstrap, saved watchdogs, backup promises, orphan cleanup, spools and failed accounting fences. Those tests passed. - The last full local `pnpm test:run` stopped in the general-server phase with 16,658 passing tests and five skill/connector fixture failures caused by an ancestor workspace skill directory. That full local run preceded this rebase and has not been repeated for the import conflict. Full current-head CI passed: 53 successful checks and two conditional skips, with no failures. - All eight review findings are fixed and their threads resolved, including accepted upgrade authentication, detached startup writes and the POST conflict contract. Current-head Apex review is 5/5 with no new findings; all eight review threads remain resolved. - No live provider stop or production change was performed. ## Risks - The hosting controller must use the owned protocol and hold external admission through its final validation and provider stop. A legacy drain cannot authorize idle sleep. During an idle hold, new requests receive 503 with `Retry-After: 1`; the host must handle queueing or retry before enabling this path. - This is a single-process protocol. An unexpected restart after the last validation can race an external provider stop. The host must serialize deploy/wake/sleep operations and bind the validation to the instance it stops. Multiple replicas need shared fencing. - Some retained state conservatively prevents sleep, including every enabled plugin. This PR does not promise that every inactive instance becomes eligible. - The HTTP adapter uses Express 5 router layers. Real Express tests cover nested routes, errors and disconnects. New routes must register before tracking is installed. Detached work must have durable state or explicit work tracking. - Unrecoverable in-memory cleanup debt keeps the instance awake until reconciliation. This change does not make such debt survive an unplanned process crash. No schema migration or provider configuration change is included. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use and code execution. The exact deployment variant/model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; the full local fixture limitation and clean full CI result are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
941a3fa991 |
Pin Copilot native dependencies with the maintained lockfile refresh (#15572)
Pin the three optional GitHub Copilot 1.0.88 native packages for Runner and server using the maintained lockfile workflow. Synchronize the package contract and bound initial render readiness in the deliberately throttled browser fixture. Current-head CI and focused checks pass. Co-Authored-By: Dotta <cryppadotta@users.noreply.github.com> Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
71cd0a2621 |
fix(skills): honor the current run harness checkout (#15548)
## Thinking Path > - Paperclip manages work for AI agents. > - The runtime claims eligible assigned tasks before it starts an agent. > - The wake tells the agent when the runtime already holds that claim. > - The legacy skill still requires another checkout in every case. > - This PR makes the skill honor the current task and run claim. > - Manual checkout and server ownership checks remain in place for other cases. ## Linked Issues or Issue Description **Where is the issue?** `skills/paperclip/SKILL.md`, in the scoped wake procedure and Step 5. **What's wrong?** The wake can say that the harness already checked out the issue. The skill still tells the agent that it must call checkout. These instructions conflict. **Suggested fix** Skip the second checkout only when the runtime wake explicitly confirms the claim for this issue and run. Retain manual checkout when that statement is absent or the agent selects another task. Refs #14948 for the existing shared prompt reduction. ## What Changed - Honor the explicit runtime claim in the scoped wake procedure and Step 5. - Keep context reads, status writes, deliverable handling and conflict rules. - Add checks for normal and resumed wake text and excluded automatic claims. - Retain successful checkout HTTP activity for legacy stock-task evals. Bind each receipt to the exact company, task, agent and run. Keep this observation separate from the original task grades. ## Verification - Checkout observation calibration: nine tests pass. - Focused skill, wake and database ownership tests: in progress. - Full repository build, typecheck and tests: in progress. - Planned live comparison: the existing assigned-skill document case on legacy Codex and Claude. One attempt per variant and profile. No automatic retries. The baseline and candidate share the observation code and task oracle. - Live results are pending. This draft does not claim behavioral qualification. ## Risks - Agents may misread prompt guidance. The API still enforces ownership; the text grants no new authority. - The exception is specific to the current issue and run. It does not remove ordinary legacy completion writes or authorize another task. - Activity measures successful checkout HTTP calls. Failed attempts require separate run-log inspection. Missing or mismatched observations cannot count as zero calls. - One trial per profile cannot establish general reliability, speed or cost trends. ## Model Used OpenAI Codex, GPT-6 family. The exact model build and context window are not exposed in this session. Used code editing, shell tools and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d66acb7ac1 |
feat: automate Slack bot app setup and installation (#15413)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors give each agent a customer-owned bot and task-backed conversations. > - Manual Slack setup requires app creation and copying durable credentials. > - Operators need a shorter setup that an assisting agent can use safely. > - This pull request creates the app through Slack's Manifest API and installs it through OAuth. > - Durable registration state supports recovery without creating another app. > - A four-screen wizard, automatic avatar upload, and OAuth account linking reduce setup work. > - Connector settings and per-turn tool guidance support daily use after installation. ## Linked Issues or Issue Description **Subsystem affected** Native Slack bot setup, company secret storage, chat connector management, and agent tool guidance. **Problem or motivation** New Slack bots require manual app creation and copying a signing secret and bot token. Interrupted setup can create duplicate apps. The setup and management screens contain unnecessary controls. Agents also need guidance for native questions, files, thread replies, and governed Slack actions. **Proposed solution** Use a temporary app-configuration access token to create a customer-owned app. Save durable secrets in the vault. Bind OAuth to the initiating actor, company, endpoint, registration revision, scopes, and configured origins. Preserve manual and existing-app recovery. Link the installing user's account, send a welcome DM, and advance from saved server evidence. Keep request-URL recovery instructions available if automatic connection detection waits. **Alternatives considered** The Slack CLI adds installation requirements. Socket Mode changes transport. A shared Paperclip-owned app changes app ownership. These alternatives are outside this change. **Roadmap alignment** This extends existing chat connectors and secrets capabilities. Related public work: #14037 and #13954 cover Slack MCP prerequisites and user OAuth. No duplicate bot-registration PR was found. ## What Changed - Share one reviewed manifest builder between automatic registration and manual setup. - Add replay-safe migration 0318 and company-bound registration state with vault references and uncertain-creation recovery. - Add registration, installation, callback, and resume APIs with short-lived, single-use OAuth state. - Save installation credentials before downstream checks and preserve bot identity constraints. - Reduce automatic setup to four screens. Keep advanced app details, manual recovery, and existing-app setup. - Upload the agent avatar with the Paperclip dark background. Link the OAuth installer's account and send setup DMs. - Show agent and connector-owner avatars. Simplify settings, access, and conversation screens. - Discover joined Slack channels and enable them by default. Start a task from a bare mention and admit same-thread follow-ups. - Refresh Slack tool guidance each turn. Add native-form, file, approval, and delivery regressions plus manual model probe definitions and sanitized acceptance records. - Update deployment/database docs, OpenAPI, redaction, removal cleanup, production Storybook stories, and provider browser tests. - Merge current master and move the registration migration after its latest migration without rewriting published commits. The completed Slack success view intentionally has a single centered **Done** action and no **Save & exit**, as explicitly requested by the product owner. `DESIGN.md` records this exception; unfinished setup steps retain the aligned wizard footer. ## Verification - Passed after the master merge: repository typecheck, full build, Storybook build, design-token gates, module-boundary gates, and migration generation. - Passed: all 352 focused Slack deterministic tests and all 14 affected provider browser tests. Browser tests use controlled provider fixtures and a separate throwaway instance. - Passed on current head `c5d01e0e2`: the complete GitHub test matrix (general server, chat, all workspaces, serialized server, and Runner), all eight browser shards, typecheck/release registry, build, canary dry run, security checks, and policy gates. There are 52 passing checks and no pending or failing checks. - Greptile completed on the exact current head with 5/5 and no actionable findings or open review threads. - Local repair verification passed 93 focused tests, including same-app reinstall after revocation and rejection of consent started before revocation, the AgentMail browser journey, and repository typecheck. Local build and Storybook build also passed. The redundant local full-suite rerun was stopped after the complete current-head CI matrix passed. - Real Slack setup and agent replies were exercised in the authorized isolated test drive during the setup iteration. - The ten additional model probes were attempted with legacy `codex_local`, `gpt-5.6-sol`: five passed, two failed, and three were partly verified. Native runtime is not qualified. See `server/src/services/connectors/slack/evals/2026-10-08-acceptance.md` for evidence and limits. - Passing model probes cover native forms, downloaded file bytes, bare mentions with thread replies, explicit posts/reactions, and saved approval denial. - The controlled uncertain-write probe found wrong delivery-check IDs. The canvas fallback attempt used an invented tool name. Search pagination/native search, a private-source denied-tool receipt, and distinct board/webhook origins remain unqualified. Reviewer path: enable Chat connectors, start Slack chat setup, select an agent, enter an app-configuration access token, and approve Slack installation. Send a message to the bot and confirm that setup advances to success. Inspect settings and allowed channels. See `doc/connections/SLACK-AUTOMATIC-SETUP.md` for deployment and recovery. ## Risks - Slack app creation has no provider idempotency guarantee. A timeout after dispatch stays uncertain until the operator checks Slack. - OAuth needs a stable public HTTPS board origin. Webhook ingress may use a separate configured HTTPS origin. Workspace policy can delay installation. - Migration 0318 can replay safely on instances that applied the earlier development migration. - OAuth installation now links the installer to the initiating Paperclip user. Identity checks and company access rules still apply. - Joined channels now enable bot responses by default. Linked-user authorization and per-action approval rules still apply. - Model behavior has the documented delivery-check and canvas fallback failures. A passing CI run does not establish that every model probe passed. - Removing the connection does not delete the customer's Slack app. No new first-party telemetry is added. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser verification. The runtime does not expose a more specific authoring model ID or context-window size. The live bot probes used OpenAI `gpt-5.6-sol` through `codex_local` in legacy mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d0f69670db |
fix(runner): recover saved execution prompts after upgrades (#15518)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs save an immutable execution context for restart recovery. > - The context includes the prompt text, revision, and content hashes. > - The parser required that saved prompt to match the current release. > - A server upgrade could reject a valid saved run before provider recovery. > - This pull request validates and preserves the saved prompt snapshot. > - Routine prompt changes no longer need a catalog of past strings. ## Linked Issues or Issue Description **What happened?** A hot restart selected a dead native runner for same-run recovery. Reading its saved v5 execution input failed with `input.runtimeContext.prompt must match the fixed Paperclip prompt revision`. The new controller accepted only v6. **Expected behavior** Recovery uses the saved prompt and validates its content hashes. It preserves the same run and provider session without starting a duplicate turn. **Steps to reproduce** 1. Start a native run and save its execution input and provider checkpoint. 2. Change the fixed execution prompt in the server release. 3. Stop the runner and recover the saved run with the new controller. 4. Observe that the old parser rejects the saved prompt before provider recovery. **Paperclip version or commit** The v5-to-v6 prompt change was introduced in #15446. The defect also reproduces on current master before this fix. **Deployment mode** Source-built server with the native runner. Related work: #15446 added task-monitor guidance. The held prompt-size experiment in #15489 changes prompt wording but does not add recovery compatibility. ## What Changed - Read the prompt text and revision from the saved execution snapshot. - Treat the revision as non-empty metadata and preserve the exact saved bytes. - Validate the prompt SHA-256 and the aggregate context digest. - Keep fresh-run builders on the current prompt constants. - Test arbitrary saved prompts, malformed fields, altered text, stale hashes, and aggregate drift. - Test recovery parsing for Codex input versions v3-v5 and OpenCode, ACPX Pi, and Dot v6 inputs. - Extend the real-process restart suite with both the incident's v5 wire fixture and a prompt unknown to this release. - Run the restart recovery suite in the existing Rust-equipped PR lane, where its runner and fake-provider binaries are built. Verify complete, non-overlapping test coverage for PR, release, and local callers. - Document recovery from saved snapshots without a historical prompt catalog. ## Verification - Red: the new contract regressions fail against the catalog-based parser with the original prompt-validation error. - Green: 53 focused contract and materialization tests pass. - Red: the real-process unknown-prompt regression fails with master's original parser at the saved-input recovery read after process loss. - Green: all 15 real-process restart tests pass locally on the final branch. The saved-prompt cases keep the run and provider session, replace the PID, and record one `turn/start`. - Local repository `pnpm -r typecheck` and `pnpm build` passed after rebase on `89f09dad723766e5351953f0731b9aa5daada28d`. The 53 focused tests also passed on that head. - Red: the new test-roster checks fail against the old CI placement. - Green: all 26 test-scheduling checks pass after moving the restart suite. - CI ran all 15 restart recovery tests with no skips on final head `6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`. [Runner test job](https://github.com/paperclipai/paperclip/actions/runs/37764782334/job/113271464966). - Greptile scored 5/5 on that exact head with no actionable findings. - The complete CI matrix passed on final head `6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`: general/workspace tests, serialized server suites, both runner Vitest lanes, Rust and static checks, all browser shards, typecheck, build, and the canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37764782334). - There are no unresolved review threads or merge conflicts. - Reproduce focused tests with `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/contracts/runtime-context.test.ts src/contracts/native-execution.test.ts src/drivers/runtime-context-materializer.test.ts`. - Reproduce restart tests with `pnpm --filter @paperclipai/paperclip-runner build:rust` followed by `pnpm exec vitest run server/src/services/native-runtime/native-runner-restart-recovery.integration.test.ts`. They use temporary PostgreSQL, real runner processes, and a fake Codex provider. They do not use paid inference. ## Risks - The parser now accepts internally consistent saved prompt text that is absent from the current source. Inputs must come from trusted server persistence. Content hashes verify consistency; they do not authenticate authorship. - Existing execution-schema, ownership, checkpoint, provider, permission, and session-compatibility checks still apply. - This change validates the saved base prompt. It does not make all additional code-generated instruction strings versioned. - The process-level recovery proof uses Codex. Other provider coverage verifies the shared input parser and retained provider configuration. ## Model Used OpenAI Codex, GPT-6 family, with repository inspection, code editing, and test tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b5342febe5 |
fix(runner): require current-turn completion after connection continuations (#15514)
## Thinking Path > - Paperclip manages agents, tasks, permissions, and execution budgets. > - Native tasks can resume after a connection decision in the same provider conversation. > - Each turn still needs an accepted completion report. > - Compact continuation messages did not explain that reports from earlier turns cannot finish the new turn. > - The connection evaluator could also grade before the final reply was stored or reject valid unavailable-access wording. > - This PR clarifies the current-turn report requirement and fixes those observation boundaries. > - The original connection instructions and strict native completion gate stay in place. ## Linked Issues or Issue Description Refs #15489. The reduction remains draft while this separate repair is qualified. Refs #15471 for the earlier connection continuation work. ## What Changed - Add a current-turn completion reminder to compact continuation inputs. - Keep final prose insufficient for completion. Preserve permissions and retry policy. - Wait for the final successful task run's saved, attributed decline reply within the existing deadline. - Use one bounded explanation matcher for both decline checks. - Wait for a recorded tool-action rejection to dispatch its bound continuation, with strict company, task, agent and source-run checks. - Retain the exact grading input before later API refreshes. - Add failure and delay regressions and update the Runner and evaluator docs. ## Verification The fresh comparison has **15/15 original passes on each variant**: 15 unchanged pass pairs, zero new failures, and no pending pair. There are 30 case attempts and **65 actual agent runs** (baseline 33; candidate 32). All runs succeeded. All 30 cleanups passed. No model attempt was retried. | Profile | Baseline | Candidate | | --- | --- | --- | | Native Codex `gpt-5.6-sol` | 5/5 | 5/5 | | ACPX Claude `claude-sonnet-5` | 5/5 | 5/5 | | OpenCode `openrouter/deepseek/deepseek-v4-flash-0731` | 5/5 | 5/5 | Each profile covers service approval, service decline, connection decline, provider decline, and selection of the second provider. The saved replies, approved briefings, decisions, fixture observations, final task states, and native completion records were inspected. All 18 saved decline-grade snapshots match their original captured inputs and checks. Result, API snapshot, and final ledger run sets agree. - [Candidate campaign](https://github.com/paperclipai/paperclip/actions/runs/37711658378) · [public candidate report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711658378-1/index.html) - [Baseline campaign](https://github.com/paperclipai/paperclip/actions/runs/37711675579) · [public baseline report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37711675579-1/index.html) - Candidate source: `79905343bba280d462765faad19a26e7f179259e`. Baseline source: `7c5e120158f1385a1fc5f20be41f66c58a605534`. Both use master context `fc6304dfe5f2e446e09bd052a7b45f51e930f250`. The trusted workflow source is separately frozen at `dd777f4b7343305c4e6f44c422f44a1d78e12e4f`. - Both variants use the same evaluator, fixtures, models, permissions, 720-second cell deadline, and 1,000-cent company and agent hard stops. The only production difference is the compact continuation reminder. - Suite hash: `aee30b74b4d38ada08777798db0932fbc64e368bb427d29caa8cfd86c7f59747`. Definition hash: `ea9e17f54fe0af1acbb2337ddfab8a7e92db8488bd0f59510a63095ca0229b60`. - Provider-free transport capture: startup/resume input stays at 55,726 bytes. Compact continuation input grows from 53,334 to 53,535 bytes. The 42-tool catalog stays unchanged. These are Paperclip input bytes, not complete vendor prompt tokens. - Focused evaluator tests: 31 pass. Native contract, transport delivery, and session tests: 194 pass. Evaluator support: 1,819 TypeScript tests and 128 Node tests pass, with one intentional skip. - Full build, workspace typecheck, and evaluator typecheck pass. Current-head CI passes all required gates. The current rollup has 51 successful check runs, two intentional Storybook skips, and a successful Snyk status. Review is 5/5 with zero unresolved threads. - CI attempt 1 had one initial runtime-fixture health timeout. The exact test and its full 164-test file pass locally. One CI shard retry passed. The original CI failure, its dependent verify failure, and the retry remain visible in [CI history](https://github.com/paperclipai/paperclip/actions/runs/37711235060). - The broad local `pnpm test:run` attempt was interrupted after about 49 minutes (exit 130). It recorded one failure in the unchanged Zep memory-connector disabled-setting test. That test and the full 388-test tool-access file pass in separate local checks; current-head CI also passes. The local cause is not established, and this broad local attempt is **not** claimed as passing. Three earlier local failures also pass in their isolated checks; their original logs remain retained. - The first two setup admissions were cancelled before provider jobs to include the review correction. They made no provider calls. The completed campaigns above are the first and only model attempts for these corrected variants. Cost evidence stays separate from behavior. Original result summaries report only OpenCode amounts: baseline $0.039192096 and candidate $0.063221620. Final run ledgers also retain estimates for Claude (baseline $1.232439000; candidate $1.333611200) and Codex (baseline $2.340324400; candidate $2.129335600). These estimates do not replace the original summaries. Local and GitHub runtime are unmetered here. Invoices are unknown. This is not a cheaper or faster claim. ## Risks - A single matched trial cannot prove general equivalence or causation. The reminder is an instruction change, not a new completion enforcement rule. - Candidate OpenCode service-decline finished within one continuous run; its baseline used two. That pair passed the task outcome, but it does not qualify the reminder on a resumed decline turn. No extra paid run was used to replace it. - The text matcher is bounded evidence of an explanation. It does not prove reasoning or consumption of feedback. Bounded stdout excerpts do not prove that every extra attempted tool call is absent. - Missing saved replies or continuations still fail at the original deadline. Failed native completion remains a failure even when final prose is correct. - The unresolved local-suite discrepancy above remains a validation limit. Full remote CI and both focused local reproductions pass. - The original connection instructions stay in place. These results do not qualify the reduction in #15489. Its original 11/15 versus 12/15 grades and two new failing pairs remain unchanged. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, test inspection and eval analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the focused local tests listed above and they pass; the interrupted broad local run is disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e4b39da6f6 |
Isolate ACPX admission deadline tests from lease ports (#15527)
Apply the reviewed change for Isolate ACPX admission deadline tests from lease ports. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ed0c6958f8 |
Explain refused local Hermes gateway connections (#15520)
Apply the reviewed change for Explain refused local Hermes gateway connections. Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
dd777f4b73 |
feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path
> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.
## Linked Issues or Issue Description
**Subsystem affected**
Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.
**Problem or motivation**
An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.
**Proposed solution**
Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.
**Roadmap alignment**
This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.
## What Changed
- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.
## Verification
- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head
|
||
|
|
fc6304dfe5 |
feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path > - Paperclip manages AI agents, tasks, permissions, and execution budgets. > - Paperclip Runner gives each provider the same admitted task and tool authority. > - OpenAI Dot runs outside the local process tree and needs asynchronous work delivery. > - The merged MCP gateway supplies OAuth consent and signed event delivery. > - A personal assistant grant cannot safely stand in for an assigned agent. > - This pull request adds a separate Dot agent connection and a durable Rust Runner bridge. > - The operator can assign work to Dot and inspect its accepted work, tool receipts, and result. ## Linked Issues or Issue Description **Agent or provider** OpenAI Dot, as an experimental provider of the existing Paperclip Runner adapter. **Why this adapter is useful** An operator can assign normal Paperclip tasks to an existing Dot. Dot can read its mailbox, request work on an assigned task, use admitted task tools, and submit a result. Paperclip keeps company scope, checkout, approvals, known budget limits, and activity attribution. **How the agent is invoked** A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the assignment. The Rust Runner owns the durable turn and operation receipts. The first release supports self-hosted instances with a local Runner controller. **Additional context** This extends the merged public MCP gateway from #14846 and the assistant invitation and device-consent work from #14933. This also integrates the merged assistant tool and configuration expansion in #15380. Dot retains its dedicated agent resource and cannot receive personal configuration permission. The public assistant connection remains a personal connection. ## What Changed - Add a durable Rust Dot provider and its TypeScript Runner driver. - Add closed PRP v3 external-provider operations and native execution input v6. - Add company-scoped pairing, mailbox, assignment, and operation records. - Reuse merged browser/device consent, client metadata verification, webhook admissions, refresh, secret rotation, and warm-standby gates. - Keep Dot scopes, issuer, grants, event workers, and tool access separate from personal assistant access. - Add Dot configuration, pairing, readiness, and consent UI. Keep agent grants out of the personal Connections entry. - Regenerate the Dot-only migration after master. Preserve published gateway migrations. Make the new migration safe to reapply. - Document setup, recovery, accounting limits, evidence, and remaining account qualification. - Reverify reconnect callbacks and wake outstanding work with a fresh mailbox reference; preserve the existing assignment and operation receipts. - Clean up Dot bindings and waiting runs on OAuth revoke and refresh-token replay. Old grants cannot revoke replacement bindings. - Restore the pairing reference when an unsaved agent form is reopened; document board-only pairing routes in OpenAPI. - Accept a clean Rust exit after the acknowledged shutdown receipt. Unexpected exits still require recovery. - Clear the cached binding after a successful revoke so a failed connection refresh cannot restore it. - Add production-component Storybook states and screenshots for pairing and connection review. All preview account data is synthetic. - Persist normalized completion, serialize Dot turns and durable work admission, and poll subscription readiness. - Serialize mailbox writes and cursor reads; retain paused fence acknowledgement without task authority. - Authorize admitted review runs without changing the worker assignee. Include the fenced assignment ID in production stop notices. ## Verification - This PR integrates master `4a8178e9c`. Dot migration `0317_messy_famine.sql` follows the published history and is safe to reapply. The merge preserves the reserved migration connection, batch-commit handling, private task checks, task monitors, and native accounting. - Local workspace typecheck, full build, and UI token gates pass. The server typecheck passes after the review fixes. Database and native executor regressions pass. - All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot driver tests pass. The tests cover native document writing and finalization, durable replay, queue admission, mailbox ordering, admitted reviews, stale authority, production stop references, and paused acknowledgements. - Current head `d0e7e0626` passes all 57 checks: 53 pass and four are intentionally skipped. This includes full typecheck, build, tests, Rust Runner verification, browser E2E, release verification, and Canary Dry Run. Greptile rates this exact head 5/5. All review threads are resolved. - The full local root test run is slower than the sharded CI run and has not completed. The full CI test gates pass on the current commit. Focused local regressions pass. - Real-account pairing and event delivery on this base commit remain unqualified. Live account and setup proof are recorded in the follow-up #15414. The following screenshots use synthetic preview data. They show the production pairing component and do not qualify a real account or the full agent setup journey.   ## Risks - This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP and Paperclip Runner. The separate experimental-settings follow-up in #15414 replaces this environment flag with saved operator settings. - Dot does not expose provider token usage or cost. The operator must acknowledge external billing. Known Paperclip budget gates still apply. - Cancellation fences Paperclip authority. It does not confirm that Dot stopped all external activity. - Assigned skill files and third-party MCP bindings are unsupported and reject admission. There is no mounted workspace, model selector, or provider thread identifier. - Hosted agent-broker and remote controller deployments are not qualified. - The new migration follows the merged master history. Existing prototype databases still need the normal master migration history before this Dot-only migration. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window size are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, and test inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a7a244ab33 |
feat: add company decision models with permission and cost controls (#15473)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Optional product features need small, typed model decisions. > - Each company needs to choose which shared API connection pays for those decisions. > - Calls must retain user and task permissions, budget limits, and cost attribution. > - This pull request adds a managed decision service, setup UI, and request history. > - Features can check availability cheaply and keep their existing behavior when decisions are unavailable. ## Linked Issues or Issue Description **Subsystem affected** Server services, shared contracts, database accounting, Company Settings, and Costs. **Problem or motivation** Paperclip has no common decision-model service. Adding provider calls within each feature would duplicate credential access, permission checks, and billing rules. **Proposed solution** Let a connection manager configure one company decision model. Support OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test and metadata-only history. Default company-sponsored background decisions to on during setup, and preserve a saved off setting. **Alternatives considered** Per-feature credentials would duplicate existing connection management. Personal overrides and provider fallback chains add permission and billing complexity; they remain deferred. **Roadmap alignment** Reviewed ROADMAP.md and searched open PRs. This extends existing connection access and budget accounting. Product features that call the service remain outside this change. No matching decision-model service PR was found. ## What Changed - Add company settings, an internal `decisionModelService`, local availability checks, and trusted human, agent/run, and system contexts. - Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and ordered-score decisions for both providers. Bound requests and time; disable paid retries. - Add durable invocation metadata and agentless decision ledger charges. Preserve fractional cents, pricing evidence, dispatch identity, and unresolved billing holds. - Share the company accounting lock and apply company, agent, and project budgets. Settle charges once, retain unknown holds, and recover interrupted calls without resubmission. - Reuse connection setup and management UI. Add a Decisions view under Costs, production-component Storybook coverage, database migration, and service documentation. ## Verification - Passed 166 current-code tests covering the decision service/provider, setup component, Costs, OpenAPI, and every failure from the earlier broad run. Coverage includes native SDK wire formats, refusals, billed malformed responses, permission and secret-rotation races, identity changes, concurrent budget admission, unresolved holds, agent/task deletion, and stale setup feedback. - Passed 170 existing connection, cost, budget, heartbeat-accounting, and profile regression tests. - Passed repository typecheck, production build, and design token gates after integrating master. Verified the generated migration on a fresh test database and upgraded the populated preview database from the branch's earlier migration without losing settings or usage. - Ran the required full `pnpm test:run`: its general phase completed with 16,354 passed and 10 failures across five files while this branch was still being updated. Every reported failure passes in the current-code rerun; the serialized phase did not run after that failure. The full GitHub CI suite passed on `b1b856a88`: general and serialized tests, browser shards, runner checks, typecheck, production build, packaging/canary, and policy gates. [CI evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360). - Passed the full-shell Storybook setup-to-history interaction test again after integrating master. Greptile rates the final revision 5/5 with zero unresolved threads. - Walked through the running app: empty setup, add each provider, save, reload, run all three sample questions, inspect fractional charges in history, switch provider while sponsorship is off, disable, and reconnect. Checked mobile settings. These tests used the actual UI, vault, server, SDKs, and database with simulated upstream responses. - Live paid setup tests remain unverified: this environment has no authorized OpenAI/OpenRouter credentials available. No mocked test is presented as live provider evidence. Reviewer journey: Company Settings → General → Decision model. Add/select a shared API connection, save, run the billed sample, open View usage, then disable decisions and verify Run test is disabled after reload. ## Risks - The SDK decision interface is experimental. Pinned versions and wire-format tests limit upgrade drift. - The migration allows agentless service charges and reservations. Existing agent cost-reporting APIs still require an agent, and decision receipts stay separate from run reconciliation. - Timeouts can have unknown provider charges. Holds remain until an audited accounting correction resolves them. - OpenAI prices use a versioned Decisions rate snapshot; OpenRouter costs use provider receipts. Unknown pricing is retained as unknown. - A configured company authorizes background spending by default. Setup explains this, and managers can turn it off. - Live provider account/model availability still needs the two credentialed acceptance checks. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser tools. The exact serving revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1c44b7b56 |
Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs. Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
cd40ebb95b |
Preserve ACP bridge terminal evidence before disconnect (#15478)
Drain terminal frames through backpressure and record bounded terminal evidence without waiting for log persistence. Guard later events and input failures against writes after socket end. Validated with 217 focused tests, full typecheck/build, independent review and green CI with Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
fd8c6b920a |
fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native tasks can pause while a human decides whether to allow a connection action. > - The next turn needs both the saved decision and complete accounting for the previous turn. > - Cancellation could discard final usage, and complete direct Claude API receipts could remain unpriced. > - A stale blocked or review report could also request approval again after the original card was declined. > - This pull request retains shutdown accounting and rejects approval waits bound to an already resolved action. > - The benefit is reliable continuation with the existing budget and approval controls. ## Linked Issues or Issue Description Refs #15420. Related: #15312 addresses requester ownership during dispatch. This change addresses receipt capture and final-response validation. ## What Changed - Retain usage events after native cancellation. Continue to reject late provider messages and work. - Drain same-turn accounting and terminal events for at most fifteen seconds after a durable governed wait. Keep incomplete accounting blocked. - Estimate complete, unpriced, direct Anthropic API receipts for the exact `claude-sonnet-5` model. Record the rate version and assumptions. Use the one-hour cache-write rate when the receipt lacks cache TTL. - Bind stale approval reports to exact interaction, action-request, or invocation IDs in the same company, task, agent, and run. Cover blocked, review, and response-wake reports. Keep independent reviews valid. - Fence checkpoint and result writes after a controller detaches for restart, including operations waiting for a database lock. Reject stale successful returns before certifying accounting. - Allow bounded subscription teardown only after retaining an actual provider terminal. - Journal the exact governed-wait trigger and disposition before provider interruption. Recover that wait independently of a later saved answer, replay retained accounting, and reject mismatched or unproven terminal evidence. - Give settling governed turns a bounded window before shutdown detaches their controller. - Keep fuzzy external app matches alongside installed capability matches instead of forcing an unrelated provider question for a generic query. - Return up to twenty exact active catalog tool names after an invalid request, after eligibility checks; still reject the request without granting access or creating an approval. - Require retained provider terminal proof before settling a governed wait, including when complete usage arrives before stream closure/error/timeout. Retain harmless numbered cancellation events so restart replay stays contiguous. - Isolate accounting-test OpenCode config from the host plugin directory. - Add regression coverage and document the accounting, restart and connection-search behavior. ## Verification Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`; later review fixes are verified separately and do not relabel those runs. - Repository typecheck and build pass. - Current runner runtime and cancellation suites: 191 tests pass. Six new regressions cover stream end/error/timeout without provider-stop proof and contiguous cancellation acknowledgement/request replay; all six failed before the fix. The existing bounded cleanup case now explicitly supplies terminal proof. All 31 adapter accounting tests pass with isolated fixture config. - New checkpoint-rebinding and approval-criterion suites: 65 tests pass; runner HTTP integration: 29 tests pass. Unchanged executor/control-plane suites: 640 tests pass; database-backed connection suites: 68 tests pass. - Current retained evidence verification covers 222 file hashes across all fifteen original result artifacts. The three-case [published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html) and all eight screenshot hashes verify. The [three-case recovery report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html) and all six screenshots also verify; the nine-result campaign did not publish. - Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node tests pass. Eval typecheck and catalog discovery pass. - Full local repository unit run was interrupted before the follow-up edits after three tool-access failures and one runner HTTP failure. Those failures pass in isolation; the 09a full run was interrupted after one rapid Slack callback-ordering failure and seven skill-service failures. All eight pass both isolated and with full-runner environment settings, and all 74 skill-service tests pass together; the subsequent full run reported two 15-second OpenCode accounting timeouts and was stopped with exit 130 to apply review fixes. The timeouts reproduce while copying this host’s 61 MB OpenCode config. All 31 tests pass after isolating config inside each fixture without increasing timeouts or changing assertions. A complete local full-suite pass is not claimed. Repository-wide CI also passes on the final review-fix head in [run 37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952). Earlier database-skipped diagnostics and the older ENFILE run are retained and are not full-suite passing evidence. - Three-case live campaign [37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525) passes all three original grades on frozen source 09a: 55/55 checks, eight succeeded run records, complete accounting receipts, matching checkpoint identities and no pending approvals. Campaign [37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353) adds nine original passes (173/173 checks, eighteen succeeded records) on the identical source. Its other three jobs failed before runner assignment or any step while GitHub could not load the paid environment; those original infrastructure failures are retained. Campaign [37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356) completes only those unstarted cells: all three original grades pass (57/57 checks, six succeeded run records). All fifteen exact cases now pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new overall failures and zero pending pairs. Total current evidence: 285/285 checks and thirty-two succeeded run records, complete accounting receipts, matching checkpoint identities, no pending approvals or retry records. Earlier campaigns retain forty-seven additional run records and two known same-run recovery attempts; actual provider-call counts and invoices remain unknown. The nine-result campaign skipped publication and its public URL returns 403; original artifacts remain retained. Previous ba9 campaign [37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840) completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline pass, while OpenCode resumes but times out searching for exact tool names and never creates the access card. That failure and incomplete cancelled-run accounting remain preserved; the new catalog error guidance targets this observed dead end. Campaign [37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312) remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker leak and unknown-criterion approval gap. The original baseline remains 7 PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS / 3 FAIL. All fifteen selected cases are qualified by their original grades in this bounded trial. These live grades belong to 09a. Its twelve governed-wait checkpoints retain matching same-turn terminal fingerprints, but passing artifacts omit detailed event journals; the later six adversarial regressions qualify the new terminal-proof and replay guards separately. Final-head repository CI passes. [Fresh Greptile review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780) is 5/5, confirms both findings are fixed, and reports no new actionable issues. All review threads are resolved and the PR has no merge conflicts. ## Risks - Governed cancellation drains accounting for up to fifteen seconds. Restart detachment gives a settling batch up to twenty seconds to finish. An incomplete receipt or unproven provider terminal still prevents successful qualification. - Claude prices are estimates, not invoices. The estimate assumes standard global API pricing and uses a conservative cache-write rate. Unsupported models, billers, and billing modes remain unpriced. - Approval identity matching must remain scoped to the current run and the requested approval. It does not authorize execution of a declined call. - Original baseline and final candidate have different merged master context. Exact-case outcomes are before/after observations, not isolated causal attribution to this repair. - No schema migration, fixture, oracle or grader change. Search-result guidance now treats fuzzy external matches as suggestions. Existing app authorization and provider-consent checks remain required. ## Model Used - OpenAI GPT-6 through Codex, with code editing, terminal tools, and test execution. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (191 current runtime tests and 31 accounting tests; interrupted full-suite history and CI coverage are disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0514e8abd5 |
Keep Cursor managed-runtime test installation offline (#15484)
Intercept the Cursor installer in its local shell fixture and guard against accidental curl downloads. Preserve real managed-home archive operations and the default timeout; change no production code. Verified 46 related tests, safe negative mutation, independent review and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
465140596f |
Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials. Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ae6f95ed7a |
feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults. Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
06484b3c41 |
fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Recovery schedules bounded retries after a provider disconnects. > - A stopped run can still hold its environment lease while cleanup runs. > - Retrying before that lease is released cancels the new run before it starts and spends another retry. > - Restoring the task can also send the worker repair instructions from an already resolved recovery action. > - This pull request preserves the waiting retry and removes settled recovery instructions from later wakes. > - The task can continue after cleanup without an operator repairing the same incident again. ## Linked Issues or Issue Description **What happened?** A legacy conversation run disconnected while it was doing ordinary work. Its environment cleanup took longer than the retry delay. Two retries were cancelled before dispatch with `execution_reconciliation_required`. Those cancellations exhausted the failure budget. After an operator restored the task, the wake still told the original worker to repair the runtime and hand the task back to itself. **Expected behavior** Cleanup waits preserve the pending attempt. The same retry can continue after ownership is released, subject to all current gates. Once a recovery action is resolved or cancelled, subsequent task wakes omit its repair instructions. **Steps to reproduce** 1. Fail a legacy conversation run while its environment lease remains in `pending_cleanup`. 2. Schedule a bounded retry and run promotion before cleanup releases that lease. 3. Repeat the scheduler sweep. Before this fix, retries promote and then cancel without starting. 4. Resolve a stranded-task recovery action and build the restored task wake with that action ID. Before this fix, the wake still includes the settled repair instructions. **Paperclip version or commit** Reproduced with database regressions against `ceabc3bc880` on master. Related: #15019 restores a skipped assignment handoff after lease release. #15235 filters stale handoff evidence in the recovery sweep. This change preserves an existing scheduled conversation retry and corrects restored wake content. It does not create a new handoff wake. ## What Changed - Keep an unstarted legacy conversation retry on the same durable row while prior execution ownership remains active. Recheck after 30 seconds without increasing retry accounting. - Return a queued retry to scheduled state if it encounters that hold at the claim gate. Retain its issue claim and publish the status change. - Record one local lifecycle diagnostic per blocking run. Remove that wait marker on promotion. - Include recovery action metadata only while the referenced action is active or escalated. - Return an explicit `waiting` response and the saved schedule when Retry now meets cleanup. Show the wait inline without a false success or disabled button. - Add database, rendered-prompt, route, and UI regressions. Document the execution and run-log contracts. ## Verification - Five cleanup and restored-wake regressions fail against the original production code. The Retry now route and UI regressions also fail before their correction. - Related retry, dispatch, stale-queue, and recovery suites: 321 tests pass across seven files. All ten focused cleanup/restored-wake cases pass after rebase. The final dispatch adjustment passes all 46 adapter tests. - Retry now routes and affected UI suites: all 46 tests pass. The tests cover repeated clicks, the saved schedule, unchanged accounting, promotion after release, and no false success or error state. - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm check:token-gates` pass. - Full local `pnpm test:run` was started and then stopped after the final commit passed all GitHub CI test shards. No complete local full-suite result is claimed; CI supplies the complete test result for the final commit. - Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub checks are green or intentionally skipped. Apex review is 5/5 after two reviews, with no unresolved threads. The PR has no merge conflicts. ## Risks The wait applies only to unstarted legacy conversation retries. Native runs and non-conversation execution keep their existing recovery rules. Cleanup must actually release ownership before execution can resume. The wait does not fix a cleanup service that never finishes. Promotion and dispatch still enforce cancellation, reassignment, pause, budget, and reconciliation gates. No schema change is required. The Retry now response adds a `waiting` outcome; the shared contract and all three UI controls handle it. ## Model Used OpenAI Codex based on GPT-6, with repository analysis, tool use, and local code execution. The runtime does not expose the precise serving model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; complete final-head test coverage in CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9fb955e9a2 |
fix(codex): gate models on the Codex CLI floor before the ChatGPT backend rejects them (#15399)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Codex adapter and the native runner run Codex in local hosts and in managed sandboxes, and the model catalog lists the models an operator can select > - The Codex backend accepts a new model with ChatGPT sign-in only from a recent enough Codex CLI. An older CLI fails every turn with `The '<model>' model is not supported when using Codex with a ChatGPT account.` > - After #14942 added `gpt-6.1-sol`, an operator selected it for an agent in a managed Daytona sandbox. The sandbox image still shipped Codex 0.156.0. The connection test and runs failed with that sentence, which reads like an account problem > - Paperclip only checked the Codex compatibility window (`>=0.149.0 <0.161.0`), so 0.156.0 passed and the failure surfaced from the backend without a cause > - This pull request records the verified Codex CLI floor for each gated model, compares the installed `codex --version` with that floor in the environment Test and in the remote runner, and names the backend rejection when it still happens > - The benefit is a precise, actionable message ("gpt-6.1-sol requires Codex CLI 0.159.0 or newer; detected 0.156.0; promote a sandbox image with Codex 0.160.0") instead of an opaque 400, and no doomed hello probe ## Linked Issues or Issue Description Refs #14942 **Bug: selecting `gpt-6.1-sol` in a managed sandbox fails with a ChatGPT account error.** **Steps to reproduce** 1. Promote a sandbox image that ships Codex CLI 0.156.0. 2. Create a Codex agent with a ChatGPT sign-in connection, select `gpt-6.1-sol`, and run the environment Test or a task in that sandbox. **Expected behavior** The Test or the run tells the operator that the Codex CLI in the sandbox is older than the model needs. **Actual behavior** The Test reports `codex_hello_probe_failed` with `{"type":"error","status":400,"error":{"type":"invalid_request_error","message":"The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account."}}`. A run fails with "The selected model is not supported by the current ChatGPT connection." **Evidence** - Codex 0.156.1 and older are rejected for `gpt-6.1-sol` with ChatGPT sign-in: https://github.com/openai/codex/issues/49396 - Codex 0.159.0 and 0.159.2 are accepted: https://github.com/openai/codex/issues/49464 and https://github.com/decolua/9router/issues/4471 (same error, fixed by raising the client version identity from 0.155.0 to 0.159.0) - Codex 0.157.0 release notes: "Add GPT-6 Sol and Luna to the model catalog": https://github.com/openai/codex/releases/tag/rust-v0.157.0 - GPT-6.1 Sol is included for Plus, Pro, Business, Enterprise, and Edu with ChatGPT sign-in: https://learn.chatgpt.com/docs/models ## What Changed - `packages/adapters/codex-local/src/index.ts`: add `minimumCodexCliVersionForModel` (`gpt-6.1-sol` → 0.159.0, `gpt-6-sol` and `gpt-6-luna` → 0.157.0, with the sources above), `parseCodexCliVersionOutput`, `codexCliVersionAtLeast`, and `CODEX_CHATGPT_MODEL_REJECTION_RE`. Models without a verified floor return `null`, so nothing probes them. The agent configuration doc describes the floors. - `packages/adapters/codex-local/src/server/cli-version.ts` (new): run `codex --version` where the run would execute it (local, SSH, or sandbox), compare it with the model floor, and build the `codex_cli_version_compatible` / `codex_cli_version_incompatible` checks. Map the backend rejection sentence to `codex_hello_probe_model_rejected` with the detected CLI version and a hint that separates a stale CLI from a plan that does not include the model. - `packages/adapters/codex-local/src/server/test.ts` (CLI lane Test): run the version check before the hello probe when the model has a floor. Skip the hello probe when the CLI is too old (`codex_hello_probe_skipped_cli_version`). When the probe still fails with the backend sentence, report `codex_hello_probe_model_rejected` instead of the generic `codex_hello_probe_failed`. The new codes do not match the auth-failure patterns, so a managed connection is not invalidated. - `packages/adapters/codex-local/src/server/acp.ts` (ACP lane Test): for remote targets, run the same version check against the shared `codex` the ACP server spawns. - `server/src/services/native-runtime/native-session-executor.ts`: after the compatibility-window check, compare the remote Codex with the configured model's floor and fail with `runner_remote_provider_artifact_incompatible: <model> requires Codex <floor> or newer with ChatGPT sign-in, received <version> from the sandbox image; promote a sandbox image with Codex 0.160.0 or configure PAPERCLIP_RUNNER_REMOTE_CODEX_NPM_SPEC=...`. When a preinstalled Codex fails this check and an npm spec is configured, the existing fallback installs the pinned release. - Tests: adapter metadata, CLI-lane Test (too old, compatible, no floor, backend rejection), ACP-lane Test (too old, compatible, no floor), and the remote runner floor (nine version/model cases). - `doc/adapter-model-audit-2026-10-02.md`: record the floors and the reason. No catalog entry, default model, saved agent configuration, pin, lockfile, or image definition changes. ## Verification Local (Node 25.9, pnpm 9.15.4 via corepack): ```sh pnpm --filter @paperclipai/adapter-codex-local typecheck pnpm --filter @paperclipai/adapter-codex-local exec vitest run # 81 passed pnpm --filter @paperclipai/server exec vitest run \ src/services/native-runtime/native-session-executor.test.ts \ src/services/native-runtime/codex-runtime-compatibility.test.ts -t Codex # 84 passed ``` Also run locally: the full `native-session-executor.test.ts` file (533 passed) and the full Codex adapter suite (81 passed). Follow-up commit `9fd4d3b25` (Greptile P2: the ACP-lane version probe ignored the agent's configured env): the probe now receives the adapter's string-valued `env` entries, the same ones `buildCodexAcpConfig` hands to remote ACP runs, so a `PATH` override selects the same `codex` for the Test as for the run. New test covers a `PATH` + `CODEX_HOME` override and a dropped non-string entry. Re-run on that head: `pnpm --filter @paperclipai/adapter-codex-local typecheck` clean; full Codex adapter suite 503 passed (31 files). Not run here: `pnpm -r typecheck` for `server` (the direct `tsc --noEmit` was killed by the sandbox memory cap; the adapter package typecheck passes and CI covers the server), `pnpm build`, browser suites, and a live sandbox probe. No UI files changed, so token gates do not apply. Manual check for a reviewer: set an agent's Codex model to `gpt-6.1-sol`, point its environment at a sandbox whose `codex --version` prints `codex-cli 0.156.0`, and run the environment Test. The result contains `codex_cli_version_incompatible` with `Detected Codex CLI 0.156.0.` and no hello probe check. With Codex 0.160.0 the Test contains `codex_cli_version_compatible` and runs the hello probe. ## Risks - A floor is enforced for all authentication modes, because the runner does not know the credential method at verification time. An OpenAI API key on Codex 0.157.0 or 0.158.0 with `gpt-6.1-sol` is now rejected before launch. Paperclip pins Codex 0.160.0 everywhere, so this only affects installs that lag the pin, and the message names the fix. - The 0.159.0 floor for `gpt-6.1-sol` is the oldest stable release verified to work; 0.157.0 and 0.158.0 were not verified either way. If OpenAI accepts an older client, the floor can be lowered in one table. - The `codex --version` probe adds one short process run to the Test only for models with a floor. - Rollback: revert this pull request. No data or configuration migrates. ## Model Used Claude Fable 5.1 (`claude-fable-5-1`), Anthropic, with extended thinking, tool use, and web research, operating as a Paperclip agent through Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Bender (Fable) <noreply@paperclip.ing> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
f669194298 |
fix(runner): propagate configured environment to future turns (#15451)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runners start provider processes and enforce a separate tool environment policy. > - The server resolves task environment bindings at each run boundary. > - Fixed launch allowlists dropped custom variables after resolution. > - Warm process reuse and a blanket fingerprint exclusion also hid configuration changes. > - This pull request carries a bounded, server-selected list of task variable names through each process boundary. > - The benefit is that later turns and tool commands receive configured values while ambient host secrets stay excluded. ## Linked Issues or Issue Description **What happened?** A native Codex agent could not use configured task credentials. Adding the variables in Settings did not fix the next turn. The provider sanitizer, runner launch, Rust provider launch, and shell policy each used fixed allowlists. Effective config fingerprints also excluded all variables with the `PAPERCLIP_` prefix. **Expected behavior** Explicitly bound task variables must reach provider processes and tool commands. Added, changed, and removed values must take effect at the next run boundary. A running turn keeps its original configuration. **Steps to reproduce** 1. Start a native Codex task without a custom environment binding. 2. Add a fake `PAPERCLIP_PAGE_BUCKET` value and a fake Pages credential binding in Settings. 3. Continue the task and inspect the tool environment. 4. Before this fix, those values are absent even from a fresh runner launch. **Paperclip version or commit** Reproduced at `b31558064`. The fix is rebased on current master. **Deployment mode** Self-hosted server with a native process runner. Related changes: [the legacy Codex MCP environment fix](https://github.com/paperclipai/paperclip/pull/13321) and [ambient server-secret exclusion](https://github.com/paperclipai/paperclip/pull/12870). These affect different launch paths. This change preserves their credential boundaries. ## What Changed - Capture scoped task bindings after resolution, before managed provider credential injection. Mint and validate the names-only projection at native dispatch before host inheritance. Legacy adapters retain their previous environment limits. - Strip user-supplied projection markers from agent, environment, project, and routine config. - Carry selected values through the Codex, ACPX, OpenCode, runnerd, and Rust subprocess launch boundaries. - Add selected names to native and ACPX Codex shell include lists. Keep selected values out of command arguments, including selected bootstrap values. - Reject malformed projections, reserved authority and loader names, missing values, null bytes, and oversized input. - Replace a retained native process when projected values change. Keep unchanged processes reusable. - Fingerprint custom namespaced variables while excluding known generated runtime variables. - Document next-run behavior and add regression coverage. ## Verification - Red: the permanent reproduction failed at four launch/tool boundaries and the namespaced fingerprint check. Two control checks passed. - Green: the initial regression plus existing Codex environment and shell tests passed (39 tests). - Server config resolution, fingerprints, and native-session suites passed (653 tests), including addition, rotation, removal, and unchanged warm-session reuse. - Rust regression tests passed. A real shell child received added and rotated values, then lost them after removal. Unselected host variables stayed absent. - `pnpm -r typecheck` and `pnpm build` passed after rebase. The full `pnpm test:run` was attempted but could not complete: fresh embedded PostgreSQL databases fail during bootstrap on this macOS host. An isolated suite and a disposable native `initdb` probe reproduced the failure before test execution. `shmget` reports `No space left on device` because the host has exhausted shared-memory IDs. This is not disk exhaustion. The run was stopped after confirming the external setup failure. CI results will be recorded separately. - Runner boundary suites: 361 tests passed. Two process-launch errors during concurrent binary staging passed on isolated rerun. - Rust Codex provider and process supervisor integration suites: 97 passed, 2 intentionally ignored. - ACPX shell and selected-bootstrap argv regressions: 3 failed before the fix, then all 55 relevant tests passed. - Review compatibility regression: 129-variable and large-value legacy configurations failed before the correction and passed after moving native-only validation to dispatch. - Final review head `3f11e8d25`: repository typechecks and production build passed. CI completed its implementation checks; one general-server shard hit SQL `40P01` in `heartbeat-runtime-skills.test.ts` during its `beforeEach` table truncate (1,199 tests passed in that shard). The shard passed on its single rerun. All CI checks for this head are green. Greptile reviewed this head at 5/5 with no new actionable findings; the compatibility thread is resolved. - No live provider credentials or model calls are required by these tests. ## Risks - Configured task credentials now reach the tools they were configured for. The controller selects names only after existing scope and secret-binding authorization. - TypeScript and Rust validate the same bounded projection. Their reserved-name rules must stay aligned. - A changed projection replaces an idle provider process. Unchanged values preserve reuse. Active turns retain their original environment. - No database migration or API schema change. > This fixes existing runner configuration behavior. It does not add a new roadmap capability. ## Model Used - OpenAI GPT-6 (Codex), with reasoning, local code execution, and repository tools. The session does not expose a more specific API model identifier or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cfedd7df8 |
fix(runner): restore task monitors and durable timed waits (#15446)
## Thinking Path > - Paperclip manages AI agents and their task execution. > - Agents need a durable way to return to work after a delayed check. > - The issue monitor scheduler already provides a one-shot wake for an assignee. > - Native runners reject generic execution-policy writes and had no bound monitor tool. > - Scheduling alone is insufficient because native completion also needs to accept a timed wait. > - This pull request adds an authorized monitor tool and connects it to completion and the existing scheduler. > - An agent can now schedule its next check, end the run, and resume on the same task. ## Linked Issues or Issue Description **What happened?** A native runner could not set its own task monitor. `call_api` correctly rejected execution-policy writes, while `schedule_wake` had no production binding. `paperclip_finish` also rejected monitor waits. **Expected behavior** A standard native run can set a one-shot monitor on its current task or another accessible task assigned to the same agent. After a confirmed schedule on the current task, it can yield. The scheduler later delivers `issue_monitor_due`. **Steps to reproduce** 1. Start a standard native task. 2. Ask the agent to check the task again later and end its current run. 3. Inspect available tools and try the generic issue execution-policy update. 4. Observe the missing native tool and the lifecycle-write denial. Related PRs: #14680 concerns monitor notes in the shared wake prompt. #11919 changes attempt-limit scope. This PR adds native scheduling and completion authority and retains the existing cumulative attempt bounds. It does not depend on either PR. ## What Changed - Add provider-neutral `set_task_monitor` with a default current-task target, future timestamp, required notes, existing bounds, and explicit clearing. - Check company, task visibility, ownership, runtime permissions, work mode, and active-run authority. Preserve review-only restrictions. Reject the reserved server-owned quota-recovery name before saving or accepting a native wait. - Commit the monitor, audit event, and retry receipt together. Retry receipts survive a successor run without re-arming cleared or consumed timers. - Permit `paperclip_finish` to yield to a persisted monitor. Recheck ownership and the schedule when committing final disposition. Release execution without an immediate continuation. - Preserve due monitors during native execution. Fence wake admission and consumption against replacement, clearing, reassignment, and completion. Preserve unrelated review policy. - Expose scheduled and consumed monitor instructions in task context. Update provider schemas, Rust validation, generated contracts, and execution documentation. - Add an opt-in live Codex smoke script with isolated data and explicit run/session/runner/process evidence. ## Verification - Repository `pnpm -r typecheck` and `pnpm build` passed after rebase. Server typecheck passed again after review fixes. All CI test shards pass on `3def77b1b`, including runner TypeScript/Rust, server, serialized server, workspace, and browser tests. All CI gates are green, including the canary dry run. Greptile is 5/5 on the same commit with zero unresolved threads. - The local monolithic `pnpm test:run`, started before the rebase, was interrupted after current-head CI test coverage passed. It is not counted as a standalone full-suite pass; the focused local regression suites passed. - Targeted server tests cover scheduling, replacement, clearing, policy preservation, cumulative bounds, cross-run retries, permissions, provider-neutral discovery, review restrictions, completion authority, and scheduler/finalizer races. - Runner contract/catalog/semantic tests and Rust terminal-tool tests cover the new operation and monitor completion. - Live Codex test passed twice (latest live run on `e0bcd63e6`) in a temporary database and workspace, with a 300,000 ms warm window. First run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at `2026-10-07T13:15:03.274Z`. Second run `97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`, received `issue_monitor_due`, and completed the same task. Exactly one monitor wake was recorded. - Both live runs used native session `7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner `e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session `01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same process start time. This proves warm reuse for that local Codex test, not only successful scheduling. - Reproduce the paid live test with `node --import ./server/node_modules/tsx/dist/loader.mjs server/scripts/smoke-native-task-monitor.ts --run`, with the installed Codex binary on `PATH` and a valid local login. ## Risks - The scheduler now defers monitor dispatch while the task has an active native run. A stuck run still depends on the existing recovery lifecycle. - Idempotency uses the existing run ledger; no table or migration is added. - Other providers share the tested tool and completion contracts. Only Codex received a live model test. - Existing `call_api` lifecycle restrictions remain enforced. Monitor waits do not bypass task blockers, reviews, or approvals. ## Model Used OpenAI Codex, GPT-6 family, with tool use, code execution, and TypeScript/Rust editing. The session does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
abbd88007f |
fix(hermes): keep managed instructions out of resumed user turns (#15439)
Deliver managed instructions through Hermes's native system overlay while keeping current wake and runtime identity in user turns. Add fresh/resumed regression coverage and include Hermes in the default CI test roster. Fixes #15385 Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3a726e676f |
feat: enforce private task permissions across execution and data (#10633)
Enforce private task and project access across direct reads, search, execution, files, plugins, live delivery, and sharing mutations. Preserve downward-only sharing, current responsible-user authorization, and audited emergency access. Bind historical draft assets with migration 0314. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b67db12d90 |
feat: add durable storage for private tasks (#14717)
Add private task and project ownership, downward access grants, and immutable run/workspace provenance. Apply migration 0313 with bounded batch commits and concurrent indexes. Keep the enforcement and sharing changes in their dependent PRs. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1ead554bd1 |
fix(claude-local): resume sessions across agent file working copies (#15437)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Claude CLI adapter resumes task sessions between runs. > - Each run now receives a private copy of the agent files. > - The adapter put that copy's path in the cached system instructions. > - The new run ID changed the prompt bundle even when all instructions and skills stayed the same. > - This pull request sends the current file location with each run's prompt and keeps it out of the stable bundle. > - Agents can resume unchanged task sessions and use the current working copy. ## Linked Issues or Issue Description Fixes: #15373 Refs #14420, which introduced the per-run agent directory copies. Related PRs checked: #5699 changes the fingerprint algorithm, and #12034 handles saved sessions without a bundle key. Neither fixes this working-copy path regression. ## What Changed - Keep instruction and skill contents in the cached system prompt. Supply the current instruction path and relative-file base in every run prompt, including resumed turns and fresh retries. - Report a cwd or execution-target mismatch only when that value differs. A bundle mismatch no longer produces a false cwd warning. - Add a four-run regression: initial run, relocated copy, changed instructions, and changed skill contents. Check the CLI arguments, bundle keys, current file guidance, and reset logs. - Cover and explain remote-to-local execution resets, even when the working directory matches. - Update the agent-file documentation and existing resume/fallback assertions. ## Verification - **Red:** With only the new regression test added, the second run fails because the CLI arguments do not contain `--resume`. - **Green:** All 64 tests pass in the command below. This uses a fake Claude subprocess and real adapter execution, file caching, and session serialization; it does not call a paid model. ```sh pnpm exec vitest run server/src/__tests__/claude-local-execute.test.ts packages/adapters/claude-local/src/server/execute.remote.test.ts packages/adapters/claude-local/src/server/execute.acp-fallback.test.ts server/src/__tests__/adapter-session-codecs.test.ts ``` - Repository-wide `pnpm -r typecheck` and `pnpm build`: passed. The adapter typecheck, build, and all 64 focused tests also pass after the review fix. - Local `pnpm test:run` was stopped after it reported a failure in the unchanged native-session recovery database orchestration test. That test passes in isolation with PostgreSQL enabled (1 passed, 47 filtered out). The full local run did not complete; this is not a clean local full-suite result. All CI checks pass on `a71d23a39f1cc874d23a8715cf29bdea6edbb8ff`, including the full test shards. ## Risks - A session saved with the old path-bearing bundle starts fresh once after upgrade. Later runs resume when instruction and skill contents stay unchanged. - The current location now travels in the run prompt. Stable system guidance directs relative file references to that location, and each turn explicitly replaces earlier locations. - No database, API, authentication, permission, or UI change. ## Model Used OpenAI Codex (GPT-6), with reasoning, repository inspection, code execution, and test tools. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d0db8820db |
fix(connections): repair native baseline and approval continuations (#15420)
## Thinking Path > - Paperclip manages AI agents and the tools they may use. > - Connection setup separates provider preference from permission to use a tool. > - The first native connection baseline could not exercise its intended decisions. > - The browser used mutable task titles, and the provider fixture already granted access. > - Native provider-choice instructions also disagreed with the preferred question format. Schema rejection gave no field guidance. > - This pull request repairs those test preconditions and native guidance, then fixes restart/approval defects exposed by the corrected baseline. It also restores missing OpenCode tool-error evidence. > - The benefit is an inspectable baseline before any further instruction reduction. ## Linked Issues or Issue Description Refs #15407. The original 15-cell baseline remains 0 PASS / 15 FAIL. Ten cells stopped on stale titles, two Codex cells had schema denials, two OpenCode cells used already-granted tools, and one Claude cell returned no native result. No intended user decisions were submitted. The exact invalid Codex field and underlying Claude failure cause remain unknown. [Original campaign](https://github.com/paperclipai/paperclip/actions/runs/37562577199) · [Original report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37562577199-1/index.html) ## What Changed - Match the browser's task route and visible identifier instead of a title the agent can change. - Start native provider-choice fixtures with no agent tool access. Verify the public effective-access records. - In the positive case, select Arcade, then grant its exact HubSpot tool through the real access card. Require both saved decisions and exactly one observed call. - Return a canonical `providerQuestionSet` for native input and retain the equivalent legacy `providerQuestion`. - Keep invalid input rejected. Return bounded schema locations and required field names without submitted values. - Preserve Claude's exact session/content identity while allowing authenticated registered instruction-copy paths to rotate on a new run. - Reject duplicate approval reports for an existing exact tool-action card before they create another human review. - Wait for a recorded service-approval continuation within the existing deadline; retain missing or failed continuation grades. - Forward OpenCode tool activity through the runner facade, preserving bounded errors and execution-part identity without inventing host-call joins or exposing arguments. - Preserve original grades, costs, scope limits and diagnoses in the dated repair report. ## Verification - Eval typecheck passes. Support suite: 1,801 PASS, one intentional skip; Node checks: 128 PASS. - Connection/schema tests: 51 PASS. Real-server public fixture setup: one PASS with zero providers. - Browser support regression: five PASS, including renamed and wrong tasks. - Focused Rust safe-feedback test: one PASS. - Repository typecheck and build pass before the latest master replay. Post-replay connection/shared/real-server fixture checks: 52 PASS; eval typecheck passes. The browser review fix additionally passes all five browser checks and seven suite checks. - The full local repository run was interrupted incomplete after about 45 minutes, with five integration failures retained. All five pass in a separate targeted invocation (1,250 unrelated tests skipped). No full local-suite pass or root cause for the initial local failures is claimed. - Corrected frozen source `162cc90fdabe7f505b88ae095044531b82784c92`: **10 PASS / 5 FAIL** across the [passing Codex canary](https://github.com/paperclipai/paperclip/actions/runs/37575158761) and [remaining 14 cells](https://github.com/paperclipai/paperclip/actions/runs/37576261807). The canary passes all 17 checks. Claude's two provider-choice continuations fail on restart, Claude service approval exposes an early evaluator rejection, Codex service approval creates a duplicate approval, and OpenCode provider-second times out after both decisions with no HubSpot call. No original result is regraded. - Final ledgers count 31 actual runs: 27 succeeded, two failed, two cancelled during cleanup. All 15 cleanup/budget checks pass. The late Claude continuation is absent from its earlier workflow snapshot; it remains in the result/API/final ledger. Original evidence retains 279 hashes. Recorded LLM subtotal $0.04553787 is incomplete billing, not actual total cost; local runtime is unmetered. - New repair regressions reproduce the Claude attach failure, duplicate approval acceptance and dropped OpenCode tool events before their respective fixes. Nine Rust attachment checks, 127 ACPX host/adapter tests, 33 completion/control-plane checks, nine eval deadline tests, 59 OpenCode proxy/driver tests, one Rust tool-error/redaction check, and TypeScript/Rust composer parity pass. Eval typecheck, repository typecheck and build pass. Existing support coverage is 1,802 PASS plus 128 Node PASS, one intentional support skip; two additional deadline tests also pass. - New-source full CI/review and live canaries are pending. The next bounded selection is Claude provider-decline, Codex service-approve and one OpenCode provider-second diagnostic with repaired event evidence. No broader campaign or instruction-reduction qualification is claimed. - Initial corrected campaign [37574251834](https://github.com/paperclipai/paperclip/actions/runs/37574251834) was cancelled during shared build after review found the breadcrumb whitespace assumption. Its matrix job has zero steps and no provider execution. The real adjacent-span browser regression now reproduces the old failure and passes after the fix. ## Risks - The corrected baseline remains 10/15. The new restart/approval fixes require live qualification; OpenCode evidence forwarding does not itself establish or fix its prior behavioral failure. - The positive provider case now expects three runs, including separate access approval. Its new results are distinct from the original invalid fixture. - The old Claude missing-result cause and rejected Codex field are unknown. These repairs do not retroactively explain or erase either failure. - Path rotation must preserve prompt, custom instruction, skill/content identity and protected provider settings; regression checks reject stale or changed content. No connection authorization, JSON schema, budget, cleanup, or final-result requirement is relaxed. Historical Everyday prompts and gateway setup remain unchanged. ## Model Used OpenAI Codex, GPT-6. The exact deployment variant and context window are not exposed in this session. Used repository inspection, code editing, test execution and retained-evidence analysis. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
799e4d556f |
fix: make accounting durable and synchronize cost reporting (#14997)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
99a9de9940 |
fix(mcp): personalize assistant connections and hide revoked grants (#15411)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Assistant connections let people use their organization from another assistant. > - Setup begins with an invitation and returns to a list of connected assistants. > - Client-only labels obscure the authorizing person, while revoked rows clutter that list. > - This pull request shows the person’s avatar, names connections by owner and client, and hides revoked rows. > - It also shortens the copied invitation while keeping the approval instructions. ## Linked Issues or Issue Description Related: #14933 and #15380. **What existing behavior does this improve?** Invitation copy and assistant connection management, in the organization’s Connections screen and the account-wide management page. **Subsystem affected** Cross-cutting: shared MCP connection types, server profile projection, UI and Storybook. **Current behavior** Connections are labeled only with a client name such as “Codex.” Revoked connections remain visible. The invitation includes an extra sentence about agent identity. **Proposed behavior** Show the authorizing person’s avatar and use names such as “Dotta’s Codex connection.” Hide revoked rows after successful revocation and when loading retained revoked grants. Failed revocation leaves the connection visible. Remove the extra identity sentence from invitation copy. **Reason and benefit** Make connection identity clear and keep the list focused on usable connections. **Breaking changes** The connection response adds optional `user` metadata with name and image. Older servers remain usable. Names and revoked-row visibility change in the UI; OAuth client identity, authorization and audit retention remain unchanged. ## What Changed - Shorten the shared invitation text. - Project the authorizing person’s name and avatar through the user-scoped connection endpoint, without returning email or credentials. - Reuse the existing Identity component and owner naming conventions across the connection page, catalog card and account-wide list. - Hide revoked grants and remove a successfully revoked row from the shared cache, even if the subsequent refresh fails. - Update documentation, regression tests and production-page Storybook fixtures and revocation journeys. ## Verification - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook` and `pnpm check:token-gates` pass. Final UI type checks also pass. - Focused consent and connection UI tests: 31 pass, including company filtering, legacy metadata, custom client names, user-initial fallbacks, failed revocation, retained revoked rows and refresh failure after successful revocation. - Existing MCP regression suite: 6 pass; 71 database checks are skipped locally because embedded PostgreSQL cannot start on this machine. The added database check verifies user-profile isolation and retained revocation history; CI runs these checks. - Browser verification with Storybook fixtures: owner avatar and name render; revocation removes the selected row in both production pages, leaves other connections visible, and restores the empty state after the last revocation. - All 54 current-head CI checks pass, with two optional Storybook jobs skipped. CI includes database, browser, runner, typecheck, build and clean-install canary coverage. - Greptile reviewed commit `1368d79e1066b418712224378d89d64c2b11cb86`: 5/5, no actionable findings or unresolved threads. - The full local `pnpm test:run` was stopped after complete CI passed. Local database coverage remains unavailable because embedded PostgreSQL cannot start; no full local-suite pass is claimed. ## Risks Low risk. The additive profile field is optional for compatibility. Revoked grants are filtered only from management UI and retained for audit. Revocation failure does not hide an active connection. No authorization scopes, token handling or schema changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, command execution and browser verification. The exact deployment ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a9a20fb5c6 |
feat(security): add read-only customer-success inspection APIs (#15405)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a persistent identity and a verified active run for governed access. > - Customer-success inspection needs broad reads without tenant writes or secret access. > - Ordinary board login and database credentials give more authority than this task needs. > - This pull request adds a dedicated inspection API and strict managed-run authority. > - Cloud owns short grants, human approval, replay protection, and audit records. > - The benefit is inspectable access that an operator can disable immediately. ## Linked Issues or Issue Description Refs: #15352. This change reuses the persistent Ed25519 identity from that PR. **Subsystem affected** Server authentication, pure resource readers, and the shared wire contract. **Problem or motivation** One internal Paperclip agent must inspect customer onboarding work. It must not receive owner login or database credentials. Reads must not create customer sessions, memberships, activity, or read receipts. **Proposed solution** Add disabled-by-default run authority and versioned tenant inspection endpoints. Require strict instance-bound managed-run JWTs on the home instance. Require exact-operation, single-use Cloud permits on tenants. Execute a reviewed company-scoped catalog in read-only transactions. Cloud applies seven-day stack-age eligibility and human exceptions. **Roadmap alignment** This is access support for Cloud deployments and governed agent identities. Bot creation, scheduling, scoring, and reports are separate work. The maintainer requested this implementation. ## What Changed - Reuse existing public identity reads and managed private-key injection. Reject unprovisioned keys, paused agents, ended runs, legacy signatures, and wrong instances. - Mount `/api/customer-success/v1` before actor/session synchronization. Verify Cloud permits and consume them centrally before reading. - Add explicit company-scoped database readers and bounded instruction, skill snapshot, run log, workspace, and asset reads. Preserve existing redactions and file protections. - Add protocol, security, database immutability, and managed-agent qualification tests. Add deployment and rollback documentation. ## Verification - Full `pnpm -r typecheck` and `pnpm build` passed. Server typecheck passed after review fixes. - The broad local `pnpm test:run` recorded 14,277 passes and four failures in unchanged suites: two timeouts and two PR-metadata mock assertions. All three affected suites passed on isolated reruns (36 tests). The complete CI matrix passes at the final head, including every test lane, typecheck, build, runner checks, canary dry run, and the security scan. - Focused inspection, JWT, and existing identity tests pass. The catalog test compares every public database table before and after reads. - Inspection and route-contract tests: 22 passed. The coordinated test runs a real managed process agent against separate home/customer PostgreSQL databases and a PostgreSQL broker over HTTP. It proves wake through the existing controller, bounded binary file reads, single challenge consumption across replicas, concurrent grants with a two-connection pool, scoped SQL audits, append-only runtime auditing, one-year retention, and unchanged tenant data/files. - Run the coordinated test with `PAPERCLIP_INSPECTION_CLOUD_DIST` pointing at the sibling Cloud build. Normal unit runs skip that optional private integration. - Final-head Greptile is 5/5 with no unresolved findings. - No production deployment or customer inspection occurred. ## Risks - This adds an authentication boundary. Keep both feature flags disabled until coordinated staging and canary qualification. - Cloud support must deploy after this API. Unsupported tenants fail closed. There is no owner-login or database fallback. - Existing redactions remain the content boundary. Arbitrary pasted secrets in readable prose or files may remain. - Remote files and suppressed provider traces remain unavailable. Wake can cause normal startup/background writes; test those separately. - Disable Cloud policy first during rollback. Preserve existing identity material and Cloud audit history. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, and browser testing. The session does not expose a more specific deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; broad-run flakes are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a6306ba606 |
feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
faa8e452c7 |
fix(tasks): stop repeated reminders for historical questions (#15392)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task questions remain saved so a person can answer them later. > - The composer moves an old question into history after a newer human message. > - Native completion still treated every pending question as a required response. > - This caused agents to demand an old answer after the person moved work forward. > - This pull request shares the historical-question rule across context and task execution. > - Agents can finish verified work while the original question remains answerable. ## Linked Issues or Issue Description Refs #15229, Refs #14613, Refs #13130. **What happened?** An agent repeatedly asked a person to answer a question that had moved into feed history. Completion feedback explicitly told the agent to request a response. The saved pending row also blocked task completion. **Expected behavior** A question before newer human direction remains answerable in the feed. Its pending state alone must not require another reminder or stop completed work. A new input blocker, an approval, or a configured review stage must keep its gate. **Steps to reproduce** 1. Let a task agent create an ordinary question. 2. Dismiss the question and send a newer task message. 3. Let the agent finish the requested work and submit its completion report. 4. Observe a demand to answer the old question and a retained completion gate. ## What Changed - Add one company-scoped predicate for historical questions. Only later human comments count. Exclude agent attribution, run attribution, system notices, and untrusted source data. - Apply the predicate to completion feedback, native waits, finalization, commit validation, retry validation, blocked routing, and successful-run handoff. - Include question classification and guidance in heartbeat context and both native task-context tools. Add the guidance to fresh and resumed task prompts. - Replace automatic reminders for current ordinary questions with instructions to assess the real blocker, continue independent work, and withdraw obsolete questions through the existing API. - Preserve historical question rows during completion while cancelling their live native source runs through the existing post-commit and recovery paths. Let an authorized human answer them after completion without reopening work or creating a response wake. Preserve cancellation, current-input, approval, permission, credential, connection, and review gates. Add no dismissal storage or migration. - Document the rule in the execution contract and agent skill. Refresh generated capability source anchors. Add database-backed status, context, attribution, and governance regression tests. ## Verification - `pnpm build` passed. The server rebuild also passed after the lifecycle fix. Generated capability contract and inventory checks passed after the agent documentation update. - `pnpm -r typecheck` passed. Final `pnpm --filter @paperclipai/server exec tsc --noEmit` also passed after the last test additions. - Lifecycle and interaction regressions passed: 224 tests in 3 suites. Context and prompt tests also passed. Final historical-question cases passed (33 tests), native cancellation/recovery cases passed (6 tests), and the existing interaction/confirmation suites passed (72 tests). The full local `pnpm test:run` was attempted and stopped after more than two hours with unrelated fixture/hook timeout failures; it did not pass. All 52 successful GitHub checks are green on the latest commit, including the complete test matrix; no checks are pending or failing. Greptile is 5/5 and both review threads are resolved. - Regression cases cover the old-question/new-human-message sequence, final task status, answering after completion with no wake, live-run cancellation and crash recovery, current input blockers, both task-context tools, API context, timestamp precision, attribution boundaries, and protected gates. ## Risks - A later human task message makes an earlier ordinary question historical even if its input is still missing. The agent must identify the current blocker and ask only for information that still prevents work. - Browser dismissal remains a local preference. Dismissal without a later human message is not recorded by this change. - No schema change or data migration. Completion retains ordinary historical questions; cancellation still expires them. A completed task accepts historical answers only from an authorized human and creates no response-delivery outbox row. Governed requests retain their gates. ## Model Used OpenAI Codex, GPT-6, with repository editing, code execution, and browser diagnostics. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2d0c138122 |
Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also use assistants in Codex, Claude, and other MCP clients. > - The existing assistant connection can read work and create tasks or comments. > - It cannot edit tasks, exchange files, or manage normal agent and project settings. > - These operations must retain the person's permissions and Paperclip's execution rules. > - This pull request adds an explicit operation registry and separately consented configuration access. > - Assistants can manage work without receiving credentials or runner authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: assistant MCP, domain routes, consent UI, storage, and Product E2E. **Problem or motivation** A connected assistant cannot update tasks, maintain documents, attach files, or configure existing agents, projects, and skills. Users must leave the assistant for these routine actions. **Proposed solution** Add named tools and a restricted API registry to direct connections. Require separate configuration consent. Reuse domain routes and retry receipts. File uploads save the attachment when the byte transfer succeeds. **Alternatives considered** Arbitrary REST forwarding would expose administration and credential operations. Runner impersonation would bypass execution ownership. Separate upload completion calls add unnecessary client state. **Roadmap alignment** Checked ROADMAP.md and related MCP pull requests. This extends the human-authorized connection from #14933. It does not replace the runner or introduce agent impersonation. Companion Cloud routing and directory isolation: https://github.com/paperclipai/paperclip-cloud/pull/678. ## What Changed - Add task editing, finish/block, documents/revisions, deliverables, agent settings/instructions, projects/repositories, and skills/files. - Add an allowlisted API search/call registry with identical field restrictions, scopes, and retry identities. - Add unchecked configuration consent. Existing write grants retain their current authority. - Add hashed, expiring file transfer tickets and atomic upload receipts. No completion call is required. - Preserve company boundaries, human attribution, active native execution ownership, and execution review gates. - Add protocol/domain tests, consent stories, and eight paid Product E2E workflows. - Repair two CI fixture races: await cold route setup before assertions, and wait for asynchronously loaded connection copy. Both fixture suites pass (24 + 48 tests). ## Verification - Consent revision: one write-access checkbox controls requested work and configuration permissions in browser and device flows. All 16 consent tests, UI typecheck/build and token gates pass. Updated interactive stories cover default approval, opt-out and viewer restrictions. The paid browser helper uses the new exact label. Real GPT-5.4 Mini Product E2E passes 2/2 at `64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic retries, cleanup passed; $0.04149375 estimated assistant cost plus unpriced worker usage. Raw results, usage and source fingerprints are retained in the worktree. UI and Product E2E typechecks pass. - Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks pass; two optional Storybook checks skip. Greptile 5/5 on that head, no unresolved review threads. Final consent head `64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing and Greptile 5/5 with no unresolved threads. The unchanged Cursor sandbox test had one 10-second timeout, passed in local isolation, and passed its single CI rerun; the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37554106934). The existing chat retry-denial browser test had one visibility failure; its single rerun passes, and the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37542735691). - Full workspace `pnpm -r typecheck` and `pnpm build` pass at final runtime source `b2196fae1`. UI token gates pass. - 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests pass, including one-connection PostgreSQL OAuth and concurrent upload retries. - Paid Product E2E: all eight expanded cases qualified across Mini, Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell timed out before application startup. Final affected-case qualification passes 9/9 on all three models with grader v16, including the failed cell. Automatic retries disabled; failures, costs, source hashes and independent durable-state/file assertions are retained in [the verification record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md). - Actual Codex CLI, Claude Code and OpenCode clients completed local reads/mutations. Codex wrote a report, Claude updated it in a later conversation, and OpenCode uploaded/downloaded a file with matching SHA-256 and registered the attachment. Revoking the CLI grant rejects subsequent bridge initialization. - Butter staging is verified on final runtime `b2196fae1` ([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)). A fresh OpenCode workspace fetched the copied invitation, configured remote MCP, started OAuth and reached real consent with configuration unchecked. Invalid transfer tickets return 403 through Cloud. Human approval for the new persistent staging grant is pending; hosted task/file success is not yet claimed. The final transaction fix is deployed. - Full local `pnpm test:run` passed 15,614 general-server tests but stopped on two macOS timeouts. The heartbeat test passed in isolation; the existing 40,000-file Git stress fixture timed out again. Its Linux CI lane passes. Later local full-suite phases did not run after the timeout; this is not an all-green local full-suite claim. - Instructions and security limits are in `doc/public-mcp.md`; the saved plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`. ## Risks - This expands the experimental direct MCP surface. Explicit schemas and domain permissions must stay synchronized. - Migration 0311 adds transfer tickets and upload receipts. Expired orphan cleanup must not remove committed attachments. - Configuration requires a new consent request containing that scope; the single write-access choice controls it alongside work mutations. Refreshing an old grant does not add it. - The public directory keeps its original ten tools through the companion Cloud change. - Hosted consent/work proof remains the final delivery gate. The PR stays draft while approval of the new staging grant is pending; code checks and review are green. Merging is a separate action. ## Model Used OpenAI Codex (GPT-6, tool use and code execution). The exact serving model ID and context window are not exposed in this session. Paid evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001; claude-sonnet-4-6. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite macOS limitation disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
892b0b3606 |
fix: scope quota reports to authorized subscription accounts (#14994)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
eab93fd4a0 |
fix: checkpoint adapter usage and preserve unknown prices (#14991)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |