mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
5716fe907e596ce73501408fc6efdb19fb61edf2
1644
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fdf8c8464d |
feat(runner): add managed provider backends (#12699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner provides durable, provider-neutral agent execution. > - The current stack supports qualified local providers but omits the managed provider paths from the integration branch. > - Claude Managed Agents and AWS AgentCore need explicit profile qualification, durable recovery, usage accounting, and cleanup controls. > - This pull request adds those managed backends as the third part of the Runner parity stack. > - The benefit is managed execution without weakening the default-off Runner rollout gate. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: Runner, server orchestration, database profiles, CLI, and adapter configuration UI. **Problem or motivation** The current Runner stack cannot select or execute the managed Claude Agents API or AWS Bedrock AgentCore Harness backends. It also lacks qualified profile storage and recovery checks for those remote resources. **Proposed solution** Add qualified managed and remote profiles, API and CLI management, exact provider selection, durable lifecycle handling, cumulative usage accounting, bounded cleanup, and retention acknowledgement. Keep `enableNativeRunner` default-off. **Alternatives considered** A direct copy of the old integration branch was rejected because its provider contracts, model values, credential flow, and migration history no longer match the current base. A single large parity pull request was also rejected because stacked review keeps each subsystem bounded. **Roadmap alignment** This continues the existing Runner architecture and rollout work. It does not introduce a separate execution system. **Additional context** This pull request is based on the merged #12691 and #12685 stack. It also closes the delayed security-review findings reported on #12691 by binding qualified ACPX and OpenCode launch artifacts to the bytes actually executed. A GitHub search for managed agent, AgentCore, and Claude managed work found no duplicate public issue or pull request. ## What Changed - Add Claude Managed Agents and AWS AgentCore provider executors to runnerd. - Add qualified managed and remote profile storage, routes, OpenAPI contracts, CLI commands, and migration 0237. - Validate profile ownership, enabled state, exact qualified revision, model, agent version, and secret binding before persistence and recovery. - Persist durable provider session and owned skill state for restart-safe cleanup. - Reconcile uncertain create responses and delete remote sessions before owned skills. - Track cumulative provider usage and enforce positive session spend caps. - Recover interrupted AgentCore usage at the next turn boundary by charging the prior invocation ceiling exactly once; keep the session gated until an explicit monotonic budget raise. - Isolate AgentCore AWS configuration from host profiles and credential-process/SSO configuration while preserving workload identity. - Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact paths; remove the ambient executable override. - Snapshot and content-verify ACPX and OpenCode commands, scripts, and provider executables before launch. Linux executes sealed inherited descriptors; macOS uses authenticated private snapshots with retry-safe rematerialization at the spawn boundary. - Persist canonical ACPX and OpenCode launch-profile digests, reject drift across fresh recovery, and make recovery failures sticky. - Close and journal unsafe ACPX active-turn recovery before any provider bootstrap or reconnect. - Add managed provider fields to the Runner configuration UI and permission projection. - Preserve the default-off `enableNativeRunner` experimental flag. ## Verification - `pnpm -r typecheck` - `pnpm build` - Focused managed server, database, CLI, Runner TypeScript, Rust, Claude, AgentCore, ACPX, OpenCode, process-supervisor, and durable-recovery tests passed. - `cargo test -p paperclip-runner-core --lib --locked` (160 tests) - `cargo check --workspace --all-targets --locked` - Native Codex integration tests passed (60 tests); native provider tests passed (7 tests); server native-runtime tests passed (87 tests). - Verified-launch replacement, nested-spawn retry, exact-version, profile-drift, sticky-failure, and no-bootstrap active-recovery tests passed. - `git diff --check` - The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust workspace lockfile adds the approved `rustix` dependency used for safe descriptor handling while `#![forbid(unsafe_code)]` remains enabled. ## Risks - The provider APIs can change while they are in beta. Exact qualification and fail-closed recovery checks limit drift. - Remote cleanup can fail after a partial create. Durable ownership inventories and retry-safe deletion preserve recovery state. - Migration 0237 adds profile tables. The generated migration and snapshot pass the repository migration checks. - Managed execution can incur provider cost. Positive default spend caps and explicit retention acknowledgement limit accidental use. - An interrupted AgentCore invocation without final metadata is conservatively charged to its active session ceiling. This can overstate cost, but cannot undercount it; later work requires an explicit budget increase. - Linux qualified launches use sealed memory descriptors. macOS lacks executable-descriptor APIs, so the runner uses owner-only private snapshots and minimizes linked-path lifetime; hostile same-UID processes remain outside the documented local-host trust boundary. - The global Runner feature remains default-off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
84bedd4ca1 |
feat(runner): activate qualified OpenCode and ACPX providers (#12691)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner is the experimental native runtime for governed agent work. > - The runtime contracts already describe Codex, OpenCode, and ACPX providers. > - The merged control plane still rejected OpenCode and ACPX for new runner agents. > - Runnerd also selected only the Codex provider implementation. > - This pull request activates the qualified OpenCode and ACPX paths from the form to runnerd. > - The benefit is one durable runner path with provider-specific permissions and recovery. ## Linked Issues or Issue Description Refs #12685 **Subsystem affected** This change affects the runner package, server orchestration, adapter configuration, and UI configuration. **Problem or motivation** Paperclip Runner stores provider contracts for OpenCode and ACPX. New agents cannot select those providers. Runnerd cannot execute those stored provider descriptors. The UI also shows only Codex. **Proposed solution** Accept the qualified OpenCode 1.18.17 profile and the fixed ACPX Claude and Codex profiles. Route them through runnerd. Keep provider selection, model selection, permissions, credentials, events, and recovery inside closed provider-specific boundaries. **Alternatives considered** One option was to keep the contracts dormant. That option leaves stored configuration and runtime behavior out of sync. Another option was to enable every ACPX agent. That option is not safe because Pi does not yet have the same verified launch path. **Roadmap alignment** This change supports the completed cloud and sandbox agent milestone. It also supports self-healing runs and governed agent execution. It does not add a new roadmap surface. ## What Changed - Add one server profile resolver for Codex, OpenCode, and qualified ACPX descriptors. - Keep `adapterConfig` as the provider and permission authority for fresh runs. - Add Paperclip Runner provider, ACPX agent, and provider-specific permission controls to the UI. - Reset the model to a compatible qualified value when the provider changes. - Route Codex, OpenCode, and ACPX through the durable runnerd provider selector. - Add a durable ACPX executor with bounded state, recovery, events, tool receipts, and identity checks. - Remove Codex labels from OpenCode events, results, evidence, and recovery diagnostics. - Pass only provider-specific credential names to child processes. - Keep ACPX Pi unavailable and reject it before process launch. - Keep the existing Paperclip Runner experimental flag unchanged. ## Verification - `pnpm exec vitest run packages/paperclip-runner/src/backends/native-backend-factory.test.ts packages/paperclip-runner/src/live/runnerd-codex-transport.test.ts packages/adapters/codex-local/src/ui/build-config.test.ts ui/src/adapters/codex-local/config-fields.test.tsx server/src/__tests__/adapter-registry.test.ts server/src/__tests__/adapter-routes.test.ts server/src/__tests__/agent-adapter-validation-routes.test.ts server/src/__tests__/company-portability.test.ts server/src/services/native-runtime/runtime-mode.test.ts server/src/services/native-runtime/native-session-executor.test.ts server/src/services/heartbeat-runner-provider-config.test.ts` - The focused TypeScript, server, and UI suites passed 274 tests. - `cargo test -p paperclip-runner-core --test native_provider_backend` - The executable native provider integration suite passed 4 tests. - `cargo test -p paperclip-runner-core --lib` - The Rust unit suite passed 91 tests. - `pnpm -r typecheck` - `pnpm check:token-gates` - `pnpm build` - `git diff --check codex/runner-parity-task-runtime...HEAD` ## Risks - This changes provider process selection and durable recovery. The experimental flag still gates every fresh Paperclip Runner run. - OpenCode requires a model in `provider/model` form and stays pinned to version 1.18.17. - ACPX accepts only exact Claude and Codex profile versions and models. Pi stays unavailable. - ACPX steering stays unavailable and reports that limit through the driver capabilities. - Child processes receive explicit environment allowlists. They do not inherit the full server environment. - This pull request has no database migration. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
72b9f92d76 |
fix(runner): restore task runtime parity (#12685)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task view shows a running agent and lets an operator guide that agent. > - The merged runner stack lost parts of the accepted task experience. > - Native event errors could hide current reasoning from the operator. > - Queued message steering had no server route on `master`. > - This pull request restores the task-runtime behavior and keeps the runner experimental gate. > - The benefit is a visible and steerable native run with durable fallback behavior. ## Linked Issues or Issue Description **What happened?** The task view could stop showing current runner reasoning. The steering action also failed because the server route was absent. Runner instruction files were not declared as supported. **Expected behavior** The task view must show current provider activity. It must use the live log when durable native events are empty or unavailable. The operator must be able to steer a queued message into the active native turn. **Steps to reproduce** 1. Enable the Paperclip Runner experimental setting. 2. Start a native runner task. 3. Open the task view while the run emits reasoning. 4. Queue a message and select the steering action. **Paperclip version or commit** The regression reproduces on `24a674f8858060e77ea1beb50689d26473e91431`. **Additional context** Related closed work: Refs #12592. ## What Changed - Restored the queued-comment steering route for active native sessions. - Added durable and queue-bound steering acknowledgements for safe retries. - Restored runner instruction bundle support. - Added live-log fallback when native events are empty or unavailable. - Restored the compact live reasoning ticker in the task view. - Added a visible temporary-unavailable state when both activity sources fail. - Kept the unified Paperclip Runner experimental gate unchanged. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm check:token-gates` - `pnpm build` - Seven focused test files passed with 140 tests. - The final steering regression file passed with 12 tests. - The broad local test run reached unrelated workspace, port, and shared database failures. The changed-area tests remained green. ## Risks - The steering route changes queue and run records in one transaction. Tests cover stale targets, unavailable sessions, lost responses, and wrong-queue acknowledgements. - Native events remain the primary transcript source. The live log is used only when event data is absent or its poll fails. - The experimental gate still hides and rejects the runner when the setting is off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — no documentation change is required for this regression repair - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b4f302d040 |
feat(grok-local): copy a refreshed sandbox credential back to the host (#12696)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agents through provider-specific adapters in local and remote environments > - A remote Grok run can refresh its credential inside its sandbox > - The host copy can become stale when teardown discards that refreshed credential > - This pull request copies the refreshed credential back through a locked, fail-closed teardown path > - The benefit is that later Grok runs can use the refreshed host credential without another login ## Linked Issues or Issue Description Refs: #12618 **Agent or provider** Grok local adapter. **Why this adapter is useful** A remote Grok run can refresh its access token during a run. Copying the refreshed credential back to the host keeps later runs ready to use. **How the agent is invoked** Paperclip invokes the Grok local adapter through its remote subscription run path. The adapter stages the company Grok home as a sandbox asset. The change adds a copy-out step on the teardown path. ## What Changed - `grok-auth-merge-decision.cjs` adds a host predicate in its own process. It compares the whole `<issuer>::<uuid>` identity key of the two files. It reads `expires_at` as an ISO-8601 string, an epoch-seconds number, or an epoch-milliseconds number. It exits 10 to use the source, 20 to keep the destination, 21 when the expiry shape is unreadable, and 22 when the source expiry sits more than 400 days after the host clock. It fails closed in every unclear case: an unusable side, a different identity, an absent expiry, a tie, an unreadable expiry, and an implausible expiry all keep the destination. - `grok-auth-merge-decision.ts` adds a wrapper that runs the predicate and maps the exit code to a typed result. - `grok-auth-copyback.ts` adds `copyBackGrokAuth({ hostHomeDir, readSandboxAuth, log, env })`. It locks on `hostHomeDir` with `withDirectoryMergeLock`, stages the sandbox bytes into a private `0600` temporary file, runs the predicate, and installs the file with an atomic rename in the same directory. It keeps no backup of the displaced credential. It leaves no temporary file on the success path, the keep path, or an error path. On an error it logs the `errno` code only, then re-throws. - `execute.ts` adds a `restore` callback to the Grok `home` asset. The callback takes the destination from `resolveManagedGrokHomeDir(process.env, agent.companyId)`, never from `env.GROK_HOME`. A copy-out failure does not fail the run. - `package.json` updates the `build` script to copy `grok-auth-merge-decision.cjs` into `dist/server/`, because `tsc` does not copy a `.cjs` file. **The credential shape this predicate reads** A redacted sample of a real vendor credential answered four structural questions. The answers hold no credential bytes, no account identifier, no file path, and no timestamp value. 1. `expires_at` is present. 2. `expires_at` sits inside the value object, under the `<issuer>::<uuid>` key. It is not a top-level field. 3. `expires_at` is an ISO-8601 string. It carries UTC time with a trailing `Z` and six fractional-second digits. 4. A normal run rewrites `auth.json`. The value object carries a `refresh_token` next to `expires_at`, so the client refreshes the access token and rewrites the file. ## Verification - [x] `pnpm vitest run packages/adapters/grok-local` — 114 tests in 12 files pass. - [x] `pnpm --filter @paperclipai/adapter-grok-local typecheck` — clean. - [x] `pnpm --filter @paperclipai/adapter-grok-local build` — succeeds, and `dist/server/grok-auth-merge-decision.cjs` exists after the build. - [x] Continuous integration is green on every check. ## Risks The predicate keeps the host credential when identity, expiry, file access, or freshness data is unclear. The copy-out path can log an error and leave the run successful when it cannot install the refreshed credential. The atomic rename and directory lock protect the host file from partial writes and concurrent copy-out actions. ## Model Used OpenAI GPT-5, current deployment. The exact runtime version and context window are not exposed to this agent. The model used tool calls and code inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f9f850c20 |
fix: limit plan-to-auto transition to plan confirmation (#12695)
<!-- This pull request uses ASD-STE100 Simplified Technical English. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue thread controls plan review and agent work modes. > - A user can accept a full plan or confirm a smaller checkbox action. > - Only full plan acceptance must start automatic agent work. > - The current transition did not check the interaction kind. > - This pull request limits the transition to an accepted plan confirmation. > - The benefit is a safe and clear start of agent work after plan approval. ## Linked Issues or Issue Description **What happened?** An accepted confirmation that targeted a plan could change an issue from planning mode to standard mode. This included a checkbox confirmation. A checkbox action is not approval of the full plan. **Expected behavior** Only acceptance of a current full-plan confirmation starts automatic agent work. Other interaction kinds and rejected confirmations keep the current work mode. **Steps to reproduce** 1. Put an issue in planning mode. 2. Create a checkbox confirmation that targets the current plan revision. 3. Accept the checkbox confirmation. 4. Observe that the issue enters standard mode before this fix. **Paperclip version or commit** The problem was present on `master` before this change. **Deployment mode** The problem is in the core server logic and is not deployment-specific. ## What Changed - Require a full `request_confirmation` interaction before plan acceptance starts automatic work. - Add service tests for acceptance, rejection, stale interaction kinds, and unchanged standard-mode behavior. - Check the route activity log for the planning-to-standard mode change. - Document the plan acceptance transition in the V1 contract. ## Verification - `pnpm exec vitest run server/src/__tests__/issue-thread-interactions-service.test.ts server/src/__tests__/issue-thread-interaction-routes.test.ts` passes 140 tests. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` was also started. Unrelated workspace-runtime tests failed because fixed local runtime ports were occupied or offset on the shared host. The same failures reproduce alone. The changed test files pass alone. ## Risks - Risk is low. The change adds one interaction-kind guard to the existing transition. - A full accepted plan confirmation still changes planning mode to standard mode and an eligible review issue to todo in one transaction. - Checkbox confirmations, questions, rejection, and standard-mode issues keep their previous behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5, reasoning, tool use, and code execution. The runtime does not expose the exact model suffix or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b6de5327e |
Remove cheap model profiles (#12683)
## Thinking Path > - Paperclip manages agents that use different model providers and adapters. > - Paperclip must keep agent execution rules clear and predictable. > - The cheap-model profile added a second execution mode across adapters, task recovery, APIs, and the UI. > - That mode increased configuration and recovery complexity. > - This pull request removes the cheap-model profile as a product feature. > - The benefit is one model-selection path for normal work and recovery work. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change simplifies model selection across agent configuration, task execution, recovery, and adapter capabilities. **Current behavior** Paperclip exposes cheap-model profiles in adapter metadata, agent runtime configuration, task overrides, recovery rules, APIs, and the board UI. Recovery work can select a different model profile from the agent's configured model. **Proposed behavior** Paperclip uses the agent's configured model for normal work and recovery work. Status-only recovery stays limited to coordination work. The API rejects legacy model-profile configuration. A migration removes stored model-profile values from existing agent, issue, and historical revision records. **Reason and benefit** One model path reduces configuration, API, UI, and recovery complexity. It also prevents status recovery from becoming a separate product-level model-routing feature. **Breaking changes** This change removes model-profile fields and adapter capability metadata. Existing stored model-profile values are removed by an idempotent migration. The validators reject new legacy profile values with clear errors. ## What Changed - Removed model-profile types, adapter capabilities, API fields, and model selection logic. - Removed cheap-model controls from agent and task UI surfaces. - Kept status-only recovery limited to coordination context while normal continuations use the configured agent model. - Added an idempotent migration that removes stored model-profile values from agents, issues, and configuration revisions without changing issue update timestamps. - Updated tests and product documentation for the single-model behavior. ## Verification - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` completed with 5,607 passing tests and 8 environment-sensitive failures in unrelated fixed-port and database-deadlock suites. The same failures repeated in an isolated rerun. CI is the final clean-room result. ## Risks - This is an intentional breaking change for clients that send model-profile fields. - The migration changes legacy agent, issue, and configuration-revision JSON. It is idempotent and preserves unrelated fields and issue update timestamps. - The change is cross-cutting because the removed feature existed in adapters, shared contracts, the server, plugins, and the UI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. Reasoning and tool use were enabled. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1ab159d3a7 |
feat(apps): consolidate connector management (#12684)
Completes the post-managed-OAuth connector lifecycle, Paperclip Cloud provisioning defaults, governed test flows, and consolidated Apps UI.\n\nCo-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
141f202e40 |
Clean up experimental settings features (#12681)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - Instance settings control optional product features and developer tools. > - The experimental settings page mixed active experiments, internal tools, and old recovery controls. > - Some workspace links also used the selected company instead of the workspace owner. > - These problems made settings hard to scan and could send users to the wrong company route. > - This pull request removes old controls, groups developer settings, and resolves workspace links from workspace data. > - The benefit is a smaller settings surface and correct workspace navigation. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the instance experimental settings page, task watchdog controls, dependency wake recovery, and execution workspace routes. **Current behavior** The settings page shows old recovery controls and mixes product experiments with internal developer settings. Task watchdogs require an extra feature flag. Some direct workspace links use the current company prefix instead of the company that owns the workspace. **Proposed behavior** Remove the old task recovery experiment and its unused API surface. Make task watchdog controls available without the removed flag. Put worktree execution and managed environment controls in the developer section. Resolve direct workspace links from the workspace owner and reject a company prefix that does not own the workspace. **Reason and benefit** The smaller settings page is easier to understand. The server keeps only the dependency wake backstop that it still uses. Workspace links open under the correct company route. **Breaking changes** This removes the experimental issue graph recovery preview and run endpoints. It also removes the task watchdog feature flag. Task watchdog data and dependency wake behavior remain available. ## What Changed - Removed the old task watchdog and issue graph recovery feature flags. - Removed the old issue graph recovery preview, run controls, API contracts, and unused recovery implementation. - Kept resolved dependency wakes as the scheduler backstop. - Grouped product experiments and Paperclip developer settings on the instance settings page. - Made task watchdog controls available without an extra experimental flag. - Added owner-aware redirects and company checks for execution workspace routes. - Hid the false stopped-state badge while a workspace has no active runtime state. - Updated focused server and UI tests for the new behavior. ## Verification - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run` completed with 5,620 passing tests and four failures in unchanged workspace runtime port tests. The same four failures repeat when the two files run alone. - The complete GitHub CI matrix passed, including all server, serialized server, build, canary, and end-to-end jobs. ## Risks - Clients that call the removed experimental recovery endpoints must stop calling them. - The route checks depend on workspace detail access. An unknown or cross-company workspace returns the global not-found page. - There are no database migrations, lockfile changes, workflow changes, or design image changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The exact deployment ID and context window are not exposed. Reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed3559dd21 |
feat(server): split the Sentry DSN into front-end and backend variables (#12678)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip reports server and browser errors through optional Sentry monitoring > - One environment variable sends both error types to one Sentry project > - Operators need separate control for browser and server error data > - This pull request adds specific variables and keeps the existing variable as a fallback > - The benefit is separate monitoring without breaking current deployments ## Linked Issues or Issue Description **What existing behavior does this improve?** The Sentry configuration for server and browser monitoring uses one environment variable. **Subsystem affected** Cross-cutting (multiple of the above) **Current behavior** `SENTRY_DSN` supplies the server and browser clients. Both clients therefore report to the same Sentry project. **Proposed behavior** `SENTRY_DSN_FRONTEND` supplies the browser client. `SENTRY_DSN_BACKEND` supplies the server process. `SENTRY_DSN` remains a fallback for either component. **Reason and benefit** Operators can send browser and server errors to separate Sentry projects. Operators can also activate only one component. **Breaking changes** None. Existing deployments can continue to use `SENTRY_DSN`. ## What Changed - Add `resolveSentryDsns(env)` and use it in the server and browser configuration paths. - Add precedence, empty-string, fallback, and route tests. - Update the README, observability guide, and stale code comments. - Log one warning when the server uses the legacy fallback without exposing a DSN value. ## Verification - `pnpm vitest run --project server sentry-dsn` — 8 tests pass. - `pnpm vitest run --project server auth-routes` — 21 tests pass. - The earlier run of the three targeted suites passed 40 tests. - `tsc --noEmit` passes for the files in this diff. - All required GitHub Actions checks pass, including the full continuous-integration suite. ## Risks The main risk is an incorrect environment variable precedence rule. Unit tests cover specific values, empty strings, and legacy fallback behavior. The existing `SENTRY_DSN` path remains compatible. ## Model Used OpenAI Codex — GPT-5, current runtime, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
86ebdf842e |
fix(runner): keep agents running when app connections expire (#12670)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can receive governed access to connected apps through the runtime MCP gateway. > - A connected app can become unavailable when its sign-in expires or its health state needs attention. > - The native runner treated that optional app state as a fatal runtime setup error. > - One unavailable app could therefore stop all unrelated agent work. > - This pull request removes the fatal dependency and keeps the available app assignment immutable. > - The benefit is that an agent can continue its work while the stream tells the user which app needs reconnection. ## Linked Issues or Issue Description **What happened?** An agent could not start a native run when one assigned app connection was disabled, degraded, failed, or missing its secret. Runtime context creation or MCP delivery threw an error before the agent could do unrelated work. **Expected behavior** The run must continue without the unavailable app. Healthy assigned apps must remain available. The stream must explain which app needs reconnection. A changed assignment must not give a native run new access after its immutable context is captured. **Steps to reproduce** 1. Assign an MCP app connection to a Paperclip Runner agent. 2. Set the connection to a state that needs attention, such as `degraded`. 3. Start a task run for that agent. 4. Observe that native runtime setup fails before the agent starts. **Paperclip version or commit** Reproduced from `ee2a19062`. The branch is rebased on `dda4dff64`. **Deployment mode** Local development from source with embedded Postgres. No matching public issue or open pull request was found in the GitHub search. ## What Changed - Filter unavailable assigned app connections from the immutable native runtime MCP snapshot. - Keep healthy assigned connections and their tools in the snapshot. - Replace the fatal native MCP availability check with an optional stream warning callback. - Withhold MCP delivery when the current assignment digest does not match the captured native context. - Prevent a warning delivery failure from stopping the agent run. - Add regression tests for disabled, degraded, mixed healthy and unavailable, and assignment-drift cases. ## Verification - `pnpm exec vitest run server/src/services/native-runtime/runtime-context.test.ts server/src/__tests__/heartbeat-runtime-mcp-servers.test.ts` passes with 8 tests. - `pnpm -r typecheck` passes. - `pnpm check:token-gates` passes. - `pnpm build` passes. - `pnpm test:run` was attempted. Unrelated workspace runtime and port-exposure tests failed on this macOS host. The same files also failed when run without the changed MCP tests. The changed MCP tests remained green. Clean GitHub CI is the final full-suite check. ## Risks - Low migration risk. This change has no schema or API contract migration. - An unavailable app is absent from the run MCP surface until it is reconnected and a later run captures it again. - Assignment drift fails closed. The agent keeps running, but the changed gateway is not delivered. - This pull request does not auto-block the issue before the agent decides that the app is required. It emits reconnect guidance in the stream. The existing connection-request interaction remains the path for a required app. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5.6-sol`, with high reasoning, repository tools, code execution, and browser automation. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
14c7efa068 |
fix(workspaces): enable UI hot reload by default (#12612)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed worktrees can run a Paperclip development server for each task > - The managed runtime used the built UI when its service did not set the UI development middleware option > - This made new UI source changes require a manual build instead of a hot reload > - The runtime must supply the development default while it must keep an explicit operator choice > - This pull request enables the UI development middleware for new managed Paperclip development services > - The benefit is that UI edits appear in the managed worktree browser without a manual build ## Linked Issues or Issue Description **What happened?** A new managed Paperclip development worktree served the built UI by default. An operator had to set `PAPERCLIP_UI_DEV_MIDDLEWARE=true` before UI source changes could hot reload. **Expected behavior** New managed Paperclip development worktrees must enable the UI development middleware by default. An explicit `PAPERCLIP_UI_DEV_MIDDLEWARE=false` value must continue to disable it. **Steps to reproduce** 1. Start a managed Paperclip development service without `PAPERCLIP_UI_DEV_MIDDLEWARE`. 2. Open its UI. 3. Change a UI source file. 4. Observe that the browser does not receive the change until the UI is built again. **Paperclip version or commit** This was reproduced on `317394456` from `master`. **Deployment mode** Local development with a managed worktree runtime. ## What Changed - Set `PAPERCLIP_UI_DEV_MIDDLEWARE=true` for managed `paperclip-dev` services when the service does not set a value. - Keep explicit service values, including `false`. - Add a regression test and document the default and the opt-out. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/workspace-runtime.test.ts -t "enables UI dev middleware by default"` - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run` completed with 5,397 passing tests. Four existing runtime-port tests could not use ports `42000` and `52000` because a live managed runtime owns those ports on this host. The new regression test passed separately. ## Risks - Risk is low. The change applies only to managed services named `paperclip-dev`. - A service can keep the built UI by setting `PAPERCLIP_UI_DEV_MIDDLEWARE=false`. - There is no database or API contract change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, hosted Codex context window, high reasoning, tool use, code execution, and multi-file repository editing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ee2a190626 |
Unify Paperclip Runner experimental controls (#12666)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner is an experimental execution adapter. > - The adapter and its required sandbox ingress had separate settings. > - A user could enable one setting and still have an unusable runner configuration. > - The runtime already makes one durable native or legacy decision for each run. > - This pull request uses that runtime decision for ingress authorization. > - The benefit is one clear opt-in with safe recovery for existing native runs. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the experimental settings and transport authorization for Paperclip Runner. **Subsystem affected** Cross-cutting. This change affects the React settings UI, shared settings contracts, adapter utilities, and server runtime selection. **Current behavior** Settings shows separate Paperclip Runner and Runner Preview Ingress controls. A user can enable the runner but leave required sandbox ingress disabled. **Proposed behavior** Settings shows only Paperclip Runner. Its native runtime decision also authorizes provider WebSocket ingress when the execution target requires it. A persisted native run keeps its recovery transport after the setting is disabled. **Reason and benefit** Paperclip Runner is one experimental capability. One opt-in removes an invalid partial configuration and makes the rollout boundary easier to understand. **Breaking changes** The Runner Preview Ingress card is removed. The old `enableRunnerPreviewIngress` key remains accepted in stored settings and managed configuration, but it has no server runtime effect. The public adapter-utils input remains compatible through a deprecated alias. **Additional context** Refs: #12638, #12641, #12656. ## What Changed - Removed the separate Runner Preview Ingress card from Experimental Settings. - Made resolved native runtime selection authorize required provider ingress. - Preserved ingress recovery for persisted native runs after the rollout flag is disabled. - Kept the old settings key and adapter-utils input as deprecated compatibility contracts. - Added focused UI, runtime policy, transport, stored-settings, and managed-config regression tests. - Updated deployment documentation and feature descriptions. ## Verification - GitHub Actions will run typecheck, tests, build, policy, and browser shards. - Focused tests cover the single settings control, runtime authorization, fail-closed transport selection, the deprecated public input, and old managed configuration. - No local tests were run, per the maintainer request to use GitHub Actions for verification. - `git diff --check` passes. ## Risks Low to moderate risk. The effective ingress gate changes from a separate stored flag to the resolved native run decision. Fresh runs still require `enableNativeRunner`. Persisted native runs remain recoverable. Legacy adapters never receive ingress authorization. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1955b0e2d8 |
Gate Paperclip Runner setup behind an experimental flag (#12656)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters control how Paperclip starts and resumes an agent runtime. > - Paperclip Runner is an experimental Rust runtime and must stay opt-in. > - The server already rejected new runner selections when the flag was off. > - Some setup and onboarding views did not enforce the same boundary. > - This pull request exposes the existing flag and applies it to every new setup path. > - The benefit is a safe rollout with unchanged legacy onboarding and recoverable existing native runs. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves experimental adapter selection in Settings, onboarding, new-agent setup, invite setup, and company import. **Subsystem affected** Cross-cutting: the React UI and the server onboarding seed service. **Current behavior** The server defaulted Paperclip Runner to off, but Settings did not expose the flag. First-run onboarding could show the runner after opt-in. A direct new-agent URL and some setup pickers could also reveal native runner configuration before the availability check completed. **Proposed behavior** Settings has a default-off Paperclip Runner toggle. Explicit agent configuration shows the runner only after the server reports that the flag is enabled. First-run and invite onboarding always use legacy adapters. Existing native agents and runs remain readable and recoverable. **Reason and benefit** This keeps the experimental runtime out of normal onboarding. It also gives administrators one clear opt-in before users can create a native runner agent. **Breaking changes** None. Legacy adapter selection and execution stay unchanged. Existing native records remain available. ## What Changed - Added the Paperclip Runner opt-in to Experimental Settings. - Refreshed adapter availability after the setting changes. - Kept UI and server-seeded onboarding on legacy adapters. - Made native runner choices fail closed in new-agent, invite, and import setup. - Preserved edit and recovery behavior for existing native agents and runs. - Added focused regression tests for flag-off and flag-on behavior. ## Verification - GitHub Actions will run the repository test, typecheck, build, and policy gates. - Focused tests cover Settings, onboarding, agent creation, invite setup, import setup, and server-seeded onboarding. - No local test suite was run, per the maintainer request to use GitHub Actions for verification. - `git diff --check` passes. ## Risks Low risk. The change narrows new adapter selection only. The server remains the final enforcement point. Existing native records do not depend on the current flag value for read or recovery behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1ed29abaa6 |
fix(runner): harden dormant provider boundaries (#12654)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner currently enables only the Codex production path. > - The package also contains dormant OpenCode and ACPX provider boundaries. > - Dormant boundaries must still fail safe before later activation work. > - Provider children must not inherit unrelated server secrets or host homes. > - Permission defaults must require interaction instead of broad automatic approval. > - This pull request hardens those boundaries without activating them. > - The benefit is a safer base for later provider-specific runnerd work. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the inactive OpenCode and ACPX provider boundary in Paperclip Runner. **Subsystem affected** The adapter permission contract, Runner provider environment, and native execution input builder. **Current behavior** Dormant OpenCode code can inherit the full server environment. Its default permission mode allows operations. ACPX also defaults to broad approval. The provider guard can accept inherited object property names. **Proposed behavior** Use exact provider identifiers. Use interactive defaults. Allow only required OpenCode environment keys. Reject invalid proxy permission modes. **Reason and benefit** This reduces accidental authority and secret exposure before future provider activation. **Breaking changes** No production provider is activated. Codex runtime selection and Codex credential-home discovery do not change. Dormant OpenCode and ACPX callers that omit permission modes now receive safer defaults. ## What Changed - Change dormant OpenCode and ACPX permission defaults to interactive modes. - Reject prototype property names as provider identifiers. - Default dormant ACPX input to the qualified Codex agent profile. - Add an explicit OpenCode runner environment allowlist. - Exclude host homes, server credentials, database values, and Node injection options. - Add a fail-closed OpenCode proxy permission parser. - Add focused tests for defaults, filtering, and invalid values. ## Verification GitHub Actions must run: - Adapter utility tests. - Paperclip Runner tests, type checks, and build. - Server native runtime tests. - Repository test, type-check, build, policy, and security gates. No local test command was run. The repository owner requested GitHub-only verification. ## Risks Future OpenCode credential providers must add required variables to the allowlist through review. The safer defaults can pause dormant internal scenarios that relied on implicit broad approval. Production Codex behavior is unchanged. ## Model Used OpenAI Codex with the GPT-5 agent model. The work used high reasoning, repository inspection, tool use, and parallel security review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
131f5c4065 |
feat(runner): add administration and observability (#12641)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Administrators need bounded controls for experimental native execution. > - The lower stack adds remote Codex execution and the task workspace. > - Operators need to configure Codex safely and inspect provider traces. > - Unsupported providers must not appear as runnable choices. > - This pull request adds Codex-only administration and observability. > - The benefit is a default-off operational surface for production diagnosis. ## Linked Issues or Issue Description Refs #12640. Refs #12616. Refs #12352. **Subsystem affected** Agent configuration, instance experimental settings, run ledger, provider trace inspector, and administrator actions. **Problem or motivation** The native runner lacks one safe operator surface for Codex permissions, lifecycle, raw trace capture, and run inspection. The integration branch also contains provider choices that the production backend cannot execute yet. **Proposed solution** Expose only the qualified Codex controls. Keep Paperclip Developer Mode and runner preview ingress off by default. Gate raw trace actions by administrator access and existing trace authorization. **Alternatives considered** Exposing unfinished providers would create configurations that fail at runtime. Always-on tracing would increase sensitive data and storage risk. **Roadmap alignment** This work supports governed Cloud and Sandbox agents and production diagnostics. ## Stack - Base PR: #12640. - Lower PRs: #12639 and #12638. - This PR contains only its 54-file administration and observability delta. - This is the final feature PR in the Codex production stack. ## What Changed - Added Codex-only Paperclip Runner permission and lifecycle controls. - Added bounded warm idle configuration. - Kept the provider field fixed to Codex. - Added administrator-only one-run raw trace requests. - Added a persistent future-run raw trace toggle. - Added trace status, metadata, ledger, and canonical runner inspection. - Added JSON-RPC request-origin grouping and finalization lineage. - Restored the stateful PRP transcript parser and focused projection tests required by trace inspection. - Added default-off Paperclip Developer Mode. - Added Honeycomb run links for authorized developer mode. - Disabled the legacy operational skill for `paperclip_runner`. - Did not expose OpenCode, ACPX, Pi, Claude Managed, or AWS runner choices. - Did not change migrations, workflows, dependencies, or `pnpm-lock.yaml`. ## Verification - GitHub Actions will run UI tests, server tests, repository typecheck, build, browser tests, security, and policy gates. - Tests cover Codex configuration defaults and bounds, administrator trace actions, persistent settings, ledger inspection, trace lineage, and Honeycomb links. - Existing server trace authorization and retention tests remain the backend authority. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check runner/task-workspace-experience...HEAD` passes. - The delta contains 54 files. ## Risks - Raw provider traces can contain sensitive provider data. - Existing server authorization controls access, reveal, download, retention, and deletion. - The UI gates trace actions by administrator access and developer mode. - All new instance settings remain off by default. - Fresh Paperclip Runner configuration remains Codex-only. - Direct adapters and legacy task behavior do not change in this PR. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0a422fda52 |
feat(runner): add remote execution substrate (#12638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner gives native runs a durable and governed execution path. > - The current native path runs on the control-plane host. > - Remote environments need an authenticated execution-target contract. > - The contract must not change direct adapters or enable new runtimes by default. > - This pull request adds the remote execution substrate and Daytona ingress. > - The benefit is a bounded base for later remote runner transport work. ## Linked Issues or Issue Description Refs #12616. Refs #12352. **Subsystem affected** Cross-cutting. This change touches runner transport, server orchestration, plugin contracts, and shared settings. **Problem or motivation** Native execution cannot resolve an authenticated runner ingress through a remote environment. The server also lacks one provider-neutral contract for remote execution targets. **Proposed solution** Add a default-off runner preview ingress capability. Add transport-neutral runner connectivity. Add remote execution target and lifecycle handling. Add a Daytona ingress implementation with redacted credentials. **Alternatives considered** A provider-specific server path would duplicate orchestration and authorization. A public endpoint without an environment contract would weaken the trust boundary. **Roadmap alignment** This work supports the Cloud and Sandbox agents milestone. It also supports self-healing runs and governed tool access. ## What Changed - Added execution-target traits for local, SSH, and sandbox environments. - Added plugin RPC contracts for runner ingress endpoints. - Added authenticated Daytona preview ingress. - Added transport-neutral PRP outbound connections. - Added remote runner artifact verification and fail-closed provider selection. - Added bounded native session resume, cancellation, and lifecycle recovery. - Preserved Codex-only selection for fresh experimental runner starts. - Preserved all direct adapter execution and finalization paths. - Removed stale Pi provider-pack requirements that security review rejected. - Kept the rollout controls off by default. - Did not change pnpm-lock.yaml, Cargo, database migrations, or GitHub workflows. ## Verification - GitHub Actions will run the repository test, typecheck, build, security, and policy gates. - Focused tests cover ingress validation, redaction, execution targets, remote lifecycle, cancellation, resume, and legacy adapter selection. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check origin/master...HEAD` passes. - The diff contains 52 files. ## Risks - Remote execution crosses a trust boundary. - The implementation validates target capabilities, artifact digests, provider-pack pins, and connection metadata. - The feature remains default-off. - Fresh native selection remains Codex-only. - Existing direct adapters remain on the legacy path. - This PR does not yet make remote Codex runnable. The next PR adds the Rust WSS and TLS transport. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
51ad751e0b |
feat(runner): integrate Codex native execution (#12616)
## Thinking Path > - Paperclip is the open source control plane for teams of AI agents. > - Agent runs currently use direct adapters and their established finalization paths. > - The new runner package needs one production integration before it can execute a real provider through the server. > - That integration must not change direct adapters or expose unsupported providers. > - The rollout must also preserve native runs that were already recorded when the feature flag changes. > - This pull request adds a default-off, Codex-only native execution path and its authority boundary. > - The benefit is a recoverable production vertical slice with explicit compatibility guards. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting server orchestration and adapter selection. **Problem or motivation** The runner package exists, but the server cannot yet start and recover a governed Codex run through it. A careless integration could also route existing direct adapters into the native runtime or lose cancellation and finalization state. **Proposed solution** Add a hidden `paperclip_runner` adapter for Codex. Keep it behind the default-off instance flag. Bind native execution, resume, cancellation, semantic tool authority, and finalization to the recorded company, issue, run, and coordinator identities. Leave every direct adapter on its existing path. **Alternatives considered** A multi-provider launch was rejected because only Codex has the complete production bridge in this series. Replacing direct adapter execution was rejected because the runner remains experimental. **Roadmap alignment** This work supports governed tool access, action attribution, and self-healing runs. It keeps the integration narrow and default-off. ## What Changed - Add the Codex-only native session executor and persisted resumption path. - Add run-scoped semantic tool projection, authorization, receipts, and idempotency. - Add audited native cancellation with durable issue and coordinator binding. - Add result fencing so a recorded result cannot reacquire the provider and run twice. - Reject fresh runner starts when the rollout flag is off while preserving recorded native recovery. - Keep direct adapters outside native status, cancellation, record creation, and finalization. - Add focused conformance, recovery, cancellation, status, portability, and compatibility coverage. ## Verification - GitHub Actions is the authoritative test environment for this large stack. - The PR policy and lightweight stack checks run while this is a middle PR. - The full required suite runs when this PR becomes the lowest unmerged or top PR. - Greptile will review this exact delta after the branch is pushed. ## Risks - The main risk is routing a legacy adapter into native execution. Runtime selection and heartbeat tests cover that boundary. - The next risk is stale or cross-company cancellation. Durable binding checks and transactional audit persistence cover it. - The adapter remains hidden and default-off. Only Codex is admitted. - There are no database migration, lockfile, or GitHub workflow changes in this PR. ## Stack 1. [Runner package, SDK, and developer tools](https://github.com/paperclipai/paperclip/pull/12608) 2. This PR: Codex production server integration 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617) ## Model Used OpenAI Codex with GPT-5, extended reasoning, repository tools, and parallel review agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
560e7e48b5 |
feat(runner): add SDK and developer tooling (#12608)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0b73ebb86c |
Fix managed OAuth catalog activation (#12623)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Apps system gives agents governed access to external services > - The managed OAuth callback discovers provider tools before it activates a grant > - Fresh managed connections kept every discovered tool in quarantine > - The Apps page also counted disabled tools as available actions > - This pull request makes setup activation atomic and keeps later catalog changes quarantined > - The benefit is a usable catalog after consent without weakening reauthorization safeguards ## Linked Issues or Issue Description N/A — no public GitHub issue exists for this follow-up. Related merged work: [#12619](https://github.com/paperclipai/paperclip/pull/12619) and [paperclip-cloud #319](https://github.com/paperclipai/paperclip-cloud/pull/319). **What happened?** A fresh Paperclip-managed OAuth connection discovered the correct Google Workspace tools, but it left every allowed tool in quarantine. The Apps page then reported zero actions for write profiles and counted disabled actions for read profiles. **Expected behavior** A fresh or revived managed connection must remain disabled until Paperclip stores credentials, discovers the catalog, reviews the profile allowlist, installs default policies, and activates the connection. A later reauthorization must preserve user choices. A later catalog change must quarantine new or changed tools. **Steps to reproduce** 1. Connect a managed Google Workspace write profile. 2. Complete provider consent and return through the instance callback. 3. Open the connection in Apps. 4. Observe that the connection is active but the allowed actions remain quarantined. **Paperclip version or commit** Commit `c7ebc089c` from merged pull request #12619. **Deployment mode** Self-hosted local development through Tailscale HTTPS. The same callback logic applies to Cloud-hosted instances. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Not adapter-specific. This change affects the core Apps and tool-access paths. ## What Changed - Kept fresh and revived managed connections in the draft state until catalog finalization succeeds. - Added a managed-draft refresh option that quarantines discovery results without changing generic draft behavior. - Activated reviewed profile tools, created bindings, and installed ask-first policies in the existing finalization transaction. - Preserved custom profiles, bindings, archived state, and policies during ordinary reauthorization. - Kept new or changed tools quarantined after activation and kept out-of-profile tools disabled. - Counted only active catalog entries as available actions in the Apps page. - Added retry, revival, reauthorization, policy, profile, and UI regression coverage. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` — passed, 208 tests. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/apps/AppDetail.test.tsx` — passed, 52 tests. - Server and UI typechecks passed. - `pnpm build` — passed on the final tree. - `pnpm check:token-gates` — passed. - `git diff --check` — passed. - Browser walkthrough — passed for all 16 enabled Google Workspace profiles through the staging Cloud broker and a self-hosted Tailscale HTTPS instance. Every final connection became active and exposed at least one allowed action. - `pnpm test:run` — the changed suites passed. The shared live-QA environment caused unrelated workspace-runtime concurrency and cleanup failures, so hosted CI is the clean-environment authority for the full suite. ## Risks - The change affects managed OAuth only. Customer-owned OAuth setup keeps its current behavior. - A failed initial finalization now leaves a safe draft that the callback can retry. - An ordinary active reauthorization does not rebuild defaults, so existing user policy remains intact. - There are no database migrations and no public API changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.6, with reasoning, repository tools, browser control, code execution, test execution, and parallel subagent review. The effective context window was managed by the Codex task runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and the changed suites pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c7ebc089cb |
fix(apps): complete managed Google Workspace rollout (#12619)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Apps system gives agents governed access to external services > - The managed Google Workspace connector uses separate profiles for each app and access level > - Several app definitions and callback paths did not enforce the same profile contract > - The default method could also select a customer OAuth setup when a managed read profile was available > - This pull request aligns the profile contracts, setup guidance, default selection, and activity attribution > - The benefit is a consistent managed connection flow for every Google Workspace app ## Linked Issues or Issue Description N/A — no public GitHub issue exists for this follow-up. Related merged work: [#12600](https://github.com/paperclipai/paperclip/pull/12600) and [#12609](https://github.com/paperclipai/paperclip/pull/12609). This change follows the merged Paperclip Cloud managed OAuth broker work. It does not add a new broker or provider client. **What happened?** The Google Workspace connection definitions could drift from the shared connector profile registry. The callback activity always named Gmail. Google Sheets did not show the Developer Preview requirement. The setup flow could select a customer-owned write method when Cloud advertised only a managed read profile. The tool-access service had no non-Gmail managed callback test. **Expected behavior** Each managed Google Workspace profile must use its exact app slug, MCP URL, scopes, ownership, risk tier, and write-tool policy. Callback activity must name the correct app and profile. Every Google Workspace card must show the same Developer Preview prerequisite. An available managed method must be the default within the selected capability. The customer-owned method must remain available as a fallback. **Steps to reproduce** 1. Advertise only the `gmail.read` managed profile. 2. Open the Gmail connection setup. 3. Observe that the customer-owned draft method becomes the default. 4. Complete a managed Google Drive callback. 5. Observe that the activity row names Gmail instead of Google Drive. 6. Open the Google Sheets setup. 7. Observe that it does not show the Google Developer Preview prerequisite. **Paperclip version or commit** Current `master` at the start of this follow-up. **Deployment mode** Local development. The same connector definitions apply to Cloud-hosted and self-hosted instances. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Not adapter-specific. This change affects the core Apps and tool-access paths. ## What Changed - Added one table-driven invariant for all 16 Google Workspace profiles. - Verified each profile against its app slug, MCP URL, exact scopes, capability, ownership, grant kind, risk tier, and write-tool allowlist. - Kept the Google Chat write profile least-privilege because its only enabled write tool is `send_message`. - Added the Google Developer Preview prerequisite to Google Sheets. - Preferred an available Paperclip-managed method before a customer-owned method. - Preserved explicit capability selection and the customer OAuth fallback. - Switched managed-profile availability from the anonymous global capability document to the enrolled instance's signed status response, so internal-pilot profiles cannot be enabled locally without an authorized instance binding. - Replaced the Gmail callback activity constant with the validated app slug and connector profile. - Added connector and route coverage for signed per-instance capabilities, including inactive and malformed responses. - Added a Google Drive callback test that covers the signed profile request, personal vault refs, encrypted secret rows, catalog filtering, and non-sensitive activity details. ## Verification - `pnpm -r typecheck` — passed across all workspaces before the signed-capability follow-up; final targeted shared and server typechecks also passed after it. - `pnpm exec vitest run server/src/services/paperclip-cloud-connector.test.ts` — passed, 9 tests. - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts` — passed, 19 tests. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts` — passed, 205 tests. - `pnpm test:run` — incomplete after the general server group reported five failures in `server/src/__tests__/workspace-runtime.test.ts`. The failures are outside the changed files. The run was stopped before the remaining serialized suites because the shared worktree was needed for a follow-up edit. - `pnpm build` — passed on the final tree. - `git diff --check` — passed on the final tree. The five full-suite failures were: - `records teardown and cleanup operations when a recorder is provided` - `does not accept an occupied allocated port when listener ownership is unavailable` - `backfills a pre-existing HTTP-only managed worktree runtime to verified HTTPS in place` - `re-adopts a live service whose shell command differs from the surviving process argv` - `reuses a registered legacy worktree that already has the branch checked out` ## Risks - The default setup method changes when at least one Paperclip-managed method is available. Explicit read, write, or draft choices still stay within the selected capability group. - Managed method availability now depends on Paperclip Cloud's signed enrolled-instance status. A Cloud outage or an inactive enrollment hides managed methods while leaving customer-owned OAuth available. - The callback activity schema gains a non-sensitive `profile` value. It does not include tokens, account identifiers, emails, tenant identifiers, or provider error text. - The profile invariant is strict. A future Google scope or tool change must update the shared registry and the matching app definition together. - There are no database migrations and no public API changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.6, with reasoning, repository tools, code execution, and test execution. The effective context window was managed by the Codex task runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3475e33fc7 |
fix(connector): isolate broker enrollment targets (#12609)
## Thinking Path > - Paperclip lets operators connect provider accounts for their agents. > - Self-hosted instances can enroll with the Paperclip Cloud OAuth broker. > - A local identity can remain on disk when an operator changes the broker from production to staging. > - The old code combined the new target with the old identity and produced a verification link that Cloud could not accept. > - The final enrollment callback also returned to an unscoped Apps path. > - This pull request binds each identity to one broker target and returns to the correct company route. > - The benefit is a fail-closed enrollment flow that works across Cloud targets and company-prefixed routes. ## Linked Issues or Issue Description Refs: #12600 **What happened?** A self-hosted instance with a saved connector identity could switch its broker base URL and environment. Paperclip then used the saved identity with the new target. The enrollment page received an unknown draft. A successful callback also opened an unscoped Apps path, which the UI treated as a company prefix. **Expected behavior** Paperclip must use one atomic identity and broker target. A target change must never mix old keys with a new broker. A completed enrollment must return to the initiating company's Connections page. **Steps to reproduce** 1. Start a self-hosted Paperclip instance and create a pending connector enrollment against the production broker. 2. Set the connector base URL and environment to staging. 3. Start enrollment again and open the returned verification URL. 4. Complete enrollment and inspect the final browser route. **Paperclip version or commit** `300a89ec1` **Deployment mode** Local dev (`pnpm dev`) through private HTTPS. ## What Changed - Resolve the connector broker and environment as one target. - Rotate a non-active identity when an administrator explicitly starts enrollment for a different target. - Reject active target changes and broker/environment mismatches. - Treat managed environment identity fields as one atomic tuple. - Require the Cloud verification URL to contain only the exact enrollment identifier. - Validate the configured target again before the instance redeems an enrollment callback. - Return successful enrollment callbacks to the company-prefixed Connections page. - Add regression tests for target isolation, managed identity precedence, URL validation, and the return path. ## Verification - `pnpm exec vitest run server/src/services/paperclip-cloud-connector-enrollment.test.ts server/src/services/paperclip-cloud-connector.test.ts server/src/routes/tool-access-connection-intent.test.ts` (31 tests passed) - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm build` - `git diff --check` - Completed a staging Cloud enrollment from a local Paperclip instance through private HTTPS. - Confirmed the instance reports an active staging enrollment for its exact HTTPS origin. - Confirmed the company-prefixed Connections route renders the enrollment success state. - `pnpm test:run` also reached five unrelated macOS harness failures. Two compare `/var` with `/private/var`. Three expect listener-fixture failures that do not occur on this host. The same five failures reproduce when the two workspace-runtime suites run alone. ## Risks - Low risk. The change affects only Paperclip Cloud connector identity selection and the enrollment return route. - An active identity now fails closed when an operator changes its broker target. The operator must restore the original target or perform a new enrollment flow. - Starting a new target replaces a non-active draft, so its previous one-time approval link no longer works. - This change does not alter provider tokens, grants, catalogs, or tool calls. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5.6 Sol. The work used high-reasoning mode, repository tools, code execution, browser automation, and parallel review subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and the scoped tests pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
300a89ec13 |
Detect the qualifier-less Claude usage-limit message in quota classification (#12475)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The heartbeat runtime classifies adapter run failures, and the
recovery service uses that classification to decide between automatic
retry, a timed provider-quota wait, and a board escalation
> - The Claude CLI changed its subscription-limit stop message to
"You've hit your limit · resets 2:30am (UTC)", and no quota matcher
knows this qualifier-less wording
> - A limit-hit run therefore classifies as `adapter_failed` (or
`claude_auth_required`), recovery burns its continuation retries against
a hard limit, and the issue blocks with the opaque "No live execution
path" notice instead of waiting for the reset and retrying automatically
> - This pull request teaches the adapter and the recovery service the
new wording, and titles stranded-escalation notices from the classified
run error code so operators see the cause at a glance
> - The benefit is that usage-limit stops self-heal at the provider
reset time, and the notices that do post say "Error: usage limit
reached" or "Error: not logged in to Claude" instead of a generic title
## Linked Issues or Issue Description
No public issue exists; the underlying problem follows the bug template:
**What happened?**
On a staging deployment, an assigned `in_progress` issue hit the Claude
subscription usage limit. The run recorded the error `Claude run failed:
subtype=success: You've hit your limit · resets 2:30am (UTC)`. The
automatic continuation retry failed the same way in 34 seconds with
`errorCode: adapter_failed`. Terminal-run recovery then escalated: the
issue moved to `blocked` with the notice "No live execution path" and a
board-owned recovery action. The notice gave the operator no indication
that the cause was a usage limit with a known reset time.
**Expected behavior**
A usage-limit stop classifies as `provider_quota` with the reset clock
parsed into `retryNotBefore`. The recovery service takes its
provider-quota wait path: a system-owned recovery action that waits for
the reset time and retries the original assignee automatically. If an
escalation notice does post, its title names the classified cause.
**Steps to reproduce**
1. Run a `claude_local` agent on an issue until the Claude subscription
limit is hit, so the CLI result is "You've hit your limit · resets
\<time\> (UTC)".
2. Let terminal-run recovery retry the continuation.
3. Observe the issue block with the "No live execution path" notice
instead of a timed quota wait. `classifyAdapterFailureForRecovery`
returns `null` for the recorded error text; `CLAUDE_PROVIDER_QUOTA_RE`
and `PROVIDER_QUOTA_ERROR_RE` both fail to match it.
## What Changed
- `CLAUDE_PROVIDER_QUOTA_RE` and `CLAUDE_EXTRA_USAGE_RESET_RE`
(claude-local adapter) accept "you've hit your limit" with no qualifier,
alongside the existing "session"/"usage" wordings, so the run classifies
as `provider_quota` and the reset clock lands in `retryNotBefore`.
- `PROVIDER_QUOTA_ERROR_RE` and `isProviderQuotaRecovery` (recovery
service) accept the same wording, so runs recorded before the adapter
fix (errorCode `adapter_failed` with the limit text in the error) also
route to the quota wait.
- `parseProviderQuotaClockReset` parses the "resets 2:30am (UTC)" clock
shape alongside the existing "try again at" shape.
- `buildStrandedRecoveryEscalationNotice` titles the notice from the
source run's classified error code when one is mapped: `provider_quota`
→ "Error: usage limit reached", `claude_auth_required` → "Error: not
logged in to Claude", `acpx_auth_required` → "Error: agent login
required". The raw failure text stays withheld from the issue thread;
only the server-classified code is surfaced. Unmapped codes keep the
existing seed/cause titles.
## Verification
- `pnpm vitest run
packages/adapters/claude-local/src/server/parse.test.ts
server/src/services/recovery/provider-failure-classification.test.ts
server/src/services/recovery/stranded-notice.test.ts` — 72 tests pass,
including 5 new cases that use the exact new CLI message.
- `pnpm vitest run server/src/__tests__/issue-recovery-actions.test.ts
server/src/__tests__/heartbeat-retry-scheduling.test.ts
server/src/__tests__/heartbeat-process-recovery.test.ts` — 189 tests
pass (no reroute regressions from the widened matchers).
- `tsc --noEmit` clean for `@paperclipai/adapter-claude-local` and
`@paperclipai/server`.
## Risks
- Low risk. The regex widenings are additive; every previously matched
wording still matches, and the existing negative test ("Workspace
storage capacity limit reached." stays unclassified) still passes.
- Behavioral shift, intended: an `adapter_failed` run whose error text
is the new limit wording now routes to the silent system-owned quota
wait instead of a board escalation. This matches how the older limit
wordings already behave.
- The notice title change only affects escalations whose source run
carries one of the three mapped error codes; all other notices render
exactly as before.
## Model Used
Claude Fable 5 (`claude-fable-5`, Claude Code CLI, extended thinking
with tool use).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
25cf079ec5 |
feat(runner): add Codex-native application integration (#12591)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package is useful only when the application can start, observe, and recover a native Codex run safely. > - Existing direct adapters must keep their current execution and finalization paths. > - The application boundary therefore needs additive persistence, authorization, coordination, and recovery behind an explicit experimental adapter. > - This pull request adds that Codex-only boundary without activating generalized providers, remote environments, or the later task/SDK surfaces. ## Linked Issues or Issue Description **Subsystem affected** Shared contracts, database persistence, adapter utilities, server native-runtime services, and the experimental Paperclip Runner adapter. **Problem or motivation** The already-landed runner package has a qualified Codex path, but the application needs durable native-run state, guarded runtime selection, authenticated coordination, tool security, finalization, and recovery before the experimental adapter can be exercised safely. **Proposed solution** Add a Codex-only `paperclip_runner` application path behind the existing default-off native-runner setting. Bind native state and coordination to company/run identity, preserve persisted-run recovery, and leave every direct adapter on its existing legacy execution path. **Alternatives considered** The earlier stack boundary introduced a generalized executor and remote-environment lifecycle here. That made this PR depend on implementations in higher PRs and changed reusable sandbox behavior globally. Those pieces are now deferred together to #12592. **Roadmap alignment** ROADMAP.md does not list a conflicting native-runner integration project. This change adds the application boundary for the existing Runner architecture. ## What Changed - Added native run/result/finalization/provider-trace persistence, shared validators, and idempotent migration/replay coverage. - Added guarded Codex-only runtime selection, authenticated PRP coordination, recovery, finalization, and interaction services. - Added run/company-bound tool-gateway authorization, credential redaction, SSRF protections, and replay-safe behavior. - Added the explicit `paperclip_runner` adapter behind the default-off rollout setting. - Preserved legacy answered-question wake projection and direct-adapter execution/finalization paths. - Hardened cancellation so only owned in-memory child processes are signaled; persisted recycled PIDs/process groups are never trusted. - Retained the narrow Claude ACPX isolated-context security follow-up discovered after #12590. - Deferred the generalized executor, provider ingress, remote lifecycle, SDK/lab/eval work, release-process changes, and lockfile. ## Verification - Changed-file delta against `master`: 133 files. - GitHub Actions is the authoritative verification environment for this PR. - Full CI, security, and Greptile review will run on this lowest unmerged stack PR. - Local tests/build/typecheck were not run because this checkout is resource constrained. - Static diff/reference checks pass, and `pnpm-lock.yaml` is unchanged. ## Risks - This touches central heartbeat and agent-route code, so legacy compatibility is the primary risk. - Runtime selection remains Codex-only and explicit; direct Codex, Claude, OpenCode, process, HTTP, and plugin adapters remain on their existing paths. - Fresh native starts fail closed while the rollout flag is off; persisted native records remain readable and recoverable. - Cancellation, company/run binding, tool calls, status decisions, and completion writes are guarded or replay-safe. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI and security gates are green - [ ] Greptile is 5/5 with no open actionable findings - [x] I will address all Greptile and reviewer comments before merge ## Stack - Position: 3 of 5 overall; lowest of 3 currently unmerged - Base: `master` - Previous: [#12590](https://github.com/paperclipai/paperclip/pull/12590), qualified Claude ACPX runtime — merged - Next: [#12592](https://github.com/paperclipai/paperclip/pull/12592), generalized Codex executor, task experience, and developer SDKs --------- Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
a7e6b818e9 |
feat(apps): add Paperclip Cloud managed OAuth connector (#12600)
## Thinking Path > - Paperclip lets operators give governed tools to AI agents. > - Connected Apps already support provider OAuth and personal connection grants. > - Some providers require one stable callback and do not support dynamic client registration. > - Self-hosted Paperclip instances can run at private or changeable origins. > - Paperclip Cloud can provide the stable callback while each instance keeps its durable provider credentials. > - This pull request adds the instance side of that managed OAuth protocol and keeps customer-created clients available. > - The benefit is a safe path to one-click Workspace connections for hosted and enrolled self-hosted instances. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change updates the server, Apps UI, shared app definitions, and connection documentation. **Problem or motivation** Some OAuth providers require a pre-registered callback and provider-owned client. An arbitrary self-hosted Paperclip origin cannot use that client callback directly. Paperclip ID must also stay limited to product identity instead of resource authorization. **Proposed solution** Use the existing Paperclip Cloud application as the fixed callback broker. Enroll each instance to an exact origin and separate Ed25519 and X25519 keys. Bind every request and sealed envelope to the instance, environment, user, company, provider, profile, and exact scope set. Store durable provider credentials only in the originating instance vault. **Alternatives considered** Customer-created OAuth clients remain available as the independent fallback. A generic redirect relay was rejected because it would allow caller-selected destinations and scopes. Paperclip ID was rejected as the broker because it is the identity boundary. A new service was rejected because the existing Cloud application already owns customer login and the public callback origin. **Roadmap alignment** This work extends the shipped MCP Tool Gateway and Apps milestone. It also supports the Connected Apps and Cloud deployments roadmap items. Companion Cloud implementation: https://github.com/paperclipai/paperclip-cloud/pull/312 The duplicate search found no related open Paperclip PR or issue. ## What Changed - Add a `paperclip_cloud_connector` client with signed requests, exact profile and scope bindings, and X25519-sealed credential handling. - Add explicit self-hosted enrollment with owner-only instance key storage and exact HTTPS origins. - Route managed Google Workspace setup through Paperclip Cloud and preserve customer-created OAuth clients. - Keep broker claims retryable until the local vault transaction commits. - Keep managed Google per-profile removal local-only to avoid client-wide provider revocation. - Add setup status to the Connections page and retain the Paperclip ID names as compatibility aliases. - Document the trust boundaries, enrollment, callback, refresh, removal, and rollout flows. ## Verification - `pnpm -r typecheck` - `pnpm --filter @paperclipai/shared exec vitest run src/app-definitions.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/services/paperclip-cloud-connector.test.ts src/services/paperclip-cloud-connector-enrollment.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts -t 'brokered Gmail OAuth|brokered OAuth state'` - `pnpm --filter @paperclipai/ui exec vitest run src/pages/apps/Connections.test.tsx` - `pnpm check:token-gates` - `pnpm build` - The full stable test runner also reproduced existing macOS workspace, skill-discovery, and listener fixture failures outside the changed paths. GitHub Linux CI is the authoritative full-suite result. ## Risks - The managed flow depends on https://github.com/paperclipai/paperclip-cloud/pull/312. Real provider profiles stay disabled until Cloud deploys that protocol and the provider approves the managed client. - A Cloud outage blocks new authorization and refresh. Existing access tokens continue to work until expiry. - Managed Google profile removal only deletes the local grant. This avoids invalidating the user's other profiles that share the managed Google client. - Legacy `paperclip_id_connector` records require a reconnect after their current access tokens expire. Old Paperclip ID keys and refresh tokens are not sent to Paperclip Cloud. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5.6 (Codex). Agentic coding, tool use, code execution, and subagents were enabled. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2e5a24e177 |
feat(runner): add qualified OpenCode runtime (#12588)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner provides a durable execution boundary for supported providers. > - The current production runtime supports Codex but cannot execute OpenCode sessions. > - OpenCode needs a qualified transport, strict input mapping, and normalized events. > - This pull request adds the OpenCode runtime as one isolated provider unit. > - The benefit is a reviewable provider expansion that does not weaken the existing Codex path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: packages/paperclip-runner and the Codex-local adapter configuration contract. **Problem or motivation** Paperclip Runner has provider-neutral contracts, but the production backend factory cannot start a qualified OpenCode session. This blocks OpenCode from using the durable runner path. **Proposed solution** Add the qualified OpenCode app-server proxy, driver, MCP bridge, backend, fixtures, and factory wiring. Keep existing Codex behavior unchanged. **Alternatives considered** Keeping OpenCode only on the direct adapter path would avoid this runtime work, but it would not provide durable runner recovery or normalized provider events. **Roadmap alignment** ROADMAP.md does not list a conflicting provider-runtime project. This change extends the existing Paperclip Runner architecture. ## What Changed - Added the qualified OpenCode app-server proxy and input queue. - Added collaboration-mode and provider-event normalization. - Added the OpenCode MCP bridge and native session backend. - Added strict fixtures and focused unit coverage. - Added only the package exports and adapter configuration required by this runtime. - Kept deferred SDK, lab, eval, and public package surfaces out of this change. ## Verification - GitHub Actions is the authoritative verification environment for this PR. - Run the package type checks and focused OpenCode tests in CI. - Run repository typecheck, test, build, security, and policy gates through the stack-aware workflow. - Local tests were not run because this checkout is resource constrained. ## Risks - OpenCode protocol changes could affect event normalization or recovery. - The driver fails closed on malformed input and unsupported runtime behavior. - Existing Codex selection remains unchanged unless the stored provider is OpenCode. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack - Position: 1 of 4 - Base: master - Next: additional qualified provider runtimes |
||
|
|
8610e7934e |
fix(release): bundle vendored runner ACPX runtime (#12582)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The npm release includes the Paperclip server and a vendored runner. > - The vendored runner imports ACPX when it starts a Codex agent. > - The server package did not include the ACPX version that the runner needs. > - A fresh canary install therefore stopped with `ERR_MODULE_NOT_FOUND` after onboarding. > - This pull request bundles the patched ACPX runtime with the server package. > - The benefit is that a fresh npm install can load the vendored runner. ## Linked Issues or Issue Description No public issue exists for this bug. A GitHub search found no duplicate or related pull request. **What happened?** A fresh `npx paperclipai@canary onboard` command completed onboarding. The server then failed to start. Node could not resolve `acpx` from the vendored Paperclip runner. **Expected behavior** The server should start after onboarding from a fresh npm cache and a temporary data directory. **Steps to reproduce** 1. Run `npx paperclipai@canary onboard --data-dir "$(mktemp -d /tmp/paperclip-canary.XXXXXX)"`. 2. Select Quickstart. 3. Start Paperclip. 4. Observe `ERR_MODULE_NOT_FOUND` for `acpx`. **Paperclip version or commit** `paperclipai@2026.831.0-canary.6` **Deployment mode** Other: local trusted Quickstart through `npx`. **Installation method** npm through `npx`. **Agent adapter(s) involved** Codex. **Database mode** Embedded PGlite. **Access context** Board operator during onboarding. **Node.js version** Node.js 26.4.0. **Operating system** macOS. **Relevant logs or output** ```shell Cannot find package 'acpx' imported from .../node_modules/@paperclipai/server/dist/vendor/paperclip-runner/drivers/acpx/codex-runtime-adapter.js ``` **Relevant config (if applicable)** No custom configuration was required. **Additional context** The published adapter utilities contain a nested `acpx@0.12.0`. Node cannot resolve that nested package from the sibling vendored runner. Installing `acpx@0.13.1` at the clean package root makes the failing runner import succeed. **Privacy checklist** The log excerpt contains no user path, token, company name, or other private value. ## What Changed - Added `acpx@0.13.1` as a bundled server runtime dependency. - Added a version-specific patch check for the ACPX versions used by the server and adapter utilities. - Added release-package coverage for the server ACPX bundle. ## Verification - `pnpm test:release-registry` passed 98 tests. - `node --test scripts/acpx-patch-packaging.test.mjs` passed 12 tests. - `pnpm exec vitest run server/src/__tests__/server-package-build-script.test.ts` passed 4 tests. - `node --test scripts/release-package-map.test.mjs` passed 12 tests. - `pnpm -r typecheck` passed. - A clean extracted server tarball contained the patched `acpx@0.13.1` runtime. - The previously failing vendored runner module imported from that clean tarball. - The repository-wide test suite was stopped before completion at the maintainer's request because it takes too long for this urgent packaging fix. ## Risks - Risk is low. - The server tarball grows because it now contains ACPX and its production dependencies. - The release stager now uses a version-specific marker to verify the ACPX patch. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`. - The model used reasoning mode, tool use, and code execution. - The context window size was not disclosed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7a3abb88a0 |
feat(runner): authorize server Codex tools (#12385)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The hidden native coordinator already computes a run-scoped semantic tool projection. > - The durable Codex backend now accepts and enforces that projection. > - The server did not include the projection in its `run.prepare` command. > - Codex therefore received no production semantic tools even when the server authorized them. > - This pull request adds the deterministic wire projection and sends it to runnerd. > - The benefit is one fail-closed authorization catalog from the server through Codex. ## Linked Issues or Issue Description Refs #12384 **What existing behavior does this improve?** This improves the existing flagged Paperclip Runner Codex path. **Current behavior** The server creates a run-scoped list of authorized read tools. It does not pass that list to runnerd, so the production Codex session starts with no tools. **Proposed behavior** The server maps the authorized definitions to the versioned runner contract. It computes a cross-language catalog digest. It includes that immutable contract in `run.prepare`. **Reason and benefit** Runnerd and the server now enforce the same catalog identity. Unknown, duplicate, changed, or malformed tool contracts fail before Codex can use them. **Breaking changes** None. Direct adapters are unchanged. A native run with an empty server projection still starts with no dynamic tools. ## What Changed - Add a deterministic semantic-definition to runner-authorization projection. - Match the Rust canonical digest with a shared test vector. - Include the server coordinator projection in the native Codex `run.prepare` command. - Extend the native Codex vertical slice to require and execute a semantic tool. - Verify the production prepare payload in a host-independent server test. ## Verification - `pnpm --filter @paperclipai/paperclip-runner test:typescript` (354 tests pass) - `pnpm --filter @paperclipai/server exec vitest run src/services/native-runtime/native-codex-runner.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm build` - The embedded-Postgres vertical slice is present for CI. This local host reports that embedded Postgres is unavailable, so Vitest skips that host-dependent test locally. - Confirmed that the PR changes 7 files against `runner-codex-durable-tools`. - Confirmed that `pnpm-lock.yaml` is unchanged. ## Risks The main risk is a catalog digest mismatch between TypeScript and Rust. Both implementations use canonical JSON. They share the same fixed digest vector. Runnerd also recomputes the digest and rejects a mismatch. The rollout flag and the existing native runtime selection rules remain unchanged. ## Model Used OpenAI Codex with GPT-5 and repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bc9ba7cd26 |
feat(runner): project native runs into task threads (#12321)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The experimental Paperclip Runner can execute a guarded Codex run and persist provider-neutral events. > - The task page still reads direct-adapter transcripts and cannot present those native events. > - Structured runner questions must also use the existing task interaction experience. > - Runtime selection must use the persisted run mode, not an adapter name or a current feature flag. > - This pull request projects native events and questions into the existing task thread. > - Direct adapters keep their existing transcript, composer, interaction, and finalization paths. > - The benefit is a complete native Codex task thread without a behavior change for existing adapters. ## Linked Issues or Issue Description Refs #12202. This pull request replaces that stale implementation on current `master`. **What happened?** The server persists native runner events and structured input requests. The task page only consumes direct-adapter transcripts. A native run therefore cannot present a complete transcript, usage, or question flow through the normal task experience. **Expected behavior** Native runs project persisted provider-neutral events into the existing task thread. Native structured questions use the existing interaction card. Direct adapters retain their current behavior. **Steps to reproduce** 1. Enable the experimental runner. 2. Start a native Codex run that emits progress, usage, a structured question, and a final reply. 3. Open the task page. 4. Observe that the direct-adapter transcript path cannot project the native event records. **Paperclip version or commit** `master` at `67f9867bc`. ## What Changed - Add the canonical structured-question validator and shared contract exports. - Materialize native input requests as existing task interactions. - Validate native answers and deliver them through the durable question-response receipt. - Resume the original PRP request with an idempotent `request.resolve` command. - Project native messages, tool activity, cumulative usage, and final replies into the existing transcript model. - Propagate persisted `runtimeMode` to the task page and select native handling only for `runtimeMode: "native"`. - Expire pending interactions through the shared issue service on every terminal transition, including decisions, stalled reviews, tree control, and pipeline retry cleanup. - Queue native run cancellation while a transaction is open and execute it only after the owning transaction commits. - Keep nonterminal and non-runner issue paths on their existing service call shapes and behavior. ## Verification - `pnpm --filter @paperclipai/server typecheck` — passed, including the Rust runner release build and protocol/catalog drift gates. - Focused native-thread and lifecycle suites — 18 files and 481 tests passed during review. - `issue-execution-policy-routes.test.ts` — 19/19 passed after the final transactional-queue expectation update. - `issue-agent-mutation-ownership-routes.test.ts` — 87/87 passed in the final isolated compatibility rerun. - GitHub Actions — policy, build, canary, typecheck/release registry, 5 serialized server shards, 8 general-test shards, 3 browser shards, and both aggregate gates passed on `7793f3193`. - Security — Snyk, Socket Project Report, Socket PR Alerts, and Superagent passed. - Greptile — 5/5 on `7793f3193`; all actionable review threads resolved. - `git diff --check` — passed. - Diff against `master`: 44 files. ## Compatibility Boundary - Native transcript polling only runs when the persisted run reports `runtimeMode: "native"`. - Missing or legacy runtime modes continue through `useLiveRunTranscripts`. - Legacy questions keep the existing optional free-text choice. - Native closed select sets can suppress that legacy fallback. - Terminal cleanup uses the same issue service for native and legacy interactions; only a bound native question schedules a native run cancellation. - Native cancellation happens after transaction commit, so failed or rolled-back writes do not cancel a still-valid run. - The durable delivery service checks the original native request before it considers a continuation run. - This pull request adds no migration, dependency, workflow, manifest, or lockfile change. ## Risks The main risk is routing a direct-adapter task through native handling or changing terminal issue behavior. The implementation selects the native path only from persisted runtime facts, retains the existing nonterminal call shape, and schedules native cancellation only for a validated bound native question after commit. Focused and repository-wide tests cover both paths. Native requests remain bound to the company, issue, run, and agent; answers are validated, durable, and idempotent across reconnects. ## Model Used OpenAI Codex, GPT-5 family. The client does not expose the exact deployment ID or context window. Agentic reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected local tests and they pass - [x] I have added or updated tests where applicable - [x] I have updated the compatibility notes for this change - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I addressed all Greptile and reviewer comments before requesting merge |
||
|
|
4310b0c947 |
refactor(server): remove unreachable task-drain compensation paths (#12511)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task-drain service controls when task execution can start and stop. > - The service had compensation paths for states that its validators or recovery process already handle. > - These paths added rollback state and a stuck-claim marker without improving normal drain behavior. > - This pull request removes the unreachable TTL clamp, audit rollback, generation counter, and double-fault marker. > - The result keeps input validation, audit ordering, atomic release, and orphan recovery. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task-drain service and its routes manage drain state, audit rows, and execution locks. **Subsystem affected** server/ — REST API and orchestration services. **Current behavior** The service clamps a validated TTL value. The routes mutate drain state before audit writes and then restore state after a failed write. Claim release also tracks a second durable-write failure with an in-memory marker. **Proposed behavior** The validator remains the single TTL policy. The routes write audit rows before they mutate drain state. Claim release logs a failed write and lets the orphan reaper release the issue lock. **Reason and benefit** The removed paths cannot handle a valid API request that reaches them. The rollback can lose the original start time. The marker can keep a drain non-quiescent until process restart. The simpler flow keeps state consistent and uses the existing recovery path. **Breaking changes** None to the public API. A failed claim release keeps the issue lock until the next orphan-reaper cycle. ## What Changed - Remove the service-layer TTL clamp because the shared validator rejects values above the limit. - Write task-drain audit rows before drain mutation and remove the rollback helpers. - Remove the rollback generation counter and its unused state. - Remove double-fault stuck-claim tracking and keep the atomic release path. - State that the quiescent flag describes work in this process. - Keep the orphan reaper as the recovery path after a failed claim release. ## Verification - Run `pnpm --filter @paperclipai/server test server/src/__tests__/heartbeat-task-drain-admission-release.test.ts`. - Run `pnpm --filter @paperclipai/server test server/src/__tests__/heartbeat-task-drain.test.ts`. - Run `pnpm --filter @paperclipai/server test server/src/__tests__/instance-settings-routes.test.ts`. - Run `pnpm --filter @paperclipai/server test server/src/__tests__/heartbeat-scheduling-suppression.test.ts`. - Run `pnpm --filter @paperclipai/server test server/src/__tests__/execution-lock-orphan-cleanup.test.ts`. - The five affected test files pass with 70 tests. - Confirm the full pull request checks pass before merge. ## Risks The issue lock remains held until the orphan reaper runs after a failed claim release. This uses the existing recovery path for interrupted runs. The change does not alter the public API or database schema. ## Model Used OpenAI Codex, GPT-5, 400K context window, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a560b48d6d |
feat(apps): refine Postman and Shopify setup (#12357)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give those agents governed access to external tools. > - Provider catalogs must match each provider's current protocol and credential contract. > - Postman method labels and API-key placement were outdated. > - Shopify now offers a UCP commerce endpoint that needs a managed agent-profile argument. > - This pull request updates both providers and documents the complete connection-authoring workflow. > - The benefit is accurate setup, safer runtime defaults, and a repeatable provider review process. ## Linked Issues or Issue Description Refs #11965 This is stack 10 of 11. It depends on stack 9 and preserves the final catalog work recovered from #11965. Related: #5904 covers Shopify skill routing. This pull request covers the Apps connection contract instead. ## What Changed - Update Postman hosted MCP methods, capability choices, default selection, and bearer-token placement. - Add Shopify UCP commerce and Storefront compatibility methods with public-store prerequisites. - Inject the reviewed Shopify UCP agent profile at runtime and remove that managed field from user input schemas. - Classify Shopify checkout completion and cancellation as destructive actions. - Expand the connection authoring runbook from provider research through verification and pull request handoff. - Add focused shared, server, and UI coverage. - Make the approved-execution waiter phase-aware so slow preparation cannot consume the provider execution timeout and grace period. - Settle legacy pre-execute-on-approve requests and invocations as failed, clear their stale idempotency key, and allow a fresh governed approval instead of leaving work stuck in `executing`. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts server/src/__tests__/tool-access-service.test.ts ui/src/pages/apps/AppsConnect.test.tsx -t "Postman|Shopify|normalizeConnectionMethodConfig|classifyRisk"` (16 passed) - `pnpm exec vitest run server/src/services/approved-execution-wait.test.ts` (4 passed) - `pnpm exec vitest run server/src/__tests__/tool-gateway.test.ts -t "enforces policy, approvals, retries, rate limits, and company boundaries for connected remote MCP calls"` (1 passed) - `pnpm exec vitest run server/src/__tests__/tool-gateway-service.test.ts` (21 passed; includes legacy approval settlement and fresh-approval recovery) - `pnpm --filter @paperclipai/server typecheck` - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` ## Risks - Shopify UCP calls now include a Paperclip-managed agent profile that overrides caller input at the same path. - Postman EU credentials now use the hosted MCP server's bearer-token contract instead of the general REST API header. - The catalog generator and checked-in definitions change together to prevent regeneration drift. - Approved execution preparation has an explicit two-minute bound; provider execution retains its own 65-second timeout and persistence grace starting from durable provider start. - Legacy approvals created before execute-on-approve are intentionally terminalized and must be requested again under the current signed contract. > I checked `ROADMAP.md`. This provider update does not duplicate planned core work. The related open Shopify PR addresses skill routing, not Apps connections. ## Model Used OpenAI Codex, GPT-5. The runtime exact model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d387cc0ff0 |
feat(connections): add managed external MCP connectors (#12346)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connection intents need secure provider implementations to complete setup. > - Some providers use managed OAuth or external credential brokers. > - Those tokens must stay out of durable Paperclip state and fail closed when refresh fails. > - This pull request adds managed connector backends and the required storage contract. > - The benefit is safer provider setup with governed credential lifecycles. ## Linked Issues or Issue Description Refs #11965 This is stack 8 of 11. It depends on stack 7 and replaces another reviewable part of #11965. ## What Changed - Add managed Google Workspace and external connector backends. - Add Vercel Connect support without storing provider bearer tokens. - Add replay-safe migration 0232 and its generated snapshot. - Fail closed and clear stale token bindings when organization OAuth refresh needs reauthorization. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 194 tests passed. - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` - `pnpm exec vitest run --project @paperclipai/server server/src/services/remote-url-credentials.test.ts` (5 passed, including URL userinfo vault extraction) ## Risks - Broker metadata errors can block provider setup. - OAuth refresh failure disables the shared organization connection until reauthorization. - Migration 0232 is generated, ordered after 0231, and safe to replay. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the public source pull request with `Refs #` - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b3343dbd64 |
feat(connections): add self-serve intent runtime (#12345)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a governed way to request app connections during issue work. > - The catalog now describes the available providers and setup methods. > - A request must become a durable, company-scoped intent before an operator acts on it. > - This pull request adds that intent runtime across server, agent, CLI, and shared contracts. > - The benefit is a safe bridge from agent need to operator-approved setup. ## Linked Issues or Issue Description Refs #11965 This is stack 7 of 11. It depends on stack 6 and replaces another reviewable part of #11965. ## What Changed - Add connection intent types, validation, service logic, and routes. - Add agent runtime tools and CLI support for connection requests. - Add issue-thread interaction support for connection intents. - Add runtime, route, adapter, and contract tests. - Hold the final resolved-continuation row lock through asynchronous adapter preparation until an actual process spawn, so parking or reassignment cannot cross that boundary. - Report Hermes Gateway's first remote run request through the shared dispatch hook so the resolved-intent lock is released at the true dispatch boundary. - Revalidate the addressed user's live non-viewer membership and connection-management authority for every intent mutation, including OAuth completion. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 176 tests passed. - `pnpm build` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-stale-queue-invalidation.test.ts` (32 passed; includes non-process dispatch lock-release coverage) - `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/connection-intents-service.test.ts -t "addressed-user mutation"` (1 passed) - `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/tool-access-service.test.ts -t "binds OAuth callback completion to the initiating board session"` (1 passed) - `pnpm --filter @paperclipai/hermes-paperclip-adapter test -- src/gateway/server/execute.test.ts` (23 passed; includes dispatch-hook ordering and exactly-once coverage) - `pnpm --filter @paperclipai/hermes-paperclip-adapter typecheck` ## Risks - A malformed intent could create an unusable operator request. - Validators and company checks reject invalid or cross-company requests. - The final continuation gate holds the issue row lock through adapter preparation until process or remote dispatch; later operator changes use the normal active-run interruption path. - The change does not add a database migration. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the public source pull request with `Refs #` - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fcb2e99e8f |
feat(apps): expand the self-serve connection catalog (#12344)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A useful app store needs accurate and selectable provider definitions. > - Local brand assets now cover the expanded provider set. > - Provider methods differ in transport, authentication, ownership, and required scope. > - This pull request expands the catalog and encodes those provider contracts. > - The benefit is a larger self-serve store with explicit setup choices. ## Linked Issues or Issue Description Refs #11965 This is stack 6 of 11. It depends on stack 5 and replaces another reviewable part of #11965. ## What Changed - Add and update provider definitions for the self-serve catalog. - Add Google Workspace connection methods and capability profiles. - Add catalog generation, ingestion, URL matching, and contract tests. - Update legacy key tests to use a provider that still uses header credentials. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 164 tests passed. - `pnpm build` ## Risks - An incorrect provider definition can offer the wrong setup method. - Contract tests verify transport, authentication, and provider URL behavior. - The change does not add a database migration. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Refs #` or (b) described the issue in this pull request - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6244e4cf32 |
feat(apps): add Composio and Gmail connectors (#12342)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - App connections need both direct providers and managed provider hubs. > - The grant layer now defines safe credential ownership. > - Composio needs parent and child connection lifecycle rules, and Gmail needs governed setup. > - This pull request adds both connector families on the grant foundation. > - The benefit is broader app access without weakening credential isolation. ## Linked Issues or Issue Description Refs #11965 This is stack 4 of 11. It depends on stack 3 and replaces another reviewable part of #11965. ## What Changed - Add Composio parent and child connection support. - Add Gmail connection setup and governance. - Preserve credential paths and remove duplicate binding declarations. - Cascade Composio pause and restore actions to child connections. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 164 tests passed. - `pnpm build` ## Risks - Parent lifecycle changes can affect every Composio child. - The service restores only children whose provider accounts remain active. - Credential binding paths are normalized before secret resolution. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
20ccf3f476 |
feat(apps): add connection grants and delegated identities (#12341)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - External tools need explicit identity and access boundaries. > - Shared connection credentials cannot represent every user-scoped use case. > - Grants must stay company-scoped and support safe delegation. > - This pull request adds connection grants, identity rules, and their database contract. > - The benefit is durable control over which identity an agent may use. ## Linked Issues or Issue Description Refs #11965 This is stack 3 of 11. It depends on stack 2 and replaces another reviewable part of #11965. ## What Changed - Add company and user connection grants. - Add delegated identity and membership rules. - Synchronize database, shared, server, and UI contracts. - Register the grant-member replacement route in the OpenAPI surface in the same layer that mounts it. - Add migration 0231 with replay-safe guards and coverage. ## Verification - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/openapi-routes.test.ts` (5 passed) - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` ## Risks - Incorrect grant selection could expose the wrong credential scope. - The service enforces company and subject boundaries before credential use. - Migration 0231 is generated, ordered after 0230, and safe to replay. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked a public issue or pull request with `Refs #` - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b51112798f |
feat(apps): improve gateway and workspace connection UX (#12340)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - App connections must work in both the operator UI and agent tool gateway. > - The first stack layer adds secure remote connections. > - Operators still need clear setup, test, and recovery states. > - This pull request adds the gateway behavior and the workspace connection experience. > - The benefit is a connection flow that is easier to understand and recover. ## Linked Issues or Issue Description Refs #11965 This is stack 2 of 11. It depends on stack 1 and replaces another reviewable part of #11965. ## What Changed - Improve remote tool gateway connection behavior. - Add clearer app setup, test, and recovery states. - Add focused server and UI tests for the new paths. - Keep the diff isolated from later identity and catalog work. - Stabilize DNS-pinned remote HTTP protocol fixtures and the managed-runtime public-origin fixture for this independently tested layer. ## Verification - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` (150 passed) - `pnpm test:run` - `pnpm check:token-gates` - `pnpm build` ## Risks - Gateway errors now surface through new user-facing states. - A stale connection can require a new setup attempt. - The change does not add a database migration. - The injected HTTP transport and public URL are test-only fixtures; production DNS pinning and runtime behavior are unchanged. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cabc9146d0 |
feat(apps): add secure remote MCP and PostHog setup (#12339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give those agents governed access to external tools. > - Remote MCP setup needs secure endpoint validation and durable credentials. > - PostHog needs both browser sign-in and personal API key setup paths. > - This pull request adds the shared remote MCP foundation and the PostHog definition. > - The benefit is a secure and reusable base for later app connection work. ## Linked Issues or Issue Description Refs #11965 This is stack 1 of 11. It replaces the first reviewable part of #11965. ## What Changed - Add guarded remote MCP setup and credential handling. - Add PostHog OAuth and API key connection methods. - Add focused server, shared contract, and UI coverage. - Keep the migration replay-safe and idempotent. - Give the late-close security regression the same 10-second CI headroom as the adjacent real-timer handshake test. - Synchronize fake-timer handshake tests at the exact ensure-session boundary so real filesystem setup cannot race the fake deadline. - Drive PTY overflow coverage only after listener registration so scheduling cannot reorder the test fixture. ## Verification - pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts server/src/__tests__/plugin-worker-manager.test.ts (220 passed; affected cases also passed five focused stress repetitions) - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never leaks a sandbox-provided value from a late close rejection into logs or the result"` (1 passed) - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never promotes a late ensureSession resolution|closes a late-resolving real handle exactly once"` (2 passed) - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` ## Risks - Remote endpoint validation can reject configurations that previously passed without checks. - OAuth configuration errors can block setup until the operator corrects the provider settings. - The migration uses guarded statements so repeated execution is safe. - The test-only synchronization changes do not affect runtime behavior; they remove filesystem/fake-clock and listener-registration races observed under parallel CI load. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6154e00f26 |
feat(server): add a task-drain admission hold to the instance API (#12485)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server admits agent work through heartbeat scheduling and execution paths > - Operators need to stop new work before maintenance or a graceful shutdown > - A process restart alone does not provide a reusable admission control primitive > - This pull request adds an instance API that holds new task admission and reports process quiescence > - The benefit is a small, auditable control that lets operators wait for active work without a restart ## Linked Issues or Issue Description **Problem or motivation** Operators cannot hold new task admission without restarting the Paperclip process. A restart can interrupt maintenance flows and does not provide a status signal for active work. **Proposed solution** Add `GET /instance/task-drain`, `POST /instance/task-drain`, and `DELETE /instance/task-drain`. The server keeps the drain state in process memory, applies it to every scheduling suppression path, supports an optional TTL up to 24 hours, and reports active wake and run counts. **Alternatives considered** A timer would clear the drain after its TTL, but it could keep the Node.js event loop open during shutdown. A database row would add storage and query work for process-local state. The implementation uses lazy expiry and process memory instead. **Roadmap alignment** The change supports the roadmap goal for enforced outcomes and safe recovery actions. It does not duplicate a listed roadmap item. **Additional context** This is a server and shared-package change. It adds no user interface and no database migration. ## What Changed - Add process-local task-drain state with lazy TTL expiry. - Add task-drain admission suppression to the shared heartbeat resolver. - Add instance routes to read, start, and stop a task drain. - Add validation for positive TTL values and the shared 24-hour maximum. - Add activity records for drain mutations and tests for status, access control, validation, and suppression. ## Verification - Run `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/heartbeat-task-drain.test.ts server/src/__tests__/instance-settings-routes.test.ts server/src/__tests__/heartbeat-scheduling-suppression.test.ts`. - Run `pnpm --filter @paperclipai/shared exec tsc --noEmit`. - Run `pnpm --filter @paperclipai/server exec tsc --noEmit` and compare its known pre-existing errors with the base commit. - Confirm that pull request CI reaches a terminal green state. ## Risks The drain state exists only in process memory, so a restart clears it. This behavior matches the process-local design. A drain without a TTL remains active until an operator calls the delete route. The status route reads in-memory activity sets and does not query stale database rows. ## Model Used OpenAI Codex, GPT-5, extended reasoning with tool use and code execution. The exact runtime context window is not exposed by the execution environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a20a4944ec |
feat: add Grok device login to the sandbox login panel (#12469)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses adapters to connect agents and model providers to its control plane > - The sandbox login panel supports displayed-code login for selected adapters > - Grok users need the same login path and a private credential home for later runs > - This pull request adds Grok support to the shared device-login path and preserves the existing Codex path > - The benefit is one secure login flow for both adapters with company-scoped credential storage ## Linked Issues or Issue Description **Agent or provider** Grok Local needs displayed-code login support in the sandbox login panel. **Why this adapter is useful** This change lets users sign in to Grok from the sandbox login panel. It also gives later Grok runs access to the stored credential. **How the agent is invoked** The Grok local adapter uses its login command through the shared displayed-code login flow. Later runs receive the managed home through `GROK_HOME`. **Additional context** The change uses adapter-scoped login lifecycle handling. It stores the credential in a company-scoped directory with mode `0700`, and it stores the credential file with mode `0600`. ## What Changed - Rename the shared device-login modules to adapter-neutral names. - Scope the shared login lifecycle to a closed adapter set. - Return the device-login URL that the provider prints. - Add the Grok prompt parser, login command, capability, and login panel entry. - Store the Grok credential in a private, company-scoped home directory. - Pass `GROK_HOME` to later Grok runs. - Add tests for the Grok adapter, the Daytona sandbox provider, the server login path, and the user interface. ## Verification - Run `pnpm vitest run packages/adapters/grok-local/src/server/adapter-auth-promotion.test.ts`. - Run the Grok adapter package suite. - Run the Daytona sandbox provider suite. - Run the server device-login suites. - Run the user interface suite. - Confirm the full CI suite passes. ## Risks The change extends shared login lifecycle code to another adapter. A regression could affect Codex login. The credential path uses explicit `chmod` calls to keep the directory at mode `0700` and the file at mode `0600`. ## Model Used OpenAI Codex, GPT-5. The runtime used tool calls and code review support. The runtime did not provide a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cec675ffca |
test(server): load the agent-permissions route module graph once per file (#12471)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server test suites verify agent permissions and route behavior > - The agent-permissions route suite rebuilt its full module graph for every test > - CPU load made that repeated work exceed the test timeout and caused intermittent failures > - This pull request loads the route module graph once for the file and resets each mock before every test > - The benefit is faster, stable test execution with the same test coverage and isolation ## Linked Issues or Issue Description **What happened?** `server/src/__tests__/agent-permissions-routes.test.ts` failed intermittently in continuous integration. The suite reset modules and re-imported the route module graph for every test. Under CPU load, one import took seconds instead of milliseconds and caused an expected response to become an HTTP 500. **Expected behavior** The suite should run all 54 cases without intermittent timeout failures. Each test should keep isolated mock state. **Steps to reproduce** 1. Run `npx vitest run server/src/__tests__/agent-permissions-routes.test.ts`. 2. Repeat the file run under high CPU load. 3. Compare the failure rate and run time before and after this change. **Paperclip version or commit** `c4d1af4216f174a92823ca3a20e0717c54371dd5` **Deployment mode** Built from source. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Not adapter-specific. This change affects a server test suite. **Database mode** Not database-related. ## What Changed - Load the route module graph one time for the describe block with the existing `hoistModuleGraph` helper. - Make `createApp` synchronous and read the hoisted graph. - Remove per-test `vi.resetModules()` and the 26 `vi.doUnmock(...)` calls. - Keep stable mock objects and reset each route-facing mock before every test. - Keep all 44 `it` blocks and 54 parameterized cases. ## Verification - `npx vitest run server/src/__tests__/agent-permissions-routes.test.ts` passes and reports 54 tests. - `npx tsc --noEmit -p server/tsconfig.json` passes with 0 errors. - Under 32 concurrent CPU-bound loops, the file passed 15 of 15 runs after this change, with 54 of 54 cases on each run. - The same test failed 1 of 15 runs before this change. - Per-run wall-clock time changed from about 29–34 seconds to about 8–10 seconds. ## Risks - Low risk. The change affects test setup only. - The hoisted mock objects keep stable identity, and `beforeEach` resets every route-facing mock. - The registration step arms no mock implementations, and the test adapter still unregisters in a `finally` block. ## Model Used OpenAI Codex, GPT-5. The runtime provided tool use and code execution. The runtime did not provide a context window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad474abece |
fix(server): load the mocked module graph once in the closed-workspace issue route suite (#12470)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server tests cover issue routes and execution workspace state > - The closed-workspace route suite reloads a mocked module graph before each test > - CPU contention can bind one test to the real service and hide a 500 response > - This pull request loads the mocked graph once and checks exact success statuses > - The benefit is a stable suite that detects route failures instead of accepting them ## Linked Issues or Issue Description **What happened?** The closed-workspace issue route suite reloaded and unmocked the module graph before each test. Under CPU contention, a route could bind to the real execution-workspaces service. The request then returned `500`, while a weak assertion accepted the result. **Expected behavior** The suite must use the configured service mocks for every test. Each success case must assert its exact expected HTTP status. **Steps to reproduce** 1. Run the closed-workspace route suite many times in parallel. 2. Use CPU contention during the run. 3. Observe intermittent failures or weak assertions that accept `500` responses. **Paperclip version or commit** Commit `0d5686b942f1d1366112cb922f95e5c8622bb9a6`. **Deployment mode** Built from source with the server test runner. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific. **Database mode** Not database-related. **Additional context** This pull request relates to [#11472](https://github.com/paperclipai/paperclip/pull/11472). It does not change product code. ## What Changed - Load the mocked service graph once per file with `hoistModuleGraph`. - Remove the per-test module reset and unmock cycle. - Assert exact success statuses for the three affected responses. - Add the missing `refreshReopenPendingConsumption` mock. - Use one named timeout constant for every `vi.waitFor` call. - Replace `setImmediate` barriers with fake-timer advances. ## Verification - `npx vitest run --project @paperclipai/server server/src/__tests__/issue-closed-workspace-routes.test.ts` passes 12 of 12 tests locally. - A 200-sample sweep with 20 parallel copies passed with zero failures. - An independent 100-sample sweep with 20 parallel copies passed with zero failures. - `npx tsc --noEmit -p server` reports the same 61 pre-existing errors with and without this change. - GitHub Actions must run the full server suite for final verification. ## Risks Low risk. This pull request changes one test file and does not change product code. The stricter assertions can expose a real route failure that the old suite hid. ## Model Used Codex, GPT-5, tool use and code review assistance. The exact context window and reasoning mode are not available in this handoff. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
64b7dce0ad |
refactor(adapter-utils): replace the process-wide byte ledger with route-local byte bounds (#12465)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The adapter layer carries sandbox requests to host processes. > - The HTTP/2 bridge used one process-wide byte ledger for all routes. > - One busy route could exhaust that shared budget and move another route to file transport. > - This pull request gives each host retention site a fixed byte bound and limits concurrent HTTP/2 streams. > - The benefit is local protection: one route cannot consume the byte budget of another route. ## Linked Issues or Issue Description **What happened?** The HTTP/2 bridge used one aggregate byte ledger for retained bytes across all routes. A busy route could exhaust the shared budget and force an unrelated route to use file transport. **Expected behavior** Each route should protect its own retained bytes. A reset on one HTTP/2 stream should cancel only that stream's host forward. **Steps to reproduce** 1. Start the HTTP/2 bridge with multiple sandbox routes. 2. Send enough retained data through one route to reach the aggregate byte limit. 3. Send a request through a sibling route. 4. Observe that the sibling route can fall back to file transport because the first route used the shared ledger. **Paperclip version or commit** `47639e227e78e3c5e0dd1a3c0e2d792fe86895a3` **Deployment mode** Built from source with the adapter-utils and server test suites. ## What Changed - Bound each host retention site with a fixed local byte limit. - Limited concurrent live HTTP/2 streams with one built-in stream limit. - Bound each host forward and response-body read to its own HTTP/2 stream lifetime. - Removed the process-wide byte ledger, its environment override, its metrics, and its file-transport fallbacks. - Added tests for the stream limit, host body budget, and sibling-stream cancellation. ## Verification - Run `pnpm vitest run --project adapter-utils`. - Confirm that 996 adapter-utils tests pass. - Confirm that `test_live_forward_work_never_passes_the_stream_limit` passes. - Confirm that `test_the_host_body_budget_matches_the_stream_limit` passes. - Confirm that the sibling-stream cancellation test passes. - Run `pnpm tsc --noEmit`. - Confirm that all pull request checks pass. ## Risks The bridge no longer uses a process-wide byte ledger. A local bound or stream limit that is too low can reject or delay valid work. The tests cover the new limits and stream cancellation behavior. ## Model Used OpenAI GPT-5 Codex. Runtime model ID: GPT-5. The model used code execution and repository tools. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8da49d6ea7 |
fix(ci): make runner image tests deterministic (#12452)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Pull request checks protect the runtime and authorization boundaries > - The same checks must produce the same result on GitHub and RunsOn Ubuntu images > - One runtime test assumed that a shell PID always owns the listening socket > - One watchdog test denied issue reads while it tried to test assignment denial > - These assumptions caused image-sensitive failures during the AWS runner canary > - This pull request tests the production contracts directly > - The benefit is a reliable CI result across both runner images ## Linked Issues or Issue Description Refs #12350 ## What Changed - Verify stale service ownership through the existing process-group ownership helper. - Allow normal issue reads in the watchdog reassignment fixture. - Assert that the watchdog reassignment reaches and denies the `tasks:assign` guard. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts -t "does not reuse a stopped auto-port service port while another process owns it"` - `pnpm exec vitest run server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts -t "still enforces normal assignment guards for watchdog reassignment"` - Ran the complete issue agent mutation ownership suite six times. All 522 test executions passed. - Ran the runtime regression case eight times. All eight test executions passed. ## Risks - Low risk. This pull request changes test fixtures and assertions only. - The process-group assertion matches the ownership rule that the runtime already uses. - The watchdog fixture still denies `issue:mutate` and `tasks:assign`. ## Model Used - OpenAI Codex with GPT-5.6 (`gpt-5.6-sol`). The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes, Closes, or Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run relevant tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes. No documentation change is required for this test-only fix. - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d9449e636e |
feat(onboarding): sign in to an agent provider during onboarding (#12440)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - New organizations create their first agent through the onboarding wizard > - The wizard does not show provider sign-in when a host credential is absent or unknown > - The create step also gives unclear feedback when the provider needs authentication > - This pull request adds a safe auth signal and a provider sign-in step for sandbox drivers > - The benefit is a clearer onboarding path with no token or account data in the signal ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (server API, shared types, and UI) **Problem or motivation** The onboarding wizard can fail when the selected provider needs authentication. It does not tell the person how to complete sign-in. **Proposed solution** Add a status-only provider auth signal. Show the sign-in panel for sandbox drivers when the signal says `absent` or `unknown`. Apply a stored Claude login to the new agent and block creation when the adapter test reports missing authentication. **Alternatives considered** The wizard could hide the sign-in panel when the signal read fails. This would hide a needed action, so this pull request shows the panel when the signal is unknown. **Roadmap alignment** The change supports the roadmap goal for scoped and audited credential bindings. **Additional context** The auth signal returns only `present`, `absent`, or `unknown`. It never returns a token, identifier, or account name. ## What Changed - Add `GET /api/companies/:companyId/adapters/:type/auth-signal` with company and permission checks. - Add shared auth-signal types and the UI query path. - Apply a stored Claude login by reference without reading its token. - Show the provider sign-in panel only for sandbox drivers with interactive terminal support. - Block agent creation when the provider test reports missing authentication. - Add route, wizard, and end-to-end test coverage. ## Verification - `pnpm --filter @paperclipai/server test adapter-auth-signal-routes` passes 50 tests. - `pnpm --filter @paperclipai/ui test OnboardingWizard` passes 69 tests. - `pnpm --filter @paperclipai/ui exec tsc --noEmit` exits with code 0. - The `e2e_shards` lane runs `tests/e2e/onboarding.spec.ts`. ## Risks The route reads a host-local readiness signal. It returns `unknown` on read errors and never exposes credential data. The UI may add a sign-in step when the signal is unavailable. ## Model Used OpenAI Codex, GPT-5, extended reasoning, tool use, and code execution. The exact context window was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5f1e25c112 |
test(server): load the company-skills route module graph once per file (#12426)
## Thinking Path > - Paperclip is an open source app that helps people manage AI agents for work. > - The server provides company-scoped routes for company skills and their test runs. > - The authorization tests for these routes fail intermittently in continuous integration. > - The failure returns HTTP 500 instead of the expected HTTP 403. > - The test file resets and rebuilds its module graph before each test. > - This rebuild imports an unmocked issue service and raises a TypeError. > - This pull request loads the module graph once per describe block. > - The change keeps the test result stable and preserves all 53 tests. ## Linked Issues or Issue Description **What happened?** The company-skill test-run authorization tests failed intermittently in continuous integration. One test returned HTTP 500 instead of HTTP 403. **Expected behavior** Each unauthorized request must return HTTP 403. The test file must keep all 53 tests and skip none. **Steps to reproduce** 1. Run `npx vitest run server/src/__tests__/company-skills-routes.test.ts` before this change. 2. Repeat the run in continuous integration. 3. Observe the intermittent HTTP 500 result in an authorization case. **Paperclip version or commit** Commit `651d26a96f6e24811d336759d3e67ff3abb5ec29`. **Deployment mode** Built from source. Continuous integration runs the test suite. **Agent adapter(s) involved** Not adapter-specific. This issue affects server test module setup. **Database mode** Not database-related. ## What Changed - Load the mocked route module graph once for each describe block. - Use the existing `hoistModuleGraph` helper, as the cost service test does. - Remove twelve per-test `vi.doUnmock` calls and the redundant module rebuild. - Keep the change in `server/src/__tests__/company-skills-routes.test.ts` only. ## Verification - Run `npx vitest run server/src/__tests__/company-skills-routes.test.ts`. - Confirm that 53 tests pass and 0 tests skip. - Confirm that the diff changes only `server/src/__tests__/company-skills-routes.test.ts`. - Confirm that all Paperclip continuous integration checks pass. ## Risks Low risk. The change affects test setup only. It does not change production code, route behavior, or assertions. ## Model Used OpenAI GPT-5 (Codex), exact model ID `gpt-5`, tool use and code execution enabled. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8316ceb0b9 |
Add the Better Auth issuer column so signup and sign-in work (#12396)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A self-hosted install in `authenticated` mode signs users in with Better Auth, mounted at `/api/auth` over a hand-written Drizzle `account` table in `packages/db` > - Better Auth 1.7.0 added a required `issuer` field to that `account` model, plus a unique index on `(issuer, accountId)` > - The dependency bump in #11886 changed only `server/package.json` and the lockfile, so the Drizzle table never grew the column > - The Drizzle adapter checks the model against the schema on every write, so `linkAccount` throws and sign-up answers 500 with an empty body; a fresh install cannot create its first user, and an upgraded install locks out every existing user > - This pull request adds the `issuer` column and its unique index, and migrates the column in with a backfill that covers every existing row > - The benefit is that sign-up and sign-in work again, on a new install and after an upgrade ## Linked Issues or Issue Description No existing issue. Describing it inline, following `.github/ISSUE_TEMPLATE/bug_report.yml`. Refs #11886 (the dependency bump that introduced the required field). Refs #12269 (an earlier attempt at this fix; its backfill covers only `provider_id = 'credential'`). **What happened?** Sign-up fails on a self-hosted install. `POST /api/auth/sign-up/email` answers HTTP 500 with a zero-byte body. The server log carries: ``` [Better Auth]: The field "issuer" does not exist in the "account" Drizzle schema. # SERVER_ERROR: [BetterAuthError: The field "issuer" does not exist in the "account" Drizzle schema.] ``` The request writes the `user` row and then fails on the `account` row. The address is stuck after that: a second sign-up answers 422 `USER_ALREADY_EXISTS`, sign-in answers 401, and password reset answers 400 `RESET_PASSWORD_DISABLED` because the account that would hold the password does not exist. An upgraded install is worse. `sign-in/email` matches the credential account on `account.issuer === 'local:credential'`. Rows written before the upgrade have no issuer, so every existing user is locked out. **Expected behavior** `POST /api/auth/sign-up/email` answers 2xx and writes both the `user` row and its credential `account` row. `POST /api/auth/sign-in/email` then answers 2xx and sets a session cookie. An install that upgrades keeps its existing users. **Steps to reproduce** 1. Start a server from `master` with `PAPERCLIP_DEPLOYMENT_MODE=authenticated` against an empty database. 2. `curl -X POST http://127.0.0.1:<port>/api/auth/sign-up/email -H 'Content-Type: application/json' -H 'Origin: http://127.0.0.1:<port>' --data '{"name":"A","email":"a@example.com","password":"a-long-password"}'` 3. The response is HTTP 500 with an empty body. **Paperclip version or commit** `master` at |
||
|
|
7895f7f2b0 |
Install the declared Sentry server package into the hosted image (#12330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip supports opt-in Sentry error monitoring for server and browser errors. > - The hosted image must include the server package when an operator sets SENTRY_DSN. > - The server package is an optional peer in the source tree, so the image did not include it. > - This pull request installs the declared server package in the hosted image and checks the result. > - The benefit is a hosted tenant can send server errors without a manual package install. ## Linked Issues or Issue Description No public issue exists for this change. **What happened?** The hosted image did not include the declared @sentry/node server package. A hosted tenant could set SENTRY_DSN, but the server could not load the package from the image. **Expected behavior** The hosted image must include the exact @sentry/node version from server/package.json. The self-hosted image must remain without this optional package. **Steps to reproduce** 1. Build or pull the hosted image. 2. Resolve @sentry/node from the server package path. 3. Compare its version with server/package.json. 4. Confirm that the tsx loader path still resolves. **Paperclip version or commit** Commit b6ff556a33ebdbe764b7f495951cd59009776608. **Deployment mode** Docker hosted image. ## What Changed - Add a cloud-server-deps Docker stage that installs the declared @sentry/node version in isolation. - Copy the isolated package into the cloud image without changing the production image. - Add a probe that checks the tsx loader and the resolved Sentry version. - Run the probe after the hosted image push in the Docker workflow. - Add server tests and update the observability documentation. ## Verification - Run `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/cloud-image-sentry.test.ts`. - Confirm that the changed test passes in CI. - Confirm that all pull request checks pass. - Note that the Docker workflow does not run for pull requests. It runs after a push to master, for configured tags, or after manual dispatch. ## Risks - Low risk. The production image body stays unchanged. - The cloud image adds the declared Sentry package and a small dependency tree. - The workflow probe fails if the image loses the tsx loader or resolves a different Sentry version. ## Model Used OpenAI GPT-5; exact model version supplied by the execution service; tool use and code execution; context window not specified. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b7fd6c59b6 |
fix(server): load the costs-service route module graph once per file (#12375)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server test suite checks budget and cost routes > - The costs-service test rebuilt the full route module graph before every test > - Synchronous graph rebuilds caused long stalls under CPU load > - This pull request loads the mocked graph once for each describe block and keeps per-test mock setup > - The benefit is a stable 17-test file without production code changes ## Linked Issues or Issue Description **What happened?** The costs-service route test rebuilt its full mocked module graph before every test. Under CPU load, a rebuild sometimes stalled a test past the 15-second timeout. **Expected behavior** The test file should load its mocked route graph once for each describe block while each test keeps isolated mock behavior. **Steps to reproduce** 1. Run the costs-service route test under synthetic CPU load. 2. Repeat the file test 30 times. 3. Observe intermittent test timeouts before this change. **Paperclip version or commit** The test used the current master branch at the time of this change. **Deployment mode** Built from source. ## What Changed - Add `hoistModuleGraph` to load the mocked route graph once for each describe block. - Keep per-test mock setup in `beforeEach` so test isolation stays unchanged. - Keep all 17 tests and their assertions. - Remove the module graph rebuild from the per-test path. ## Verification - Run `npx vitest run src/__tests__/costs-service.test.ts` from `server/`. - Confirm that the file reports 17 tests and zero skipped tests. - Confirm that 30 runs under the same synthetic CPU load report 0.0% failure after the change, compared with 10.0% before the change. - Confirm that mutation checks still fail when each authorization guard is broken. ## Risks This change affects test setup only. The main risk is weaker test isolation if a mock keeps state between tests. Each test still re-arms its mock behavior in `beforeEach`, and the full assertion set remains. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This bug fix does not add a core feature. ## Model Used OpenAI GPT-5. The model used tool calls and code execution. The exact context window and reasoning configuration were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (the submitting engineer ran the file before handoff) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
036600d922 |
fix(db): close test database clients before the embedded Postgres cluster stops (#12335)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses database clients and embedded PostgreSQL test fixtures > - A fixture stopped its embedded PostgreSQL cluster while clients still held connections > - The postgres.js driver then scheduled a write on a stopped connection > - That write escaped the timer callback and caused a test process to exit with an error > - This pull request closes registered clients before the fixture stops its cluster > - The benefit is stable test teardown and clear failure reporting in continuous integration ## Linked Issues or Issue Description Refs: #10869 **What happened?** An embedded PostgreSQL test fixture stopped its cluster while database clients still held open connections. The postgres.js driver then scheduled a deferred write on a dead connection. The write caused an unhandled error after the test shard reported success. **Expected behavior** The fixture closes all live clients for its cluster before it stops the embedded PostgreSQL cluster. Tests then finish without a deferred write on a dead connection. **Steps to reproduce** 1. Run the database regression test with the embedded PostgreSQL fixture. 2. Stop the fixture while its database client still has an open connection. 3. Observe the deferred write and the process exit status. **Paperclip version or commit** Branch base: |
||
|
|
666f5a6e69 |
fix(server): make the wake-claim lease test deterministic (#12331)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server test suite checks question response delivery and wake claims > - One test used wall-clock time and could start a second delivery under load > - The second delivery reused one promise resolver and could hang for 15 seconds > - This pull request uses the injected clock and one resolver for each wakeup > - The benefit is a deterministic test that fails at once if a second wakeup occurs ## Linked Issues or Issue Description **What happened?** The wake-claim lease test slept for 70 milliseconds before it ran the pending sweep. Under load, the lease could look stale during that interval. The sweep then started a second delivery. The second delivery reused one promise resolver, so the test hung until the 15-second suite timeout. **Expected behavior** The test must control the time used by the service. One wakeup must use one resolver. An unexpected second wakeup must fail at once. **Steps to reproduce** 1. Run the question response delivery test under CPU load. 2. Let the test sleep before the pending sweep. 3. Observe that a second delivery can start and the test can reach the 15-second timeout. **Paperclip version or commit** Commit `0dd735e53a5cc9f6d3395834b826dfb0b1da2ea9`. **Deployment mode** Built from source with the server test suite. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific (core test issue). **Database mode** Not database-related. **Additional context** The change affects one test file. It does not change production source code. ## What Changed - Drive the test with the service's injected clock. - Give each wakeup call its own promise resolver. - Assert that lease renewal advances the last attempt time. - Assert that the wakeup runs one time and the attempt count stays at 1. ## Verification - Run `pnpm exec vitest run server/src/services/__tests__/question-response-delivery.test.ts`. - The changed file reports 29 passing tests. - Run the changed file 25 times, including 5 runs under CPU load. ## Risks Low risk. The change affects one test file and test setup only. It does not change production behavior. ## Model Used OpenAI GPT-5, exact runtime model ID supplied by the Paperclip agent environment, tool use and code review assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |