mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 20:34:57 +02:00
bac60d9d3164c2dee10df7c96d2cbabbde2fa686
296
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bac60d9d31 |
fix: preserve GitHub sign-in and show connected repository access (#12993)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - GitHub connections give agents an account with selected repository access. > - Fresh local instances enroll with production Paperclip Cloud. > - Enrollment could finish while the GitHub OAuth profile remained disabled. > - Setup then switched to a personal access token form without explanation. > - This change preserves sign-in intent and shows the connected account and repositories. ## Linked Issues or Issue Description Related: #12907, #12943, #12947. Existing open GitHub connection work was checked. No duplicate was found. **What happened?** After Cloud enrollment, a fresh test-drive asked for a GitHub key. Production did not advertise the managed GitHub profile. Staging did. The permissions page also omitted the authenticated username and repository names. **Expected behavior** Continue with GitHub OAuth when available. Explain unavailable sign-in and allow retry otherwise. Show the GitHub username and complete accessible repository list. **Steps to reproduce** Start a fresh test-drive. Choose GitHub and complete instance enrollment while the Cloud GitHub profile is disabled. Open an existing GitHub connection's permissions page. ## What Changed - Preserve managed sign-in intent when the gallery omits its profile. - Refresh the selected gallery entry on retry without resetting the audience. - Fetch all pages of GitHub installations and repositories. - Store only repository IDs, full names, and installation IDs in grant metadata. - Show the GitHub username, repository list, management link, and refresh action. - Discard the repository snapshot after newer installation lifecycle events. Preserve snapshots verified after delayed events. - Lock and re-read grant metadata when applying installation events or saving refreshed access. Patch only webhook fields for other events. Reject snapshots if access changed during the external fetch, using unique access revisions even when timestamps collide. - Show repository installation recovery for managed OAuth even when the app also offers an advanced PAT method. - Update tests and the GitHub connection runbook. No SQL migration is required. ## Verification - Local typecheck, build, and token gates passed. All latest-head CI gates passed, including the complete test matrix and browser suites. Greptile is 5/5 with no unresolved findings. - All 382 focused setup, permissions, metadata, service, and webhook tests passed across final runs. One socket-hang-up test passed on rerun with the full service suite. Final service, metadata, and webhook checks passed all 230 tests. - The broad local suite was stopped after failures. Seven workspace-runtime exposure and control-conflict failures reproduce on base commit `54a99d884`. The broad run also overlapped local iteration; final focused tests and clean-checkout CI are tracked separately. - Browser: a fresh production-backed instance completed enrollment, retried after profile enablement, reached GitHub consent, recovered from a missing installation, and completed OAuth. - Browser: the permissions page showed the authenticated username and the selected private test repository. A real `get_me` call returned the same account. Reading the selected repository passed; reading an unselected private repository failed with 404. - Browser: a second fresh instance completed enrollment and OAuth without a PAT form or unavailable state. Its username and repository list survived reload and refresh. A real get_me call on the final code returned the displayed account. ## Risks - Repository names are now stored in company-scoped grant metadata and shown with that credential. They are display data, not authorization data. - Large selections require more GitHub API calls. A failed later page rejects the refresh rather than reporting a partial list. - Older grants and webhook-invalidated snapshots require Refresh access to load the list. - Cloud profile enablement is separate deployment configuration. This PR does not change OAuth scopes or GitHub App permissions. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) via Codex. Reasoning, code execution, and browser tools were used. The exact context window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0ffc091473 |
feat(connections): add durable GitHub identities and webhooks (#12843)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents need source control access for repository work > - A shared token cannot preserve the responsible person's identity or an agent's dedicated identity > - GitHub App tokens also need durable refresh, repository access checks, and webhook delivery > - Paperclip already has managed connections, encrypted grants, run secret leases, and merge-confirmation behavior > - This pull request extends those systems with GitHub identities instead of adding a parallel credential system > - The benefit is durable GitHub access with explicit identity, repository, runtime, and webhook boundaries ## Linked Issues or Issue Description No public GitHub issue describes this connection change. This description follows the feature request template. **Subsystem affected** Connected Apps, connection grants, secret resolution, native Git runtime setup, webhook processing, and the Apps UI. **Problem or motivation** Users need to connect GitHub once and let agents use the correct GitHub identity. A run should use a dedicated agent account when one exists. Otherwise, it should use the responsible person's account. The connection must survive token expiry, repository access changes, and temporary instance downtime. **Proposed solution** Add user-owned and agent-owned GitHub grants to the existing connection model. Resolve one identity for MCP, Git, `gh`, health checks, and webhook bindings. Store provider tokens in the existing encrypted secret system. Refresh expiring token pairs under the existing lease and compare-and-swap path. Register signed Cloud webhook bindings and process normalized pull request and installation events through a durable local inbox. **Alternatives considered** An organization-wide GitHub token would lose person and agent attribution. Environment variables alone would bypass the managed connection and grant model. A new GitHub-only credential store would duplicate the existing secret and access systems. GitHub App installation tokens and private-key custody remain outside this first version. **Roadmap alignment** This change implements the Connected Apps direction. It also extends the shipped MCP Tool Gateway, per-agent secret access, and action-attribution systems. It does not add a repository catalog. The open repository catalog work in [#11234](https://github.com/paperclipai/paperclip/pull/11234) is related and complementary. ## What Changed - Added agent-owned connection grants and a per-agent credential policy with company and subject constraints. - Added a managed GitHub App method while keeping the personal access token method as an advanced fallback. - Added durable access-token and refresh-token handling with proactive rotation and one automatic recovery after a provider `401`. - Added GitHub identity and installation summaries without storing repository-name lists. - Added signed Cloud webhook binding, event lease, acknowledgement, local idempotency, pull request merge processing, and installation access handling. - Added one identity resolver for MCP, native Git, `gh`, checkout, health checks, and webhook bindings. - Added a class-3 run projection for `GH_TOKEN`, `GITHUB_TOKEN`, a `github.com`-only credential helper, SSH-to-HTTPS rewrite, and GitHub noreply commit attribution. - Added personal and dedicated-agent setup choices plus identity, repository, continuity, and webhook status in the Apps UI. - Added schema migrations, tests, and connection documentation. ## Verification - The current head is fully green in GitHub CI, including build, typecheck, all serialized/general server shards, all browser shards, policy, canary dry run, review, and security checks. - Live staging proof completed with a non-expiring GitHub App user token, selected-repository installation, repository add/remove refresh, managed MCP, native `gh`, HTTPS clone/push/delete, GitHub noreply commit attribution, signed merged-PR webhook acceptance, durable Cloud-to-instance delivery, and installation-access event processing. Temporary branches and temporary repository access were removed afterward. - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed before and after the rebase onto `origin/master`. - `pnpm build` passed. - The focused connector suite passed 285 tests after the rebase. - The full stable suite passed 5,790 tests and failed 22 tests across 8 general server files. The failures reproduced as shared-runner environment issues. They included `/tmp` versus `/private/tmp`, closed database connections, and invalid high ephemeral ports. The focused connection tests pass in isolation. ## Risks - Migrations add agent grant subjects and a durable connection-event inbox. Migration numbering and safety checks pass. - A raw GitHub user token enters the agent process for Git and `gh`. Per-tool Ask-first controls cannot limit those shell operations. The UI warns users about this boundary. - GitHub App user tokens can be non-expiring. Paperclip performs a continuity check every 30 days, but provider revocation still requires a reconnect. - The webhook path accepts only signed and bounded payloads. It stores a minimal normalized record and no raw provider payload. - GitHub repository permissions remain authoritative. Removed access can make a cached repository count temporarily stale, but runtime access fails immediately. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`, extended reasoning, tool use, code execution, browser control, and multi-file repository editing. The context window size was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7b094724e6 |
fix(runner): recover native sessions across restarts (#12845)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner keeps durable run and provider state outside one server process. > - A server restart can leave that runner alive or can interrupt it after a provider checkpoint. > - The old startup path used handoff intent and PID evidence, but it did not reconstruct native ownership. > - That gap could block the issue, create a replacement run, or start duplicate provider work. > - This pull request adds durable same-run recovery for coordinated and uncoordinated restarts. > - The benefit is exact recovery of the run, runner, session, provider, steering, and finalization state. ## Linked Issues or Issue Description Refs #9628. That pull request added earlier local-adapter hot-restart work. This change adds native PRP authority reconstruction and same-run provider resume. Refs #10935. That pull request handles missing hot-restart snapshots. This change also supports hard restarts with no snapshot. Refs #11624. That pull request prevents unsafe retry after an adopted legacy process exits. This change reconciles native terminal evidence before provider recovery. Refs #12070. That pull request improves process liveness checks. This change also binds recovery to a process-start fingerprint and fails closed on ambiguity. **What happened?** The server could record hot-restart intent, but startup did not rebuild native runner ownership. A live runner could not re-register its PRP authority. A dead runner could not resume the exact native and provider session on the same heartbeat run. Generic recovery could then block the issue or create replacement work. **Expected behavior** A live native runner must reconnect with the same PID and logical identities. A dead runner must resume the same durable session and heartbeat run with only a new operating-system PID. A proposed or terminal result must finalize once before any provider turn starts. Ambiguous process or session evidence must stay blocked without a signal or duplicate spawn. **Steps to reproduce** 1. Start a Paperclip Runner heartbeat and wait for an active provider turn. 2. Restart only the Paperclip server, with or without a hot-restart marker. 3. Observe that the old startup path does not reconstruct the native control-plane authority. 4. Kill both the server and runner after a provider checkpoint. 5. Observe that the old path cannot resume the exact native session on the original heartbeat run. **Paperclip version or commit** The defect was reproduced from commit `1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto the current `master`. **Deployment mode** Local development and self-hosted server deployments that use the local Paperclip Runner. ## What Changed - Added correlated hot-restart requests and version-compatible native handoff fields. - Added controller boot identity, process-start identity, controller generation, recovery state, request id, and bounded history to the native finalization ledger. - Added transactional recovery claims for live-runner reattach, dead-runner resume, and incomplete bootstrap. - Added fail-closed ownership takeover rules and process identity validation. - Added live runner adoption to the local runner transport without a duplicate spawn. - Added same-run provider checkpoint resume and legacy retry-row compatibility. - Reconciled proposed and terminal results before runner or provider recovery. - Bound the HTTP and PRP listener before startup recovery and delayed scheduling and generic reapers until classification completes. - Added restart-aware health diagnostics, run-log recovery transitions, durable runner diagnostics, and bounded shutdown finalizer draining. - Moved restart-survivable diagnostics into runner-owned, pre-redacted bounded writes; raw stdout and stderr are never persisted. - Added process-start fencing for controller, runner, and provider PIDs; startup classifies every candidate without an implicit cap. - Added crash-recoverable, contention-safe development restart-request coordination and failed-startup listener cleanup. - Added a credential-free real-process restart suite for eight restart, scale, and identity scenarios. - Documented native restart operation, persistence, diagnostics, and verification. ## Verification - The documented native restart commands passed. They ran eight real-process/database recovery scenarios and the live runner adoption transport test. - Native executor tests passed: 111 tests. - Heartbeat recovery tests passed: 124 tests. - Hot restart, health, and shutdown tests passed: 52 tests. - The broader affected server suite passed: 350 tests. - Focused native recovery and startup tests passed: 49 tests. - Runner transport and control-plane tests passed: 63 tests. - Runner-owned diagnostic tests passed for write-time bounding, credential redaction, private file modes, and raw stream non-persistence. - Development restart coordination tests passed: 11 tests. - Database migration checks and the partial-application/replay regression test passed. - Server, database, and Paperclip Runner typechecks passed. - `git diff --check` passed. - Full Paperclip PR CI passed, including build, canary, all five general server shards, all five serialized server shards, all three browser E2E shards, workspace suites, and release-registry verification. - Greptile completed at 5/5 with no outstanding findings, recommendations, follow-ups, or open review threads. ## Risks - Moderate risk. This changes startup ordering and ownership transfer for active native runs. - The migration adds nullable columns and does not rewrite existing rows. - Recovery fails closed when process or durable session identity is incomplete or contradictory. - The first implementation supports the local Paperclip Runner. Remote targets keep their existing behavior. - The real-process suite covers cleanup and asserts that no runner or provider process survives each test. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime did not expose a more specific model revision or context-window size. Repository editing, shell execution, database tests, and real-process test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a0028d7e1b |
chore(deps-dev): bump vitest from 4.1.10 to 4.1.11 (#12262)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.10 to 4.1.11. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v4.1.11</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Revive global concurrency limit for test lifecycle [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> and <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10992">vitest-dev/vitest#10992</a> <a href="https://github.com/vitest-dev/vitest/commit/5146df80b"><!-- raw HTML omitted -->(5146d)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Encode iframeId in tester iframe URL [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a>, <strong>Pduhard</strong> and <strong>Claude Opus 4.8</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10955">vitest-dev/vitest#10955</a> <a href="https://github.com/vitest-dev/vitest/commit/10b2cd201"><!-- raw HTML omitted -->(10b2c)<!-- raw HTML omitted --></a></li> <li>Trigger playwright/chromium gc on lower disk availability [backport to v4] - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>OpenCode</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10951">vitest-dev/vitest#10951</a> <a href="https://github.com/vitest-dev/vitest/commit/9851dbc41"><!-- raw HTML omitted -->(9851d)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>mocker</strong>: <ul> <li>Restrict redirect mocks to the fs allowlist [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10974">vitest-dev/vitest#10974</a> <a href="https://github.com/vitest-dev/vitest/commit/fe5a11d3c"><!-- raw HTML omitted -->(fe5a1)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v4.1.10...v4.1.11">View changes on GitHub</a></h5> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/9bd8d464e6328c567c2dbcd8fdd977d57a9425c2"><code>9bd8d46</code></a> chore: release v4.1.11 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10995">#10995</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/9851dbc41c286a30abfb6b29cce65f3e5b7b40a1"><code>9851dbc</code></a> fix(browser): trigger playwright/chromium gc on lower disk availability [back...</li> <li>See full diff in <a href="https://github.com/vitest-dev/vitest/commits/v4.1.11/packages/vitest">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
fdf8c8464d |
feat(runner): add managed provider backends (#12699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner provides durable, provider-neutral agent execution. > - The current stack supports qualified local providers but omits the managed provider paths from the integration branch. > - Claude Managed Agents and AWS AgentCore need explicit profile qualification, durable recovery, usage accounting, and cleanup controls. > - This pull request adds those managed backends as the third part of the Runner parity stack. > - The benefit is managed execution without weakening the default-off Runner rollout gate. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: Runner, server orchestration, database profiles, CLI, and adapter configuration UI. **Problem or motivation** The current Runner stack cannot select or execute the managed Claude Agents API or AWS Bedrock AgentCore Harness backends. It also lacks qualified profile storage and recovery checks for those remote resources. **Proposed solution** Add qualified managed and remote profiles, API and CLI management, exact provider selection, durable lifecycle handling, cumulative usage accounting, bounded cleanup, and retention acknowledgement. Keep `enableNativeRunner` default-off. **Alternatives considered** A direct copy of the old integration branch was rejected because its provider contracts, model values, credential flow, and migration history no longer match the current base. A single large parity pull request was also rejected because stacked review keeps each subsystem bounded. **Roadmap alignment** This continues the existing Runner architecture and rollout work. It does not introduce a separate execution system. **Additional context** This pull request is based on the merged #12691 and #12685 stack. It also closes the delayed security-review findings reported on #12691 by binding qualified ACPX and OpenCode launch artifacts to the bytes actually executed. A GitHub search for managed agent, AgentCore, and Claude managed work found no duplicate public issue or pull request. ## What Changed - Add Claude Managed Agents and AWS AgentCore provider executors to runnerd. - Add qualified managed and remote profile storage, routes, OpenAPI contracts, CLI commands, and migration 0237. - Validate profile ownership, enabled state, exact qualified revision, model, agent version, and secret binding before persistence and recovery. - Persist durable provider session and owned skill state for restart-safe cleanup. - Reconcile uncertain create responses and delete remote sessions before owned skills. - Track cumulative provider usage and enforce positive session spend caps. - Recover interrupted AgentCore usage at the next turn boundary by charging the prior invocation ceiling exactly once; keep the session gated until an explicit monotonic budget raise. - Isolate AgentCore AWS configuration from host profiles and credential-process/SSO configuration while preserving workload identity. - Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact paths; remove the ambient executable override. - Snapshot and content-verify ACPX and OpenCode commands, scripts, and provider executables before launch. Linux executes sealed inherited descriptors; macOS uses authenticated private snapshots with retry-safe rematerialization at the spawn boundary. - Persist canonical ACPX and OpenCode launch-profile digests, reject drift across fresh recovery, and make recovery failures sticky. - Close and journal unsafe ACPX active-turn recovery before any provider bootstrap or reconnect. - Add managed provider fields to the Runner configuration UI and permission projection. - Preserve the default-off `enableNativeRunner` experimental flag. ## Verification - `pnpm -r typecheck` - `pnpm build` - Focused managed server, database, CLI, Runner TypeScript, Rust, Claude, AgentCore, ACPX, OpenCode, process-supervisor, and durable-recovery tests passed. - `cargo test -p paperclip-runner-core --lib --locked` (160 tests) - `cargo check --workspace --all-targets --locked` - Native Codex integration tests passed (60 tests); native provider tests passed (7 tests); server native-runtime tests passed (87 tests). - Verified-launch replacement, nested-spawn retry, exact-version, profile-drift, sticky-failure, and no-bootstrap active-recovery tests passed. - `git diff --check` - The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust workspace lockfile adds the approved `rustix` dependency used for safe descriptor handling while `#![forbid(unsafe_code)]` remains enabled. ## Risks - The provider APIs can change while they are in beta. Exact qualification and fail-closed recovery checks limit drift. - Remote cleanup can fail after a partial create. Durable ownership inventories and retry-safe deletion preserve recovery state. - Migration 0237 adds profile tables. The generated migration and snapshot pass the repository migration checks. - Managed execution can incur provider cost. Positive default spend caps and explicit retention acknowledgement limit accidental use. - An interrupted AgentCore invocation without final metadata is conservatively charged to its active session ceiling. This can overstate cost, but cannot undercount it; later work requires an explicit budget increase. - Linux qualified launches use sealed memory descriptors. macOS lacks executable-descriptor APIs, so the runner uses owner-only private snapshots and minimizes linked-path lifetime; hostile same-UID processes remain outside the documented local-host trust boundary. - The global Runner feature remains default-off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4b6de5327e |
Remove cheap model profiles (#12683)
## Thinking Path > - Paperclip manages agents that use different model providers and adapters. > - Paperclip must keep agent execution rules clear and predictable. > - The cheap-model profile added a second execution mode across adapters, task recovery, APIs, and the UI. > - That mode increased configuration and recovery complexity. > - This pull request removes the cheap-model profile as a product feature. > - The benefit is one model-selection path for normal work and recovery work. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change simplifies model selection across agent configuration, task execution, recovery, and adapter capabilities. **Current behavior** Paperclip exposes cheap-model profiles in adapter metadata, agent runtime configuration, task overrides, recovery rules, APIs, and the board UI. Recovery work can select a different model profile from the agent's configured model. **Proposed behavior** Paperclip uses the agent's configured model for normal work and recovery work. Status-only recovery stays limited to coordination work. The API rejects legacy model-profile configuration. A migration removes stored model-profile values from existing agent, issue, and historical revision records. **Reason and benefit** One model path reduces configuration, API, UI, and recovery complexity. It also prevents status recovery from becoming a separate product-level model-routing feature. **Breaking changes** This change removes model-profile fields and adapter capability metadata. Existing stored model-profile values are removed by an idempotent migration. The validators reject new legacy profile values with clear errors. ## What Changed - Removed model-profile types, adapter capabilities, API fields, and model selection logic. - Removed cheap-model controls from agent and task UI surfaces. - Kept status-only recovery limited to coordination context while normal continuations use the configured agent model. - Added an idempotent migration that removes stored model-profile values from agents, issues, and configuration revisions without changing issue update timestamps. - Updated tests and product documentation for the single-model behavior. ## Verification - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` completed with 5,607 passing tests and 8 environment-sensitive failures in unrelated fixed-port and database-deadlock suites. The same failures repeated in an isolated rerun. CI is the final clean-room result. ## Risks - This is an intentional breaking change for clients that send model-profile fields. - The migration changes legacy agent, issue, and configuration-revision JSON. It is idempotent and preserves unrelated fields and issue update timestamps. - The change is cross-cutting because the removed feature existed in adapters, shared contracts, the server, plugins, and the UI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. Reasoning and tool use were enabled. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
25cf079ec5 |
feat(runner): add Codex-native application integration (#12591)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package is useful only when the application can start, observe, and recover a native Codex run safely. > - Existing direct adapters must keep their current execution and finalization paths. > - The application boundary therefore needs additive persistence, authorization, coordination, and recovery behind an explicit experimental adapter. > - This pull request adds that Codex-only boundary without activating generalized providers, remote environments, or the later task/SDK surfaces. ## Linked Issues or Issue Description **Subsystem affected** Shared contracts, database persistence, adapter utilities, server native-runtime services, and the experimental Paperclip Runner adapter. **Problem or motivation** The already-landed runner package has a qualified Codex path, but the application needs durable native-run state, guarded runtime selection, authenticated coordination, tool security, finalization, and recovery before the experimental adapter can be exercised safely. **Proposed solution** Add a Codex-only `paperclip_runner` application path behind the existing default-off native-runner setting. Bind native state and coordination to company/run identity, preserve persisted-run recovery, and leave every direct adapter on its existing legacy execution path. **Alternatives considered** The earlier stack boundary introduced a generalized executor and remote-environment lifecycle here. That made this PR depend on implementations in higher PRs and changed reusable sandbox behavior globally. Those pieces are now deferred together to #12592. **Roadmap alignment** ROADMAP.md does not list a conflicting native-runner integration project. This change adds the application boundary for the existing Runner architecture. ## What Changed - Added native run/result/finalization/provider-trace persistence, shared validators, and idempotent migration/replay coverage. - Added guarded Codex-only runtime selection, authenticated PRP coordination, recovery, finalization, and interaction services. - Added run/company-bound tool-gateway authorization, credential redaction, SSRF protections, and replay-safe behavior. - Added the explicit `paperclip_runner` adapter behind the default-off rollout setting. - Preserved legacy answered-question wake projection and direct-adapter execution/finalization paths. - Hardened cancellation so only owned in-memory child processes are signaled; persisted recycled PIDs/process groups are never trusted. - Retained the narrow Claude ACPX isolated-context security follow-up discovered after #12590. - Deferred the generalized executor, provider ingress, remote lifecycle, SDK/lab/eval work, release-process changes, and lockfile. ## Verification - Changed-file delta against `master`: 133 files. - GitHub Actions is the authoritative verification environment for this PR. - Full CI, security, and Greptile review will run on this lowest unmerged stack PR. - Local tests/build/typecheck were not run because this checkout is resource constrained. - Static diff/reference checks pass, and `pnpm-lock.yaml` is unchanged. ## Risks - This touches central heartbeat and agent-route code, so legacy compatibility is the primary risk. - Runtime selection remains Codex-only and explicit; direct Codex, Claude, OpenCode, process, HTTP, and plugin adapters remain on their existing paths. - Fresh native starts fail closed while the rollout flag is off; persisted native records remain readable and recoverable. - Cancellation, company/run binding, tool calls, status decisions, and completion writes are guarded or replay-safe. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI and security gates are green - [ ] Greptile is 5/5 with no open actionable findings - [x] I will address all Greptile and reviewer comments before merge ## Stack - Position: 3 of 5 overall; lowest of 3 currently unmerged - Base: `master` - Previous: [#12590](https://github.com/paperclipai/paperclip/pull/12590), qualified Claude ACPX runtime — merged - Next: [#12592](https://github.com/paperclipai/paperclip/pull/12592), generalized Codex executor, task experience, and developer SDKs --------- Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
d387cc0ff0 |
feat(connections): add managed external MCP connectors (#12346)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connection intents need secure provider implementations to complete setup. > - Some providers use managed OAuth or external credential brokers. > - Those tokens must stay out of durable Paperclip state and fail closed when refresh fails. > - This pull request adds managed connector backends and the required storage contract. > - The benefit is safer provider setup with governed credential lifecycles. ## Linked Issues or Issue Description Refs #11965 This is stack 8 of 11. It depends on stack 7 and replaces another reviewable part of #11965. ## What Changed - Add managed Google Workspace and external connector backends. - Add Vercel Connect support without storing provider bearer tokens. - Add replay-safe migration 0232 and its generated snapshot. - Fail closed and clear stale token bindings when organization OAuth refresh needs reauthorization. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 194 tests passed. - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` - `pnpm exec vitest run --project @paperclipai/server server/src/services/remote-url-credentials.test.ts` (5 passed, including URL userinfo vault extraction) ## Risks - Broker metadata errors can block provider setup. - OAuth refresh failure disables the shared organization connection until reauthorization. - Migration 0232 is generated, ordered after 0231, and safe to replay. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the public source pull request with `Refs #` - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b3343dbd64 |
feat(connections): add self-serve intent runtime (#12345)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a governed way to request app connections during issue work. > - The catalog now describes the available providers and setup methods. > - A request must become a durable, company-scoped intent before an operator acts on it. > - This pull request adds that intent runtime across server, agent, CLI, and shared contracts. > - The benefit is a safe bridge from agent need to operator-approved setup. ## Linked Issues or Issue Description Refs #11965 This is stack 7 of 11. It depends on stack 6 and replaces another reviewable part of #11965. ## What Changed - Add connection intent types, validation, service logic, and routes. - Add agent runtime tools and CLI support for connection requests. - Add issue-thread interaction support for connection intents. - Add runtime, route, adapter, and contract tests. - Hold the final resolved-continuation row lock through asynchronous adapter preparation until an actual process spawn, so parking or reassignment cannot cross that boundary. - Report Hermes Gateway's first remote run request through the shared dispatch hook so the resolved-intent lock is released at the true dispatch boundary. - Revalidate the addressed user's live non-viewer membership and connection-management authority for every intent mutation, including OAuth completion. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 176 tests passed. - `pnpm build` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-stale-queue-invalidation.test.ts` (32 passed; includes non-process dispatch lock-release coverage) - `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/connection-intents-service.test.ts -t "addressed-user mutation"` (1 passed) - `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/tool-access-service.test.ts -t "binds OAuth callback completion to the initiating board session"` (1 passed) - `pnpm --filter @paperclipai/hermes-paperclip-adapter test -- src/gateway/server/execute.test.ts` (23 passed; includes dispatch-hook ordering and exactly-once coverage) - `pnpm --filter @paperclipai/hermes-paperclip-adapter typecheck` ## Risks - A malformed intent could create an unusable operator request. - Validators and company checks reject invalid or cross-company requests. - The final continuation gate holds the issue row lock through adapter preparation until process or remote dispatch; later operator changes use the normal active-run interruption path. - The change does not add a database migration. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the public source pull request with `Refs #` - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6244e4cf32 |
feat(apps): add Composio and Gmail connectors (#12342)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - App connections need both direct providers and managed provider hubs. > - The grant layer now defines safe credential ownership. > - Composio needs parent and child connection lifecycle rules, and Gmail needs governed setup. > - This pull request adds both connector families on the grant foundation. > - The benefit is broader app access without weakening credential isolation. ## Linked Issues or Issue Description Refs #11965 This is stack 4 of 11. It depends on stack 3 and replaces another reviewable part of #11965. ## What Changed - Add Composio parent and child connection support. - Add Gmail connection setup and governance. - Preserve credential paths and remove duplicate binding declarations. - Cascade Composio pause and restore actions to child connections. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 164 tests passed. - `pnpm build` ## Risks - Parent lifecycle changes can affect every Composio child. - The service restores only children whose provider accounts remain active. - Credential binding paths are normalized before secret resolution. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
20ccf3f476 |
feat(apps): add connection grants and delegated identities (#12341)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - External tools need explicit identity and access boundaries. > - Shared connection credentials cannot represent every user-scoped use case. > - Grants must stay company-scoped and support safe delegation. > - This pull request adds connection grants, identity rules, and their database contract. > - The benefit is durable control over which identity an agent may use. ## Linked Issues or Issue Description Refs #11965 This is stack 3 of 11. It depends on stack 2 and replaces another reviewable part of #11965. ## What Changed - Add company and user connection grants. - Add delegated identity and membership rules. - Synchronize database, shared, server, and UI contracts. - Register the grant-member replacement route in the OpenAPI surface in the same layer that mounts it. - Add migration 0231 with replay-safe guards and coverage. ## Verification - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/openapi-routes.test.ts` (5 passed) - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` ## Risks - Incorrect grant selection could expose the wrong credential scope. - The service enforces company and subject boundaries before credential use. - Migration 0231 is generated, ordered after 0230, and safe to replay. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked a public issue or pull request with `Refs #` - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cabc9146d0 |
feat(apps): add secure remote MCP and PostHog setup (#12339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give those agents governed access to external tools. > - Remote MCP setup needs secure endpoint validation and durable credentials. > - PostHog needs both browser sign-in and personal API key setup paths. > - This pull request adds the shared remote MCP foundation and the PostHog definition. > - The benefit is a secure and reusable base for later app connection work. ## Linked Issues or Issue Description Refs #11965 This is stack 1 of 11. It replaces the first reviewable part of #11965. ## What Changed - Add guarded remote MCP setup and credential handling. - Add PostHog OAuth and API key connection methods. - Add focused server, shared contract, and UI coverage. - Keep the migration replay-safe and idempotent. - Give the late-close security regression the same 10-second CI headroom as the adjacent real-timer handshake test. - Synchronize fake-timer handshake tests at the exact ensure-session boundary so real filesystem setup cannot race the fake deadline. - Drive PTY overflow coverage only after listener registration so scheduling cannot reorder the test fixture. ## Verification - pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts server/src/__tests__/plugin-worker-manager.test.ts (220 passed; affected cases also passed five focused stress repetitions) - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never leaks a sandbox-provided value from a late close rejection into logs or the result"` (1 passed) - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never promotes a late ensureSession resolution|closes a late-resolving real handle exactly once"` (2 passed) - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` ## Risks - Remote endpoint validation can reject configurations that previously passed without checks. - OAuth configuration errors can block setup until the operator corrects the provider settings. - The migration uses guarded statements so repeated execution is safe. - The test-only synchronization changes do not affect runtime behavior; they remove filesystem/fake-clock and listener-registration races observed under parallel CI load. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a3b78f9d34 |
fix(cli): make embedded-Postgres tests survive runner contention (#12466)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The CLI test suite uses embedded Postgres for worktree checks. > - Several tests start a cluster and run migrations on each test. > - Runner contention can make this cost class exceed small hand-written time budgets. > - A timed-out test can also leave the cluster alive while cleanup removes its data directory. > - This pull request gives the cost class one measured timeout and registers cleanup for the timeout path. > - The benefit is stable tests and fewer orphaned embedded Postgres processes. ## Linked Issues or Issue Description No public GitHub issue tracks this defect. Related PRs #10987 and #10961 address wider embedded-Postgres test setup issues. This pull request addresses the CLI timeout and cleanup path described below. **What happened?** Five CLI tests used different hand-written time budgets while they started embedded Postgres and ran migrations. Runner contention made one test exceed its budget. A timeout also skipped the `finally` cleanup path and left a Postgres process alive. **Expected behavior** Each embedded-Postgres test gets a budget that covers measured runner contention. Cleanup stops the cluster before the test removes its data directory, including after a timeout. **Steps to reproduce** 1. Run `npx vitest run cli/src/__tests__/worktree.test.ts` on a contended runner. 2. Observe the embedded-Postgres tests take longer than their small hand-written budgets. 3. Inspect the process list after a timeout and observe an orphaned Postgres process. **Paperclip version or commit** The defect reproduces on `master` before this pull request. **Deployment mode** Built from source with Vitest. **Installation method** Built from source with pnpm. ## What Changed - Export `EMBEDDED_POSTGRES_TEST_TIMEOUT_MS` from the embedded-Postgres test helper. - Apply the shared timeout to every test defined by `itEmbeddedPostgres`. - Replace the five hand-written timeout values in `worktree.test.ts`. - Register temporary-directory removal and cluster stop with `onTestFinished` in the correct order. - Restore the working directory during timeout cleanup. ## Verification - `npx vitest run cli/src/__tests__/worktree.test.ts` passes 63 tests. - `npx vitest run cli/src` passes 424 tests across 59 files. - `tsc --noEmit` in `cli/` adds no new error in the touched files. - CI must pass on the pull request head. - Greptile must report 5/5 with no unresolved blocking thread. ## Risks - Low risk. The change affects test helpers and test cleanup only. - The shared budget can lengthen a failing test before Vitest reports the failure. - The cleanup order depends on Vitest callback order, which the tests now use explicitly. ## Model Used OpenAI Codex, GPT-5. The model used tool calls and code review support. The context window size and internal reasoning mode were not exposed in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8316ceb0b9 |
Add the Better Auth issuer column so signup and sign-in work (#12396)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A self-hosted install in `authenticated` mode signs users in with Better Auth, mounted at `/api/auth` over a hand-written Drizzle `account` table in `packages/db` > - Better Auth 1.7.0 added a required `issuer` field to that `account` model, plus a unique index on `(issuer, accountId)` > - The dependency bump in #11886 changed only `server/package.json` and the lockfile, so the Drizzle table never grew the column > - The Drizzle adapter checks the model against the schema on every write, so `linkAccount` throws and sign-up answers 500 with an empty body; a fresh install cannot create its first user, and an upgraded install locks out every existing user > - This pull request adds the `issuer` column and its unique index, and migrates the column in with a backfill that covers every existing row > - The benefit is that sign-up and sign-in work again, on a new install and after an upgrade ## Linked Issues or Issue Description No existing issue. Describing it inline, following `.github/ISSUE_TEMPLATE/bug_report.yml`. Refs #11886 (the dependency bump that introduced the required field). Refs #12269 (an earlier attempt at this fix; its backfill covers only `provider_id = 'credential'`). **What happened?** Sign-up fails on a self-hosted install. `POST /api/auth/sign-up/email` answers HTTP 500 with a zero-byte body. The server log carries: ``` [Better Auth]: The field "issuer" does not exist in the "account" Drizzle schema. # SERVER_ERROR: [BetterAuthError: The field "issuer" does not exist in the "account" Drizzle schema.] ``` The request writes the `user` row and then fails on the `account` row. The address is stuck after that: a second sign-up answers 422 `USER_ALREADY_EXISTS`, sign-in answers 401, and password reset answers 400 `RESET_PASSWORD_DISABLED` because the account that would hold the password does not exist. An upgraded install is worse. `sign-in/email` matches the credential account on `account.issuer === 'local:credential'`. Rows written before the upgrade have no issuer, so every existing user is locked out. **Expected behavior** `POST /api/auth/sign-up/email` answers 2xx and writes both the `user` row and its credential `account` row. `POST /api/auth/sign-in/email` then answers 2xx and sets a session cookie. An install that upgrades keeps its existing users. **Steps to reproduce** 1. Start a server from `master` with `PAPERCLIP_DEPLOYMENT_MODE=authenticated` against an empty database. 2. `curl -X POST http://127.0.0.1:<port>/api/auth/sign-up/email -H 'Content-Type: application/json' -H 'Origin: http://127.0.0.1:<port>' --data '{"name":"A","email":"a@example.com","password":"a-long-password"}'` 3. The response is HTTP 500 with an empty body. **Paperclip version or commit** `master` at |
||
|
|
036600d922 |
fix(db): close test database clients before the embedded Postgres cluster stops (#12335)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses database clients and embedded PostgreSQL test fixtures > - A fixture stopped its embedded PostgreSQL cluster while clients still held connections > - The postgres.js driver then scheduled a write on a stopped connection > - That write escaped the timer callback and caused a test process to exit with an error > - This pull request closes registered clients before the fixture stops its cluster > - The benefit is stable test teardown and clear failure reporting in continuous integration ## Linked Issues or Issue Description Refs: #10869 **What happened?** An embedded PostgreSQL test fixture stopped its cluster while database clients still held open connections. The postgres.js driver then scheduled a deferred write on a dead connection. The write caused an unhandled error after the test shard reported success. **Expected behavior** The fixture closes all live clients for its cluster before it stops the embedded PostgreSQL cluster. Tests then finish without a deferred write on a dead connection. **Steps to reproduce** 1. Run the database regression test with the embedded PostgreSQL fixture. 2. Stop the fixture while its database client still has an open connection. 3. Observe the deferred write and the process exit status. **Paperclip version or commit** Branch base: |
||
|
|
bdd8f1bedb |
Repair the drizzle snapshot so generate emits no spurious migration (#12333)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip keeps its state in PostgreSQL, and `packages/db` owns that schema through Drizzle > - `drizzle-kit generate` writes a new migration by diffing `packages/db/src/schema/` against the newest snapshot in `packages/db/src/migrations/meta/`, so the snapshot must describe the schema that the migrations produce > - Snapshot `0228` recorded the new `error_count` column on the wrong table, and snapshot `0229` inherited the error, so the newest snapshot no longer matched the schema > - Because of that, `generate` on `master` folded the drift into any new migration: it emitted an `ADD COLUMN` for a column that migration `0228` already creates, which fails on a fresh database, plus an out-of-scope `DROP COLUMN` > - This pull request moves the column entry to the correct table in both snapshots and adds a test that repeats the diff `generate` performs > - The benefit is that the next person who generates a migration gets only their own change, and CI fails if the snapshot drifts again ## Linked Issues or Issue Description No existing issue. The description below follows `.github/ISSUE_TEMPLATE/bug_report.yml`. **What happened?** `drizzle-kit generate` on `master` emits a wrong migration. The newest snapshot, `packages/db/src/migrations/meta/0229_snapshot.json`, disagrees with the schema in two places. It omits `issue_question_response_deliveries.error_count`, which `0228_nasty_grim_reaper.sql` creates. It also carries `decision_archive_notification_outbox.error_count`, which no migration ever creates and the Drizzle schema never declared. Snapshot `0228` introduced both halves of the error: it added the new `error_count` column to `decision_archive_notification_outbox` instead of the table that the same migration creates. Snapshot `0229` copied it forward. Any new migration therefore starts with two statements that do not belong to it: ```sql ALTER TABLE "issue_question_response_deliveries" ADD COLUMN "error_count" integer DEFAULT 0 NOT NULL; ALTER TABLE "decision_archive_notification_outbox" DROP COLUMN "error_count"; ``` The `ADD COLUMN` fails on a fresh database, because migration `0228` already creates that column. The `DROP COLUMN` targets a column that does not exist on any deployment. **Expected behavior** `drizzle-kit generate` reports "No schema changes, nothing to migrate" on a clean checkout of `master`, and a new migration contains only the author's own schema change. **Steps to reproduce** 1. Check out `master` at commit `bc1a21564`. 2. Run `pnpm install`. 3. Run `pnpm --filter @paperclipai/db generate`. 4. Read the emitted `packages/db/src/migrations/0230_*.sql`. It contains the two statements above, and no schema file was changed. **Paperclip version or commit** `master` at `bc1a21564`. The drift entered in #12307 (snapshot `0228`) and was carried forward by #12291 (snapshot `0229`), which worked around it by building its snapshot by hand. **Deployment mode** Not deployment specific. It affects anyone who generates a migration, and it affects any fresh database that would later run the bad migration. **Database mode** All PostgreSQL modes: embedded, local Docker, and hosted. ## What Changed - Moved the `error_count` column entry from `decision_archive_notification_outbox` to `issue_question_response_deliveries` in `packages/db/src/migrations/meta/0228_snapshot.json` and `packages/db/src/migrations/meta/0229_snapshot.json`. Both files keep their `id` and `prevId`, so the snapshot chain is unchanged. - Added `packages/db/src/migration-snapshot-drift.test.ts`. It reads the newest snapshot named by `_journal.json`, serializes the schema modules with `generateDrizzleJson`, and asserts that `generateMigration` returns no statements. This is the same diff that `generate` performs. - Documented the snapshot rule in `doc/DATABASE.md` under a new "Migration snapshots" section. No migration SQL was added, renumbered, or edited. No schema file changed. The database is correct as it is; only the snapshot was wrong. Why both snapshots and not only the newest one: `0229` is the file that `generate` reads, so repairing it is what fixes the bug. `0228` holds the same error, and `drizzle-kit drop` removes the last migration and its snapshot, which would promote `0228` back to newest and bring the drift back. Repairing both removes that trap. Snapshots are never applied to a database, so neither edit changes any deployment. ## Verification Commands run from the repository root. - `pnpm --filter @paperclipai/db generate` — "No schema changes, nothing to migrate 😴". It writes no SQL file, no snapshot, and no journal entry. `git status` stays clean. Before the fix, the same command wrote `0230_fast_caretaker.sql` with the two spurious statements. - The repaired `0229_snapshot.json` is byte-identical to the snapshot that a real `generate` run produced, except for the `id` and `prevId` that keep the chain intact. - Chain check: the repaired `0228` and `0229` snapshots now differ by exactly the two columns that `0229_drop_company_brand_color_and_attachment_max_bytes.sql` drops, `companies.brand_color` and `companies.attachment_max_bytes`, and by nothing else. - Database check: applied all 229 migrations in order to an embedded PostgreSQL, then compared the live schema with the repaired snapshot. 179 tables and 2687 columns match, with no missing column, no extra column, and no nullability difference. The same comparison against the pre-fix snapshot reports exactly two problems: `column only in snapshot: decision_archive_notification_outbox.error_count` and `column only in database: issue_question_response_deliveries.error_count`. This harness was a scratch script and is not part of the pull request. - `pnpm --filter @paperclipai/db typecheck` — pass. It runs `check:migrations`, which is `check-migration-numbering` and `check-migration-safety`. - `npx vitest run --root packages/db` — 28 files, 102 tests, all pass. This includes the new test. - New test, negative case: with the pre-fix `0229_snapshot.json` restored, `migration-snapshot-drift.test.ts` fails and prints exactly the two spurious statements, plus the instruction to run `generate`. It passes on the repaired snapshot. It takes about 1.2 seconds and needs no database. - `node scripts/check-forbidden-tokens.mjs` and `node scripts/check-no-git-push.mjs` — pass. ## Risks Low risk. A Drizzle snapshot is a build-time record for `drizzle-kit generate`. It is never applied to a database, so this change cannot alter any deployment, and no operator action is needed. Databases that already ran migrations `0228` and `0229` are correct today and stay correct. The proof is a clean `generate`: the command that produced the wrong migration now reports "No schema changes, nothing to migrate" and writes nothing. Two smaller notes: - The new test depends on `drizzle-kit/api`, which is already a dev dependency of `packages/db`. If a future `drizzle-kit` upgrade changes that surface, the test fails loudly at import rather than passing silently. - The test imports every module in `packages/db/src/schema/`, which is the same set that `drizzle.config.ts` points the CLI at. It deduplicates by object identity, because the barrel re-exports the same table objects and `drizzle-kit` rejects a table it sees twice. ## Model Used Claude (Anthropic), Claude Opus, agentic tool use via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bc1a21564f |
Remove the company brand color and per-company attachment limit (#12291)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - A company is the top-level container, and the company General page
holds its settings
> - Two of those settings did almost nothing: the brand color only
tinted the generated company icon, and the attachment size limit sat
under the deployment-level `PAPERCLIP_ATTACHMENT_MAX_BYTES` cap that
already bounded every upload
> - A setting that changes one icon hue, and a setting that can only
lower a limit the operator already set, are not worth the page space or
the code that carries them
> - This pull request deletes both settings from the UI, the validators,
the API contract, the server, and the database
> - With the deployment cap as the only limit left, the message a person
sees when an upload is rejected has to name that limit in terms they can
act on, so the raw byte count becomes a human-readable size
> - The benefit is a shorter company General page for every deployment,
one attachment limit instead of two, and less code between an upload and
its ceiling
## Linked Issues or Issue Description
No existing issue. The description below follows
`.github/ISSUE_TEMPLATE/enhancement.yml`.
**What existing behavior does this improve?**
The company General page (`/company/settings`), the `PATCH
/api/companies/{companyId}` and `PATCH
/api/companies/{companyId}/branding` request contracts, and the
attachment upload limit on task, case, and company-import uploads.
**Subsystem affected**
Cross-cutting: `ui/`, `server/`, `packages/shared`, `packages/db`.
**Current behavior**
The company General page shows an "Appearance" section with three
controls: Logo, Brand color, and Attachment size limit. The brand color
is a hex value that feeds one thing — the hue of the generated company
pattern icon. Companies that never set one already get a hue derived
from the company name. The attachment size limit is a per-company byte
count stored on `companies.attachment_max_bytes`. Every upload path
clamps it against the deployment-level `PAPERCLIP_ATTACHMENT_MAX_BYTES`
cap, so the per-company value can only lower a limit the operator
already chose.
**Proposed behavior**
The Appearance section keeps the Logo control only. The company pattern
icon always derives its hue from the company name. Every attachment path
reads the deployment cap directly, so `PAPERCLIP_ATTACHMENT_MAX_BYTES`
is the single limit. An upload rejected by that limit says so in human
units — "File is larger than the 10 MB limit" rather than a raw byte
count. The `companies.brand_color` and `companies.attachment_max_bytes`
columns are dropped, and both fields leave the company API contract.
**Reason and benefit**
Both settings ask an operator to make a decision that changes almost
nothing. The brand color moves one icon hue on a page that also lets you
upload a real logo, which overrides the icon entirely. The attachment
limit reads as a real control but cannot raise anything, so it is a
second place to look when an upload is rejected. Removing both shortens
the page every deployment sees, removes a company-scoped read from the
task attachment upload path, and leaves one attachment limit to reason
about instead of two.
**Breaking changes**
The company API responses no longer include `brandColor` or
`attachmentMaxBytes`, and `GET /api/invites/{token}` no longer includes
`companyBrandColor`. `PATCH /api/companies/{companyId}/branding` is
strict, so a request that sends `brandColor` now returns 400; the
non-strict `PATCH /api/companies/{companyId}` schema strips it. Company
packages exported by older versions still import: the portability
company manifest schema is non-strict, so the retired keys are stripped
and ignored rather than rejected. Companies that stored a brand color
lose it — their icon reverts to the name-derived hue that every company
without a color already used.
## What Changed
- Removed the "Brand color" and "Attachment size limit" fields from the
company General page, along with their state, dirty checks, save
payload, and Save-button gating.
- Removed `brandColor` and `attachmentMaxBytes` from
`createCompanySchema`, `updateCompanySchema`, and
`updateCompanyBrandingSchema`, and deleted the now-orphaned
`DEFAULT_COMPANY_ATTACHMENT_MAX_BYTES` and
`MAX_COMPANY_ATTACHMENT_MAX_BYTES` constants.
- Removed both fields from the `Company` type, the portability manifest
type and schema, and the `companiesApi.update` payload allowlist.
- Dropped `brandColor` from `CompanyPatternIcon` and its callers, so the
icon hue always comes from the company name. Deleted the now-unused
`hexToHue` helper and the now-unused `pickTextColorForSolidBg` export.
- Stopped emitting `brandColor` from the company service selection and
from the invite-summary and invite-branding payloads in
`server/src/routes/access.ts`.
- Replaced `normalizeIssueAttachmentMaxBytes` with the deployment cap:
task attachments, case attachments, and company import now use
`MAX_ATTACHMENT_BYTES` directly. The helper is deleted.
- Added `formatAttachmentSize()` next to `MAX_ATTACHMENT_BYTES` and
routed every over-limit message through it, so a rejected upload names
the limit in human units instead of raw bytes: `Image exceeds 10485760
bytes` becomes `Image is larger than the 10 MB limit`. Enforcement is
unchanged — the same single cap, the same multer limits, the same status
codes and response shapes.
- Added migration
`0229_drop_company_brand_color_and_attachment_max_bytes.sql` and removed
both columns from the Drizzle `companies` schema.
- Kept legacy imports working: the portability company manifest schema
is non-strict, so older packages carrying the retired keys still import
with the keys ignored.
- Updated the skill API reference and the implementation spec, and
pruned the token-extraction allowlist entries that the removed code made
stale.
## Verification
Commands run from the repository root:
- `pnpm --filter @paperclipai/shared typecheck` — pass
- `pnpm --filter @paperclipai/db typecheck` — pass (includes
`check:migrations`, which validates the new migration number and journal
entry)
- `pnpm --filter @paperclipai/ui typecheck` — pass
- server typecheck via `node_modules/.bin/tsc --noEmit` in `server/` —
pass. `pnpm --filter @paperclipai/server typecheck` could not run
locally because it builds the Rust runner first and `cargo` is not
installed on this machine; the TypeScript step it wraps is the command
above.
- `npx vitest run packages/shared/src/validators/company.test.ts` — 6
passed
- `npx vitest run server/src/__tests__/company-portability.test.ts` — 90
passed
- `npx vitest run server/src/__tests__/attachment-types.test.ts
server/src/__tests__/assets.test.ts
server/src/__tests__/issue-attachment-routes.test.ts
server/src/__tests__/company-portability.test.ts
server/src/__tests__/cases-routes.test.ts` — 165 passed (the
human-readable limit messages)
- `npx vitest run server/src/__tests__/company-branding-route.test.ts
server/src/__tests__/issue-attachment-routes.test.ts
server/src/__tests__/invite-summary-route.test.ts
server/src/__tests__/openclaw-invite-prompt-route.test.ts
server/src/__tests__/companies-route-cross-company-authz.test.ts` — all
passed
- `npx vitest run cli/src/__tests__/company.test.ts
cli/src/__tests__/company-delete.test.ts` — 27 passed
- `npx vitest run` in `ui/` — 4425 passed, 1 pre-existing failure
unrelated to this change (`OnboardingWizard.test.tsx` "renders instead
of throwing when the browser denies storage access", which also fails on
`master`)
- `npx vitest run` in `server/` — see the note below
- `node scripts/check-token-gates.mjs` — no new violations; the only
reported violations are the pre-existing `PillGuy.tsx` ones present on
`master`
New tests added:
- `packages/shared/src/validators/company.test.ts` — the create and
update schemas strip the retired keys, the strict branding schema
rejects `brandColor`, and the portability manifest schema accepts a
legacy entry carrying both keys and drops them.
- `server/src/__tests__/company-branding-route.test.ts` — `PATCH
/api/companies/{companyId}/branding` returns 400 for `brandColor` and
does not call the company service.
- `server/src/__tests__/company-portability.test.ts` — a legacy package
that declares `brandColor` and `attachmentMaxBytes` imports
successfully, and neither key reaches `companies.create`.
- `server/src/__tests__/issue-attachment-routes.test.ts` — the effective
task attachment limit is the deployment cap, and the route no longer
loads the company to size an upload.
- `server/src/__tests__/attachment-types.test.ts` —
`formatAttachmentSize()` renders the default cap as `10 MB`, keeps one
decimal place for fractional sizes and drops a trailing `.0`, falls back
to KB and bytes for small caps, steps up to GB, and never emits `NaN`
for a degenerate input.
- `server/src/__tests__/assets.test.ts` — the asset-image and
company-logo routes both return the human-readable limit message on an
over-cap upload.
## Merge with master
`master` moved while this was open, and the merge needed two
resolutions:
- **`ui/src/pages/CompanySettings.tsx`.** #12243 reworded the
user-facing
copy from "company" to "organization", and that rewording landed inside
the "Brand color" and "Attachment size limit" hints — the two fields
this change deletes. Both fields are removed, so the conflicted block is
dropped whole. The Logo field and every other copy change from #12243
are
kept.
- **Migration renumbered 0228 -> 0229.** #12307 landed
`0228_nasty_grim_reaper`, so this migration is now
`0229_drop_company_brand_color_and_attachment_max_bytes`. Its snapshot
is
rebuilt from master's `0228_snapshot.json` with only the two `companies`
columns removed, and `meta/_journal.json` is master's journal plus a
single `idx: 229` entry. `pnpm --filter @paperclipai/db
check:migrations`
passes.
The snapshot was rebuilt by hand rather than taken from `drizzle-kit
generate`, because master's `0228_snapshot.json` has drifted from
master's
own schema: `issue_question_response_deliveries.error_count` is created
by
master's 0228 SQL but missing from its snapshot, and the snapshot still
carries `decision_archive_notification_outbox.error_count`. Regenerating
folds both into this migration, and the resulting `ADD COLUMN
error_count`
would fail on a fresh database where master's 0228 already created that
column. Rebuilding from master's snapshot leaves that drift exactly
where
it is and keeps this migration to the two column drops. The drift is
pre-existing on master and is not addressed here.
## Risks
- **The migration is a destructive column drop.**
`0229_drop_company_brand_color_and_attachment_max_bytes.sql` removes
`companies.brand_color` and `companies.attachment_max_bytes`. It is safe
because both features are removed in the same change and nothing reads
either column after it. The statements use `DROP COLUMN IF EXISTS`,
matching the convention of the recent drop migrations in this
repository. The drop is not reversible: a downgrade after this migration
loses any stored values.
- **Stored brand colors are lost.** A company that had set a color now
renders the name-derived icon hue that every company without a color
already used. No other surface changes, and an uploaded logo still
overrides the icon.
- **API response shape narrows.** `brandColor` and `attachmentMaxBytes`
leave the company payloads, and `companyBrandColor` leaves the invite
summary payload. A client reading those fields now sees `undefined`. The
bundled UI and CLI are updated in this change.
- **Legacy imports are covered.** Packages exported by older versions
still carry both keys. The manifest schema is non-strict, so the keys
are stripped rather than rejected, and a test locks that in.
- **The over-limit message strings changed.** Anything matching on the
old `... exceeds N bytes` text — a test, a script, or a client that
string-matches `body.error` — needs updating. The status codes (422) and
response shapes are unchanged, so structured clients are unaffected.
- **Attachment limits can only widen.** A deployment that had lowered a
company below the deployment cap now allows uploads up to the cap for
that company. Lower `PAPERCLIP_ATTACHMENT_MAX_BYTES` if a smaller
ceiling is needed.
- **Storybook visual baselines shift** for the `CompanyPatternIcon`
matrix story, because those fixtures had brand colors. That workflow
runs only on a PR labeled `storybook-visual`, so it does not gate this
PR; regenerate the baselines if the label is added.
## Model Used
Claude (Anthropic), Claude Opus, agentic tool use via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
67f9867bc6 |
fix(interactions): deliver question answers durably (#12307)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can pause a task and ask the user structured questions. > - The answer is durable in the issue interaction, but delivery to the next run is not durable. > - A process restart can therefore leave an answered interaction without a continuation attempt. > - Native runners also need a provider-neutral question contract before the task page can consume native events safely. > - This pull request adds a content-free delivery outbox and an optional native steering seam. > - Direct adapters keep their existing heartbeat continuation path. > - The benefit is reliable answer delivery without changing runtime selection or task-page behavior. ## Linked Issues or Issue Description Refs #12202. This pull request replaces the question-delivery foundation from that stale task-thread pull request. The task-thread projection will follow in a smaller pull request. **What happened?** Question answers were stored in the issue interaction. The server then made one in-memory continuation wake. A server stop between those operations could leave the answer stored but not delivered. The combined native task-thread pull request also made this behavior hard to review separately from UI changes. **Expected behavior** The answer and its delivery receipt must commit in one transaction. The server must retry pending receipts after a restart. Existing direct adapters must keep the current wake path. A native runtime may use the optional steering seam, but this pull request does not enable native steering in production. **Steps to reproduce** 1. Create an `ask_user_questions` interaction. 2. Answer the interaction. 3. Stop the server before the continuation wake completes. 4. Start the server again. 5. On current master, no durable record tells the server to retry the answer delivery. **Paperclip version or commit** Current `master` at `4d82f5eae`. ## What Changed - Add the `issue_question_response_deliveries` table and migration. - Store only routing state, a correlation ID, and a payload digest in the delivery row. The answer remains in the existing interaction result. - Commit an answered interaction and its pending delivery row in one transaction. - Add bounded claims, retry recovery, cumulative terminal state, and content-free activity records. - Keep every built-in direct adapter and external adapter on the existing heartbeat wake path. - Add an optional native steering seam. No production caller supplies that seam in this pull request. - Retain the provider-neutral `paperclip.question_set.v1` presentation on recovered interactions. - Run delivery immediately after an answer and sweep pending rows at startup and on the existing server interval. - Add focused database, service, route, startup, adapter-matrix, digest, and duplicate-delivery tests. ## Compatibility Boundary - This pull request does not change adapter selection. - This pull request does not start runnerd. - This pull request does not create native run records. - Direct adapters never call the native steering seam. - The existing interaction result stays authoritative for answer content. - The migration is additive and does not rewrite existing rows. - This pull request has no UI, dependency, workflow, package-manager, or lockfile changes. - The diff has 19 files. ## Verification - `pnpm exec vitest run server/src/__tests__/question-response-delivery.test.ts server/src/services/issue-thread-interactions.test.ts server/src/__tests__/issue-thread-interaction-routes.test.ts server/src/__tests__/server-startup-feedback-export.test.ts` — 4 files and 120 tests passed. - `pnpm -r typecheck` — passed for all applicable workspaces. This includes Cargo format and check, protocol drift checks, and migration safety. - `pnpm build` — passed. This includes the Rust release binary, server build, and UI production build. - `git diff --check` — passed. - Secret patterns were not present in the changed text files. - The repository token gates currently report violations from unchanged files on `master`. This pull request does not change those files. ## Risks The main risk is routing a direct-adapter answer into a native session. The service checks the persisted runtime mode, and the adapter matrix proves that all direct adapters use only the existing wake path. The new table is additive. It has foreign keys, unique correlation constraints, bounded attempts, and status checks. Activity records omit question and answer content. ## Model Used OpenAI Codex, GPT-5 family. The client does not expose the exact deployment ID or context window. Agentic reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes:` / `Refs:` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket ID or instance-derived details - [x] I have run the affected tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the new contracts and compatibility boundary - [x] I have considered and documented compatibility and security risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8ef39febd7 |
refactor(db): drop the vendored postgres teardown patch (reverts #12227) (#12234)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - #12227 vendored a pnpm patch of the `postgres` driver to stop a teardown race (`nextWrite` firing after the socket is nulled) from crashing the process and failing green CI shards > - Maintainer call: carrying a vendored driver patch is not worth it for a CI flake — the patch adds a maintenance obligation on every future driver upgrade > - The race is an upstream bug in `postgres@3.4.9`; the plan is to wait for an upstream release that fixes it and bump the dependency instead > - This pull request reverts #12227 in full: the patch file, its `package.json` registration, and the regression test that exercised the patched behavior > - The benefit is an unmodified dependency graph; the known flake signature returns and is retried when it bites ## Linked Issues or Issue Description Reverts #12227. **What existing behavior does this improve?** Dependency hygiene: `postgres@3.4.9` is consumed unmodified again, with no `pnpm.patchedDependencies` entry to re-evaluate on every driver upgrade. **Current behavior** The repo carries `patches/postgres@3.4.9.patch` (null-socket guard in the driver's deferred write flush, plus an `execute()` refusal on socketless connections) and a regression test for it. **Proposed behavior** Plain upstream `postgres@3.4.9`. The teardown race stays an upstream bug: a green test shard can occasionally fail with `Vitest caught 1 unhandled error` and `TypeError: Cannot read properties of null (reading 'write')` at `Immediate.nextWrite`; the remedy is retrying the shard until an upstream driver release fixes the race and we bump. **Reason and benefit** A vendored driver patch is a standing maintenance cost that outweighs the flake it suppressed. ## What Changed - Reverts #12227 (`6c7c0fd1f`) in full: removes `patches/postgres@3.4.9.patch`, its `pnpm.patchedDependencies` registration in root `package.json`, and `packages/db/src/postgres-driver-teardown.test.ts`. No lockfile involvement — the merged commit never touched `pnpm-lock.yaml` and the refresh bot had not yet recorded the patch. ## Verification - `pnpm install` on the reverted tree is coherent; the full `packages/db` suite passes (26 files / 100 tests). - `git revert` applied cleanly with no conflicts. ## Risks - Low. This restores the exact pre-#12227 state. The known flake signature returns; it fails jobs whose tests all passed and is cleared by retrying the shard. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic coding session with tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
6c7c0fd1f2 |
fix(db): stop the postgres driver from crashing the process on a write/close race (#12227)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The server and its test suites talk to PostgreSQL through the
`postgres` (postgres.js) driver, and tests routinely tear their
databases down while connections still carry traffic
> - The driver flushes small buffered frames from a `setImmediate`, and
that deferred flush calls `socket.write()` without checking that the
socket still exists; a reserved connection whose backend died keeps
accepting queries, so the flush can fire with a null socket
> - The resulting `TypeError` escapes from a timer callback with no
try/catch above it, crashing the process — in CI this fails suites whose
tests all passed ("Vitest caught 1 unhandled error"), and the same crash
is reported against the driver in the wild after ECONNRESET
> - The latest driver release (3.4.9) still has the bug, so this pull
request adds a pnpm patch guarding the flush and normalizing timer state
on close, plus a deterministic regression test
> - The benefit is CI that no longer fails randomly on a teardown race,
and production processes that survive a database connection dying at the
wrong moment
## Linked Issues or Issue Description
No public issue exists; the underlying problem follows the bug-report
template.
**What happened?**
CI jobs fail with all tests passing: vitest reports `Vitest caught 1
unhandled error during the test run` with `TypeError: Cannot read
properties of null (reading 'write')` at `postgres/src/connection.js`
`Immediate.nextWrite`. The attribution points at whichever test file
happened to be running (e.g. `native-codex-runner.integration.test.ts`),
because the throw comes from a process-level timer callback, not from a
test. The identical crash is reported against the upstream driver by
other projects after `ECONNRESET` (e.g. immich-app/immich#25098).
**Expected behavior**
A connection dying between a write being scheduled and its deferred
flush must settle the affected queries through the driver's normal
connection-error path, never throw from a bare timer callback.
**Steps to reproduce**
Run the new `packages/db/src/postgres-driver-teardown.test.ts` with the
patch removed: reserve a connection (`sql.reserve()` — the same surface
`sql.begin()` uses), destroy the backend socket, wait for the client to
process the close, then issue one query on the reserved connection. The
deferred flush fires one tick later with `socket === null` and crashes
the process with exactly the CI signature.
**Paperclip version or commit**
master `198fc8b28`, `postgres@3.4.9` (latest release; bug still present
on the driver's master branch).
## What Changed
- `patches/postgres@3.4.9.patch` (new, wired via
`pnpm.patchedDependencies`): `nextWrite` returns without writing when
`socket === null`, dropping the buffered bytes — the close path has
already settled every in-flight query, so those bytes have nowhere to
go. The `closed()` and `terminate()` handlers additionally reset
`nextWriteTimer`/`chunk` after `clearImmediate`, so a stale cleared
handle cannot silently block a future reconnect's first flush. All three
shipped builds (`src`, `cjs`, `cf`) get the identical change.
- `packages/db/src/postgres-driver-teardown.test.ts` (new):
deterministic reproduction against a minimal in-process fake wire server
(startup auth + an empty result for the `fetch_types` bootstrap).
Asserts the late query settles with `CONNECTION_DESTROYED` through
`sql.end()` instead of crashing the process.
## Verification
- The regression test fails against unpatched `postgres@3.4.9` with the
exact CI signature (verified by running the same scenario against an
unpatched checkout) and passes with the patch.
- Full `packages/db` suite: 27 files / 101 tests pass.
- Spot-checked server suites that exercise the database through the
patched driver.
## Risks
- Low. The behavioral change activates only in a state that previously
crashed the process (write flush with no socket). Dropping the buffered
bytes matches what the connection's close path already promised callers:
every in-flight query has been settled with a connection error.
- The timer/chunk reset in `closed()`/`terminate()` prevents a
theoretical stale-handle hang after reconnect; on the normal path both
were already reset by `nextWrite`.
- The patch pins to `postgres@3.4.9`; a future driver upgrade will
surface the patch for re-evaluation (pnpm fails loudly on version
mismatch), and the guard can be dropped if the fix lands upstream.
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic
coding session with tool use (driver source analysis, wire-protocol fake
server, local test execution).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
0cedb45df3 |
build(deps-dev): bump typescript from 5.9.3 to 7.0.2 (#11880)
Bumps [typescript](https://github.com/microsoft/TypeScript) from 5.9.3 to 7.0.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/microsoft/TypeScript/releases">typescript's releases</a>.</em></p> <blockquote> <h2>TypeScript 7.0.2</h2> <p><a href="https://devblogs.microsoft.com/typescript/announcing-typescript-7-0/">https://devblogs.microsoft.com/typescript/announcing-typescript-7-0/</a></p> <p>This tag was originally released at: <a href="https://github.com/microsoft/typescript-go/releases/tag/typescript%2Fv7.0.2">https://github.com/microsoft/typescript-go/releases/tag/typescript%2Fv7.0.2</a></p> <h2>TypeScript 6.0.3</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.2%22">fixed issues query for TypeScript 6.0.2 (Stable)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.3%22">fixed issues query for TypeScript 6.0.3 (Stable)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.2%22">fixed issues query for TypeScript 6.0.2 (Stable)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0.1 RC</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0-rc/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0 Beta</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0-beta/">release announcement</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22+is%3Aclosed+">fixed issues query for Typescript 6.0.0 (Beta)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/microsoft/TypeScript/commit/1e4744d68260a7cb91b62b12edc3f6a2187faaf1"><code>1e4744d</code></a> Merge branch 'main' into ts7-release</li> <li><a href="https://github.com/microsoft/TypeScript/commit/a5a219c3b5da0db4fa0ecf6c0b1f588c9af9c669"><code>a5a219c</code></a><code>microsoft/typescript-go#4558</code></li> <li><a href="https://github.com/microsoft/TypeScript/commit/ecfe30dce91368d52c9a49b6095bb0b673a238f8"><code>ecfe30d</code></a> Update status localization</li> <li><a href="https://github.com/microsoft/TypeScript/commit/5de25b5f8fec2ca35eadaed041f1f06d2e214895"><code>5de25b5</code></a> Hide executable name in TypeScript status</li> <li><a href="https://github.com/microsoft/TypeScript/commit/d7ce74a75da2b80e8201506a1599c06549432b93"><code>d7ce74a</code></a> Show bundled TypeScript version for packaged servers</li> <li><a href="https://github.com/microsoft/TypeScript/commit/29be66a607707f90d7a53103a4469bb3015a4d54"><code>29be66a</code></a> Correct TS 7 release version to 7.0.2</li> <li><a href="https://github.com/microsoft/TypeScript/commit/ed2bd1bfa4aac5211ce4bc58fcd1313c7eddc8ff"><code>ed2bd1b</code></a> Merge branch 'main' into ts7-release</li> <li><a href="https://github.com/microsoft/TypeScript/commit/887307575c58ea640dbeba3b4e8fdb6347cd3044"><code>8873075</code></a> Bump the github-actions group across 1 directory with 3 updates (microsoft/ty...</li> <li><a href="https://github.com/microsoft/TypeScript/commit/9427131ae2d4e230a90ee8a09daac4e75da3e311"><code>9427131</code></a> Set up stable / nightly extension split, other prep (microsoft/typescript-go#...</li> <li><a href="https://github.com/microsoft/TypeScript/commit/d4eaca5460a1f5f02a829e62706794b0a6fb903e"><code>d4eaca5</code></a><code>microsoft/typescript-go#4549</code></li> <li>Additional commits viewable in <a href="https://github.com/microsoft/TypeScript/compare/v5.9.3...v7.0.2">compare view</a></li> </ul> </details> <details> <summary>Maintainer changes</summary> <p>This version was pushed to npm by <a href="https://www.npmjs.com/~microsoft1es">microsoft1es</a>, a new releaser for typescript since your current version.</p> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4d2af732ae |
feat(runner): add native persistence contracts (#12169)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need durable records so Paperclip can explain results and final status changes. > - The current heartbeat tables support direct adapters, but they do not model native runner evidence. > - The runner transport and server coordinator must share a strict finalization contract before they write production data. > - This pull request adds that contract and its additive database boundary. > - It does not select the Paperclip Runner or change any existing adapter execution path. > - The benefit is a reviewable persistence layer that preserves all current behavior and supports later guarded integration. ## Linked Issues or Issue Description Refs #11962 Refs #12129 ## What Changed - Add native run result, finalization, completion, assessment, status decision, and status effect tables. - Add inert native metadata to heartbeat runs and events. Keep `legacy` as the default runtime mode. - Bind each evidence relationship to one company, issue, run, contract, result, assessment, and decision with composite constraints. - Add a strict `paperclip.native_finalization.v1` shared type and validator. - Preserve database functions, triggers, and the unique indexes required by foreign keys in JavaScript backups. - Add migration, backup, mixed-owner denial, validator, and direct-adapter compatibility tests. - Document the new records and their ownership rules. ## Verification - Run `pnpm -r typecheck`. - Run `pnpm build`. - Run `pnpm db:generate`. The schema output and migration safety checks remain current. - Run `PAPERCLIP_PSQL_PATH=/Applications/Postgres.app/Contents/Versions/latest/bin/psql pnpm exec vitest run packages/shared/src/validators/native-finalization.test.ts packages/db/src/client.test.ts packages/db/src/backup-lib.test.ts server/src/__tests__/heartbeat-workspace-busy.test.ts server/src/__tests__/heartbeat-comment-wake-batching.test.ts`. All 52 tests pass. - The full local `pnpm test:run` run completed 4,688 tests. It found 30 existing macOS test-environment failures. A serial rerun with the canonical `/private/tmp` path reduced those failures to six existing listener-diagnostics and skill-browser cases. None of those suites use files in this change. - The full Linux GitHub Actions matrix passes. This includes all general-server, serialized-server, workspace, browser, build, typecheck, canary, and aggregate verification jobs. - Greptile passes at 5/5. Contributor trust, Superagent, Socket, and Snyk pass with no finding from this change. - Storybook visual regression skips by path because this pull request has no UI or Storybook change. - Confirm that the diff contains 25 files. Confirm that it contains no workflow or `pnpm-lock.yaml` changes. ## Risks - The migration adds tables, columns, indexes, a function, a trigger, and ownership constraints. It does not remove or rename existing data. - Composite foreign keys reject mixed-company, mixed-issue, and mixed-run evidence even when each ID exists. - The status-version trigger runs only when an issue status changes. Backup tests confirm that restore retains this trigger and its dependencies. - Native source identifiers are unique when present. Legacy event rows remain unchanged. - This change does not add a unique run sequence constraint. The later native writer must allocate its sequence atomically before that invariant can be safe. - Existing adapters keep their current execution and finalization paths. New heartbeat runs default to `legacy` mode. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment ID and context-window size are not exposed. The model used agentic reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
38d8f37172 |
fix(build): enforce Node 24 across Paperclip (#11792)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs across the CLI, server, adapters, plugins, CI, and container images. > - These surfaces declared different Node.js versions from 20 through 24. > - A newer `@types/node` major can expose APIs that the supported runtime does not provide. > - Node.js 20 is no longer a suitable project baseline, and Node.js 24 is the current LTS line. > - This pull request sets Node.js 24.11.0 as one repository-wide baseline, adds a drift check, and gives users actionable startup guidance when their runtime is too old. > - The benefit is one clear runtime contract for development, release, installation, and published packages. ## Linked Issues or Issue Description Refs #2734 Refs #11727 Refs #739 ## What Changed - Require Node.js 24.11.0 or newer in all 42 package manifests and runtime checks. - Use Node.js 24 in GitHub Actions, Docker images, smoke images, sandbox setup, portable installs, and esbuild targets. - Align every direct `@types/node` declaration on `^24.0.0`. - Prevent Dependabot from opening major `@types/node` upgrades without a matching runtime decision. - Add `.nvmrc` and a CI policy check for Node version drift. - Update ACP version gates, tests, and user documentation for the new minimum. - Print a non-blocking warning on CLI and server startup when Node is unsupported, with remediation through a version manager or the documented downloaded `install.sh` workflow. - Deduplicate that warning when `paperclipai run` boots the CLI and server in the same process. ## Verification - `node scripts/check-node-version-policy.mjs` - `node --check scripts/check-node-version-policy.mjs` - `node --check cli/esbuild.config.mjs` - `node --check scripts/generate-npm-package-json.mjs` - `bash -n scripts/install.sh scripts/test-install-sh-docker.sh scripts/e2e-install-lifecycle.sh` - Parsed all 42 package manifests and confirmed `engines.node` is `>=24.11.0`. - `git diff --check` - `vitest run packages/adapter-utils/src/sandbox-install-command.test.ts` passed with 3 tests. - `vitest run cli/src/node-version.test.ts` passed with 4 tests. - Directly exercised the shared warning helper for unsupported-version messaging and same-process deduplication. - The focused exe.dev suite could not resolve the locally unbuilt plugin SDK from this isolated worktree. A full offline workspace install was also blocked because the package-manager signature verifier requires registry access. The full suite was not run locally; draft CI performs a clean install and evaluates the wider impact. ## Risks - This is a breaking runtime change for users, plugins, and deployments that still use Node.js 20 or 22. - Published workspace packages will now produce an engine warning or failure in strict package managers on older Node.js releases. - Node.js 24 can reveal dependency, native module, Playwright, or agent CLI compatibility issues in CI. - The bootstrap installer now installs Node.js 24 when the current runtime is older than 24.11.0. - The portable sandbox fallback is pinned to Node.js 24.11.0 and depends on that upstream tarball remaining available. - Unsupported runtimes continue booting after a warning, so a later incompatibility can still fail at its point of use. - The CLI and server share the warning policy through the published `@paperclipai/shared` package; packaging checks must keep that subpath export available. - This PR does not commit `pnpm-lock.yaml` because repository policy assigns lockfile generation to CI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact deployment ID and context window are not exposed in this session. Reasoning, repository tools, shell execution, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cb0009b097 |
fix: preserve recovery retries across restarts (#11817)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane must keep each active issue on a clear execution or recovery path. > - A missing issue disposition can require more than one bounded repair attempt. > - A server restart could lose that repair path or move source ownership to the recovery owner. > - A parked or expired retry could also make the user interface show a false healthy state. > - Concurrent recovery loops must not schedule the same repair attempt twice. > - This pull request keeps retry state durable, makes scheduling atomic, and keeps source ownership stable. > - The benefit is that recovery continues after a restart and operators see the correct state. ## Linked Issues or Issue Description **What happened?** A run that ended without a valid issue disposition could lose its repair path after a server restart. Manager recovery could also change the source owner. In addition, a parked or expired retry could make the issue look healthy when no active work existed. Concurrent reconciliation could also schedule the same repair attempt twice. **Expected behavior** Paperclip must keep bounded source and manager repair attempts across restarts. Recovery ownership must stay separate from source issue ownership. The server and user interface must report only a live retry as active work. Each repair attempt must be scheduled at most once per company. **Steps to reproduce** 1. Start an agent run on an issue. 2. End the run without a valid issue disposition. 3. Let the first repair attempt schedule a retry. 4. Restart the server, let the retry time pass without a live run, or start two reconciliation loops together. 5. Observe that the repair path can stop, the issue can show a false healthy state, or duplicate retries can be created. **Paperclip version or commit** The problem existed on `master` before candidate head `d8e620fe86bade7df18decac332007f5821ae04f`. **Deployment mode** The problem affects self-hosted servers and local builds that use automatic recovery. ## What Changed - Persist bounded source-owner and manager repair lineages with stable fingerprints and retry limits. - Resume incomplete disposition repairs after a server restart. - Keep recovery ownership separate from source issue ownership and enforce source mutation authority. - Project live retry evidence into issue and blocker summaries. - Show recovery owner, return owner, attempt count, and retry state in the board user interface. - Treat expired or parked retries as attention states unless a queued or running attempt exists. - Atomically deduplicate disposition-repair wake requests with a company-scoped partial unique index. - Reuse the winning run when concurrent reconciliation loses the uniqueness race, without duplicate scheduling activity. - Honor disabled on-demand wake policy before recovery scheduling and again before delayed retry promotion. - Keep the new index migration safe for lagging seeded databases that already contain the index. - Add server and user interface tests for recovery, restart, ownership, retry, concurrency, and blocker states. - Update the implementation and execution semantics documents. ## Verification - Focused server recovery and ownership suites: 282 tests passed on the repaired base candidate. - Focused user interface recovery suites: 128 tests passed on the repaired base candidate. - Atomic-deduplication schema and recovery suites: 111 tests passed on the first Greptile repair. - Recovery and scheduled-retry wake-policy suites: 126 tests passed at `d8e620fe86bade7df18decac332007f5821ae04f`. - The exact lagging-source migration-order test passed after the index migration became idempotent: 1 test passed and 62 unrelated tests were skipped. - `@paperclipai/db` and `@paperclipai/server` typechecks passed at the current head. - Migration generation and migration safety checks passed for migration `0226_tan_colossus.sql`. - `pnpm check:token-gates` passed on the repaired base candidate. - `pnpm -r typecheck` passed on the repaired base candidate. - `pnpm build` passed on the repaired base candidate. - `pnpm test:run` passed 4,540 tests on the repaired base candidate. Four fixed-port cases met listeners that already existed on the host. - The two unchanged fixed-port files passed in an isolated network namespace: 129 tests passed and 27 tests were skipped. - Independent Security and QA reviews approved `63c0423aab54c66f2293a20b0fb3f3b013ee3ba8`; exact-head re-review is required after automated checks settle on `d8e620fe86bade7df18decac332007f5821ae04f`. ## Risks - Recovery orchestration affects issue liveness and ownership. The new paths use bounded attempts, stable fingerprints, row locks, authority checks, and database uniqueness. - A conservative attention state can show more warnings when a scheduled retry has no queued or running attempt. It does not hide stopped work. - Migration `0226_tan_colossus.sql` creates a partial unique index on a known-large table. Migrations run transactionally, so `CONCURRENTLY` is unavailable. The matching disposition-repair key namespace is introduced by this release, so deployed databases have no matching rows before the index is added. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex from the GPT-5 model family used agentic reasoning, tool use, and code execution. The runtime did not expose the exact model ID or context window. - Anthropic Claude Opus 5 used a 1M context window, tool use, and code execution for part of the user interface repair, as recorded in the commit history. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e0b64529b3 |
feat(auth): normalize agent login in the sandbox onto one session table and a capability contract (#11730)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox agents need a safe login path for each supported adapter > - Codex device login and Claude setup-token login used separate session stores and route logic > - Separate stores made session lookup, expiry, and login capability checks harder to keep consistent > - This pull request unifies both flows on one session table and one capability contract > - The benefit is one company-scoped login model with public session identifiers and shared lifecycle rules ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Codex and Claude sandbox login used separate session stores and different route paths. This split increased the risk of inconsistent company scoping, session lookup, and cleanup. **Proposed solution** Use `adapter_auth_sessions` for both login flows. Use public session identifiers for API access. Select login behavior from projected adapter capability data. Share the route spine, lease arguments, runner lifecycle, and reaper rules. **Alternatives considered** Keep two session tables and add matching fixes to both routes. This keeps duplicate logic and does not provide one capability contract, so this pull request uses shared infrastructure. **Roadmap alignment** This change supports the shipped Cloud / Sandbox agents milestone in `ROADMAP.md`. ## What Changed - Unify Codex device login and Claude setup-token login on `adapter_auth_sessions`. - Return and look up sessions with company-scoped public session identifiers. - Enforce one active session for each company, owner, and adapter. - Share the login route spine, sandbox lease arguments, runner lifecycle, and missing-auth check. - Add a standalone setup-token reaper with adapter-specific row selection. - Add optional login capability projection for adapters and drive route and UI selection from that data. - Rename the provider flag to `supportsLoginPty` and validate its deprecated alias. - Remove the old Claude setup-token session table and add the required migrations. ## Verification - Server typecheck passed with `tsc`. - Database typecheck passed. - UI typecheck passed with `tsc -b`. - Codex login service and route suites passed. - Setup-token session, route, and reaper suites passed. - Adapter session schema, plugin validator, capability projection, UI render, and Daytona suites passed. - GitHub Actions must confirm the complete CI gate after pull request creation. ## Risks - The migrations remove short-lived in-flight login rows during deployment. A login that spans the migration can continue until its provider lease expires. - The Codex credential store remains company-scoped. A cross-owner credential race remains a documented, board-accepted risk. - API clients that use internal session row identifiers no longer work. The API accepts only public session identifiers. ## Model Used Codex, GPT-5, exact runtime model ID not exposed in this handoff, large context window, reasoning, and repository tool use. The implementing engineer produced the code with AI assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
98298cb7ff |
build(deps-dev): bump tsx from 4.23.1 to 4.23.12 (#11722)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.23.1 to 4.23.12. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/privatenumber/tsx/releases">tsx's releases</a>.</em></p> <blockquote> <h2>v4.23.12</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.11...v4.23.12">4.23.12</a> (2026-08-10)</h2> <h3>Bug Fixes</h3> <ul> <li>shim <code>import.meta</code> when tokens are split by comments or newlines (<a href="https://redirect.github.com/privatenumber/tsx/issues/829">#829</a>) (<a href="https://github.com/privatenumber/tsx/commit/ed9d33046a135de13a35fdfce12368b79d1b1518">ed9d330</a>), closes <a href="https://redirect.github.com/privatenumber/tsx/issues/828">#828</a></li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.12"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.11</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.10...v4.23.11">4.23.11</a> (2026-08-07)</h2> <h3>Bug Fixes</h3> <ul> <li>preserve async ESM require fallback (<a href="https://github.com/privatenumber/tsx/commit/55cbecef8ebe839c7110e8c141a1c3bc4da326cd">55cbece</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.11"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.10</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.9...v4.23.10">4.23.10</a> (2026-08-07)</h2> <h3>Bug Fixes</h3> <ul> <li>support nyc coverage discovery (<a href="https://redirect.github.com/privatenumber/tsx/issues/710">#710</a>) (<a href="https://github.com/privatenumber/tsx/commit/ec1bcd5f711e5159b67cb0aea211f06cf2cfce8a">ec1bcd5</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.10"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.9</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.8...v4.23.9">4.23.9</a> (2026-08-06)</h2> <h3>Bug Fixes</h3> <ul> <li>map Node test locations (<a href="https://github.com/privatenumber/tsx/commit/2f55884195a8c745fbe64a0288de69bc062ed876">2f55884</a>)</li> <li>support data URLs in tsImport (<a href="https://github.com/privatenumber/tsx/commit/b94f46f6b6a7e6dc575624b0ecc7124318723056">b94f46f</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.9"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.8</h2> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/privatenumber/tsx/commit/ed9d33046a135de13a35fdfce12368b79d1b1518"><code>ed9d330</code></a> fix: shim <code>import.meta</code> when tokens are split by comments or newlines (<a href="https://redirect.github.com/privatenumber/tsx/issues/829">#829</a>)</li> <li><a href="https://github.com/privatenumber/tsx/commit/651f5bec70d9a116d1fe1706000b7ca011a516fc"><code>651f5be</code></a> test: cover CommonJS TypeScript import.meta paths</li> <li><a href="https://github.com/privatenumber/tsx/commit/bd3bc6448e957c1172eb91a0584ec7fec2d6a7ad"><code>bd3bc64</code></a> test: cover CommonJS loader source fallback</li> <li><a href="https://github.com/privatenumber/tsx/commit/55cbecef8ebe839c7110e8c141a1c3bc4da326cd"><code>55cbece</code></a> fix: preserve async ESM require fallback</li> <li><a href="https://github.com/privatenumber/tsx/commit/6c5ba85f7a1e57f06bcb760718d6ba3b978e5a05"><code>6c5ba85</code></a> docs: document CommonJS default interop</li> <li><a href="https://github.com/privatenumber/tsx/commit/ec1bcd5f711e5159b67cb0aea211f06cf2cfce8a"><code>ec1bcd5</code></a> fix: support nyc coverage discovery (<a href="https://redirect.github.com/privatenumber/tsx/issues/710">#710</a>)</li> <li><a href="https://github.com/privatenumber/tsx/commit/b6e5b48a7b0fa4639e67119c8b02150ec8c6cef7"><code>b6e5b48</code></a> docs: clarify CommonJS default imports</li> <li><a href="https://github.com/privatenumber/tsx/commit/2f55884195a8c745fbe64a0288de69bc062ed876"><code>2f55884</code></a> fix: map Node test locations</li> <li><a href="https://github.com/privatenumber/tsx/commit/de935d588b6c0959ac8af71e2cf40d1a61e659e7"><code>de935d5</code></a> docs: document Node source-map stack formatting</li> <li><a href="https://github.com/privatenumber/tsx/commit/b94f46f6b6a7e6dc575624b0ecc7124318723056"><code>b94f46f</code></a> fix: support data URLs in tsImport</li> <li>Additional commits viewable in <a href="https://github.com/privatenumber/tsx/compare/v4.23.1...v4.23.12">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
e1df4c6068 |
fix(workspaces): keep deferred seed databases reliable (#11706)
## Thinking Path > - Paperclip manages agent work in isolated execution workspaces. > - A workspace depends on a valid database seed before it can run. > - Deferred seed failures were hidden behind a successful provision status. > - The seed restore also had two possible owners for the embedded PostgreSQL process. > - That allowed the target database to stop while the restore was still running. > - This pull request makes seed failures visible and gives the seed process sole lifecycle ownership. > - The benefit is that workspace provisioning reports the real result and does not stop its own target database. ## Linked Issues or Issue Description Related: #11684 **What happened?** Initial worktree provisioning could report success before its deferred database seed completed. The seed restore could also reuse a target embedded PostgreSQL process with another shutdown owner. This could stop the target database during the restore. **Expected behavior** Workspace status must show a failed deferred seed as a failure. The seed restore must own the target embedded PostgreSQL process until restore, migration, and validation finish. **Steps to reproduce** 1. Provision a worktree with deferred database seeding. 2. Make the seed manifest end in a failed state while the command exits with code 0. 3. Observe that the provision status remains successful on `master`. 4. Start a seed restore against an already-running target embedded PostgreSQL process. 5. Observe that another lifecycle owner can stop the target during restore. **Paperclip version or commit** `51a843e135` **Deployment mode** Local dev with execution workspaces and embedded PostgreSQL. ## What Changed - Add a first-class `workspace_seed` operation for deferred database seeds. - Require terminal, verified seed evidence before the seed operation succeeds. - Surface the seed phase and failure metadata in workspace status and UI state. - Give the seed process exclusive lifecycle ownership of the target embedded PostgreSQL process. - Suppress imported embedded-Postgres exit hooks without removing existing host listeners. - Record a credential-safe shutdown diagnostic in failed seed manifests. ## Verification - The original deferred-seed commit passed 4 server tests, 24 workspace-status UI tests, shared/server/UI typechecks, and the UI token gate. - The original PostgreSQL-lifecycle commit passed 3 lifecycle tests, 3 ownership/diagnostic tests, 1 real embedded-Postgres seed integration, and the affected package typechecks. - No local tests were rerun after the clean cherry-pick because the operator requested the shortest landing path. - Review the automatic PR checks for the clean `origin/master` replay. ## Risks - A live target database now causes an early error instead of being reused. The error includes recovery guidance. - Workspace consumers must handle the new `workspace_seed` operation type. Shared types and UI state handling are updated in this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, high-reasoning mode, with repository, shell, and GitHub tool use. The runtime does not expose a more specific deployment suffix or context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a2bf936f9a |
feat(workspaces): sign the workspace login handoff and gate readiness (#11671)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed worktree services run isolated Paperclip instances with cloned databases. > - A reachable service was reported as ready even when its database, runtime identity, or login path was not usable. > - The first candidate added verified database seeding and managed repair in #11665. > - This pull request consolidates that candidate with signed login handoff and a complete readiness contract. > - Post-QA fixes close five defects in repair identity, repair responses, UI retry, seed journal handling, and seed-source trust. > - The benefit is a workspace that either opens safely or reports one accurate recovery action. ## Linked Issues or Issue Description No public GitHub issue exists for this work, so the problem is described here. **What happened** Managed workspace URLs could return HTTP 200 and report ready while login failed. QA also found cases where repair used the wrong instance identity, returned a generic error, left the UI stuck, rejected a safe journal lag, or trusted a mutable workspace manifest. **Expected behavior** Opening a ready workspace signs the board user in to the correct isolated instance. Provisioning and repair use a registered source and report a structured recovery state. **Actual behavior** Entry depended on a password copied into the clone. Several failure paths could publish stale readiness, hide the repair precondition, or trust state that the workspace could modify. **Additional context** This pull request includes the commits first published in #11665. That pull request keeps the original base head for review history. This consolidated pull request is the merge candidate. Related open readiness work includes #11575 and #11621. ## What Changed - Adds a short-lived, signed, single-use login ticket. It binds the user, workspace, instance, and runtime origin. - Exchanges the ticket through Better Auth. It creates the session and cookie through the supported adapter path. - Adds protected workspace readiness fields for the database, clone data, login handoff, seed phase, and runtime identity. - Fails readiness closed when the guest has no company or execution-workspace binding. - Binds ticket issuance to the exact cloned user and active company membership selected for the handoff. - Verifies every current active board identity through the exact-user handoff before publication or reuse. - Gates managed runtime publication on the readiness contract and the recorded worktree instance identity. - Refreshes runtime work products from the live runtime row after a port change. - Adds one workspace access card with ready, degraded, repairing, and failed states. - Uses the runtime response identity for repair. It returns structured repair precondition errors. - Lets a valid source journal lag converge during provisioning. - Binds seed and repair manifests to a source registered outside the agent-writable worktree. - Clears recovered UI errors so a successful retry can open the workspace. - Makes runtime tests register canonical sources and avoid ports owned by live host listeners. - Keeps Vitest on source suites when compiled `dist` trees exist. - Isolates CLI and adapter tests from ambient AWS and runtime API environment variables. - Preserves a 404 response for cross-company workspace ID lookups before runtime authorization. - Makes concurrent single-flight coverage independent of path-canonicalization scheduling order. ## Verification The following checks passed on the integrated head: ```sh pnpm -r typecheck pnpm build pnpm check:token-gates pnpm --filter @paperclipai/db check:migrations ``` - The server source lane passed 420 files and 4,953 tests. Five tests were skipped. - The CLI lane passed 57 files and 385 tests. - The database lane passed 26 files and 97 tests. - The shared package passed 58 files and 506 tests. - The adapter utility lane passed 640 tests. Four tests were skipped. - The Claude adapter passed 220 tests. One test was skipped. - The Codex adapter passed 323 tests. - The OpenClaw adapter passed 13 tests. - The OpenCode adapter passed 42 tests. - The plugin SDK passed 45 tests. - The workspace runtime suite passed 124 tests. - The caller-scoped readiness and handoff suite passed 52 tests. - The workspace provisioning shell suite passed 7 tests. - The runtime exposure suite passed 17 tests while live host mappings occupied fixed test ports. - `git diff --check` passed and the worktree is clean. The serialized route lane will run in GitHub CI with its normal shards. No deployment or active-workspace migration was performed. ## Risks - This is a medium-risk authentication and runtime-readiness change. - The login ticket uses exact origin, workspace, instance, and user binding. It has a short expiry and a one-time nonce. - Runtime publication is stricter. A real readiness, identity, per-user handoff, or control-plane database disagreement now blocks publication. - This pull request supersedes #11665 as the merge candidate. Close #11665 after this pull request merges. - No new database migration is included. The lockfile and workflow files are unchanged. - Deployment and active-workspace migration are intentionally outside this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Opus 5 (`claude-opus-5[1m]`), 1M context, extended thinking, tool use, and code execution produced the main candidate. OpenAI GPT-5 (`gpt-5`) through Codex, with agentic reasoning, tool use, and code execution, integrated the post-QA fixes and hardened the test gates. The Codex context-window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6b8e42168e |
Add governed secret alias confirmation cards (#11486)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need scoped secret bindings to use external services safely. > - Agents could not request an existing secret under a new config name without an internal secret identifier. > - Existing binding proposals were only visible in Settings and did not create an issue-thread approval path. > - A confirmation card could record acceptance without proving that the binding was created. > - This pull request extends the existing secret proposal system with safe source references and governed issue-thread confirmation cards. > - The benefit is a one-click flow that creates the binding or shows a clear failure without exposing secret material. ## Linked Issues or Issue Description Related prerequisite: #11482. **Subsystem affected** Cross-cutting: server REST APIs, shared interaction contracts, database proposal schema, and issue-thread UI. **Problem or motivation** An agent can need an existing bound secret under a second config name. The agent cannot safely discover the internal secret identifier. The existing proposal is also easy for the operator to miss because it only appears in Settings. A generic confirmation can record acceptance without executing the binding. **Proposed solution** Let an agent create a binding proposal from one of its existing config paths. Mint a server-owned, human-only confirmation card on the checked-out issue. Recheck the operator's target-agent permission under the proposal row lock. Execute the existing proposal transaction after card acceptance. Store an `executed` or `failed` result on the card. Render the complete lifecycle in the issue thread and attention resolver. **Alternatives considered** A new alias subsystem would duplicate proposal quotas, expiry, authorization, and binding synchronization. A text-only issue comment would not provide a governed action or an execution result. An agent-supplied card payload would permit metadata smuggling. This change uses the existing proposal transaction and a server-owned payload instead. **Roadmap alignment** This change extends the completed "Secrets Manager with per-agent access" roadmap item. It preserves scoped bindings and audited resolution. The required GitHub search found no other open duplicate issue or pull request. ## What Changed - Added safe source-config-path binding proposals and preserved user-secret ownership checks. - Added a proposal-to-interaction link and an idempotent database migration. - Minted human-only `request_confirmation` cards with server-owned `secretProposal` metadata. - Rejected agent-supplied governed metadata and agent addressees. - Rechecked `agent_config:update` authority under the proposal lock before execution. - Recorded `executed` or `failed` results and posted a failure comment when no binding was created. - Settled failed accepted proposals atomically and mirrored rejection, withdrawal, and expiry in both directions. - Emitted `secret.binding.created` for new agent binding writes. - Added a dedicated issue-thread card for pending, executed, failed, rejected, withdrawn, and expired states. - Showed only the source label, target agent, config path, skeptical justification, expiry, and safe failure code. - Replaced resolved attention-query entries immediately with the stitched server result. - Added focused server, database, UI, and state-transition tests. - Added Storybook fixtures for every review state and documented the API and agent behavior. ## Verification - `pnpm exec vitest run ui/src/components/IssueThreadInteractionCard.test.tsx ui/src/components/AttentionInteractionResolver.test.ts` — 58 passed. - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm build-storybook` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/db check:migrations` - `NODE_ENV=test pnpm exec vitest run server/src/__tests__/issue-thread-interaction-routes.test.ts server/src/__tests__/secret-proposals-routes.test.ts server/src/__tests__/secrets-routes.test.ts server/src/__tests__/agents-service-secret-bindings.test.ts` — 142 passed. - `NODE_ENV=test pnpm --filter @paperclipai/db exec vitest run src/company-secret-proposals-migration.test.ts --silent` — 1 passed. - `pnpm -r typecheck` - `pnpm test:run` — server 4,175 passed, UI 4,109 passed; the CLI AWS-doctor case passes 8/8 with runtime-injected static AWS credential variables unset. - `pnpm build` - `git diff --check origin/master...HEAD` ## Risks - Migration `0221` adds one nullable foreign key and one index. It uses idempotent guards. - The accept route performs a governed write after it records card acceptance. A failed write is visible and settles the proposal as rejected. - Concurrent proposal and card resolution must use proposal-before-interaction lock order. A race test covers direct approval against card rejection. - The new audit event increases activity rows for newly added agent bindings. It does not include secret values or fingerprints. - The card includes only safe proposal metadata. It does not include secret value, fingerprint, version, or internal secret identifiers. - The UI uses the stitched resolution result. Focused tests cover immediate cache replacement and every terminal state. > This work extends an existing completed roadmap capability. The GitHub duplicate search returned no other open related work. ## Model Used - OpenAI Codex with model ID `gpt-5`. The runtime did not expose its context-window size. Reasoning, repository tools, code execution, database integration tests, UI rendering, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c1c46f1e4e |
feat: Claude login on the new-agent page before agent creation (#11347)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Claude local adapter supports subscription login through a sandbox > - The new-agent page must show login before the user creates an agent > - Test results must not expose raw sandbox diagnostics or secret values > - This pull request adds the login UI to both Test lanes and closes the diagnostic boundary > - The branch also adds durable cleanup recovery for failed sandbox teardown > - Reusable sandboxes must retain both their recorded teardown configuration and a valid lifecycle path until destruction succeeds > - The benefit is a usable login flow with fixed public checks, redacted server logs, and recoverable sandbox cleanup ## Linked Issues or Issue Description Related public work: [#9488](https://github.com/paperclipai/paperclip/pull/9488) adds first-class recognition for `CLAUDE_CODE_OAUTH_TOKEN` in headless and remote runs. Related public issue: [#2681](https://github.com/paperclipai/paperclip/issues/2681) requests Claude Code subscription support. This pull request adds the login transport and new-agent UI flow that those changes do not provide. **Subsystem affected:** Claude local adapter, server login probes, sandbox provider setup, cleanup recovery, and the new-agent UI. **Problem or motivation:** The Test lanes did not show the sandbox login panel in all supported cases. Test results also exposed raw probe diagnostics, and JSON escapes could end secret redaction early. **Proposed solution:** Surface the login capability through the bundled provider manifest. Prepare the same probe runtime in the ACP lane. Send diagnostics only to redacted server logs. Keep Test checks on fixed public messages. Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. Consume JSON escapes during redaction. Preserve failed sandbox cleanup state across retries and restarts, and prevent deletion from severing the lifecycle context of a live reusable sandbox. **Alternatives considered:** Keep raw diagnostics in Test checks or trust login URL text from the sandbox. Both choices increase information exposure. Keep separate probe behavior in the ACP lane. That choice would leave the two Test lanes inconsistent. ## What Changed - Surface the sandbox login panel on both Test lanes. - Reconcile the bundled Daytona plugin manifest so `supportsSetupTokenLogin` reaches the UI capability gate. - Prepare the ACP Test lane with the same probe runtime as the CLI Test lane. - Add the `claude_acp_login_probe_unavailable` warning when the ACP probe cannot run. - Send raw sandbox diagnostics only to redacted server logs. - Keep Test checks on fixed public messages in the ACP, managed-config, and CLI paths. - Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. - Redact JSON and escaped-JSON secret values, including escaped quotes and backslashes. - Preserve orphan cleanup records across provider failures, restarts, and unavailable plugins. - Atomically block environment deletion while a live reusable sandbox lease still depends on it. - Verify pending cleanup destroys plugin sandboxes with the provider configuration recorded on the lease, even after the current environment configuration changes. ## Verification - Head under review: `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241`. - Focused environment route/service/runtime coverage passes: 196 tests across 3 files. - `pnpm -r typecheck` passes. - `pnpm build` passes. - The full Vitest run completed with 4,754 passing and 28 failing tests. All 23 source-test failures reproduce unchanged on parent head `58cfe61a33191ce03d965d65085d26064b4888ba`; the other 5 are duplicate executions from stale `server/dist` output. The failures are unrelated macOS path/listener and scheduler-fixture failures, so there is no new bad commit for bisect to localize. - All required CI checks pass for the current head, including build, typecheck/release registry, all server and workspace shards, serialized server suites, canary, and e2e. - A fresh Greptile review for `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241` reports 5/5, “safe to merge,” with no blocking failure remaining. ## Risks - A probe or redaction change could hide useful server diagnostics. - An allowlist change could reject a valid Claude login URL. - Cleanup recovery changes could affect provider teardown ordering. - An environment with a live reusable sandbox can no longer be deleted until the owning issue or execution workspace completes teardown. - The implementation keeps public Test messages fixed and sends detail to redacted server logs. ## Model Used OpenAI GPT-5 via Codex — exact model ID: GPT-5; tool use and code execution enabled; extended reasoning enabled. The implementation author used AI-assisted development. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and documented the result - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation or confirmed no separate documentation change is needed - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4802b1bbc |
feat(runtime-exposure): least-privilege Tailscale HTTPS broker, shared contract, and persisted exposure state (#11524)
<!-- Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts and supervises managed runtime services for a project's execution workspaces, so an agent's branch can be previewed while it works > - Those services only listen on plain loopback HTTP. A person on another device, or on a phone, cannot open the preview > - A Tailscale HTTPS mapping solves this, but `tailscale serve` needs host privileges that the Paperclip server process must not hold > - This pull request adds the foundation only: a separate least-privilege host broker, the shared exposure contract, and the database columns that hold exposure state > - Nothing calls the broker yet, so there is no behavior change. The benefit is that the privileged surface is small, reviewable, and isolated before any lifecycle code depends on it ## Linked Issues or Issue Description No public GitHub issue exists. The change follows the feature request template. **Subsystem affected** Managed workspace runtime services, the shared type and validator package, and the database schema. **Problem or motivation** A managed runtime service binds to loopback only. There is no supported way to reach that preview from another device. Adding HTTPS directly to the server would mean the server process runs `tailscale serve`, which needs privileges far wider than the task requires. A compromised or buggy server could then map any port to the tailnet. **Proposed solution** Split the privileged work into a separate broker process with a narrow protocol, and define one shared contract that the server, the UI, the runtime, and the broker all read. Land this foundation first, with no caller, so the privileged code can be reviewed on its own. **Alternatives considered** - Call `tailscale serve` from the server process. This was rejected because it gives the server unrestricted mapping authority. - Use `sudo` for single `tailscale` commands. This was rejected because the argument list is the only guard, and it is easy to widen by accident. - Use a generic reverse proxy. This was rejected because it does not remove the need for a privileged Tailscale mapping step. **Roadmap alignment** This supports the existing managed workspace runtime capability. It adds no new product surface on its own. **Additional context** The broker is the security boundary of the feature, so it is deliberately the first slice. Three later pull requests build on it: the server exposure lifecycle, the runtime lease and recovery integration, and the leased-port mediator. ## What Changed - Add the `@paperclipai/tailscale-https-broker` workspace package. The broker listens on a unix socket, authorizes each peer with `SO_PEERCRED`, and answers a small request protocol. - Restrict what the broker will map. It accepts only same-number HTTPS-to-loopback pairs inside the Paperclip port range, refuses protected ports, and confirms that the loopback port belongs to a Paperclip-owned listener. - Parse every request with a strict JSON reader that rejects duplicate keys, prototype keys, and unknown fields. - Write an append-only audit record for each broker decision. - Add the shared exposure contract in `@paperclipai/shared`: the `RuntimeExposureConfig`, `RuntimeExposureState`, and `RuntimeExposureStatus` types, their zod validators, the app and HMR port rules, and the loopback-bind helpers. - Persist exposure state on `workspace_runtime_services` with the new `exposure` column, plus the server-private `exposure_handle` and `backend_url` columns that are never serialized to API clients. - Add the `execution_workspace_runtime_leases` table that the later lease slice uses. - Extend the runtime read-model test fixture for the three new columns. ## Verification Focused checks, all run on this branch: - `pnpm --filter @paperclipai/tailscale-https-broker test` — 12 files, 82 tests pass. This covers peer credentials, port policy, protected ports, the serve config writer, the strict JSON reader, argv parsing, and the socket server. - `pnpm --filter @paperclipai/tailscale-https-broker typecheck` — clean. - `npx vitest run --root packages/shared src/runtime-exposure src/validators/runtime-exposure.test.ts` — 3 files, 40 tests pass. - `pnpm --filter @paperclipai/db typecheck` — runs `check:migrations` first. Migration numbering and migration safety both pass. - `pnpm --filter @paperclipai/shared typecheck` — clean. - `pnpm --filter @paperclipai/ui typecheck` — clean. - `npx vitest run --root server src/services/workspace-runtime-read-model.test.ts` — 3 tests pass. - `npx tsc --noEmit -p server/tsconfig.json` — 139 errors, which is exactly the count on `master` before this branch. All 139 come from the unbuilt `@paperclipai/plugin-sdk` package. To confirm the exposure state is inert, start a managed runtime service as usual. The new columns stay null and the service behaves as it does today. ## Risks - Migration risk is low. Both migrations only add a table and three nullable columns. No column is backfilled and no existing column changes. The migration safety check passes. - Behavior risk is low. No code path calls the broker in this pull request, and the shared exposure fields are optional. - The broker is privileged, so it is the real risk surface. It is mitigated by peer-credential authorization, a fixed port range, a protected-port deny list, same-number pair enforcement, listener-ownership checks, strict JSON parsing, and an audit trail. Reviewers should read `packages/tailscale-https-broker/src/authorization.ts` and `src/port-policy.ts` closely. - The broker requires a `tailscale` version floor, which its README records. An older host CLI makes the broker refuse to start rather than map incorrectly. - `pnpm-lock.yaml` changes because a new workspace package is added. The diff is the new importer block, plus one duplicate `tinyexec` entry that pnpm removed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. ## Model Used Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
10d0555189 |
fix(interactions): authorize resolvers consistently (#11376)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue interactions give agents and people a structured decision record. > - Resolver routes used different authorization rules. > - Some routes blocked valid agents, including task watchdogs with normal issue access. > - The API did not show who could resolve a pending interaction. > - This pull request gives every interaction kind one resolver policy evaluator. > - The benefit is a clear decision path with consistent governance and company isolation. ## Linked Issues or Issue Description Fixes: #8087 Refs: #7403 Related PR: #11082 proposes board-only confirmation rules. This change keeps human-only review as an explicit policy. **What happened?** Agents could create issue interactions. Some resolver routes still required board access. This left valid agent confirmations pending. Task watchdogs could see the same problem without board identity. **Expected behavior** Every interaction kind must use one resolver policy contract. The contract must support `anyone`, `not_creator`, and `human_only`. It must also apply all normal governance controls. **Steps to reproduce** 1. Create a `request_confirmation` interaction as an agent. 2. Resolve it with another authorized agent. 3. Observe the board-only denial. **Paperclip version or commit** The problem exists on `master` before this change. **Deployment mode** Local development with `pnpm dev`. ## What Changed - Add canonical policies for `anyone`, `not_creator`, and `human_only`. - Use one server evaluator for every interaction kind. - Apply named addressees, company limits, review rules, and task watchdog scope. - Charge cross-issue resolutions to the existing per-run action limit. - Return the effective resolver audience in attention and interaction data. - Show the audience, governance choices, and denial reasons in the board UI. - Add telemetry, API documents, product documents, and regression fixtures. - Add migration provenance for safe legacy behavior. - Make migration `0218` safe for complete replays and partial prior runs. ## Product Rules - An interaction records a response. It does not grant authority for the next action. - `anyone` lets any authorized issue participant respond. - `not_creator` requires a responder other than the interaction creator. - `human_only` requires an authorized person. - A named addressee, company policy, or governed action can narrow the audience. - These controls cannot widen the audience. - A task watchdog uses the same rules as an ordinary agent. - A task watchdog does not receive board authority. - An agent resolution on another issue uses the shared cross-issue action limit. - Legacy pending interactions keep their earlier restrictions. - The UI shows the effective audience and a permanent denial reason. ## Verification - `pnpm --filter @paperclipai/db check:migrations` - `pnpm --filter @paperclipai/db typecheck` - `pnpm exec vitest run packages/db/src/issue-thread-interaction-resolver-policy-migration.test.ts` - The focused PostgreSQL test applies migration `0218` twice. - The test also completes a partial prior run and preserves existing provenance. - The latest GitHub head has 29 successful checks. - The opt-in Storybook visual check skipped as expected. - Greptile reports 5/5 with no open comments. ## Risks - New interaction writes use `anyone` by default. - Callers must select `not_creator` or `human_only` when they need stricter review. - Legacy pending interactions keep the old creator and human restrictions. - Migration `0218` fills only missing provenance fields during recovery. - Cross-issue resolutions can reach the existing action limit. - The shared evaluator affects every interaction kind. - Route, service, database, shared contract, and UI tests cover these rules. > This work matches the Agent Reviews and Approvals direction in `ROADMAP.md`. It does not duplicate a planned item. ## Model Used OpenAI Codex, GPT-5. The runtime does not expose the exact deployment ID or context window. The agent used reasoning, repository tools, shell commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked public issues or described the issue with the required labels - [x] I have not referenced internal Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented the risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open comments - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
44694328a3 |
fix(issues): make DELETE /api/issues/:id succeed for issues with dependents (#11331)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server provides issue APIs and the database stores issue child rows > - The issue delete endpoint removes the parent issue before dependent rows > - Several issue foreign keys had no delete policy, so PostgreSQL returned a foreign-key error > - This pull request adds safe cascade and set-null policies and a clear conflict response > - The benefit is reliable issue deletion with a useful error when a restricted audit row still blocks deletion ## Linked Issues or Issue Description Fixes #7728 Fixes #4660 Fixes #7991 Fixes #4627 Fixes #5086 **What happened?** `DELETE /api/issues/:id` returned HTTP 500 when dependent comments, thread interactions, read states, inbox archives, feedback votes, or ledger rows referenced the issue. The database raised SQLSTATE 23503 because several foreign keys had no delete policy. **Expected behavior** The endpoint must remove dependent rows that have no meaning without the issue. It must keep ledger rows with a null issue reference. It must return HTTP 409 when a restricted decision audit row still references the issue. **Steps to reproduce** 1. Create an issue. 2. Add a comment or thread interaction that references the issue. 3. Send `DELETE /api/issues/:id`. 4. Observe the HTTP 500 response. **Paperclip version or commit** Commit `1f8f456f8340823fe2bd891ae8933d942f190b7b`. **Deployment mode** Local dev with embedded PGlite or external PostgreSQL. ## What Changed - Add `CASCADE` to five issue child foreign keys. - Add `SET NULL` to the finance and cost event issue foreign keys. - Keep decision audit references restricted. - Map SQLSTATE 23503 from the issue delete service to HTTP 409. - Add migration 0217 for the seven changed tables. - Add regression tests for cascade deletion and restricted decision references. ## Verification - Run `pnpm --filter @paperclipai/db typecheck`. - Run `pnpm --filter @paperclipai/server typecheck`. - Run `npx vitest run src/__tests__/issue-remove-cascade.test.ts` from `server/`. - The regression test applies migration 0217 to a fresh embedded PostgreSQL database. ## Risks - Migration 0217 changes only seven foreign keys that reference `issues.id`. - Cascade deletion removes child rows that cannot exist without the parent issue. - Set-null preserves finance and cost ledger rows. - Decision audit rows remain protected, so the endpoint can return HTTP 409. ## Model Used Codex, based on GPT-5, with tool use and code-review support. The implementation author used an AI coding agent. This PR handoff uses the same model family to validate the commit and manage the pull request. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f0e6c0f549 |
feat(server): receive and apply the Paperclip Cloud onboarding seed (#11098)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud provisions a dedicated tenant stack for each
customer. During signup it asks for a mission, a name and role for the
first agent, and a first task.
> - Cloud pushes those answers into the new stack at activation, as
`POST /api/companies/:companyId/onboarding-seed`.
> - No route served that path. The tenant answered 404, so Cloud
recorded the push as unacknowledged and retried on every portfolio
fetch.
> - The failure was soft. The answers stayed durable in Cloud and the
stack still activated. But the stack opened on the empty first-run
wizard, and it asked the customer again for what they had already given.
> - This pull request adds the receiving endpoint. It validates the
seed, applies it, and acknowledges it.
> - The benefit is that a seeded stack opens with the mission, the agent
and the first task already in place.
## Linked Issues or Issue Description
No public GitHub issue covers this. The problem is described in-PR,
following the feature template.
**Subsystem affected**
server/ — Express REST API and orchestration services. Also
`packages/db` (one new table) and `packages/shared` (one new validator).
**Problem or motivation**
Paperclip Cloud collects onboarding answers at signup and pushes them to
the tenant stack at activation. The tenant had no route for that
request. It answered 404. Cloud treats a non-2xx as "not yet applied",
so it kept the answers and retried, but the stack itself stayed
unseeded. A customer who had already named their mission, their first
agent and their first task arrived at an empty first-run wizard that
asked for all three again.
**Proposed solution**
Serve `POST /api/companies/:companyId/onboarding-seed`. Validate the
body, apply it to the company, then acknowledge it.
The seed is customer free text, so it is bounded and validated in
`packages/shared` and read from the JSON body only. It is never read
from an `x-paperclip-cloud-*` header. That header set is the trusted
identity envelope: every member is derived server-side from the host
plus verified domain records, and that is exactly what makes it
trustworthy. Mixing user content into it would remove the property. A
test plants a mission on a cloud header and asserts that the body value
wins.
Application reuses the shapes the first-run wizard already produces, so
a seeded stack and a manually onboarded one look the same afterwards:
- The mission becomes the company-level goal. A multi-line mission
splits into a title and a description, as the wizard does.
- The agent becomes the company's first hire. Its free-text role ("Chief
of Staff") lands on `title`. The structural `role` stays `ceo`, which is
what the org chart and the default-instructions lookup read.
- The first task becomes an issue in the Onboarding project, assigned to
that agent.
Cloud retries until it gets a 2xx, and it reads any 2xx as "the tenant
holds this content". So the endpoint is idempotent per `revision`. A new
`company_onboarding_seeds` table records the applied revision together
with the goal, the agent and the issue it produced. A replay of a
revision that already matches is a successful no-op. A later revision —
the customer edited their answers — updates those three rows in place
instead of creating a second agent and a second task. The record is
written last, after every other write has landed, so a partial
application cannot present itself as acknowledged.
Everything is applied before the 200 is sent. This is an ordering
guarantee, not eventual consistency. The tests read the database
immediately after the response, with no waiting and no polling, so a
lazy receiver fails them on a fast machine as well as a slow one. That
matters because the redirect into the tenant dashboard is gated on this
acknowledgement.
**Alternatives considered**
Store the seed and let the tenant UI apply it on first load. Rejected:
the dashboard redirect is gated on the acknowledgement, so a background
apply would let the dashboard open before the agent and the task exist.
The whole point is that it must not.
Reuse `POST /companies/:companyId/agents` and `POST
/companies/:companyId/issues` over HTTP from Cloud. Rejected: it needs
three round trips with no shared idempotency key, and it moves the "did
all of it land?" decision to the caller.
**Roadmap alignment**
This completes an existing Cloud-to-tenant contract. It does not add a
new user-facing surface.
## What Changed
- Add `POST /api/companies/:companyId/onboarding-seed` in
`server/src/routes/onboarding-seed.ts`. It authenticates exactly as
`POST /api/companies/:companyId/logo` does, through
`assertCompanyAccess`.
- Add `server/src/services/onboarding-seed.ts`. It applies the mission,
the agent and the first task, and records the applied revision last.
- Add the `company_onboarding_seeds` table: schema, migration `0216`,
and journal entry. It holds the applied revision and the ids of the
goal, agent and issue the seed produced.
- Add `applyOnboardingSeedSchema` in `packages/shared`. It bounds
mission to 2000, agent name to 80, agent role to 120, task title to 200,
and task details to 2000 — the same limits Cloud enforces before it
sends.
- Mount the router in `server/src/app.ts` and register the path in the
OpenAPI document.
- Add `server/src/__tests__/onboarding-seed-route.test.ts` with 13
tests.
- The seeded agent is created on `claude_local`. This mirrors the
teams-catalog default for agents created server-side, where no human
runs an environment test first. `PAPERCLIP_ONBOARDING_SEED_ADAPTER_TYPE`
overrides it.
## Verification
```sh
pnpm typecheck # whole workspace, passes
npx vitest run \
server/src/__tests__/onboarding-seed-route.test.ts \
server/src/__tests__/openapi-routes.test.ts # 15 passed
```
The suite runs against embedded Postgres with migrations applied, so
migration `0216` is exercised by every test.
The route tests cover:
- the happy path — mission, agent and task all applied, read immediately
after the 200
- replay of the same revision — no second agent, no second task, no
second goal, no second project
- a later revision — the goal, agent and task are updated in place
- a multi-line mission splitting into a goal title and description
- a revision-only seed
- the activity log entry written once, and not again on a replay
- a caller without access to the company — 403, and nothing written
- a body with no revision — 400
- each field bound past its limit — 400
- a mission planted on an `x-paperclip-cloud-*` header — ignored, body
wins
- an existing Onboarding project — reused, not duplicated
Not verified here: the full Cloud-to-tenant walk against a live stack.
That needs a deployed Cloud and a provisioned tenant together, which is
separate staging work.
## Risks
Migration `0216` creates one new table. It adds no column to an existing
table, rewrites nothing, and backfills nothing, so it is safe to apply
online. The migration safety check passes.
The endpoint writes to a company. Access is enforced by
`assertCompanyAccess`, the same gate the company logo write uses, and a
test covers the denial.
Behavioral note for stacks that already hold data. If a company already
has a non-built-in `ceo` agent, a first seed updates that agent's name
and title rather than creating a second lead. Likewise a seed adopts an
existing company-level goal rather than adding a parallel one. This is
deliberate: the seed is the customer's own stated answer from signup,
and two competing missions or two leads would be worse than one updated
in place. In the intended case — a stack that Cloud has just activated —
none of these exist yet.
The seeded agent is created on `claude_local` with an empty adapter
config. It is idle and needs the usual credential setup before it runs.
Seeding it does not start it.
## Update — rebased onto master + review hardening
Master moved on after this PR was cut, so it was **rebased onto
`master`** and
the seed migration was **renumbered from `0212` to `0216`** (the merged
#11101
took `0212_onboarding_first_task_unique`); the drizzle journal was
re-stitched
and `check:migrations` passes.
Two things landed on top of the original receiver:
- **Mission-only walk contract (PAP-67 r17.4).** The tenant now owns the
first
agent and the first task via #11101's server-owned onboarding path,
which
stamps `ONBOARDING_FIRST_TASK_ORIGIN_KIND` and races safely on the
partial
unique index `issues_onboarding_first_task_uq`. A comment in the apply
path
documents why this receiver leaves the first task to that path on the
cloud
walk, and a paperclip-cloud `node:test`
(`src/onboarding/walk-seed.test.ts`)
asserts the walk's seed carries no `agent`/`firstTask`. The receiver
retains
the agent/first-task code for its documented body contract, kept inert
on the
cloud path by the mission-only seed.
- **Three Greptile P1 fixes** (`95622fa37`): concurrent application is
now
serialized under a per-company `pg_advisory_xact_lock` (no duplicate
goal/agent/project/task on overlapping pushes); a revised first task
carries
its resolved `assigneeAgentId`/`goalId`; and the
`company.onboarding_seed_applied`
audit write is best-effort so a logging failure can't leave the entry
permanently absent. Two new regression tests cover the first two.
## Model Used
Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking,
with tool use and code execution. Used for the original codebase
investigation, the implementation, and the tests. The rebase, migration
renumber, mission-only contract, and the three P1 fixes were done with
Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, with tool use
and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
1e07d5b9aa |
fix(db): give the last two embedded-Postgres migration tests a timeout (#11313)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `@paperclipai/db` package owns the database schema and its migrations > - Some migration tests start an embedded Postgres server and replay a migration against it > - An embedded Postgres server needs 7 to 12 seconds to start on a CI runner > - Vitest stops a test after 5 seconds unless the test sets its own timeout > - Two of these tests do not set a timeout, so they fail on CI before they assert anything > - This pull request gives both tests a 30 second timeout > - The benefit is that unrelated pull requests stop failing on a test they did not change ## Linked Issues or Issue Description No public issue exists for this. The problem follows. **What happened?** The test `packages/db/src/company-secret-proposals-migration.test.ts` fails on CI. The error is `Test timed out in 5000ms`. The test never reaches its assertions. The suite reports `1 failed | 104 passed`. The failure is not caused by the branch under test. It appeared on three different branches in a few hours: | Run | Head | Failing jobs | | --- | --- | --- | | 31630781317 | `95622fa3` | `General tests (workspaces-b)`, `verify`, `e2e shard (2/3)`, `e2e` | | 31652020976 | `feba90c9` | `General tests (workspaces-b)`, `verify`, `e2e shard (3/3)`, `e2e` | | 31651467721 | `f2115207` | `General tests (workspaces-a (1/2))`, `verify` | The `verify` job reads the result of the general tests. One timeout therefore turns into two red checks. A reviewer sees two failures and reads them as a regression. **Expected behavior** The test starts an embedded Postgres server, replays the migration, and asserts the schema. It must pass on a normal CI runner. **Steps to reproduce** 1. Open any pull request against `master`. 2. Wait for the job `General tests (workspaces-b)`. 3. Read the failure. The test times out after 5000 ms. The failure needs a slow runner. A fast development machine starts embedded Postgres in less than 5 seconds, so the test passes there. **Paperclip version or commit** `master` at `a09d7dcc0`. **Deployment mode** CI only. GitHub Actions, `ubuntu24` runner image. ## What Changed - `packages/db/src/company-secret-proposals-migration.test.ts` — the test now uses a 30 second timeout. The migration suites in this package already use 20 to 60 seconds. 30 seconds is the most common value. - `packages/db/src/status-card-migrations.test.ts` — the same change. This test has the same defect. It does not fail yet because it replays fewer statements. A fix to only one test moves the problem instead of removing it. - Both tests get a comment. The comment tells the next author why the 5 second default is too short. These two tests were the only embedded-Postgres migration tests in the package without a timeout. ## Verification - Run `pnpm vitest run src/company-secret-proposals-migration.test.ts src/status-card-migrations.test.ts` in `packages/db`. Both tests pass. - These suites skip themselves when the Postgres binaries are absent. A pass alone therefore proves nothing. Run the command with `--reporter=verbose`. The output contains Postgres `NOTICE` messages, for example `relation "status_cards" already exists, skipping`. These messages prove the tests ran real SQL. - Run the same command with `--testTimeout=1`. Both tests still pass. This proves the per-test timeout overrides the global timeout. Before this change, the same command fails immediately. - All CI jobs on this pull request pass. The job `General tests (workspaces-b)` passes. This job failed on the three runs listed above. Not done: no attempt to reproduce the timeout on a development machine. A fast machine starts embedded Postgres in less than 5 seconds, so the failure does not occur there. ## Risks Low risk. The change adds two timeout arguments to tests. It changes no source code, no schema, and no dependency. A longer timeout cannot hide a regression here. The tests assert the same conditions as before. A migration that truly hangs now fails after 30 seconds. Before, it failed after 5 seconds with a message that pointed at the wrong cause. The `e2e` failures on the runs above have a different cause. The spec `mcp-user-stories.spec.ts › US-9` fails with `502 — fetch failed` and `fetch failed: bad port`. These errors come from MCP tool-connection health checks. The failures hit different shards on different runs. This pull request does not change that behavior. `e2e shard (2/3)` passes here, which supports the view that those failures are unstable infrastructure. To revert, remove the two timeout arguments. ## Model Used Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking enabled. Tool use enabled: file read and edit, shell command execution for the local test runs, and the GitHub CLI to read the failing CI logs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e5a7fd7038 |
Add sandbox device-login for the Codex adapter (#11237)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Agent adapters connect Paperclip to tools such as the Codex command line tool. > - A sandboxed Codex agent may start without a credential. > - The operator needs a safe sign-in flow that does not expose credentials to the shared package or the sandbox. > - This pull request adds a company-scoped device-login flow with a temporary Daytona sandbox. > - The flow promotes the credential only after readiness checks pass and removes the temporary sandbox after use. > - The result lets an operator sign in to a sandboxed Codex agent from the agent form. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** A Codex adapter that runs in a sandbox cannot authenticate when the company has no pre-provisioned Codex credential. **Proposed solution** Add a company-scoped device-login session. Start a temporary sandbox, run `codex login --device-auth`, stream the code and URL, verify readiness, promote the credential, and delete the sandbox. **Alternatives considered** Pre-provisioning a credential does not support first-time sandbox login. Keeping the credential in the login sandbox does not provide a durable company credential. **Roadmap alignment** This supports the roadmap item for cloud and sandbox agents. **Additional context** The flow uses a five-minute cleanup reaper, compare-and-set status changes, and a PostgreSQL advisory lock to protect promotion and cleanup. ## What Changed - Add the adapter login-session contract, database table, and migration. - Add company-scoped server routes and a service for sandbox device login. - Add credential promotion, readiness checks, and cleanup after login. - Add restart-safe cleanup for abandoned login sandboxes. - Add sandbox login controls to the agent creation and edit forms. - Keep device-login and vendor identifiers out of public shared and adapter UI symbols. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local exec vitest run` passed with 310 tests at the submitted commit. - The server login route, service, and reaper tests passed with 45 tests at the submitted commit. - The agent form render tests passed with 26 tests at the submitted commit. - The public-symbol leak check passed at the submitted commit. - A live Daytona sign-in flow still requires confirmation by a user with a live sandbox. ## Risks The migration adds a new company-scoped table. A promotion or cleanup race could remove a credential or leave a sandbox active, so the service uses claims, compare-and-set transitions, and an advisory lock. The live Daytona flow needs operator confirmation because local tests do not provide a real browser sign-in. ## Model Used OpenAI Codex, GPT-5, tool use and code execution, extended reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
67001ec6eb |
chore(db): treat Drizzle migration snapshots as binary in diffs (#11254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip stores its state in a Postgres database, managed in `packages/db`. > - The schema uses Drizzle. `drizzle-kit` writes a full-schema snapshot to `packages/db/src/migrations/meta/` for every migration. > - Each snapshot is a large generated JSON file. One snapshot is about 39k lines. > - PR #11240 marked these files `linguist-generated=true`. That collapses the file view and drops the files from language stats. > - But `linguist-generated` does not change the PR line-count badge. Git still counts every snapshot line, so a PR that adds one migration shows a ~39k-line badge (two migrations show ~78k). PR #11237 shows +84,665 for this reason. > - This pull request adds `-diff` to the same files, so git treats them as binary and their lines leave the +/- count. > - The benefit is a pull request badge that reflects the real code change, not generated snapshot noise. ## Linked Issues or Issue Description No existing issue. This is a small follow-up to merged PR #11240. Description follows the enhancement template: **Problem / Motivation** PR #11240 marked `packages/db/src/migrations/meta/**` as `linguist-generated=true`. That attribute collapses the diff in the Files-changed view and removes the files from language stats, but it does not remove their lines from the PR additions/deletions badge. Each Drizzle snapshot is a full copy of the schema (~39k lines), so any PR that adds a migration still shows a huge line count. PR #11237 shows +84,665, of which ~78k are two generated snapshots. **Proposed Solution** Add `-diff` to the same glob. Git then treats the snapshots as binary. `git diff --numstat` reports `-` for these files, so their lines leave the +/- badge and GitHub shows "Binary file not shown" in place of the full JSON. **Alternatives Considered** - Keep only `linguist-generated`: leaves the misleading ~39k/78k badge on every migration PR. - Use the `binary` macro (`-diff -merge -text`): also disables EOL normalization. This repo has open CRLF/LF work, so `-text` is left unset on purpose. ## What Changed - Added `-diff` to `src/migrations/meta/**` in `packages/db/.gitattributes`. - Kept `linguist-generated=true` (language stats) and `-merge` (no auto-merge of generated snapshots). - Left `-text` unset on purpose, so end-of-line normalization stays intact. - Left the `.sql` migration files untouched, so their diffs stay visible for review. ## Verification Check the attributes and confirm the snapshot is now treated as binary: ``` git check-attr linguist-generated diff merge -- \ packages/db/src/migrations/meta/0031_snapshot.json \ packages/db/src/migrations/0009_fast_jackal.sql git diff --numstat origin/master...origin/feat/adapter-sandbox-login -- \ packages/db/src/migrations/meta/0214_snapshot.json ``` Expected: ``` packages/db/src/migrations/meta/0031_snapshot.json: linguist-generated: true packages/db/src/migrations/meta/0031_snapshot.json: diff: unset packages/db/src/migrations/meta/0031_snapshot.json: merge: unset packages/db/src/migrations/0009_fast_jackal.sql: linguist-generated: unspecified packages/db/src/migrations/0009_fast_jackal.sql: diff: unspecified packages/db/src/migrations/0009_fast_jackal.sql: merge: unspecified - - packages/db/src/migrations/meta/0214_snapshot.json ``` The snapshot reports `-` in numstat (binary, not counted). The `.sql` migration keeps normal diff behavior. GitHub reads the rule from the PR tree, so the badge drops on the next PR that touches these files. ## Risks Low risk. The change only affects how git and GitHub render and count generated files. It does not touch application code, the schema, or any migration. - With `-diff`, GitHub and local `git diff` no longer show a text diff for a snapshot. This is intended; the files are generated and are not reviewed by hand. The raw file is still viewable. - `-text` is left unset, so this change does not affect the CRLF/LF handling that other PRs (for example #8922) address. ## Model Used Claude Opus 4.8 (Anthropic), model ID `claude-opus-4-8`, used through Claude Code with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — not applicable; this change touches no code path, only git/GitHub file handling - [x] I have added or updated tests where applicable — not applicable; `.gitattributes` behavior is verified with `git check-attr` and `git diff --numstat` (see Verification) - [x] I have updated relevant documentation to reflect my changes — not applicable; the `.gitattributes` file documents its own rules inline - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0aa743fc30 |
build(db): clean dist before drizzle generate (#11241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Database migrations are generated by drizzle-kit, which reads the schema from the db package's built `dist/schema/*.js` > - `tsc` never deletes stale outputs, so a long-lived checkout keeps compiled schema files whose sources were deleted long ago > - A `generate` run in such a checkout sees those ghost tables and sweeps phantom `CREATE TABLE` statements into an unrelated migration > - This pull request makes `generate` clean `dist` before building, so drizzle always diffs against exactly the current schema sources > - The benefit is that no contributor can accidentally resurrect deleted tables inside a new migration ## Linked Issues or Issue Description **What happened?** Running `pnpm --filter @paperclipai/db generate` in a months-old checkout produced a migration re-creating `cloud_upstream_connections`, `cloud_upstream_runs`, and `company_secret_pools` — tables whose schema sources were deleted in #10507. The compiled copies were still in `dist/schema/`, and `drizzle.config.ts` reads the schema from `dist`, so drizzle treated them as new tables missing from the snapshot. **Expected behavior** `generate` diffs the current schema sources only; deleted tables can never reappear in a generated migration. **Steps to reproduce** 1. Build the db package, then delete a schema source file without cleaning `dist`. 2. Run `pnpm --filter @paperclipai/db generate`. 3. The generated migration re-creates the deleted table. ## What Changed - `packages/db/package.json`: the `generate` script runs `pnpm run clean` before `tsc`, so the drizzle-kit input is always a fresh build of the current sources. ## Verification - In a checkout carrying stale `dist/schema/cloud_upstreams.js` / `company_secret_pools.js` artifacts, `generate` produced a phantom migration before this change; after a clean build it reports "No schema changes, nothing to migrate". With this change the clean happens inside `generate` itself. ## Risks - Low risk: `generate` is a developer-only script; the change only adds the existing `clean` step ahead of the existing build, at the cost of a full rebuild per generate. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c0bdf26633 |
chore(db): collapse Drizzle migration snapshot diffs and block auto-merge (#11240)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip stores its state in a Postgres database, managed in `packages/db`. > - The schema uses Drizzle. `drizzle-kit` writes a full-schema snapshot to `packages/db/src/migrations/meta/` for every migration. > - Each snapshot is a large generated JSON file. One snapshot is over 200 KB. > - GitHub shows these files as 30k+ line diffs in a pull request. The diffs add no review value, because a human never edits the files. > - The files also merge badly. `drizzle-kit` computes the `id`/`prevId` chain and the full schema state, so a line-level merge of two snapshots produces a file that no real `generate` run creates. > - This pull request adds a `.gitattributes` file that marks the snapshot directory as generated and blocks its auto-merge. > - The benefit is clean pull request diffs and a loud conflict that forces the correct fix when two branches add a migration. ## Linked Issues or Issue Description No existing issue. This is a small repository-hygiene change. Description follows the enhancement template: **Problem / Motivation** Every migration adds a full-schema snapshot JSON under `packages/db/src/migrations/meta/`. These files are large and generated. GitHub renders them as 30k+ line diffs in pull requests, which buries the real change (the `.sql` migration) in noise. The files also have no meaningful line-level merge: `drizzle-kit` computes each snapshot's `id`/`prevId` chain and full schema state. **Proposed Solution** Add `packages/db/.gitattributes`: - `linguist-generated=true` on `src/migrations/meta/**` — GitHub collapses the diff and drops the files from language stats. - `-merge` on the same glob — git refuses the line-level merge and raises a conflict instead of fabricating an invalid snapshot. **Alternatives Considered** - `-diff` / `binary`: hides the diff completely and blocks text merge, but also blocks any local `git diff` and gives a worse conflict experience. `linguist-generated` keeps the file expandable and text-based, so it is the lighter option. - Do nothing: leaves the noisy diffs and the risk of a silent bad merge. ## What Changed - Added `packages/db/.gitattributes`. - Marked `src/migrations/meta/**` as `linguist-generated=true` to collapse the snapshot and journal diffs on GitHub. - Set `-merge` on the same files so git raises a conflict instead of auto-merging generated snapshots. - Left the `.sql` migration files untouched, so their diffs stay visible for review. ## Verification Run `git check-attr` against the affected files and a control `.sql` file: ``` git check-attr linguist-generated merge -- \ packages/db/src/migrations/meta/0031_snapshot.json \ packages/db/src/migrations/meta/_journal.json \ packages/db/src/migrations/0009_fast_jackal.sql ``` Expected output: ``` packages/db/src/migrations/meta/0031_snapshot.json: linguist-generated: true packages/db/src/migrations/meta/0031_snapshot.json: merge: unset packages/db/src/migrations/meta/_journal.json: linguist-generated: true packages/db/src/migrations/meta/_journal.json: merge: unset packages/db/src/migrations/0009_fast_jackal.sql: linguist-generated: unspecified packages/db/src/migrations/0009_fast_jackal.sql: merge: unspecified ``` The snapshot and journal files carry both attributes. The `.sql` migration keeps its normal diff and merge behavior. GitHub applies the rule from the pull request tree, so the collapse shows on the next pull request that touches these files. ## Risks Low risk. The change only affects git and GitHub display and merge behavior for generated files. It does not touch application code, the schema, or any migration. - `-merge` leaves the current-branch version in the working tree on conflict and marks the file conflicted. It does not insert conflict markers into the JSON. The correct resolution stays "renumber the later migration and regenerate", then commit. - Related open pull requests #879 and #8922 also add `.gitattributes` rules for migration files, but for CRLF/LF hash mismatches on Windows. If either lands, a follow-up can merge the rules into one file. There is no functional overlap with this change. ## Model Used Claude Opus 4.8 (Anthropic), model ID `claude-opus-4-8`, used through Claude Code with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — not applicable; this change touches no code path, only git/GitHub file handling - [x] I have added or updated tests where applicable — not applicable; `.gitattributes` behavior is verified with `git check-attr` (see Verification) - [x] I have updated relevant documentation to reflect my changes — not applicable; the `.gitattributes` file documents its own rules inline - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
23a1b025c2 |
feat(server): chunked resumable company import transfers (#11223)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import moves large packages into an instance, and since the upload cap rose to 1 GB, the transport is the weak point: one HTTP request, buffered fully in memory, with no resume > - A dropped connection at 90% of an 800 MB upload starts the whole transfer over, and a server restart loses all progress > - This pull request adds the server side of chunked resumable import transfers: a durable run ledger and routes that accept the same import zip as verified ~32 MB parts spooled to disk > - An interrupted transfer resumes from the parts already uploaded — across dropped connections, page refreshes, and server restarts — and peak upload memory drops from the whole package to one part > - The benefit is that large imports become reliable on real-world connections instead of all-or-nothing ## Linked Issues or Issue Description **What happened?** Large company imports travel as a single HTTP upload. On a slow or flaky connection, any interruption discards all progress and the upload restarts from zero. The server buffers the entire compressed package in memory during upload. A server restart mid-upload loses the transfer entirely. With the upload cap now at 1 GB, these failure modes govern exactly the imports the cap was raised for. **Expected behavior** A large import upload survives interruptions: already-transferred data is kept and verified, only the missing remainder is re-sent, and the server's memory use during upload is bounded by a part, not the package. **Steps to reproduce** 1. Import a multi-hundred-MB company package over a connection that drops mid-upload. 2. The upload fails; retrying starts from byte zero. 3. Repeat on an unstable connection and the import may never complete. ## What Changed - New `company_transfer_runs` table (drizzle schema + migration) and `companyTransferRunService`: one row per transfer with a content-derived idempotency key, per-part completion recorded atomically and idempotently, resume scoped to actor and direction, completed runs short-circuiting retries of identical content. - New transfer routes beside the existing import routes, same authorization: declare a sliced zip (`POST /import/transfers` — validates cap, 64 MB part ceiling, contiguity, size sums, sha256 format), upload parts (`PUT .../parts/:n` — raw body, hash-and-size verified before an atomic write to a disk spool under the instance root; re-uploads are no-op successes), poll resume state (`GET .../:id` — missing parts recomputed from disk), and apply (`POST .../:id/apply` — requires all parts, re-verifies the assembled zip against the whole-file hash fail-closed, then feeds the existing import pipeline through factored helpers rather than duplicated logic). - Hourly sweep fails and cleans spools idle for 24 h; a swept transfer honestly reports all parts missing on resume. - Strict UUID gating on run ids before any filesystem path construction. - The existing single-shot upload path is untouched; clients arrive in the follow-up PR. ## Verification - Transfer route suite (embedded Postgres): create/upload/status/apply round-trip with a real imported company, out-of-order parts, wrong-hash part rejected and unrecorded, re-upload no-op, apply-with-missing-parts rejection, resume after failure with prior progress intact, assembled-hash mismatch failing closed with spool deletion, actor scoping 404s, async-job apply, sweep followed by honest resume. - Ledger suite (embedded Postgres): part idempotency, actor/direction scoping, completed-run short-circuit, cancelled runs staying cancelled. - Existing portability route suite unchanged and green; server + db typechecks clean. Exact counts in the PR checks. ## Risks - New routes are additive; the existing import path is untouched. The transfer routes carry the same board authorization as the import routes they sit beside. - Disk spool: bounded by the existing upload cap per transfer, cleaned on success, failure, hash mismatch, and by the 24 h sweep. Spool paths are strict-UUID-gated. - The apply step still materializes the assembled zip in memory once (same profile as today's single-shot import at apply time); upload-time memory drops to one part. - Known limitation, deliberate: transfers are keyed on content alone, so identical package content cannot be imported twice without re-exporting (surfaced explicitly to the caller). Acceptable for v1; noted for review. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
815e49bb7c | feat: make chat-style tasks the default experience (#11101) | ||
|
|
0a511ed1b0 |
feat(apps): support multiple provider connections (#11060)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps subsystem connects company tools through governed provider connections. > - A company can need more than one account for the same provider. > - The current database constraint and Apps flow assume one named connection per company. > - New quarantined actions also need an explicit review decision before activation. > - This pull request supports multiple provider connections and complete action review decisions. > - The benefit is safer access control and a clear multi-account Apps workflow. ## Linked Issues or Issue Description Refs: #11040 **Subsystem affected** Cross-cutting. This change affects the Apps UI, the tool access API, the shared request contract, and the database schema. **Problem or motivation** The connection name constraint prevents a company from keeping more than one connection for a provider. The Apps UI also reuses an existing OAuth connection when a user asks to connect another account. Action review can enable selected entries without recording a decision for every quarantined action. **Proposed solution** Remove the company and connection name uniqueness constraint. Let users open, count, edit, and create multiple provider connections. Require the finish request to cover every quarantined action exactly once before the server activates reviewed entries. **Alternatives considered** The UI could generate unique internal names and keep the database constraint. This would preserve a one-connection assumption in the data model and would make display names part of identity. The server could also infer review decisions from enabled actions. This would not distinguish a reviewed disabled action from an action that the user did not review. **Roadmap alignment** This change extends the completed MCP Tool Gateway and Apps milestone. It also supports the Connected Apps roadmap item. It follows the navigation and connection management work in #11040. ## What Changed - Remove the company-scoped connection name uniqueness index with an ordered and idempotent migration. - Add a reviewed action list to the finish-app contract and reject incomplete or duplicate review decisions. - Activate reviewed entries and keep unreviewed quarantined entries blocked. - Enable a completed connection and preserve the company and connection scope in all updates. - Show provider connection counts and open the provider setup page from Browse. - Let users edit existing connections or connect another account without reusing an active OAuth connection. - Update focused server and UI coverage for multiple connections and action review. ## Verification - Ran the focused Apps UI suite. All 116 tests passed in 11 files. - Ran the focused server and CLI suite. All 276 tests passed in 3 files. - Ran `pnpm --filter @paperclipai/db check:migrations`. The migration safety check passed. - Ran `pnpm -r typecheck`. All projects passed. - Ran `pnpm build`. All projects built successfully. - Ran `pnpm test:run`. It passed 3,735 tests and skipped 4 tests. One worktree-safety assertion failed because the execution workspace reloads its worktree marker. The same test passed with an isolated non-worktree marker. - Ran `pnpm check:token-gates`. It reports 12 existing violations in the unchanged `PaperclipOrbit3D.tsx` file from the target branch. - Started the six affected Playwright specifications. Chromium could not start because the host does not provide `libatk-1.0.so.0`. The GitHub e2e jobs will verify these specifications. - GitHub Actions passed every final-head CI gate, including all three e2e shards and the aggregate `e2e` and `verify` jobs. - Greptile reviewed final commit `9af9200426` at 5/5 with zero review threads. ## Risks - Removing the name uniqueness index permits duplicate display names. Stable connection IDs and UIDs remain unique within a company. - The finish-app endpoint accepts the new review field as optional for backward compatibility. When clients send it, the server requires a complete decision for all quarantined actions. - Multiple OAuth connections depend on the explicit new-connection route flag. Focused tests cover active and draft connection reuse. - The migration is ordered after migration 0210. Its `DROP INDEX IF EXISTS` statement is safe to repeat. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with the `gpt-5.6-sol` model assisted this change. The agent used repository tools, code execution, test execution, and agentic reasoning. The Codex runtime manages the context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
52b8741b8e |
perf(server): cut steady-state DB hot paths in dashboard, attention, and productivity sweeps (#10992)
<!-- ASD-STE100 --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server keeps fleet health with periodic sweeps and shows a dashboard with run activity > - The Paperclip instance became slow again after the first round of recovery-sweep indexes landed > - Live profiling found four steady-state hot paths that read much more data than they use > - This pull request bounds the dashboard recursion, adds the missing taskKey index, and narrows two wide reads > - The benefit is a large drop in constant database load and a responsive server ## Linked Issues or Issue Description **Describe the bug** The server becomes slow while agents work. Live query sampling shows four hot paths: 1. The dashboard run-activity recursive CTE reads every run a company ever had on each call. One call takes 2.85 seconds. The UI calls it after almost every fleet event through the dashboard and sidebar-badges routes. 2. The productivity-review sweep runs each 30 seconds. Its run-scope filter is `issueId OR taskId OR taskKey` on the run context JSONB. No index exists for `taskKey`. The planner must detoast every run snapshot for the agent. One query takes 444 ms and the sweep makes one for each of ~152 candidate issues. 3. The attention failed-run section selects the full `context_snapshot` for every run newer than the oldest exhausted run. That fetch moves 29 MB for each feed build. 4. The retention sweep pages the attention feed with a cursor. Each page makes a full feed rebuild. **Expected behavior** Periodic sweeps and dashboard queries read only the data they use, and use indexes. **Actual behavior** The database stays saturated. Users see a slow server. ## What Changed - `server/src/services/dashboard.ts`: bound both arms of the `recovered_runs` recursive CTE to the chart window. A retry is always newer than the run it retries, so the bound cannot change visible chart data. Live time went from 2,852 ms to 54 ms. - `packages/db/src/migrations/0210_heartbeat_context_taskkey_index.sql`: add the `taskKey` expression index that completes the issueId/taskId/taskKey trio. With all three, the planner uses a BitmapOr. Live time for the productivity run-scope query went from 444 ms to 1.9 ms. - `packages/db/src/schema/heartbeat_runs.ts`: mirror the new index in the Drizzle schema. - `server/src/services/productivity-review.ts`: select only the seven run fields the evidence code reads. Before, the query pulled full rows with `result_json` (up to 43 kB per row, 100 rows per issue). - `server/src/services/attention.ts`: project `issueId`/`taskId` text fields instead of the full `context_snapshot` in the failed-run newer-runs query (29 MB per feed build before). - `server/src/index.ts`: the retention sweep now builds the attention feed once per company with `all: true` instead of one full rebuild per cursor page. - `packages/db/src/heartbeat-context-snapshot-index-migration.test.ts`: cover the new index and re-run migration 0210 statements to prove idempotency. ## Verification - `pnpm --filter @paperclipai/db typecheck` (includes migration numbering and safety checks) — pass. - `npx tsc --noEmit` in `server/` — pass. - `npx vitest run packages/db/src/heartbeat-context-snapshot-index-migration.test.ts` — pass (embedded Postgres, full migration chain, planner assertions, idempotent re-run of 0209 and 0210). - `npx vitest run` on attention, dashboard, productivity-review, decision-retention, issue-blocker-attention, and issue-review-attention test files — 72/72 pass. - Live EXPLAIN ANALYZE before/after numbers are in the What Changed list. ## Risks - Migration 0210 builds one btree index without CONCURRENTLY inside the transactional migration runner. The table is not in the large-table bucket. The 0209 twin built in seconds on a 100k-row live table. - The CTE bound excludes retry ancestors that are older than the chart window. Those rows are not visible to the chart query, so chart output does not change. - The attention projection changes JSONB scalar handling in one edge case: a non-string `issueId`/`taskId` value now casts to text instead of reading as absent. These keys are always strings in practice. - The retention sweep now holds one full feed in memory per company. The cursor loop already accumulated all pages into one array, so peak memory is unchanged. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic, Mythos-class tier, extended thinking + tool use) via Paperclip agent runtime. - [x] I searched existing PRs and issues and this change is not a duplicate. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e591e75b8c |
perf(db): index context_snapshot/payload issueId lookups used by recovery sweeps (#10969)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server runs a recovery sweep every 30 seconds. The sweep makes
sure assigned issues do not stall.
> - The sweep reads the latest heartbeat run for each candidate issue.
It filters `heartbeat_runs` on `context_snapshot ->> 'issueId'`. No
index covers this expression.
> - Each lookup scans ~26k rows and detoasts each row's JSONB context
(~645 ms per query, measured with EXPLAIN ANALYZE). With ~460 candidate
issues, one sweep schedules ~500 seconds of database work every 30
seconds.
> - The database saturates. Users see the whole server as slow.
> - This pull request adds expression indexes so these lookups become
index scans.
> - The benefit is that the recovery sweep drops from ~500 seconds of
database work per tick to milliseconds, and the server becomes
responsive again.
## Linked Issues or Issue Description
**What happened?**
The server became slow for all users. `pg_stat_activity` sampling showed
the same two recovery-sweep queries active continuously (15/15 samples).
Each `getLatestIssueRun` call filtered `heartbeat_runs` on the unindexed
expression `context_snapshot ->> 'issueId'`, scanned ~26k rows, and
detoasted a 2.1 GB TOAST region. `heartbeat_runs` accumulated 10.8
billion sequential tuples read. `hasActiveExecutionPath` also scanned
`agent_wakeup_requests` (1.8M rows) on the unindexed expression `payload
->> 'issueId'`.
**Expected behavior**
Per-issue run lookups in the recovery sweep complete in milliseconds.
The sweep finishes well inside its 30-second interval. Background
maintenance does not degrade interactive latency.
**Steps to reproduce**
1. Run a board with several hundred issues in `todo` / `in_progress` /
`in_review` and a large `heartbeat_runs` table with big
`context_snapshot` payloads.
2. Let the heartbeat scheduler run its 30-second recovery sweep.
3. Observe `pg_stat_activity`: the per-issue `heartbeat_runs` lookups
run continuously; `EXPLAIN ANALYZE` shows a filter on `context_snapshot
->> 'issueId'` removing tens of thousands of rows per call.
**Paperclip version or commit**
master (
|
||
|
|
f554d67377 |
fix(server): add explicit review verdict policies (#10931)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Issues use `in_review` to request a final decision from an authorized writer > - The server rejected an assignee agent that tried to close its own review, even when the issue had no independent-review rule > - This rejection stopped the default agent workflow and did not represent the configured execution-stage rules > - Paperclip needs an open default and explicit issue-level constraints for teams that require an independent or human verdict > - This pull request removes the unconditional rejection and adds `anyone`, `not_creator`, and `human_only` review policies > - The benefit is a working default path with opt-in, authenticated verdict controls ## Linked Issues or Issue Description Refs #10635, #4429, and #10671. The related public work covers execution-stage independence, self-approval fallback behavior, and durable review paths. This change is distinct. It controls who can resolve an issue review verdict. It keeps configured execution stages active. ## What Changed - Added a nullable `review_policy` issue column. Null has the same meaning as `anyone`. The migration does not backfill existing issues. - Added shared create, update, response, and compact issue contracts for `anyone`, `not_creator`, and `human_only`. - Removed the unconditional agent self-approval rejection for `in_review` issues. - Added one reusable verdict-actor check for terminal status changes and pending interaction accept or reject actions. - Used the authenticated principal type for `human_only`. Agent keys and run tokens remain agent principals. - Used the latest transition into `in_review` to identify the requester for `not_creator`. - Added actionable 403 responses that name the policy, the allowed actor, and the next step. - Kept the configured execution-stage transition and signoff behavior. - Added focused contract, helper, status-route, interaction-route, and execution-stage regression tests. - Updated the implementation specification for the new issue field. ## Verification - `pnpm exec vitest run packages/shared/src/validators/issue.test.ts server/src/__tests__/issue-review-policy.test.ts server/src/__tests__/issue-stalled-review-decision-routes.test.ts --reporter=dot` passed: 42 tests. - `pnpm --filter @paperclipai/shared typecheck` passed. - `pnpm --filter @paperclipai/db typecheck` passed, including migration numbering and safety checks. - `pnpm --filter @paperclipai/server typecheck` passed. - `pnpm run typecheck:build-gaps` passed across server, CLI, plugin SDK/examples, plugin wiki, and UI. - `git diff --check origin/master...HEAD` passed. - SecurityEngineer review approved the authenticated-principal checks and accepted policy-relaxation tradeoff with no required changes. - Greptile reviewed the latest head at 5/5 with zero inline comments or follow-ups. - The latest-head GitHub rollup passed build, typecheck, server/workspace tests, serialized suites, canary, e2e, and external security checks. ## Risks - The migration adds one nullable text column. It has no default and no backfill. - `not_creator` reads the latest recorded transition into `in_review`. It denies the verdict when it cannot identify the requester. - Agents can change or relax `reviewPolicy` when they have issue write access. This is intentional for this issue-level control. - Null and `anyone` do not add a database query to the verdict path. - Configured execution-stage checks still run after the issue-level policy check. > This work aligns with the completed "Agent Reviews and Approvals" and "Enforced Outcomes" roadmap items. It does not add a new roadmap capability. ## Model Used - OpenAI Codex, GPT-5. The exact deployment ID and context-window size are not exposed to the agent. The run used reasoning, repository tools, code execution, and GitHub CLI access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e43f187cad |
feat(secrets): add human-approved secret proposals (#9934)
## Thinking Path > - Paperclip is the control plane people use to manage AI-agent companies. > - Agents can encounter credentials during work. > - Directly creating live secrets or bindings would bypass human governance. > - Proposal records must remain inert and separate from live secret resolution until an authorized human approves them. > - Approval must reuse the existing secret-create and protected agent-config write paths. > - This pull request adds the propose, review, approve, and reject lifecycle. > - The benefit is that agents can safely hand credentials into Paperclip without exposing plaintext or gaining authority to activate them. ## Linked Issues or Issue Description Follow-on to #9921, which established run-bound agent secret access. **Problem / motivation:** Agents can receive credentials during work. There is no governed way for them to propose a credential or binding without exposing plaintext in work artifacts or immediately creating live access. **Proposed solution:** Store agent-authored proposals outside live secret tables. Encrypt each proposed value and register exact-value redaction when Paperclip receives it. Require an authorized human to approve or reject each proposal. Approval executes through the normal write paths as the human approver. Binding proposals can target only the proposer or its downward reporting chain under the restrictive V1 policy. **Alternatives considered:** We rejected live secrets with a `proposed` status. That design would put untrusted rows in resolver, list, and sync paths. It would also allow uniqueness squatting. We rejected direct agent binding writes because a binding is an agent-config write and must keep the existing human permission gate. **Roadmap alignment:** This change extends the run-bound agent secret-access foundation in #9921 with a governed proposal workflow. ## Security Verdict Q0 SecEng verdict: **PASS-with-required-changes**. The review accepted the separate proposal-table design and required the implementation to: - fail closed unless both encryption and exact-value run redaction registration succeed; - scrub ciphertext idempotently on reject, withdraw, and expiry, with audit-visible state; - treat agent justification as hostile input and foreground action, target, provenance, and approver permissions; - snapshot and re-check the target agent plus reports-to chain at approval to prevent org-chart laundering; - make cascade approval atomic and fail closed if either secret creation or binding authorization fails; - deny low-trust, `skill_test`, `task_bridge`, and non-run-bound sources consistently; and - execute approval through the normal human secret/config write paths, including protected-change gates. Those requirements are implemented and covered by focused service, route, and UI tests. Residual V1 risk remains the accepted 14-day encrypted retention window. Proposal-time redaction also cannot clean a value that leaked before the propose call. ## What Changed - Added `company_secret_proposals`, migration `0207`, shared proposal contracts, and a state-machine service for create, approve, reject, withdraw, cascade, expiry, and ciphertext scrubbing. - Added run-bound agent proposal routes and board review routes. The routes derive provenance from authentication and enforce source restrictions, company isolation, chain-of-command checks, approval-as-approver, wake-on-resolution, and dual audit trails. - Added durable per-run exact-value redaction registration so proposal values remain redacted on later read surfaces. - Added the Secrets **Proposals** tab and agent configuration **Proposed access** rows. The UI shows fingerprint and length only. It also frames agent justification as untrusted input, runs permission preflight, supports approve and reject actions, and confirms cascades. - Updated OpenAPI, agent skill guidance, API reference documentation, and focused server and UI regression coverage. - Rebased the branch onto current `master` and renumbered the proposal migration after `0206`. ## QA Acceptance Results Q5 QA verdict: **PASS — 9/9 acceptance criteria met**, with one Minor non-blocking follow-up. - **AC1:** proposed values never echo, never appear in live lists/resolvers, and expose only fingerprint + length to board reviewers. - **AC2:** restrictive `self_and_reports` matrix passes: self/downward allowed; upward/lateral denied. - **AC3:** secret approval uses the normal create path, honors rename overrides, records proposer/approver provenance, and scrubs ciphertext. - **AC4:** approved bindings materialize and resolve through the target agent's runtime list/fetch routes. - **AC5:** pending-secret bindings require cascade; cascade succeeds atomically and permission failures leave nothing applied. - **AC6:** reject, withdraw, dependent rejection, and expiry paths scrub ciphertext and preserve reasons/audit state. - **AC7:** token/source and approver denial matrix passes through live checks plus focused route tests. - **AC8:** proposal lifecycle events and reused `secret.created`/config-write events form the required dual audit trail; origin-issue notification and wake are queued. - **AC9:** both review surfaces render and execute correctly; UI approval materializes the binding. QA also confirmed zero plaintext occurrences for all exercised proposal values in server logs. The single finding is that the company-level `bindingTargetPolicy` toggle is not wired yet. V1 is hardcoded to the restrictive `self_and_reports` policy. The matrix is correct and the follow-up is tracked separately, so QA classified it as non-blocking. ## Verification - Focused server proposal and redaction suite: 83 tests pass. - Focused proposal review UI suite: 54 tests pass. - Embedded-Postgres migration reapply test: 1 test passes with the documented 30-second timeout. - `pnpm --filter @paperclipai/db typecheck` passes, including migration numbering and safety checks. - `pnpm --filter @paperclipai/shared typecheck` passes. - `pnpm --filter @paperclipai/ui typecheck` passes. - `pnpm check:token-gates` passes with all gates clean. - Q5 exercised the complete propose, review, approve, bind, and runtime-resolve flow over real HTTP, JWT, and database paths. It verified 9/9 acceptance criteria. ## Risks - Proposal ciphertext is retained encrypted for up to 14 days while pending. Terminal-state and expiry scrub paths reduce but do not remove server-compromise risk during that window. - The V1 target policy is restrictive but not yet company-configurable. A separate follow-up owns that change. - A new migration can require another renumber if another migration lands before maintainers merge this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent. The exact runtime model ID and context-window size are not exposed. The agent used reasoning, repository editing, terminal execution, Paperclip API, and GitHub CLI capabilities. Q3 UI work also records Claude Opus 4.8 assistance in its commit trailers. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
678728f650 |
feat: maintained in_review review-path contract + stalled-review actions (#10675)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents move issues to `in_review` and rely on a "review path" (an interaction, an approval, a monitor, or a named reviewer) to tell them who decides next. > - That review path can silently disappear. A user comment supersedes the pending interaction, a monitor is exhausted, or a run ends without restoring a path. The issue then sits in `in_review` with nobody reviewing it and no visible action. > - Such issues become invisible zombies. Nobody knows a decision is owed, so the work stalls forever. > - This pull request makes the review path a maintained invariant, exposes a `reviewAttention` surface, and gives every stalled review three inline actions in the UI. > - The benefit is that an `in_review` issue always shows who reviews it, or shows an amber "nobody is reviewing this" notice with one-click Approve, Request changes, and Send back to work. ## Linked Issues or Issue Description This pull request describes the problem inline. The tracking issue is internal. **Subsystem affected** The review and attention loop that agents and humans share: the `in_review` status, the `reviewAttention` surface, the /decisions attention feed, and the issue-page review panel. **Problem or motivation** Agent-owned issues in `in_review` can lose their last review path. A user comment supersedes the pending interaction. A monitor is exhausted. A run ends without restoring a path. The issue then sits in `in_review` with no reviewer and no visible action. It becomes an invisible zombie and the work never progresses. **Proposed solution** Maintain the review path as a server invariant. Expose a `reviewAttention` field that says what is under review, who decides, and since when. Render a persistent review panel on the issue page and inline actions on the /decisions feed. Keep human PATCHes into `in_review` ungated, but record the requesting user so the panel never renders empty. **Alternatives considered** A pure background auto-recovery sweep. This stays opt-in and is not enough on its own, because it is invisible to the human. A bare status banner. This is rejected, because it gives no action to resolve the stall. **Roadmap alignment** This improves the core review and attention loop that both agents and humans use every day. ## What Changed - **Server — maintained review-path invariant:** when an issue enters or sits in `in_review`, the server derives and persists a review path (interaction, approval, monitor, or the requesting user) and recovers a stale path with one bounded wake instead of leaving the issue pathless. - **Server — `reviewAttention` surface:** a new field describes what is under review (bound target with links), who decides, since when, and whether the review is stalled. Stalled agent-assigned reviews are now included in the attention feed. - **Server — inline stalled-review decisions:** secured routes let a permitted responder Approve (→ `done`), Request changes (→ `todo` + wake carrying the note), or Send back to work (→ `todo` + wake) directly from the attention feed. - **Server — resume-intent wake:** an `in_review -> todo` transition now wakes the assigned agent so a resumed review is not dropped. - **Server — user-entry symmetry:** user PATCHes into `in_review` stay ungated (no 422 for humans) and record the requesting user, who becomes the named responder when no other path exists. - **UI — review panel:** a persistent `IssueReviewPanel` renders above the thread whenever status is `in_review`. The covered state shows the bound target, responder, and outcomes and hoists the pending interaction/approval card. The stalled state shows the amber notice plus the three actions. - **UI — decisions card actions:** the same three actions render inline on the /decisions `AttentionQueueRow`. - **UI — responsive fix:** the stalled action row stacks to full-width buttons at phone width and returns to a horizontal row at `sm` and up. New 390px stories capture the phone layout. ## Verification - `cd ui && npx vitest run src/components/IssueReviewPanel.test.tsx src/components/AttentionQueueRow.test.tsx src/lib/attention.test.ts src/api/issues.test.ts` — 91 tests pass. - Server suites added and updated: `issue-review-attention`, `issue-stalled-review-decision-routes`, `review-path-recovery`, `recovery-observability`, and related route/liveness tests (run by CI). - A designer reviewed the UI at 390px and desktop in light and dark themes on both the issue-page panel and the /decisions card. The stalled action row stacks cleanly at phone width with no overlap and keeps the horizontal row on desktop. ## Risks - **Migration:** adds migration `0200` (next after master `0199`, no renumber). It extends the agent-wakeup-requests schema and is additive. - **Behavioral shift:** `in_review -> todo` now dispatches a wake. This is intended (resume intent) and covered by tests. - **Authz:** the inline decision routes are permission-gated. Only a permitted responder sees and can trigger the actions. - Overall risk is moderate and contained to the review and attention loop. ## Model Used - Claude, Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f91a6e27c0 |
feat(issues): contain cross-issue agent side effects (#10837)
## Thinking Path > - Paperclip is the control plane that coordinates autonomous agent work. > - Agents need to collaborate on issues beyond their current assignment. > - Cross-issue comments and updates are useful, but an unbounded run can create cascading side effects. > - The control plane must preserve company-wide collaboration while containing each run's influence. > - Comment attribution must also show the responsible user and the acting agent in audits. > - This pull request adds run-bound cross-issue containment, attribution, and agent-class wake rules. > - The benefit is safer collaboration without restoring issue-assignee ownership restrictions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent-authenticated issue comments, updates, reopen behavior, and assignee wake routing. **Subsystem affected** Cross-cutting: server routes and services, shared contracts, database schema and migration, and implementation documentation. **Current behavior** An authenticated agent can collaborate across company issues, but one heartbeat run has no per-run side-effect boundary. Comment records also do not persist the responsible user separately from the acting agent. **Proposed behavior** Require a valid heartbeat run for agent cross-issue comments and updates. Audit each attempt and cap a run at 20 cross-issue effects. Keep the cap in log-only mode until it automatically changes to enforcement at 2026-08-11 00:00 UTC. Preserve same-issue writes. Use agent-class wakes for agent comments. Keep same-run completion comments from reopening completed work. Record the responsible user on agent-authored comments and activity. **Reason and benefit** Agents can collaborate on other issues without an assignment gate, while each run has an atomic and inspectable side-effect limit. Operators can identify both the acting agent and the responsible user. **Breaking changes** After 2026-08-11 00:00 UTC, the twenty-first cross-issue comment or update from one heartbeat run returns a containment error. Agent cross-issue writes without valid run context are rejected. The migration is additive and backfills existing agent-authored comment attribution where the source data is available. ## What Changed - Added an atomic per-run counter for cross-issue agent comments and updates. - Added audit events for allowed and rejected cross-issue effects. - Added the automatic log-only to enforcement flip at 2026-08-11 00:00 UTC. - Added responsible-user attribution to agent-authored comments, activity records, shared types, and validators. - Added an additive migration and migration coverage for existing comments. - Updated reopen, resume, and wake behavior so agent comments create agent-class wakes and same-run completion comments remain inert. - Updated the implementation specification and regression coverage. ## Verification - `pnpm exec vitest run server/src/__tests__/cross-issue-influence-limit.test.ts server/src/__tests__/issue-comment-attribution-audit-routes.test.ts server/src/__tests__/issue-comment-reopen-routes.test.ts packages/db/src/issue-comment-on-behalf-migration.test.ts` — 97 tests passed. - `pnpm -r typecheck` — passed, including migration safety checks. - `pnpm test:run` — server batch: 3,364 passed and 2 skipped; UI batch: 3,504 passed. One unrelated CLI doctor test warned because this agent runtime injects static AWS credentials. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` — 8 tests passed and confirmed the CLI failure was ambient-environment sensitive. - `pnpm build` — passed. ## Risks - The fixed enforcement timestamp changes production behavior automatically on 2026-08-11 00:00 UTC. Audit logs before that time provide rollout visibility. - The per-run counter serializes on the heartbeat-run row. This prevents concurrent attempts from racing past the cap but adds a small lock scope for cross-issue writes. - Existing comments can only be backfilled when their acting run or agent attribution is recoverable. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 in the Codex agent runtime. The runtime did not expose a context-window size. Reasoning, shell tools, code editing, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ded813ad6f |
feat(interactions): add governed agent addressees (#10252)
## Thinking Path > - Paperclip is the control plane that lets humans govern companies of AI agents. > - Issue-thread interactions are the structured handoff point for confirmations, questions, suggested tasks, and other governed decisions. > - Those interactions previously assumed that only board users could resolve them, preventing one agent from explicitly addressing another agent for a response. > - Agent resolution needs company-level governance, auditable resolver identity, safe terminal-state handling, and attention routing so authorization is enforced server-side rather than inferred from UI behavior. > - This pull request adds governed agent resolution, withdrawal and terminal expiry semantics, explicit agent addressees, lifecycle reconciliation, and attention-feed filtering. > - The benefit is that agents can participate in structured decisions without weakening board control, company isolation, wake behavior, or audit invariants. ## Linked Issues or Issue Description ### Subsystem affected Issue-thread interactions across database, shared contracts, server authorization/services, adapter callbacks, agent skill guidance, API docs, and UI governance surfaces. ### Problem or motivation Structured interactions were board-only, had no explicit agent addressee, and lacked durable withdrawal/terminal-expiry semantics. That made peer-agent decisions impossible to authorize and audit safely. ### Proposed solution Persist requested/effective resolver policy and addressee identity, enforce company governance and eligible agent resolution, reconcile addressee lifecycle changes, expose withdrawal and terminal expiry, and route attention to the intended active agent with board fallback. ### Alternatives considered Implicitly authorizing the issue assignee or mentioned agents was rejected as ambiguous and difficult to audit. Using comments alone was rejected because it loses structured outcomes and continuation behavior. ### Roadmap alignment Supports the ROADMAP direction for lightweight leadership-agent communication that still resolves into governed decisions and work objects. ### Additional context Public GitHub issue/PR search found no duplicate implementation; open PR search for interaction resolver governance and agent addressees only returned this PR. ## What Changed - Add company-scoped interaction resolver governance contracts and persistence. - Add requested/effective resolver policy, resolver identity, withdrawal, and terminal-expiry behavior. - Add explicit `addresseeAgentId` validation, authorization, persistence, lifecycle reconciliation, API documentation, and skill guidance. - Route pending addressed interactions to the intended invokable agent and fall back to board attention when that agent becomes ineligible or is deleted. - Preserve sandbox callback identity fields required by governed resolution paths. - Add migrations `0193` and `0194` plus route, service, attention, adapter, CLI, and UI coverage. - Add governance state and company settings UI, including responsive mobile behavior and distinct withdrawn/expired audit presentation. ## Verification - `pnpm check:token-gates` — passed. - `pnpm -r typecheck` — passed, including migration numbering and safety checks. - `pnpm test:run` — feature/server and UI workspace suites passed; one unrelated CLI AWS doctor test observed injected static AWS credentials and warned instead of passing. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts --project paperclipai` — 8 tests passed, confirming the failure was environment-sensitive. - `pnpm build` — passed. - Latest rebased head `e24cece6be9f1877bdbac7691bcb44fd583c0161` completed all GitHub CI jobs successfully. ## Risks - Migrations add interaction and company-governance fields; numbering is conflict-free on current `master`, additive statements are idempotent, and migration safety checks pass. - Agent authorization behavior expands beyond board-only resolution, but defaults remain board-only and coverage exercises company boundaries, resolver eligibility, lifecycle invalidation, wake behavior, withdrawal, expiry, and attention fallback. - Attention routing depends on current agent invokability; reconciliation and read-time filtering prevent stale addressees from retaining visibility or resolution authority. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.6-sol` with reasoning, terminal tool use, code execution, Git/GitHub integration, and Paperclip control-plane tools. Context-window metadata was not reported by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
cb52f0b750 |
db: env-configurable client options; parallelize attention feed queries (#10795)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The server stores all state in PostgreSQL through Drizzle and the
postgres.js driver
> - Self-hosted installs run Postgres on localhost, so per-query latency
is near zero; hosted installs often attach Postgres over a network,
sometimes through a transaction-mode pooler
> - The DB client passes no options to the driver, so operators cannot
disable prepared statements or tune the pool without a source edit, and
the deploy docs told them to edit `client.ts`
> - The attention feed also runs its related-data lookups one after
another, so its latency grows as queries × network round trip
> - This pull request adds optional environment configuration for the DB
client and batches the independent attention-feed lookups with
`Promise.all`
> - The benefit is that network-attached deployments get correct pooler
support and a much faster attention feed, while self-hosted behavior
does not change
## Linked Issues or Issue Description
No public issue exists for this; description follows the bug report
template:
**What happened?**
On deployments where PostgreSQL is network-attached (managed providers,
pooled endpoints), the attention feed endpoint is slow:
`attentionService.list()` awaits ~15–20 queries strictly in sequence, so
a 70ms round trip turns into more than one second of pure network wait
per call. Separately, connecting through a transaction-mode pooler
(pgbouncer, Supavisor port 6543, Neon `-pooler` hosts) requires
disabling prepared statements, and the only documented way was to
hand-edit `packages/db/src/client.ts` — which `doc/DATABASE.md` itself
tells operators not to do.
**Expected behavior**
The DB client is configurable from the environment (prepared statements,
pool size, timeouts) with driver defaults when unset, and hot read paths
do not multiply network latency by issuing independent queries
sequentially.
**Steps to reproduce**
1. Run the server with `DATABASE_URL` pointing at a Postgres instance
with ~70ms round-trip latency.
2. Open the attention feed (`GET /companies/:companyId/attention`) and
measure response time — it exceeds one second even with little data.
3. Try to connect through a transaction-mode pooler: there is no
supported configuration to disable prepared statements.
## What Changed
- `packages/db/src/client.ts`: `createDb` accepts a
`DatabaseClientOptions` argument and reads optional env config —
`DATABASE_PREPARED_STATEMENTS`, `DATABASE_POOL_MAX`,
`DATABASE_IDLE_TIMEOUT_SECONDS`, `DATABASE_CONNECT_TIMEOUT_SECONDS`.
When nothing is set, no option is passed to the driver and behavior is
identical to the previous bare `postgres(url)`.
- `packages/db/src/client-options.test.ts` (new): env parsing and
driver-option mapping tests, including malformed-value rejection.
- `server/src/services/attention.ts`: the independent related-data
lookups in each feed section now run under `Promise.all` (issue
summary/image/plan-document maps, decision bundle titles, blocked-issue
maps, the newer-runs scan). Section order, item assembly, and query
shapes are unchanged.
- `doc/DATABASE.md` and `docs/deploy/database.md`: the edit-source
pooling instruction is replaced with the env toggle, plus a short
client-tuning reference.
## Verification
- `pnpm --filter @paperclipai/db exec vitest run
src/client-options.test.ts` — 6 tests pass.
- `pnpm --filter server exec vitest run
src/__tests__/attention-service.test.ts` — 22 tests pass.
- `pnpm --filter server exec vitest run
src/__tests__/decisions-service.test.ts
src/__tests__/decision-training.test.ts` — 45 tests pass; this covers
the call path that runs `attentionService.list()` inside
`db.transaction`, where postgres.js serializes queries on the reserved
connection.
- `tsc` reports no errors in the changed files.
## Risks
- Low risk for self-hosted installs: with no env vars set,
`postgres(url, {})` receives an empty options object, which postgres.js
treats the same as no options — driver defaults throughout.
- The `Promise.all` batches only group queries that had no data
dependency on each other; on the transaction call path the driver still
executes them one at a time on the reserved connection, so transactional
semantics are unchanged.
- Malformed env values now fail fast at startup with a clear message
instead of being silently ignored; this is intentional and only affects
operators who set the new variables.
## Model Used
Claude Fable 5 (`claude-fable-5`), Anthropic — via Claude Code CLI,
extended thinking enabled, tool use (test execution, live latency
measurement against a network-attached Postgres to size the problem).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (searched "prepared statements", "pgbouncer", "pool",
"attention feed", "lockfile" — closest matches are #10573/#10787
lockfile chores, unrelated to this change)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|