mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
8676795188cf7c623f7d115733a0945035c0451e
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8676795188 |
ci: split e2e PR lane into three shards (#10629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The pull request workflow protects changes with a Playwright e2e lane. > - That lane already uses a weighted file partition so slow specs do not cluster by test count. > - Recent green PR runs showed the two e2e shard jobs were slower than the next slow required lane. > - The largest spec is indivisible, so a third shard lets that spec run alone and lets the rest split by duration. > - This pull request changes only the PR e2e shard matrix and the guard test. > - The benefit is a shorter expected PR critical path while the required `e2e` aggregate check name stays stable. ## Linked Issues or Issue Description Refs #9923 **What existing behavior does this improve?** The `pull_request` workflow Playwright e2e lane. **Subsystem affected** Cross-cutting: GitHub Actions CI and test scripts. **Current behavior** The PR workflow runs the weighted Playwright e2e partition across two jobs. Recent green runs showed those jobs as the slowest required checks. **Proposed behavior** The PR workflow runs the same e2e spec set across three weighted jobs. The aggregate required check stays named `e2e`. **Reason and benefit** The third shard lets the slow smoke-lab spec run alone while the rest of the catalog stays balanced. This should shorten the PR critical path. The win is bounded by fixed per-job setup time. **Breaking changes** None. The required aggregate check contract is preserved. ## What Changed - Change the PR e2e shard matrix from two entries to three entries. - Update the shard guard test to expect three shards. - Floor the balance bound at the largest single spec weight. - Assert that the workflow does not define more shard indexes than `SHARD_COUNT`. ## Verification - `node --test ./scripts/__tests__/e2e-shard.test.mjs` passes with 6 tests. - The recorded-weight partition is complete and non-overlapping: 168.0s, 116.5s, and 114.4s. - I checked `ROADMAP.md` and found no overlapping roadmap-level core feature. - I searched public GitHub PRs and issues for related e2e shard work. I found related PR #9923 and no open duplicate for this branch or change. ## Risks - This adds one extra GitHub Actions runner to the PR e2e lane. - The wall-clock win is bounded by fixed per-job setup. - Behavior risk is low because the aggregate required check remains named `e2e`. ## Model Used OpenAI Codex, GPT-5, tool-enabled coding agent in this repository. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> |
||
|
|
6401f4f78c |
fix(sandbox): bundle git copy-back against the merge-base so diverged/reset workspaces still import (#10601)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - An agent that runs in a sandbox has its workspace copied back to the host when the run ends, so its work persists — the copy-back ships a git bundle of the sandbox's commits > - The bundle is created as a thin delta, `git bundle create HEAD --not <baseSha>`, which records `baseSha` (the host workspace HEAD captured at export) as a prerequisite the host must already hold > - That assumption breaks when the sandbox HEAD has diverged from `baseSha`, or when a shared host workspace no longer holds `baseSha` at import time — then `git fetch` on the host hard-fails and the entire run is lost even though the agent finished its work > - This pull request bundles against the merge-base of `baseSha` and the sandbox HEAD (with a full-bundle fallback), which the host can satisfy in those cases > - The benefit is that copy-back no longer discards a completed run's work over a base the host can't reconcile ## Linked Issues or Issue Description **What happened?** A sandbox agent run completed its work, then failed during workspace finalize: ``` git -C <host workspace> fetch --force <git-delta.bundle> refs/…/export:refs/…/imported error: Could not read <baseSha> fatal: revision walk setup failed error: git-delta.bundle did not send all necessary objects ``` The run is reported as `adapter_failed` even though the agent produced output. The copy-back bundle names the host workspace's recorded HEAD (`baseSha`) as a prerequisite, but the host cannot satisfy it. **Steps to reproduce** Two independent triggers, both reproduced in tests: 1. The sandbox's HEAD has diverged from `baseSha` — e.g. the sandbox carries a local-only branch that forked from an older commit than the host's current HEAD. 2. The shared host workspace no longer holds `baseSha` at import time (it was reset / re-realized between export and import). In either case `git fetch` of the thin bundle fails with a missing prerequisite. **Expected behavior** Copy-back imports the sandbox's work as long as the host holds any common ancestor, instead of hard-failing and discarding the run. **Paperclip version** Current `master`. **Deployment mode** Any deployment running agents in sandbox environments with workspace sync (notably shared-workspace clones and custom images that carry a local-ahead branch). ## What Changed - `buildRemoteGitDeltaBundleScript` now computes `bundle_base = git merge-base <baseSha> HEAD` and bundles `HEAD --not <bundle_base>`. The merge-base is an ancestor of `baseSha`, so any host that holds `baseSha` (or an ancestor of it — e.g. after a reset) can satisfy the prerequisite, and the bundle stays a delta rather than a full-history transfer. - When `baseSha` is absent from the sandbox, or no merge-base exists, it falls back to a full, self-contained bundle (no prerequisites) so the import can always complete. - The existing empty-bundle no-op (no new commits) and the ordinary fast-forward path are unchanged; the `cat-file` base check no longer aborts the script under `set -e`. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — new cases: a diverged sandbox HEAD imports when the host holds only the merge-base (not `baseSha`), and the full-bundle fallback imports into a host that shares no history; existing thin-delta and empty-bundle cases still pass. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`. - Standalone shell repro confirmed the old thin bundle fails with "Repository lacks these prerequisite commits" in both trigger cases, and the merge-base bundle imports successfully. ## Risks - Low. For the common case (sandbox HEAD descends from `baseSha`) the merge-base is `baseSha`, so the bundle is byte-for-byte the same delta as before. The change only alters behavior when the old code would have hard-failed. - This makes the copy-back import succeed on a diverged base; the subsequent reconciliation of divergent histories (`integrateImportedGitHead`) is unchanged and still owns how the imported head is merged into the host branch. Where a workspace's history has genuinely diverged (e.g. a stale custom image carrying a local-only branch), a clean re-clone/re-capture is still the right operational fix — this change prevents work loss, it does not reconcile intentional divergence. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, shell repro, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ee9d907d01 |
fix(codex): do not inject a duplicate --skip-git-repo-check for sandbox runs (#10595)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A Codex agent runs `codex exec`, and the adapter assembles its argument vector from the agent's config plus execution-context options > - For sandbox execution the adapter injects `--skip-git-repo-check`, because a headless remote workspace has no git trust prompt to answer > - The adapter also appends the operator's `extraArgs` verbatim, so an agent that already lists `--skip-git-repo-check` in its config gets the flag twice on a sandbox run > - `codex exec` rejects a repeated `--skip-git-repo-check` and exits with code 2, which the adapter surfaces as `adapter_failed` before any work runs > - This pull request skips the sandbox injection when the operator's args already carry the flag > - The benefit is that a common, harmless-looking config no longer crashes every sandbox run ## Linked Issues or Issue Description **What happened?** A `codex_local` agent configured with `extraArgs: ["--skip-git-repo-check"]` fails on every sandbox run: ``` error: the argument '--skip-git-repo-check' cannot be used multiple times Usage: codex exec [OPTIONS] [PROMPT] ``` The adapter reports `stopReason: "adapter_failed"` (Codex exited with code 2). The flag appears twice in the argv: once injected by the adapter for sandbox execution, once from the operator's `extraArgs`. **Steps to reproduce** 1. Configure a `codex_local` agent with `extraArgs: ["--skip-git-repo-check"]` (or the legacy `args` field). 2. Point it at a sandbox environment. 3. Start a run — `codex exec` aborts immediately on the duplicate flag. **Expected behavior** The run launches with a single `--skip-git-repo-check`. An operator listing the flag the adapter already injects should be a no-op, not a hard failure. **Paperclip version** Current `master`. **Deployment mode** Any deployment running Codex agents in sandbox environments. ## What Changed - `buildCodexExecArgs` no longer pushes the sandbox `--skip-git-repo-check` when the resolved args (`extraArgs`, or the legacy `args` fallback) already contain it. The operator's copy stands; the argv carries the flag exactly once. Non-sandbox runs and configs without the flag are unchanged. ## Verification - `cd packages/adapters/codex-local && pnpm vitest run src/server/codex-args.test.ts` — new cases: `extraArgs` already carrying the flag (single occurrence), the legacy `args` field carrying it (single occurrence), and the operator's flag preserved when the sandbox injection is not requested. Existing "adds --skip-git-repo-check when requested" case unchanged. - `cd packages/adapters/codex-local && pnpm vitest run` — full package suite (218 tests). - `pnpm run typecheck` in the package. ## Risks - Low. The change only suppresses a duplicate of a single, idempotent flag; it never removes an operator-supplied argument and never adds one that was not already going to be present. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c0b875c46c |
fix(codex): let sandbox runs use the sandbox image's own Codex login (#10582)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Codex agents can run inside sandbox environments, and operators can bake a Codex login into the sandbox image during interactive image setup > - Two credential gates (the control plane's pre-dispatch configuration-incomplete gate and the adapter's execute-time fail-fast) required host-side Codex credentials — a usable `auth.json` in the managed home or a configured `OPENAI_API_KEY` — regardless of where the run executes > - On managed cloud hosts a local Codex login never exists, so every sandbox run of a Codex agent failed immediately with "configuration incomplete: no Codex credentials available for managed home …", even though the adapter's inbound auth merge already supports the image-login case end to end > - This pull request makes the execute-time gate probe the sandbox for its own `~/.codex/auth.json` before failing, and exempts sandbox-destined runs from the pre-dispatch host check > - The benefit is that a sandbox image signed in to Codex is a first-class credential source, matching what the auth-merge, precedence-warning, and copy-back machinery were already built for ## Linked Issues or Issue Description **What happened?** Running a `codex_local` agent in a sandbox environment whose image carries a Codex login failed instantly with `configuration incomplete: no Codex credentials available for managed home "…/codex-home". Sign in to Codex on the host with a ChatGPT subscription, or bind a per-agent OPENAI_API_KEY secret for this agent.` The host has no Codex login and never will on a managed cloud deployment; the sandbox's own login was never consulted. **Steps to reproduce** 1. Configure a sandbox environment and capture a custom image after signing in to Codex inside the interactive image setup. 2. Create a `codex_local` agent that uses that environment, on a host with no Codex login and no `OPENAI_API_KEY` bound. 3. Start a run: it fails pre-dispatch with the configuration-incomplete blocker above. **Expected behavior** The run launches and Codex authenticates with the sandbox image's own login, the same way the adapter's host↔sandbox auth merge already keeps the sandbox credential when the host ships none. A run should only fail fast when neither the host, a bound `OPENAI_API_KEY`, nor the sandbox has credentials. **Paperclip version** Current `master` (cloud image deployments). **Deployment mode** Managed cloud stacks (any deployment where the server host has no local Codex login). ## What Changed - Extracted the adapter's execute-time gate into `assertCodexCredentialsLaunchable`: when host readiness fails and the target is a sandbox, it probes `~/.codex/auth.json` in the sandbox (same command the auth-precedence warning uses) and proceeds with a log line naming the credential source; when the sandbox has no login either, the error now names all three remediation options (sandbox image sign-in, per-agent `OPENAI_API_KEY`, host sign-in). Non-sandbox targets keep today's strict behavior byte-for-byte. - The control plane's pre-dispatch gate in `resolveExecutionRunAdapterConfig` now takes the selected environment's driver and skips the host-credential check for sandbox-destined runs — only the adapter can probe the sandbox once it is up, so the execute-time gate is the authority there. Non-sandbox runs keep the early, well-attributed configuration-incomplete blocker. - The codex Test flow needed no change: it already seeds host credentials only when they exist and otherwise leaves the sandbox's `CODEX_HOME` alone; this aligns the run path with it. ## Verification - `cd packages/adapters/codex-local && pnpm vitest run` — 210 tests, including new gate cases: sandbox login present (proceeds + logs source), sandbox and host both credential-less (fails with the extended message), non-sandbox target (strict host requirement kept, no sandbox probe), per-agent API key (no probe at all). - `cd server && pnpm vitest run src/__tests__/heartbeat-project-env.test.ts src/__tests__/codex-local-adapter-environment.test.ts` — includes the new sandbox-exemption case next to the existing blocker tests. - `pnpm run typecheck` in `server` and `packages/adapters/codex-local`. ## Risks - Sandbox-destined misconfigurations (no credentials anywhere) now surface at adapter execute time instead of pre-dispatch, so they read as an adapter failure with a precise message rather than a configuration-incomplete blocker. The trade-off is deliberate: the sandbox must be up to know whether credentials exist, and the failure message names the exact remediations. - The sandbox probe adds one short (5s-capped) shell command to sandbox runs whose host has no credentials; runs with host credentials or a bound key are untouched. - Self-hosted behavior is unchanged for local and SSH targets. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
90ead239a8 |
feat(ui/server): name cross-company environment secret refs instead of calling them missing (#10577)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environment configs (sandbox providers, SSH) can bind stored company secrets through `format: "secret-ref"` fields, picked in the environment editor's secret picker > - Environments are instance-scoped and shared by every company on an instance, but the picker lists only the current company's secrets, so a ref pointing at another company's secret renders as "Missing secret (…)" in destructive styling > - That state is indistinguishable from a genuinely deleted secret, so operators "fix" a healthy binding by creating a duplicate secret in their own company — the exact sequence that used to corrupt bindings before #10576 > - This pull request adds an instance-gated metadata endpoint for an environment's secret refs and teaches the picker to name a cross-company secret and its owner honestly > - The benefit is that operators can tell a healthy cross-company binding from a broken one, and stop creating duplicate secrets ## Linked Issues or Issue Description **Is your feature request related to a problem? Please describe.** In the environment editor, a secret-ref field that points at a secret owned by a different company shows "Missing secret (22095402…)" in red, with "The previously selected secret is no longer available. Pick another or remove the binding." The binding is actually healthy — the current company's picker just cannot list the other company's secrets. Operators react by creating a duplicate secret and re-pointing the field. **Describe the solution you'd like** The editor should know the referenced secret's name, status, and owning company (metadata only, never the value) and present a cross-company ref neutrally, a deleted secret as deleted, and only an unknown id as missing. Related: #10576 (fixes the binding corruption this UI state used to trigger). ## What Changed - New `GET /environments/:id/secret-refs` returns `{ refs: [{ configPath, secretId, name, status, companyId, companyName }] }` for the environment's config-derived secret refs. Values are never returned. The route sits behind `assertCanAccessInstanceEnvironments`, the same gate as environment editing. - New `secretService.describeSecretRefs` loads that metadata across companies; unknown ids are omitted. - `SecretBindingPicker` reads an optional `SecretRefHintsContext` (keyed by secret id). With a hint, a ref the company list cannot show renders as `NAME — Owning Company` with neutral styling and the note "Owned by the … company. The binding keeps working; selecting a secret from this list re-points it here." A hint with `status: "deleted"` reports the secret as deleted. Without hints, behavior is byte-identical to before — agent editors and other picker users are unaffected. - `CompanyEnvironments` fetches descriptors for the environment being edited and provides them through the context. ## Verification - `cd server && pnpm vitest run src/__tests__/environment-routes.test.ts src/__tests__/secrets-service.test.ts` — new endpoint happy path, agent 403 (descriptors never computed), and embedded-Postgres coverage proving cross-company names resolve and unknown ids drop out. - `cd ui && pnpm vitest run src/components/SecretBindingPicker.test.tsx src/components/JsonSchemaForm.test.tsx src/pages/CompanyEnvironments.test.tsx` — hinted cross-company rendering, hinted deleted secret, and unchanged no-hint fallback. - `pnpm run typecheck` in `server` and `ui`. - Manual: edit an environment whose secret-ref field references another company's secret; the field names the secret and its owning company instead of "Missing secret". ## Risks - The endpoint exposes secret names and company names across companies to instance-level environment editors. Those actors already manage instance-shared environments (and instance admins are implicit members of every company), so this reveals no secret material and no new reach; the service method documents that callers must sit behind an instance-level gate. - UI change is additive and context-gated; pickers without a provider render exactly as before. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f51cba33fa |
fix(server): keep environment secret bindings consistent when re-pointing config secrets (#10576)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can run inside environments (SSH boxes, sandbox providers); a sandbox environment's config can reference stored company secrets (for example a provider API key) through `format: "secret-ref"` fields > - Environments are instance-scoped and shared by every company on an instance, but `company_secret_bindings` rows are company-scoped, and the environment routes synced config-derived bindings under one guessed "context company" resolved from the environment's existing bindings > - When a save re-pointed a secret-ref field at a secret owned by a different company, the binding sync threw after the config row had already been persisted: the config referenced the new secret, the binding still pointed at the old one, every later lease acquisition failed with `Secret is not bound to environment:<id> at apiKey`, and the stale cross-company binding made every later save fail with a company-context conflict — with no route-level way to recover > - This pull request makes config-derived bindings follow the company that owns each referenced secret, and makes the environment write and its binding syncs atomic > - The benefit is that environment saves can no longer strand an environment in a half-updated state that breaks all of its runs ## Linked Issues or Issue Description Refs #10577 (companion UX change: the editor state that nudges operators into this sequence). **What happened?** Saving an environment whose secret-ref config field points at a secret owned by a different company than the environment's existing binding partially applied: the config row updated, the binding sync failed server-side, and the environment was left referencing a secret it has no binding for. Every run that leased the environment then failed with `lease_acquire_failed: ... Secret is not bound to environment:<id> at apiKey`, and every later save of the environment returned 409 `Environment secret bindings already use a different company context.` — with no route-level way to recover. **Steps to reproduce** 1. On an instance with two companies, create a sandbox environment from company A with a picker-bound API-key secret owned by A (the binding lands in A). 2. From company B, create a new secret and re-point the environment's API-key field at it, then save. 3. The save persists the config but the binding sync throws, so no binding for B's secret exists. 4. Run any agent that uses the environment, or try to save the environment again. **Expected behavior** The save either fully applies (config and bindings consistent) or fully fails. Re-pointing a config secret ref to a secret owned by another company moves the binding with the secret. **Paperclip version** Reproduced on current `master` (also present on recent release images). **Deployment mode** Multi-company server deployment (any mode with more than one company on the instance). ## What Changed - New `secretService.replaceSecretRefsForInstanceTarget`: writes each config-derived binding under the company that owns the referenced secret, replaces all non-`env.*` bindings of the target across every company, and validates every ref (secret exists, not deleted, config-path and projection-class rules) before any row is written. `env.*` env-var bindings stay company-scoped and untouched. - The environment create and update routes now run the environment write and its binding syncs inside one `db.transaction`, threading the transaction through new optional executor seams on `environmentService.create/update` and the existing `SecretBindingDb` seam pattern, so an invalid ref rolls the whole save back instead of leaving a half-updated environment. - `resolveEnvironmentSecretContextCompanyId` no longer lets existing bindings veto the caller's context (the 409s above); it now only picks where new raw-pasted secrets are created and how env-var bindings and probes resolve: explicit route/query company first, then the single company the bindings live in, then the actor's company. ## Verification - `cd server && pnpm vitest run src/__tests__/environment-routes.test.ts src/__tests__/environment-instance-routes.test.ts src/__tests__/secrets-service.test.ts src/__tests__/environment-custom-image-routes.test.ts` (165 tests, includes new coverage below) - New embedded-Postgres tests prove: a re-point moves the binding to the new secret's company and deletes the stale row; refs across several companies each bind under their own secret's company; an unknown secret ref rejects without touching existing bindings; `env.*` rows survive config-ref replacement. - New route tests prove: a cross-company re-point that previously 409'd now saves, with the update and binding replacement on the same transaction executor; a failing ref surfaces as 422. - `cd server && pnpm run typecheck` ## Risks - Behavioral shift: environment saves no longer 409 on a company-context mismatch between the caller and existing bindings; bindings follow the referenced secret's company instead. Environment routes are instance-admin gated, and instance admins already had access to every company's secrets by passing the company explicitly, so this removes an ordering trap rather than widening access. - Runtime lease resolution is unchanged: a run still resolves environment secrets under the run's own company, so an environment referencing company B's secret still only leases for company B runs (fail-closed as before). - The delete route's per-company binding cleanup is unchanged. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ea0dd3917e |
build: make image layer caching actually hit (#10571)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Published container images are the deployable unit for self-hosted and managed instances, so merge-to-image latency bounds every deploy iteration > - The docker workflow configures BuildKit caching, but builds still ran ~12+ minutes essentially cold > - Two causes: the most expensive layer (four CLI toolchains + apt) is ordered after the always-changing app copy so it can never cache, and the type=gha cache's 10GB repo cap means the two multi-arch mode=max jobs evict each other > - This pull request reorders the tool layer above the app copy (with a weekly epoch so @latest tools keep advancing) and switches both jobs to registry-backed cache in ghcr > - The benefit is that warm builds shrink to roughly the app build + push, targeting the sub-5-minute range together with the amd64-only cloud variant ## Linked Issues or Issue Description No existing public issue — inline description following the feature request template: **Subsystem affected** CI / release publishing (docker workflow, Dockerfile) **Problem or motivation** Despite `cache-from/cache-to` being configured, image builds run effectively cold: (1) the production stage installs four CLI toolchains + apt packages *after* `COPY --from=build /app /app`, and since the app copy changes every commit, that most-expensive layer rebuilds every build, per arch; (2) the `type=gha` BuildKit cache is capped at 10GB per repository, and two multi-arch `mode=max` jobs overflow and evict each other's entries. **Proposed solution** Order the tool/OS layer before the app copy (it references nothing from `/app`), refresh it weekly via a `CLI_TOOLS_CACHE_EPOCH` build arg so the `@latest` tools don't freeze in the cache, and move both jobs to registry-backed BuildKit cache (`:buildcache` / `:buildcache-cloud` refs in ghcr, no size cap, separate refs so the parallel jobs don't clobber each other). **Alternatives considered** Pinning CLI tool versions instead of the weekly epoch — more deterministic, but adds a version-bump chore; the weekly epoch preserves current freshness semantics with bounded staleness. Keeping type=gha with `mode=min` — smaller cache but loses intermediate-stage reuse, which is where most of the win is. **Roadmap alignment** Not on ROADMAP.md; CI/publishing speed improvement only. ## What Changed - `Dockerfile`: the production stage's tool/OS `RUN` (npm --global CLIs, apt, `/paperclip` setup) moves above `COPY --from=build /app /app`; new `CLI_TOOLS_CACHE_EPOCH` arg consumed by that layer. The `cloud` stage is unaffected — it only layers plugin dists on top of the finished production stage. - `.github/workflows/docker.yml`: both jobs stamp the ISO week into `CLI_TOOLS_CACHE_EPOCH`, and both switch `cache-from/cache-to` from `type=gha` to `type=registry` with per-job refs. - Includes the one-line amd64-only cloud-variant commit from #10570 so the two PRs can't conflict; if #10570 merges first, this PR rebases down to a single commit automatically. ## Verification - Image content is unchanged by layer reordering: the moved `RUN` references nothing from `/app`, and Docker layer ordering only affects caching, not the final filesystem (tool installs and app copy touch disjoint paths). - The cache ref is written only by this workflow — `docker.yml` runs on master/tag pushes, never on PRs — so the workflow's existing "no shared caches into build inputs" supply-chain stance is unchanged (BuildKit layer cache was already accepted via type=gha; the registry backend has the same writer trust). - Runtime proof lands with the first two master builds after merge: the first warms the cache, the second should show the tool layer and deps/build stages as CACHED in the build log, with wall clock dropping accordingly. I'll be watching those as part of managed-deploy work. ## Risks - Low. Worst case the registry cache misses (cold-build behavior, same as today). The weekly epoch means CLI tools update at most a week late inside images; a release built mid-week ships the tools from that week's first build. Cache refs add two small artifacts to ghcr. ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, via Claude Code with tool use and code execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (no test-affecting changes) - [x] I have added or updated tests where applicable (n/a — build config and layer ordering only) - [x] I have updated relevant documentation to reflect my changes (in-file comments document both mechanisms) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ce40343b9c |
build: publish the cloud image variant for amd64 only (#10570)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Published container images are how both self-hosted users and managed-deployment hosts run it > - The repository publishes two variants: the self-hosted image and a cloud variant for managed deployments > - The cloud variant was built for amd64+arm64, but its only consumers are managed-deployment hosts, which run amd64 > - The QEMU-emulated arm64 half dominates the build's wall clock, delaying every merge-to-deployable-image cycle > - This pull request drops arm64 from the cloud variant only, keeping the self-hosted image multi-arch > - The benefit is roughly halving the time from merge to a deployable cloud image, with no change for any actual consumer ## Linked Issues or Issue Description No existing public issue — inline description following the feature request template: **Subsystem affected** CI / release publishing (docker workflow) **Problem or motivation** The cloud image variant builds for `linux/amd64,linux/arm64`, but the arm64 half runs under QEMU emulation and dominates the job's wall clock — while no consumer of the cloud variant runs arm64 (managed-deployment hosts are amd64). Every deploy iteration pays ~double the necessary build time. **Proposed solution** Build the cloud variant amd64-only. The self-hosted image keeps `amd64+arm64` so ARM users (Apple Silicon, ARM servers) are unaffected. **Alternatives considered** Keeping multi-arch but building arm64 on native arm64 runners with a manifest merge — faster than QEMU and worth doing for the self-hosted image if its build time becomes a pain point, but unnecessary complexity for a variant with no arm64 consumers. **Roadmap alignment** Not on ROADMAP.md; CI/publishing speed improvement only. ## What Changed - `.github/workflows/docker.yml`: the `build-and-push-cloud` job's `platforms` is now `linux/amd64` (with a comment explaining why). The self-hosted `build-and-push` job is untouched. ## Verification - Build-config-only change; the workflow runs on merge to master. The published `-cloud` manifest will be amd64-only, which its consumers already pull. - No test changes: nothing at runtime differs on any platform that actually runs the image. ## Risks - Low. If an arm64 consumer of the cloud variant ever appears (e.g. local `docker run` on Apple Silicon for debugging), it would fall back to emulation on the consumer's machine or need this reverted — a one-line change. The self-hosted image's platform matrix is unchanged. ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, via Claude Code with tool use and code execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (no test-affecting changes) - [x] I have added or updated tests where applicable (n/a — CI platform matrix only) - [x] I have updated relevant documentation to reflect my changes (in-workflow comment documents the rationale) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
521271ebb7 |
build: bake the build commit into published images (#10566)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instances commonly run from published container images, and operators need to observe which build a container actually serves > - `/api/health` now reports the running build commit, and server-info already falls back to `PAPERCLIP_BUILD_COMMIT` when git is unavailable > - But published images carry no `.git` and never received `PAPERCLIP_BUILD_COMMIT`, so containers report `commit: null` — verified live against a current image > - That leaves the new deployment-verification field inert exactly where it matters most: containerized deploys > - This pull request bakes the exact build commit into both image variants at build time, mirroring how `PAPERCLIP_BUILD_VERSION` is already stamped > - The benefit is that containers report their true commit on `/api/health`, so deploy tooling can verify a rollout actually shipped ## Linked Issues or Issue Description Companion to #10563 (which exposed the `commit` field on `/api/health`). Inline description following the bug report template: **What happened?** A container from a published image responds to `GET /api/health` with `"commit": null`. The image has no `.git` directory and the `PAPERCLIP_BUILD_COMMIT` fallback that `server-info` supports is never provided at build time, so git metadata resolves as unavailable. **Expected behavior** A container reports the commit it was built from, the same way it already reports its build version via the baked `PAPERCLIP_BUILD_VERSION`. **Steps to reproduce** Run any published image (e.g. `ghcr.io/paperclipai/paperclip:sha-c4f6264-cloud`) and `curl /api/health` — `commit` is `null` even though the build commit is known at image-build time. **Paperclip version or commit** `sha-c4f6264-cloud` (first image containing #10563). ## What Changed - `Dockerfile`: new `PAPERCLIP_BUILD_COMMIT` build arg, exported as an ENV in the production stage (the `cloud` stage inherits it), directly parallel to `PAPERCLIP_BUILD_VERSION`. Empty for local `docker build`, which keeps the normal fallbacks. - `.github/workflows/docker.yml`: both build jobs pass `PAPERCLIP_BUILD_COMMIT=${{ github.sha }}`. ## Verification - Reviewed the plumbing end-to-end: `build-commit.ts` reads `PAPERCLIP_BUILD_COMMIT` (validated as a full SHA), `server-info.ts` `readGitInfo` falls back to it when the git CLI fails, producing `available: true, fullSha` — which `/api/health` surfaces as `commit`. - Verified live that a current published image reports `commit: null`; this change repairs that on the next build. Post-merge, the first master image should report its commit — I'll be verifying that as part of managed-deploy validation. - No test changes: the fallback path is already covered by existing server-info tests; this PR only supplies the env at image build. ## Risks - Low. Two build-time stamps; no runtime code changes. A wrong SHA would only mislabel the build (same failure mode `PAPERCLIP_BUILD_VERSION` already carries), and `${{ github.sha }}` is the exact commit the workflow builds. ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, via Claude Code with tool use and code execution); diagnosis included live probes of a running container's `/api/health`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (no test-affecting changes; server suites unaffected) - [x] I have added or updated tests where applicable (n/a — build-time stamps only) - [x] I have updated relevant documentation to reflect my changes (Dockerfile comments document the arg) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
51bb41c7e3 |
Expose the running build commit on the unauthenticated health response (#10563)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instances run self-hosted or under hosting/deploy tooling, and operators need to observe what build a server is actually running > - `/api/health` carries the git SHA only inside `serverInfo`, which is gated to board/agent actors — anonymous callers get a redacted body with no version signal at all > - Deploy tooling that manages instances from outside (fleet rollouts, hosting providers, upgrade scripts) therefore cannot ground-truth that a deploy actually shipped without holding credentials > - A build commit is a plain git SHA of this public repository — it is not a secret, and gating it buys no security while blocking legitimate verification > - This pull request surfaces the running build commit as a top-level `commit` field on every `/api/health` response, including the redacted anonymous one > - The benefit is credential-free deploy verification: any operator or tool can confirm which commit an instance serves, while the fuller `serverInfo` block stays access-controlled as before ## Linked Issues or Issue Description No existing public issue — inline description following the feature request template: **Subsystem affected** Server (API, runs, routes) **Problem or motivation** An anonymous `GET /api/health` returns a redacted body with no version information; the running git SHA exists only in `serverInfo.git.fullSha`, which requires a board/agent actor. External deploy tooling (fleet rollouts, hosting providers, upgrade scripts) therefore cannot verify that an instance is actually serving the build it was just upgraded to — a rollout that silently keeps running the old image is indistinguishable from a successful one at the health endpoint. **Proposed solution** Surface the running build commit as a top-level nullable `commit` field on every `/api/health` response shape, including the redacted anonymous one, while keeping the fuller `serverInfo` block access-controlled as before. A build commit is a plain git SHA of this public repository — exposing it costs nothing and enables credential-free deploy verification, like the `version` endpoints on most server software. **Alternatives considered** Authenticating deploy tooling as a board actor to read `serverInfo` — rejected: it forces credential plumbing into infrastructure that only needs a public SHA, and adds a whole class of auth-misconfiguration failure to deploy verification. **Roadmap alignment** Not on ROADMAP.md; a small operational observability improvement, no overlap with planned core work. ## What Changed - `server/src/routes/health.ts`: derive `commit` from the server info snapshot (`serverInfo.git.fullSha` when git metadata is available, else `null`) and include it as a top-level field on every `/api/health` response shape — the redacted anonymous body, the full-details body, the no-db body, and the 503 database-unreachable body. - `serverInfo` itself remains gated to full-details responses exactly as before; only the bare commit is newly public. - `server/src/__tests__/health.test.ts`: updated exact-shape assertions to include `commit`, and added an assertion that `commit` is `null` (not omitted) when git metadata is unavailable. The redacted-response tests now pin that anonymous callers receive the commit. ## Verification - `pnpm vitest run src/__tests__/health.test.ts` in `server/` — 13 tests pass, including the redacted-anonymous shapes (which now pin the `commit` field) and the git-unavailable `null` case. - `tsc -p server/tsconfig.json --noEmit` — clean. - Manual: `curl -s https://<instance>/api/health` as an anonymous caller returns `"commit": "<full sha>"` alongside the existing redacted fields. ## Risks - **Version disclosure:** anonymous callers can now fingerprint the exact running commit. This is a deliberate trade-off: the builds are of a public repository (the SHA reveals no private code), the endpoint already responds to anonymous callers, and the operational value — verifying deploys actually shipped — outweighs the marginal fingerprinting surface. Operators who consider this sensitive are typically fronting `/api` with their own access controls already. - Otherwise low risk: no behavioral change to any gated field, no schema or API-surface removal; `commit: null` keeps the field shape stable when git metadata is absent (e.g. non-git installs). ## Model Used Claude Opus 4.8 (`claude-opus-4-8`, extended thinking, via Claude Code with tool use and code execution) authored the change and tests; finalized and PR'd under Claude Fable 5 (`claude-fable-5`). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (none needed beyond code comments — health endpoint has no standalone doc) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
dd1a7f5290 |
Ensure app-home ownership before the privilege drop, not only on remap (#10530)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Docker image persists all instance state (project checkouts, worktrees, run logs, uploads) under `PAPERCLIP_HOME`, and deployments mount a volume there for durability > - The entrypoint starts as root and drops privileges to the `node` user, but it fixes `PAPERCLIP_HOME` ownership only when it remaps the user's UID/GID > - A freshly mounted volume arrives root-owned and shadows the image's build-time `chown`, so a default-UID boot drops privileges onto an unwritable home and the server crashes on its first `mkdir` > - This pull request makes the entrypoint probe the home's ownership and chown whenever it does not match the runtime user, before the privilege drop > - The benefit is that the image works out of the box on any platform-managed volume, with the common already-correct boot staying chown-free ## Linked Issues or Issue Description No public issue exists — describing the bug inline (per the bug report template). **What happened?** Running the image with a freshly created volume mounted at `/paperclip` (a Docker named volume, a Kubernetes PV, or any platform-managed volume) and the default `USER_UID`/`USER_GID` crashes on boot: `Error: EACCES: permission denied, mkdir '/paperclip/instances/default/logs'`. **Expected behavior** The container boots and initializes its instance tree on the mounted volume, exactly as it does when `/paperclip` is the image's own (build-time chowned) directory. **Steps to reproduce** 1. `docker volume create paperclip-data` 2. `docker run -v paperclip-data:/paperclip ghcr.io/paperclipai/paperclip:<any current tag>` 3. Observe the EACCES crash on the first `mkdir` under `/paperclip`. **Root cause** `scripts/docker-entrypoint.sh` chowns `/paperclip` only inside its UID/GID remap branch (`changed=1`). A fresh volume mount is root-owned and shadows the image's build-time `chown node:node /paperclip`; with the default 1000:1000 no remap happens, so no chown happens, and `gosu node` drops onto an unwritable home. **Paperclip version or commit:** reproduces on `master` and any published image. **Deployment mode:** any; observed on managed-cloud volume mounts and reproducible with plain Docker named volumes. **Installation method:** Docker image (`ghcr.io/paperclipai/paperclip`). **Related PRs (dedup search):** no open or merged PR touches the entrypoint ownership logic; the entrypoint's privilege-handling tests were added previously and this extends them. No duplicate found. ## What Changed - `scripts/docker-entrypoint.sh`: the remap-conditional `chown` is replaced by an ownership probe — after any UID/GID remap, the entrypoint stats `PAPERCLIP_HOME` (default `/paperclip`) and runs `chown -R node:node` only when the owner does not match the runtime user, before `exec gosu node`. Covers fresh root-owned mounts and trees written under a previous UID mapping; the already-correct boot performs no chown. The unprivileged (non-root start) branch is unchanged. - `server/src/__tests__/docker-entrypoint.test.ts`: `stat` stub added to the harness; new cases for the fresh root-owned mount with default UID/GID and for `PAPERCLIP_HOME`-relative probing; the remap case now models the post-remap ownership mismatch. ## Verification - `pnpm vitest run server/src/__tests__/docker-entrypoint.test.ts` — 7 passed (5 existing behaviors unchanged, 2 new). - Live on a managed deployment: a container that crash-looped with the EACCES above boots cleanly once the home is chowned before the drop (the same effect this entrypoint change produces; forced there by a UID remap as an interim workaround). ## Risks - Low. Behavior changes only for boots where `PAPERCLIP_HOME` exists with mismatched ownership — exactly the boots that crash today. `chown -R` on a large previously-mismatched tree adds one-time boot latency; correctly-owned homes skip it entirely. Kubernetes restricted / OpenShift non-root starts keep the existing exec-directly path untouched. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic; Claude Code CLI with extended thinking and tool use; tests executed locally via Vitest). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (no duplicates; extends the existing entrypoint privilege tests) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
075951f6bd |
Fix import completion UX: inbox flood, false-failure message, stale company list (#10538)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company Import/Export (#10507, hardened in #10523 and #10531) now imports a large company end to end via an async job > - A real 1,418-issue import succeeded, but three rough edges showed up in that success > - Imported issues flooded the inbox, a completed import surfaced a false "failed" message after its in-memory result expired, and the new company didn't appear in the switcher until a manual refresh > - This pull request keeps imported issues out of the inbox, treats an expired-but-completed import as success, and refreshes the company list on completion > - The benefit is that a successful import looks and feels successful, and doesn't bury the user's inbox in historical tasks ## Linked Issues or Issue Description - Refs #10507 / #10523 / #10531 (Import/Export and its hardening). No open issue; three post-import bugs described above. ## What Changed - **Imported issues no longer flood the inbox.** The inbox "mine" tab is a query: an issue is "touched" if the user authored a comment on it, and import re-attributes bundled user comments to the importing user — so every imported issue appeared. Import now seeds a per-user `issue_inbox_archives` row for each imported issue (via a batched `issues.archiveImportedInbox`), the exact table the inbox visibility query excludes. Gated on an actor user id, so agent/system imports and normal issue creation are untouched; genuine new activity still resurfaces the issue. - **A completed import no longer shows a false failure.** The in-memory job's terminal retention was 5 minutes, so a poll after that 404'd and the UI showed "failed." Retention is extended to 60 minutes — the real mitigation for a user who steps away during a long import. `watchImportJob` additionally treats a *server-confirmed* success whose full result is no longer retained (a `succeeded` status carrying only the compact summary — a cloud tenant job, or a board job whose full in-memory result aged out) as a soft success ("import completed — open the company"), navigating by the summary's company id. A 404 while the job is still being watched is *not* treated as success: a running job is never dropped by the retention sweep, so its disappearance means a restart mid-import that may not have finished, and it surfaces the honest "may have restarted while the import ran" error. A first-poll 404 (the id never existed) is likewise a real error. - **The imported company appears without a refresh.** `onSuccess` now invalidates the companies/switcher query unconditionally (covering both the full-result and expired-but-completed paths) and navigates by the job's company id. ## Verification - shared/server/ui typechecks clean; 15 UI tests in the touched spec green, plus the embedded-Postgres import batching and portability-routes suites. - New tests: embedded-Postgres test that imported touched issues are archived for the actor and excluded from the inbox query while a normally-created issue still appears; job resolvable at the old window+1 and only 404s past 60 min; UI soft success on a server-confirmed `succeeded` job without a retained full result (no error, list invalidated, navigates by company id), a running-then-gone job → honest error (restart mid-import), and a first-poll 404 → error. ## Risks - Low and import-scoped: the inbox archive only affects imported issues for the importing user; normal issue creation and non-user (agent/system) imports are unchanged. Retention extension is a constant; the async job store remains in-memory by design. A restart mid-import still 404s and is surfaced honestly as a possible failure (never masked as success); only a server-confirmed success whose full result has expired is reported as a soft success. ## Model Used - Implementation: Claude Fable 5 (`claude-fable-5`, Anthropic). Review hardening (the confirmed-success narrowing): Claude Opus 4.8 (`claude-opus-4-8`, Anthropic). Both via the Claude Code CLI with extended thinking + tool use; root-caused against the live import. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5ec7ce76e5 |
Upload company import packages as compressed zip uploads (fix large-company imports through Cloud) (#10531)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company Import (#10507, hardened in #10523) lets a user upload a company package on the Import page > - The page expanded the user's `.zip` into a files map and POSTed it as ONE inline JSON body — ~40MB for a real company because attachment blobs get base64-inflated > - On Paperclip Cloud that body travels browser → harness proxy → tenant, where it truncated in transit → body-parser 400 → the browser saw "Failed to fetch", and nothing imported > - Two compounding causes: the giant inline body itself, and the board async opt-in riding an `x-paperclip-cloud-*` header that the Cloud harness strips as anti-spoofing (so async never engaged and the import held one fragile synchronous connection) > - This pull request uploads the raw compressed `.zip` as a multipart request (about a third the size, already compressed) parsed server-side into the same bundle the importer consumes, and moves the async opt-in to a proxy-safe `?async=1` > - The benefit is that a large-company import actually completes through Cloud: a small compressed upload, a real async job that survives dropped connections ## Linked Issues or Issue Description - Refs #10507 / #10523 (Import/Export and its hardening). No open issue; problem described above (large-company browser import through a proxy: inline JSON body truncates → 400 → "Failed to fetch"; async opt-in header stripped by the front door → async never engages). ## What Changed - **Multipart zip transport.** The Import page uploads the raw `File` as `multipart/form-data` (field `package`, import options in a JSON `meta` field); the server unzips it into `{ rootPath, files }` and runs the exact existing preview/import logic. The `application/json` inline path is byte-identical for CLI/programmatic callers. Bare `application/zip` (meta via `?meta=`) is also accepted for programmatic use. - **Shared node zip reader.** `packages/shared/src/portability-zip.ts` (node-only subpath, not re-exported to the browser bundle — same pattern as `portability-hash.ts`); the CLI's `zip.ts` becomes a thin re-export. Identical codec (STORE + DEFLATE via `inflateRawSync`, rejects data descriptors/zip64). - **Proxy-safe async signal.** `wantsAsyncImport` = `?async=1` (board browsers, survives the harness) OR the existing `x-paperclip-cloud-async-import` header (cloud tenants, set server-side). The UI async client now uses `?async=1`. Backward compatible. - **Size + preflight.** New `PORTABLE_ZIP_UPLOAD_LIMIT_BYTES = 128MB`; the inline 56MB preflight no longer gates the zip path (it shows the compressed size instead). Async submit/poll/resume, the duplicate-guard fingerprint (now over the resolved bundle), pause-on-import, progress/error panels, and activation all apply to the multipart path. - OpenAPI documents json + multipart + zip bodies and the `async` query param. ## Verification - Full typecheck chain (shared, server, ui, cli) clean. - 152 tests across 8 files: new `portability-zip.test.ts` (STORE/DEFLATE/base64-blob byte-exact round-trip, truncation throws, data-descriptor rejection); `company-portability-routes.test.ts` +7 (multipart import+preview equals the inline bundle; async multipart 202→poll→success; board async via `?async=1` with no cloud header; cloud-tenant async via header; sync fallback with neither; truncated-zip 400, nothing imported); `CompanyImport.test.tsx` asserts the local zip sends the raw File and the inline preflight no longer blocks; `openapi-routes.test.ts` green. - NOT yet measured: the end-to-end browser upload through the live Cloud harness — verified on staging after deploy before closing out. ## Risks - Import semantics unchanged — only transport changed; the JSON inline path is byte-identical, the cloud-tenant header async path untouched. Multipart parsing is server-side (memory-bound: a ~13MB zip → ~30MB files map, fine on the server). - The bare `application/zip` path is programmatic-only and covered by content-type dispatch but not a dedicated route test (the multipart path is). ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), Claude Code CLI, extended thinking + tool use; root-caused against live logs/DB and the harness proxy source. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
276ae3a75d |
Harden company import: durable UI, async jobs, integrity guard, batched inserts (#10523)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company Import/Export (#10507) moves whole companies between instances as portability bundles > - Real-world use on a large company (1,418 issues, ~10.6k comments) surfaced a cluster of related failures: the import took hours and the browser connection died while the server kept running, a retry silently produced a second partial import, the progress/error UI gave no durable signal, and a cloud-tenant user couldn't even open the companies afterward > - Root cause of the slowness: importBundle inserted every issue, comment, and document as a separate round-trip to a network Postgres — an N+1-over-network pattern > - This pull request hardens the whole import path: durable progress/error UI, an async server-side job so imports survive dropped connections (with a duplicate-submit guard), a fail-closed guard against incomplete payloads, and batched inserts that cut a large import from hours to minutes > - The benefit is that migrating a real, large company actually completes, is legible while it runs, and can't half-import twice ## Linked Issues or Issue Description - Refs #10507 (the Import/Export feature this hardens). Supersedes #10513 (the progress/error-UI piece, folded in here). No open issue; problem described above (large-company import: slow, connection-fragile, silently duplicable, opaque UI). ## What Changed - **Batched inserts (perf):** importBundle pre-generates entity ids in JS and inserts in chunked multi-row statements, so children no longer wait on parents' generated ids. A 1,418-issue import drops from ~15,600 insert statements to **82** (190×); benchmark below. Import semantics — collision handling, pause-on-import, label/blocker/monitor/attachment/embedded-asset handling, blob sha verification — are unchanged (full portability suite green). - **Async import jobs for board sessions:** the existing cloud-tenant async job path opens to board sessions with per-actor job keys; the import page submits, polls, and resumes watching after a reload or dropped connection instead of holding one fragile request. A non-terminal job blocks a duplicate submit (409 returns the running job), preventing the double-import. - **Fail-closed completeness guard:** an optional `expectedFileCount` on inline imports; the server rejects (422 `import_payload_incomplete`) a body carrying fewer files than declared, so a re-framed/short payload fails loudly instead of half-importing. - **Durable progress/error UI (was #10513):** persistent progress panels with size-aware copy, persistent error panels with retry guidance, and inline explanation when the preview button is disabled; request-lifecycle guards so stale previews/imports can't publish or detach. ## Verification - `pnpm -r` typechecks (shared, server, ui) clean. - `company-portability.test.ts` (76) + `company-portability-routes.test.ts` (30) green — the import correctness net — plus new `CompanyImport.test.tsx` async/resume/409 coverage and a new batching regression test (a 50-issue import issues <50 issue-insert statements; rows land unchanged). - **Batching benchmark (embedded Postgres):** at 1,418 issues × 7 comments × 1 doc — 82 insert statements vs ~15,598 one-per-row (190×), ~1s wall-clock; a row-verifying run at that scale imports all 1,418 issues / 9,926 comments / 1,418 documents with unique identifiers and no warnings (no rows dropped by chunking). Over a network DB the round-trip reduction is the hours→minutes lever. - What is NOT directly measured here: wall-clock against a real network Postgres (that happens on a staging deploy); the local timing is network-free. ## Risks - Batching is the load-bearing change: it rewrites the import write path. Mitigated by the unchanged 106-test correctness suite, a new scale/row-integrity test, and per-writer transactions (a failure rolls back its table group; not a single outer transaction across writers — noted, correctness preserved). - Async jobs are in-memory (lost on server restart → pollers 404 and can resubmit); matches the pre-existing cloud-tenant job semantics. - `expectedFileCount` is optional (older callers unaffected); over-count is allowed, only under-count fails closed. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), Claude Code CLI, extended thinking + tool use; implementation across Fable 5 subagents with live diagnosis against a running instance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1c52f02d34 |
Let cloud tenant sessions reach companies they hold memberships in (#10524)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - On Paperclip Cloud, each stack authenticates its users to the tenant app through trusted headers (`resolveCloudTenantActor`), which seed a primary company for the stack > - That actor was pinned to exactly one company — the seeded primary — regardless of any other companies the user actually holds a membership in > - Companies created later (via the import flow, or company creation) write real membership rows for the user, but the pinned actor ignored them, so those companies showed up in listings yet returned "User does not have access to this company" when opened > - This pull request unions the pinned primary with the user's own active membership rows, exactly as a locally authenticated session already does > - The benefit is that a Cloud user can reach every company they belong to — most visibly, a company they just imported ## Linked Issues or Issue Description - Refs #10507 (Import/Export — imported companies were unreachable on Cloud stacks). No open issue; bug described above (companies visible in listing but unreachable; expected: reachable when the user holds an active membership). ## What Changed - Extracted the session path's own active-membership query into `loadActiveUserCompanyMemberships(db, userId)` (single-sourced; the session path now calls it too). - `resolveCloudTenantActor` unions its result with the pinned primary: `companyIds = [primary, ...others]`, memberships likewise, primary first. Strictly per-user; a membership-read failure degrades to primary-only (mirrors the existing fail-closed owner-elevation pattern). No change to owner instance-admin elevation, grant seeding, the stale instance-admin purge, or trusted-header validation. - Grants are seeded at membership creation across all flows (company create, invite/join, import), not per request — so no extra seeding was added here. ## Verification - `@paperclipai/server` typecheck clean. - `cloud-tenant-actor.test.ts` (+ union / other-user-excluded / inactive-excluded / no-rows-identical cases), `auth-session-route.test.ts` (route-level: trusted headers reach a unioned company through `assertCompanyAccess`), plus agent-auth, authz-company-access, cross-company-authz, portability-routes — 83 tests green. ## Risks - Low and tightly scoped: only widens a Cloud actor's reachable companies to those it already holds active memberships in; users without extra memberships, other users' rows, and owner elevation are all unaffected. Read failure fails closed to primary-only. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), Claude Code CLI, extended thinking + tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
187a90b7bc |
Add opt-in in-flight run-log mirroring with graceful-shutdown flush (#10512)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The run-log store records each agent run's output and can mirror completed logs to S3-compatible object storage > - The mirror uploads only on finalize, so a server restart mid-run loses the whole in-flight log > - Deployments and crashes are routine on ephemeral hosts, and lost run output makes failed runs impossible to debug > - This pull request adds an opt-in throttled mirror for still-running logs plus a graceful-shutdown flush > - The benefit is that a restart mid-run keeps the log tail up to the last mirror interval, and an orderly restart keeps everything ## Linked Issues or Issue Description No public issue exists — describing the feature inline (per the feature request template). **Subsystem affected** server/ — REST API & orchestration services **Problem or motivation** `RUN_LOG_S3_BUCKET` gives finished run logs durability, but the mirror uploads only on finalize. A run that is still writing when the server restarts leaves nothing in object storage. On hosts with ephemeral disks the local file is gone too, so the run's output is lost end to end and failed runs cannot be debugged. **Proposed solution** Mirror the in-flight log to the same object key on a throttled cadence (`RUN_LOG_S3_INFLIGHT_MIRROR_SECONDS`), and flush dirty tails during graceful shutdown. Keep it opt-in so existing deployments see zero new upload traffic unless they ask for it. **Alternatives considered** Per-append uploads (rejected: one PUT per output chunk is hostile to S3 endpoints and run latency). Chunked part objects with read-time stitching (rejected: complicates the read path, and S3 multipart minimum part sizes do not fit small tails). Persistent volumes (rejected upstream already: the data dir is deliberately an emptyDir in hardened cloud_tenant deployments). **Roadmap alignment** Not on ROADMAP.md; extends the existing run-log durability mirror without changing any default behavior. **Additional context** Ranged reads already serve partial objects like a live tail, so the read path needs no change; finalize overwrites the mirror with the complete file. **Related PRs (dedup search):** the finalize-only S3 mirror landed previously and this extends it; no duplicate or competing PR found for in-flight run-log mirroring. ## What Changed - `server/src/services/run-log-store.ts`: new opt-in `inflightMirrorMs` on the S3 options (`RUN_LOG_S3_INFLIGHT_MIRROR_SECONDS` env). When set, appends schedule at most one upload of the current file per interval, to the same key finalize uses. Ranged reads already serve that key, so a partial object behaves like a live tail and needs no read-path change. Finalize retires the in-flight bookkeeping and waits out an upload already on the wire, so a stale partial can never overwrite a finalized log. Upload failures warn, re-mark the tail dirty, and retry at most once per interval. - `server/src/services/run-log-store.ts`: new `flushInflightMirrors()` on the store and a module-level `flushInFlightRunLogMirrors()` for the shutdown path. Both are no-ops when the mirror is off. - `server/src/index.ts`: graceful shutdown flushes dirty in-flight tails after the heartbeat run drain, so runs the drain did not finalize (timeouts, the hot-restart skip path) still persist their output. - `server/src/services/run-log-store.test.ts`: five new tests — off-by-default (no uploads before finalize), tail preserved after a wipe without finalize, throttle coalescing with a single flush upload, finalize superseding the in-flight mirror and retiring its timer, and upload failures never breaking appends with recovery on the next flush. ## Verification - `pnpm vitest run server/src/services/run-log-store.test.ts` — 13 passed (8 existing + 5 new). - `pnpm vitest run server/src/__tests__/heartbeat-run-log.test.ts server/src/__tests__/heartbeat-active-run-output-watchdog.test.ts` — 21 passed (consumers of the store, unchanged behavior). - `pnpm -C server run typecheck` — clean. - Self-hosted behavior is unchanged unless `RUN_LOG_S3_INFLIGHT_MIRROR_SECONDS` is set: with the variable unset there are zero new uploads and the finalize-only mirroring is byte-identical (asserted by the off-by-default test). ## Risks - Low. The feature is opt-in; unset env preserves today's behavior exactly. When enabled, worst case is one extra PUT per interval per active run, and every upload is best-effort — a failing endpoint warns and never breaks appends, finalization, or shutdown. The finalize path awaits any in-flight upload before writing the complete file, closing the only overwrite race the design introduces. Timers are `unref`ed so the mirror never keeps the process alive. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic; Claude Code CLI with extended thinking and tool use; tests executed locally via Vitest). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (no duplicates found for in-flight run-log mirroring; the finalize-only mirror landed previously and this extends it) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
916c13501f |
Replace host-to-host Cloud Sync with full-fidelity company Import/Export (#10507)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A company accumulates real state — issues, labels, blockers, documents, work products, monitors, attachments, agents, routines — and people need to move that state between instances: self-hosted to cloud, cloud back to self-hosted, or plain backups > - The experimental, flag-gated Cloud Sync transport (#6548) tried to solve this host-to-host: the source pushed into a receiver over HTTPS with a cross-instance consent/token handshake, which required the destination to be publicly reachable and broke for common self-hosted topologies (plain-HTTP LAN/VPN origins); the receiver half never landed upstream at all > - Meanwhile the portability bundle and the existing export/import pages already move companies offline with none of those networking constraints — but silently dropped labels, blockers, issue documents, work products, monitors, and every attachment > - This pull request removes the host-to-host transport and makes Import/Export the single data-movement path: the pages become first-class company-settings destinations, exports declare exactly what they do not carry, and bundle schemaVersion 6 now carries all of the above, with attachments as content-addressed sha256 blobs verified before a single row is written > - The benefit is a migration and backup flow that works between any two instances with no reachability requirements, no cross-instance auth, and no silent data loss ## Linked Issues or Issue Description - Refs #6548 — the original Cloud Sync sender this PR supersedes and removes. - Related, not duplicates: #1697 (goals in the portability manifest — orthogonal field addition), #954 (an earlier import/export + skill-visibility proposal predating the current portability bundle). - No open issue describes this directly, so in brief (feature-request shape): **Problem** — moving a company between instances silently lost labels (imports with label references actually hard-failed), blocker relations, issue documents, work products, monitor state, and all attachments, and the alternative Cloud Sync transport required the destination to be publicly reachable over HTTPS plus a consent handshake, which failed for typical self-hosted setups. **Desired behavior** — one Import/Export flow in company settings that produces a portable bundle carrying all of that data, tells the operator up front what it cannot carry, imports with automations paused, and offers real one-click activation afterwards. ## What Changed - New export fidelity report (`GET /api/companies/:companyId/export/fidelity`) + an "Export fidelity" panel on the Export page listing anything a bundle will not include (now only: approvals, cost history, activity history) - Imports accept `pauseAutomations`; imported agents and routines land paused, the import result reports created routines, and the Import page ends in an activation panel that actually resumes selected agents/activates routines - Export and Import pages promoted into the company-settings nav; the Cloud Upstream wizard, ux-lab page, and API client removed; the old settings route redirects to Export - Host-to-host transport removed: upstream-sync/receiver-client routes and services, CLI `cloud connect`/`cloud push` + keypair store, the shared upstream transfer contract, and the `enableCloudSync` flag; migration `0196` drops the two experimental `cloud_upstream_*` sender tables - Bundle schemaVersion 6: labels (definitions + per-task names, remapped by name on import), blocker relations (`blockedBy` slugs, cycle-tolerant), issue documents (`tasks/<slug>/documents/<key>.md`), work products (system refs nulled), monitors (notes/scheduledBy restored, imported un-armed) - Attachments travel as content-addressed `blobs/<sha256>` entries (deduped; comment-scoped attachments re-link via comment index); every blob is hash-verified **before any write**, so a corrupted bundle cannot leave a partially imported company; both zip codecs now round-trip extensionless/binary entries byte-exactly; the Import page preflights the inline body limit and offers continue-without-attachments - v5 (and older) bundles still import, with an informational warning; bundles newer than v6 are rejected cleanly - Docs: board-operator import/export guide, CLI README, README/ROADMAP updated ## Verification - `pnpm -r` typechecks (shared, db incl. migration numbering/safety checks, server, ui, cli) and `pnpm check:token-gates` — clean - Vitest: full server + shared sweep 4,888 passed / 1 skipped, with the only 3 failures being pre-existing on `master` (2× heartbeat-workspace-branch-containment, 1× workspace-runtime auto-port; reproduced identically with this change stashed); ui + cli suites green; the embedded-Postgres export-fidelity suite applies the full migration chain including the new `0196` against a fresh database - Live end-to-end on a scratch instance: seeded a company with labels, a blocker pair, an issue document, a work product, a monitor, an agent, a routine, and two binary attachments (one comment-scoped) → export → import into a fresh company → labels remapped to new ids, blocker edge and document restored, monitor un-armed with notes intact, attachments byte-identical (sha256-compared through the API), agents/routines paused → activation panel resumed them; a v5-shaped bundle imported with only the info warning; flipping one byte in a blob made the import 422 with **zero** rows created - Reviewer repro: create a company with a labeled issue + attachment → Settings → Export → download → Settings → Import on another company/instance → watch the preview, apply with "start paused", then activate ## Risks - Migration `0196` drops `cloud_upstream_connections`/`cloud_upstream_runs` — experimental tables behind a default-off flag; their connection/run history is intentionally discarded - Breaking removals are all of experimental, flag-gated surface: `/api/upstream-sync/*` + `/api/cloud-upstreams/*` routes, `paperclipai cloud connect|push`, and the `enableCloudSync` flag (stale keys in stored instance settings parse harmlessly) - Import remains non-atomic on mid-apply errors generally (pre-existing behavior); the new blob verification specifically moved ahead of all writes so tampered bundles cannot create partial state - GitHub-sourced imports do not fetch `blobs/*` and skip attachments with a warning ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), via Claude Code CLI with extended thinking, tool use, and subagent orchestration; implementation and review split across Fable 5 subagents, with live end-to-end verification against a running instance ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5c5366d0c1 |
fix(server): honor explicit plugin RPC timeouts (#10460)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Plugin workers connect Paperclip to external runtimes and sandbox
providers.
> - Some adapter heartbeats run a full sandbox session inside one
`environmentExecute` RPC.
> - The worker manager limited every RPC timeout to 15 minutes, even
when the caller gave a longer timeout.
> - This pull request keeps the normal default timeout behavior but
honors explicit caller timeouts.
> - The benefit is that long sandboxed agent sessions can continue past
15 minutes while other safety guards still bound hung work.
## Linked Issues or Issue Description
No public GitHub issue exists for this bug.
### Bug Report
Pre-submission checklist:
- Searched existing open and closed issues and did not find a duplicate.
- Confirmed the bug is reproducible on `master` from the current source
tree.
- Confirmed the error starts in Paperclip timeout handling, not in an
adapter provider or local configuration.
What happened?
- A sandbox-backed adapter heartbeat can run a full agent session inside
one `environmentExecute` plugin RPC.
- The plugin worker manager capped every RPC timeout at 15 minutes.
- The cap also applied when the caller passed a longer explicit timeout
for an execute-style call.
- A long sandbox command could fail before the adapter budget expired.
Expected behavior:
- Ordinary plugin RPC calls should keep the normal 30-second default
timeout.
- The default timeout path should still have a 15-minute maximum.
- A caller-supplied positive finite timeout should be honored, including
values above 15 minutes.
Steps to reproduce:
1. Use a plugin environment driver that calls `environmentExecute` with
an explicit timeout above 15 minutes.
2. Run a command that stays active longer than 15 minutes and remains
inside the adapter budget.
3. Observe that the worker manager times out the RPC at 15 minutes
before this fix.
4. Run the same path after this fix and observe that the explicit
timeout is used.
Paperclip version or commit:
- Reproduced from the current `master` line before this change.
Deployment mode:
- Local dev or self-hosted server with sandbox-backed execution.
Installation method:
- Built from source.
Agent adapter(s) involved:
- Codex.
- Custom or external plugin adapter.
- Core plugin worker timeout handling.
Database mode:
- Not database-related.
Access context:
- Agent execution context.
Node.js version:
- Not version-specific.
Operating system:
- Not OS-specific.
Relevant logs or output:
```shell
RPC call "environmentExecute" timed out after 900000ms
```
Relevant config:
- Not config-related.
Additional context:
- Execute-style sandbox calls already have adapter inactivity monitors,
platform silent-run checks, and provider command timeouts. This PR
removes the unintended worker-manager clamp only for explicit positive
finite caller timeouts.
Privacy checklist:
- Reviewed all pasted output for PII, user paths, API keys, tokens,
company names, and internal instance links.
Duplicate search:
- Searched open PRs and open issues in `paperclipai/paperclip` for
`environmentExecute timeout`, `MAX_RPC_TIMEOUT_MS`, and
`plugin-worker-manager timeout`.
- Searched the same terms in `HenkDz/paperclip`.
- Found no matching open PRs or issues.
- Compared this patch-id against my open PRs in `paperclipai/paperclip`;
no match was found.
## What Changed
- Added `resolveRpcCallTimeoutMs()` to keep explicit positive finite
timeouts intact.
- Kept the 15-minute maximum only for the default timeout path.
- Updated `callInternal()` to use the new resolver.
- Added unit tests for explicit long timeouts, default timeout clamping,
fractional values, and invalid explicit values.
- Clarified why notification invocation scopes still use the 15-minute
TTL.
## Verification
- `corepack pnpm install --frozen-lockfile`
- `corepack pnpm --filter @paperclipai/plugin-sdk ensure-build-deps`
- `corepack pnpm --filter @paperclipai/server exec vitest run
src/__tests__/plugin-worker-manager.test.ts`
- `corepack pnpm --filter @paperclipai/server exec tsc --noEmit`
- `git diff --check
|
||
|
|
7a5a217d60 |
fix(server): gate sandbox/ssh execution targets by shared remote-managed adapter capability (#10459)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip can run an agent in a remote-managed environment, such as a sandbox provider or an SSH host. > - The server resolves an execution target for each run. The resolver kept its own hardcoded list of allowed adapters. > - The shared capability metadata in `packages/shared/src/environment-support.ts` already defines which adapters support remote-managed environments. The environment selector and the capabilities API use it. > - The two lists drifted. The UI offered sandbox environments to Grok Build (`grok_local`) agents, but the resolver refused them at run time. > - This pull request makes the resolver use the shared capability check for both the sandbox gate and the SSH gate. > - The benefit is one source of truth. The UI and the runtime now agree on which adapters can use remote-managed environments. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Inline description per the bug report template: **What happened?** A Grok Build (`grok_local`) agent was assigned a sandbox environment (a Daytona provider). The UI allowed the assignment. Every run and primary-model test then failed with the warning: `Adapter "grok_local" is not allowed in "<environment>" environments.` **Expected behavior** An adapter that the environment selector offers for a sandbox environment must also pass the runtime gate. The Grok Build run must start in the sandbox. **Steps to reproduce** 1. Create a sandbox environment (for example, with a Daytona provider plugin). 2. Create an agent that uses the `grok_local` adapter. 3. Set the agent's environment to the sandbox environment. The UI accepts this. 4. Run the agent, or run the primary-model test. The run fails with the adapter-not-allowed warning. **Paperclip version or commit** Reproduced on `master` at `0edb742f8d`. **Deployment mode** Local instance with a remote sandbox provider plugin. The same gate also applies to SSH environments. ## What Changed - `resolveEnvironmentExecutionTarget` in `server/src/services/environment-execution-target.ts` now gates the sandbox path with the shared `adapterSupportsRemoteManagedEnvironments()` helper. Before, it used a hardcoded six-adapter list that did not include `grok_local`. - The SSH path in the same file now uses the same shared helper. - New regression tests in `server/src/__tests__/environment-execution-target.test.ts`: sandbox target resolution for every remote-managed adapter (including `grok_local`), SSH target resolution for `grok_local`, and the null path for an adapter without remote-managed support. ## Verification - Run `node_modules/.bin/vitest run server/src/__tests__/environment-execution-target.test.ts`. All 10 tests pass, including the 3 new ones. - Confirm `grok_local` is in the `REMOTE_MANAGED_ADAPTERS` set in `packages/shared/src/environment-support.ts`. The resolver now reads the same set. - On a live local instance with this fix, a `grok_local` agent assigned to a Daytona sandbox environment no longer produces the adapter-not-allowed warning. ## Risks Low risk. The change routes two hardcoded checks through existing shared capability metadata. Behavior changes only where the lists had drifted: `grok_local`, and any future adapter added to the shared set, can now resolve sandbox and SSH execution targets. Adapters outside the shared set still return `null`. ## Model Used Claude Fable 5 (`claude-fable-5`) by Anthropic, with extended thinking and tool use, running in Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
78f8c6c3d4 |
Recover managed bundled plugin workers on demand (#10429)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed deployments can auto-provision bundled sandbox-provider plugins so cloud or remote execution environments appear in the board UI > - In a multi-service deployment, several server processes can share one database and boot concurrently > - A sibling process can create a bundled plugin row while the web process sees it before it reaches `ready` > - The web process correctly avoids clobbering the existing row, but its startup `loadAll()` can miss the plugin and never start that worker locally > - The environments capabilities route then filters out the sandbox provider because the plugin is ready in the database but not running in the web process > - This pull request adds a narrow managed-bundle recovery path that lazily starts the missing worker when the capabilities route sees a ready managed bundled plugin > - The benefit is that the sandbox provider becomes visible after the install finishes, without requiring a web-process restart ## Linked Issues or Issue Description - No public GitHub issue found for this exact deployment race. - Related broad plugin runtime context: Refs #432. Bug description: - What happened: in a managed multi-service deployment with shared database state and bundled plugin auto-install enabled, the API-serving process can skip a plugin row while it is still `installed`, run startup plugin loading before that row becomes `ready`, and then permanently omit the sandbox provider from environment capabilities. - Expected behavior: once the managed bundled plugin row reaches `ready`, the API-serving process should be able to start the plugin worker and include its sandbox provider without a restart. - Steps to reproduce: boot a web process and a sibling worker process concurrently; have the sibling create the bundled plugin row and transition it to `ready` after the web process has already skipped auto-install and run `loadAll()`. - Deployment mode: managed multi-service deployment with shared database state and `plugins.autoInstall` configured. ## What Changed - Added a managed bundled plugin worker recovery helper that single-flights lazy `loadSingle()` starts and only allows configured managed bundled plugin keys. - Passed the managed recovery hook into the environments capabilities route. - Updated `listReadyPluginEnvironmentDrivers()` to attempt bounded recovery for ready managed bundled plugins whose worker is missing in the current process, and only for plugins that actually declare a `sandbox_provider` environment driver. - Made request-time recovery use `loadSingle(id, { markErrorOnFailure: false })` so a local activation failure in one process never transitions the shared plugin row to `error` (a sibling process may be running the plugin successfully). - When error writes are suppressed and activation fails after the worker was spawned, the loader now tears down the partially-registered local runtime (scheduler registration, event subscriptions, agent tools, worker process) instead of leaving a half-activated worker lingering; the teardown steps are factored out of `unloadSingle()` into a shared helper. - A failed recovery attempt now discards the crashed/stopped handle it left registered in the worker manager (a worker that dies during initialize is killed without a scheduled restart), so later capability requests can retry recovery instead of being blocked by the handle-presence gate until a process restart. Handles in starting/running/backoff states are left to the worker manager's own lifecycle; recovery only ever starts when no handle existed, so no pre-existing worker can be affected. - Added a regression test suite covering the installed-to-ready race, allowlist behavior, the driver-kind gate, existing worker handles, concurrent single-flight recovery, bounded slow recovery attempts, suppressed shared error-state writes, partial-runtime teardown on late activation failure, and retry after a dead handle is discarded. ## Verification - `pnpm vitest run src/__tests__/plugin-environment-driver-ready-recovery.test.ts` (in `server/`) passed: 10 tests. - `pnpm --filter @paperclipai/server typecheck` passed. ## Risks - Low risk for self-hosted single-process deployments because lazy recovery is only wired when managed plugin auto-install config is present; with no managed config the capabilities route takes the exact pre-change code path. - The capabilities route can wait briefly while attempting recovery; the attempt is bounded and defaults to 2 seconds. - Failed recovery keeps the prior behavior of omitting the provider until a later successful worker start, and now also cleans up any partially-started local worker so retries begin from a clean slate. ## Model Used - Initial implementation: OpenAI GPT-5 via Codex local coding agent, with repository tool use and command execution. - Review-feedback follow-ups (driver-kind gate, partial-runtime teardown, expanded regression tests): Claude Fable 5 (claude-fable-5) via Claude Code, with repository tool use and command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4c8d92f086 |
fix(ui): leave agent detail after termination (#10451)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The agent detail page lets operators inspect one agent and run lifecycle actions from that context > - Terminated agents are removed from the normal active-agent surface, so the current detail route may no longer be fetchable after termination > - Before this change, terminating an agent from its detail page invalidated agent queries while leaving the browser on the now-stale detail route > - That refetch could surface an "Agent not found" error even though the terminate action itself succeeded > - The shared action button already owns the terminate mutation, so it can notify detail-page callers when termination succeeds > - This pull request redirects the detail page back to the agents list after a successful terminate action > - The benefit is that operators land on a valid route and the Back button does not return them to the stale terminated-agent detail route ## Linked Issues or Issue Description No direct public GitHub issue or PR was found for this detail-page termination flow. ### What happened? After terminating an agent from its detail page, the UI could remain on that agent's detail route and show an "Agent not found" error after query invalidation/refetch. ### Expected behavior Once termination succeeds, the operator should leave the now-stale detail page and land on a valid agents view. ### Steps to reproduce 1. Open a non-built-in agent detail page. 2. Use the overflow actions menu to terminate the agent. 3. Observe the post-termination route/error state. ### Paperclip version or commit Current `master` before this PR. ### Deployment mode Browser UI behavior, independent of a specific deployment mode. Duplicate search: searched public GitHub issues and PRs for `agent not found terminate`, `terminate agent detail`, and `Agent not found`; no direct duplicate or viable in-flight PR was found. ## What Changed - Added an optional `onTerminateSuccess` callback to `AgentActionButtons`, fired only after the shared terminate mutation succeeds. - Wired `AgentDetail` to replace-navigate to `/agents/all` after successful termination. - Extended `AgentActionButtons` coverage for the terminate success path, including API args, callback payload, and query invalidations. ## Verification - `corepack pnpm exec vitest run ui/src/components/AgentActionButtons.test.tsx` - `corepack pnpm check:token-gates` - `git diff --check origin/master..HEAD` - Local diff scan for obvious tokens, credential filenames, and email addresses found no matches. ## Risks Low risk. The new callback is optional, only fires for successful terminate actions, and preserves existing behavior for other `AgentActionButtons` callers. ## Model Used OpenAI Codex, GPT-5-based coding agent (`gpt-5`), tool use enabled for repository inspection, editing, local verification, and GitHub CLI operations. Context window details were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dc12197cce |
fix: prevent duplicate built-in agents and self-heal reconciliation (#10223)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Every company is auto-provisioned a set of built-in agents (e.g. the
Summarizer), and a startup reconciler keeps that set correct across
every company on boot.
> - Provisioning marks these agents with
`metadata.paperclipBuiltInAgent.key`, but nothing in the database
enforced one active agent per `(company, key)` —
`provision()`/`ensure()` did a check-then-insert with no guard.
> - Two concurrent server processes (e.g. a `tsx watch` double-boot)
could both read "no summarizer exists" and both insert, leaving a
company with duplicate built-in agents plus paired orphan pending
`hire_agent` approvals.
> - That data blemish then became a recurring outage: `findSingleAgent`
throws on >1 marked row, and because the throw escaped
`reconcileBuiltInAgentsOnStartup`'s sequential loop, **every company
after the affected one was silently skipped** on each boot — no
auto-provisioning, no default grants — until manual DB surgery.
> - This pull request closes the race at the database level and makes
reconciliation self-healing and fault-isolated.
> - The benefit is that concurrent provisioning can no longer create
duplicates, and even pre-existing duplicates are resolved automatically
instead of bricking startup reconciliation for unrelated companies.
## Linked Issues or Issue Description
- [x] I searched the GitHub PR list (open and recently closed) for
similar PRs and confirmed this is not a duplicate.
No public GitHub issue exists; describing the bug in-PR (bug-report
shape):
**What happened**
A dev instance booted with two concurrent server processes. Both ran
built-in agent provisioning for the same company at the same time, and
the check-then-insert in `provision()`/`ensure()`
(`server/src/services/built-in-agents.ts`) let both writers see "no
summarizer exists" and each create one — the company ended up with two
identical Summarizer agents (identical `paperclipBuiltInAgent` markers)
plus two paired pending `hire_agent` approvals.
From then on, **every** server boot logged:
```
ERROR: startup reconciliation of built-in agents failed
Multiple built-in agents found for summarizer (built_in_agent_duplicate_instance)
```
because `findSingleAgent` throws on >1 marked row rather than resolving
the duplicate. Worse, `reconcileBuiltInAgentsOnStartup` loops companies
sequentially and the throw escaped the loop, so every company *after*
the affected one was silently skipped on every boot.
**Expected behavior**
1. Concurrent provisioning must not create duplicate built-in agents
(there was no DB uniqueness constraint on the marker key per company).
2. Reconciliation should be resilient: if duplicates exist anyway,
self-heal (keep the oldest row, terminate the newer dupe, cancel its
orphan pending `hire_agent` approval), and never let one bad company
abort reconciliation for the rest.
**Steps to reproduce**
- Race two `provision(companyId, "summarizer")` calls for a company with
board approval for new agents enabled (or simulate a double-boot); both
insert.
- Restart the server → startup reconciliation error fires, companies
later in the loop are never reconciled.
## What Changed
**Part 1 — stop creating duplicates**
- Migration `0192_built_in_agent_unique_marker` adds a **partial unique
index** on `(company_id, metadata->'paperclipBuiltInAgent'->>'key')`
where the marker exists and `status != 'terminated'`. It first resolves
any pre-existing duplicates (keep oldest by `created_at`, terminate
newer dupes, cancel their orphan pending `hire_agent` approvals, revoke
their API keys) so the index can be created on already-affected
instances.
- `provision()`/`ensure()` now catch the losing race's `23505` unique
violation (walking the driver's wrapped cause chain) and re-resolve to
the winning row instead of surfacing the error.
**Part 2 — resilient reconciliation**
- `findSingleAgent` self-heals: keeps the oldest marked row, terminates
the newer duplicates, and cancels each one's orphan pending `hire_agent`
approval (idempotent) instead of throwing.
- `reconcileBuiltInAgentsOnStartup` isolates per-company failures in
both loops so one bad company can't abort reconciliation for the rest;
it surfaces a `companyFailures` count in the startup log.
- Adds `approvalService.cancel()` for system-initiated cancellation of
an orphan approval.
## Verification
- `pnpm --filter @paperclipai/db run check:migrations` → numbering +
safety checks pass.
- `packages/db` migration test (real embedded Postgres) — seeds
pre-index duplicate state, runs the migration, asserts dupes resolved +
index enforced: **1 passed**.
- `server` `built-in-agents.test.ts` — self-heal, concurrent races
(plain and board-gated), and startup
self-heal-without-aborting-later-companies: **34 passed**.
```
pnpm --filter @paperclipai/db exec vitest run src/built-in-agent-unique-marker-migration.test.ts
pnpm --filter @paperclipai/server exec vitest run src/__tests__/built-in-agents.test.ts
```
## Risks
- **Migration safety**: the migration mutates data (terminates duplicate
rows, cancels their orphan pending approvals, revokes their API keys)
before creating the index. It keeps the oldest row per `(company, key)`
and only touches non-terminated marked rows; the destructive step is
covered by the migration test and the safety-check baseline. On a clean
instance it is a no-op cleanup followed by `CREATE UNIQUE INDEX IF NOT
EXISTS`.
- Otherwise low risk: the unique index is partial (excludes terminated
rows, so re-provisioning after a termination stays possible), and the
conflict handling degrades gracefully to re-resolving the existing
winner.
## Model Used
Claude Opus 4.8 (`claude-opus-4-8`), 1M context window, extended
thinking, with tool use.
|
||
|
|
9f5af4ea5d |
fix(server): accept secret_ref binding objects in sandbox provider environment config (#10355)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents execute in environments; sandbox provider plugins (Daytona,
Modal, e2b, …) declare their config via a JSON schema, with credentials
marked `format: "secret-ref"`
> - The environments UI renders those fields with a secret picker that
submits `{ type: "secret_ref", secretId, version }` binding objects,
while the server-side environment config paths only understood raw
string values and bare secret-id strings
> - The binding object reached the plugin worker's
`environmentValidateConfig` untouched; plugins parse non-string config
values as absent, so saving or testing an environment with a
picker-bound secret always failed validation (e.g. "Daytona sandbox
environments require an API key in config or DAYTONA_API_KEY.", "Modal
sandbox environments require tokenId and tokenSecret.")
> - Worse, an environment first saved with raw pasted values becomes
uneditable: the stored value is a secret reference, the edit form
re-submits it as a binding object, and every subsequent save fails the
same way
> - This pull request canonicalizes binding objects to the bare secret
id before plugin validation, and teaches the persistence/runtime/probe
secret-ref resolvers to accept the object shape defensively
> - The benefit is that picker-bound secrets work for every
schema-driven sandbox provider — create, edit, and Test — with no plugin
changes required
## Linked Issues or Issue Description
Fixes #10105
The same failure reproduces with the Daytona provider: Settings →
Instance settings → Environments → New, driver sandbox, provider
daytona, bind Api Key to an existing secret via the picker → Save fails
with "Daytona sandbox environments require an API key in config or
DAYTONA_API_KEY."
## What Changed
- `server/src/services/json-schema-secret-refs.ts`: new
`parseSecretRefBindingObject()` that recognizes the `{ type:
"secret_ref", secretId, version? }` shape the secret picker submits
(version defaults to `"latest"`; malformed objects return null).
- `server/src/services/plugin-environment-driver.ts`:
`validatePluginSandboxProviderConfig()` now canonicalizes binding
objects at the driver schema's `format: "secret-ref"` paths to the bare
secret id (the persisted shape) before invoking the plugin worker's
`environmentValidateConfig`. Pinned numeric versions are rejected with a
clear 422, since sandbox provider references always resolve the latest
version — silently resolving a different version would be worse.
- `server/src/services/environment-config.ts`: the persistence, runtime,
and probe secret-ref resolvers plus `collectEnvironmentSecretRefs()`
accept the binding-object shape defensively, so any previously persisted
object-shaped refs (from providers whose validation tolerated them)
resolve instead of being silently skipped; the missing-companyId runtime
guard also now fails closed for object-shaped refs.
## Verification
- `npx vitest run server/src/__tests__/json-schema-secret-refs.test.ts
server/src/__tests__/plugin-sandbox-provider-config-validation.test.ts
server/src/__tests__/environment-routes.test.ts
server/src/__tests__/environment-config.test.ts` — 82 tests pass,
including new coverage: binding-object canonicalization before plugin
validation, pinned-version rejection, raw-string pass-through, and a
route-level create with a picker-submitted binding object persisting the
bare secret id without minting a duplicate secret.
- `npx vitest run server/src/__tests__/environment-runtime.test.ts` — 24
tests pass against embedded Postgres, including a new test that persists
an object-shaped ref and verifies runtime resolution produces the
plaintext credential for the plugin worker.
- `pnpm typecheck` in `server/` — clean.
## Risks
- Low. The canonical persisted shape (bare secret-id string) is
unchanged, so existing saved environments and lease-resume fingerprints
are unaffected; raw pasted values and bare-id strings take exactly the
same code path as before.
- New behavior only triggers where a save/probe previously failed 422
(binding objects at secret-ref paths) or where an object-shaped ref was
previously skipped silently at runtime (now resolved, or failed closed
without a companyId).
- Pinned binding versions at sandbox-provider paths are now an explicit
422 instead of an accidental validation failure; no UI submits pinned
versions today (`allowVersionSelector={false}`).
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic) with extended thinking and
agentic tool use (Claude Code harness): source diagnosis, fix, and
tests.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
doc surface changed)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
c274f10abc |
feat(server): computed owner instance-admin elevation for cloud-managed instances, behind platform floors (#10343)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Cloud-managed instances authenticate tenant users through a trusted-header path (`resolveCloudTenantActor`) that deliberately never grants `instance_admin`, so every tenant user is company-scoped > - On a dedicated (single-owner) managed instance that leaves the paying owner unable to administer their own instance: instance settings, the environments admin surface, and the custom sandbox image flow are all instance-admin gated (the environments UI can't even show the provider/image of the platform sandbox because the restricted read view blanks `config` entirely) > - Re-granting the old blanket `instance_user_roles` row would repeat the mistake the shared-pool hardening fixed: DB rows go stale, resurrect via restores, and elevate through every auth path > - This pull request elevates only the stack `owner`, computed per request at the trusted-header boundary behind a new managed-tier feature flag, and ships that elevation together with code floors on the platform-owned surfaces an instance admin must not control on a managed instance > - The benefit is that dedicated-stack owners can administer their own instance while platform credentials, execution policy, backups, and runtime code-install stay platform-owned, and self-hosted behavior is unchanged ## Linked Issues or Issue Description No public GitHub issue exists for this change; the underlying issue is described here following the feature-request template. ### Problem or motivation - On cloud-managed instances, tenant users resolved from trusted headers are always company-scoped. For dedicated instances with a single paying owner, the owner cannot reach any instance-admin surface of their own instance (instance settings, environments administration, custom image setup), and the restricted environment read view hides even structural fields like the sandbox provider and image. - The previous hardening intentionally removed blanket elevation (and purges stale `instance_user_roles` rows on every trusted-header authentication). That protection must not regress for shared multi-tenant pools. ### Proposed solution - Owner-only, computed, flag-gated elevation plus code floors on platform-owned surfaces, in one PR so the elevation can never ship without the floors. ### Alternatives considered - Re-inserting an `instance_user_roles` row for owners (the pre-hardening model): rejected — DB rows go stale, survive restores, and elevate through every auth path; #7525 removed exactly this. - Widening only the environments read view without any elevation: rejected — it fixes one screen but still leaves a dedicated-instance owner unable to administer instance settings or custom images. - Elevating additional stack roles (`member`/`admin`/`support`): rejected — only the owner has an ownership claim over the whole instance; other roles stay company-scoped. ### Roadmap alignment - Extends the shipped "Cloud deployments" roadmap work (multi-tenant isolation, company-scoped cloud tenants, managed-instance bootstrap) without overlapping planned core items, and leaves self-hosted behavior unchanged. ## What Changed - **New feature key** `enableOwnerInstanceAdmin` (`packages/shared`): boolean flag in `instanceExperimentalSettingsSchema`, catalog tier `managed`, `cloudDefault: true`, `selfHostedDefault: false`. Inert on self-hosted instances — the elevation path only exists behind the cloud tenant trust token. - **Computed elevation** (`server/src/middleware/auth.ts`): `resolveCloudTenantActor` now returns `isInstanceAdmin: true` only when the trusted-header stack role is `owner` **and** the flag is enabled. The flag is resolved through the instance-settings service so the managed-config overlay applies (the control plane can disable elevation fleet-wide without touching tenant databases; a DB row edit or restore cannot resurrect it). Resolution fails closed on settings read errors. The `instance_user_roles` never-insert and the per-request stale-row purge are byte-identical. `member`/`admin`/`support` stack roles stay company-scoped. - **Authorization guard split** (`server/src/services/authorization.ts`): the blanket-allow now trusts the actor's *computed* `isInstanceAdmin` flag (only the attested resolver can set it for `cloud_tenant` actors) while keeping the `instance_user_roles` DB lookup excluded for `cloud_tenant` — a stale or hand-inserted role row still elevates nothing. - **Floor F1 — platform environment credentials** (`server/src/routes/environments.ts`): on cloud-managed instances, platform-provisioned environment rows (`managedByPaperclip` marker, plus the legacy managed-Kubernetes marker) use a single floored view for every reader on all environment routes (list, get, create, update, delete responses): `envVars` are never echoed and credential-shaped `config` keys (reusing the managed-config `SECRET_LIKE_CONFIG_KEY_PATTERN`) are dropped — for **all** actors including instance admins — while structural config (provider, image, template, region, …) and the managed markers stay visible. This also fixes the environments UI for managed sandboxes, which previously lost the provider/image entirely in the restricted view. The floor also covers writes: `PATCH /environments/:id` and `DELETE /environments/:id` on a platform-provisioned row are rejected (403, `environment_platform_managed`) for every actor including instance admins, and the guard binds to the persisted row's markers so a patch cannot strip the managed marker to lift the floor. The one recovery path is a metadata-only PATCH that solely clears the marker keys (null/false), for rows stamped through the old unrestricted API before the markers became reserved — and it never applies to a row whose slot markers are live platform state: the single local row (`environments_local_driver_idx`), which `ensureLocalEnvironment` adopts and stamps on cloud-managed instances from every caller (company creation, the heartbeat, run orchestration), and the single marked sandbox row (`environments_managed_sandbox_idx`) while a managed-sandbox bootstrap path is configured (managed-config `environments` section or `PAPERCLIP_EXECUTION_MODE=kubernetes`) and the provisioner therefore adopts and refreshes it on every boot. Clearing a live slot row's markers would let the next write reclassify it as tenant-managed and bypass the floor; conversely, when no sandbox provisioning path is configured the platform holds no claim on any sandbox row, so a platform marker there is stale by definition and the recovery patch applies. Every marker outside a live slot is clearable, so no legacy row is ever locked permanently. Custom-image setup and probes on the platform sandbox stay available to instance admins — those are the owner-facing flows this elevation exists for. The marker keys themselves are reserved: client create/update payloads that set `managedByPaperclip` or `managedKubernetesSandbox` are rejected (422, `environment_platform_marker_reserved`) on cloud-managed instances, so a tenant row can never be stamped platform-provisioned through the API and self-locked behind the write floor (the provisioner writes markers at the service layer, not through these routes). Tenant-created environments are otherwise unaffected. - **Floor F2 — executionMode** (`server/src/routes/instance-settings.ts`): on cloud-managed instances, `PATCH /instance/settings/general` rejects writes that would change `executionMode` (403, `execution_mode_platform_managed`). Same-value echoes pass so settings forms that submit the full general-settings object keep working. The boot-time execution-policy bootstrap path is untouched (it calls the service directly). - **Floor F3 — manual database backups** (`server/src/routes/instance-database-backups.ts`): floored off on cloud-managed instances (403, `database_backups_platform_managed`); backups are platform-owned there, and the result would also echo a server-side filesystem path. - **Floor F4 — adapter code install** (`server/src/routes/adapters.ts`): `POST /adapters/install` and `POST /adapters/:type/reinstall` are floored off on cloud-managed instances (403, `adapter_install_platform_managed`). Adapter packages execute in the server process, so a runtime install would let an instance admin read the platform trust anchors out of the process environment. This mirrors the existing bundled-only plugin install floor; adapter code on managed instances comes bundled with the platform image. ## Instance-admin surface audit Before widening who can hold `isInstanceAdmin`, every instance-admin-gated surface in `server/src` was enumerated and reviewed for whether its response or side effects could echo process environment values or platform credentials (tenant trust token, JWT signing keys, database connection strings, provider API keys): 29 distinct gate definitions covering ~90+ call sites, in four groups — sole instance-admin gates (12), instance-admin-or-company-permission gates (10), response-shaping/scope-widening sites (6), and the central `allow_instance_admin` short-circuit in the authorization service (58 `decide()` call sites). Findings and dispositions: - **Environment read/write responses** exposed platform sandbox `envVars`/credential-shaped config to instance admins → closed by floor F1. - **Manual backup trigger** echoed a server filesystem path and triggers a platform-owned operation → closed by floor F3. - **Adapter install/reinstall** loads externally fetched code into the server process (indirect, complete env exposure) → closed by floor F4. The sibling plugin-install path already had a bundled-only floor on managed instances and needed no change. - **Token-minting surfaces** (gateway tokens, custom-image terminal/connection tokens) mint credentials scoped to the instance's own resources, not platform trust anchors → acceptable for an owner-admin of a dedicated instance; unchanged. - All remaining gated surfaces return ordinary instance-scoped business data; none echo `process.env` or platform secrets directly. OAuth client secrets are referenced by env-var *name* only; SSH private keys are stored as secret refs before persistence and are not echoed. Operational note for managed platforms: this model assumes the process environment of a managed instance holds only that instance's own credentials. Platform operators should keep provider credentials per-instance (never fleet-shared) since an instance admin ultimately controls in-process code on their own instance. ## Verification - `pnpm vitest run server/src/middleware/cloud-tenant-actor.test.ts` — resolver matrix: owner × flag on/off, flag via managed overlay (on-over-DB-off and off-over-DB-on), member/admin/support × flag on, no-token self-hosted, fail-closed settings read, purge still runs and no role row is ever inserted (14 tests). - `pnpm vitest run server/src/__tests__/authorization-service.test.ts` — computed flag elevates a `cloud_tenant` actor; a stale `instance_user_roles` row still never does; `session` actors unchanged (full suite, embedded Postgres). - `pnpm vitest run server/src/__tests__/environment-routes.test.ts` — F1: no secret echo to admins on get/list, structural config visible to restricted readers, platform-row PATCH/DELETE rejected for admins (including a marker-stripping patch), marker-clear recovery allowed for stale legacy rows and for a marked sandbox row when no provisioning path is configured, but refused on the managed local row and on the sandbox slot row under a managed-config `environments` entry or the forced kubernetes execution mode, client marker-stamping creates/patches rejected, tenant rows still readable and writable, self-hosted read+write regression (60 tests). - `pnpm vitest run server/src/__tests__/environment-service.test.ts` — `ensureLocalEnvironment` adopts a pre-existing local row on cloud-managed instances (marker stamped, other metadata preserved, idempotent — no rewrite on re-ensure) and leaves self-hosted rows untouched (22 tests, embedded Postgres). - `pnpm vitest run server/src/__tests__/instance-settings-routes.test.ts server/src/__tests__/instance-database-backups-routes.test.ts` — F2 change-vs-echo matrix incl. self-hosted regression; F3 floor for both admin shapes (32 tests). - `pnpm vitest run server/src/__tests__/adapter-routes-authz.test.ts` — F4 floor; self-hosted install/reinstall behavior unchanged (existing cases). - `pnpm vitest run server/src/__tests__/first-admin-claim.test.ts server/src/__tests__/bootstrap-claim-routes.test.ts server/src/__tests__/managed-config.test.ts server/src/__tests__/health.test.ts server/src/__tests__/instance-settings-managed-overlay.test.ts server/src/services/managed-environments.test.ts server/src/services/execution-policy-bootstrap.test.ts` — first-admin bootstrap gate and managed-config behavior unchanged (91 tests). - `pnpm vitest run packages/shared/src/feature-catalog.test.ts` — catalog/schema sync tests cover the new key (selfHostedDefault must equal the schema default). - `pnpm run typecheck` — all 31 workspace projects clean. ## Risks - Self-hosted behavior is unchanged: every floor binds to `isCloudManagedInstance()` (tenant trust token present), the new flag defaults off with no elevation path, and regression tests pin the self-hosted branches. - The elevation is fail-closed and stateless: turning the flag off (managed overlay or DB) de-elevates on the next request; there is no role row to clean up and restores cannot resurrect elevation. - On a cloud-managed instance a pre-existing unmarked local row is adopted (stamped `managedByPaperclip`) by the next ensure and becomes platform-owned — the intended managed-product semantic: the platform owns the single local slot. Self-hosted instances are untouched. - F1 widens restricted readers' view of platform-provisioned rows from fully blanked `config`/`metadata` to structural-only `config` plus markers. Platform-delivered config is guaranteed secret-free by the managed-config contract (secret-shaped keys fail startup), and the floor re-drops secret-shaped keys defensively. - One extra instance-settings read per trusted-header request for owner-role actors (the resolver already performs several queries per request). ## Model Used Claude Fable 5 (Anthropic) — model id `claude-fable-5`, extended thinking enabled, agentic tool use via Claude Code; read-only explore subagents on the same model were used for the surface audit sweep. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
273315a4d0 |
feat(server): provision managed sandbox environments from the managed config (#10324)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Hosted/managed deployments configure instances entirely from the control plane: `PAPERCLIP_MANAGED_CONFIG` already delivers feature flags and `plugins.autoInstall` (bundled sandbox provider plugins), parsed fail-closed at boot > - A plugin alone is not usable for execution: runs need an instance-level `driver: "sandbox"` environment row pointing at the provider, and today only Kubernetes has a boot path for that (`PAPERCLIP_EXECUTION_MODE` → `ensureKubernetesEnvironment`); every other provider requires a manual product-API call the control plane cannot make on a managed instance > - Adding one `ensureXxxEnvironment` per provider would multiply near-identical boot hooks and env-var surfaces > - This pull request generalizes the existing Kubernetes machinery: the managed-config document gains an optional `environments` section that declares a sandbox environment for any bundled provider, ensured idempotently at boot by a provider-agnostic service function (the Kubernetes hook becomes a thin wrapper over it) > - The benefit is that a managed fleet can provision any sandbox provider (Daytona, Modal, E2B, …) purely from configuration — no per-provider code, no manual API calls, no secrets in the document — while self-hosted behavior is untouched ## Linked Issues or Issue Description No public issue exists. Refs #10157 (the cloud image variant that bundles sandbox provider plugins — this PR is the configuration half that makes an installed provider usable). **Problem (feature-request shape):** on a managed instance the control plane can auto-install a bundled sandbox provider plugin via `PAPERCLIP_MANAGED_CONFIG.plugins.autoInstall`, but cannot create the environment row that makes the provider schedulable. The only boot-time environment provisioning is Kubernetes-specific (`PAPERCLIP_EXECUTION_MODE=kubernetes` + `PAPERCLIP_K8S_*`). A generic, config-driven path is needed so any bundled provider can be provisioned without per-plugin code or manual API calls. ## What Changed - `server/src/services/managed-config.ts`: optional `environments` top-level section — `[{ name, description?, provider, config? }]` — validated fail-closed: unknown keys, more than one entry (the DB permits exactly one Paperclip-managed sandbox row, `environments_managed_sandbox_idx`), a `provider` not present in `plugins.autoInstall`, `config.provider`, or secret-looking config keys at any depth (`api_key`/`token`/`secret`/`password`/`credential`) all refuse startup. Absent section ⇒ `environments: []`, so pre-section documents keep booting newer builds. - `server/src/services/environments.ts`: new provider-agnostic `ensureManagedSandboxEnvironment({ name, description?, provider, config?, extraMetadata? })` — idempotently owns the single managed sandbox row: refreshes name/description/config each call, adopts the slot across provider switches (dropping the stale `managedKubernetesSandbox` marker), adopts a same-name unmanaged sandbox row (stamping it managed) instead of colliding on `environments_name_idx` every boot, and falls back to keeping the current name if the desired name belongs to a different row. `ensureKubernetesEnvironment` is now a thin wrapper that pins `provider: "kubernetes"` and stamps the legacy marker. - `server/src/services/managed-environments.ts` (new): `applyManagedEnvironments` boot step — no-op for self-hosted/empty; throws (fail startup) when `PAPERCLIP_EXECUTION_MODE` is also set, since both would own the same managed sandbox row; otherwise ensures each declared environment fail-safe per entry (log + continue boot, matching bundled-plugin provisioning posture). - `server/src/index.ts`: runs the new boot step right after the execution-policy bootstrap, before the heartbeat resumes queued runs. - `server/src/services/index.ts`: exports `applyManagedEnvironments` and `ManagedEnvironmentSpec`. - Secrets stay out of the document by construction: provider credentials reach managed instances only as process env vars (each provider's documented fallback, e.g. `DAYTONA_API_KEY` for the Daytona plugin). ## Verification ```sh cd server pnpm exec tsc --noEmit -p tsconfig.json pnpm exec vitest run \ src/__tests__/managed-config.test.ts \ src/services/managed-environments.test.ts \ src/services/execution-policy-bootstrap.test.ts \ src/__tests__/environment-service.test.ts \ src/__tests__/environment-instance-routes.test.ts \ src/__tests__/environment-routes.test.ts \ src/__tests__/plugin-install-guard.test.ts \ src/__tests__/environment-execution-target.test.ts \ src/__tests__/instance-settings-managed-overlay.test.ts \ src/__tests__/bundled-plugins.test.ts ``` All pass locally (typecheck clean; environment-service suite runs against embedded Postgres and exercises the refactored Kubernetes wrapper plus the new generic ensure: create/refresh, provider switch, unmanaged-row adoption, name-conflict fallback). New tests cover the parser (12 cases incl. secret-key rejection at depth) and the boot step (no-op, mutual exclusion, pass-through, fail-safe). ## Risks - **Self-hosted: none intended.** Without `PAPERCLIP_MANAGED_CONFIG` nothing new executes; the `PAPERCLIP_EXECUTION_MODE=kubernetes` path is regression-covered by the existing bootstrap/service/route suites (all green). - **Behavioral shift in `ensureKubernetesEnvironment` (deliberate):** it now also refreshes `name`/`description` to their managed defaults each boot (desired-state semantics, same as config today) and adopts a `managedByPaperclip` sandbox row that lacks the Kubernetes marker — previously that state made the ensure throw every boot. - **New startup failure modes are all explicit misconfigurations** (malformed section, provider not auto-installed, secret in config, execution-mode conflict) and fail with precise errors; DB-side ensure failures never block boot (fail-safe per entry, logged). - No migrations; no API surface changes. ## Model Used Claude Fable 5 (Anthropic, model ID `claude-fable-5`) with extended thinking and tool use, driving the change end-to-end inside a Claude Code / agent-harness session (code, tests, and verification runs). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (the managed-config module header is the contract doc) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d1b9448b57 |
fix(server): stamp the real build version into images instead of the package.json placeholder (#10257)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work; it ships as a Docker image that self-hosters and managed deployments run. > - The server resolves its own version at runtime in `server/src/version.ts` (`resolveServerVersion()`), which feeds analytics and the server debug panel. > - That resolver derives the real version from `git describe`, and falls back to `server/package.json`'s `version` when git isn't available. > - But `server/package.json`'s version is a static placeholder — CI only stamps the real CalVer at publish, so in source it is never the real version (currently `0.3.1`). > - A Docker image has no `.git` (it's dockerignored), so `git describe` can't run inside it. Every image therefore falls back to the placeholder and reports `0.3.1` in analytics and the debug panel, regardless of which commit it was built from. > - This PR computes the real version once on the CI build runner (where `.git` and tags exist), bakes it into the image, and has `resolveServerVersion()` prefer that stamp when `git describe` is unavailable. > - The benefit: self-hosted and cloud images report their true version instead of a misleading placeholder, with no change to dev checkouts, `git describe`-based resolution, or local `docker build`. ## Linked Issues or Issue Description No public issue exists — describing the bug inline (per the bug report template). **What happened?** Docker images built from `master` (and release tags) report the server version as the `0.3.1` placeholder in analytics and the server debug panel, instead of the real version of the commit the image was built from. **Expected behavior** An image reports the real version of its build commit (e.g. `2026.722.0+51.git.<sha>`), so operators can tell which build is running. **Steps to reproduce** 1. Build the server Docker image from any `master` commit (the `Docker` workflow, `production` target). 2. Run the image and open the server debug panel (or inspect the version reported to analytics). 3. Observe the version is `0.3.1` rather than the commit's real version. **Root cause** `resolveServerVersion()` derives the real version from `git describe`, but the image has no `.git` (dockerignored), so it falls back to `server/package.json`'s `version` — a static placeholder CI only replaces with the real CalVer at publish time. Nothing bakes the real version into the image. **Paperclip version or commit:** reproduces on `master` (`4c55f0d8`) and any published image. **Deployment mode:** self-hosted and managed (both the `production` and `-cloud` images). **Installation method:** Docker image (`ghcr.io/paperclipai/paperclip`). **Related PRs (dedup search):** #9103 (merged — added the `git describe`-based source-install resolution this builds on) and #9637 (closed). Neither bakes a version into the image; this PR closes that gap. No duplicate found. ## What Changed - **`.github/workflows/docker.yml`** — checkout with full history + tags (`fetch-depth: 0`), and a new `Compute build version` step that runs `git describe --tags --match 'v*' --long --dirty` on the pristine runner checkout. The result is passed as a `PAPERCLIP_BUILD_VERSION` build-arg to both the `production` and `-cloud` image builds. - **`Dockerfile`** — the `production` stage takes an `ARG PAPERCLIP_BUILD_VERSION` (default empty) and bakes it into the runtime `ENV`; the `cloud` stage inherits it via `FROM production`. - **`server/src/build-version.ts`** (new) — `readBuildVersion()` / `parseBuildVersion()`, mirroring `build-commit.ts`: reads `PAPERCLIP_BUILD_VERSION` (or a `.paperclip-build-version` file) as a single-token stamp. - **`server/src/version.ts`** — `resolveServerVersion()` prefers the baked build version when `git describe` is unavailable, parsing it with the same rules as a live checkout (`parseGitDescribeVersion`), and falling through to the existing `build-commit` stamp and package version when unset. A live checkout's `git describe` still wins over any stamp. - Tests for the new behavior and the precedence. ## Verification - `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && tsc --noEmit` in `server/` — clean. - `vitest run server/src/__tests__/version.test.ts server/src/__tests__/build-version.test.ts` — **23 tests pass**, covering: stamped version used when git describe fails, stamp parsed to real CalVer, stamp preferred over the build-commit fallback, on-tag stamp collapses to the release version, a pre-resolved stamp used verbatim, and a live git describe still winning over a stamp. - `git describe --tags --match 'v*' --long` for this commit → `v2026.722.0-51-g<sha>`, which `resolveServerVersion()` reports as `2026.722.0+51.git.<sha>` — no longer `0.3.1`. - Not run locally: the full multi-arch image build (CI-only). The workflow change is verified by inspection; the version is computed on the pristine checkout before any lockfile refresh, so it carries no spurious `-dirty`. ## Risks Low. Additive and image-only: - No runtime behavior changes for dev checkouts (git describe still primary and wins over any stamp) or for local `docker build` (empty arg → server keeps its existing fallbacks). - Not a breaking change; no schema or API surface. The stamp is informational (version reporting only). - `fetch-depth: 0` makes the release-image checkout fetch full history/tags — a modest cost on a workflow that already runs at release cadence with a 60-minute budget. - Rollback: revert the commit; images simply return to reporting the placeholder. ## Model Used Claude Opus 4.8 (`claude-opus-4-8`, 1M-context variant), extended thinking, with tool use / code execution — agentic edits, `tsc` + `vitest` runs, and a `git describe` resolution check. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work (bugfix, not core feature work) - [x] I have searched GitHub for duplicate or related PRs and linked them above (#9103, #9637 — related, not duplicates) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`fix/build-version-stamp`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing docs affected; behavior is documented inline in `version.ts` / `build-version.ts` and the workflow/Dockerfile) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
965a827ee7 |
feat(docker): publish a cloud image variant with built bundled plugins (#10157)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed (cloud-hosted) deployments configure instances through `PAPERCLIP_MANAGED_CONFIG`, including a `plugins.autoInstall` key list that the boot-time installer resolves against the bundled plugin catalog > - The installer requires each bundled plugin's `dist/manifest.js` (`server/src/services/bundled-plugins.ts`), but the published image only ships the sandbox providers' *source* — they are intentionally excluded from the pnpm workspace, and the Dockerfile never builds them > - Every managed auto-install therefore logs `bundled plugin bundle not present; skipping auto-install` and no sandbox provider can be provisioned through managed config > - Baking built plugins into the single published image would fix it but makes every self-hosted pull carry the providers' `node_modules` for a managed-only mechanism > - This pull request adds a `cloud` Dockerfile target extending `production` with built bundled plugins — parameterized by build arg and currently just `daytona` — published alongside the default image with a `-cloud` tag suffix > - The benefit is working plugin auto-provisioning for managed deployments while the self-hosted image stays byte-identical and the cloud variant only carries what is actually deployed ## Linked Issues or Issue Description Fixes #10158 (filed for this problem; no prior issue existed — searched for duplicate/related PRs and issues around bundled plugins, docker image variants, and auto-install). Summary: **What happened:** on a managed instance with `plugins.autoInstall: ["daytona"]` delivered via `PAPERCLIP_MANAGED_CONFIG`, boot logs `bundled plugin bundle not present; skipping auto-install` with `pluginPath: /app/packages/plugins/sandbox-providers/daytona`, and the plugin is never installed. **Expected:** the advertised bundled-catalog keys are installable from the published image. **Why:** the image ships plugin source without `dist/` — nothing in the Dockerfile builds the workspace-excluded sandbox providers. ## What Changed - `Dockerfile`: new `cloud-plugins` stage (based on `build`, so devDependencies are available for `tsc`) that installs and builds each provider named in the `CLOUD_BUNDLED_PLUGINS` build arg standalone (`pnpm install --ignore-workspace --no-lockfile && pnpm build`, exactly as the providers' READMEs prescribe), asserting `dist/manifest.js` exists per plugin and failing loudly on unknown names; new `cloud` stage = `production` + the built plugin tree. The arg defaults to `daytona` — the only provider managed deployments auto-install today; every entry adds its `node_modules` to the image, so the list grows only with actual need (a one-line workflow change). - `.github/workflows/docker.yml`: the existing build step is pinned to `target: production` (without this, the new trailing stage would silently become the default build target — this pin is what keeps the self-hosted image identical); new metadata + build-push steps publish the `cloud` target (with `CLOUD_BUNDLED_PLUGINS=daytona`) under the same tag set with a `-cloud` suffix (`sha-<short>-cloud`, `latest-cloud`, `<version>-cloud`), same schema labels, reusing the GHA layer cache ## Verification - All seven sandbox providers build standalone from a clean checkout with the exact commands the new stage runs, each producing `dist/manifest.js` — so the current `daytona` default works and future list additions are known-good - The stage's shell loop was dry-run against the checkout (directory existence + per-plugin assertion logic) - Workflow YAML lints clean - **Not run:** a full multi-arch `docker build` (no local docker daemon). The `cloud` stage is additive and the default target is pinned, so the risk is contained to the new build step; the first master build after merge proves it end-to-end ## Risks - Self-hosted behavior: unchanged. The default image build is pinned to the `production` target, which produces the same layers as before this change; the `cloud` stages run only for the new build step. - The plugin installs in the `cloud-plugins` stage use `--no-lockfile` (the providers are workspace-excluded and lockfile-less by design), so plugin dependency resolution is not pinned at image-build time. This mirrors the existing Plugins-page install path, which resolves from npm at install time. - CI cost: one additional build-push per master push. It reuses the layer cache from the production build, so the marginal work is the single plugin's build layers. - An unknown name in `CLOUD_BUNDLED_PLUGINS`, or a provider that stops producing `dist/manifest.js`, fails the cloud build loudly rather than publishing a broken variant. ## Model Used Claude (Anthropic), model ID `claude-fable-5[1m]` via Claude Code CLI — extended thinking and tool use (code edits, standalone plugin build verification, workflow lint). ## Checklist - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] Self-hosted behavior unchanged (default build target pinned to `production`) - [x] One clear change: publish a cloud image variant with built bundled plugins |
||
|
|
14f20be92b |
ci: harden Docker image build workflow against lockfile drift and runner disk exhaustion (#10142)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Docker image publish workflow (`.github/workflows/docker.yml`) builds and pushes the multi-arch `ghcr.io` image on every master push, so users pulling the container get the latest code > - The two newest master runs of that workflow failed, so no images have been published past a recent master commit > - The failures had two distinct causes: run [30054330748](https://github.com/paperclipai/paperclip/actions/runs/30054330748) hit `ERR_PNPM_LOCKFILE_CONFIG_MISMATCH` (committed `pnpm-lock.yaml` drifted from `patchedDependencies` in package metadata), and run [30050197392](https://github.com/paperclipai/paperclip/actions/runs/30050197392) hit `no space left on device` during the multi-arch buildx export > - This pull request hardens the publish job against both failure modes: it refreshes the lockfile (lockfile-only, guarded) before the build, and frees runner disk space before buildx setup > - The benefit is that image publishing keeps working through routine lockfile drift and the growing multi-arch build footprint, so `ghcr.io` images stay current with master ## Linked Issues or Issue Description - Refs #8286 — same class of Docker-build lockfile mismatch failure - Refs #8827 — pnpm 9.15.x pin / lockfile regeneration discussion - Note: the immediate lockfile drift on master was fixed by #10132; the refresh step here prevents the *next* drift from breaking image publishing again ## What Changed - Added a pnpm + Node setup and a **"Refresh lockfile for Docker build context"** step to the image job in `.github/workflows/docker.yml`: runs `pnpm install --lockfile-only --ignore-scripts --no-frozen-lockfile`, exits cleanly if nothing changed, and **fails the job if anything other than `pnpm-lock.yaml` was modified** by the refresh - Added a **"Free runner disk"** step (before buildx setup) that prunes the pnpm store, apt caches, preinstalled toolchains (`/usr/share/dotnet`, Android SDK, Swift, Boost, PowerShell, GHC, CodeQL/PyPy/Ruby toolcache), and dangling Docker state, logging `df -h` before/after - No changes outside the workflow file (54 added lines, nothing removed) ## Verification - Pulled the logs of both failed master runs and matched each failure to the step that addresses it: [30054330748](https://github.com/paperclipai/paperclip/actions/runs/30054330748) failed with `ERR_PNPM_LOCKFILE_CONFIG_MISMATCH`, [30050197392](https://github.com/paperclipai/paperclip/actions/runs/30050197392) failed with `no space left on device` during the buildx export - Confirmed pnpm `9.15.4` in the new setup step matches the repo `packageManager` field and the version used in the Dockerfile, so the refreshed lockfile is generated by the same pnpm the image build consumes - Validated the workflow YAML parses cleanly - The workflow triggers on master pushes / manual dispatch; the definitive check is the first master run after merge — reviewers can also `workflow_dispatch` it from this branch if desired ## Risks - The lockfile refresh runs with `--ignore-scripts` and a guard that aborts on any non-lockfile change, so it cannot silently pull unexpected code into the image; worst case it fails the job with a clear diff - The published image could be built from a refreshed lockfile that differs from the committed one when drift exists — that keeps publishing alive but can mask drift on master, which still needs the committed lockfile fixed (as #10132 did) - Disk cleanup removes preinstalled toolchains only on the ephemeral runner for this job; other jobs/workflows are unaffected - Low risk overall: additive steps in a single workflow file ## Model Used - Claude (Anthropic) — `claude-fable-5` (Claude Code agent harness, extended thinking, tool use). Used to diagnose the failing CI runs from logs, author the workflow changes, and prepare this PR. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (no runtime code touched; workflow YAML validated — see Verification) - [ ] I have added or updated tests where applicable (n/a — CI workflow change) - [x] I have updated relevant documentation to reflect my changes (none needed) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm once checks run) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review pass) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
176a9e8230 |
fix(built-in-agents): allow first-time setup of a needs_setup built-in under board-approval policy (#10129)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Companies can require **board approval for new agents**; built-in agents (e.g. the Reflection Coach / Briefs) are provisioned through the `built-in-agents` service `provision()` > - Some built-in agents are *auto-provisioned* as a hire that, once approved, resolves to an idle agent row whose `adapterConfig` is still empty — status `needs_setup` > - When the board operator then opens that agent's setup dialog and submits the adapter config, `provision()` saw `adapterType`/`adapterConfig` on an already-existing row and classified it as a **reconfiguration**, throwing a dead-end 409: *"Built-in agent adapter changes require board approval before they can be applied."* > - The operator *is* the board, so there was no one left to grant an approval they already implicitly hold — setup could never be completed > - This pull request treats first-time adapter setup of a `needs_setup` built-in as the first-time configuration it actually is, applying it directly while still gating genuine reconfiguration of a live agent > - The benefit is the board can finish setting up an auto-provisioned built-in agent without hitting an unsatisfiable approval wall ## Linked Issues or Issue Description <!-- No public GitHub issue exists; describing the underlying bug in-PR following the bug_report template. --> **What happened?** With "require board approval for new agents" enabled, completing the adapter setup of an auto-provisioned but unconfigured built-in agent (status `needs_setup`, e.g. the Reflection Coach) failed with a 409 — *"Built-in agent adapter changes require board approval before they can be applied."* — even for the board user. Because the operator *is* the board, no additional approver existed, so setup was permanently blocked. Root cause: in `builtInAgentService.provision()`, any request carrying `adapterType`/`adapterConfig` against an existing row was treated as a reconfiguration and gated, regardless of whether that row had ever completed its initial adapter setup. An auto-provisioned hire resolves to an idle row with an empty `adapterConfig` (`needs_setup`), so its very first configuration was misclassified. **Expected behavior** The board can complete first-time setup of an already-sanctioned built-in agent without a fresh approval, matching the behavior when board approval is not required. Genuine reconfiguration of an already-configured (`ready`/`paused`) agent should still require approval. **Steps to reproduce** 1. In a company with `requireBoardApprovalForNewAgents` enabled, have a built-in agent auto-provisioned so its row exists but its adapter is unconfigured (status `needs_setup`). 2. As the board user, open that agent's setup dialog and submit an adapter type + config. 3. Observe the 409 "Built-in agent adapter changes require board approval before they can be applied." with no way for the board to grant the approval. **Deployment mode** Local single-instance / self-hosted (server `built-in-agents` service). ## What Changed - `server/src/services/built-in-agents.ts`: In `provision()`, when the existing built-in row has **not** yet completed adapter setup (`!hasCompleteAdapterConfig(...)`, i.e. `needs_setup`), first-time adapter configuration now applies directly via `ensure()` — the same path used when board approval is not required. The hire that created the row was already sanctioned, so no fresh approval is required. - Reconfiguration of an already-configured (`ready`/`paused`) built-in agent stays gated behind board approval exactly as before, and `pending_approval` rows are handled before the new branch. - `server/src/__tests__/built-in-agents.test.ts`: Added a regression test — under `requireApproval: true`, completing first-time setup of a `needs_setup` built-in returns `approval: null`, transitions the agent to `ready`, and creates **no** approval row. ## Verification ```bash cd server npx vitest run src/__tests__/built-in-agents.test.ts # Test Files 1 passed (1) # Tests 31 passed (31) ``` - New test `completes first-time setup of a needs_setup built-in without a fresh board approval` passes. - Full `built-in-agents.test.ts` suite (31 tests) passes, including existing tests that assert genuine reconfiguration of a configured agent **remains** gated. ## Risks Low risk. The change narrows an over-broad approval gate: it only opens the direct-apply path for rows that have never completed adapter setup (`needs_setup`), determined by the existing `hasCompleteAdapterConfig` predicate that already drives `deriveBuiltInAgentStatus`. Already-configured (`ready`/`paused`) agents, and `pending_approval` rows, are unaffected and still gated. No schema or migration changes. ## Model Used Claude Opus 4.8 (`claude-opus-4-8`), 1M context, extended thinking, with tool use / code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (searched my open PRs and compared patch-ids — no duplicate exists) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
429792f1f3 |
fix(interactions): stop wedging confirmation accept on a terminal workspace_finalize (#10099)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - When an agent finishes work in an execution workspace, the board can confirm the result through an issue-thread interaction (e.g. the "Merged" / mark-done confirmation button on a `request_confirmation`). > - That accept action is gated: it must not race a worktree sync-back (`workspace_finalize`) that is still copying the agent's commits out of the sandbox, or the board could act on a base that hasn't received them yet. > - The gate (`runWorkspaceIsFinalized`) treated the sync-back as "settled" only when the latest `workspace_finalize` op was `succeeded` — so a run whose finalize reached a terminal `failed` state, or died leaving a stale `running` op, was treated as "still syncing" forever. > - Users hit a permanent, misleading `... has not finished syncing its workspace` error and could never click "Merged", even though nothing was syncing and the run had long since ended. > - This PR fixes the settle semantics so the gate blocks only while a sync-back is genuinely pending or in flight, and treats any terminal (or stale-orphaned) finalize as done. > - The benefit is that a failed or abandoned sync-back no longer wedges the human confirmation, while a genuinely in-flight sync-back on a live run still blocks correctly. ## Linked Issues or Issue Description No public GitHub issue exists for this. Describing the bug in-PR (bug report): **What happened** Clicking the "Merged" / mark-done confirmation at the bottom of an issue thread returns an error that the workspace "has not finished syncing its workspace" — but nothing is actually syncing, and the run that created the interaction has already ended. The confirmation is permanently stuck; the only workaround is to merge and mark the task done manually. **Expected behavior** Once the source run's worktree sync-back has finished — whether it succeeded, failed, or was skipped — the confirmation should be acceptable. The gate should block only while a sync-back is genuinely still running on a live run. **Steps to reproduce** Have an agent run reach `workspace_finalize` and end without a `succeeded` finalize (e.g. the sync-back fails, or the run process dies mid-finalize leaving a `running` op). Then attempt to accept the `request_confirmation` interaction it created → 409 "... has not finished syncing its workspace" with no way to proceed. **Paperclip version or commit** Reproduced on the current `master` line (server service); root cause is in `runWorkspaceIsFinalized` in `server/src/services/issues.ts`. **Deployment mode** Local / self-hosted instance (server service). **Root cause** `runWorkspaceIsFinalized` returned `true` only when the latest `workspace_finalize` operation was `succeeded`. A terminal `failed` finalize (the sync-back ran and failed; it will not retry within that run) and a `running` finalize left behind by a dead run both left the gate closed forever. ## What Changed - `runWorkspaceIsFinalized` (server/src/services/issues.ts) now treats a sync-back as **settled** when the latest `workspace_finalize` op reached any terminal status (`succeeded`, `failed`, or `skipped`), instead of only `succeeded`. - A `workspace_finalize` still marked `running` blocks only while its owning run is alive; a `running` record left behind by a terminal/missing run is treated as stale (settled), so a dead run can no longer wedge the gate. - Preserved existing behavior for the other cases: no operations recorded at all → settled; earlier phases recorded but no `workspace_finalize` yet → still blocks (the sync-back hasn't been attempted). - Extracted the run-liveness check into a shared exported helper `heartbeatRunIsTerminalOrMissing` and reused it from the existing `isTerminalOrMissingHeartbeatRun` closure (no behavior change there). - Added a short comment at the confirmation-accept gate (server/src/services/issue-thread-interactions.ts) documenting the relaxed settle semantics. - The dependency-readiness / blocker barrier (`listPendingFinalizeBlockerIssueIds`) is deliberately left unchanged: an automated dependent must not proceed onto a base that never received a blocker's synced-back commits, so a failed finalize keeps that gate closed. Only the human-driven confirmation accept is relaxed. - Added regression tests for: failed finalize, stale `running` finalize on a dead run, and a genuinely `running` finalize on a live run (must still block). ## Verification - `cd server && node_modules/.bin/vitest run src/__tests__/issue-thread-interactions-service.test.ts -t "accept"` → 17 passed (includes the 3 new regression tests), 21 unrelated tests skipped by the name filter. - Manual reasoning walkthrough of `runWorkspaceIsFinalized` for each op-history shape (no ops / earlier-phase-only / terminal finalize / running-on-dead-run / running-on-live-run) confirms the intended block-vs-settle outcome. ## Risks - Low risk and narrowly scoped to the human confirmation-accept gate. The only behavioral change is that a terminal (`failed`/`skipped`) or stale-orphaned `running` finalize now settles the gate instead of blocking forever. - A genuinely in-flight sync-back on a live run still blocks (covered by a regression test), so the accept cannot race commits that are actively being synced back. - The blocker/dependency barrier for automated dependents is unchanged, so no dependent will be advanced onto a base missing a failed blocker's commits. ## Model Used - Provider/model: Claude (Anthropic), **Opus 4.8**, model ID `claude-opus-4-8`, 1M context window. - Capabilities used: extended thinking, tool use (repo inspection, local test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0ef3b320c7 |
Harden POST /plugins/install: canonicalize localPath for all instances; bundled-only floor for managed instances (#10067)
**Builds on** #10058 — managed detection keys off the *presence* of the `PAPERCLIP_MANAGED_CONFIG` env var that PR introduces, deliberately never its parsed body. **Summary.** Two layered hardenings of the plugin install route. (1) For **all** instances: `localPath` installs previously skipped the package-name validation entirely; the path is now null-byte-checked, resolved absolute, `realpath`'d (collapsing `..` traversal and symlinks), and required to be an existing directory before the loader ever sees it. (2) For instances running under a managed hosting control plane (detected by the *presence* of `PAPERCLIP_MANAGED_CONFIG` — deliberately never its body, so a corrupted document cannot widen the surface): registry/npm installs return 403, and `localPath` installs must canonicalize to inside the bundled plugin catalog root (`packages/plugins`) — a positive allowlist enforced in code at the route, independent of any flag value. Self-hosted behavior is otherwise unchanged. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The plugin system lets instance admins install plugins from a registry or from a local filesystem path, and plugin installation is code execution on the host > - The `localPath` branch of `POST /plugins/install` skips the validation applied to registry installs; the raw path reaches the plugin loader without canonicalization > - Separately, instances operated by a managed hosting control plane must constrain installs to the bundled plugin catalog, because there the host belongs to the operator, not the tenant > - This pull request canonicalizes and validates `localPath` for all instances, and adds a bundled-only install floor for managed instances > - The benefit is a smaller install-route attack surface everywhere, and a positive code-enforced allowlist where the operator owns the machine ## Linked Issues or Issue Description No public issue exists; `bug_report` template fields for the validation gap this PR fixes: - **What happened:** `POST /plugins/install` with `localPath` set bypasses the package-name validation entirely; the un-canonicalized path (relative segments, symlinks, no existence check) is handed straight to the plugin loader. - **Expected behavior:** path installs are validated like registry installs — null-byte-checked, resolved absolute, `realpath`'d, and required to be an existing directory before the loader sees them. - **Steps to reproduce:** as an instance admin, call `POST /plugins/install` with a `localPath` containing `..` traversal or a symlink pointing outside any plugin directory; observe the loader receives the raw path. Exploitability is bounded (the route already requires instance admin), so this is hardening of an admin-only surface rather than an open exploit. - **Version:** current `master`. The managed-instance bundled-only floor layered on top is new behavior (motivation: on managed hosting, arbitrary plugin install is arbitrary code execution on operator infrastructure), aligned with the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed - New `server/src/services/plugin-install-guard.ts` — three pure primitives: managed detection (presence-based), path canonicalization (null-byte check → absolute resolve → `realpath` → must be an existing directory), and segment-based containment in the bundled plugin catalog root. - Route enforcement in `server/src/routes/plugins.ts`: npm/registry installs return 403 on managed instances; `localPath` installs are canonicalized on every instance and, on managed instances, must land inside the bundled catalog root. - The plugin loader now receives the canonical path instead of the raw request string. ## Verification - 15 guard unit tests (`server/src/__tests__/plugin-install-guard.test.ts`): traversal, symlink escape, null byte, file-vs-directory, string-prefix sibling root. - 13 route security tests (`server/src/__tests__/plugin-install-route-security.test.ts`): 403 matrix on managed instances + self-hosted happy paths. - 36 existing plugin route authz tests green (`server/src/__tests__/plugin-routes-authz.test.ts`). - Server `tsc --noEmit` clean. ```bash cd server pnpm vitest run src/__tests__/plugin-install-guard.test.ts src/__tests__/plugin-install-route-security.test.ts src/__tests__/plugin-routes-authz.test.ts pnpm exec tsc --noEmit ``` ## Risks - Managed instances: npm/registry installs and out-of-catalog `localPath` installs now return 403 — intended new behavior, enforced in code rather than configuration. - All instances: `localPath` installs that previously pointed at nonexistent paths or non-directories now fail with 400 before reaching the loader (previously the loader failed later, less safely). Symlinked deployment layouts are handled by canonicalizing both sides of the containment check. - Self-hosted npm install path is unchanged. Low residual risk. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
04e070bf45 |
docs(skill): require agents to claim only monitors they actually scheduled (#10064)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work, and the shared `skills/paperclip/SKILL.md` is the behavioral
contract every managed agent follows each heartbeat.
> - Issue continuation between heartbeats depends on real, persisted
state: an issue only auto-resumes when it has a scheduled **issue
monitor** (`monitorNextCheckAt` + an execution-policy `monitor` block)
that the server's `tickDueIssueMonitors` scheduler polls and re-wakes
via `issue_monitor_due`.
> - A run/heartbeat is an ephemeral execution window — nothing keeps
"watching" after it exits — but the skill never said this, so agents
narrated a "watcher in this run" as if a live subscription existed.
> - That gap produced a concrete user-facing failure: an agent claimed
"a watcher in this run wakes me when CI + Greptile complete," then the
run ended with no monitor scheduled and nothing ever resumed, leaving
the user unsure whether a watcher existed at all.
> - This PR closes the gap by documenting what a monitor actually is and
adding hard rules so agents only claim a watcher they have actually
scheduled, describe it in checkable terms, and never imply a live
watcher on a task they mark `done`.
> - The benefit is that agent narration stays consistent with the
disposition guard and recovery classifier that already enforce these
paths in state, so users get accurate expectations about whether and
when a task will resume.
## Linked Issues or Issue Description
This is a documentation-only change to a shared agent skill, so no code
issue is required. The underlying problem it addresses:
**Problem or motivation** — Agents were telling users that a "watcher in
this run" would wake them when external checks (CI, Greptile) finished,
when no persisted issue monitor had been scheduled. Because a heartbeat
is ephemeral, no such watcher exists after the run exits, so the task
silently never resumed and the user was left confused about what would
happen next.
**Proposed solution** — Document, in the shared skill, exactly what an
issue monitor is (durable `monitorNextCheckAt` + execution-policy
`monitor` block, polled by `tickDueIssueMonitors`, re-woken via
`issue_monitor_due`) and add rules that agents may only claim a
watcher/monitor after actually scheduling one, must describe it in
checkable terms (kind / next check / timeout / attempts), and must never
imply a live watcher on a task being marked `done`.
**Alternatives considered** — Enforcing purely in server state (the
disposition guard and recovery classifier already reject
`in_review`/parked issues without a real wake path). That enforcement
exists but does not stop an agent from *narrating* a non-existent
watcher in a comment; aligning the skill guidance with the existing
state enforcement is the missing piece.
## What Changed
- Added a **"Monitors and Watchers (say only what you actually
scheduled)"** subsection to `skills/paperclip/SKILL.md` explaining that
a watcher does not live inside a run, and that only a persisted issue
monitor can auto-resume an issue (with the concrete fields and the
`tickDueIssueMonitors` / `issue_monitor_due` polling path).
- Added three behavioral rules: only claim a monitor after scheduling
one (and how to schedule/confirm it via `PATCH /api/issues/{id}` and
`monitor/check-now`); describe monitors in checkable terms; never imply
a live watcher on a task marked `done`.
- Cross-referenced the rule from the **Critical Rules** list.
- Tightened the final-disposition checklist so `in_review` /
`in_progress` continuation requires a real, non-null
`monitorNextCheckAt` rather than a merely described one.
## Verification
- Docs-only change to `skills/paperclip/SKILL.md`; no code paths are
affected.
- Confirmed every identifier referenced in the new text is real in the
codebase: `monitorNextCheckAt`, `monitorScheduledBy`,
`executionPolicy.monitor`, `tickDueIssueMonitors`, and the
`issue_monitor_due` wake reason.
- Rendered the Markdown to confirm the new subsection and the Critical
Rules bullet display correctly and links resolve within the document.
- `git diff` confirms the change is limited to the single skill file (14
insertions, 2 deletions).
## Risks
Low risk. This is guidance text in a shared agent skill with no runtime
or schema impact. Worst case is stylistic wording that can be refined in
a follow-up; it cannot break builds, migrations, or behavior. It
strengthens (never loosens) the existing disposition guarantees.
## Model Used
Claude Opus 4.8 (model id `claude-opus-4-8`, 1M-context variant) running
in an agent harness with extended thinking and tool use (file edit,
shell, git, GitHub CLI).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [ ] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
62e367c1b2 |
fix(ui): rewrite recovery and blocked-notice copy in plain language (#10065)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The issue-detail UI shows recovery cards and blocked/parked notices when a task loses its next step — a run finished with no disposition, a task is stranded, work is blocked behind other tasks, or an assigned item sits in the backlog > - That copy was written in the scheduler's internal vocabulary — "Corrective wake queued", "Graph Liveness", "lost a live action path", "the responsible" — which describes Paperclip's internals rather than the user's situation > - Users seeing these cards report having no idea what the card means or what they are supposed to do > - This pull request rewrites the user-facing copy in plain language and adds explicit calls to action that match the options in the card's Resolve menu > - The benefit is that a non-expert operator can read a recovery or blocked notice and immediately understand what happened and which action to take next ## Linked Issues or Issue Description No public GitHub issue exists for this; describing the problem here (bug-report format): - **What happened:** Recovery action cards and blocked notices render internal jargon, e.g. the headline "Paperclip detected this task lost a live action path. A recovery owner needs to act.", the chip "Corrective wake queued", the kind label "Graph Liveness", and phrases like "Comments still wake the responsible". Status values also appear as raw code literals (`in_progress`, `todo`). - **Expected behavior:** These notices should tell a normal user, in plain language, what happened and what to do next (retry the task, mark it done, send it for review, or record a blocker). - **Impact:** Operators stall on tasks that only need a simple disposition because the UI doesn't tell them that's what is being asked. Related prior work: #9417 (merged) made the reopen-suppressed blocked message explicit; this PR extends the same plain-language treatment to the rest of the recovery and blocked-notice copy. ## What Changed - Recovery card headlines for `missing_disposition`, `stranded_assigned_issue`, and `issue_graph_liveness` now say what Paperclip found and name the concrete next steps ("try the task again, mark it done, or send it for review") matching the card's Resolve menu. - The `issue_graph_liveness` kind label "Graph Liveness" is now "Task Needs Next Step", and the "Wake" metadata row is now "Follow-up". - Wake-policy chips describe actual behavior: "An agent will be asked to choose the next step" (was "Corrective wake queued"), "Board will decide", "Manual follow-up needed", "Repair needed before retry", "Check scheduled". - Blocked/waiting/parked notices say "the assignee" instead of "the responsible" / "responsible agent", and "notify" instead of "wake". - The still-needs-a-next-step notice drops raw `in_progress` code literals and keeps a plain-language option list (mark done or cancelled, send for review, record what is blocking it, delegate follow-up). - Parked-backlog notice renders "To do / In progress" as plain labels instead of code literals. - Component tests updated to pin the new copy and the successful-run example options. ## Verification - `cd ui && npx vitest run src/components/IssueRecoveryActionCard.test.tsx src/components/IssueBlockedNotice.test.tsx src/components/IssueAssignedBacklogNotice.test.tsx src/components/IssueChatThread.test.tsx` — 4 files, 126 tests, all passing. - Copy-only review: the diff touches display strings, one label map entry, and test assertions; no control flow, props, or identifiers change. ## Risks - Low risk — user-facing strings and test updates only. No behavior, API, or schema changes. The only functional surface is that anything keying off the displayed text (e.g. screenshots, external docs) will show the new wording. ## Model Used - Claude (Anthropic) — Claude Fable 5, model ID `claude-fable-5`, extended thinking enabled, running in Claude Code with agentic tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no docs reference this copy) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0f898f95af |
Render managed experimental settings as locked, with a 'Managed by Paperclip Cloud' badge (#10061)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The experimental settings page renders one interactive toggle per feature from the settings API > - On managed instances some values are enforced by the hosting control plane, and the API now reports those keys as managed > - Rendering enforced values as live toggles misleads users: the click appears to work, and the value silently snaps back > - This pull request renders managed keys as locked toggles with a "Managed by Paperclip Cloud" badge and guards the handlers so no PATCH can be emitted > - The benefit is UI honesty on managed instances, with self-hosted responses rendering exactly as before ## Linked Issues or Issue Description Builds on #10058 — renders the per-key `managedKeys` metadata #10058 adds to settings responses (typing shared from #10058 at rebase). No public issue exists; `feature_request` template fields: - **Problem or motivation:** on managed instances users see fully interactive toggles for settings the control plane enforces; changes appear to apply and never do, with no explanation. - **Proposed solution:** disabled toggle + badge + guarded handler driven by the settings response's managed-key metadata; the ~17 uniform setting cards are extracted into one shared component with copy, aria-labels, and patch payloads preserved verbatim. - **Alternatives considered:** hiding managed settings entirely (users lose sight of the effective value and why it is fixed); tooltip-only hints on still-active toggles (doomed PATCHes are still emitted and stripped server-side). - **Roadmap alignment:** supports the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed - When the settings API reports a feature key as managed (`managedKeys` from the managed-config overlay), the experimental settings page renders that toggle disabled with a badge and a guarded handler, so a click can never emit a PATCH. Previously, managed-instance users saw fully interactive toggles they could never actually change. Self-hosted responses (no `managedKeys`) render exactly as before. - The ~17 copy-pasted uniform setting cards are extracted into one `ExperimentalToggleCard` component with titles, descriptions, footnotes, aria-labels, and patch payloads preserved verbatim; the two bespoke cards get inline managed handling (the managed auto-recovery toggle also cannot open its preview dialog). - `ui/src/api/instanceSettings.ts` response typing now uses the shared `InstanceExperimentalSettingsWithManaged` / `ManagedSettingMetadata` types from #10058; `ui/src/pages/InstanceExperimentalSettings.tsx` locked rendering + card extraction; tests. ## Verification - 24 page tests (20 existing unmodified + 4 new: locked badge with no PATCH while unmanaged keys stay editable; managed auto-recovery opens no dialog; an open recovery preview closes with no PATCH when a refresh marks auto-recovery managed; self-hosted unaffected): `pnpm --filter @paperclipai/ui exec vitest run src/pages/InstanceExperimentalSettings.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` clean ## Risks - Low risk. UI-only change; no server or API behavior changes. Self-hosted responses carry no `managedKeys`, so the page renders exactly as before there. The card extraction preserves copy, aria-labels, and patch payloads verbatim, covered by the 20 pre-existing page tests passing unmodified. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c4cdcc4826 |
Generalize bundled plugin provisioning: ensureBundledKubernetesPlugin → ensureBundledPlugins (#10063)
**Builds on** #10058 — reads `plugins.autoInstall` from the parsed managed-config contract #10058 introduces (the interim `readManagedPluginAutoInstall` shim is retired at rebase). **Summary.** Boot-time bundled-plugin provisioning becomes catalog-driven. A new bundled-plugin catalog lists the sandbox providers shipped in-tree (keys like `kubernetes`, `daytona` → plugin key + path under the catalog root). Managed instances read `plugins.autoInstall` from `PAPERCLIP_MANAGED_CONFIG`; unknown keys or paths escaping the catalog root (symlinks resolved) **throw before listen** — a managed instance refuses to start rather than boot half-provisioned. Installation keeps today's mechanism: an in-process, fail-safe `loader.installPlugin({ localPath })` under a system actor — no HTTP route, no user, no role widening. Self-hosted boot is unchanged (kubernetes bundle only, existing env override honored, install failures still log-and-continue). **Semantics.** A plugin already present in any non-uninstalled state is skipped, so an operator-disabled plugin is never silently re-enabled; managed mode reinstalls soft-uninstalled bundles (the control plane owns provisioning); removal from the autoInstall list never auto-uninstalls. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox-provider plugins ship in-tree, but boot-time provisioning is hard-coded to exactly one of them (Kubernetes) via a bespoke function > - On managed hosting, tenant users have no install privileges, so any bundled plugin that is not provisioned at boot is unusable > - Widening install routes or granting roles to fix that would trade a provisioning gap for a security regression > - This pull request generalizes the existing boot installer into a catalog-driven `ensureBundledPlugins`, fed by `plugins.autoInstall` from `PAPERCLIP_MANAGED_CONFIG` > - The benefit is that managed tenants get working bundled plugins out of the box, through the same in-process, role-free mechanism the codebase already trusts, while self-hosted boot is unchanged ## Linked Issues or Issue Description No public issue exists; `feature_request` template fields: - **Problem or motivation:** on managed instances tenant users cannot install plugins (by design they never hold instance admin), so even plugins shipped with the product are unusable; boot provisioning currently knows only the Kubernetes bundle. - **Proposed solution:** a bundled-plugin catalog plus `ensureBundledPlugins(keys)` driven by the managed config; same in-process `loader.installPlugin({ localPath })` under a system actor; unknown keys or catalog-escaping paths fail startup; already-present plugins are skipped so operator-disabled plugins are never silently re-enabled. - **Alternatives considered:** granting tenant users install privileges (widens secrets/adapters/settings access to solve a one-button problem); a separate non-admin install route for bundled plugins (new authz surface; provisioning removes the need for any install action at all). - **Roadmap alignment:** supports the in-progress "Cloud deployments" milestone and builds on the shipped sandbox-provider milestone in `ROADMAP.md`. Refs #10058. ## What Changed - New `server/src/services/bundled-plugins.ts`: the bundled-plugin catalog, the fail-to-start resolver (`resolveBundledPluginInstalls`, positive allowlist + catalog-root containment with symlinks resolved), and the fail-safe installer (`ensureBundledPlugins`). - `server/src/app.ts`: replaces the hard-coded `ensureBundledKubernetesPlugin` boot hook with resolver + installer wiring, with test hooks (`managedPluginAutoInstall`, `bundledPluginCatalogRoot` options). - `server/src/index.ts`: passes `plugins.autoInstall` from the single fail-closed `PAPERCLIP_MANAGED_CONFIG` startup parse (#10058) into `createApp`; absent env means self-hosted and changes nothing. ## Verification - 24 new tests in `server/src/__tests__/bundled-plugins.test.ts` (catalog resolution, containment incl. symlink and `..` escapes, skip/reinstall matrix, self-hosted invariants, installer error paths) — all green. - 85 adjacent startup/plugin-route/auto-build/managed-config tests green (`managed-config`, `instance-settings-managed-overlay`, `plugin-install-autobuild`, `plugin-routes-authz`, `server-startup-feedback-export`). - Server `tsc --noEmit` clean. ```bash cd server npx vitest run src/__tests__/bundled-plugins.test.ts npx vitest run src/__tests__/managed-config.test.ts src/__tests__/instance-settings-managed-overlay.test.ts src/__tests__/plugin-install-autobuild.test.ts src/__tests__/plugin-routes-authz.test.ts src/__tests__/server-startup-feedback-export.test.ts npx tsc --noEmit ``` ## Risks - Managed instances with a malformed or unknown `plugins.autoInstall` entry now **refuse to start** (fail closed, by design) instead of booting half-provisioned; harness misconfiguration surfaces as a precise startup error. - Self-hosted behavior is unchanged (kubernetes bundle only, `PAPERCLIP_KUBERNETES_PLUGIN_PATH` honored without containment, install failures log-and-continue), so the default deployment path carries low risk. - No uninstall path exists in this module; removal from the autoInstall list can leave a previously provisioned plugin installed (intentional v1 semantics, documented in code). ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
216d3d2680 |
Managed-instance config: fail-closed PAPERCLIP_MANAGED_CONFIG parsing and read-time settings overlay (#10058)
**Builds on.** #10055 — the `catalogVersion` this config document pins is the feature-catalog artifact #10055 emits. **Summary.** Instances operated by a managed hosting control plane can now receive instance configuration through a single environment variable, `PAPERCLIP_MANAGED_CONFIG` (versioned JSON: `mode`, `catalogVersion`, `features`, `plugins.autoInstall`). When the variable is absent the instance is self-hosted and nothing changes. When present, parsing is strict and **fail-closed**: blank value, malformed JSON, unknown feature key, a feature key this build's feature catalog does not mark tier `managed`, missing required section, or unsupported version refuses startup with a precise error — a typo that silently does nothing is how a security control quietly fails. Managed feature values are overlaid **at read time** inside the instance settings service (never persisted), so a DB restore or manual row edit cannot resurrect a disabled capability; responses expose per-key `managedKeys` metadata (`managed: true`, `managedBy`) so clients can render locked state. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs both self-hosted and under managed hosting, where an operator's control plane owns instance configuration > - Today instance feature settings live only in the tenant database; a hosting control plane has no way to enforce a configuration that tenant-side writes or restores cannot undo > - Managed configuration will carry security posture, so delivery must be atomic and parsing must fail closed — a typo that silently does nothing is how a security control quietly fails > - This pull request adds strict parsing of one `PAPERCLIP_MANAGED_CONFIG` env var and overlays its feature values at read time inside the settings service, never persisting them > - The benefit is a minimal, auditable managed-hosting contract: absent var ⇒ self-hosted instances are byte-for-byte unchanged; present ⇒ deterministic, locked configuration surfaced to clients via per-key managed metadata ## Linked Issues or Issue Description Refs #966 — this PR delivers that issue's "managed config injection" hook, via a strict env-var contract rather than the config-file path it sketches; the issue's other hooks (identity header, health, usage webhook, lifecycle, external secrets, IAM auth) are out of scope, so the PR refs rather than closes it. *Mechanism differs from #966's proposal, so the `feature_request` fields are also filled in:* - **Problem or motivation:** managed hosting deployments need to centrally enable/disable instance features; DB-stored settings can be edited, restored, or migrated back to permissive values, and nothing marks a value as operator-enforced. - **Proposed solution:** one versioned JSON env var; fail-closed parse at startup; read-time overlay in the settings service (precedence: managed value over stored value over schema default); `managedKeys` metadata in settings responses so clients can render locked state. - **Alternatives considered:** per-feature env vars (non-atomic across a half-updated env set, unbounded env surface); seeding the DB at boot (persisted values can be edited or restored over, and cannot express "forced"); lenient warn-and-drop parsing (fails open — unacceptable for a security-bearing control). - **Roadmap alignment:** supports the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed - New `server/src/services/managed-config.ts` (pure parser over the env record) - Startup parse ordered before the first `instanceSettingsService` construction in `server/src/index.ts` - Read-time merge + `managedKeys` in the settings service - Shared validator updates ## Verification - 29 parser/overlay tests (fail-closed matrix incl. blank/whitespace env, missing sections, catalog-tier mismatch, empty-section happy path): `pnpm vitest run src/__tests__/managed-config.test.ts src/__tests__/instance-settings-managed-overlay.test.ts` (from `server/`) - 40 existing settings route/service tests green: `pnpm vitest run src/__tests__/instance-settings-routes.test.ts src/__tests__/instance-settings-service.test.ts` (from `server/`) - 15 shared validator tests: `pnpm vitest run src/validators/instance.test.ts` (from `packages/shared/`) - Server `tsc --noEmit` clean: `pnpm typecheck` (from `server/`) ## Risks - Self-hosted instances (no `PAPERCLIP_MANAGED_CONFIG` set) are byte-for-byte unchanged — the parser only runs when the variable is present. - For managed instances, a malformed document now refuses startup by design (fail-closed). This is an intentional behavioral guarantee, not a regression: the control plane owns the variable and a precise startup error is the contract. - Overlay values are never persisted, so no migration or data-shape risk. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad74fb5450 |
Add a feature catalog build artifact derived from the experimental settings schema (#10055)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instances expose ~23 experimental feature settings, all declared in one shared zod schema and toggled per instance > - Deployment tooling and hosting control planes have no machine-readable list of those feature keys for a given release — the schema is only reachable from code that imports the package > - Any external system that references feature keys therefore does so as free text, and typos drift silently > - This pull request derives a versioned `feature-catalog.json` build artifact from the schema, with a compiler-checked metadata map so the schema stays the single source of truth > - The benefit is a stable contract external tooling can validate feature-key references against, with zero runtime behavior change ## Linked Issues or Issue Description No public issue exists; `feature_request` template fields: **Problem or motivation:** External deployment tooling cannot enumerate or validate an instance's feature keys per release; free-text references fail silently when keys are renamed or removed. **Proposed solution:** A metadata map keyed by the settings schema's own keys (compiler flags drift) plus a build step emitting `feature-catalog.json` (keys, tiers, defaults, `catalogVersion`) as a release artifact. **Alternatives considered:** A hand-maintained catalog file (drifts from the schema); serving the schema from a runtime API (requires a running instance at validation time — a build artifact works offline and pins to a release). **Roadmap alignment:** Supports the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed Adds a metadata map (title, description, tier, cloud/self-hosted defaults) keyed by the keys of `instanceExperimentalSettingsSchema`, so the schema stays the single source of truth and the compiler flags any drift. A new build step (`build:feature-catalog --version <v>`) emits `feature-catalog.json` — all 23 feature keys, their tiers, and a `catalogVersion` — as a release artifact that managed-hosting control planes can validate feature-flag writes against. No runtime behavior changes. - New `packages/shared/src/feature-catalog.ts`: per-flag metadata map keyed by a type derived from the settings schema (adding/removing/renaming a flag without updating the map is a compile error), plus `featureCatalogArtifactSchema` and `buildFeatureCatalogArtifact`/`renderFeatureCatalogArtifact` for the artifact - New `scripts/generate-feature-catalog.ts` wired as `pnpm build:feature-catalog --version <v>` - `scripts/create-github-release.sh` generates the artifact and uploads it as a GitHub Release asset (with a dry-run preview line) - Tests in `packages/shared/src/feature-catalog.test.ts` ## Verification - `vitest run packages/shared/src/feature-catalog.test.ts` — 9 tests: schema-key coverage, drift detection, artifact shape - `pnpm --filter @paperclipai/shared typecheck` - Artifact generation run end-to-end: `pnpm build:feature-catalog --version 0.0.0-test` emits 23 keys with `catalogVersion` ## Risks Low risk — no runtime behavior changes; the change is metadata, a build script, and a release-artifact emission step only. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1944c86153 |
fix(ci): preserve required e2e check for sharded runs (#9923)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The pull request workflow is the main merge gate for changes to that app. > - The Playwright e2e lane is expensive because every spec shares one isolated server and runs serially. > - Splitting that lane across runners shortens the critical path, but the public required-check contract still needs a check named exactly `e2e`. > - This pull request shards the real e2e work while preserving a fast aggregate `e2e` job for branch protection. > - The benefit is a faster PR workflow without making otherwise-good PRs unmergeable because a legacy required check disappeared. ## Linked Issues or Issue Description No public GitHub issue exists for this CI follow-up. Related prior CI work: - Refs #8360 - Refs #9168 - Refs #9516 Bug report: ### What happened? Sharding the PR e2e lane directly at the workflow job level changes the emitted check names to shard-specific names, while existing branch protection expects a check named exactly `e2e`. ### Expected behavior The PR workflow should be able to run e2e specs across multiple runners while still emitting a stable aggregate check named `e2e`. ### Steps to reproduce 1. Open a PR against `master`. 2. Run the PR workflow with the e2e lane split only as a matrix job. 3. Observe that the shard checks complete, but a required check named exactly `e2e` never appears. ### Paperclip version or commit Current `master`. ### Deployment mode GitHub Actions pull request workflow. ## What Changed - Added `scripts/e2e-shard.mjs`, which partitions default Playwright e2e specs by recorded per-spec duration. - Added `scripts/e2e-shard-durations.json` with measured e2e spec durations so the slow smoke-lab spec does not dominate one runner. - Split the PR workflow e2e lane into two `e2e_shards` matrix jobs and added a fast aggregate job named exactly `e2e`. - Added `scripts/__tests__/e2e-shard.test.mjs` to lock the shard partition, ignored-spec sync, manifest coverage, and aggregate required-check contract. ## Verification - `node --test scripts/__tests__/e2e-shard.test.mjs` - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check upstream/master..HEAD` - Searched GitHub for duplicate or related e2e-shard / required-check PRs and issues before opening this PR; no direct duplicate was found. ## Risks Low risk. The main risk is that the duration manifest can drift as specs are added or runtimes change; missing specs fall back to the median known duration, and the focused shard test catches empty, overlapping, or badly imbalanced partitions. ## Model Used OpenAI GPT-5 via Codex CLI coding agent, with shell/tool execution and repository inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2f42a4968d |
Treat cloud-managed instances as bootstrapped in the health gate (#9912)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip instances can be self-hosted, or provisioned and managed by a cloud control plane that authenticates users through trusted headers validated against `PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN` (`resolveCloudTenantActor`) > - In `authenticated` deployment mode, the health route reports `bootstrapStatus: bootstrap_pending` until at least one `instance_admin` exists, and the UI locks everyone out at the "waiting on its first admin" claim screen until then — correct for self-hosted instances, where a human operator must claim the instance > - But the cloud-tenant trust middleware, by deliberate security hardening, never grants `instance_admin` and actively purges legacy grants — so a cloud-managed instance can never leave `bootstrap_pending`: the gate demands a role the middleware forbids > - Every control-plane-provisioned instance is therefore permanently locked at the claim screen even though its users and memberships exist > - This pull request makes the gate cloud-aware: when the tenant server token is configured, the instance is considered bootstrapped, because the control plane owns identity and there is no operator claim step > - The benefit is that cloud-managed instances become usable while self-hosted behavior stays byte-for-byte identical, now pinned by a previously missing regression test ## Linked Issues or Issue Description Refs #2927 (introduced the browser-native first-admin bootstrap flow this gate feeds). No existing public issue for the deadlock; inline description per the bug report template: - **What happened?**: an instance configured with `PAPERCLIP_DEPLOYMENT_MODE=authenticated` and `PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN` reports `bootstrapStatus: bootstrap_pending` forever. All users — including ones created via the trusted-header path with owner-level company membership — are locked out at the "This Paperclip is waiting on its first admin" screen. - **Expected behavior**: a control-plane-managed instance has no first-admin claim step; users arriving with control-plane identity should reach the app. - **Steps to reproduce**: 1. Run the server with `PAPERCLIP_DEPLOYMENT_MODE=authenticated` and a `PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN` set 2. Create users only through trusted cloud headers (the middleware upserts them but never grants `instance_admin`, and purges any legacy grants) 3. `GET /api/health` → `bootstrapStatus` stays `bootstrap_pending`; the UI shows the claim screen for every visitor, and no supported path exists to create the `instance_admin` the gate requires - **Paperclip version or commit**: reproducible on `master` as of 2026-07-20; present since the cloud-tenant `instance_admin` purge hardening landed. ## What Changed - `server/src/middleware/auth.ts`: new exported `isCloudManagedInstance()` predicate beside the trust middleware that defines the tenant-token contract. - `server/src/routes/health.ts`: the authenticated-mode first-admin gate is skipped when the instance is cloud-managed; `bootstrapStatus` reports `ready`. - `server/src/__tests__/health.test.ts`: two new tests — authenticated without the token → `bootstrap_pending` (previously untested regression baseline), and with the token → `ready` despite zero instance admins. ## Verification - `pnpm vitest run src/__tests__/health.test.ts` in `server/` — 13/13 - `pnpm vitest run src/middleware/cloud-tenant-actor.test.ts` — 6/6 - Manual: with the env vars from the repro steps set, `GET /api/health` now returns `bootstrapStatus: "ready"`; without the token, behavior is unchanged ## Risks - None for self-hosted deployments: without the env var the gate is the prior behavior, now pinned by the new regression test. - For cloud-managed instances the claim screen and `bootstrapInviteActive` flow no longer appear — intended; browser-based claim was already disabled in that configuration. ## Model Used - Claude (Anthropic) — model id `claude-fable-5`, via the Claude Code CLI harness with tool use (shell, file edits, test execution). Diagnosis and change agent-assisted, human-directed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (n/a — behavior documented in code comments and pinned by tests) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fa4f900d1a |
ci: label Docker images with their bundled schema migration set (#9908)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip publishes Docker images from this repo that self-hosters and orchestration tooling deploy; the server refuses to start when its database is missing any schema migration the build bundles (`ensureMigrations`) > - Anything that deploys these images therefore needs to know an image's schema expectations *before* deploying it — today that requires pulling the image or checking out the matching commit, both heavyweight for tooling that just wants to answer "will this image boot against a database migrated to N?" > - Getting this wrong is expensive: an image ahead of the applied schema crash-loops at startup and fails healthchecks after deployment resources are already created > - This pull request labels every published image with its bundled migration set (last migration file and count), computed at build time from `packages/db/src/migrations` in the same tree the image is built from > - The benefit is image/schema compatibility verification with two cheap registry requests (manifest + config blob), no pull, and no drift risk between label and image contents ## Linked Issues or Issue Description No existing public issue; inline description per the feature request template: - **Subsystem affected**: Docker image publishing (`.github/workflows/docker.yml`), `packages/db` migrations - **Problem or motivation**: deployment tooling cannot cheaply determine which schema migrations a published image expects; the only options are pulling the image or checking out the matching commit. Deploying an image whose bundled migrations exceed the applied schema makes the server refuse to start, so this check is needed *before* resources are created. - **Proposed solution**: OCI labels (`io.github.paperclipai.schema.last-migration`, `io.github.paperclipai.schema.migration-count`), computed from the migrations directory at build time via the existing `docker/metadata-action` step. Keys use org-based reverse-DNS (the GitHub org) so the label contract survives product-domain migrations. - **Alternatives considered**: a schema manifest published beside the image (second artifact to keep in sync — rejected); encoding schema info in tags (tags already carry semver/sha meaning — rejected). - **Roadmap alignment**: checked `ROADMAP.md` — no overlap with planned core work; this is build metadata only. ## What Changed - `.github/workflows/docker.yml`: a `Compute schema migration labels` step (`ls` + `sort` over `packages/db/src/migrations/*.sql`) feeding two custom labels into the existing `docker/metadata-action` step. ## Verification - Workflow YAML validated locally. - The label-computation commands run against the current tree produce `last=0181_decision_training_retention_policy.sql`, `count=180`. - After merge, verify with: fetch the image config blob for a fresh `sha-*` tag from ghcr and confirm both `io.github.paperclipai.schema.*` labels are present. ## Risks - Low. Labels are metadata only; no change to image contents. If the migrations directory ever moves, the label step fails the workflow loudly (`ls` exits non-zero) rather than publishing wrong labels. ## Model Used - Claude (Anthropic) — model id `claude-fable-5`, via the Claude Code CLI harness with tool use (shell, file edits). Change authored and verified agent-assisted, human-directed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (YAML validation; label commands — no app code changed) - [x] I have added or updated tests where applicable (n/a — CI metadata only) - [x] I have updated relevant documentation to reflect my changes (n/a — workflow comment documents the labels) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ce7dedf33d |
perf(ci): balance general-server test shards by recorded suite duration (#9516)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its PR CI runs the general-server vitest lane pinned to `maxWorkers=1` and sharded across 3 runners (introduced in #8360) > - Suites were assigned to shards round-robin by sorted file index, so shard test time was unbalanced: a recent PR run split 73s / 153s / 115s, and the heaviest shard made "General tests (server 2/3)" the slowest check in the whole workflow at 314s wall > - The slowest shard sets the lane's wall time, so unbalanced partitions waste the other two runners and stretch the PR critical path > - This pull request replaces the round-robin assignment with a deterministic longest-processing-time partition weighted by a checked-in per-suite duration manifest > - The benefit is near-even shard weights (projected 113s / 113s / 113s with the current manifest), taking roughly 40s off the PR critical path with no reduction in coverage ## Linked Issues or Issue Description - Refs #8360 (introduced the 3-way general-server sharding this PR rebalances) - No public issue exists. Problem: the general-server test lane's round-robin shard assignment ignores per-suite duration, so one shard can carry multiple 30s+ suites while another finishes in half the time; the slowest shard alone determines the check's wall time. ## What Changed - `scripts/general-server-shard.mjs` (new): manifest loader and deterministic LPT (longest-processing-time) partitioner; suites missing from the manifest get the median recorded weight, and a missing or malformed manifest degrades to uniform weights so the lane never fails on stale data - `scripts/general-server-shard-durations.json` (new): per-suite duration manifest sampled from a real PR run (240 suites); the `$comment` field documents how to regenerate it - `scripts/run-vitest-stable.mjs`: both shard-selection sites (run and `--dry-run`) now use the balanced partition instead of index round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: 6 new tests covering skew-balance vs round-robin, determinism, median fallback for unlisted suites, malformed-manifest degradation, manifest coverage of the current suite set, and real-partition balance - `server/src/__tests__/heartbeat-issue-rewake-throttle.test.ts`: hardened the `afterEach` sweep — post-run bookkeeping (run-event records, follow-up wake scheduling) can still insert rows briefly after a run reaches a terminal status, and a late insert landing between the `agent_wakeup_requests` and `agents` deletes failed teardown with a foreign-key violation on the first CI attempt of this PR; the sweep now retries so a late background write cannot take down the shard - `release-verify.yml` shares the same runner script and inherits the balancing with no workflow change ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass (run against current master) - `npx vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6/6 pass against embedded Postgres with the hardened teardown - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass - `node scripts/run-vitest-stable.mjs --dry-run` with each shard flag shows every suite assigned exactly once across the 3 shards, with projected weights ~113s each ## Risks - Low risk: partition changes which runner executes which suite, not what runs; a completeness test asserts every suite is assigned to exactly one shard - The duration manifest will drift as suites are added/changed; unlisted suites get the median weight and a coverage test flags when the manifest covers less than half the suite set, so drift degrades balance gracefully rather than breaking the lane > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking enabled, agentic tool use (file edits, shell, test execution) via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude (Paperclip SWE) <noreply@paperclip.ing> |
||
|
|
b49d178c46 |
fix(ui): experiments auto-recovery dialog leaves UI dimmed and locked after enabling (#9513)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Operators tune instance behavior through Settings → Experiments,
where experimental features are toggled on and off
> - The task graph liveness auto-recovery experiment shows a
confirmation dialog (preview of what would be recovered) before it is
enabled
> - After confirming with "Enable only" or "Enable and run", the
dialog's Radix overlay and the `pointer-events: none` body lock were
left behind, dimming the page and blocking all interaction until a
refresh
> - The dialog was unconditionally mounted and only closed inside the
mutation's `onSuccess`, so the overlay teardown depended on the mutation
outcome and could race or never happen
> - This pull request closes the dialog before the mutation fires in
both confirm flows, clears the pending preview alongside the open flag,
and mounts the dialog conditionally so its overlay fully unmounts
> - The benefit is that enabling an experiment behaves like every other
settings change: the dialog goes away, the page stays interactive, and
errors surface in the page-level error banner instead of a dead UI
## Linked Issues or Issue Description
No public GitHub issue exists for this bug; description follows the bug
report template. Refs #4587 (the PR that introduced the configurable
liveness auto-recovery controls this dialog belongs to).
**What happened?** In Settings → Experiments, toggling on "Task graph
liveness auto-recovery" and confirming via "Enable only" left the whole
UI dimmed and unclickable. The dialog content disappeared, but the modal
overlay and the `pointer-events: none` lock on `<body>` remained until a
full page refresh.
**Expected behavior:** Confirming (or dismissing) the auto-recovery
dialog should close it completely and return the page to a fully
interactive state, with the toggle reflecting the new setting.
**Steps to reproduce:**
1. Open Settings → Experiments.
2. Toggle on "Task graph liveness auto-recovery"; the confirmation
dialog with the recovery preview appears.
3. Click "Enable only".
4. The dialog content disappears but the page stays dimmed and nothing
is clickable; refreshing restores the UI and shows the setting was
applied.
**Paperclip version or commit:** master @
|
||
|
|
49d1abc458 |
fix(ui): detail the env unsaved-changes banner and guard against draft loss (#9391)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instance settings include an Environments section where operators configure execution environments, each with an environment-variables editor for run-time bindings > - The editor showed a bare "Unsaved changes" banner that never said which variables changed, sometimes appeared the moment a saved config was opened (a lossy round-trip through the editor's emit rules made clean values look dirty), and the environment form let you navigate away without any confirmation, silently dropping the draft > - Operators could not tell what was unsaved, distrusted the phantom banner, and lost half-finished environment edits to a stray click — the agent configuration page already confirms before discarding, so environments behaved inconsistently > - This pull request lists the new/edited/removed variable names under the banner, normalizes both sides of the dirty comparison so saved values no longer look dirty on open, and confirms before cancel, in-app navigation, or tab unload while the form has unsaved changes > - The benefit is that the banner is trustworthy and specific, and unsaved environment edits can no longer be lost without an explicit confirmation ## Linked Issues or Issue Description Related (not fixed by this PR): #8930 introduced the current environment-variables editor and its unsaved-changes banner; #9386 moved environment create/edit from a modal to routed pages, which this PR's navigation guard builds on. No existing public issue for the defects themselves; described per the bug report template: **What happened?** The environment-variables editor in Environments settings showed a bare "Unsaved changes" banner with no indication of which variables changed. For some saved configurations (names with surrounding whitespace, incomplete secret references, duplicate names differing only by whitespace) the banner appeared immediately on opening the edit form, before any user input. Navigating away from the environment form — cancel, an in-app link, or closing the tab — silently discarded the draft with no confirmation. **Expected behavior** The banner should say which variables are new, edited, or removed; a freshly opened saved configuration should show no banner; and leaving the form with unsaved changes should require an explicit confirmation, consistent with the agent configuration page. **Steps to reproduce** 1. Open Settings → Instance settings → Environments and edit an environment whose saved config round-trips lossily (e.g. an env var name stored with trailing whitespace) — the "Unsaved changes" banner appears with no user edits. 2. Add or edit a variable — the banner gives no hint of what is unsaved. 3. With a dirty draft, click any in-app link or Cancel — the draft is dropped with no confirmation. **Deployment mode** Self-hosted (local development instance), reproducible on `master`. ## What Changed - The unsaved-changes banner in `EnvironmentVariablesEditor` now renders a change summary line — `New: … · Edited: … · Removed: …` — showing up to three names per group with a `+N more` overflow and the full list in a `title` tooltip. A rename shows as one addition plus one removal. - Dirty detection normalizes both the committed value and the draft through the same rules the editor uses when emitting values (trimmed names, incomplete secret refs dropped, last-writer-wins on trimmed duplicates), so a saved config that round-trips lossily no longer shows a phantom banner on first open. - The editor exposes an `onDirtyChange` callback and warns via `beforeunload` while its local draft is dirty. - The environment create/edit page (`CompanyEnvironments`) tracks a payload-level baseline fingerprint of the form as initialized and treats the page as having unsaved changes when the current form differs from it or the editor draft is dirty. While dirty it confirms ("Discard unsaved environment changes?") on Cancel, intercepts same-origin in-app link clicks, and warns on tab unload. ## Verification - `node_modules/.bin/vitest run ui/src/pages/CompanyEnvironments.test.tsx ui/src/components/environment-variables-editor/EnvironmentVariablesEditor.test.tsx` — 48 tests pass, including new coverage for: the change-summary banner text, no phantom banner for lossy round-trip values, beforeunload only while dirty, cancel confirmation on the edit page, and unload/link-click warnings after edits are staged into the form. - `tsc -b` in `ui/` passes. - Manual: edit an environment, add/edit/remove variables, observe the summary line; click Cancel or an in-app link and observe the confirmation; save and observe navigation proceeds without prompting. ## Risks - Low risk, UI-only. The click interceptor is scoped to same-origin anchor navigation while the environment form page has unsaved changes and is removed on cleanup; modified-key/middle-button clicks and external links are left alone. - The dirty-normalization intentionally ignores differences the editor could never persist (incomplete secret refs, untrimmed duplicate names); those were previously reported as unsaved changes that could not be saved away. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code, with extended thinking and tool use (code editing, test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9cde4e128c |
feat: run ACP sessions in sandbox execution targets (#9390)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters (Claude, Codex, Gemini) default to the ACP engine lane, which needs a live bidirectional stdio session with the agent process > - Sandbox execution targets only exposed one-shot command execution, so every ACP-capable adapter refused remote targets and fell back to the CLI lane with a "supports only the local Paperclip host" warning > - Running agents in sandboxes is a core deployment mode, and losing ACP there means losing streaming updates, structured events, and default-lane parity with local runs > - This pull request adds a provider-agnostic process-session bridge that relays the ACP stdio session into the sandbox over the existing sandbox runner contract, and updates the adapters to use it > - The benefit is that the default ACP lane now behaves the same on the local host and in any sandbox provider, with CLI fallback reserved for targets that genuinely cannot host a bidirectional session ## Linked Issues or Issue Description No existing public issue covers this; inline description following the feature request template: **Problem or motivation** Configuring an ACP-capable adapter (e.g. Claude) with a sandbox environment made every run fall back to the CLI lane with the warning "Claude ACP currently supports only the local Paperclip host, but this run targets a remote environment." The ACP engine only knew how to spawn a local subprocess, while sandbox providers only expose one-shot command execution — so there was no way to hold the bidirectional stdio session ACP requires. **Proposed solution** Add a process-session bridge in `adapter-utils`: a local ACPX-spawnable proxy script connects to a token-authenticated loopback TCP server, which relays JSON-framed stdin/stdout/stderr events to and from a small relay script executed inside the sandbox via the provider's ordinary runner. Claude/Codex/Gemini adapters now treat sandbox targets with a runner as ACP-capable, resolve agent commands against the remote target, and fall back to CLI only when the sandbox exposes no bidirectional path. The sandbox callback bridge injects a run-scoped API endpoint and bridge token so the agent inside the sandbox can reach Paperclip (including work-product handoffs) without ever receiving the host run JWT. **Alternatives considered** A provider-specific lane was prototyped first: Daytona minting SSH access metadata at lease time, converted into an SSH execution target. It was dropped because it only worked for providers able to advertise SSH, added per-provider surface area, and left every other sandbox provider on the CLI fallback. The merged design rides the one-shot runner contract all providers already implement; a regression test pins that sandbox targets stay on the bridge lane even when lease metadata advertises SSH access. **Roadmap alignment** Directly advances the "Cloud / Sandbox agents" roadmap item — agents running in remote and sandboxed environments keep the same control-plane behavior as local ones. No overlap with other planned core work. ## What Changed - `packages/adapter-utils/src/execution-target.ts`: new `startAdapterExecutionTargetProcessSessionBridge()` plus helpers — writes a token-authenticated local proxy script (spawnable by ACPX) and a remote relay script synced into the sandbox, with a loopback TCP server streaming JSON-framed stdio between them; events emitted before the ACP client attaches are buffered so none are lost. - `packages/adapter-utils/src/acpx-engine/execute.ts`: the ACP engine can execute against remote sandbox targets through the bridge instead of requiring a local subprocess, including remote cwd/env shaping. - `packages/adapter-utils/src/sandbox-callback-bridge.ts`: sandbox-scoped API bridging extended to allow work-product handoffs; the sandbox payload env carries a bridge token, never the host run JWT. - `packages/adapters/claude-local`, `codex-local`, `gemini-local` (`src/server/acp.ts`): default-lane selection no longer rejects all remote targets; command resolution is remote-aware (`ensureAdapterExecutionTargetCommandResolvable`, `resolveAdapterExecutionTargetCwd`); the fallback reason is now scoped to sandboxes that expose only one-shot execution. - `server/src/__tests__/environment-execution-target.test.ts`: pins that sandbox targets resolve to the bridge lane, including when lease metadata advertises SSH access. - Non-sandbox remote targets (e.g. SSH) keep the CLI lane: the ACP engine's remote transport is sandbox-only, so default-lane selection falls back for those targets across all three adapters, and tests covering CLI-specific remote behavior pin `engine: "cli"` explicitly. - The bridge authenticates loopback connections before they can own the session or receive buffered output (token required, idle unauthenticated peers dropped), and remote event writes are serialized so the exit event always lands after stdout/stderr have drained. - Daytona plugin: formatting-only residue from the earlier iteration; no functional change. ## Verification - `vitest run` over the touched suites — `packages/adapter-utils/src/acpx-engine/execute.test.ts`, `packages/adapter-utils/src/execution-target-sandbox.test.ts`, `packages/adapter-utils/src/sandbox-callback-bridge.test.ts`, the three adapter `acp.test.ts` files, and `server/src/__tests__/environment-execution-target.test.ts` — 102 tests pass. - End to end: with a Claude agent configured on a Daytona sandbox environment, the primary-model test now selects the default ACP lane (no fallback warning), and the full round trip (wake → sandbox execution → API bridge → comment post) was exercised twice from inside a live sandbox. ## Risks - Behavioral shift: adapters that previously always fell back to CLI on sandbox targets now default to ACP there; `engine=cli` still pins the CLI lane explicitly. - The bridge relays stdio as JSON lines over loopback TCP guarded by a per-session random token; the remote relay runs inside the sandbox under the provider's runner. Providers with slow one-shot execution will see higher session startup latency — the CLI fallback remains for genuinely incapable targets. - No schema or migration changes. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) — extended thinking enabled, agentic tool use via the Claude Agent SDK harness; implementation iterated with local Vitest verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no shipped docs describe the old local-only ACP limitation) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.local> |
||
|
|
0f9b1d399c |
fix(ui): make environment edit a routed page (#9386)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instance settings include an Environments section where operators configure sandbox/SSH/local execution environments, including interactive custom-image setup sessions with a browser terminal > - The environment create/edit form was rendered inside a modal dialog, so pressing Escape anywhere — including inside the embedded SSH terminal while capturing a snapshot — closed the whole modal and destroyed the in-progress session > - Environment editing is a heavyweight, long-lived flow; losing it to a reflexive Escape keypress is destructive and surprising > - This pull request converts environment create/edit from a modal into routed standalone pages, so Escape no longer dismisses the form > - The benefit is that terminal sessions and half-completed edits survive Escape, and the flow gets shareable URLs and normal back/forward navigation ## Linked Issues or Issue Description No existing public issue; described per the bug report template: **What happened?** While editing an environment's sandbox snapshot in the embedded SSH terminal, pressing Escape (e.g. to exit a mode inside the terminal) closed the entire environment edit modal, discarding the setup session and any unsaved form state. **Expected behavior** Escape inside the terminal or form should not dismiss the environment editor. A heavyweight flow like environment configuration should be a standalone page where Escape behaves as expected within the focused widget. **Steps to reproduce** 1. Open Instance settings → Environments and edit a sandbox environment 2. Start a custom image setup session and focus the browser terminal 3. Press Escape 4. The modal closes and the session context is lost ## What Changed - Converted the environment create/edit dialog in `CompanyEnvironments.tsx` into routed pages at `/company/settings/instance/environments/new` and `/company/settings/instance/environments/:environmentId/edit` - Registered the new routes in `App.tsx` and wired breadcrumbs for the list/create/edit states - Form state now initializes from the route (create vs edit) instead of dialog open/close state, and successful saves navigate back to the environments list - Updated `CompanyEnvironments.test.tsx` and `CompanySettings.test.tsx` to render through a router with the new routes and assert against the routed form page instead of a dialog ## Verification - `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx ui/src/pages/CompanySettings.test.tsx` — 22/22 passing - `tsc --noEmit` on the `ui` package — clean - Behavioral coverage: the updated tests exercise the routed create/edit pages end to end (open edit via the list, interact with the setup-session controls on the form page, save navigates back to the list); with the form no longer in a dialog there is no Escape-close handler to trigger ## Risks - Low risk; UI-only routing change. Deep links into the old modal state do not exist (the modal had no URL), so no redirects are needed - The edit page resolves the environment from the route param; a stale/unknown id falls back to the environments list ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Fable 5), extended thinking enabled, agentic tool use via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Cody <cody@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ac66fd65cb |
Fix Cody default model adapter test config (#9365)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent configuration includes adapter-specific model settings and a
built-in adapter test action so operators can verify runtime
configuration before saving changes.
> - Cody/Codex-style local adapters can use an adapter default model
when the user clears the explicit model field.
> - The adapter test path still passed an object containing `model:
undefined` in some create/edit flows, which is different from omitting
the model and can break default-model behavior.
> - The previous fix was reverted because it also included an unrelated
skill documentation edit.
> - This pull request reapplies only the UI default-model test-config
fix, with no doc or skill changes.
> - The benefit is that testing Cody/Codex adapter settings with the
default model follows the same contract as saving default model
settings: no explicit model key is sent.
## Linked Issues or Issue Description
Bug report:
- Summary: Testing a Cody/Codex local agent after selecting the default
model could send an adapter config with an undefined model value instead
of omitting the model key.
- Expected behavior: Clearing the model to use the adapter default
should test with `adapterConfig: {}` unless another model is explicitly
selected.
- Actual behavior: The UI test-config path could preserve `model:
undefined`, causing the adapter test to fail instead of exercising the
default model.
- Related PRs: Reapplies the UI-only portion of #9361 after #9363
reverted the original PR.
## What Changed
- Exported and reused `omitUndefinedEntries` so adapter test config
payloads drop undefined adapter config entries before calling the test
endpoint.
- Hardened the current model display value so create-mode values that
are nullish or non-string do not crash the model selector/test flow.
- Added render coverage for editing a Codex agent back to the default
model and for testing a create form with the default model.
## Verification
- `pnpm exec vitest run
ui/src/components/AgentConfigForm.render.test.tsx`
- `pnpm check:token-gates`
- Confirmed `git diff origin/master --name-only` contains only:
- `ui/src/components/AgentConfigForm.render.test.tsx`
- `ui/src/components/AgentConfigForm.tsx`
- `ui/src/lib/agent-config-patch.ts`
## Risks
Low risk. The change only removes `undefined` adapter config entries
from the UI adapter-test payload and adds focused render coverage.
Explicit model values and other adapter config fields are preserved.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex, GPT-5 coding agent, tool-use enabled. Context window size
not exposed in this runtime.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
d1f6a6850a |
Fix agent detail URL after agent rename (#9340)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The agent detail page uses route references that can be based on an agent's URL key. > - Renaming an agent can change that URL key while the browser is still on the old route. > - After save or rollback, refetching the stale route reference can render an "Agent not found" state even though the agent still exists. > - This pull request redirects the detail page to the updated canonical route when the saved agent's route reference changes. > - The benefit is that agent renames keep users on the same configuration workflow without landing on a stale URL. ## Linked Issues or Issue Description - Refs #1848 - Related public search performed for agent rename/not-found issues and PRs; no closer in-flight PR was found. - Bug context: after saving a renamed agent or rolling back to a revision with a different name-derived URL key, the agent detail page could continue using the old URL and show "Agent not found". ## What Changed - Added a small route-sync helper that compares the previous and updated agent route refs after mutations. - Redirects the agent detail page with `replace: true` when a save or rollback changes the canonical route ref. - Removes the stale detail-query cache entry so the old route reference is not refetched after a rename. ## Verification - Local outgoing patch scan for common secrets, private paths/emails, and internal issue/link references: no matches. - `corepack pnpm install --frozen-lockfile` - `corepack pnpm --dir ui run typecheck` - `corepack pnpm --dir ui exec vitest run src/pages/AgentDetail.progress.test.ts src/App.test.tsx` - `corepack pnpm check:token-gates` ## Risks - Low risk: the redirect only runs when the updated agent resolves to a different route ref than the current agent. - If a future mutation response omits both URL key and name, the existing route-ref fallback behavior still applies. ## Model Used OpenAI Codex, GPT-5 coding agent via the local Codex adapter, with tool-assisted repository inspection, shell execution, and GitHub API use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
17dde9d3f2 |
fix(sandbox): keep custom-image snapshots applied to config tests, probes, and saves (#9385)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox environments can capture reusable custom images (provider snapshots) so agents boot with pre-installed tools and CLI logins > - The custom-image runtime fingerprint check included provider secret-ref paths (e.g. the Daytona `apiKey`), while capture-time fingerprinting excluded them, so any config carrying a credential never matched its captured snapshot > - As a result, agent config tests and environment probes silently booted the provider base image instead of the snapshot, test sandboxes were deleted before operators could inspect them, and any environment save orphaned the snapshot without warning > - The UI compounded the confusion by displaying an internal template id that matches nothing in the provider dashboard > - This pull request aligns runtime fingerprints with capture-time exclusions, re-stamps fingerprints on saves that cannot affect the snapshot (warning when they can), archives test/probe sandboxes instead of deleting them, and surfaces the provider snapshot ref in the UI > - The benefit is that custom images actually apply to config tests and probes, survive unrelated config edits, and are debuggable against the provider dashboard ## Linked Issues or Issue Description No public GitHub issue exists for this; describing it in-PR per the bug template. Related: Refs #9329 (saved-environment probe company context — this branch carries an equivalent fix), Refs #8794 (introduced reusable sandbox custom images). **What happened?** With a Daytona environment whose provider config stores the API key as a secret reference and an active captured custom-image snapshot: - Agent config tests and environment probes booted the provider base image (`daytonaio/sandbox:0.8.0`) instead of the captured snapshot, so CLI upgrades/logins baked into the snapshot were missing and the probe reported "login required" and an outdated CLI. - The environment card showed an internal template id (e.g. `b5be03e1-ca5…`) that does not correspond to any snapshot name in the provider dashboard, making the active image impossible to correlate. - Test/probe sandboxes were deleted immediately after the run, so the sandbox a test used could not be inspected afterwards. - Saving the environment config (even fields unrelated to the image) changed the stored fingerprint, silently detaching the snapshot with no warning. **Expected behavior** Config tests and probes boot the captured snapshot when one is active; the UI shows the provider-facing snapshot/template ref; test sandboxes stay inspectable for a short window; unrelated config edits keep the snapshot linked, and edits that genuinely invalidate it produce an explicit warning. **Steps to reproduce** 1. Configure a sandbox environment on Daytona with the API key stored as a company secret reference. 2. Capture a custom image snapshot from the environment page and mark it active (e.g. after installing/logging into a CLI in the setup sandbox). 3. Run the agent config test or an environment probe: the sandbox boots the base image, not the snapshot, and the sandbox is deleted immediately after the test. 4. Save the environment config with an unrelated field change: the snapshot silently stops applying. **Paperclip version or commit** `master` at the merge-base of this branch. **Deployment mode** Self-hosted local instance (macOS, pnpm dev server) with the Daytona sandbox provider plugin. ## What Changed - Runtime custom-image fingerprint checks now exclude provider secret-ref paths, matching capture-time exclusions, so configs carrying credentials match their captured snapshots (`environment-custom-image-runtime.ts`). - Agent config tests and saved-environment probes force fresh, non-reused sandboxes and pass company context so lease-backed probes can resolve company secrets and boot the real snapshot (`environment-probe.ts`, `routes/agents.ts`, `routes/environments.ts`). - Test/probe sandboxes are released by archiving (stop + 60-minute provider-side auto-delete) instead of immediate deletion, so operators can inspect the exact sandbox a test used (Daytona plugin). - On environment PATCH save, changes that cannot affect the captured snapshot re-stamp the template's source fingerprint so the snapshot stays linked; boot-source or provider-identity changes (new manifest field `templateIdentityPaths`) mark the template detached and the save response reports it (`environment-custom-images.ts`, shared plugin types/validators). - The custom-image overview exposes `activeTemplateMatchesConfig`; the environments UI shows the provider snapshot/template ref (internal id moved to a tooltip), warns via toast when a save detaches the snapshot, and shows a persistent "Not in use" warning when the active template no longer matches the saved config (`CompanyEnvironments.tsx`, `api/environments.ts`). ## Verification - `pnpm vitest run server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/environment-probe.test.ts server/src/__tests__/environment-routes.test.ts server/src/__tests__/agent-test-environment-routes.test.ts` — server coverage for fingerprint exclusions, re-stamp/detach on save, probe company context, and fresh-sandbox test behavior. - `pnpm vitest run packages/plugins/sandbox-providers/daytona/src/plugin.test.ts` — archive-on-release and snapshot ref handling. - `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx` — snapshot ref display, detach toast, and "Not in use" warning. - Manually verified end-to-end on a live self-hosted instance against real Daytona: config test boots the captured snapshot (CLI login and version persist), the test sandbox remains visible in the provider dashboard as archived, and saving unrelated fields keeps the snapshot applied. ## Risks - Fingerprint exclusion widening: a provider credential rotation alone no longer detaches a captured snapshot; that is the intended behavior (the snapshot content does not depend on the credential), and provider-identity fields (e.g. Daytona `apiUrl`) still detach via `templateIdentityPaths`. - Archived test sandboxes consume provider-side resources for up to their auto-delete window instead of being freed immediately; bounded (60 minutes) and only for test/probe sandboxes. - New optional manifest field `templateIdentityPaths` is backward-compatible; providers that omit it keep current matching behavior. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
e84731af70 |
Revert "Fix default model adapter test config" (#9363)
Reverts paperclipai/paperclip#9361 |
||
|
|
ebd62ca5ae |
Fix default model adapter test config (#9361)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board UI lets operators create and edit agent adapter configuration, including a primary model field and an adapter test action. > - For Cody and similar adapter forms, selecting the default model means the model value is intentionally unset so the adapter can use its default. > - The adapter test path still allowed `model: undefined` to survive in the generated adapter config, which could send an invalid test payload instead of omitting the field. > - This pull request normalizes create/edit adapter test config so default-model selections omit `model` entirely. > - The benefit is that testing an agent configured to use the adapter default model exercises the same clean config shape that should be saved and run. ## Linked Issues or Issue Description No public GitHub issue was found for this local UI bug, so the problem is described inline. Bug description: - What happened: using the adapter test action after choosing the default model could include `model: undefined` in adapter config and surface a UI/runtime error instead of testing with the adapter default. - Expected behavior: choosing the default model should omit the `model` field from adapter config so the adapter default is used. - Steps to reproduce: edit a Codex/Cody-style agent with a concrete model, switch the model selector to Default, then run the adapter Test action. - Paperclip version/commit: current `master` before this PR. - Deployment mode: board UI, deployment-mode independent. ## What Changed - Sanitized adapter test config assembly so undefined adapter config entries are omitted before the test request is sent. - Made create-mode current model display resilient when the model is unset for adapter defaults. - Added regression coverage for editing an existing agent from a concrete model back to Default and testing it. - Added regression coverage for create-mode testing with an unset/default model. - Hardened the developer skill wording used by the existing server skill utility contract test. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/AgentConfigForm.render.test.tsx` - `pnpm exec vitest run server/src/__tests__/paperclip-skill-utils.test.ts` - GitHub PR checks on this branch are green, including Typecheck + Release Registry, Build, General tests, e2e, verify, security scans, and Greptile Review. ## Risks - Low risk: this only removes undefined values from adapter test config payloads, which aligns with the existing persisted patch behavior. - Low risk: default-model display now treats unset create-mode model values as an empty string. - No database, API schema, migration, auth, or adapter runtime contract changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected - check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-enabled software-engineering session. Exact context window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <cody@paperclip.ing> |
||
|
|
a4993a72a6 |
Fix live run streaming text readability (#9330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue thread UI renders live agent output from adapter run logs and transcript parsing. > - Some adapter streams emit many small or repeated token chunks, and live UI updates can expose partial words, duplicated slices, or transient markdown placeholders. > - That makes active run updates look like gibberish even when the underlying agent output is valid. > - The fix needs to preserve raw logs while making the live thread view stable, readable, and ordered. > - This pull request adds monotonic run-log sequencing, safer live transcript dedupe/order handling, markdown placeholder hiding, and readable live text stabilization. > - The benefit is a live issue thread that updates smoothly without showing confusing partial parser artifacts. ## Linked Issues or Issue Description No public GitHub issue exists yet, so this PR includes the bug details inline. ### What happened? Live run updates in the issue thread can show confusing repeated or partial text while an adapter is streaming. The visible text appears to lose parsing boundaries during active updates, especially with ACP-style token deltas, so the live output can briefly render duplicated chunks, incomplete words, or HTML-comment placeholders. ### Expected behavior Live text should remain readable while preserving the underlying run output for raw inspection. ### Steps to reproduce 1. Start a live agent run whose adapter emits small stdout token deltas. 2. Watch the issue thread while the run is still active. 3. Observe transient duplicated chunks, incomplete words, or markdown placeholder artifacts in the live rendered text. ### Paperclip version or commit Reproduced against current `master` before this PR branch. ### Deployment mode Local dev issue-thread UI with live local adapter runs. ### Additional context GitHub PR search for `live run streaming text markdown transcript` found one broad merged PR, `#252` (“Dotta updates - sorry it's so large”), but no targeted duplicate for this live streaming readability bug. ## What Changed - Added per-run monotonic sequence numbers to persisted and live run-log chunks. - Dedupe and order live transcript chunks by sequence before falling back to timestamp ordering. - Hide markdown HTML comment placeholder text from rendered markdown output. - Smooth live issue-thread text updates so partial additions reveal at readable word boundaries and sliding-window removals do not produce gibberish. - Added coverage for run-log ordering/deduping, markdown comment hiding, live issue-thread stabilization, and Greptile-reviewed edge cases where overlap rewrites could synthesize text or no-boundary additions could stay hidden. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts src/components/MarkdownBody.test.tsx src/components/transcript/useLiveRunTranscripts.test.tsx` passed before the review fix: 3 files, 86 tests. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts` passed after the review fix: 1 file, 30 tests. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts src/components/transcript/useLiveRunTranscripts.test.tsx` passed after the final Greptile overlap fix: 2 files, 40 tests. - `pnpm check:token-gates` passed. - Local PII/secret scan of touched files found only expected code/test words such as `secret`, `token`, and redaction-related strings; no literal credentials found. - `pnpm -r typecheck` passed after restoring declared dependencies with `CI=1 pnpm install --frozen-lockfile` and running with a short `TMPDIR` because `tsx` IPC sockets fail under the long sandbox temp path. - `pnpm build` passed with existing Vite CSS/font/chunk warnings. - GitHub PR checks passed on head `4c052dfe86aecb5feb73504e6b48843f68fce813`: build, typecheck/release registry, server and workspace test shards, serialized server suites, e2e, canary dry run, policy, review, Socket, Superagent, Snyk, and verify. - Greptile review passed on head `4c052dfe86aecb5feb73504e6b48843f68fce813` with confidence score 5/5 and no blocking issues found. - `pnpm test:run` failed in unrelated server workspace tests on this macOS local environment: - `server/src/__tests__/heartbeat-workspace-branch-containment.test.ts`: two assertions compare `/tmp/...` with `/private/tmp/...`. - `server/src/__tests__/heartbeat-worktree-suppression.test.ts`: expected one heartbeat run but observed two, followed by cleanup fallout in the full run. - Isolated rerun of those two server suites reproduced the same three failures. ## Risks - Low product risk for the UI changes: the readable smoothing only affects active live-run display stabilization, not stored comments or raw run logs. - Moderate verification risk: local full Vitest did not pass because of unrelated server workspace tests. Targeted tests for this change, typecheck, token gates, and build passed. - Run-log sequence fields are optional for compatibility with older log rows that do not include `seq`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-using local workspace execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
05973b2073 |
Enable sandbox environments for Grok local adapter (#9338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local CLI adapters can run against Paperclip-managed execution environments instead of only the host filesystem. > - The environment picker and environment capability API derive sandbox support from shared adapter capability lists. > - The Grok Build adapter is implemented as a local CLI adapter, but it was missing from those shared environment capability lists. > - That made Grok agents look local-only even when sandbox environments were configured. > - This pull request registers `grok_local` in the shared adapter constants and remote-managed environment support path. > - The benefit is that Grok Build agents can select the same local, SSH, and sandbox environment overrides as other local CLI adapters. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Inline bug report follows. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on `master`. - [x] I have confirmed the error originates in Paperclip itself, not in the Grok adapter provider or local configuration. ### What happened? When configuring a Grok Build local agent in the board UI, Paperclip did not expose configured sandbox environments as selectable environment overrides. The shared environment capability helper treated `grok_local` as local-only because it was missing from the remote-managed local adapter allowlist. ### Expected behavior Grok Build should behave like other local CLI adapters: when environments are enabled and a runnable sandbox environment exists, the agent configuration form should show the environment override selector and allow the sandbox to be selected. ### Steps to reproduce 1. Enable environments in instance experimental settings. 2. Configure at least one runnable sandbox environment. 3. Open the agent configuration form for a Grok Build local agent. 4. Observe that the sandbox environment is not offered as an override before this fix. ### Paperclip version or commit `master` before this PR. ### Deployment mode Local dev (`pnpm dev`). ### Installation method Built from source (`pnpm dev` / `pnpm build`). ### Agent adapter(s) involved - Grok Build local adapter. - Core bug in shared environment capability logic. ### Database mode Not database-related. ### Access context Board human operator. ### Relevant logs or output No runtime error is emitted; the issue is a missing UI option caused by shared capability metadata. ### Relevant config No secret-bearing config required. Reproduction only needs environments enabled and a runnable sandbox environment configured. ### Privacy checklist - [x] I have reviewed all pasted output for PII, usernames, file paths, API keys, tokens, company names, and redacted where necessary. ## What Changed - Added `grok_local` to the shared built-in adapter type list. - Added `grok_local` to the remote-managed adapter set used by environment capability helpers. - Added shared regression coverage for Grok local sandbox provider and driver support. - Added a UI render regression test that confirms Grok Build agents show the environment override when a runnable sandbox exists. ## Verification - `git diff --check` - Changed-file secret scan with `rg` for common token/key patterns. - `pnpm --filter @paperclipai/shared exec vitest run src/environment-support.test.ts` - `pnpm --dir ui exec vitest run src/components/AgentConfigForm.render.test.tsx` ## Risks - Low risk. This expands environment support for an existing local adapter to match the local CLI adapter behavior already used by Claude, Codex, Gemini, OpenCode, Cursor, and Pi. - Operators still need at least one configured runnable sandbox environment before a Grok agent has a sandbox option to select. - This PR was created from a Paperclip execution workspace branch whose name is runtime-provided; the PR body intentionally avoids internal issue identifiers or instance-local links. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 coding agent, tool-enabled terminal/code execution environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bc85b456a1 |
fix(ui): keep agent names visible on mobile agents index (#9236)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Agents index (`ui/src/pages/Agents.tsx`) is the main roster view, with both a list layout and an org-tree layout > - On mobile viewports the roster could render rows with no visible agent names because shrink-resistant trailing controls competed with the fixed-width title cell > - A roster where names are unreadable is unusable on phones, and the org-tree view had a related narrow-width overflow risk for deeply indented names > - This pull request lets the list title flex below `xl`, hides nonessential row controls on mobile, keeps Join/Leave reachable, and truncates org-tree names safely > - The result is a readable agents index on mobile while preserving desktop meta-column alignment and truncation ## Linked Issues or Issue Description No existing public GitHub issue; describing the bug per the bug report template: **What happened?** On a mobile-width viewport, the agents index showed rows where agent names could become unreadable or disappear because trailing controls consumed the available row width. In the org view, deeply indented names could overflow the row. **Expected behavior** Agent names remain visible on every viewport, Join/Leave stays reachable, and desktop rows continue to align and truncate as before. **Steps to reproduce** Open Paperclip in a browser at a narrow viewport, navigate to the Agents index, and observe rows with long names or left-membership controls. **Paperclip version or commit** Reproducible on `master` prior to this fix. **Deployment mode** Local dev instance. The bug is viewport-width dependent, not deployment-mode dependent. ## What Changed - `EntityRow` now lets callers control title text, subtitle text, and the meta spacer classes while keeping the default truncation behavior unchanged. - Agents list rows now use `flex-1 xl:flex-none xl:w-56`, so names get mobile width while desktop meta columns keep their aligned fixed title column. - Long list-view names/subtitles wrap below `xl` but return to truncation at `xl` and above. - Join/Leave stays visible on mobile; run/status/star row controls stay hidden on mobile to avoid squeezing names. - Org-tree rows use `min-w-0 truncate`, so deep indentation shortens names with an ellipsis instead of overflowing. - Regression tests cover mobile name visibility, left-membership dimming, responsive title/meta behavior, and mobile Join/Leave reachability. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/EntityRow.test.tsx src/pages/Agents.test.tsx` — 2 files, 18 tests passed. - `pnpm check:token-gates` — all gates clean. - `git diff --cached | rg -n "(AKIA|ASIA|SECRET|TOKEN|PASSWORD|PRIVATE KEY|BEGIN RSA|BEGIN OPENSSH|api[_-]?key|bearer|paperclip_api_key|DATABASE_URL|postgres://|sk-[A-Za-z0-9]|xox[baprs]-)"` — no matches before push. ## Risks - Low risk: UI-only change scoped to the agents index. The main behavior change is that nonessential row controls remain hidden on mobile while Join/Leave remains available there and all controls remain available on larger screens/detail pages. ## Model Used - OpenAI GPT-5 via Codex (`codex_local` adapter), agentic code editing and local/GitHub verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eedc7ddef2 |
Make ACP the default engine for local adapters (#9238)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter packages are the bridge between the control plane and local agent harnesses such as Claude Code, Codex, and Gemini CLI. > - ACP support was concentrated in a separate `acpx_local` adapter, which made ACP feel like a separate agent choice instead of an execution capability of the harness adapters. > - Claude, Codex, and Gemini now have ACP-capable harnesses, so the native adapter should own ACP selection, fallback, config, transcript parsing, and environment diagnostics. > - The standalone ACPX adapter still needs a compatibility path for existing rows, but it should not be offered as an active adapter for new agents. > - This pull request moves the shared ACP runtime into `@paperclipai/acpx-engine`, wires Claude/Codex/Gemini local adapters to prefer ACP when prerequisites are available, and retires `acpx_local` to a tombstone. > - The benefit is one adapter per harness, richer ACP transcripts by default where possible, and a migration path for existing Claude/Codex ACPX agents. ## Linked Issues or Issue Description Closes #5932 — the broken default `acpx_local` Claude path is replaced by native `claude_local` ACP support, existing Claude/Codex ACPX rows migrate to native adapters, and new agents no longer choose the standalone ACPX adapter. Refs #4893 — original merged ACPX local adapter runtime that this PR replaces with native per-harness ACP engines. Refs #6590 — prior ACPX-Claude seamlessness work folded into the new native Claude ACP path. Refs #197 — related open generic ACP/Kiro adapter work; this PR does not close it because Kiro/custom generic ACP remains a separate adapter decision. Refs #7018 — related Kimi-specific `acpx_local` shell failure; this PR retires the built-in standalone adapter but does not add a native Kimi adapter. Refs #8864 — related ACPX prompt/API guidance PR; this PR moves runtime guidance into the shared/native ACP engine path instead of the old standalone adapter. Refs #8881 — related `acpx_local` POSIX shell failure from the old `acpx` pin; this PR updates ACP dependencies but does not claim custom/OMP ACP support as a first-class native adapter. Refs #8964 — related open `acpx_local` stderr cleanup PR; this PR makes the old runtime path obsolete for new agents but keeps it as a non-closing reference. Problem description: - The standalone `acpx_local` adapter duplicates Claude/Codex agent choices that already have first-class local adapters. - ACP should be an execution engine capability of each harness adapter when the underlying harness supports ACP. - Existing `acpx_local` agents should either migrate to native harness adapters or fail with an explicit retirement message instead of silently falling back to the process adapter. ## What Changed - Added `@paperclipai/acpx-engine` as the shared ACP execution, session-codec, CLI formatter, and UI parser package. - Wired `claude_local`, `codex_local`, and `gemini_local` to auto-select ACP by default when prerequisites pass, with `engine=cli` opt-out and `engine=acp` strict mode. - Added ACP config schema/UI fields, environment checks, session-codec preservation, transcript parsing, and adapter capability metadata for the native adapters. - Retired `acpx_local` to a server tombstone, removed its UI/package/runtime image surface, and added a migration for existing Claude/Codex ACPX agents. - Updated package manifests, lockfile, release tooling, docs, Kubernetes sandbox defaults, and tests. ## Verification - `corepack pnpm --filter @paperclipai/acpx-engine typecheck` - `corepack pnpm --filter @paperclipai/adapter-claude-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-codex-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `corepack pnpm --filter @paperclipai/acpx-engine exec vitest run` - `corepack pnpm --filter @paperclipai/adapter-claude-local exec vitest run src/server/acp.test.ts src/server/execute.acp-fallback.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-codex-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-gemini-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts src/ui/parse-stdout.test.ts` - `corepack pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && corepack pnpm --filter @paperclipai/server exec tsc --noEmit` - `corepack pnpm --filter @paperclipai/server exec vitest run src/__tests__/adapter-routes.test.ts src/__tests__/adapter-session-codecs.test.ts src/__tests__/adapter-models.test.ts` - `corepack pnpm --filter @paperclipai/ui typecheck` - `corepack pnpm --filter @paperclipai/ui exec vitest run src/adapters/metadata.test.ts src/adapters/adapter-display-registry.test.ts src/components/AgentConfigForm.test.ts src/components/AgentConfigForm.render.test.tsx src/components/transcript/RunTranscriptView.test.tsx` - `node --test scripts/bootstrap-npm-package.test.mjs scripts/release-package-map.test.mjs scripts/verify-release-registry-state.test.mjs` Note: the server typecheck script calls `pnpm` internally; this dev shell exposes pnpm through Corepack only, so I ran the two script steps manually with `corepack pnpm`. ## Risks - Migration changes existing `acpx_local` Claude/Codex agents to native adapter types and clears old ACPX task sessions/runtime state. - Custom ACP commands remain on the retired tombstone and will need a separate future adapter/plugin path. - ACP auto-selection depends on local Node and ACP server command prerequisites; remote and unsupported environments fall back to CLI unless `engine=acp` is explicit. - `@paperclipai/acpx-engine` is a new public package and needs npm trusted-publishing bootstrap before release automation can publish it. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent. Exact hosted model build and context-window size are not exposed in this runtime. Tool use included shell execution, repository editing, GitHub CLI operations, and local test/typecheck execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cc17f29e7e |
fix(timeouts): raise sandbox wall-clock backstop to 4h and make acpx_local timeouts self-describing (#9232)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runs execute through adapters (e.g. `acpx_local`), which can run locally, over SSH, or inside sandbox execution targets, each with a wall-clock execution timeout > - Sandbox-backed runs defaulted to a 30-minute wall-clock backstop (`DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC = 1800`), which kills healthy long agent runs that are still making progress — long before the recovery watchdog's 4h critical threshold would even consider them stuck > - On top of that, `acpx_local` resolved its timeout directly from `adapterConfig.timeoutSec` instead of the shared execution-target resolver, and its timeout failures surfaced as a bare `Timed out after Ns` — giving operators no clue which timer fired or which knob raises it > - This pull request raises the sandbox backstop to 4h (aligned with the recovery watchdog), routes `acpx_local` through the shared timeout resolver, logs the effective timeout and its source at run start, and makes every timeout error message self-describing > - The benefit is that long-running sandbox agent runs no longer die at 30 minutes, and when a wall-clock timeout does fire, the run log states exactly which timer fired and how to configure it ## Linked Issues or Issue Description Refs #4535 (related: wall-clock execution timeouts killing agent runs that are still making progress — that issue covers a different hardcoded 600s timer, but the operator pain is the same). No exact public issue exists for this one, so describing it in-PR: **Bug:** A long sandbox-backed `acpx_local` agent run was killed with a bare `Timed out after 1800s` even though the agent was actively working. - **What happened:** The run hit the 30-minute `DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC` backstop. `acpx_local` never consulted the shared execution-target timeout resolution (it read `adapterConfig.timeoutSec` directly, default 0), so on sandbox targets the sandbox-provider default applied with no adapter-level say. The resulting error named neither the timer that fired nor the knob that controls it. - **Expected:** Healthy long runs should not be killed by a 30-minute wall-clock backstop when the recovery watchdog only treats runs as critically stuck after 4h of output silence; and any timeout error should say which timeout fired and how to raise it. - **Impact:** Long, legitimate agent runs in sandboxes fail mid-work; operators waste time reverse-engineering which of several timers produced "Timed out after Ns". ## What Changed - `packages/adapter-utils/src/execution-target.ts` - `DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC` raised from `1_800` to `14_400` (4h), with a comment explaining it intentionally matches the recovery watchdog's `ACTIVE_RUN_OUTPUT_CRITICAL_THRESHOLD_MS` (4h) so the adapter backstop never fires before the watchdog path. Output-inactivity monitors remain the primary hang detectors. - New `resolveAdapterExecutionTargetTimeout(target, configuredTimeoutSec)` returns `{ timeoutSec, source }` where `source` is `configured` / `sandbox_default` / `unlimited`. The existing `resolveAdapterExecutionTargetTimeoutSec` is preserved as a thin wrapper, so current callers are unaffected. - New `formatAdapterExecutionTimeoutErrorMessage(resolution)` and `formatAdapterExecutionTimeoutStartLogLine(resolution)` produce self-describing messages that name the timer that fired and the `adapterConfig.timeoutSec` knob that controls it. - `packages/adapters/acpx-local/src/server/execute.ts` - `buildRuntime` now resolves the wall-clock timeout through the shared resolver: sandbox targets default to the 4h backstop, local/SSH keep the historical "0 = no adapter timeout", and a configured `adapterConfig.timeoutSec` always wins. - The executor logs the effective timeout and its source at run start (`[paperclip] Adapter execution timeout: …`), so a later timeout is diagnosable from the run log alone. - All three bare timeout messages (timer cancel reason, turn result `errorMessage`, catch-path `messageOverride`) now use the self-describing format. - `packages/adapters/acpx-local/src/index.ts` — the adapter configuration doc for `timeoutSec` states the sandbox default and that the output-inactivity monitor remains the primary hang detector. - Tests: `packages/adapter-utils/src/execution-target-sandbox.test.ts` and `packages/adapters/acpx-local/src/server/execute.test.ts` (see Verification). ## Verification - `pnpm --filter @paperclipai/adapter-utils typecheck` — passes - `pnpm --filter @paperclipai/adapter-acpx-local typecheck` — passes - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapters/acpx-local/src/server/execute.test.ts` — 2 files, 40 tests, all pass - New/updated test coverage: - sandbox default resolves to 4h (and the constant is asserted to be `4 * 60 * 60`) - `resolveAdapterExecutionTargetTimeout` reports `configured` / `sandbox_default` / `unlimited` sources with the correct precedence (configured > sandbox default; local/SSH stay unlimited) - exact wording of the self-describing error message and the start-of-run log line - `acpx_local` runtime picks up the sandbox default into `timeoutMs`, keeps the unlimited local default, honors configured-over-default precedence, emits the start-of-run log line, and surfaces the self-describing `errorMessage`/cancel reason when the wall-clock timer kills a turn ## Risks - **Behavioral shift:** sandbox-backed adapter runs that previously hit the 30-minute backstop now run up to 4h before the adapter kills them. Genuinely hung runs are still caught much earlier by the adapters' output-inactivity monitors and by the recovery watchdog; the wall-clock timer is a last-resort kill switch. Operators who relied on the 30-minute default can restore it explicitly via `adapterConfig.timeoutSec`. - **Error-message consumers:** any tooling that pattern-matched the exact `Timed out after Ns` string from `acpx_local` will see the new self-describing message instead. - No API or schema changes; `resolveAdapterExecutionTargetTimeoutSec` keeps its exact signature and behavior (modulo the raised sandbox default). > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic) via Claude Code CLI — model ID `claude-fable-5`, extended thinking enabled, agentic tool use (file edits, shell, test execution) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e68ee09809 |
perf(release): batch npm registry version queries (#9202)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The release subsystem publishes the public workspace packages and also powers release-related CI validation. > - The release flow currently asks npm for package versions one package at a time in multiple places. > - That serial registry latency slows the PR Canary Dry Run path and real release invocations even though the checks are independent. > - This pull request batches npm registry version lookups with bounded concurrency and reuses the result for version calculation. > - The benefit is shorter non-build release-script time while preserving the fresh target-version existence check before publishing. ## Linked Issues or Issue Description - No public GitHub issue exists for this release-script performance cleanup. ### Problem or motivation Release validation spends avoidable time on repeated serial `npm view` calls across the public package set. The slow path affects PR release validation and real release invocations because version discovery waits on independent registry reads one at a time. ### Proposed solution Fetch package version maps concurrently with bounded parallelism, reuse that map for stable/canary version calculation, and keep a fresh parallel absence check for the target publish version. ### Alternatives considered Keeping the existing serial shell loop is simpler, but it preserves the CI latency cost. Caching the final target-version existence check was rejected because release publish safety should still query npm freshly before publishing. ### Roadmap alignment This is a small release-tooling performance improvement. It does not duplicate any planned core product work found in `ROADMAP.md`. ## What Changed - Added `scripts/release-registry-versions.mjs` to fetch npm package version maps and assert target-version absence with bounded parallelism. - Updated `scripts/release.sh` to prefetch package versions once and to batch the final target-version absence check. - Updated `next_stable_version` and `next_canary_version` to use the prefetched version map when present, with the existing per-package npm fallback preserved. - Added release-registry helper coverage and included it in `pnpm run test:release-registry`. - Hardened the release publish helper tests so their fake `pnpm`/`npm` fixture PATH is preserved under non-login shell execution. ## Verification - `node --test scripts/release-registry-versions.test.mjs` - `pnpm run test:release-registry` - `bash -n scripts/release.sh scripts/release-lib.sh` - `git diff --check` - Safety scan before push: searched changed files for common key/token/password patterns and PII markers; only benign script-name text matched (`secrets:migrate-inline-env`). - Remote PR checks on the latest head passed, including `Typecheck + Release Registry`, `Canary Dry Run`, build, tests, e2e, policy, security scans, and commitperclip review. - Greptile reviewed the latest head with Confidence Score 5/5 and no blocking issues. ## Risks - Low risk. The release version helpers keep their original npm fallback when no prefetched version map is supplied. - The existence check remains fresh and uncached before publish, but now reports all matching package/version pairs from a parallel check. - If npm has transient failures during the prefetch step, missing or failed packages still map to an empty version list, matching the old helper behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent using GPT-5, with shell/tool execution in the local repository. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
bdffd26ad3 |
fix(ui): pass company context to custom image setup (#9028)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environment settings let board users configure sandbox execution targets for agent runs > - Sandbox custom-image setup is company-scoped because the backend resolves provider secrets and access through company context > - The browser UI already knows the selected company, but the custom-image setup calls were not passing it to routes whose OpenAPI contract includes `companyId` > - In multi-company authenticated deployments, the backend cannot safely infer company context and returns a `companyId query parameter is required` error > - This pull request passes the selected company id through the custom-image overview, setup, rollback, and disable UI paths > - The benefit is that custom-image setup works consistently in multi-company instances while preserving backend company-boundary checks ## Linked Issues or Issue Description - No public issue is filed for this regression. - Related prior work: #8911 - Bug summary: after the browser SSH terminal custom-image setup flow shipped, opening custom-image setup in an authenticated multi-company instance could show `companyId query parameter is required for environment customImage setup` instead of starting the setup session. - Expected behavior: the environment settings UI should send the selected company context to company-scoped custom-image endpoints. - Reproduction shape: run Paperclip in authenticated/private mode with more than one company, open a sandbox environment's edit dialog, and use the custom-image setup controls. ## What Changed - Added company-id query construction for custom-image overview, setup, rollback, and disable calls in the UI environment API wrapper. - Passed the selected company id into the custom-image panel used by the environment edit dialog. - Updated the environment page tests so the regression fails if company context is dropped again. ## Verification - `pnpm exec vitest run ui/src/pages/CompanyEnvironments.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - `git diff --check` - Also applied the same patch to a local dev checkout serving port 3100 and confirmed `/api/health` still returns `ok` after the dev watcher reload. ## Risks - Low risk: this only adds the selected company id to UI calls for endpoints that already declare or require company context. - If the selected company id is stale or invalid, the existing server-side company access checks still reject the request. ## Model Used - OpenAI Codex CLI using GPT-5, with tool-enabled repository inspection, editing, shell command execution, and local test execution. Context window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a328ec953a |
Fix inherited workspace reuse fallback (#8963)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent heartbeats provision execution workspaces before invoking local or sandboxed adapters. > - Some follow-up issues intentionally request `reuse_existing` so they continue in an inherited execution workspace. > - The heartbeat provisioning path treated missing or archived workspace rows as if no explicit reuse request existed. > - That could silently realize and persist a fresh project/default workspace over an explicit inherited-workspace binding. > - This pull request keys explicit reuse off the issue preference and workspace id, then either restores that workspace or fails with a structured workspace validation error. > - The benefit is that intentional workspace inheritance remains auditable and does not silently degrade into unrelated fallback workspaces. ## Linked Issues or Issue Description Refs #8058 Refs #6036 Refs #2203 This fixes a narrower heartbeat provisioning bug around explicit `reuse_existing` issue runs: if the target inherited execution workspace is missing, archived, or fails restore, provisioning now reports the reuse failure instead of replacing the issue's workspace binding with a freshly realized fallback. ## What Changed - Added explicit helpers for resolving workspace reuse requests and deciding whether reuse should restore, refresh metadata, or keep prior replacement-class drift visible. - Changed heartbeat workspace provisioning so explicit `reuse_existing` requests go through restore-or-fail behavior instead of falling back to `realizeExecutionWorkspace` when the stored workspace row is unavailable. - Added structured `workspace_validation_failed` details for inherited workspace reuse failures. - Added regression coverage for replacement-class drift, restore errors, missing rows, archived rows, and restore misses. ## Verification - `pnpm install --frozen-lockfile` - `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-workspace-session.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check origin/master...HEAD` - Scanned the branch diff and commit messages for credentials, tokens, private URLs, PII-style values, and internal issue links before pushing; no unsafe hits remained. ## Risks - Explicit reuse requests whose stored workspace cannot be restored now fail the run instead of opportunistically creating a replacement workspace. That is intentional, but it may surface stale or archived workspace rows as visible provisioning failures that require repair. - Non-reuse workspace provisioning still uses the existing realization path, so the behavior shift is scoped to issues that explicitly request existing workspace reuse. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex local agent, with shell/tool use enabled for repository inspection, code editing, verification, git, and GitHub CLI operations. Runtime context-window details were not exposed by the adapter. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
85e36aaefb |
Show heartbeat progress in run logs (#8965)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The agent detail page is where operators inspect heartbeat runs and watch live execution output. > - Heartbeat runs can publish progress events while longer operations are happening. > - The live log viewer already appended streamed log and structured run events, but it ignored progress events for the same run. > - That made useful progress text invisible in the run log until another event type arrived or the operator inspected other surfaces. > - This pull request renders run progress events as system log lines in the live agent run viewer. > - The benefit is clearer live feedback during long-running heartbeat operations without changing the backend event contract. ## Linked Issues or Issue Description No public GitHub issue exists, so this PR describes the issue inline following the bug report template. ### What happened The agent detail run log subscribed to company live events and handled `heartbeat.run.log` plus structured `heartbeat.run.event` payloads, but it ignored `heartbeat.run.progress` events for the active run. ### Expected behavior When a heartbeat run emits a progress message, the active run log should show that message immediately as operator-visible system output. ### Steps to reproduce 1. Open an agent detail page for a live heartbeat run. 2. Trigger a run operation that emits `heartbeat.run.progress` events with a `message` and optional `phase`. 3. Watch the live log viewer. Before this change, the progress event was ignored by the log viewer. After this change, it appears as a system log line, prefixed by `[phase]` when a phase is present. ### Paperclip version / deployment mode Current `master`; local development and normal board UI deployments. ### Related work search Searched public GitHub issues and PRs in `paperclipai/paperclip` for `heartbeat.run.progress AgentDetail` and `run progress log viewer`; no duplicate issue or PR was found. ## What Changed - Added live handling for `heartbeat.run.progress` events in `AgentDetail`'s run `LogViewer`. - Render progress messages as `system` log lines for the matching run. - Include the optional progress phase in the displayed line as `[phase] message`. - Prefer the event's `updatedAt` timestamp when provided, falling back to the live event timestamp. - Added a replay key for progress log lines so WebSocket reconnect replay does not duplicate the same rendered progress line. - Added focused formatter/key tests covering phased progress, unphased progress, empty messages, and replay-key output. ### Visual output example The rendered log text is covered by the new formatter test: ```text [workspace] Syncing issue history Preparing workspace ``` No layout or styling changes are included; this PR only makes existing log-line UI receive one more live event type. ## Verification - `pnpm install --frozen-lockfile` — completed; emitted non-fatal bin-link warnings for the unbuilt plugin SDK dev CLI. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/AgentDetail.progress.test.ts src/context/LiveUpdatesProvider.test.ts` — passed, 2 files / 26 tests. - `git diff --check` — passed. - Local sensitive-content scan over the PR diff using patterns for API keys, tokens, secrets, passwords, auth headers, private keys, localhost/private paths, internal ticket ids, agent links, and tailnet markers — no findings. ## Risks Low risk. This is a UI-only live-event handling change for an existing event type. The replay guard is intentionally scoped to progress lines and uses the rendered timestamp, stream, and chunk as the key, so repeated progress events with distinct timestamps or messages still appear. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent with repository tool use, shell execution, GitHub CLI access, and local test execution. Context window size was not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bcac517f3b |
Add browser SSH terminal for custom image setup (#8911)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Environment sandboxes already support custom image creation and refresh through a temporary SSH setup session. > - The existing workflow makes operators copy an SSH command into an external terminal before they can install packages or make image changes. > - That extra context switch is slower, easier to get wrong, and less integrated with the setup session Paperclip already tracks. > - This pull request adds an embedded browser SSH terminal for custom image setup, so operators can start working in the target sandbox directly from the environment configuration flow. > - The implementation uses short-lived websocket attachment tokens, session-lifetime SSH host-key pinning, and server-managed terminal cleanup so the feature fits the existing setup-session boundary. > - The benefit is a smoother custom image creation and refresh experience without asking users to leave Paperclip for routine sandbox setup work. ## Linked Issues or Issue Description No public GitHub issue exists. ### Subsystem affected Cross-cutting: `server/` custom image setup APIs and websocket handling, `ui/` environment configuration UI, and shared custom image contracts. ### Problem or motivation Custom image creation and refresh require an operator to open a separate SSH client, paste the command shown by Paperclip, perform setup work, then return to the browser to finish the image flow. This is functional but awkward for a setup process that already starts and tracks a temporary sandbox session. ### Proposed solution Embed an SSH terminal in the custom image setup UI. When a setup session exposes an SSH payload, Paperclip should open a browser terminal backed by a server-side websocket session, let the operator run setup commands in-place, and then close the terminal when setup is finished, cancelled, expired, or disconnected. ### Alternatives considered - Keep the existing copy/paste SSH command workflow. This remains a fallback, but it does not streamline the common path. - Put SSH credentials directly into websocket URLs. This was avoided so terminal authentication can happen in an explicit first websocket auth frame rather than in logged URLs. - Trust the SSH host blindly for every reconnect. This PR instead pins the observed host-key fingerprint for the setup-session lifetime. ### Roadmap alignment This fits the roadmap theme of making agent workspaces usable in more remote and sandboxed environments while preserving Paperclip's control-plane model. ### Additional context Public GitHub search did not find a duplicate issue or PR for `custom image terminal ssh` in `paperclipai/paperclip`. ## What Changed - Added server-side terminal session tracking for custom image setup sessions, including connect-token issuance, websocket attachment, expiry, resize, input, and shutdown handling. - Added an embedded browser terminal to the custom image creation and refresh flow when a setup session provides SSH connection details. - Moved terminal token authentication out of the websocket URL and into the first websocket JSON auth frame. - Added SSH host-key SHA-256 pinning for each terminal session and documented the provider convention for username-embedded SSH credentials. - Updated the custom image environment API and UI so the setup terminal can open, reconnect, show status, authenticate, resize, and remain active for the setup-session lifetime once attached. - Kept custom image setup routes company-scoped and closed active terminal sessions on setup finish/cancel. - Added focused unit/integration/UI coverage for token expiry, setup-session expiry, websocket close paths, host-key pinning, and terminal session lifecycle behavior. - Removed the generated lockfile delta from the PR; CI owns temporary lockfile regeneration for manifest-changing PRs. ## Verification - `pnpm exec vitest run server/src/__tests__/server-startup-feedback-export.test.ts server/src/__tests__/environment-custom-image-terminal-ws.test.ts server/src/services/environment-custom-image-terminal-sessions.test.ts server/src/__tests__/environment-custom-image-routes.test.ts packages/shared/src/environment-custom-images.test.ts ui/src/pages/CompanyEnvironments.test.tsx` - 6 test files passed - 58 tests passed - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server build` - `pnpm --filter @paperclipai/ui build` - `pnpm run typecheck:build-gaps` - `git diff --check` - Local sensitive-content scan over the PR diff using patterns for API keys, private keys, private hostnames, local paths, token fields, and credential-like strings. - Findings were limited to removed URL-token code and synthetic test placeholders such as `ssh-token-secret` and `terminal-token-terminal-token-123456`. - No real credentials, private hostnames, local filesystem paths, or instance-local links were found. - Remote PR checks were green after the implementation commit, including Build, Typecheck + Release Registry, General tests, serialized server suites, e2e, verify, Socket, Snyk, Superagent, and Greptile 5/5. - Post-merge PR hardening on July 3, 2026: merged `origin/master` at `47448721e` into the branch, resolved the `CompanyEnvironments.tsx` import conflict, reran focused tests, server/UI typechecks, server/UI builds, `pnpm run typecheck:build-gaps`, and `git diff --check`, scanned the final diff for sensitive content, pushed `4b43558cc`, and confirmed all remote checks plus Greptile 5/5 were green. - PR metadata correction on July 3, 2026: changed the title/body framing from bug-fix language to feature-request language. No source files changed for this metadata-only update. ## Risks - Moderate surface area because this adds websocket routing, setup-session runtime state, package dependencies, and a new custom image UI path. - New websocket attachments still require valid short-lived tokens; established terminal sessions remain bounded by setup-session expiry, explicit finish/cancel, client close, or server shutdown. - The terminal-session store is in-memory, so active terminal websocket tokens and host-key pins do not survive server restarts. - SSH host-key verification uses session-lifetime TOFU pinning because the current provider payload does not expose a trusted host-key fingerprint. - The external SSH command remains important as a fallback if a browser, proxy, or network environment cannot sustain the websocket terminal. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent with shell/tool execution. Context window size was not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c48feee190 |
Improve live agent feedback during sandboxed runs (#8915)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A core part of that experience is watching active agent runs without dropping into raw logs first > - Local and sandbox-backed adapters already record useful run output, progress, and tool activity > - But active issue threads could sit visually stale while the agent was syncing workspaces, tailing sandbox output, or emitting incremental tool-call updates > - Operators need timely, human-readable progress while preserving the raw transcript underneath > - This pull request streams sandbox run-log progress into runtime status, keeps visible issue threads refreshed, and folds repeated ACPX tool updates into stable transcript cards > - The benefit is that long-running agent work becomes easier to supervise without changing the task/comment control-plane model ## Linked Issues or Issue Description No public GitHub issue exists for this exact change. Problem/motivation: - During long-running sandboxed agent work, the issue UI can appear idle even though the agent is actively syncing, running tools, or producing incremental output. - Operators need realtime feedback at the issue-thread layer, not only after opening raw logs or waiting for the final heartbeat result. - Related public context: #1808 previously added live-run status dots to Projects; #4362 touches heartbeat wakeup behavior but is not a duplicate of this runtime/UI feedback change. ## What Changed - Added sandbox run-log streaming support and defaulted sandbox-capable local adapters into the richer live-feedback path. - Surfaced environment/sandbox sync progress through heartbeat runtime status with bounded, redacted snippets. - Added live issue-thread cache patching so visible active runs update as progress events arrive. - Folded repeated ACPX `tool_call` updates into one transcript card instead of stacking duplicate cards. - Updated adapter docs and added focused regression coverage for sandbox log streaming, runtime status, ACPX parsing, live updates, transcript rendering, and issue chat messages. ## Verification - `pnpm install --frozen-lockfile` - `pnpm exec vitest run ui/src/context/LiveUpdatesProvider.test.ts` - `pnpm exec vitest run server/src/services/heartbeat-run-runtime-status.test.ts server/src/__tests__/heartbeat-runtime-state.test.ts ui/src/context/LiveUpdatesProvider.test.ts` - `pnpm exec vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts server/src/services/heartbeat-run-runtime-status.test.ts server/src/__tests__/agent-live-run-routes.test.ts server/src/__tests__/heartbeat-runtime-state.test.ts packages/adapters/acpx-local/src/ui/parse-stdout.test.ts ui/src/context/LiveUpdatesProvider.test.ts ui/src/components/transcript/RunTranscriptView.test.tsx ui/src/lib/issue-chat-messages.test.ts ui/src/components/IssueChatThread.test.tsx` - GitHub PR workflow on head `8397953e7b41ccd42e5d9457ee7e4dfb996e4ec5`: `verify`, build, typecheck/release-registry, e2e, general shards, serialized server shards, and canary dry run passed. - Greptile Review on head `8397953e7b41ccd42e5d9457ee7e4dfb996e4ec5`: Confidence Score 5/5, no unresolved review threads. ## Risks - Live issue-thread cache patching could miss an edge case for a route shape not covered by tests. - Surfacing active-run snippets needs continued care around redaction; this PR keeps snippets bounded and adds redaction-focused coverage. - More frequent active-run UI refreshes could expose performance issues on very large issue threads, though updates are scoped to visible run/query caches. ## Model Used OpenAI GPT-5 via Codex, operating as a tool-enabled coding agent with shell, git, and repository-editing capabilities. Context window size is not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bf982c8c83 |
Normalize adapter display labels (#8913)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter names are part of the board-facing agent setup and management experience. > - The product now treats adapters as harnesses, while execution environments are modeled separately. > - Several built-in adapter labels still carried legacy local wording from the older harness-by-environment model. > - That wording makes the UI noisier and implies a distinction users no longer need to reason about. > - This pull request normalizes adapter display labels while keeping persisted adapter type identifiers unchanged. > - The benefit is clearer adapter selection and management copy without a database migration. ## Linked Issues or Issue Description No public GitHub issue was found for this exact cleanup. Related public PRs: - Supersedes #8910, an earlier branch for the same cleanup that did not include the later docs/gateway/Cursor alignment. - Refs #8819, which is related display-registry work for external multi-segment adapter labels, but not a duplicate of this built-in label cleanup. Feature request details: - Subsystem affected: Cross-cutting (`ui/`, `packages/adapters`, and docs). - Problem or motivation: user-facing adapter names include legacy local qualifiers even though adapters map to harnesses and environments are first-class elsewhere. - Proposed solution: remove the legacy local wording from built-in display labels, keep machine-readable adapter type ids unchanged, and keep gateway disambiguation where it is useful. - Alternatives considered: changing persisted adapter type ids was ruled out because it would create migration and compatibility risk; one-off UI replacements were ruled out because the display registry is already the correct central label boundary. - Roadmap alignment: this is small adapter UX polish, not a new roadmap-level core feature. ## What Changed - Updated the adapter display registry so known adapter labels are final and no built-in local adapter renders a legacy local suffix. - Preserved clean derived labels for unknown plugin local types while keeping gateway disambiguation for unknown gateway types. - Updated `AdapterManager` to prefer registry labels when the server reports raw adapter type ids for built-ins. - Removed legacy local wording from built-in adapter metadata labels in UI and adapter packages. - Aligned Cursor adapter metadata with the central display registry label. - Updated adapter docs and Storybook fixtures to match the new display names. - Added focused registry coverage for built-in labels and unknown plugin suffix behavior. ## Verification - `pnpm check:tokens` - `git diff --check origin/master...fix/adapter-display-labels` - Patch-addition scan for added secrets, private paths, and internal links: no matches. - GitHub duplicate search for open adapter-label/local-suffix issues and PRs; #8910 was identified as the older superseded public PR. - `pnpm exec vitest run ui/src/adapters/adapter-display-registry.test.ts` - `pnpm --filter @paperclipai/ui typecheck` - Stale-label scan found no remaining user-facing display-label suffixes; remaining local wording is operational/test terminology such as adapter ids, docs about running locally, and test descriptions. ## Risks Low risk. The change is display-label and documentation focused, and adapter type ids remain unchanged. The main risk is ambiguous gateway naming, mitigated by keeping explicit gateway labels where variants need disambiguation. ## Model Used OpenAI GPT-5 via Codex, tool-enabled coding agent in a local repository workspace. Context window size is not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
936687ca55 |
fix(workspace): restore clean branch drift on finalize (#8914)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs can execute inside reusable, runtime-created git worktree execution workspaces. > - Those managed worktrees record the expected branch so later dispatches do not accidentally run an agent in the wrong checkout. > - Successful run finalization already checked branch coherence, but it treated every unrecorded branch switch as fatal. > - A common publishing flow can briefly switch a clean worktree to a PR/publish branch that points at the same commit as the recorded issue branch, leaving no divergent work to protect. > - This pull request keeps the strict finalization guard for unsafe drift, but lets finalization restore the recorded branch when same-commit repair is provably safe. > - The benefit is fewer false failed runs after harmless branch switches while preserving hard failures for divergent or dirty worktrees. ## Linked Issues or Issue Description No public issue exists for this exact finalization failure. Related public worktree-recovery context: #3087 and #3056, but those address different worktree realization/reuse recovery paths rather than successful-run finalization branch repair. Bug report details: **What happened?** When an adapter run succeeded after switching a managed git worktree from its recorded issue branch to a publish/PR branch, finalization failed with a managed worktree branch mismatch even when the publish branch and recorded branch pointed at the same commit and the worktree was clean. **Expected behavior** Finalization should restore the recorded branch only when it can prove the worktree is clean, registered, and the recorded branch points at the current `HEAD`. If the actual branch has different commits or unsafe state, finalization should continue to fail with bounded validation evidence. **Steps to reproduce** 1. Create a runtime-managed `git_worktree` execution workspace for an issue run. 2. During the adapter run, create and check out a new publish branch without committing new changes. 3. Return adapter success and let heartbeat finalization run. 4. Before this change, finalization records a failed branch check and fails the run even though the branches point at the same commit. 5. With this change, finalization records the repair operation, restores the recorded branch, and records a successful finalize row. 6. Repeat with a commit on the publish branch; finalization still fails because the branch heads differ. **Paperclip version or commit** Reproduced against `master` at `bac7307ec`; fixed by this PR at `64ec605cf`. **Deployment mode** Local dev / built from source. **Agent adapter(s) involved** Not adapter-specific. This is core heartbeat/workspace finalization behavior. **Database mode** Embedded test Postgres in the focused server test. **Access context** Agent run finalization. **Node.js version** `v25.6.1` **Operating system** `Darwin 24.6.0 arm64` **Relevant logs or output** The new focused test intentionally exercises both outcomes: ```text Test Files 1 passed (1) Tests 3 passed (3) ``` **Relevant config** Runtime-created `git_worktree` execution workspace. **Additional context** The unsafe divergent branch case still fails with `workspace_validation_failed` and `git_worktree_branch_incoherence` evidence. **Privacy checklist** Reviewed; this description avoids internal task links, local workspace paths, credentials, and instance-specific URLs. ## What Changed - Reused the existing guarded branch-coherence repair helper during heartbeat finalization when the final branch inspection finds clean same-commit branch drift. - Recorded repair metadata in the `workspace_finalize` operation so reviewers/operators can audit whether finalization repaired branch drift. - Preserved failure behavior for divergent branch heads and surfaced the bounded workspace validation evidence from the repair helper. - Added focused server coverage for safe finalization repair and unsafe divergent branch failure. - Updated execution semantics docs to describe the narrower finalization rule. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` ## Risks Low to medium risk. The change affects successful-run finalization for runtime-created git worktree execution workspaces. The repair path is constrained to clean, registered, same-commit branch drift, and the focused test confirms divergent branch heads still fail instead of being restored silently. ## Model Used OpenAI Codex, GPT-5-based coding agent. Exact hosted model ID was not exposed in the runtime; tool use and local shell execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bac7307ec4 |
fix(adapter-utils): improve sandbox restore failure diagnostics (#8903)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters share managed-runtime helpers from `@paperclipai/adapter-utils` so local and sandboxed runs can prepare, execute, and restore workspaces consistently. > - The sandbox managed runtime syncs workspaces in both directions by creating tar archives on the local host and inside the remote sandbox. > - A prior fix avoided archiving `.` during upload because tar self-entries can force chmod/utime on a directory the command user does not own. > - The restore/download path still archived the remote workspace as `.`, and failed managed commands only surfaced stderr. > - This pull request applies entry-based tar creation to the sandbox restore path and includes stdout in managed command failure diagnostics. > - The benefit is that restore failures become visible to operators and sandbox workspace restore avoids the same directory-metadata failure class already fixed for upload. ## Linked Issues or Issue Description No existing public issue found; describing in-PR (bug). - **What happens:** sandbox workspace restore can fail while creating `workspace-download.tar` from the remote workspace when the tar command includes a `.` self-entry and the command user cannot update metadata on the workspace directory. If the failing command writes its diagnostic to stdout, the managed runtime error can collapse to a generic failed shell command without the useful tar message. - **Expected behavior:** restore should archive the workspace entries without a `.` self-entry, and failed managed runtime commands should include useful stdout/stderr diagnostics. - **Where:** `packages/adapter-utils/src/sandbox-managed-runtime.ts` restore/download path and `packages/adapter-utils/src/command-managed-runtime.ts` command error formatting. - **Related public context:** #7836 fixed the upload side of the same tar self-entry failure class. ## What Changed - Added stdout-aware failed-command formatting in the command managed runtime, keeping diagnostics bounded to the tail of stdout/stderr. - Added remote workspace tarball creation that names top-level entries explicitly instead of archiving `.` during sandbox restore. - Preserved empty-workspace restore support by creating a valid empty tarball when the remote workspace has no entries. - Added regression coverage for stdout diagnostics, restore tar members, and empty workspace restore tarballs. ## Verification - `pnpm vitest run packages/adapter-utils/src/command-managed-runtime.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — 2 files, 17 tests passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - `git diff --check origin/master..HEAD` — passed. - Local public-hygiene/PII scan of the committed diff checked for internal ticket refs, local/private URLs, common token/key patterns, private-key blocks, and email-like values — passed. - GitHub duplicate/related search found no public issue or PR for `workspace-download.tar Permission denied` or `sandbox restore tar permission denied`; #7836 is linked as related prior work. - PR CI on commit `3ad4c8d` — all GitHub Actions lanes, security scans, and aggregate `verify` passed. - Greptile Review on commit `3ad4c8d` — Confidence Score 5/5; the prior P2 thread is resolved with no open P2s, recommendations, or follow-ups. ## Risks Low. This is limited to shared adapter runtime error formatting and sandbox restore archive construction. Archive contents should remain equivalent apart from the removed `.` self-entry, and the new diagnostics are bounded to avoid dumping unbounded command output. ## Model Used OpenAI Codex, GPT-5, tool-enabled coding agent with shell and GitHub CLI access. Context window size was not reported by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (n/a; internal adapter-runtime behavior only) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b4815bf964 |
Scope environment custom images to instance environments (#8850)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments are now managed as instance-level runtime resources rather than per-company rows > - The custom environment image setup tables were introduced with their own `company_id` columns and route query parameters > - That split made one saved environment image state depend on an extra company context even though the environment itself is the durable owner > - It also made saved-environment probes harder because applying the active custom image template could require a company context when no secret-backed config needed one > - This pull request scopes custom image templates and setup sessions directly to the saved environment > - The benefit is that reusable environment images follow the same instance-scoped model as environments while secret resolution still uses company context only when secrets require it ## Linked Issues or Issue Description No matching public GitHub issue was found. Bug report: ### What happened? saved environment custom-image routes and persistence required a `companyId` even though environments are instance-scoped, and saved sandbox probes did not opt into active custom-image template application unless a company context was present. ### Expected behavior custom-image templates and setup sessions should be owned by the saved environment, and saved sandbox probes should apply the active template while still requiring a company context only for secret-backed runtime config. ### Steps to reproduce 1. Configure an instance-scoped sandbox environment with custom-image setup support. 2. Start or inspect a custom-image session or template for that saved environment. 3. Probe the saved environment without a custom-image-specific `companyId` query parameter. ### Paperclip version or commit current `master` after the environment custom-image template migration. ### Deployment mode Local dev (pnpm dev) or authenticated local Paperclip instance. ### Installation method Built from source (pnpm dev / pnpm build). ### Agent adapter(s) involved Not adapter-specific (core bug). ### Database mode Embedded PGlite/Postgres dev database. ### Access context Board human operator. ### Privacy checklist No logs, secrets, tokens, private URLs, or local machine paths are included. Duplicate search performed: - `gh search prs "environment custom image companyId repo:paperclipai/paperclip" --state open --limit 20` - `gh search prs "custom image environment scoped repo:paperclipai/paperclip" --state open --limit 20` - `gh search issues "environment custom image repo:paperclipai/paperclip" --state open --limit 20` The returned results were unrelated adapter, Docker, auth, or stale-workspace items. ## What Changed - Removed redundant `company_id` columns from environment custom-image templates and setup sessions. - Added migration `0127_environment_custom_images_instance_scoped` to collapse duplicate active rows per environment before dropping the old company-scoped indexes/columns. - Updated custom-image services, route handlers, shared validators, and UI API/query keys to use environment-scoped custom-image state. - Kept runtime secret resolution company-aware only when secret refs or bindings require a company context. - Made saved sandbox environment probes opt into active custom-image template application. - Updated DB, shared, server, and UI tests for the new environment-scoped contract. ## Verification - `pnpm --filter @paperclipai/db run check:migrations` - `pnpm exec vitest run packages/db/src/environment-custom-images-schema.test.ts packages/shared/src/environment-custom-images.test.ts server/src/__tests__/environment-custom-image-routes.test.ts server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/environment-routes.test.ts ui/src/pages/CompanyEnvironments.test.tsx` - `pnpm -r typecheck` - `pnpm test:run` before rebasing onto latest `master`; after the rebase only the migration number changed, and the migration check plus focused suite, typecheck, and build were rerun. - `pnpm build` ## Risks - Migration safety: the migration supersedes duplicate active templates per environment and fails duplicate active setup sessions before adding environment-only unique indexes. Operators with duplicate historical active rows should review which active template is kept. - Behavior shift: plugin custom-image setup calls now receive `companyId: "instance"` when no secret binding determines a concrete company context. - Secret-backed configs still require an explicit or uniquely inferable company context; environments with secret bindings spread across multiple companies continue to fail fast. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex via the `codex_local` adapter, GPT-5-based coding model with tool-enabled repository inspection, editing, testing, git, and GitHub CLI access. Exact context-window metadata was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2eba718bef |
Fix sandbox bridge credentials and stalled review recovery (#8844)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The local adapter and heartbeat recovery systems decide whether an agent has a real control-plane mutation path. > - Sandboxed local adapters split execution between the trusted host process and the sandbox shell/tool surface. > - A host-side adapter can still reach Paperclip while the sandbox shell surface cannot, which leaves agents thinking no endpoint or credentials are configured even though the host can still post comments. > - Execution-policy review stages can also remain pending after a reviewer run finishes without recording a decision. > - This pull request makes the sandbox bridge available to the actual shell mutation surface and adds bounded recovery for terminal-but-still-pending review participants. > - The benefit is that agents get a real reachable Paperclip API path where they need it, and stalled review stages become visible recovery work instead of silently drifting. ## Linked Issues or Issue Description No exact public GitHub issue matched this combined failure. I searched for exact and related terms including `cannot reach the Paperclip control plane`, `execution_review_participant_recovery`, `sandbox callback bridge`, `review participant in_review`, and `control plane sandbox`. Related public issues: - Refs #8482 for `in_review` liveness invariant recovery. - Refs #863 for prior agent API-key reachability confusion. - Refs #248 for the broader sandboxed agent execution model. Bug summary: - What happened: a sandboxed local-adapter run could have host-side Paperclip access while the sandbox Bash/tool surface lacked a reachable API endpoint or usable run credentials. Separately, a reviewer run could finish while its execution-review stage remained pending, leaving the source issue in `in_review` with no decision and no live participant run. - Expected behavior: the mutation surface that agents actually use should receive a run-scoped Paperclip bridge, and pending review participants should get one bounded normal-model recovery wake before moving to explicit blocked/source-scoped recovery. - Steps to reproduce: run a sandbox-backed local adapter that needs Bash/curl/tooling to call Paperclip from inside the sandbox, or finish an execution-policy reviewer run without submitting the pending review decision. - Deployment mode: local/authenticated private development instance with sandbox-backed local adapters. ## What Changed - Changed sandbox callback bridge startup so bridge credentials are passed through the sandbox runner environment instead of embedded in the visible `nohup env ...` command string. - Added adapter-utils coverage proving the sandbox shell can call Paperclip through the bridge, forwards the host run JWT with `X-Paperclip-Run-Id`, and does not leak host or bridge tokens into stdout/stderr, runner command text, or runtime files. - Added one bounded execution-review participant recovery path for terminal reviewer runs whose `executionState` remains pending. - Escalated exhausted or non-invokable review participant recovery to blocked/source-scoped recovery with dedicated evidence, activity, and next-action text. - Documented the mutation-surface reachability contract in `doc/execution-semantics.md` and updated the Paperclip skill authentication guidance for sandbox bridge env vars. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts --no-file-parallelism --maxWorkers=1` - `pnpm --filter @paperclipai/adapter-utils typecheck` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` - `curl -fsS $PAPERCLIP_API_URL/api/health` returned `status: ok` on the local instance. ## Risks - Medium behavioral risk: more `in_review` issues with terminal-but-pending reviewer runs will now be retried once and then blocked explicitly instead of remaining quiet. - Low sandbox bridge risk: credential delivery moved from command text to the runner environment, which is less leaky but depends on sandbox providers honoring the env payload for startup commands. - No database migration is included. - Full repo build and CI were not run locally before opening the PR; targeted server/adapter tests and typechecks passed. ## Model Used OpenAI GPT-5 via the Codex local agent, with repository tool use and shell-based code execution. The runtime did not expose a precise context-window value to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d68c34f2cc |
Fix managed workspace branch coherence recovery (#8826)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed issue workspaces are part of the control-plane runtime boundary: the server records which git worktree and branch an agent run is allowed to use. > - Existing reuse checks validated the worktree path and cleanliness, but did not fully validate that the actual checked-out branch still matched the recorded execution workspace branch. > - That gap let an agent run switch a managed worktree onto a publishing branch without updating the execution workspace record, then later reuse or finalize the workspace as though it were coherent. > - The runtime needs a bounded repair path for provably safe mismatches and a hard validation failure for dirty, divergent, or unrecorded branch transitions. > - This pull request adds branch coherence to managed git worktree validation, records explicit recovery evidence, and prevents finalize success when a run silently changes branches. > - The benefit is that branch drift becomes either safely repaired or visibly recoverable instead of silently corrupting managed workspace state. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Bug report: - What happened: a managed agent workspace could be recorded for one branch while the underlying git worktree was actually checked out on another branch. Reuse and finalization could still treat the workspace as healthy. - Expected behavior: managed git worktrees should verify the actual branch against the recorded execution workspace branch. Safe same-HEAD clean mismatches may be repaired, while dirty, divergent, or unrecorded branch transitions should fail into explicit workspace validation recovery. - Reproduction outline: create a runtime-managed issue worktree, switch its checkout to another branch without updating the execution workspace record, then attempt reuse or run finalization. - Deployment mode: local/self-hosted Paperclip server using managed git workspaces. - Related public work: Refs #7644 and #7579. Related but not duplicate: #8275 and #5851. ## What Changed - Added managed git worktree branch inspection, formatted validation evidence, and safe same-HEAD repair logic to the workspace runtime service. - Validated recorded managed workspace branch state before reuse and during heartbeat setup. - Added finalization-time branch guards so runs that silently switch branches fail with `workspace_validation_failed` instead of recording a successful finalize. - Added recovery fingerprints and evidence for `git_worktree_branch_incoherence`, including manual-repair next actions for unsafe branch drift. - Documented branch coherence as part of runtime-created git worktree workspace coherence. - Added focused tests for safe branch repair, dirty/divergent recovery evidence, heartbeat setup validation, and finalize failure/success paths. ## Verification - `pnpm install --frozen-lockfile` - `git diff --check origin/master...HEAD` - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/issue-recovery-actions.test.ts server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` Notes: - An initial full `pnpm test:run` attempt hit a transient `socket hang up` in one `plugin-routes-authz` case. The exact case passed when rerun directly, the full `plugin-routes-authz` file passed, and the subsequent full `pnpm test:run` passed. - `pnpm build` still emits existing Vite CSS pseudo-element and chunk-size warnings unrelated to this change. ## Risks - This intentionally changes behavior for managed runs that switch branches without recording the transition: they now fail during workspace validation/finalization instead of silently proceeding. - The automatic repair path is intentionally narrow. It only repairs clean branch mismatches when both branches point at the same commit; dirty or divergent worktrees require manual recovery. - Recovery fingerprints now include workspace-validation evidence, so duplicate recovery-action grouping is more precise for branch-incoherence failures. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 Codex CLI/API coding agent, with shell/git/test execution and reasoning mode enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.ing> |
||
|
|
a8f0ebaa80 |
Refresh run config before reusing workspaces (#8797)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs are assembled by the heartbeat service from agent config, project workspaces, environment config, secret bindings, skills, and runtime session state. > - The heartbeat service intentionally reuses adapter sessions, execution workspaces, and sandbox leases when that preserves useful state. > - Reuse becomes incorrect when the effective next-run config changes after a saved session, workspace, or lease was created. > - Stale reuse can make a later run appear pinned to old agent, environment, secret, instruction, or workspace settings. > - This pull request records non-sensitive fingerprints for the effective session, workspace, and lease config at run boundaries. > - When those fingerprints drift, Paperclip refreshes persisted runtime config or starts fresh execution instead of reusing stale state. > - The benefit is predictable next-run config freshness without storing raw secret values, full env maps, provider credentials, or private path details. ## Linked Issues or Issue Description - Refs #8058 - Related PRs checked during dedup search: #4968, #4155, #84, #8480. These cover nearby workspace/session routing or model-config freshness areas, but do not duplicate this effective run config fingerprinting path. ## What Changed - Added effective run config fingerprinting for session, workspace, and lease reuse decisions, with canonicalization that ignores generated runtime noise and redacts sensitive values. - Updated heartbeat reuse logic to compare stored and next-run fingerprints, reset stale saved sessions, refresh persisted workspace config snapshots, replace stale reused workspaces when required, and avoid stale sandbox lease reuse. - Included plain environment value drift via value hashes, without storing the raw env values. - Root-bound instruction content hashing so legacy direct absolute instruction paths are represented but not read for config fingerprints. - Batched secret/version metadata lookups for environment lease fingerprinting. - Added workspace operation/run result freshness metadata so operators can inspect non-sensitive decision categories. - Surfaced config freshness labels and next-run copy in the UI and docs. - Added focused coverage for fingerprint redaction, session reset decisions, workspace refresh/replace behavior, environment lease drift, and persisted workspace restoration. ## Verification - `git diff --check` - Sensitive-data scan before push: - `git diff --unified=0 origin/master...HEAD | rg -n --pcre2 "(AWS_ACCESS_KEY_ID|AWS_SECRET_ACCESS_KEY|ghp_[A-Za-z0-9_]{20,}|github_pat_[A-Za-z0-9_]{20,}|sk-[A-Za-z0-9]{20,}|-----BEGIN (RSA |OPENSSH |EC |DSA )?PRIVATE KEY-----|AKIA[0-9A-Z]{16})"` - `git diff --unified=0 origin/master...HEAD | rg -n --pcre2 "[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Za-z]{2,}"` - `pnpm exec vitest run server/src/__tests__/effective-run-config-fingerprints.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/environment-runtime.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm --filter @paperclipai/db clean` - `pnpm test:run` - `pnpm build` - UI screenshots from Cutter: - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-01.png - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-02.png - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-03.png ## Risks - Medium: overly broad fingerprints could start fresh sessions, workspaces, or sandbox leases more often than necessary. - Medium: missing a config category would allow stale reuse to persist for that category. - Medium: legacy direct absolute instruction paths are no longer content-hashed unless they are paired with an absolute managed instructions root. - Low data risk: fingerprint metadata stores hashes and category names, not raw secrets, raw env values, provider credentials, or private path details. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex CLI / Codex coding agent, tool-enabled with shell, Git, GitHub CLI, local test execution, and code editing. The exact deployed model variant and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.ing> |
||
|
|
8d9f9fd240 |
Add reusable sandbox custom images (#8794)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A growing part of that work runs in sandboxed environments rather than on the operator's local machine. > - Today sandbox providers can start fresh workspaces and run probes, but they do not have a shared contract for capturing and reusing prepared sandbox state. > - Operators need a way to set up tools, credentials, and project dependencies once, then reuse that prepared image for later agent runs. > - This pull request adds reusable sandbox custom images across the provider contract, server runtime, and board UI. > - It also keeps probes and sandbox copy flows aligned with pre-authenticated/custom-image environments. > - The benefit is faster, more reliable sandbox runs without repeatedly rebuilding the same environment setup. ## Linked Issues or Issue Description No public GitHub issue was found for this change. Inline feature request follows. ### Problem or motivation Sandboxed agents need reusable prepared runtime state so repeated runs do not require manual setup every time. Operators often need system packages, CLIs, SDKs, dependency caches, credentials, and project tooling available before an agent can work productively. ### Proposed solution Add a provider-level custom-image capability, server-side setup/capture lifecycle, Daytona/fake provider support, and board UI controls for creating, testing, selecting, and deleting custom images. ### Alternatives considered Leaving this as provider-specific setup outside Paperclip would keep the control plane blind to image state and would not give agents consistent environment metadata. Re-running setup commands for every lease is simpler, but slower and less reliable for interactive or credentialed setup. ### Roadmap alignment Checked `ROADMAP.md`; this aligns with the Cloud / Sandbox agents roadmap area and does not duplicate any related public issue or PR found by search. Additional context: - Subsystem affected: cross-cutting (`packages/db`, `packages/shared`, `packages/plugins`, `server`, `ui`). - Duplicate search: searched GitHub for `sandbox custom image` and `sandbox template environment`; no related public issues or PRs were found. ## What Changed - Added custom-image shared types, validators, constants, API paths, and database schema/migration. - Added server services/routes for custom-image templates and setup sessions, including runtime cleanup and provider metadata handling. - Extended plugin/sandbox provider capabilities for interactive setup, template capture, and template deletion. - Implemented custom-image support in the fake sandbox provider and Daytona provider. - Updated environment runtime/config handling so active custom images flow into leases, probes, and agent execution. - Added board UI controls and API client support for custom-image setup, capture, selection, status, and error states. - Hardened sandbox copy/probe behavior for insecure clipboard contexts and pre-authenticated sandbox images. - Added targeted coverage across shared validators, DB schema, server routes/services, provider plugins, adapter probes, and UI flows. ## Verification - `pnpm install --frozen-lockfile --ignore-scripts` - `pnpm vitest run packages/adapters/claude-local/src/server/test.probe.test.ts packages/adapters/claude-local/src/server/test.ts packages/adapters/codex-local/src/server/test.remote.test.ts packages/adapters/codex-local/src/server/test.ts` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm vitest run packages/db/src/environment-custom-images-schema.test.ts packages/shared/src/environment-custom-images.test.ts packages/shared/src/validators/plugin.test.ts server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/workspace-runtime.test.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts ui/src/pages/CompanyEnvironments.test.tsx ui/src/pages/CompanySettings.test.tsx` - `pnpm -r typecheck` - `pnpm build` - `rm -rf packages/db/dist && pnpm test:run` - Public-safety scan of the final diff found no internal Paperclip issue links, private instance URLs, or real secret patterns. ## Risks - Adds a database migration and new environment runtime tables, so migration ordering and rollback need care. - Provider implementations may differ in how reliably they can capture/delete images; unsupported providers surface capability-gated UI states. - Custom-image state can contain operator-prepared tooling and credentials inside the provider image, so providers must enforce their own access controls and cleanup semantics. - Broad surface area across shared contracts, server runtime, plugins, adapters, and UI means CI and Greptile review should be watched closely. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 (`gpt-5`) via Codex CLI with tool use and code execution. Assisted with branch cleanup, conflict resolution, local verification, and PR preparation. Earlier branch implementation work was assisted by Paperclip-managed Claude/Codex agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b61c2852fb |
build(ui): drop redundant tsc type-check from ui build (#8781)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - CI runs a set of required PR checks; the Canary Dry Run job is the slowest (~4m) > - Its dominant cost is `release.sh` Step 2 `pnpm build` (~133s), whose critical path is the `ui` package > - In `ui`, `tsc -b` runs ~46s *before* `vite build` (~20s), but `ui/tsconfig.json` sets `noEmit: true`, so `tsc -b` is a pure type-check that emits nothing — vite produces all artifacts > - The PR workflow already runs a parallel Typecheck job (`typecheck:build-gaps`) that auto-type-checks any workspace package whose build script omits `tsc` > - This pull request drops `tsc -b` from the `ui` build so the type-check moves off the build/release critical path onto the slack-rich Typecheck job > - The benefit is ~46s shaved off the slowest PR check with no loss of type coverage ## What Changed - `ui/package.json`: `build` script changed from `tsc -b && vite build` to `vite build`. - Type coverage is preserved automatically: `scripts/run-typecheck-build-gaps.mjs` already type-checks any workspace whose `build` script omits `tsc` and defines a `typecheck` script. With this change, `@paperclipai/ui` (whose `typecheck` is `tsc -b`) is now picked up by the parallel Typecheck job. No script edit was required — the gap-detection is generic. ## Verification - `pnpm --filter @paperclipai/ui build` produces `ui/dist` via vite alone (no `tsc` invocation). - `pnpm typecheck:build-gaps` now type-checks `@paperclipai/ui` and exits 0. - A type error in `ui` still fails the parallel Typecheck job, so type-error coverage is unchanged. ## Risks - Low. The type-check is relocated, not removed. This is build-tooling only — no migration, no runtime behavior change. If `scripts/run-typecheck-build-gaps.mjs` ever stopped covering `ui`, a type error could slip past the build job — but its gap-detection is generic (any package with a tsc-less build plus a `typecheck` script), and `ui` qualifies. ## Model Used Claude (Anthropic), `claude-opus-4` family, extended-thinking mode with tool use, operating as an automated CI-health agent. |
||
|
|
0c2ec7deb4 |
Fix sandboxed Claude and Codex probe behavior (#8775)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters are the bridge between Paperclip's control plane and provider CLIs such as Claude Code and Codex. > - Those adapters can run either on the host machine or inside a remote/sandbox execution target. > - Sandbox probes need to validate the same auth/config path that real sandbox execution will use. > - The previous probe paths could surface misleading Claude errors, rely on host-only Codex state, or upload far more Codex home state than the probe needed. > - This pull request fixes the Claude and Codex sandbox probe/runtime behavior together while keeping provider-specific sandbox image work out of scope. > - The benefit is faster, clearer adapter health checks that better match real sandbox execution. ## Linked Issues or Issue Description No public GitHub issue was found for this exact bug during duplicate search. Bug report: **What happened?** Sandboxed Claude/Codex adapter tests could diverge from real runtime auth/config behavior. Claude sandbox probes could show the leading stream init line instead of the real final error, and Codex sandbox probes could upload full managed home state or mask a sandbox-local login with an empty uploaded `CODEX_HOME`. **Expected behavior** Sandbox probes should exercise the remote runtime contract, preserve useful sandbox credentials, avoid relying on unrelated host state, and report actionable probe failures. **Steps to reproduce** 1. Configure a remote/sandbox execution target for `claude_local` or `codex_local`. 2. Run the environment Test/probe path where host credentials differ from the sandbox's runtime credentials or the managed Codex home contains session history. 3. Observe that probe behavior can differ from the actual sandbox runtime path or surface an unhelpful Claude stream initialization line. **Paperclip version or commit** Current `master` before this PR, based on `4a2447da3`. **Deployment mode** Local development/control-plane deployment with remote sandbox execution targets. Related search performed: - Public issues: `Claude sandbox probe`, `Codex CODEX_HOME sandbox` returned no matches. - Public PRs: `Claude Codex sandbox probe`, `codex home sandbox`, `claude auth sandbox` returned no matches. ## What Changed - Made Claude sandbox Test probes materialize the same Paperclip-managed Claude config seed path used by sandbox execution. - Preserved sandbox-local Claude credentials when materializing remote Claude config and expanded auth-required detection for `/login` API-key failures. - Improved Claude hello-probe diagnostics so the final result/error is surfaced instead of the unhelpful stream init event, with transient upstream failures downgraded to warnings. - Changed Codex probe behavior to upload only minimal auth/config files instead of the full managed `CODEX_HOME`. - Let Codex sandbox probes leave `CODEX_HOME` unset when the host has no credentials, so pre-authenticated sandbox images can be tested directly. - Excluded bulky host-local Codex session/shell state from sandbox runtime home uploads. - Switched the Codex local default model away from the ChatGPT-unsupported `gpt-5.3-codex` option. - Added regression coverage for Claude parsing/probe paths, Codex adapter metadata/argument/probe behavior, and server-level Claude sandbox environment behavior. ## Verification Passed locally: - `pnpm install --frozen-lockfile` - `pnpm vitest run packages/adapters/claude-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/test.probe.test.ts server/src/__tests__/claude-local-adapter-environment.test.ts` - `pnpm vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts packages/adapters/codex-local/src/server/test.remote.test.ts` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` ## Risks - Adapter configuration behavior is sensitive to local vs sandboxed execution mode, so review should focus on environment detection, argument construction, and any state written during probe/test runs. - The Codex default-model change may affect newly created agents that rely on the adapter default instead of an explicit model. - Excluding Codex session/shell state from sandbox uploads should be safe for fresh sandbox runs, but reviewers should confirm no runtime resume path depends on that host-local state. - Provider-specific setup/capture behavior is intentionally left to separate work. ## Model Used OpenAI GPT-5 Codex via Paperclip `codex_local`; tool-enabled local coding session with terminal access. Context window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f019f54bb3 |
Fix active heartbeat run reaping (#8776)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Heartbeat monitoring is the subsystem that keeps agent execution visible and recovers work only when execution continuity is genuinely lost. > - Routes and scheduler paths can construct separate heartbeat service instances inside the same server process. > - Active adapter execution tracking was scoped to each service instance, so the periodic orphan reaper could miss a run that another service instance was actively executing. > - Remote or sandbox adapters are especially exposed because they may not persist a local process PID/group and may go quiet while the remote command is still alive. > - This pull request makes active in-process adapter execution tracking shared across heartbeat service instances and adds a regression for the cross-instance reaper case. > - The benefit is fewer false `process_lost` failures for long-running or quiet sandbox/remote agent runs. ## Linked Issues or Issue Description No public GitHub issue exists. This PR describes the bug inline using the bug report template fields. ### What happened? An actively executing heartbeat run could be finalized as `process_lost` by the orphan reaper when adapter execution was active through one `heartbeatService()` instance but the reaper ran through another instance in the same server process. ### Expected behavior The orphan reaper should skip runs that are still actively executing in-process, regardless of which `heartbeatService()` instance is doing the reaping. ### Steps to reproduce 1. Create two `heartbeatService()` instances in the same process. 2. Start an adapter run through the first instance. 3. Backdate the run row enough for orphan reaping to consider it stale. 4. Run orphan reaping through the second instance while the first instance is still awaiting adapter execution. 5. Observe that the old instance-local tracking can mark the live run as `process_lost`. ### Paperclip version or commit Reproduced against `master` before commit `44ba6d8bb4f7ae1ca3715697f750844d770d83a3`. ### Deployment mode Self-hosted/local server process with route and scheduler code paths constructing separate heartbeat service instances. Remote or sandbox adapters are the highest-risk case because they may not have local PID metadata and can be quiet while still running. ## What Changed - Moved active adapter execution tracking from the `heartbeatService()` closure to module-level process state shared by heartbeat service instances. - Added a regression test that starts a run through one heartbeat service instance and runs orphan reaping through another, proving the active run is not reaped and can finish normally. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts` passes: 61 tests. - `pnpm -r typecheck` passes. - `pnpm build` passes; existing UI build warnings remain for `::highlight(...)`, large chunks, and a mixed static/dynamic import. - `pnpm test:run` does not fully pass in this local environment: 1 unrelated existing failure in `server/src/__tests__/workspace-runtime.test.ts` for `auto-detects the default branch via symbolic-ref when origin/HEAD is set`. The fixture command fails with `git push -u origin main master` because the temp repo has no `master` ref. - Reran the isolated failing test with `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts -t "auto-detects the default branch via symbolic-ref when origin/HEAD is set"`; it reproduces the same missing-`master` ref failure. ## Risks - Low risk for single-process Paperclip servers: this only broadens in-process active run tracking across service instances. - Multi-process deployments still need persisted or distributed execution liveness to coordinate reaping across processes; this PR does not claim to solve cross-process recovery. - A run could be skipped by the reaper while its adapter promise is active, but the existing `finally` path removes the active marker after execution settles. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex, with tool use and local command execution. The runtime did not expose a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted regression, typecheck, and build pass; full `pnpm test:run` has the unrelated missing-`master` ref fixture failure documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes — no docs change needed for this internal bug fix - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip Agent <noreply@paperclip.ing> |
||
|
|
a7a73d5bc7 |
Fix stale server info debug metadata (#8753)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The experimental server info debug view helps local operators inspect what code a running dev instance is actually serving > - The view was moved into the account-menu drawer, which only mounts while that drawer is open > - That made stale health-query data easier to see after restarts, and the server was also caching the running commit at process boot > - A clean commit label alone is incomplete when the checkout has uncommitted local changes > - This pull request keeps the drawer health data fresh, refreshes git metadata on demand, and adds a path-free checkout-state summary > - The benefit is that the debug view reports restart time, running commit, and dirty-checkout state without exposing local paths, secrets, logs, or environment details ## Linked Issues or Issue Description Fixes: #8752 ## What Changed - `SidebarServerInfo.tsx`: refetch the health query whenever the drawer opens and poll every 2s while the dev server is active. - `server-info.ts`: keep `processStartedAt` stable while refreshing git HEAD through a short TTL cache instead of freezing commit metadata at module boot. - Shared health contract/OpenAPI: add `serverInfo.git.localChanges` with only staged, unstaged, and untracked counts plus safe unavailable fallbacks. - `SidebarServerInfo.tsx`: add a `Checkout state` row that renders clean/dirty/unavailable copy without file paths. - Tests: cover stale drawer refresh, interval polling, TTL commit refresh, health response shape, checkout-state count parsing, and path-free UI rendering. ## Verification - `npx vitest run server/src/__tests__/server-info.test.ts server/src/__tests__/health.test.ts ui/src/components/SidebarServerInfo.test.tsx` -> 3 files / 19 tests passing. - `pnpm install --frozen-lockfile --ignore-scripts` -> refreshed stale workspace links without lockfile/source churn. - `pnpm --filter @paperclipai/shared --filter @paperclipai/server --filter @paperclipai/ui typecheck` -> passing. - `pnpm --filter @paperclipai/ui typecheck` -> passing after the Greptile test-coverage fix. - `pnpm check:tokens` -> no forbidden tokens found. - Local diff scans for obvious secrets, credentials, private URLs, local paths, and PII patterns -> no matches. - GitHub PR checks on head `56defd446` -> all green, including `verify`, canary dry run, e2e, security scans, and Greptile Review. - Greptile latest summary -> Confidence Score 5/5, 0 new comments; the prior P2 polling-coverage thread is resolved. ## Risks Low risk. The UI remains behind the experimental `enableServerInfoDebugView` flag. The extra git status call is throttled by the existing server-info TTL and reports only counts, not paths or file names. If git status is unavailable, the commit row still works and the checkout-state row shows clear fallback copy. ## Model Used Claude Opus (claude-opus-4-8), extended thinking, with tool use / code execution assisted the original stale-metadata fix. OpenAI GPT-5 via Codex local, with tool use and code execution, added the checkout-state follow-up, Greptile fix-up, and PR verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
500a75f7ce |
Show agent environment metadata (#8671)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Agents page is where operators scan the current agent roster, runtime status, model, adapter, and last heartbeat. > - Environment selection now affects where agents execute, but the roster did not expose each agent's effective execution environment. > - Operators need that context when multiple environments are configured, especially when sandbox-backed environments use different providers. > - This pull request adds environment/provider metadata to the Agents page while keeping it hidden when the column would not add useful information. > - The benefit is better operational visibility without adding noise for single-environment or environment-disabled instances. ## Linked Issues or Issue Description No matching public GitHub issue was found after searching for related Agents environment-column issues and PRs. ## Problem or motivation Operators use the Agents page to scan active agents, but when multiple execution environments are configured there was no row-level indication of which environment each agent uses. That makes it harder to distinguish host-local agents from sandbox-backed agents and to see the sandbox provider at a glance. ## Proposed solution Show each agent's effective environment in the Agents page metadata when environments are enabled and multiple environments are configured. Resolve the value from the agent override, instance default, or local fallback, and include sandbox provider display names when available. ## Alternatives considered Always showing the column would add noise for single-environment instances. Adding filters and grouping was also considered, but this PR keeps the scope to read-only metadata until the product has more usage data. ## Roadmap alignment This is a small UI visibility improvement related to the roadmap's cloud/sandbox agents direction, without adding a new core workflow or adapter capability. ## Additional context The column is hidden when the experimental environments setting is disabled or when only one environment is configured. ## What Changed - Added environment metadata loading to the Agents page, including environment list, capabilities, and instance settings. - Resolved each agent's effective environment from the agent override, instance default, or local fallback. - Rendered environment name and sandbox provider detail in list and org views when environments are enabled and multiple environments are configured. - Hid the environment column when environments are disabled or only one environment is configured. - Added focused UI tests for display, fallback, loading, hidden-column, disabled-feature, and no filter/grouping behavior. - Addressed Greptile review feedback by preserving custom local environment names, reserving the column while environment data loads, and separating the capabilities query key. ## Verification - `pnpm install --frozen-lockfile` - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Agents.test.tsx` — 11 tests passed - `pnpm --filter @paperclipai/ui build` - Local PR-diff scan for secrets/private references before push - GitHub PR workflow passed on head `bb770acc3` - Commitperclip review check passed - Greptile Review passed with confidence score 5/5 and no inline review comments ## Risks Low risk. The change is limited to the Agents page, query keys, and tests. The main behavioral risk is extra metadata query traffic on the Agents page when environments are enabled; those queries are skipped when the experimental environments flag is disabled. ## Model Used OpenAI Codex, GPT-5 coding agent with shell/tool use and code execution in the local repository. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
765a75207a |
Add experimental server info debug view (#8676)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The dev/server UI exposes a `/api/health` endpoint and a lower-left
account drawer, but nothing surfaces *which* build the running instance
is on or when it last restarted
> - When iterating on a local dev instance it is hard to tell whether
the server you're looking at has actually restarted onto your latest
commit, or how stale the running process is
> - Developers need a lightweight, opt-in way to confirm the running
instance's identity without digging through logs or shelling into the
host
> - This pull request adds an experimental "Server Info Debug View"
setting that surfaces the running instance's last-restart time and
current commit as read-only rows in the account drawer
> - The benefit is a quick, in-UI sanity check of what the live server
is actually running, behind an experimental flag so it ships zero cost
to users who don't opt in
## Linked Issues or Issue Description
No public GitHub issue exists. Describing the underlying request inline
following the feature request template:
**Problem or motivation:**
When working against a local Paperclip dev instance there is no in-UI
way to confirm what the running server is — its current commit or when
it last restarted. You have to check logs or the host shell to know
whether the process picked up your latest build.
**Proposed solution:**
An opt-in experimental setting ("Server Info Debug View") that, once
enabled, renders a small read-only "Server" section at the bottom of the
lower-left account drawer showing **Last restarted** (the server process
start time) and **Running commit** (the current git HEAD short SHA +
subject).
**Alternatives considered:**
A separate top-right pill/overlay (like the work-life-balance plugin).
The account drawer was chosen to reuse existing menu-row styling and
avoid adding new always-present chrome.
**Roadmap alignment:**
Small, self-contained developer-experience aid gated behind an
experimental flag; does not overlap planned core roadmap work.
## What Changed
- Added `server/src/server-info.ts`: captures a `serverInfo` snapshot
once at boot — process start time and current git commit (SHA +
subject). Git is read via `execFileSync` with SHA validation and a
timeout.
- `/api/health` exposes the `serverInfo` snapshot, but only on
full-details health responses (board/agent in authenticated mode, or
local-trusted dev).
- Gated the UI surface behind a new `enableServerInfoDebugView`
experimental setting, wired through the shared instance type, validator,
settings normalizer, and OpenAPI schema.
- UI: added `SidebarServerInfo` rendering the read-only rows in the
account drawer (`BreadcrumbBar` / `SidebarAccountMenu`), plus the
experimental settings toggle and a typed `health` API client.
- Moved `ServerGitInfo` / `ServerInfoSnapshot` into
`@paperclipai/shared` so the server and UI share one definition instead
of duplicating it.
- Added unit tests for the server-info snapshot, health route exposure,
validator/normalizer, settings routes, the experimental settings page,
and the sidebar component.
## Verification
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/health.test.ts src/__tests__/server-info.test.ts
src/__tests__/instance-settings-service.test.ts
src/__tests__/instance-settings-routes.test.ts` — 32 passed
- `pnpm --filter @paperclipai/ui exec vitest run
src/components/SidebarServerInfo.test.tsx
src/pages/InstanceExperimentalSettings.test.tsx` — 9 passed
- `tsc --noEmit` on both `@paperclipai/server` and `@paperclipai/ui` —
clean
- Manual: enable **Settings → Experimental → Server Info Debug View**,
refresh the UI, open the lower-left account drawer — a "Server" section
shows Last restarted and Running commit.
## Risks
- Low risk. The UI surface is fully opt-in via an experimental flag and
defaults off.
- The `serverInfo` field on `/api/health` is access-controlled to
full-details responses only (board/agent in authenticated mode, or
local-trusted dev) — never anonymous authenticated callers — so the git
SHA is not broadly exposed.
- The only new server work is a one-time git read at boot, guarded with
SHA validation and a timeout; failures degrade gracefully (the git block
reports `available: false` rather than throwing).
## Model Used
Claude — `claude-opus-4` (Anthropic), extended thinking with tool use,
via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no
instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
fdb8b5678b |
fix(ui): restore main-content scroll position on browser back/forward (#8636)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The web UI (`ui/`) renders inside a single persistent shell (`Layout`) whose `#main-content` element is the scroll container for every route > - Because that scroll container survives route changes, browser back/forward (a history `POP`) lands on the previous page with the scroll offset of the page you just left > - Concretely: scroll deep into an issue detail, hit browser back, and the inbox renders scrolled to the issue-detail offset instead of where you left off — making it hard to find your place > - The app already reset scroll for forward (`PUSH`) navigation but deliberately did nothing on `POP`, so nothing restored the prior position > - This pull request records each history entry's scroll offset as the user scrolls and restores it on `POP`, while keeping the existing reset-to-top behavior for `PUSH`/`REPLACE` > - The benefit is that back/forward returns you to the exact place you were, so the inbox (and any scrolled page) keeps your reading position ## Linked Issues or Issue Description No public GitHub issue — describing the bug inline following the bug report template. **What happened** From the inbox, open an issue, scroll down within the issue detail, then press the browser Back button. The inbox is restored at the wrong scroll position — it shows the Y offset from the issue-detail page rather than the position you had in the inbox. **Expected behavior** Browser back/forward returns the page to the scroll position it had when you left it. **Steps to reproduce** 1. Open the inbox and scroll to a known position partway down the list. 2. Click into an issue. 3. Scroll down within the issue detail. 4. Press the browser Back button to return to the inbox. 5. Observe the inbox is scrolled to the issue-detail offset instead of where you left off. **Deployment mode** Web UI (`ui/`), any deployment — client-side scroll behavior only. Root cause: `#main-content` is a single scroll container that stays mounted across route changes. Scroll was only reset on forward (`PUSH`) navigation and left untouched on `POP`, so the stale offset from the outgoing page persisted onto the page being returned to. ## What Changed - Added `NavigationScrollMemory` (`ui/src/lib/navigation-scroll.ts`): a per-history-key map of `#main-content` scroll offsets, clamped to `>= 0`. - Added `applyMainContentScrollTop` helper to restore a saved offset onto the main content element (null-safe). - In `Layout` (`ui/src/components/Layout.tsx`): continuously record the active history entry's scroll offset on scroll, and on `POP` navigation restore the remembered offset (re-applying on the next animation frame so a late-laying-out cached page doesn't clamp the offset to a shorter interim height). Forward `PUSH`/`REPLACE` keeps the existing reset-to-top behavior. - Added unit tests covering the remember/recall logic and the restore helper. ## Verification - `ui` unit tests pass, including the new `navigation-scroll` cases (remember/recall per key, clamping, and DOM restore). - TypeScript clean on the changed files. - Manual: inbox → open issue → scroll down → browser Back returns the inbox to its previous scroll position; forward navigation still resets to top. ## Risks Low risk. Scoped to client-side scroll restoration in the web UI; no API, schema, or migration changes. The only behavioral change is that `POP` navigation now restores a saved offset instead of leaving the container untouched; `PUSH`/`REPLACE` behavior is unchanged. Memory is per-session and bounded by visited history keys. ## Model Used Claude (Anthropic), Opus-class model via Paperclip's `claude_local` adapter, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5170a9d35d |
Test cheap model during agent config check (#8632)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent configuration includes adapter model settings and optional cheap model profiles used by runtime lanes. > - The agent configuration screen already exposes a Test action for adapter environment checks. > - When a cheap model profile is configured, that Test action only exercised the primary model configuration. > - That left users able to save a cheap model profile that had not been validated by the same configuration test flow. > - This pull request makes the Test action probe both the primary model and the configured cheap model. > - The benefit is earlier feedback when the cheap model is unavailable or misconfigured. ## Linked Issues or Issue Description No exact public issue found. Problem description: - What happened: the agent configuration Test action validated the primary adapter model but did not validate an enabled cheap model profile. - Expected behavior: when a cheap model is configured and enabled, the Test action should also test that cheap model using the same environment selection. - Steps to reproduce: configure an agent with a primary model and enabled cheap model profile, then click Test in the agent configuration UI. - Paperclip version/commit: current `master` at PR creation. - Deployment mode: local development UI behavior. Related public context found during GitHub search: - #4881 added cheap model profiles for local adapters. - #6534 is related cheap-primary-model preservation work. - #6956 tracks broader agent configuration settings exposure. GitHub searches performed: - `cheap model config test` - `AgentConfigForm cheap model` - `modelProfiles cheap` - `adapter environment cheap model` ## What Changed - Updated `AgentConfigForm` so the adapter environment Test action runs the primary model check first. - Added a second cheap model check when the cheap profile is enabled and resolves to a model. - Built the cheap test payload from the resolved cheap profile config instead of a model-only override, preserving adapter-default and saved cheap-profile fields. - Merged the individual results into a single labeled result so the UI can show both outcomes together. - Preserved request/API failures so they still surface through the existing error UI instead of becoming synthetic adapter checks. - Added render coverage asserting both primary and cheap test calls, non-model cheap-profile fields, and request failure handling. ## Verification - `git diff origin/master...HEAD | rg -n "(API[_-]?KEY|SECRET|TOKEN|PASSWORD|PRIVATE[_-]?KEY|BEGIN RSA|BEGIN OPENSSH|Bearer [A-Za-z0-9._-]+|ghp_[A-Za-z0-9_]+|sk-[A-Za-z0-9]+|AIza[0-9A-Za-z_-]+|OPENAI_API_KEY|ANTHROPIC_API_KEY)" || true` produced no matches. - `git diff --check origin/master...HEAD` - `pnpm --filter @paperclipai/ui exec vitest run src/components/AgentConfigForm.render.test.tsx --reporter=dot` - `pnpm --filter @paperclipai/ui typecheck` - GitHub PR checks on head `3d9ba15d2` are green, including Build, Typecheck + Release Registry, General tests, serialized server suites, e2e, Canary Dry Run, security scans, and policy/review checks. - Greptile Review passed on head `3d9ba15d2`; all review threads are resolved. ## Risks Low risk. The change is limited to the agent configuration UI test action and its render test. The cheap-model probe now preserves adapter-default and saved cheap-profile fields, matching the runtime merge order more closely. The main residual risk is adapter-specific UI coverage for cheap-profile fields that are stored but not directly editable in this form. ## Model Used OpenAI GPT-5 Codex via the Codex local agent environment, with terminal/tool use for repository inspection, implementation, GitHub CLI operations, and local verification. Exact context window was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f90ea4dae4 |
Fix top-level secret ref binding sync (#8630)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instance environments can store provider configuration fields as Paperclip secret references > - Environment save uses secret binding sync to keep persisted config refs aligned with `company_secret_bindings` > - Top-level secret-ref fields such as `apiKey` were not deleted during sync because the cleanup only matched child paths like `apiKey.*` > - Re-saving an environment with the same top-level secret ref could therefore hit the target/path unique constraint and return a 500 > - This pull request makes the sync cleanup include the exact top-level config path before reinserting current refs > - The benefit is that saved environment provider configs can be edited repeatedly without duplicate binding failures ## Linked Issues or Issue Description No public GitHub issue found in duplicate search for this exact environment secret-binding failure. ### Bug report #### Pre-submission checklist - I searched existing open and closed issues and this is not a duplicate. - I can reproduce this on the current `master` lineage. - I confirmed the error originates in Paperclip secret-binding sync, not the sandbox provider itself. #### What happened? Saving an instance environment whose provider config contains a top-level secret-ref field can fail with a duplicate key error on `company_secret_bindings_target_path_uq`. #### Expected behavior Saving the same environment config repeatedly should update/sync bindings idempotently. #### Steps to reproduce 1. Create or edit an instance environment with a provider config that has a top-level secret-ref field such as `apiKey`. 2. Save the environment. 3. Save the environment again without moving that field under a nested object. 4. The second save can attempt to insert a duplicate binding for the same target/path. #### Paperclip version or commit Observed on local dev from current `master` lineage before this fix. #### Deployment mode Local dev (pnpm dev), authenticated private mode. #### Installation method Built from source (pnpm dev / pnpm build). #### Agent adapter(s) involved Not adapter-specific (core bug). #### Database mode Embedded local Postgres. #### Access context Board (human operator) environment settings save. #### Relevant logs or output The server returned a 500 after Postgres rejected a duplicate `company_secret_bindings` row for the same environment target and `apiKey` config path. Secret values and local paths are intentionally omitted. #### Privacy checklist I reviewed the PR description for private instance links, local paths, API keys, tokens, and company-specific secrets. ## What Changed - Updated `syncSecretRefsForTarget()` so prefix cleanup removes both the exact top-level config path and nested child paths. - Added a regression test that syncs an environment top-level `apiKey` secret ref repeatedly, then replaces it and verifies only one binding remains. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/secrets-service.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` - Local diff scan for internal issue links, local paths, bearer/session tokens, and obvious secret literals returned no matches. - GitHub duplicate searches for related environment secret-binding issues/PRs returned no matches. ## Risks Low risk. The change only broadens the existing target/path cleanup used before reinserting secret refs. It preserves the existing child-path cleanup behavior and adds the missing exact-path case. ## Model Used OpenAI GPT-5 Codex via the `codex_local` Paperclip adapter. Tool-using coding-agent session with shell, git, and repository-edit capabilities. Context window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1951c80237 |
feat(cli): surface plugin install target host + add plugin target (#8575)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The CLI (`paperclipai plugin ...`) installs and manages plugins against a Paperclip server resolved from `--api-base` / `PAPERCLIP_API_URL` / the active profile / an inferred default > - During local plugin development you can have more than one Paperclip running (a released host plus a branch build on another port), and nothing told you *which* instance a command actually talked to > - So a plugin that depends on a route or response field only present on a feature branch could be silently installed/tested against a stale host, returning `API route not found`, and look broken when the real problem was the test target > - This pull request makes the install target explicit: it probes `GET /api/health` and prints the resolved API URL + server status/version/mode/exposure before installing, and adds a `plugin target` command plus docs for running and verifying against a branch service > - The benefit is that local plugin authors can confirm they are exercising the runtime they intend to, instead of debugging phantom plugin bugs caused by hitting the wrong server ## Linked Issues or Issue Description No public GitHub issue exists, so the underlying problem is described inline following the feature-request template. **Problem or motivation** Local plugin development assumes a single Paperclip on `http://127.0.0.1:3100`. When a plugin depends on server code that only exists on a feature branch (a new scoped route, a new response field, a new managed-resource capability), installing it into a long-lived host still on older code makes the route/field missing there. The plugin falls back or errors and *looks* broken, when the real cause is that it was tested against the wrong runtime. The CLI already let you point at any server, but it never surfaced which server you ended up on — so the mistake was invisible. **Proposed solution** Make the install target explicit. Before `plugin install` runs, probe `GET /api/health` and print the resolved API URL plus server status/version/deploymentMode/exposure, so the developer can confirm which Paperclip they are installing into. Add a standalone `plugin target` command to inspect the target without installing, a `--no-verify-target` escape hatch, and docs covering how to run a branch service on its own port and verify a branch route end-to-end. **Alternatives considered** - Do nothing and rely on the existing `--api-base` / `PAPERCLIP_API_URL` resolution — rejected because the gap was never the inability to point at a branch server, it was the lack of feedback about which server was actually hit. - Fail the install when the target looks stale — rejected as too aggressive; the probe is advisory and degrades gracefully when health details are not exposed or the server is unreachable. ## Dedup Search - [x] I searched the open and recently closed GitHub PRs for similar or duplicate PRs — this is not a duplicate ## What Changed - Add `probeTargetDiagnostics` / `formatTargetDiagnostics` helpers (`cli/src/commands/client/plugin.ts`) that read `GET /api/health` and report the resolved API URL plus server `status` / `version` / `deploymentMode` / `deploymentExposure`. - `plugin install` now prints these target diagnostics before installing, so you can confirm which instance you are installing into. Skippable with `--no-verify-target`. - `plugin install --json` keeps its original flat `PluginRecord` shape (top-level `id` / `pluginKey` / `version` / `status` are unchanged); when the target was probed it gains an additional top-level `target` field. Existing automation that reads the plugin fields keeps working. - Add a standalone `paperclipai plugin target` command to inspect the install target without installing anything. - Update `doc/plugins/LOCAL_PLUGIN_DEVELOPMENT.md`: how the CLI resolves its target, how to run a branch service on its own port and point the CLI at it explicitly, an end-to-end check that the branch route is actually served, and a troubleshooting entry for the stale-target symptom. - Unit tests for the diagnostics helpers (reachable + unreachable probe, and both render paths). ## Verification - `npx vitest run cli/src/__tests__/plugin-init.test.ts` — 10/10 pass (covers `probeTargetDiagnostics` success/failure and `formatTargetDiagnostics` rendering). - CLI typecheck (`tsc --noEmit` in `cli/`) — clean. - Manual: with a server running, `paperclipai plugin target` prints `Target Paperclip: <url>` and the health line; `plugin install` prints the same block before installing and `--no-verify-target` skips it. ## Risks Low risk. The probe is read-only (`GET /api/health`) and runs before install; if the server does not expose details it degrades to `ok (no details exposed)`, and an unreachable target prints a remediation hint rather than failing the command. The `--json` output keeps its original flat shape, so existing scripts are unaffected. No server or schema changes. ## Model Used Claude Opus 4.7 (`claude-opus-4-7`), extended thinking + tool use, via Claude Code. |
||
|
|
721541c41d |
Fix sandbox restore index drift (#8595)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandboxed agents can do useful work in a remote workspace, then export that result back to a host review checkout. > - A restore can leave the host checkout's index stale even when the sandbox content and committed state are correct. > - That drift makes reviewers see misleading missing-file or dirty-state results after sandbox execution. > - This pull request repairs the host index refresh path and surfaces a clean/dirty restore signal during finalization. > - The benefit is more reliable review checkouts and clearer sandbox finalization status. ## Linked Issues or Issue Description Refs #248 No exact public GitHub issue was found for this restore/index consistency fix. The underlying bug is that a sandbox restore can update host working-tree content without leaving the host git index aligned, so local review tooling may report stale deletions or misleading dirtiness. This PR keeps the restore/export path consistent and reports the resulting clean/dirty state. GitHub search performed for related or duplicate work: `sandbox runtime status`, `sandbox restore index`, and `runtime progress`. No direct duplicate PR was found. ## What Changed - Refreshed git workspace sync/index handling after sandbox restore/export operations. - Added finalization status that reports whether the restored host checkout is clean or dirty. - Preserved sandbox-managed runtime progress callbacks around restore and finalization. - Added adapter-utils tests covering the host index drift repair and clean/dirty finalization signal. ## Verification - Local PII scan before push: high-confidence secret patterns, internal issue links, local user paths, and private URL patterns checked across all three split diffs; no real secrets or internal links found. - `git diff --check feat/sandbox-runtime-status..fix/sandbox-restore-index-sync` - `pnpm exec vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — 1 file, 10 tests passed. - `pnpm run typecheck` passed on `fix/sandbox-restore-index-sync`. - `pnpm run build` passed on `fix/sandbox-restore-index-sync`; Vite reported existing CSS `::highlight` and chunk-size warnings. - `pnpm run test:run` was attempted on `fix/sandbox-restore-index-sync`; it failed in two unrelated broad-suite tests. One depends on this host's Git default branch behavior, and one depends on local Claude model-discovery environment. The changed focused suite above passes. ## Risks - Git index refresh behavior needs to remain conservative so it does not hide real uncommitted user changes. - The clean/dirty finalization signal is diagnostic; it should not by itself mark a sandbox run successful or failed. - This PR is stacked on the runtime-status branch so finalization progress can reuse the same callback path. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent, with shell/tool execution in a local worktree. Exact context-window metadata is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
800dab1c64 |
Render sandbox runtime status in issue threads (#8594)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue thread is where operators watch active agent work and decide whether a run is healthy. > - Backend runtime-progress plumbing can expose a concise current status, but users only benefit if the issue UI renders it in context. > - The UI should show that current status as part of the active run surface without turning it into durable transcript content. > - This pull request adds the issue-thread rendering and message formatting needed for sandbox runtime status. > - The benefit is that active sandbox setup phases are visible from the normal issue detail view. ## Linked Issues or Issue Description Refs #248 No exact public GitHub issue was found for this UI rendering work. The underlying problem is that active sandboxed runs can have backend progress state without any concise issue-thread display, leaving users to infer whether setup is still moving. This PR adds the UI layer on top of the runtime-status API branch. GitHub search performed for related or duplicate work: `sandbox runtime status`, `sandbox restore index`, and `runtime progress`. No direct duplicate PR was found. ## What Changed - Included current runtime status in the heartbeat API client shape. - Rendered active run status text in the issue chat thread when available. - Updated issue-chat message formatting helpers for runtime status display. - Added focused component and formatting tests for the new active-run status behavior. ## Verification - Local PII scan before push: high-confidence secret patterns, internal issue links, local user paths, and private URL patterns checked across all three split diffs; no real secrets or internal links found. - `git diff --check feat/sandbox-runtime-status..feat/sandbox-status-ui` - `pnpm exec vitest run ui/src/components/IssueChatThread.test.tsx ui/src/lib/issue-chat-messages.test.ts` — 2 files, 91 tests passed. - `pnpm run typecheck` passed on `feat/sandbox-status-ui`. - `pnpm run build` passed on `feat/sandbox-status-ui`; Vite reported existing CSS `::highlight` and chunk-size warnings. - Visual evidence (desktop + mobile) is attached to the tracking issue rather than committed to the repo. ## Risks - This PR is stacked on the runtime-status API branch and depends on that response/event shape. - The status line is intentionally concise; long or sensitive status content should continue to be bounded/redacted by the backend. - The visual change is limited to active issue-thread runs with current runtime status data. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent, with shell/tool execution in a local worktree. Exact context-window metadata is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip CTO <cto@paperclip.local> Co-authored-by: Paperclip CTO <noreply@paperclip.ing> |
||
|
|
a27e5ad002 |
Add ephemeral sandbox runtime status plumbing (#8593)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandboxed agent runs can spend meaningful time preparing a remote workspace before the agent transcript shows useful output. > - Operators need short, current progress text for those setup phases, but that text should not become durable run history. > - The existing live-run websocket path already carries run updates to the UI, so the backend can reuse that channel instead of adding polling. > - This pull request adds an ephemeral runtime-progress contract, a process-local status store, and heartbeat integration for sandbox-managed runs. > - The benefit is a clearer active-run experience without database migrations or persistent progress rows. ## Linked Issues or Issue Description Refs #248 No exact public GitHub issue was found for this status-message plumbing. The underlying problem is that active sandboxed runs currently have setup phases, such as workspace sync and restore, where the operator cannot see concise current progress through the live run state. This PR addresses that gap for the backend/runtime layer while keeping progress messages ephemeral. GitHub search performed for related or duplicate work: `sandbox runtime status`, `sandbox restore index`, and `runtime progress`. No direct duplicate PR was found. ## What Changed - Added shared runtime-progress types and the `heartbeat.run.progress` live event type. - Added a process-local heartbeat run runtime-status store with TTL, bounded/redacted messages, and terminal cleanup. - Threaded runtime progress callbacks through heartbeat execution and active/live run serialization. - Emitted sandbox-managed runtime phase updates for sync, adapter startup, restore/export, and finalization paths. - Added backend and adapter-utils tests for ephemeral status behavior, terminal cleanup, live serialization, and sandbox progress callbacks. ## Verification - `pnpm install --frozen-lockfile` - Local PII scan before push: high-confidence secret patterns, internal issue links, local user paths, and private URL patterns checked across all three split diffs; no real secrets or internal links found. The only secret-like text is an intentional fake test fixture (`sk-test-secret`). - `git diff --check origin/master..feat/sandbox-runtime-status` - `pnpm exec vitest run server/src/services/heartbeat-run-runtime-status.test.ts server/src/__tests__/heartbeat-runtime-state.test.ts server/src/__tests__/agent-live-run-routes.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — 4 files, 23 tests passed. - `pnpm run typecheck` passed on both top stacks that include this branch: `feat/sandbox-status-ui` and `fix/sandbox-restore-index-sync`. - `pnpm run build` passed on both top stacks that include this branch; Vite reported existing CSS `::highlight` and chunk-size warnings. - `pnpm run test:run` was attempted on `fix/sandbox-restore-index-sync`; it failed in two unrelated broad-suite tests. One depends on this host's Git default branch behavior, and one depends on local Claude model-discovery environment. The changed focused suites above pass. ## Risks - Runtime progress is process-local by design, so status disappears after TTL, terminal cleanup, or server restart. - Clients that do not consume `heartbeat.run.progress` simply keep existing behavior. - Message redaction is intentionally generic; overly specific phase details should stay out of runtime-progress payloads. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent, with shell/tool execution in a local worktree. Exact context-window metadata is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip CTO <cto@paperclip.local> Co-authored-by: Paperclip CTO <noreply@paperclip.ing> |
||
|
|
51ffbb380f |
Exclude transient Codex home dirs from sandbox sync (#8581)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters can run agents against sandboxed execution targets by syncing the workspace and selected runtime assets into the sandbox. > - The Codex local adapter includes a managed Codex home asset so sandboxed Codex runs can use the expected auth, config, skills, and session state. > - That asset follows symlinks, which is useful for real Codex home content but unsafe for transient launcher directories. > - Transient `tmp` and `.tmp` directories can contain symlinks to large host binaries, so the sandbox archive can inline large executable targets instead of just the small home directory content. > - This pull request excludes transient Codex home directories from the sandbox home asset while preserving the required Codex home files. > - The benefit is a much smaller and more predictable sandbox setup upload without changing the runtime files Codex actually needs. ## Linked Issues or Issue Description No public issue was found for this exact sandbox archive-size bug. Bug description: - What happened: sandboxed `codex_local` runs sync the managed Codex home as a `home` asset with `followSymlinks` enabled. If transient Codex home dirs such as `tmp` or `.tmp` contain symlinks to a large host binary, the archive can inline that binary and make `Syncing home to sandbox` much larger than the managed home directory itself. - Expected behavior: sandbox setup should include the Codex home files needed for auth, config, skills, and session continuity, but should not archive transient launcher scratch directories. - Reproduction shape: create a managed Codex home with normal auth/config/skills files and a `tmp/arg0` or `.tmp` symlink to a large host executable, then start a sandboxed `codex_local` run. The home asset archive grows by the symlink target size. - Version/commit: observed on local `master` before this change. - Related public context: #5028 covers a different managed Codex home reliability issue around stale auth files; this PR addresses sandbox archive bloat from transient symlink targets. ## What Changed - Excluded `tmp` and `.tmp` from the Codex `home` asset that is uploaded for sandboxed runs. - Added regression coverage proving transient symlinked home dirs are excluded from the tar while required auth/config/skills files remain included. - Kept `followSymlinks` behavior for the rest of the Codex home asset so existing non-transient symlink behavior is preserved. ## Verification - `git diff --check` - Local PII/secret pattern scan over the committed diff - `pnpm exec vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts` - `pnpm --filter @paperclipai/adapter-utils typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` ## Risks Low risk. The exclusion is limited to transient Codex home scratch directories, and the regression test verifies the files needed in the sandbox are still archived. The main compatibility risk is if a user intentionally placed required persistent Codex state under `tmp` or `.tmp`; those paths are treated as volatile scratch space by this change. ## Model Used OpenAI Codex coding agent based on GPT-5, with shell, git, and GitHub CLI tool use. Exact hosted model build and context-window size were not exposed in the local adapter runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ef37203a48 |
perf(ci): build standalone public packages concurrently (#8567)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - CI runs a Canary Dry Run job that exercises `release.sh`, which builds the standalone sandbox-provider packages for publish > - That step (`scripts/build-standalone-public-packages.mjs`) built the 7 provider plugins serially — each doing `rm -rf dist && tsc` — making it the dominant cost (~49s) inside the slowest PR check (~4.9m wall) after the general-server lane was already sharded > - The packages are independent (their own `node_modules` via `--ignore-workspace`, their own `dist`), so the serial build is pure latency with no correctness benefit > - This pull request builds them with a bounded-concurrency pool sized to the runner CPU count (overridable via `STANDALONE_BUILD_CONCURRENCY`), buffering each package's output and flushing it as one block so parallel logs stay readable, and aggregating failures by original index > - The benefit is a faster Canary Dry Run / PR feedback loop without changing what gets built or published ## Linked Issues or Issue Description No public GitHub issue exists. Inline feature/perf description: ### Problem or motivation `build-standalone-public-packages.mjs` builds standalone provider packages serially, making it the largest single cost inside the slowest PR check. ### Proposed solution Run independent per-package builds through a bounded-concurrency worker pool sized to runner CPU count, with an env override and readable buffered logs. ### Alternatives considered Keep the serial build for simpler logs, but that preserves the avoidable CI latency. ### Roadmap alignment This is CI maintenance and does not overlap planned core roadmap work. ## What Changed - `scripts/build-standalone-public-packages.mjs`: replaced the serial per-package build loop with a bounded-concurrency pool (default = runner CPU count, override via `STANDALONE_BUILD_CONCURRENCY`); per-package stdout/stderr is buffered and flushed as a single block; failures are aggregated by original package index so one failure neither aborts the others mid-flight nor obscures which package broke. - `scripts/__tests__/build-standalone-concurrency.test.mjs`: new `node:test` unit suite covering the pool (limit respected, all items run, ordered failure aggregation, env-override resolution). - `.github/workflows/pr.yml`: wired the new unit test into the policy job. ## Verification - `node --test ./scripts/__tests__/build-standalone-concurrency.test.mjs` → 6/6 pass - `node ./scripts/release-package-map.mjs check` → OK (29 enabled for CI publish) - `git diff --check origin/master..HEAD` → clean ## Risks - Low risk. Build inputs/outputs are unchanged; only scheduling differs. The concurrency is bounded by CPU count and overridable; output is buffered per package so logs remain attributable. If a package fails, all failures are still reported with their package index. ## Model Used - Claude (Anthropic), `claude-opus-4-8`, extended thinking with tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b3209486ca |
fix: default Daytona sandboxes to auto-archive so they leave the disk quota (#8561)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run their work inside sandboxes, provisioned through pluggable sandbox-provider plugins; the Daytona plugin is one of them > - Daytona bills storage against an org-wide disk quota, and a *stopped* sandbox still counts against that quota — only an *archived* sandbox is moved to cold storage and stops counting > - The Daytona plugin created sandboxes without supplying any auto-stop/auto-archive/auto-delete intervals, so Daytona fell back to its own defaults (archive after 7 days, never auto-delete) > - When in-product cleanup fails or never runs (crashed runs, failed lease destroys, orphaned probes), stopped sandboxes then sit for a week at a few GiB each until the org storage quota fills and blocks all workers > - This pull request makes the Daytona plugin apply quota-safe defaults on create (auto-stop 15m, auto-archive 60m, auto-delete 7d) so every sandbox eventually leaves the disk quota on its own, even when our own cleanup fails > - The benefit is that the storage quota no longer fills from leaked/idle sandboxes, while operators keep full control to override the intervals per environment ## Linked Issues or Issue Description No public GitHub issue. Describing in-PR (bug report): ### What happened Running many agents in Daytona sandboxes eventually exhausted the org's storage quota, and Daytona returned an out-of-disk error that blocked all sandbox creation. Root cause: a *stopped* Daytona sandbox still consumes disk quota; only an *archived* sandbox is moved to cold storage and frees it. The plugin created sandboxes without specifying auto-stop/auto-archive/auto-delete intervals, so Daytona used its built-in defaults (auto-archive after 7 days, auto-delete disabled). Any sandbox that our own cleanup failed to remove — or that was orphaned by a crashed run — therefore lingered for up to a week, accumulating until the quota filled. ### Expected behavior Idle/leaked sandboxes should leave the storage quota automatically within a short window, without relying solely on in-product cleanup succeeding. ### Steps to reproduce 1. Configure the Daytona sandbox provider with no explicit auto-stop/auto-archive/auto-delete intervals. 2. Run many agent sandboxes over time (or leak some via crashed/failed runs that skip in-product cleanup). 3. Observe stopped-but-not-archived sandboxes accumulating against the org's 30 GiB storage quota until Daytona returns an out-of-disk error and blocks new sandbox creation. ### Deployment mode Self-hosted Paperclip using the Daytona sandbox-provider plugin. ## What Changed - `daytona/src/plugin.ts`: `parseDriverConfig` now defaults `autoStopInterval` → 15 min, `autoArchiveInterval` → 60 min, `autoDeleteInterval` → 7 days when the value is unset. Explicit per-environment values (including `0` and `-1`) are preserved and passed through unchanged. - `daytona/src/manifest.ts`: documents the new defaults and the quota rationale on each field, and sets the manifest `default` so the values surface in the environment configuration screen where operators can override them. - `daytona/src/plugin.test.ts`: adds tests covering the new defaults and that explicit values / disabling sentinels (`0`, `-1`) are respected. ## Verification ``` pnpm --filter @paperclipai/plugin-daytona test ``` - 24 tests pass, including the new default-behavior tests. - Confirmed an explicit `autoArchiveInterval: 0` / `autoDeleteInterval: -1` in config is still forwarded unchanged (defaults only apply when the field is absent). ## Risks Low risk. Behavior change is limited to sandboxes created with no explicit interval config: they now auto-stop/archive/delete on Daytona-managed timers instead of Daytona's longer built-in defaults. Auto-archive is reversible (resuming an archived sandbox restores it). The only persistent action is auto-delete, which is a 7-day backstop for sandboxes nobody resumes, matching prior intent. Any environment that needs long-lived warm sandboxes can override the intervals (including disabling them with `0`/`-1`). ## Model Used Claude Opus (claude-opus-4-8), via the Paperclip agent harness with tool use. Extended reasoning enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ea1e321d86 |
Fix Daytona resource overrides for image-backed sandboxes (#8564)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Remote sandbox providers are part of the execution environment layer that lets agents run outside the local host. > - The Daytona sandbox-provider plugin exposes CPU, memory, disk, and GPU settings so operators can request larger execution sandboxes. > - The plugin was forwarding resource settings through snapshot/default creation, but Daytona rejects resource overrides on that path. > - Image-backed Daytona creation does support resource settings, so Paperclip should only send those values when an image is configured. > - This pull request makes the Daytona contract explicit in validation and runtime acquisition. > - The benefit is that operators get a clear Paperclip error for unsupported snapshot/default resource configs, while image-backed sandboxes still receive the requested allocation. ## Linked Issues or Issue Description No public GitHub issue exists. ### What happened? Daytona sandbox environments can be configured with CPU/memory/disk/GPU resource settings, but snapshot/default Daytona sandbox creation rejects those resource overrides. Paperclip could pass unsupported resource fields through to Daytona and surface an opaque provider error. ### Expected behavior Paperclip should only send resource fields on Daytona creation paths that support them, and it should reject unsupported resource configurations before creating a sandbox. ### Steps to reproduce 1. Configure a Daytona sandbox environment with resource values such as `cpu: 4` and `memory: 4`. 2. Leave the environment on snapshot/default creation by not setting an image. 3. Acquire a Daytona sandbox lease. 4. Observe Daytona reject the resource override on the snapshot/default path. ### Paperclip version or commit Reproduced against the current Daytona sandbox-provider plugin behavior before this PR. This PR fixes the contract in the plugin code and tests. ### Deployment mode Local authenticated Paperclip instance using the Daytona sandbox provider. Additional scope: - Daytona provider only. - No schema changes. - No non-Daytona provider changes. Related search performed: - `gh search issues "Daytona resources snapshot repo:paperclipai/paperclip" --limit 10` - `gh search prs "Daytona resources snapshot repo:paperclipai/paperclip" --limit 10` ## What Changed - Restrict Daytona `resources` create params to image-backed sandbox creation. - Add validation/runtime guard for resource settings without an image. - Keep image-backed resource metadata and reusable-lease sentinel behavior covered by tests. - Update Daytona plugin tests for image-backed resources and snapshot/default rejection. ## Verification - `pnpm -C packages/plugins/sandbox-providers/daytona test` - `pnpm -C packages/plugins/sandbox-providers/daytona build` - Manual Daytona SDK probe in a local authenticated instance: image-backed creation with `daytonaio/sandbox:0.8.0` returned a 4 CPU / 4 GiB sandbox and deleted cleanly. - Greptile Review: 5/5, no inline comments. ## Risks - Existing Daytona environments that set resource values while relying on snapshot/default creation will now fail validation/acquisition with a clear error instead of calling Daytona and failing there. - Operators who need custom resource sizes should use image-backed creation or a provider-side resource-sized snapshot workflow. - Low migration risk: no database schema changes and no changes to non-Daytona providers. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent. Exact context window is not exposed in this environment. Tool-enabled workflow with shell, GitHub CLI, database inspection, and local test/build execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
cd38c150b0 |
feat: reuse Daytona sandbox leases (#8513)
Add opt-in reusable Daytona sandbox lease support, including retryable pending cleanup handling.\n\nPR: https://github.com/paperclipai/paperclip/pull/8513 |
||
|
|
03362b347d |
Clean up agent config environment selector (#8504)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent configuration is where operators set identity, adapter behavior, and execution environment defaults. > - The environment override control should only appear when there is a meaningful choice beyond the always-available local environment, or when an existing saved override must remain visible so it can be inspected or cleared. > - The previous UI used an "Execution" header, explanatory helper copy, and verbose inherited-default wording that made the agent config feel noisier than necessary. > - This pull request narrows the selector visibility to real alternatives and forced Kubernetes mode, then updates the visible copy to match the environment-focused mental model. > - The benefit is a cleaner agent configuration surface that only asks operators to make an environment choice when that choice exists. ## Linked Issues or Issue Description No public GitHub issue exists for this change. Public GitHub searches for related issues and PRs returned no matches: - `agent config environment selector` - `execution environment agent config` Bug-style issue description: What happened? The agent config form could surface an environment override section even when the operator had no meaningful non-local execution environment choice. The section was labeled "Execution", included helper/inheritance text, and described the inherited local default as "Inherit instance default (Local)". Expected behavior: The selector should stay hidden unless there is more than one configured environment choice, the section should be labeled "Environment", helper text should be removed, and the inherited default option should read like `Default: Local`. Existing saved non-local overrides should remain visible even if the target environment is no longer runnable, so operators can inspect or clear the stale selection. Steps to reproduce: 1. Run Paperclip from `master` with environments enabled and only the implicit local default, or Local plus one runnable non-local environment. 2. Open an agent configuration form. 3. Inspect the environment override section visibility and copy. Paperclip version or commit: `0b945f449` / current `master` before this branch. Deployment mode: Local dev (`pnpm dev`). Installation method: Built from source. Agent adapters involved: Not adapter-specific. Database mode: Not database-related. Access context: Board operator UI. ## What Changed - Shows the environment override selector only for forced Kubernetes mode, at least one runnable non-local environment, or an existing saved non-local override. - Keeps Local out of the runnable environment count so Local-only setups do not show a redundant selector. - Preserves stale saved non-local overrides in the selector options so operators can see and clear them. - Renames the section header from "Execution" to "Environment". - Removes the helper/inheritance text above the selector. - Changes the inherited option copy from `Inherit instance default (...)` to `Default: ...`. - Adds focused render tests for Local-only hiding, Local plus one runnable non-local environment showing the concise selector copy, and stale non-runnable override recovery. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/AgentConfigForm.render.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - `git diff --check origin/master...HEAD` - Checked `ROADMAP.md` for overlapping planned core work. - Searched public GitHub issues and PRs for duplicate or related work; no matches found. ## Screenshots Before, when the selector was visible: ```text Execution Environment override Inheriting the instance default: Local. [Inherit instance default (Local)] ``` After, when Local plus a runnable non-local environment exists: ```text Environment Environment override [Default: Local] [E2B · sandbox] ``` After, when only Local is configured: ```text (no Environment section is rendered) ``` ## Risks Low risk. This is isolated to the agent config UI. The main behavioral risk is hiding the selector in an environment edge case; the current branch keeps forced Kubernetes mode, runnable non-local environments, and pre-existing saved non-local overrides visible. The added render tests cover the Local-only, single non-local alternative, and stale override cases. ## Model Used OpenAI Codex local adapter, GPT-5-based Codex coding session. The exact hosted backend model ID is not exposed in the runtime. Tool use and local shell execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0b945f449b |
Make streamlined sidebar default to on (#8496)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board UI has an experimental streamlined left navigation mode that changes how projects and agents appear in the sidebar. > - Today that mode is opt-in, so users keep seeing the classic navigation unless they find and enable the experiment. > - The requested product behavior is to make streamlined navigation the default experience while keeping an explicit experiments opt-out. > - This pull request flips the shared/server default, updates UI consumers to treat only explicit `false` as classic mode, and covers both the default-on path and opt-out behavior in tests. > - The benefit is that new and legacy settings get the streamlined sidebar by default without removing the classic-sidebar escape hatch. ## Linked Issues or Issue Description Refs #7645 Related: #8430 takes the broader route of removing the classic sidebar. This PR intentionally keeps the opt-out path. ## What Changed - Default `enableStreamlinedLeftNavigation` to `true` in shared validation and server-side normalization. - Preserve explicit stored `false` as the experiments opt-out for the classic sidebar. - Render the sidebar and experimental settings toggle as streamlined-on unless the setting is explicitly `false`. - Add regression coverage for loading/default streamlined sidebar behavior and the opt-out patch from the experiments page. - Remove internal issue identifiers from newly touched source comments before publishing. ## Verification - `git diff --check origin/master...HEAD` — passed. - `pnpm -r typecheck` — passed. - `pnpm exec vitest run server/src/__tests__/instance-settings-service.test.ts ui/src/components/Sidebar.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx` — passed, 3 files / 27 tests. - `pnpm build` — passed, with existing Vite/CSS/chunk-size warnings. - `pnpm test:run` — failed in unrelated `server/src/__tests__/workspace-runtime.test.ts`: the test `auto-detects the default branch via symbolic-ref when origin/HEAD is set` creates a temp repo on `main` then runs `git push -u origin main master`; `master` does not exist in that temp repo. Summary: 1 failed, 214 passed, 1780 tests passed, 1 skipped. ## Risks Low-to-medium behavioral risk: the default sidebar changes for users who never explicitly set the experiment. Explicit `false` remains respected, so users can still opt out via experimental settings. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex local adapter, with tool use and code execution. Exact context window was not surfaced in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d1e6662ed8 |
fix(server): let codex_local agents inherit host Codex login by default (#8425)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `codex_local` adapter spawns the Codex CLI for an agent, and `server/src/routes/agents.ts` normalizes each agent's `adapterConfig.env` on create/hire/update > - PR #8272 added an isolation guard that, on every codex_local agent, force-set a per-agent `CODEX_HOME` and injected `OPENAI_API_KEY = ""`, and rejected any "shared" home > - This broke the common case: operators who deleted `CODEX_HOME` / `OPENAI_API_KEY` in the UI saw them silently re-appear on save, and every agent was forced into an isolated home instead of sharing the device's existing Codex login (`~/.codex` / `$CODEX_HOME`) > - This pull request replaces the always-on isolation guard with key-scoped isolation: a keyless agent gets no env overrides and inherits the host Codex login at runtime; we only carve out an isolated per-agent `CODEX_HOME` when the agent sets its own `OPENAI_API_KEY` > - The benefit is that env-var deletion now persists, agents on one device share the host login by default, and per-account isolation is still available by setting a per-agent key ## Linked Issues or Issue Description No public GitHub issue exists. Bug report (following `bug_report.yml`): **What happened?** In a `codex_local` agent's configuration, removing the `CODEX_HOME` and `OPENAI_API_KEY` env vars via the UI (clicking the X) appears to work, but on save they instantly re-appear. The persisted config never loses the slots. **Root cause.** `applyCodexLocalIsolationGuard` in `server/src/routes/agents.ts` (added in #8272) re-injected `CODEX_HOME` (defaulting to a per-agent home) and `OPENAI_API_KEY = ""` on every create/hire/update, and rejected shared homes outright. So a PATCH that omitted those keys had them written back server-side. **Expected behavior.** Deleting these env vars should persist. A codex_local agent with no key should inherit whatever Codex login already exists on the device. **Steps to reproduce.** 1. Open a `codex_local` agent's configuration with `CODEX_HOME` and `OPENAI_API_KEY` set. 2. Remove both env vars and save. 3. Re-open the config — both slots are back. Related (different approach): Refs #8399 (keeps per-agent isolation, fixes only the empty `OPENAI_API_KEY` slot), Refs #8272 (introduced the guard), Refs #8403 (managed-auth seeding into isolated homes). ## What Changed - `server/src/routes/agents.ts` — replaced `applyCodexLocalIsolationGuard` (+ `assertCodexLocalHomeIsNotShared` / `normalizeCodexLocalHomePath`) with `applyCodexLocalKeyIsolation`. A codex_local agent now receives **no** env overrides unless it explicitly sets `OPENAI_API_KEY`; only then is an isolated per-agent `CODEX_HOME` injected (and only if the agent has not set its own `CODEX_HOME`). Keyless agents fall back to the host Codex login at runtime. Dropped the now-unused `node:os` import and the shared-home rejection. - `server/src/__tests__/agent-adapter-validation-routes.test.ts` — updated the validation-route tests to assert env-var deletion persists, keyless agents get no overrides, and key-bearing agents still get an isolated `CODEX_HOME`. ## Verification ```bash cd server && npx vitest run src/__tests__/agent-adapter-validation-routes.test.ts ``` All 6 tests pass locally. Manually verified in the live UI: removing `CODEX_HOME` and `OPENAI_API_KEY` from a codex_local agent and saving now persists the deletion (slots no longer re-appear on reload). ## Risks - **Host-key leak for keyless agents on a host with `OPENAI_API_KEY` set.** Low/intended: the product direction here is that codex_local agents inherit the host's Codex login by default; per-account isolation is opt-in via a per-agent `OPENAI_API_KEY`, which then gets its own `CODEX_HOME`. - **Divergence from #8399.** That PR retains the per-agent isolation default and only stops injecting the empty key, so it does not address the `CODEX_HOME` re-injection half of this bug. This PR intentionally changes the default to host-login inheritance. Reviewers should pick one direction. - **No migration.** Existing agents that already store these slots are not auto-cleaned, but operators can now delete them and the deletion sticks. ## Model Used - Provider: Anthropic (Claude) - Model ID: `claude-opus-4-8` - Capabilities: tool use, code execution, extended reasoning, agentic workflow ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2c98c8e1e5 |
Fix sandbox git publishing and large workspace uploads (#8422)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox and SSH runtimes need to preserve agent work across isolated execution environments > - Git-backed workspaces were being copied mostly as filesystem archives, which breaks when `.git` points outside the mounted workspace and makes sandbox agents unable to publish their own branches > - Large ignored dependency trees could also be swept into the sandbox overlay, causing multi-GB transfers and max-string failures in some sandbox clients > - This pull request makes sandbox runtime setup use a git-backed HEAD sync plus a small dirty/untracked overlay, and bounds sandbox file transfers so large archives do not need one huge string > - The benefit is that sandbox agents can commit and push from a usable git checkout without uploading dependency trees such as `node_modules` ## Linked Issues or Issue Description Refs #8395 No public duplicate issue or PR was found after searches for `sandbox git workspace`, `git push sandbox`, and `node_modules sandbox upload`. Bug report: - What happened: sandbox-backed agent workspaces could receive a `.git` file that pointed at host-only git state, leaving the sandbox unable to run normal git workflows. The sandbox overlay upload could also include ignored dependency directories, creating very large transfers. - Expected behavior: sandbox and remote runtimes should prepare a usable git-backed workspace, copy only the necessary workspace overlay, and restore git history plus file changes without depending on a host-only `.git` path. - Steps to reproduce: 1. Run an agent in a sandbox-backed workspace whose local git checkout is a worktree. 2. Ask the agent to complete a GitHub workflow that requires commit/push access. 3. Observe that git operations can fail inside the sandbox, and ignored dependency trees can be uploaded as part of the workspace overlay. - Paperclip version or commit: reproduced against `master` before this PR, base `7aa212296eb1`. - Deployment mode: local dev / sandbox-backed runtime. - Installation method: built from source. - Agent adapters involved: local adapters using shared adapter-utils runtime preparation. - Database mode: not database-related. - Access context: agent runtime. - Local verification environment: Node.js v25.6.1, pnpm 9.15.4, macOS arm64. - Privacy checklist: all pasted output was reviewed for secrets, private hostnames, local usernames, and internal instance links. ## What Changed - Added a GitHub workflow push preflight so agent runs can detect missing push credentials when a workflow explicitly needs GitHub publishing. - Added shared git workspace sync helpers for shallow HEAD import/export and dirty/untracked overlay tracking. - Updated sandbox managed runtime setup to use git history plus a selected overlay instead of uploading the full local workspace for git-backed workspaces. - Bounded sandbox archive upload/download paths so large payloads stream or chunk instead of materializing one oversized string. - Excluded `.git` and ignored dependency trees from sandbox upload, download, and restore baselines while preserving local ignored directories during sync-back. - Added focused tests for git workspace sync, sandbox overlay selection, transfer chunking, and heartbeat push-preflight behavior. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts packages/adapter-utils/src/command-managed-runtime.test.ts server/src/__tests__/heartbeat-project-env.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts` - `pnpm --filter @paperclipai/adapter-utils typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm build` - Public-hygiene scan of the PR diff and commit messages for internal issue ids, local paths, private hostnames, and obvious token patterns. ## Risks - Medium risk: this changes sandbox runtime synchronization semantics for git-backed workspaces, especially around dirty tracked files, untracked files, deleted paths, and ignored files. - The main mitigation is focused test coverage for upload contents, restore exclusions, and git round-trip behavior. - The SSH runtime keeps the current bundle-based implementation from `master`; this PR only aligns shared excludes and sandbox behavior with that model. ## Model Used OpenAI Codex, GPT-5-based coding agent, tool-enabled shell/git/GitHub workflow, with code execution and repository inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
33353ce62b |
feat(skills): remove bundled paperclip-dev skill and retire required skill attribute (#7029)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local adapters (Claude, Codex, Cursor, Gemini, Grok, OpenCode, Pi, ACPX) ship bundled "skills" — opinionated Markdown prompt bundles materialized into the agent's runtime > - One of those bundled skills, `paperclip-dev`, existed to let agents develop Paperclip itself; it has now moved to its own external repo and no longer belongs in the core tree > - The adapter skill model also carried a `required` / `requiredReason` attribute plus a `paperclip_required` `AdapterSkillOrigin` variant, all of which only existed to mark bundled skills as non-optional in the UI and adapter sync logic > - With `paperclip-dev` gone, no bundled skill is "required" anymore, and the type / runtime surface for `required` is dead weight — but it is computed at request time and never persisted, so a clean removal is safe (no compatibility shim needed) > - This pull request deletes `skills/paperclip-dev/` and removes every trace of the `required` / `requiredReason` field and the `paperclip_required` origin across shared types, validators, adapter-utils, all eight local adapters, server routes, the company-skills service, the UI, the storybook fixtures, and the test suite > - The benefit is a smaller, simpler adapter-skill surface: one origin (`company_managed`) for managed bundled skills, `resolvePaperclipDesiredSkillNames` collapses to "just the configured desired set", and the AgentDetail skills tab no longer renders a "Required by Paperclip" section that no longer applies ## Linked Issues or Issue Description <!-- No existing public GitHub issue; describing the underlying work inline (feature_request template fields). --> **Summary** Remove the bundled `paperclip-dev` skill (now maintained in its own external repo) and retire the `required` / `requiredReason` skill attribute and the `paperclip_required` skill origin, which only existed to support it. **Problem or motivation** `paperclip-dev` is the only bundled skill that was ever marked "required". Now that it lives in a separate repository, shipping it inside the core tree is wrong, and the entire `required` surface (a type field, a validator field, a synthesized `paperclip_required` origin, UI "Required by Paperclip" section, and required-skill merging in the desired-skills calculation) becomes dead weight. The `required` value is computed at request time and never persisted, so it can be removed cleanly without a migration or compatibility shim. **Proposed solution** Delete `skills/paperclip-dev/`, drop the `required` / `requiredReason` fields and `paperclip_required` origin everywhere they are produced or consumed, collapse managed-skill origin to a single `company_managed` value, and simplify `resolvePaperclipDesiredSkillNames` to return only the configured desired set. **Alternatives considered** Keeping the `required` attribute as a no-op for forward compatibility — rejected because it is request-time only (nothing persists it), so leaving it in place is pure dead surface area with no callers. **Roadmap alignment** Internal cleanup / dead-code removal that simplifies the adapter-skill surface; it does not introduce or duplicate any planned core feature in ROADMAP.md. ## What Changed - Deleted bundled `skills/paperclip-dev/` (moved to a separate repo). - Dropped `required`, `requiredReason`, and the `paperclip_required` origin from `packages/shared/src/types/adapter-skills.ts`, `packages/shared/src/validators/adapter-skills.ts`, and `packages/adapter-utils/src/types.ts`. - In `packages/adapter-utils/src/server-utils.ts`: removed `readSkillRequired()`; dropped `required`/`requiredReason` from `listPaperclipSkillEntries()`, `normalizeConfiguredPaperclipRuntimeSkills()`, `buildPersistentSkillSnapshot()`, and `PaperclipSkillEntry`; collapsed `buildManagedSkillOrigin()` to always return `company_managed`; simplified `resolvePaperclipDesiredSkillNames()` to return only the configured desired set (signature preserved so adapter call sites are untouched). - Walked all eight local adapters (`acpx-local`, `claude-local`, `codex-local`, `cursor-local`, `gemini-local`, `grok-local`, `opencode-local`, `pi-local`) and removed every remaining `requiredReason` / `paperclip_required` reference. - `server/src/services/company-skills.ts`: dropped the `required = sourceKind === "paperclip_bundled"` synthesis when listing runtime skill entries. - `server/src/routes/agents.ts`: removed required-skill merging from the desired-skills calculation in the persist-config path and the unsupported-snapshot path (keeping the current version-aware `desiredSkillEntries` structure). - `ui/src/pages/AgentDetail.tsx`: dropped required-based filters, the required tooltip, and the entire "Required by Paperclip" section from the agent skills tab; storybook fixtures in `ui/storybook/stories/acpx-local.stories.tsx` cleaned up to match. - Tests: deleted the `required: false` case in `paperclip-skill-utils.test.ts` and the "keeps required bundled skills installed" case in every `*-local-skill-sync.test.ts`; `acpx-local-execute.test.ts`, `cursor-local-execute.test.ts`, `cursor-local-skill-sync.test.ts`, `agent-skills-routes.test.ts`, and `packages/adapter-utils/src/server-utils.test.ts` were updated to drop removed fields and map `origin: "paperclip_required"` → `"company_managed"`. - `server/src/adapters/registry.ts`: two `as unknown as ServerAdapterModule["..."]` casts on `hermesListSkills` / `hermesSyncSkills` (matching the existing `executeHermesLocal` pattern). `hermes-paperclip-adapter@0.2.0` still depends on the published `@paperclipai/adapter-utils` which keeps the retired `paperclip_required` variant; the cast bridges the workspace-vs-published type mismatch at the registry seam and can drop once hermes upgrades. ## Verification Run from the workspace root: ```sh grep -rn "skills/paperclip-dev" . grep -rn "paperclip_required" --include="*.ts" --include="*.tsx" . grep -rn "requiredReason" --include="*.ts" --include="*.tsx" . pnpm -w typecheck pnpm --filter @paperclipai/server exec vitest run paperclip-skill-utils pnpm --filter @paperclipai/server exec vitest run skill-sync ``` The first three greps return only the explanatory comment in `server/src/adapters/registry.ts` (no live `paperclip_required` / `requiredReason` usage) and zero `skills/paperclip-dev` source hits. Locally: - `pnpm -w typecheck` → all packages this PR touches pass (adapter-utils, shared, server, ui, cli, and the cursor/gemini/opencode/pi adapters). - Affected vitest suites pass: `paperclip-skill-utils`, `server-utils`, all eight `*-local-skill-sync`, `agent-skills-routes`, and the `acpx`/`cursor`/`pi` execute suites. ## Risks - Behavioral shift in the agent skills UI: the "Required by Paperclip" section disappears. No bundled skill is required anymore, so this only affects environments that previously surfaced `paperclip-dev` as a forced-on row; those installs will see the skill move into the regular "company-managed" list (and be uninstalled on next sync unless explicitly listed as desired). - Existing agents may still have the string `"paperclip-dev"` in their persisted `desiredSkills`. That entry is inert (no source for it to install from); a one-time DB cleanup is out of scope. Low risk. - Hermes adapter type bridge: two casts in `registry.ts` paper over a type-only divergence between the workspace `@paperclipai/adapter-utils` and the published version still pinned by `hermes-paperclip-adapter@0.2.0`. Runtime behavior is unaffected because the retired `paperclip_required` value is no longer produced by anything in this tree. The casts can be removed once hermes upgrades its dependency. ## Model Used - Provider: Anthropic - Model: Claude Opus 4.7 (`claude-opus-4-7`) - Capability: agent tool use via Paperclip's `claude_local` adapter ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5ebbe67ebd |
fix(workspace-runtime): base fresh worktrees on origin/master and refresh unstarted reuses (#8412)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent work runs inside git worktrees created by the workspace runtime, each branched from a configured base ref (typically the repo's default branch) > - When the local `master` is stale or ahead of `origin/master` (committed but unpushed work, or local-only commits), a freshly created worktree inherits that divergence — so an unrelated task branch silently carries commits it never intended to touch > - This surfaced as a docs-only task whose PR accidentally pulled in unrelated changes from a diverged local master > - The base for a fresh worktree should be resolved authoritatively to the remote-tracking ref (`origin/<branch>`), and an idle/unstarted reused worktree should be safely fast-forwarded — without ever destroying in-progress work > - This pull request makes both behaviors explicit in the workspace runtime > - The benefit is that task branches start from a clean, authoritative base, eliminating accidental inclusion of unrelated local changes ## Linked Issues or Issue Description This is a bug fix. No public GitHub issue exists, so describing it inline following the Bug Report template: **What happened** A task intended to change only docs produced a PR that also contained unrelated changes pulled in from `master`. The root cause: when a new worktree is created from a configured local branch (e.g. `master`), the worktree inherits whatever that local branch points at. If the local `master` has committed divergence from `origin/master` (unpushed or local-only commits), that divergence leaks into the new task branch. The leak comes from committed local-vs-`origin/master` ref drift, not uncommitted working-tree changes (each worktree has its own working tree). **Expected behavior** A fresh worktree should be based on the authoritative `origin/master` head so unrelated local commits never seed a task branch. **Steps to reproduce** 1. Have a local `master` that is ahead of `origin/master` (committed but unpushed work). 2. Create a new worktree/task branched from `master` via the workspace runtime. 3. Open a PR from that branch — it carries the unrelated local commits. **Deployment mode** Self-hosted / local workspace runtime. ## What Changed - Fresh worktrees now resolve their base ref authoritatively: a configured local branch (e.g. `"master"`) is mapped to its `origin/<branch>` remote-tracking counterpart so unpushed/ahead local commits can never seed a task branch. Remote-tracking refs, SHAs, and tags are used verbatim; an unset/`HEAD` base falls back to the detected default branch. The resolved ref is recorded (`repoRef`) so downstream drift checks stay accurate. - If a configured local branch has no matching `origin/<branch>`, the runtime warns and falls back to the local ref rather than failing. - On reuse, a *provably unstarted* worktree (no commits past base + clean tree including untracked files) is fast-forwarded to the latest `origin/master`. Started or dirty worktrees keep the prior warn-only behavior, so in-progress work is never reset. Only remote-tracking bases are eligible for the refresh. ## Verification - `cd server && npx vitest run src/__tests__/workspace-runtime.test.ts` - 5 new tests cover: local-branch→`origin/<branch>` mapping, no-remote fallback warning, unstarted-reuse fast-forward, and that started/dirty worktrees are left untouched. - Result: 65 passed. 1 pre-existing failure (`auto-detects the default branch via symbolic-ref when origin/HEAD is set`) is unrelated to this change and fails only due to the test host's git default-branch config (test setup runs `git push -u origin main master` but the local default branch is `main`); it also fails on `master`. ## Risks - Low risk. The refresh path is intentionally conservative: it only fast-forwards worktrees that are provably unstarted (zero commits past base and a fully clean tree, including untracked files) and only when the base is a remote-tracking ref. Started or dirty worktrees fall through to the existing warn-only drift behavior, so no in-progress work can be destroyed. - Behavioral shift: fresh worktrees configured against a local branch will now base on `origin/<branch>` instead of the local ref. This is the intended fix; the only case it changes is when local and remote have diverged. ## Model Used Claude — `claude-opus-4-8` (extended thinking, tool use / code execution via Claude Code). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (N/A — no UI changes) - [x] I have updated relevant documentation to reflect my changes (N/A — no doc changes needed) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI run) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a937b89a47 |
fix(daytona): valid memory input in env config form + size presets (#8389)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments are configured through a JSON-schema-driven form (`JsonSchemaForm`), and each sandbox provider (here, Daytona) supplies a manifest describing its fields > - In the Daytona environment config form, the "memory" field rendered as a free-text input; on save the value round-tripped to `0`, producing a server-side "number must be greater than zero" error even when the user typed a valid number like `2` > - Rather than patch the free-text input, memory is now a true dropdown of the sandbox sizes Daytona actually supports, so an invalid value can no longer be typed or coerced > - The dropdown leads with a blank "None" row that is selected by default (meaning "not configured — use Daytona's defaults"); `0` is never offered because it is not a valid configuration > - The benefit is that operators pick a valid memory size from a constrained list, it saves correctly as an integer, and leaving it unset cleanly omits the field instead of submitting `0` ## Linked Issues or Issue Description No public GitHub issue exists; describing the bug inline per the bug report template. ### What happened? In the Daytona environment configuration form, the memory field is a free-text input labelled "gigabytes of RAM". Entering a value such as `2` and saving coerced the value to `0`, and the server rejected the save with "the number must be greater than zero". There was no way to enter a valid memory size through the form. ### Expected behavior Selecting a valid memory size (e.g. `2`) keeps that value and saves cleanly as an integer. Leaving memory unset is valid and submits no value (Daytona defaults apply). `0` is never selectable. ### Steps to reproduce 1. Open the environment configuration form for a Daytona sandbox provider. 2. In the memory field, type `2`. 3. Click Save. 4. Observe the value becomes `0` and the form errors with "the number must be greater than zero". ## What Changed - Daytona manifest: `memory` is now an `enum` of the supported sandbox sizes `[1, 2, 4, 8]` (GiB). It stays optional, so "not configured" remains valid. `0` is not in the list. - `JsonSchemaForm` `EnumField`: optional enums now render a leading blank **None** row that is selected by default when no value is set, letting the user express "not configured" and clear a previous selection (Radix `Select` forbids an empty-string item value, so the unset state maps to a sentinel that translates back to `undefined`). - `JsonSchemaForm` `EnumField`: when every enum option is numeric, the selected value is coerced back to a number on change so the payload keeps the schema's integer type (a stringified `"2"` would otherwise fail server-side integer validation — this is what fixes the original bug for the dropdown path). - Tests: added `EnumField` coverage in `JsonSchemaForm.test.tsx` (blank row present, no `0`, blank selected by default, numeric coercion, blank → unset) and Daytona manifest coverage in the plugin test (memory enum is `[1,2,4,8]`, excludes `0`, stays optional). ## Verification - `cd ui && npx vitest run src/components/JsonSchemaForm.test.tsx src/pages/CompanyEnvironments.test.tsx` → 16/16 passing (14 + 2). - `vitest run` in `packages/plugins/sandbox-providers/daytona` → 19/19 passing (incl. 3 new manifest tests). - `cd ui && npx tsc --noEmit` → clean. - Behavior: the memory field now renders as a dropdown showing **None / 1 / 2 / 4 / 8**, with **None** selected by default. Picking `2` stores the integer `2`; picking **None** clears the field so it is omitted from the payload (no `0`, no "must be greater than zero" error). ## Risks - Low risk. The blank-row + numeric-coercion changes live in `EnumField`. The blank row is only added for **optional** enums (required enums are unaffected); numeric coercion only triggers when every option is numeric, so existing string enums (`egressMode`, `backend`, `sessionStrategy`) are unchanged. The manifest change is scoped to the Daytona provider. ## Model Used Claude (Anthropic), Opus-class model, used via the Paperclip agent workflow (author + reviewer agents) with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
254b01d2af |
docs: drop contributor screenshot requirement in favor of cutter.sh bot (#8409)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Contributors follow `CONTRIBUTING.md` and the PR template when opening pull requests > - Those docs still tell contributors to manually capture and attach before/after screenshots for UI changes > - That requirement is now redundant: the `cutter.sh` bot automatically captures and posts before/after UI screenshots to every PR > - Asking contributors to also do it by hand adds friction and produces duplicate screenshots > - This pull request removes the manual screenshot obligation from `CONTRIBUTING.md` and the PR template, and documents that `cutter.sh` handles it > - The benefit is a clearer, lower-friction contribution flow that matches how screenshots actually get posted today ## Linked Issues or Issue Description No public GitHub issue. Underlying change: Our contributor docs and PR template instructed contributors to attach before/after screenshots for any UI change. We now run a bot (`cutter.sh`) that automatically captures and posts before/after UI screenshots to the PR, so the manual step is no longer needed. This change removes the manual requirement and instead documents that the bot handles screenshots, asking contributors to simply describe the visible change. ## What Changed - `CONTRIBUTING.md`: removed the "Before / After screenshots" requirement from the Path 2 PR checklist and from the "Writing a Good PR message" section; both now explain that the `cutter.sh` bot posts UI screenshots automatically and ask contributors to describe the visible change instead. - `.github/PULL_REQUEST_TEMPLATE.md`: updated the Verification guidance to note that `cutter.sh` posts before/after screenshots automatically, and removed the "include before/after screenshots" item from the checklist. - Issue templates were intentionally left unchanged — they only offer screenshots as an optional "if helpful" field, never a requirement. ## Verification - Docs-only change; no code paths affected. - `grep -rin "screenshot" CONTRIBUTING.md .github/` confirms no screenshot *requirement* remains — only the new `cutter.sh` wording and the optional "if helpful" fields in issue templates. - Render `CONTRIBUTING.md` and the PR template to confirm the wording reads cleanly. ## Risks Low risk — documentation-only change. No build, runtime, or migration impact. ## Model Used Claude (Anthropic), model `claude-opus-4-8`, extended thinking with tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (docs-only; no tests apply) - [ ] I have added or updated tests where applicable (N/A — docs-only) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
67ca8287e8 |
fix(codex-local): seed managed auth into isolated homes (#8403)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `codex_local` adapter isolates local Codex runs by assigning managed `CODEX_HOME` state per company and per agent > - PR #8272 tightened that isolation, but it left a gap: once `CODEX_HOME` became explicit, the adapter treated it like a user-managed override and skipped auth seeding > - That meant newly isolated agents could launch with no usable `auth.json`, hit OpenAI unauthenticated, and fail with `401 Missing bearer` > - Users who had already persisted one of those broken managed homes could remain stranded even after config changes unless Paperclip repaired the home itself > - This pull request teaches Paperclip to seed managed homes correctly, backfill already-stranded managed homes on startup, and reject credential-less managed homes before they reach the provider > - The benefit is that affected managed `codex_local` agents recover automatically after upgrade and restart, without manual `CODEX_HOME` surgery ## Linked Issues or Issue Description - Fixes #497 - Refs #5028 - Related PR: #8272 - Related PR: #8399 ## What Changed - Distinguished Paperclip-managed `CODEX_HOME` paths from genuine external overrides and always seeded auth into managed homes, even when `CODEX_HOME` is explicit in config. - Wrote API-key-backed `auth.json` files for managed homes when `OPENAI_API_KEY` is configured, otherwise symlinked the shared Codex auth for subscription/OAuth flows. - Added a startup reconciliation pass that backfills already-isolated managed homes created by the broken release so upgrade plus restart repairs stranded agents automatically. - Preserved previously resolved API-key auth when the stored `OPENAI_API_KEY` binding is secret-backed and startup cannot resolve the secret value directly. - Hardened the managed-home preflight to require a credential-bearing `auth.json`, not just file presence, and documented the recovery behavior. - Added regression tests covering managed-home seeding, fail-fast behavior, and server-side startup reconciliation. ## Verification ```bash pnpm exec vitest run packages/adapters/codex-local/src/server/codex-home.test.ts packages/adapters/codex-local/src/server/execute.auth.test.ts server/src/__tests__/codex-auth-reconciliation.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/server-startup-feedback-export.test.ts pnpm --filter @paperclipai/adapter-codex-local typecheck pnpm --filter @paperclipai/server typecheck pnpm check:tokens pnpm exec vitest run server/src/__tests__/server-startup-feedback-export.test.ts ``` GitHub Actions `PR` workflow is green on latest head `e5b1e08d4`, including policy, typecheck, test shards, build, e2e, serialized server suites, and canary dry run. Greptile Review is green on latest head with 0 comments added. No UI changes. ## Risks - Startup reconciliation now mutates persisted managed Codex homes at boot. Risk is low because it only touches Paperclip-managed company/agent home paths and no-ops when a home already has usable auth. - Genuine external `CODEX_HOME` overrides remain intentionally self-managed, so those users still own repair steps inside their custom home. - Hosts with neither shared Codex auth nor an explicit per-agent API key now fail earlier with a clearer adapter error instead of surfacing a downstream `401`, which changes timing but not capability. ## Model Used - OpenAI Codex, GPT-5-based coding agent in a local Codex session; exact served model ID/context window were not exposed to the session. Tool use and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e281095cb4 |
fix(ui): make adapter test button test only (#8405)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapter configuration is where operators edit the runtime settings that control how each agent is invoked. > - The adapter form includes a test action so operators can validate the in-progress configuration before saving it. > - In edit mode, that test action changed into `Save + Test` when the form was dirty, which mixed validation with persistence. > - Operators already have an explicit Save action for committing adapter changes. > - This pull request makes the adapter test action test-only so validation never implicitly saves a dirty draft. > - The benefit is clearer operator control: Test validates the current draft, Save persists it. ## Linked Issues or Issue Description No matching public GitHub issue or in-flight PR was found, so this PR describes the bug inline. ### Pre-submission checklist - I searched existing open and closed issues and did not find a duplicate. - I can reproduce this behavior on current `master` before this branch. - I confirmed this originates in the Paperclip UI, not in a specific local agent CLI or provider configuration. ### What happened? When editing an existing agent adapter configuration, changing an adapter setting made the adapter environment test button change from `Test` to `Save + Test`. Clicking it saved dirty edits before running the environment test. ### Expected behavior The test action should always be a test action. Dirty draft values should be validated without persisting them, and the explicit Save button should be the only save path. ### Steps to reproduce 1. Open an existing agent adapter configuration. 2. Edit an adapter setting so the form becomes dirty. 3. Observe the adapter environment test action label and behavior. ### Paperclip version or commit Current `master` before this branch, `c0743482b`. ### Deployment mode Local dev/private operator UI. ### Installation method Built from source with pnpm. ### Agent adapter(s) involved Not adapter-specific. The behavior lives in the shared agent adapter config form. ### Database mode Not database-related. ### Access context Board operator UI. ### Additional context Public searches performed before opening this PR: - `adapter test save` - `Save + Test` - `AgentConfigForm testEnvironment` - `adapter configuration test button` No logs or local config are needed for this UI-only behavior report. ## What Changed - Removed the helper that converted dirty edit-mode adapter tests into `Save + Test`. - Removed the save-before-test helper path so clicking Test only invokes the adapter environment test mutation. - Kept draft testing behavior by continuing to build the test payload from the current in-progress adapter config. - Dropped unit tests that asserted the old save-and-test behavior. ## Verification - `pnpm vitest run ui/src/components/AgentConfigForm.test.ts` -> 6/6 passed. - `pnpm --filter @paperclipai/ui typecheck` was attempted locally and failed on the existing unrelated `packages/plugins/sdk/src/ui/components.ts` React module-resolution error before reaching this change. - Static review: `Save + Test`, `getAgentConfigTestActionLabel`, and `runAgentConfigEnvironmentTest` no longer appear in `ui/src`. - Before/after visual state for the dirty edit-mode adapter test action: - Before: `Save + Test` - After: `Test` - PR CI passed: policy, commitperclip review, typecheck/release-registry, build, general tests, serialized server suites, e2e, canary dry run, aggregate verify, and security scans. - Greptile passed with 5/5 confidence. The remaining screenshot/documentation thread was answered and resolved because this is a text-only button-label state already captured above. ## Risks Low risk. This is a narrow UI behavior change in the adapter config form. The main behavior change is intentional: testing a dirty adapter draft no longer persists it. If an operator expected Test to save edits, they must now use the explicit Save button. ## Model Used - Anthropic Claude via Paperclip `claude_local` produced the initial implementation commit. - OpenAI GPT-5 Codex via Paperclip `codex_local` reviewed the change, amended public commit metadata, pushed the branch, and prepared this pull request. The adapter did not expose an exact context-window value in the run payload; tool use and local shell execution were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ac143f3254 |
style(ui): simplify Company Environments screen copy (#8400)
## Thinking Path > - Paperclip is the control plane for managing AI-agent work, and the UI needs to stay legible because operators use it to steer execution. > - This change sits in the Company Environments screen inside the Paperclip app UI. > - The screen carried redundant headings and several overlapping helper sentences that added clutter without adding meaning. > - The request was presentation-only: remove redundant copy and simplify labels, with no change to environment-selection behavior. > - The safest path was a single-file edit in `CompanyEnvironments.tsx` that preserved every control and only removed/renamed text. > - This is a `style:` change (UI copy/presentation), so no behavior is altered and no new test is warranted. ## Linked Issues or Issue Description No public GitHub issue exists for this small UI cleanup, so the problem is described here instead. - The Company Environments screen had redundant headings and overlapping helper copy. - The default environment control used a verbose label and a Local option label that added noise without adding meaning. - Requested outcome: simplify the copy so the screen is easier to scan while preserving the same behavior. Related public PRs reviewed for overlap: - Refs #8380 - Refs #8391 - Refs #8398 - Refs #4902 ## What Changed - Removed the extra top-level `Environments` header inside the screen content (both enabled and disabled states). - Removed the introductory paragraph above the environment list. - Renamed `Instance default environment` to `Default` and removed its helper sentence. - Removed the duplicate helper text below the default environment select. - Changed `Local (built-in default)` to `Local`. - Removed the redundant `Saved environments` section heading. - Removed the empty-state line `No saved environments yet. Local remains the default until you add another target.` - Removed the previously added UI test, which only asserted the now-final copy and added no lasting value. - Removed the now-unused `Settings` import and `instanceDefaultEnvironment` variable. ## Verification - Ran the existing `CompanyEnvironments` vitest suite locally: 4 passed. - This is a presentation-only copy change with no API, state, or behavior changes. - Reviewer can open Company Settings -> Environments and confirm the screen matches the requested wording cleanup with no behavior change. ## Risks - Low. Copy-only UI change with no API or state-management changes. - The only real risk is removing more context than intended from the top of the screen; review should focus on clarity. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Anthropic Claude Opus 4.8 via the Claude local adapter with tool use in a Paperclip-managed workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] This is a `style:` copy-only change; no new test is warranted (test-coverage gate skips `style:`) - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c0743482bc |
fix(claude-local): tolerate sandboxes whose Claude CLI lacks --effort (#8393)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The claude-local adapter launches the Claude CLI inside execution environments (local, SSH, and ephemeral sandboxes such as Daytona) > - Newer adapter code passes `--effort` to the CLI, but the Claude binary baked into some sandbox images is older and rejects it with `error: unknown option '--effort'`, so every run in those environments fails > - This needs addressing because the failure is environment-dependent and silent from the operator's perspective — the run just dies with a CLI usage error > - This pull request probes the in-sandbox CLI for `--effort` support once, caches the result per environment, and strips the flag (with a warning) when the CLI does not support it > - The benefit is that sandboxes with older Claude CLIs keep working instead of failing, with negligible probe overhead because the capability check is cached and reused across ephemeral leases ## Linked Issues or Issue Description No public GitHub issue exists. Describing the bug inline following the bug report template: ### What happened? Runs using the `claude-local` adapter inside certain sandbox/execution environments fail with `error: unknown option '--effort'`. The adapter unconditionally appends `--effort` to the Claude CLI invocation, but the Claude CLI version present in some sandbox base images predates that flag, so the process exits with a usage error and the run dies. ### Expected behavior The adapter should detect that the target environment's Claude CLI does not support `--effort` and degrade gracefully — drop the flag and emit a warning — rather than failing the run. ### Steps to reproduce 1. Configure an execution environment (e.g. a sandbox image) whose bundled Claude CLI is old enough to predate the `--effort` option. 2. Run any claude-local task that resolves to an effort level (so `--effort` is appended). 3. Observe the run fail immediately with `error: unknown option '--effort'`. ### Paperclip version or commit `master` at the time of this PR (branch forked from current `master`). ### Deployment mode Self-hosted / local instance using execution environments (reproducible with ephemeral sandbox providers such as Daytona where `reuseLease: false`). ### Agent adapter(s) involved Claude Code (`claude-local`). ## What Changed - Add a CLI capability probe (`cli-capabilities.ts`) that runs the target Claude binary's `--help` inside the execution environment to detect `--effort` support. - Strip `--effort` from the CLI args (emitting a warning) when the probe reports the flag is unsupported; keep it otherwise. - Cache probe results keyed by `sandbox:providerKey:environmentId:command` (no lease id) so the probe is reused across ephemeral leases — important for `reuseLease: false` sandbox configs like Daytona, which would otherwise re-probe on every run. - Conservative fallback: if the probe itself can't run/parse, assume the flag is supported (preserves prior behavior). ## Verification - `node_modules/.bin/vitest run src/__tests__/claude-local-execute.test.ts src/__tests__/claude-local-adapter-environment.test.ts` → **2 files, 30 tests passed**. - Regression test issues two `execute()` calls with distinct lease ids and asserts the in-sandbox `--help` probe runs exactly once (cache reuse across leases). - Added tests covering: flag stripped when unsupported, flag retained when supported, warning emitted, and conservative fallback when the probe fails. ## Risks Low risk. The change is additive and gated behind a probe with a conservative default (assume supported on probe failure), so existing environments that support `--effort` are unaffected. Worst case for an environment where the probe is unreliable is the prior behavior (flag passed through). ## Model Used Claude Opus 4.8 (claude-opus-4-8), extended thinking, with tool use / code execution via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (N/A — no UI change) - [ ] I have updated relevant documentation to reflect my changes (N/A — no doc-facing behavior change) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
07e98d2b2c |
feat(adapter-utils): add observable sandbox sync progress (#8395)
## Thinking Path > - Paperclip is the control plane for running AI-agent companies, so long-running remote work needs to stay observable to human operators. > - Cloud / sandbox agents are an active roadmap area, and their workspace sync path is part of the runtime substrate every remote coding run depends on. > - In the sandbox and SSH execution-target flows, Paperclip logged that sync had started, then often went silent for the full transfer window. > - That made large remote syncs feel stalled and also hid a real performance problem in the command-managed sandbox upload path. > - The first part of this pull request threads a throttled progress-reporting surface through the adapter execution-target stack so sync and restore work can emit meaningful updates. > - The second part fixes the command-managed sandbox transport itself: it removes the old serial 32KB append bottleneck, but also falls back away from the single-stream path when a provider-backed sandbox runner cannot surface mid-flight stdin progress. > - The result is that sandbox and SSH transfers are both faster and more observable, including the live Daytona-style sandbox case that previously only emitted `0%` and `100%`. ## Linked Issues or Issue Description No public GitHub issue exists for this bug, so it is described inline below following the bug report template. ### What happened - Remote sandbox and SSH workspace syncs could spend a long time transferring data while only logging a start line (`Syncing workspace and runtime assets to sandbox environment`) and, at best, a terminal line. - In the command-managed sandbox path, the original upload implementation also paid a large performance cost by appending base64 data in many small sequential remote writes (thousands of serial 32KB round-trips on a large workspace). - After the initial transport rewrite, live provider-backed sandbox runs still only emitted `0%` and `100%` because the single-stream stdin RPC buffered progress until completion. ### Expected behavior - Long-running sandbox and SSH syncs should periodically report how much of the transfer is complete (a percentage and/or MB transferred) so an operator can tell the run is healthy and making progress rather than stuck. - The main sandbox upload path should not be artificially slow. - A transfer that fails partway should leave an explicit failure marker in the log rather than a dangling intermediate percentage. ### Steps to reproduce 1. Run an agent against a sandbox (command-managed) or SSH (remote-managed) execution target with a non-trivial workspace. 2. Watch the run log during the workspace/runtime asset sync phase. 3. Observe that the log shows the sync start line and then stays silent for the full transfer (live provider-backed sandbox runs only show `0%` then `100%`). ### Paperclip version or commit - Branch `PAPA-825-provide-status-updates-when-syncing-sandboxes` off `master`. ### Deployment mode - Self-hosted / local instance using sandbox (command-managed) and SSH (remote-managed) execution targets, including provider-backed sandbox runners. ## What Changed - Added shared throttled runtime progress reporting and threaded `onProgress` through the adapter execution-target surface and adapter `execute.ts` entrypoints. - Added sync and restore progress reporting for the command-managed sandbox path and the SSH/remote-managed path, including git import/export progress where totals are known. - Reworked command-managed sandbox transfer behavior so uploads use the faster single-stream path when appropriate, but fall back to chunked progress-emitting writes when the runner cannot expose mid-stream stdin progress. - Marked provider-backed environment sandbox runners as not supporting single-stream stdin progress so live sandbox runs emit meaningful intermediate updates instead of only `0%` and `100%`. - Emit an explicit terminal failure marker (`failed at NN% (x/y MB)`) when an SSH/tar transfer rejects, so a failed sync no longer leaves a dangling intermediate percentage in the log. - Run the SSH sync/restore size estimate (local directory walk / remote `du` probe) concurrently with the transfer instead of awaiting it before opening the pipe, so progress instrumentation no longer adds startup latency proportional to workspace file count. - Added and extended focused regression coverage for runtime progress throttling and the new failure marker, command-managed sandbox transfers, sandbox orchestration, SSH transfer progress, and environment execution-target wiring. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/runtime-progress.test.ts packages/adapter-utils/src/ssh-fixture.test.ts packages/adapter-utils/src/command-managed-runtime.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` - `pnpm exec vitest run packages/adapter-utils/src/command-managed-runtime.test.ts server/src/__tests__/environment-execution-target.test.ts` - `npx tsc --noEmit` for `packages/adapter-utils` ## Risks - The provider-backed sandbox fallback now prefers chunked command-managed writes when progress hooks are active, so small-to-medium uploads may trade some raw throughput for observable intermediate progress on runtimes that cannot surface true mid-stream stdin progress. - Progress percentages on tar-based transfers still depend on estimates in some cases, so operators may briefly see MB-only lines before the estimate resolves, then near-final clamping before the terminal `100%` line. - This PR changes shared execution-target behavior used by multiple adapters, so regressions would most likely appear in remote runtime setup/teardown flows rather than in a single adapter. ## Model Used - Initial implementation: OpenAI GPT-5.4 via Codex local agent (`codex_local`), high reasoning mode. - Observability follow-ups (failure marker, concurrent size estimate, added tests): Claude Opus 4.8 via Claude Code (`claude_local`). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |