mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
c9881223e49ca7486c71f79eed767ec98e5f4eb6
3250
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c9881223e4 |
fix(routines): show routines grouped by folder name (#10201)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Routines page is part of the operator UI for scheduled routine management > - The grouped-by-folder view was not presenting folder sections inline in the main pane > - That made the folder grouping mode harder to scan and hid the separation between custom folders, Unfiled routines, and built-in routines > - This pull request updates the Routines page rendering so folder groups appear as inline sections and built-in routines still split into their own section afterward > - The benefit is that grouped routines stay readable and the page matches the intended folder organization ## Linked Issues or Issue Description No corresponding public GitHub issue exists, so the problem is described directly below following the bug template. ### What happened On the Routines page, selecting Group → Folder flattened the grouped list into a single "All routines" section with only the separate built-in routines section below it. ### Expected behavior Group → Folder should render one inline section per folder, keep routines with no folder in an Unfiled section, and preserve the separate built-in routines section after the custom folder groups. ### Steps to reproduce 1. Open the Routines page. 2. Change grouping to Folder. 3. Observe the main pane. 4. The routine list is flattened instead of grouped into folder-labeled inline sections. ### Paperclip version / commit Current PR head: `a8e384c838e362de3437c7a88bc7aa38b10fd9c0` on `fix/routine-folder-grouping`. ### Deployment mode Local development workspace for the Paperclip app UI. ## What Changed - Updated the Routines page rendering so grouped folders render as inline sections instead of flattening into a single list. - Kept routines without a folder grouped under Unfiled. - Preserved the built-in routines section after custom folder groups. - Added and updated tests for the folder-grouped rendering behavior. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Routines.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` ## Risks - Low risk: the change is localized to the Routines page rendering and its test coverage. - The main behavioral risk is accidental grouping regressions if future routine-grouping logic changes without updating the tests. ## Model Used OpenAI Codex, GPT-5, tool-use enabled, 256k-context class model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.724.0-canary.9 |
||
|
|
3a16b91217 |
feat(status-cards): single-message setup drives query and update prompt (#10202)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their ongoing work. > - Status cards turn a standing question into recurring, agent-generated summaries on the board. > - The existing setup split intent across a watch prompt and separate update instructions, which made creation and later behavior harder to understand. > - A status card should have one durable source of truth for both deciding what to watch and telling the summarizer what each update must contain. > - This pull request makes the card prompt that source of truth, simplifies creation to one step, and lets operators choose the running agent immediately. > - The benefit is a smaller mental model, fewer configuration modes, and consistent update instructions throughout the card lifecycle. ## Linked Issues or Issue Description Status cards currently require operators to express the same intent in two places: the watch prompt and optional update instructions with append/replace/none modes. This feature simplifies the experimental status-card workflow so a single prompt defines both the watch query and every generated update. The create flow must also support selecting the responsible agent without a second setup step. Related prior status-card work: #10101. ## What Changed - Use the status card's single prompt to compile the watch query and directly instruct every summary update. - Add migration `0190_status_card_single_prompt` to remove `status_cards.instructions_mode` and `status_cards.instructions`. - Add `agentId` to `createStatusCardSchema`, validate company membership, and default new cards to the built-in Summarizer. - Replace the two-step create flow with one prompt-and-agent dialog and extract a shared `SummarizerAgentSelect` for create/settings surfaces. - Remove the extra-instructions settings section, reset incremental history when the prompt changes, and rename the board page to "Status". - Update the bundled `status-card-query` skill and board-operator documentation, then regenerate the skills catalog manifest. ## Verification - Server status-card suites: 29/29 passing. - UI `StatusCards` suites: 22/22 passing. - Skills catalog suite: 20/20 passing. - `tsc -b` passes for server, UI, shared, and database packages. - `pnpm check:migrations` passes. - Light and dark mode screenshots cover the new create dialog and settings tab. ## Risks - Migration `0190` intentionally drops existing separate instruction text. Existing card prompts remain and become the update instructions under the new model; status cards are experimental and feature-flagged. - Prompt edits now reset the incremental summary chain and trigger a full rebuild, which is intentional because the prompt is also the update contract. - Agent selection is company-scoped; invalid agent ids return a validation error rather than creating a misrouted card. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Implementation: Anthropic Claude via the `claude_local` adapter, agent label "Claude Fable 5"; extended reasoning, tool use, and code execution. The exact provider model id and context-window value were not retained in the task metadata. - PR preparation: OpenAI GPT-5.4 through Codex CLI, with reasoning, repository inspection, GitHub CLI, and Paperclip API tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude Fable 5 <noreply@anthropic.com>canary/v2026.724.0-canary.8 |
||
|
|
b083b173ea |
fix(db): use calendar month retention for backups (#3718)
## Thinking Path > - Paperclip includes built-in database backup and retention behavior as part of its operational reliability surface. > - The retention implementation lives in `packages/db/src/backup-lib.ts`. > - Monthly pruning is intended to keep the newest backup per retained calendar month. > - The current cutoff uses a fixed 30-day approximation, which deletes January backups too early when the current month has 31 days. > - This pull request switches the monthly cutoff to calendar-month boundaries instead of a fixed day multiplier. > - The benefit is that monthly retention now matches the documented calendar-month behavior and does not prune valid backups prematurely. ## Linked Issues or Issue Description Fixes #3713 Two other pull requests implemented the same fix and have already been closed as duplicates of this one: - #3798 — same author's later take. Anchors the cutoff to the 1st correctly, but mutates the date in local time and leaves `monthKey` on local time, and unit-tests the helper in isolation rather than end to end. - #4031 — decrements the month without anchoring to the 1st, so partial-month drift and a `setMonth` day-overflow edge case remain. No tests. ## What Changed - Replaced the fixed `30 * 24h` monthly retention cutoff with a calendar-month cutoff anchored to the first day of the earliest retained month. - Added a regression test that freezes `Date.now()` at March 31 and proves the newest January backup is retained when `monthlyMonths=2`. - Kept the rest of the pruning behavior unchanged: daily and weekly tiers still use their existing windows and bucket selection rules. ## Verification - `pnpm --filter @paperclipai/db exec vitest run src/backup-lib.test.ts` - `pnpm --filter @paperclipai/db exec tsc --noEmit` - Note: `pnpm --filter @paperclipai/db typecheck` hits an environment-specific `check:migrations` runtime failure on this host (`Cannot find module ./cjs/index.cjs from ` via Bun), so I used plain `tsc --noEmit` to validate the code changes themselves. ## Risks - Low risk. This only changes the monthly retention cutoff calculation. - The pruning buckets are still selected the same way; the fix only widens the retained month window to align with calendar-month semantics. ## Model Used - OpenAI Codex GPT-5 coding agent with terminal tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --- <sub>Edited by commitperclip during triage: added the **Linked Issues** section (`Fixes #3713`) and the duplicate-PR search line the PR template requires. The duplicate search was performed by the triage pipeline, which grouped this PR with #3798 and #4031 and selected this one as the canonical fix. Everything else is the author's original description.</sub> --------- Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>canary/v2026.724.0-canary.7 |
||
|
|
7e40ed8c43 |
feat(status-cards): add experimental status card update view (#10101)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Operators need a board-level way to monitor a changing slice of company work without repeatedly rebuilding filters or reading raw task threads. > - Existing summaries are useful snapshots, but they do not provide a dedicated query-backed card with refresh policy, change tracking, update history, and per-update cost visibility. > - The capability needs to be safe to evaluate before it becomes part of the default product surface. > - This pull request adds end-to-end experimental Status Cards, from schema and query compilation through update orchestration and operator UI. > - The entire feature is gated behind the `enableStatusCards` experimental toggle, including its route and sidebar entry. > - The benefit is a governed, inspectable way to keep focused operational rollups current while preserving explicit controls over refresh frequency and spend. ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (`packages/db`, `packages/shared`, `server/`, `ui/`, and bundled skills/docs). ### Problem or motivation Operators cannot currently define a reusable natural-language view of company work, compile it into an inspectable query, and keep its summary current as matching issues change. Rebuilding filters and rereading task threads makes board-level monitoring repetitive and hides the relationship between source changes, refresh cost, and the resulting summary. ### Proposed solution Add experimental Status Cards that compile operator intent into a query, summarize matched work, record each update, expose manual/interval/reactive refresh policies and costs, and preserve the last good result across stale, updating, paused, and error states. The capability is off by default and fully gated behind `enableStatusCards`, including its route and navigation entry. ### Alternatives considered - Extend existing one-off summaries: rejected because status cards require persistent query provenance, refresh policy, update history, and card-specific cost controls. - Add a dashboard-only filter widget: rejected because it would not provide governed background refresh, an update ledger, or an inspectable compile pipeline. - Ship the surface by default: rejected in favor of an experimental toggle while behavior and operator value are evaluated. ### Roadmap alignment This advances Paperclip’s board-level execution visibility and output-first product goals. `ROADMAP.md` was checked and no duplicate status-card initiative was found. ### Additional context No related open PR was found in the public GitHub search for status cards. The PR-only design wireframes were removed from the repository after review; the published prototype remains external to the production source tree. ## What Changed - Added company-scoped status-card schema, CRUD APIs, compile provenance, update ledger, shared contracts, validators, and OpenAPI coverage. - Added the text-to-query compile pipeline, bundled `status-card-query` agent skill, query versioning, and authorized write-back flow. - Added the experimental board, create flow, lifecycle tiles, detail/settings/debug drawers, archived view, routing, navigation, and instance setting. - Added a change-gated update engine with manual, interval, and reactive refresh policies, trigger selection, active hours, and daily token caps. - Added per-update token/cost recording, today and lifetime rollups, and policy-derived cost previews. - Added operator documentation and agent-authoring hardening for compile and update behavior. - Added PR-prep integration coverage for settings/startup wiring and replaced raw UI values with design-system tokens. - Removed the PR-only `design/pap-15023-status-cards` wireframe artifacts so the repository contains only production feature assets. ## Verification - `pnpm -r typecheck` — passes on the PR head; includes `ui` `tsc -b` passing. The UI compile gate was also independently recorded as passing at `6d7f3cf96b` on July 23, 2026. - `pnpm build` — passes. - `pnpm check:token-gates` — passes with all three gates clean. - `pnpm test:run` — 2,880 tests passed and 1 skipped; the sole failure was an unrelated 10-second `afterAll` database-cleanup timeout in `execution-workspaces-service.test.ts`. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts` — passes on immediate focused rerun (25/25). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/instance-settings-service.test.ts src/__tests__/server-startup-feedback-export.test.ts` — passes (31/31). - `pnpm --filter @paperclipai/ui exec vitest run src/pages/StatusCards/StatusCardSettingsForm.test.tsx src/pages/StatusCards/StatusCardTile.test.tsx src/pages/StatusCards/format.test.ts src/lib/status-card-state.test.ts` — passes (26/26). - Recorded pre-PR QA: compile-pipeline e2e PASS; full lifecycle and cost QA PASS; security re-review PASS after write-back hardening; UX approved. - `pnpm exec vitest run packages/db/src/status-card-migrations.test.ts` — passes; reapplies migrations `0185`–`0189` against an already-migrated embedded Postgres database. - `pnpm --filter /db check:migrations` — passes migration numbering and safety checks. - `pnpm --filter /db typecheck` — passes. - Merged current `origin/master` on July 24, 2026 with no conflicts; migrations `0185`–`0189` remain unclaimed on master. ## Risks - The feature introduces five database migrations and a new background update path; all new DDL is repeat-safe after partial application, migration numbering/safety checks pass, and update execution is company-scoped and change-gated. - Natural-language compilation can produce invalid or overly broad queries; compile provenance, query validation, debug visibility, and version history make failures inspectable and recoverable. - Reactive or interval refresh could increase spend; active hours, max refresh frequency, daily token caps, per-update cost records, and budget-paused states bound and expose that risk. - The branch name contains an internal execution identifier because it is a fixed handoff branch; it was intentionally not renamed or rebased per the release handoff instructions. - Overall rollout risk is limited because the route, navigation, services, and UI are disabled by default behind `enableStatusCards`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.5 with reasoning, repository tool use, shell execution, GitHub CLI, and test/build execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change; the fixed execution-workspace identifier is documented as an authorized handoff exception - [x] I have run tests locally and they pass, with the one cleanup timeout passing on focused rerun - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2002c4fff3 | refactor: remove redundant archived project filters (#10177) canary/v2026.724.0-canary.6 | ||
|
|
965a827ee7 |
feat(docker): publish a cloud image variant with built bundled plugins (#10157)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed (cloud-hosted) deployments configure instances through `PAPERCLIP_MANAGED_CONFIG`, including a `plugins.autoInstall` key list that the boot-time installer resolves against the bundled plugin catalog > - The installer requires each bundled plugin's `dist/manifest.js` (`server/src/services/bundled-plugins.ts`), but the published image only ships the sandbox providers' *source* — they are intentionally excluded from the pnpm workspace, and the Dockerfile never builds them > - Every managed auto-install therefore logs `bundled plugin bundle not present; skipping auto-install` and no sandbox provider can be provisioned through managed config > - Baking built plugins into the single published image would fix it but makes every self-hosted pull carry the providers' `node_modules` for a managed-only mechanism > - This pull request adds a `cloud` Dockerfile target extending `production` with built bundled plugins — parameterized by build arg and currently just `daytona` — published alongside the default image with a `-cloud` tag suffix > - The benefit is working plugin auto-provisioning for managed deployments while the self-hosted image stays byte-identical and the cloud variant only carries what is actually deployed ## Linked Issues or Issue Description Fixes #10158 (filed for this problem; no prior issue existed — searched for duplicate/related PRs and issues around bundled plugins, docker image variants, and auto-install). Summary: **What happened:** on a managed instance with `plugins.autoInstall: ["daytona"]` delivered via `PAPERCLIP_MANAGED_CONFIG`, boot logs `bundled plugin bundle not present; skipping auto-install` with `pluginPath: /app/packages/plugins/sandbox-providers/daytona`, and the plugin is never installed. **Expected:** the advertised bundled-catalog keys are installable from the published image. **Why:** the image ships plugin source without `dist/` — nothing in the Dockerfile builds the workspace-excluded sandbox providers. ## What Changed - `Dockerfile`: new `cloud-plugins` stage (based on `build`, so devDependencies are available for `tsc`) that installs and builds each provider named in the `CLOUD_BUNDLED_PLUGINS` build arg standalone (`pnpm install --ignore-workspace --no-lockfile && pnpm build`, exactly as the providers' READMEs prescribe), asserting `dist/manifest.js` exists per plugin and failing loudly on unknown names; new `cloud` stage = `production` + the built plugin tree. The arg defaults to `daytona` — the only provider managed deployments auto-install today; every entry adds its `node_modules` to the image, so the list grows only with actual need (a one-line workflow change). - `.github/workflows/docker.yml`: the existing build step is pinned to `target: production` (without this, the new trailing stage would silently become the default build target — this pin is what keeps the self-hosted image identical); new metadata + build-push steps publish the `cloud` target (with `CLOUD_BUNDLED_PLUGINS=daytona`) under the same tag set with a `-cloud` suffix (`sha-<short>-cloud`, `latest-cloud`, `<version>-cloud`), same schema labels, reusing the GHA layer cache ## Verification - All seven sandbox providers build standalone from a clean checkout with the exact commands the new stage runs, each producing `dist/manifest.js` — so the current `daytona` default works and future list additions are known-good - The stage's shell loop was dry-run against the checkout (directory existence + per-plugin assertion logic) - Workflow YAML lints clean - **Not run:** a full multi-arch `docker build` (no local docker daemon). The `cloud` stage is additive and the default target is pinned, so the risk is contained to the new build step; the first master build after merge proves it end-to-end ## Risks - Self-hosted behavior: unchanged. The default image build is pinned to the `production` target, which produces the same layers as before this change; the `cloud` stages run only for the new build step. - The plugin installs in the `cloud-plugins` stage use `--no-lockfile` (the providers are workspace-excluded and lockfile-less by design), so plugin dependency resolution is not pinned at image-build time. This mirrors the existing Plugins-page install path, which resolves from npm at install time. - CI cost: one additional build-push per master push. It reuses the layer cache from the production build, so the marginal work is the single plugin's build layers. - An unknown name in `CLOUD_BUNDLED_PLUGINS`, or a provider that stops producing `dist/manifest.js`, fails the cloud build loudly rather than publishing a broken variant. ## Model Used Claude (Anthropic), model ID `claude-fable-5[1m]` via Claude Code CLI — extended thinking and tool use (code edits, standalone plugin build verification, workflow lint). ## Checklist - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] Self-hosted behavior unchanged (default build target pinned to `production`) - [x] One clear change: publish a cloud image variant with built bundled plugins |
||
|
|
ee3ed0117e |
fix(opencode): register configured model in runtime config (#10178)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents for work > - Agent runs depend on adapters translating Paperclip configuration into each agent runtime's native configuration > - The OpenCode local adapter passes configured models through the `--model provider/model` argument > - OpenCode only resolves that argument when the model id exists in the provider's runtime `models` map > - Valid provider-served model ids missing from OpenCode's bundled catalog therefore fail locally with `Model not found` > - This pull request registers the configured model in the injected runtime configuration without overwriting explicit provider definitions > - The benefit is that uncataloged routing variants and newly released models resolve while cataloged models retain their metadata ## Linked Issues or Issue Description ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I can reproduce this on current `master`. - [x] I have confirmed the error originates in Paperclip's OpenCode adapter rather than the provider or local configuration. ### What happened? OpenCode local runs failed with `Model not found` when a configured `provider/model` id was valid at the provider but absent from OpenCode's bundled model catalog. OpenRouter routing variants such as model ids ending in `:nitro` are one example. ### Expected behavior Any configured provider-served model id should resolve when Paperclip starts OpenCode, including ids not yet present in the bundled catalog. ### Steps to reproduce 1. Configure the OpenCode local adapter with a valid provider/model id that is absent from OpenCode's bundled catalog. 2. Start an agent run. 3. Observe that OpenCode rejects the `--model` value with `Model not found` before the session starts. ### Paperclip version or commit Current `master` before this change. ### Deployment mode Local dev using the OpenCode local adapter and an existing provider API key. ### Installation method Built from source. ### Agent adapter(s) involved OpenCode local. ### Database mode Not database-related. ### Relevant logs or output `Model not found` ### Additional context Reproduced with OpenCode 1.15.5. No duplicate or related public GitHub issues or pull requests were found. ### Privacy checklist - [x] I have reviewed all pasted output for sensitive information and no secrets or PII are included. ## What Changed - Register the configured `provider/model` id as an empty custom model entry in the injected `opencode.json` provider configuration. - Preserve explicit model definitions from user configuration and `PAPERCLIP_OPENCODE_PROVIDERS`. - Skip registration for model strings that do not use the `provider/model` form. - Add focused coverage for uncataloged models, explicit definitions, and invalid model strings. ## Verification - `cd packages/adapters/opencode-local && pnpm exec vitest run src/server/runtime-config.test.ts` — 14 tests passed. - `cd packages/adapters/opencode-local && pnpm exec tsc --noEmit` — passed. - Manual reproduction with OpenCode 1.15.5: the uncataloged OpenRouter routing variant fails without the injected model entry and resolves with it. ## Risks - Low risk: the empty entry deep-merges with catalog metadata for known models, and existing explicit model definitions take precedence. - The behavior is limited to syntactically valid `provider/model` configuration values in the OpenCode local adapter. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact runtime model `gpt-5.6-sol` (context-window size not exposed by the runtime), with reasoning, tool use, terminal execution, and code-editing capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.724.0-canary.5 |
||
|
|
7f766526a6 |
feat(sandbox): add task-scoped egress grants (#10155)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Confinement providers protect agent runs with default-deny network policies > - Kubernetes environments currently apply only provider-level, namespace-wide egress allowances > - Tasks that legitimately need GitHub or package registries therefore cannot request narrow access, while network failures do not explain the governing policy or how to request a grant > - This pull request adds issue-scoped egress grants that become workload-owned, run-label-selected policies and carries the effective grant through lease audit metadata > - The benefit is that internet-dependent work can run without enabling broad egress for every concurrent task, and denied requests point operators to the exact grant path ## Linked Issues or Issue Description No public issue exists. Related but distinct: Refs #9944, which adds a provider-wide open-internet posture; this PR keeps provider defaults narrow and adds per-task grants. **Problem / motivation** Kubernetes sandbox egress is configured at the provider/tenant level. A task that needs to clone from GitHub or install from PyPI cannot request those destinations without changing the policy for every run in the tenant namespace. DNS/connectivity failures also surface as generic tool errors with no policy name or remediation path. **Proposed solution** Accept `executionWorkspaceSettings.networkEgress.allowFqdns` and `allowCidrs`, forward the setting through heartbeat environment acquisition, and create a workload-owned NetworkPolicy or CiliumNetworkPolicy selected by `paperclip.io/run-id`. Record the effective grant in lease activity/metadata, expose policy context through `PAPERCLIP_NETWORK_EGRESS_*`, and append the grant path to likely policy-related stderr failures. **Alternatives considered** A provider-wide open-internet switch is broader than required and is already covered by #9944. Mutating the existing namespace policy would leak each task's destinations to other concurrent runs. Standard Kubernetes NetworkPolicy cannot enforce FQDNs exactly, so standard mode uses the existing hardened public-IPv4 TCP 80/443 fallback only for the selected run; Cilium mode remains exact. **Roadmap alignment** This extends the existing cloud/sandbox agent roadmap capability with task-level control-plane policy and does not duplicate a planned roadmap item. ## What Changed - Added validated `networkEgress` grants to issue execution workspace settings and forwarded them through environment lease acquisition. - Added workload-owned, run-label-scoped NetworkPolicy/CiliumNetworkPolicy resources for task FQDN/CIDR grants. - Added lease audit metadata, sandbox policy environment variables, and actionable network-denial stderr guidance. - Added focused parser, manifest, policy creation, and denial-message tests plus Kubernetes provider documentation. ## Verification - `pnpm -C packages/shared exec vitest run src/validators/issue.test.ts` — 27 passed. - `pnpm -C packages/plugins/sandbox-providers/kubernetes test -- --run test/unit/network-policy.test.ts test/unit/cilium-network-policy.test.ts test/unit/scoped-network-egress.test.ts` — 21 passed. - `pnpm -C server exec vitest run src/__tests__/execution-workspace-policy.test.ts` — 15 passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-plugin-environment.test.ts server/src/__tests__/environment-runtime.test.ts` — 26 passed. - `pnpm --dir packages/db build && pnpm --dir packages/shared build && pnpm --dir packages/plugins/sdk build` — passed, including migration safety checks. - `pnpm --dir packages/plugins/sandbox-providers/kubernetes typecheck && pnpm --dir server typecheck` — passed after refreshing the worktree's frozen offline dependencies. - End-to-end cluster validation of the `build-cython-ext` benchmark remains for CI/maintainer Kubernetes infrastructure; the focused tests assert `github.com` and `pypi.org` produce a policy selected only by the granted run. ## Risks - Standard NetworkPolicy cannot express FQDNs, so an FQDN grant allows hardened public IPv4 TCP 80/443 for that run; use Cilium mode for exact hostname enforcement. - The new field is additive and absent by default, so existing runs keep the current provider-level policy. - Workload owner references garbage-collect scoped policies with the Job/Sandbox; a cluster/controller that ignores owner references could temporarily strand a policy that still selects no future run ID. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, exact model ID `gpt-5.6-sol`, high reasoning mode, tool use and code execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.724.0-canary.4 |
||
|
|
564870020b | Exclude archived projects from the default project list route (#10146) canary/v2026.724.0-canary.3 | ||
|
|
14f20be92b |
ci: harden Docker image build workflow against lockfile drift and runner disk exhaustion (#10142)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Docker image publish workflow (`.github/workflows/docker.yml`) builds and pushes the multi-arch `ghcr.io` image on every master push, so users pulling the container get the latest code > - The two newest master runs of that workflow failed, so no images have been published past a recent master commit > - The failures had two distinct causes: run [30054330748](https://github.com/paperclipai/paperclip/actions/runs/30054330748) hit `ERR_PNPM_LOCKFILE_CONFIG_MISMATCH` (committed `pnpm-lock.yaml` drifted from `patchedDependencies` in package metadata), and run [30050197392](https://github.com/paperclipai/paperclip/actions/runs/30050197392) hit `no space left on device` during the multi-arch buildx export > - This pull request hardens the publish job against both failure modes: it refreshes the lockfile (lockfile-only, guarded) before the build, and frees runner disk space before buildx setup > - The benefit is that image publishing keeps working through routine lockfile drift and the growing multi-arch build footprint, so `ghcr.io` images stay current with master ## Linked Issues or Issue Description - Refs #8286 — same class of Docker-build lockfile mismatch failure - Refs #8827 — pnpm 9.15.x pin / lockfile regeneration discussion - Note: the immediate lockfile drift on master was fixed by #10132; the refresh step here prevents the *next* drift from breaking image publishing again ## What Changed - Added a pnpm + Node setup and a **"Refresh lockfile for Docker build context"** step to the image job in `.github/workflows/docker.yml`: runs `pnpm install --lockfile-only --ignore-scripts --no-frozen-lockfile`, exits cleanly if nothing changed, and **fails the job if anything other than `pnpm-lock.yaml` was modified** by the refresh - Added a **"Free runner disk"** step (before buildx setup) that prunes the pnpm store, apt caches, preinstalled toolchains (`/usr/share/dotnet`, Android SDK, Swift, Boost, PowerShell, GHC, CodeQL/PyPy/Ruby toolcache), and dangling Docker state, logging `df -h` before/after - No changes outside the workflow file (54 added lines, nothing removed) ## Verification - Pulled the logs of both failed master runs and matched each failure to the step that addresses it: [30054330748](https://github.com/paperclipai/paperclip/actions/runs/30054330748) failed with `ERR_PNPM_LOCKFILE_CONFIG_MISMATCH`, [30050197392](https://github.com/paperclipai/paperclip/actions/runs/30050197392) failed with `no space left on device` during the buildx export - Confirmed pnpm `9.15.4` in the new setup step matches the repo `packageManager` field and the version used in the Dockerfile, so the refreshed lockfile is generated by the same pnpm the image build consumes - Validated the workflow YAML parses cleanly - The workflow triggers on master pushes / manual dispatch; the definitive check is the first master run after merge — reviewers can also `workflow_dispatch` it from this branch if desired ## Risks - The lockfile refresh runs with `--ignore-scripts` and a guard that aborts on any non-lockfile change, so it cannot silently pull unexpected code into the image; worst case it fails the job with a clear diff - The published image could be built from a refreshed lockfile that differs from the committed one when drift exists — that keeps publishing alive but can mask drift on master, which still needs the committed lockfile fixed (as #10132 did) - Disk cleanup removes preinstalled toolchains only on the ephemeral runner for this job; other jobs/workflows are unaffected - Low risk overall: additive steps in a single workflow file ## Model Used - Claude (Anthropic) — `claude-fable-5` (Claude Code agent harness, extended thinking, tool use). Used to diagnose the failing CI runs from logs, author the workflow changes, and prepare this PR. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (no runtime code touched; workflow YAML validated — see Verification) - [ ] I have added or updated tests where applicable (n/a — CI workflow change) - [x] I have updated relevant documentation to reflect my changes (none needed) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm once checks run) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review pass) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.724.0-canary.2 |
||
|
|
caae2778f0 |
fix: deliver plugin agent session turns and replies (#10137)
## Thinking Path > - Paperclip manages agent execution through heartbeat runs and adapter-specific sessions > - Plugins can open an agent session and send a conversational message through the host service > - The host previously stored that message only in opaque wake payload metadata, so local adapters never saw it in their CLI prompt > - The host also forwarded run log chunks but did not expose the persisted final assistant text as the session reply > - This pull request defines both sides of the session contract in the shared wake renderer and terminal run event > - The benefit is that local adapters receive the actual conversational turn and plugins receive one canonical final reply ## Linked Issues or Issue Description Related context: Refs #629 and Refs #2880 describe adjacent `claude_local` final-text visibility failures. They concern issue comments rather than plugin agent sessions, but exercise the same need for a canonical persisted run summary. Companion consumer change: paperclipai/paperclip-gateway#3. Bug description: - **Observed:** calling the plugin host's `agents.sessions.sendMessage()` with `prompt: "hello"` woke a `claude_local` agent, but the generated CLI prompt omitted `hello`. On completion, the session emitted log chunks and a generic `Run completed` done event, so callers could not reliably recover the assistant reply. - **Expected:** the prompt becomes the user-supplied conversational turn for that agent session, and the successful terminal event carries the run's canonical final user-facing assistant text. - **Reproduction:** create a plugin agent session for a local adapter, call `sendMessage()` with a non-empty prompt, inspect the adapter prompt and terminal session event. - **Affected baseline:** `b517b887a` on `master`, local trusted deployment with plugin host services and `claude_local`; `codex_local` shared the wake-rendering gap because both use the common Paperclip wake prompt renderer. ## What Changed - Added a typed `agentMessage` wake payload rendered by the shared adapter prompt path used by `claude_local`, `codex_local`, and other local adapters. - Labeled session content as user-supplied and explicitly non-authoritative: it cannot expand authorization, permissions, task scope, or company boundaries. - Preserved ordinary heartbeat behavior by omitting the section when no agent-session message exists. - Added canonical `finalText` to terminal heartbeat status events from the already-persisted run summary/result/message. - Defined successful `AgentSessionEvent.message` as the canonical final user-facing reply (or `null`) and forwarded it on the terminal `done` event. - Added host, wake-renderer, normal-heartbeat, and terminal-reply regression coverage. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/heartbeat-agent-session-message.test.ts server/src/__tests__/heartbeat-run-status-payload.test.ts server/src/__tests__/plugin-agent-sessions.test.ts server/src/__tests__/heartbeat-run-summary.test.ts` — 87 passed. - `pnpm -r typecheck` — passed across all 31 workspaces. - `pnpm build` — passed. - `pnpm test:run` — 2,860 passed, 1 skipped, 3 unrelated failures: two existing macOS temp-path alias assertions (`/tmp` vs `/private/tmp`) in workspace branch-containment tests and one reproducible auto-port runtime-service adoption failure. The same three failures reproduce when the two files run alone; none touch this change. - Live Slack verification intentionally remains operator-gated because it requires rebuilding/restarting the host. ## Risks - User-controlled chat text now reaches the model prompt, which is an intentional prompt-injection surface. The renderer labels it as untrusted conversational content, while the existing plugin/session company checks and caller authorization remain unchanged. - `finalText` is added to company-scoped heartbeat status events. It is derived from the same persisted summary/result/message already used for run comments; no raw stdout or secrets are added. - Consumers that ignore the new field remain compatible, and successful runs without usable final text still emit `message: null`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex (GPT-5), agentic reasoning with repository/tool use and code execution; context-window size is not surfaced in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.724.0-canary.1 |
||
|
|
0650ae970f |
chore(lockfile): refresh pnpm-lock.yaml (#10136)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.724.0-canary.0 |
||
|
|
b517b887ad | fix(acpx): decouple host proxy spawn cwd from in-sandbox remoteCwd (#10122) canary/v2026.723.0-canary.20 | ||
|
|
e41ba306c5 |
feat(secrets): thread audit actor into skip-user-secret skills routes (#10124)
## Thinking Path > - Paperclip manages agent work and needs auditable control over secret resolution > - The skip-user-secret skills routes still have to attribute access to the real actor > - These routes were calling the adapter config resolver without an access context > - That dropped actor attribution from the company `secret_ref` audit trail > - This pull request threads the existing actor-secret context helper into both skills routes > - The benefit is that audit fidelity is restored without changing `skipUserSecrets` behavior ## Linked Issues or Issue Description Refs #10115. This PR fixes a gap in the skills read/sync routes where `resolveAdapterConfigForRuntime` was being called without an audit access context, so company secret resolution could not reliably attribute the request to the acting user or agent. The change keeps `skipUserSecrets: true` intact and only restores audit fidelity. ## What Changed - Threaded `buildActorSecretContext(req, { consumerType: "agent", consumerId })` into `GET /agents/:id/skills` - Threaded the same actor context into `POST /agents/:id/skills/sync` - Updated the route tests to assert a non-`undefined` actor context reaches the resolver while `skipUserSecrets: true` stays unchanged ## Verification - `tsc --noEmit` - `agents` and `secrets` Vitest suites: 33 files / 448 tests green - Route spy assertions confirm both skills routes now pass an actor-derived context to the resolver ## Risks - Low risk: the change is limited to audit context propagation on two skills routes - If a downstream resolver assumes the third argument can be `undefined`, this makes the context explicit on these routes - The user-secret authorization behavior does not change because `skipUserSecrets` remains true ## Model Used OpenAI GPT-5 via Codex, tool-using coding agent, 256k context window ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.19 |
||
|
|
58ae799dbd |
fix(ci): regenerate lockfile for patch changes (#10132)
## Thinking Path > - Paperclip uses GitHub Actions to keep generated lockfile changes deterministic in CI > - The workflow decides when to regenerate the lockfile based on file/path changes > - Patch changes can live under a top-level `patches/` directory, and those changes also affect dependency resolution > - If the workflow misses that path, CI can skip lockfile regeneration when it should run > - This pull request adds top-level `patches/` to the trigger so patch updates participate in the existing lockfile regeneration flow > - The benefit is that patch-related dependency changes continue to get the same CI protection as the other manifest and workspace triggers ## Linked Issues or Issue Description No public GitHub issue is linked here. The underlying problem is that top-level `patches/` files are part of pnpm's dependency graph, but the PR workflow's lockfile-regeneration gate only looked at package manifests, workspace config, `.npmrc`, and `pnpmfile.*` changes. That meant patch-only edits could skip `pnpm install --lockfile-only` and leave downstream frozen-install jobs on a stale lockfile. This PR keeps the existing manual lockfile edit guard in place. The intended behavior is still: CI owns lockfile regeneration, and patch changes are allowed to trigger that regeneration without letting contributors commit `pnpm-lock.yaml` directly. ## What Changed - Added top-level `patches/` to the PR workflow's dependency-resolution trigger. - Left the manual `pnpm-lock.yaml` edit blocker unchanged so CI still owns lockfile regeneration. ## Verification - `git diff --check .github/workflows/pr.yml` - Verified the workflow path predicate matches `patches/acpx@0.12.0.patch`, `package.json`, `packages/shared/package.json`, `pnpm-workspace.yaml`, `.npmrc`, `pnpmfile.cjs`, `pnpmfile.js`, and `pnpmfile.mjs`, while excluding nested patch paths and unrelated files. ## Risks - Low risk: this only broadens the workflow trigger set for lockfile regeneration. - The main behavioral change is that patch updates at the repository root now participate in the same CI path as manifest and workspace changes. ## Model Used OpenAI Codex, GPT-5-based tool-using agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.18 |
||
|
|
176a9e8230 |
fix(built-in-agents): allow first-time setup of a needs_setup built-in under board-approval policy (#10129)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Companies can require **board approval for new agents**; built-in agents (e.g. the Reflection Coach / Briefs) are provisioned through the `built-in-agents` service `provision()` > - Some built-in agents are *auto-provisioned* as a hire that, once approved, resolves to an idle agent row whose `adapterConfig` is still empty — status `needs_setup` > - When the board operator then opens that agent's setup dialog and submits the adapter config, `provision()` saw `adapterType`/`adapterConfig` on an already-existing row and classified it as a **reconfiguration**, throwing a dead-end 409: *"Built-in agent adapter changes require board approval before they can be applied."* > - The operator *is* the board, so there was no one left to grant an approval they already implicitly hold — setup could never be completed > - This pull request treats first-time adapter setup of a `needs_setup` built-in as the first-time configuration it actually is, applying it directly while still gating genuine reconfiguration of a live agent > - The benefit is the board can finish setting up an auto-provisioned built-in agent without hitting an unsatisfiable approval wall ## Linked Issues or Issue Description <!-- No public GitHub issue exists; describing the underlying bug in-PR following the bug_report template. --> **What happened?** With "require board approval for new agents" enabled, completing the adapter setup of an auto-provisioned but unconfigured built-in agent (status `needs_setup`, e.g. the Reflection Coach) failed with a 409 — *"Built-in agent adapter changes require board approval before they can be applied."* — even for the board user. Because the operator *is* the board, no additional approver existed, so setup was permanently blocked. Root cause: in `builtInAgentService.provision()`, any request carrying `adapterType`/`adapterConfig` against an existing row was treated as a reconfiguration and gated, regardless of whether that row had ever completed its initial adapter setup. An auto-provisioned hire resolves to an idle row with an empty `adapterConfig` (`needs_setup`), so its very first configuration was misclassified. **Expected behavior** The board can complete first-time setup of an already-sanctioned built-in agent without a fresh approval, matching the behavior when board approval is not required. Genuine reconfiguration of an already-configured (`ready`/`paused`) agent should still require approval. **Steps to reproduce** 1. In a company with `requireBoardApprovalForNewAgents` enabled, have a built-in agent auto-provisioned so its row exists but its adapter is unconfigured (status `needs_setup`). 2. As the board user, open that agent's setup dialog and submit an adapter type + config. 3. Observe the 409 "Built-in agent adapter changes require board approval before they can be applied." with no way for the board to grant the approval. **Deployment mode** Local single-instance / self-hosted (server `built-in-agents` service). ## What Changed - `server/src/services/built-in-agents.ts`: In `provision()`, when the existing built-in row has **not** yet completed adapter setup (`!hasCompleteAdapterConfig(...)`, i.e. `needs_setup`), first-time adapter configuration now applies directly via `ensure()` — the same path used when board approval is not required. The hire that created the row was already sanctioned, so no fresh approval is required. - Reconfiguration of an already-configured (`ready`/`paused`) built-in agent stays gated behind board approval exactly as before, and `pending_approval` rows are handled before the new branch. - `server/src/__tests__/built-in-agents.test.ts`: Added a regression test — under `requireApproval: true`, completing first-time setup of a `needs_setup` built-in returns `approval: null`, transitions the agent to `ready`, and creates **no** approval row. ## Verification ```bash cd server npx vitest run src/__tests__/built-in-agents.test.ts # Test Files 1 passed (1) # Tests 31 passed (31) ``` - New test `completes first-time setup of a needs_setup built-in without a fresh board approval` passes. - Full `built-in-agents.test.ts` suite (31 tests) passes, including existing tests that assert genuine reconfiguration of a configured agent **remains** gated. ## Risks Low risk. The change narrows an over-broad approval gate: it only opens the direct-apply path for rows that have never completed adapter setup (`needs_setup`), determined by the existing `hasCompleteAdapterConfig` predicate that already drives `deriveBuiltInAgentStatus`. Already-configured (`ready`/`paused`) agents, and `pending_approval` rows, are unaffected and still gated. No schema or migration changes. ## Model Used Claude Opus 4.8 (`claude-opus-4-8`), 1M context, extended thinking, with tool use / code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (searched my open PRs and compared patch-ids — no duplicate exists) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.723.0-canary.17 |
||
|
|
e2068319e7 |
fix(interactions): tolerate legacy stored result outcomes so listInteractions can't fail the whole list (#10119)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents and humans coordinate on issues through interaction requests
(confirmations, decisions, task suggestions and more) that are stored
per issue and listed by both the web UI and plugin workers such as chat
gateways
> - `listForIssue` hydrates every stored interaction row by hard-parsing
its persisted `result` blob against the current Zod schema
> - Stored rows outlive code: one live row written by an older build
carried `result.outcome: "withdrawn_by_creator"`, a value no longer in
the enum, and that single row made hydration throw
> - Because the throw happened inside the list mapping, it failed the
entire issue's interaction list — the web thread errored, and every
plugin consumer of `issues.listInteractions` (notification drain, digest
confirmation sweep, pending-ledger reads) failed continuously, so
interaction cards never reached chat surfaces
> - This pull request parses stored `result` blobs tolerantly — a
`parseStoredInteractionResult` helper wrapping `safeParse`, applied to
all five interaction kinds — so an unparseable result degrades to `null`
with a warning instead of failing the whole list
> - The benefit is durable robustness at the storage→hydrate boundary:
legacy or future schema drift in a single row can no longer take down an
issue's entire interaction surface
## Linked Issues or Issue Description
No pre-existing public issue; the underlying problem is described here
following the bug-report template. Related (not a duplicate): Refs #6709
— the creator-withdraw flow it explores matches the legacy outcome value
observed in the wild; whether or not that lineage wrote the row, this PR
is defensive against any such stored-schema drift.
**What happened**
Listing interactions for an issue (`GET /api/issues/:id/interactions` on
the web, or the `issues.listInteractions` plugin RPC) fails for the
entire issue when any single stored interaction row carries a
`result.outcome` written by an older build (observed live:
`"withdrawn_by_creator"`). Downstream plugin consumers that poll this
RPC fail continuously — notification drain, digest confirmation sweep,
and pending-ledger reads.
**Expected behavior**
One legacy/unreadable stored `result` should degrade gracefully — the
interaction still lists with its result treated as absent — rather than
failing the whole issue's interaction list.
**Steps to reproduce**
1. Persist a resolved `request_confirmation` interaction whose
`result.outcome` is not in the current enum (e.g.
`"withdrawn_by_creator"`, as written by an older build).
2. Call `issues.listInteractions` (or `GET
/api/issues/:id/interactions`) for that issue.
3. The call throws `invalid_enum_value` and returns nothing, instead of
returning the remaining rows.
**Version or commit**
master @
canary/v2026.723.0-canary.16
|
||
|
|
e3f8380e70 |
feat(skills): make summarize-status actions-first (#10117)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its built-in Summarizer keeps status slots useful for people overseeing issue trees > - Those summaries need to tell the reader what they must do now to unblock progress > - The existing skill instead imposed rigid Decide:/Review:/Recent work: sections, cost commentary, and restrictive issue-fetch guidance > - This pull request rewrites the summarize-status instructions to lead with 1–3 specific, concrete unblock actions while letting the model use its judgment for the remaining context > - The benefit is a shorter, clearer summary that is immediately actionable without changing slot writes or the streaming status protocol ## Linked Issues or Issue Description Refs #9713 The built-in summarizer currently prioritizes a fixed reporting template over the reader's immediate unblock actions. Summaries should instead open with the 1–3 specific actions the reader needs to take right now, then provide only the context needed to act. This prompt-only update preserves all summary-slot mechanics and protocols. ## What Changed - Rewrote the bundled `summarize-status` skill to open with 1–3 specific, concrete, actionable items needed right now to unblock the work. - Removed the rigid Decide:/Review:/Recent work: template, the Cost discipline section, and the restrictions against fetching issue detail. - Kept slot-write mechanics and the streaming `STATUS`/sentinel protocol unchanged. - Updated all materialized copies and tests for the same skill text: the `SKILL.md` source, regenerated catalog manifest hashes, compiled fallback string, summarizer built-in `AGENTS.md` and routine, summary generation-issue instructions, and the two tests pinning those strings. - Although the diff touches eight files, every file is either the same skill text in another materialized form or a test asserting it. No behavior outside the summarizer's prompt text changes. ## Verification - `pnpm --filter @paperclipai/skills-catalog test` — 20/20 tests pass. - `pnpm exec vitest run server/src/__tests__/summary-slots.test.ts server/src/__tests__/built-in-agents.test.ts` — 46/46 tests pass. - `git diff --check origin/master...HEAD` — clean. - `pnpm exec vitest run server/src/__tests__/summary-slots.test.ts` — 16/16 tests pass after the Greptile consistency fix. - Latest-head GitHub checks — 25 terminal checks, all successful, neutral, or skipped. ## Risks - Low risk: this intentionally changes generated summary wording and prioritization, but does not change APIs, persistence, slot-write behavior, or the streaming protocol. - The branch name contains an internal task identifier because it was pre-created and pre-pushed for this assigned change; the PR title and body do not expose the internal ticket. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.6-sol`, high reasoning mode, with repository, terminal, GitHub CLI, and code-execution tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.15 |
||
|
|
148a5b11f5 |
Route blocked transitions to explicit unblock owners (#10112)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies. > - Issue status transitions determine whether work keeps moving or silently stalls. > - A blocked issue previously could rely on prose alone, leaving the intended unblock owner unstructured and unnotified. > - Existing blocker-attention classification could identify stalled chains, but the signal was not delivered to the board attention feed. > - Blocked transitions also need rollout-safe deduplication so upgrades do not notify for historical issues and repeated processing does not create notification storms. > - This pull request adds structured unblock descriptors, prospective transition timestamps, owner delivery, and board attention routing with focused authorization controls. > - The benefit is that newly blocked work has an explicit, routable next action without weakening company boundaries or allowing agents to inject arbitrary human attention items. ## Linked Issues or Issue Description Related documentation PR: #10094. ### Subsystem affected Cross-cutting: `server/`, `packages/db`, and `packages/shared`. ### Problem or motivation An issue can enter `blocked` without a machine-readable unblock path. Prose-only ownership does not reliably wake the responsible agent or surface human-owned work, while the existing `blockerAttention` classifier is not delivered to an operator-facing attention feed. ### Proposed solution Require new transitions into `blocked` to have unresolved blockers, a pending interaction/approval, or a structured `{ owner, action }` descriptor. Notify an allowed owner once per prospective transition, route human-owned cases to board attention, and leave pre-rollout blocked issues untouched. ### Alternatives considered - Keep prose-only blockers: rejected because ownership remains unroutable. - Backfill all historical blocked issues: rejected because upgrades would create notification storms. - Let agents target arbitrary users or the board: rejected after security review because it creates an attention-injection channel. ### Roadmap alignment Aligns with `ROADMAP.md` → “Enforced Outcomes (watchdogs, recovery actions, review gates)” by making blocked work carry an explicit continuation path. ### Additional context The implementation is prospective-only and deduplicated per blocked transition. Agent-authored descriptors are limited to the acting agent; board actors retain human-owner routing. ## What Changed - Added persisted unblock descriptors and prospective blocked-transition delivery timestamps with an idempotent migration. - Added shared types and validation for board, user, and agent unblock owners. - Enforced valid blocked transitions and same-company owner validation in the issue update route. - Restricted agent-authored descriptors to the acting agent itself, preventing board/user attention injection by compromised agents. - Added one-per-transition agent wake delivery and prospective-only rollout gating. - Routed human-owned blocker attention into the board attention feed. - Added focused tests for validation, prospective delivery, flap deduplication, attention routing, route authorization, and stop-relay compatibility. ## Verification - `pnpm -r typecheck` - `pnpm exec vitest run server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts server/src/__tests__/routable-blocked.test.ts server/src/__tests__/attention-service.test.ts packages/shared/src/validators/issue.test.ts` - `AWS_ACCESS_KEY_ID= AWS_SECRET_ACCESS_KEY= pnpm test:run` - `pnpm build` - `pnpm --filter @paperclipai/db check:migrations` ## Risks - Behavioral shift: new `blocked` transitions without a real blocker, pending governed action, or structured descriptor now return `422`. - Notification abuse is constrained by same-company validation, agent self-only routing, prospective rollout gating, and transition-scoped deduplication. - Migration risk is low: columns are additive, nullable, and use `IF NOT EXISTS`; historical blocked issues are not backfilled or notified. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI with GPT-5.4, reasoning-enabled tool use and code execution. The runtime did not expose a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.14 |
||
|
|
81f47e70a6 |
feat(secrets): thread the acting user into user-scoped secret resolution (#10115)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies > - Its agents and adapters need to resolve secrets through the same governed runtime path that checks ownership and company boundaries > - This change fixes a gap where user-scoped secret resolution could lose the acting-user context before adapter runtime startup > - Without that context, a required user secret could fail closed with responsible_user_missing even though an authenticated user was in scope > - This PR threads the acting user into the user-scoped secret resolution path and keeps the owner boundary explicit > - The benefit is adapter runtime setup can resolve the right credential without broadening access ## Linked Issues or Issue Description Refs #8309 (related: agent secret_ref env drift and binding context) No exact public GitHub issue for this specific behavior. ### Bug report - Problem: two agent-management routes resolved user-scoped secrets without an acting-user binding, so a required `user_secret_ref` could not be resolved before runtime. - Expected behavior: the authenticated acting user should be threaded into user-scoped secret resolution so the owning user secret can be selected safely. - Actual behavior: adapter startup paths failed closed with `responsible_user_missing` even though a user was already in scope. - Steps to reproduce: configure an adapter test-environment or login flow that depends on a user-scoped secret, then invoke it with an authenticated user context that does not carry the acting-user binding into runtime secret resolution. - Impact: the adapter test-environment probe and login path cannot start, so the runtime never reaches the work it was supposed to do. ## What Changed - Added an actor secret-context helper so the server can derive responsible-user context without inventing config-path or binding allowlists. - Added an explicit user-secret mediation mode for runtime config resolution, with an owner-scoped path that resolves by definition plus owner boundary and fails closed when an allowlist is present. - Wired the adapter test-environment route to owner-scoped mediation with an audit-only consumer and kept claude-login on the declared path with its persisted agent identity. - Added and updated tests for the factory, owner-scoped resolver mode, and adapter route coverage. ## Verification - `tsc --noEmit` clean - Factory tests: `authz-secret-context` 5/5 - Service tests: `secrets-service-user-secret-owner-scoped` 5/5, including fail-closed allowlist coverage and company-secret non-regression - Route tests: `agents-adapter-config-user-secret` 5/5, including `responsible_user_missing` and `binding_missing` coverage - Regression suites: `agents` + `secrets` 194/194 ## Risks - A regression in the owner-scoped mediation path could accidentally loosen secret access if the audit consumer or allowlist guard changes. - The change depends on the server-derived responsible user; if auth context regresses, the system should fail closed with responsible_user_missing. - The new mediation mode adds a branch in runtime config resolution, so future changes need to keep declared-mode behavior intact. ## Model Used - OpenAI GPT-5 (Codex tool-use session) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.13 |
||
|
|
f2f168f6a1 |
fix(plugins): seed proactive company scopes before worker setup() + events.subscribe resolver parity (LOOA-695) (#10113)
- [x] I searched the GitHub PR list for similar PRs (dedup search). No open PR touches the proactive `events.subscribe` ordering path; #10103 (merged) is the predecessor whose ordering bug this fixes. ## Thinking Path The gateway worker's outbound push path is permanently dead (`eventSubscriptions: 0`, `notifier.received: 0`, `decisions.delivered: 0`). The plugin loader authorizes the worker's **proactive company scopes only AFTER `startWorker` resolves**, but a proactive plugin issues its one-shot `events.subscribe` calls from `setup()` — which runs *while `startWorker` is still awaiting the worker's initialize response*. So at subscribe time `proactiveCompanyScopes` is still empty → `contextForWorkerMessage` resolves no scope → the governed-access gate rejects every subscribe with `company context is required`. The gateway subscribes once and never retries, so `eventSubscriptions` stays 0 for the worker's life. This is an **ordering bug in the #10103 fix**, not a new method — same #9557 governed-access class as `config.get` (#10092) and `state.get` (#10103). Confirmed live at the 18:21:21Z worker respawn on `3093c5e` (host log), and again at the 19:01:04Z restart (still `events.subscribe: company context is required`, `eventSubscriptions:0`). ## What Changed 1. **Loader ordering** (`plugin-loader.ts`): load `registry.listConfigs(pluginId)` in a new step 4b **before** `startWorker`, and thread the configured company set into `WorkerStartOptions.proactiveCompanyScopes` so the worker handle is authorized *before the child process issues any host call*. The same rows are reused for startup config delivery (step 5b) — no second `listConfigs` round-trip. The runtime config-change path (`routes/plugins.ts`) still refreshes scopes via `setProactiveCompanyScopes` (unchanged). 2. **Handle seed** (`plugin-worker-manager.ts`): `createPluginWorkerHandle` seeds its `proactiveCompanyScopes` set from options at creation, before spawn. 3. **Resolver/gate parity** (`plugin-worker-manager.ts`): `referencedCompanyId(method, params)` now mirrors the SDK gate `requestedCompanyScope` exactly in the functional direction — adds `events.subscribe → params.filter.companyId` (how `ctx.events.on(name, { companyId }, fn)` issues its subscribe), and declines the gate's wildcard cases (`companies.list`, `scopeKind:"company"` without `scopeId`) so proactive access only ever grants a **single explicit configured company, never "all"**. Answers LOOA-693 AC#4 (host/gate extraction parity) in the functional direction. ## Tests New `plugin-worker-manager.test.ts` cases (drive a real worker): - a `setup()`-time `events.subscribe({ filter: { companyId } })` for an options-seeded company is **admitted** (fails on prior code — no options seed, no filter parity); - an unconfigured company stays **denied**; - an unseeded worker stays **denied**. Full `plugin-worker-manager.test.ts` suite: **21 passed**. Server `tsc --noEmit`: clean. All PR CI green (typecheck, server/workspace suites, e2e, build, security scans). ## Risks - **Scope-widening risk (primary).** The change grants proactive host access keyed off configured company rows. Mitigated by: the authorized set is exactly `registry.listConfigs(pluginId).map(companyId)`; wildcard cases (`companies.list`, company-scoped key without `scopeId`) resolve to `null`, never `{ kind: all }`; empty/whitespace ids dropped; an empty config set grants zero proactive access. This is the surface SecurityEngineer must sign off (see Security gate). - **In-invocation path unchanged.** Calls carrying a host-issued `paperclipInvocationId` keep the existing strict single-company match; the proactive branch only applies when there is no invocation id — so no regression to the enforced request path. - **Blast radius.** Loader step 4b is best-effort: a `listConfigs` failure logs and proceeds with an empty seed (fails closed — no push, not a crash), matching today's behavior. ## Model Used Claude Opus 4.8 (`claude-opus-4-8`) via Claude Code (agent: CTO). ## Security gate Touches the company-scope resolution path (same surface as #10103). Routed through **SecurityEngineer review before merge** (tracked on LOOA-696) — must not widen beyond configured companies; in-invocation strict single-company match untouched; wildcard cases deliberately declined in the proactive direction. ## Verification once live - Host log clean of `events.subscribe: company context is required` at worker start - loader logs `eventSubscriptions: N>0` - beat `notifier.received` / `decisions.delivered` move on real issue/approval activity Parent: LOOA-629 (outbound push half of "gateway active"). LOOA-695. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.723.0-canary.12 |
||
|
|
3093c5e694 |
fix(plugin-worker): resolve a company scope for proactive worker→host calls (#10103)
Host authorizes the plugin's configured companies as the worker's proactive scopes, set by the loader right after the #10092 config-delivery step and refreshed on operator config-save. At the single worker→host chokepoint, a no-invocation call (notifier drain, decision reconcile, mirror drain, digest, aging, liveness beat) that references a configured company resolves to that company's scope, so the #9557 governed-access gate admits it. One change covers the full proactive surface (state.*, issues.*, approvals.*, config.get, secrets.resolve, etc.). Safety: never widens beyond configured companies (any other company stays denied); in-invocation calls keep #9557's strict single-company match untouched. Fixes the Slack gateway DM round-trip for LOOA-629. Security review PASS (LOOA-693); non-blocking LOW follow-up tracked in LOOA-694. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a17bee98f2 |
Allow trust-gated direct-parent issue reports (#10098)
## Thinking Path > - Paperclip is the control plane used to coordinate and govern AI-agent companies. > - Agent issue access must preserve company boundaries and trust-policy containment without preventing legitimate task coordination. > - Checked-out standard-trust child runs need a narrow way to report progress directly to their parent issue, but existing authorization treated that report like an arbitrary cross-boundary write. > - Low-trust review runs must remain contained, and stop propagation must not copy potentially untrusted child prose into a higher-trust parent context. > - This pull request adds an audited, one-hop direct-parent comment grant only for standard checked-out runs and a sanitized, idempotent relay for blocked or cancelled child stops. > - The benefit is restored parent/child liveness while retaining least privilege, complete mediation, and low-trust output quarantine. ## Linked Issues or Issue Description ### What happened? A standard-trust agent running a checked-out child issue could not post a progress comment to the direct parent issue because the authorization boundary treated it as an arbitrary cross-issue write. This could stall parent/child coordination. Low-trust review runs also need stop propagation without exposing quarantined child-authored prose. ### Expected behavior A standard checked-out child run may add a comment only to its direct parent issue. The grant must not allow grandparent or sibling access, issue mutation, document writes, reopening, or resuming. Low-trust runs remain denied unless separately mentioned, while blocked/cancelled stops relay only sanitized system metadata once. ### Steps to reproduce 1. Create a parent issue and a child issue assigned to different standard-trust agents. 2. Check out the child issue in a heartbeat run and authenticate as that run. 3. Post a comment to the parent issue and observe the authorization denial before this change. 4. Mark a low-trust child blocked or cancelled and observe that no bounded sanitized parent notification preserves liveness before this change. ### Paperclip version or commit Reproduces on `master` before this PR, including base commit `d36ea13e08`. ### Deployment mode Local dev (`pnpm dev`). ### Installation method Built from source (`pnpm dev` / `pnpm build`). ### Agent adapter(s) involved Not adapter-specific (core authorization and issue-routing behavior). ### Database mode External Postgres in the focused route regression suite; behavior is database-mode independent. ### Access context Agent (bearer API key associated with a checked-out heartbeat run). ### Additional context The implementation deliberately distinguishes a direct-parent report decision from general issue mutation permission and records successful grants in the activity log. ### Privacy checklist - [x] I have reviewed all pasted output for PII, API keys, tokens, company names, and private instance references. ## What Changed - Adds a distinct authorization decision for standard checked-out runs commenting on their direct parent issue. - Keeps low-trust direct-parent reports denied unless an existing explicit mention grant applies. - Forces direct-parent grants to remain comment-only even when a closed parent is unassigned or assigned to the reporting agent. - Audits successful direct-parent report grants in issue activity details. - Adds sanitized, parent-scoped, idempotent system comments and parent wakeups for blocked or cancelled child stops. - Extends the low-trust red-team route suite for allowed parent reports, forbidden upward/sibling writes, closed-parent mutation suppression, and non-laundering stop relays. ## Verification - `pnpm exec vitest run server/src/__tests__/low-trust-red-team-routes.test.ts` — 11 tests passed after the review fix. - `pnpm --filter @paperclipai/server typecheck` — passed after the review fix. - Confirmed the PR changes four files and excludes `pnpm-lock.yaml`, workflow changes, migrations, and unrelated branch commits. ## Risks - This is an authorization behavior change. An overly broad grant could enable cross-boundary writes, while an overly narrow grant could preserve the liveness failure. - The implementation constrains the grant to a standard-trust checked-out run, a direct parent target, and comments only; activity auditing and red-team coverage make regressions observable. - Stop relays intentionally contain only system-generated child identity/status metadata and are deduplicated; child-authored prose is not copied. - SecurityEngineer approval is mandatory before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.5 with reasoning, repository tool use, shell execution, and test execution. The runtime does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
429792f1f3 |
fix(interactions): stop wedging confirmation accept on a terminal workspace_finalize (#10099)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - When an agent finishes work in an execution workspace, the board can confirm the result through an issue-thread interaction (e.g. the "Merged" / mark-done confirmation button on a `request_confirmation`). > - That accept action is gated: it must not race a worktree sync-back (`workspace_finalize`) that is still copying the agent's commits out of the sandbox, or the board could act on a base that hasn't received them yet. > - The gate (`runWorkspaceIsFinalized`) treated the sync-back as "settled" only when the latest `workspace_finalize` op was `succeeded` — so a run whose finalize reached a terminal `failed` state, or died leaving a stale `running` op, was treated as "still syncing" forever. > - Users hit a permanent, misleading `... has not finished syncing its workspace` error and could never click "Merged", even though nothing was syncing and the run had long since ended. > - This PR fixes the settle semantics so the gate blocks only while a sync-back is genuinely pending or in flight, and treats any terminal (or stale-orphaned) finalize as done. > - The benefit is that a failed or abandoned sync-back no longer wedges the human confirmation, while a genuinely in-flight sync-back on a live run still blocks correctly. ## Linked Issues or Issue Description No public GitHub issue exists for this. Describing the bug in-PR (bug report): **What happened** Clicking the "Merged" / mark-done confirmation at the bottom of an issue thread returns an error that the workspace "has not finished syncing its workspace" — but nothing is actually syncing, and the run that created the interaction has already ended. The confirmation is permanently stuck; the only workaround is to merge and mark the task done manually. **Expected behavior** Once the source run's worktree sync-back has finished — whether it succeeded, failed, or was skipped — the confirmation should be acceptable. The gate should block only while a sync-back is genuinely still running on a live run. **Steps to reproduce** Have an agent run reach `workspace_finalize` and end without a `succeeded` finalize (e.g. the sync-back fails, or the run process dies mid-finalize leaving a `running` op). Then attempt to accept the `request_confirmation` interaction it created → 409 "... has not finished syncing its workspace" with no way to proceed. **Paperclip version or commit** Reproduced on the current `master` line (server service); root cause is in `runWorkspaceIsFinalized` in `server/src/services/issues.ts`. **Deployment mode** Local / self-hosted instance (server service). **Root cause** `runWorkspaceIsFinalized` returned `true` only when the latest `workspace_finalize` operation was `succeeded`. A terminal `failed` finalize (the sync-back ran and failed; it will not retry within that run) and a `running` finalize left behind by a dead run both left the gate closed forever. ## What Changed - `runWorkspaceIsFinalized` (server/src/services/issues.ts) now treats a sync-back as **settled** when the latest `workspace_finalize` op reached any terminal status (`succeeded`, `failed`, or `skipped`), instead of only `succeeded`. - A `workspace_finalize` still marked `running` blocks only while its owning run is alive; a `running` record left behind by a terminal/missing run is treated as stale (settled), so a dead run can no longer wedge the gate. - Preserved existing behavior for the other cases: no operations recorded at all → settled; earlier phases recorded but no `workspace_finalize` yet → still blocks (the sync-back hasn't been attempted). - Extracted the run-liveness check into a shared exported helper `heartbeatRunIsTerminalOrMissing` and reused it from the existing `isTerminalOrMissingHeartbeatRun` closure (no behavior change there). - Added a short comment at the confirmation-accept gate (server/src/services/issue-thread-interactions.ts) documenting the relaxed settle semantics. - The dependency-readiness / blocker barrier (`listPendingFinalizeBlockerIssueIds`) is deliberately left unchanged: an automated dependent must not proceed onto a base that never received a blocker's synced-back commits, so a failed finalize keeps that gate closed. Only the human-driven confirmation accept is relaxed. - Added regression tests for: failed finalize, stale `running` finalize on a dead run, and a genuinely `running` finalize on a live run (must still block). ## Verification - `cd server && node_modules/.bin/vitest run src/__tests__/issue-thread-interactions-service.test.ts -t "accept"` → 17 passed (includes the 3 new regression tests), 21 unrelated tests skipped by the name filter. - Manual reasoning walkthrough of `runWorkspaceIsFinalized` for each op-history shape (no ops / earlier-phase-only / terminal finalize / running-on-dead-run / running-on-live-run) confirms the intended block-vs-settle outcome. ## Risks - Low risk and narrowly scoped to the human confirmation-accept gate. The only behavioral change is that a terminal (`failed`/`skipped`) or stale-orphaned `running` finalize now settles the gate instead of blocking forever. - A genuinely in-flight sync-back on a live run still blocks (covered by a regression test), so the accept cannot race commits that are actively being synced back. - The blocker/dependency barrier for automated dependents is unchanged, so no dependent will be advanced onto a base missing a failed blocker's commits. ## Model Used - Provider/model: Claude (Anthropic), **Opus 4.8**, model ID `claude-opus-4-8`, 1M context window. - Capabilities used: extended thinking, tool use (repo inspection, local test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.723.0-canary.11 |
||
|
|
d36ea13e08 |
feat(acpx): stage once per remote ACP session - compatible resume reuses staged runtime, no cross-session credential reuse (#10089)
## Thinking Path > - Paperclip coordinates autonomous agent work across isolated company-scoped sessions > - The ACP remote lane stages workspaces and managed home state inside the sandbox so sessions can resume safely > - If a compatible resume re-staged everything every time, it would waste work and risk inconsistent session reuse behavior > - If an incompatible resume reused the wrong staged runtime, it could cross session boundaries or leak credentials > - This pull request keeps the session fingerprint as the scoping key and adds a staged-runtime cache keyed to that fingerprint > - Compatible resumes now reuse the already staged runtime while incompatible fingerprints stage fresh > - The benefit is faster safe resumes without weakening session isolation or credential separation ## Linked Issues or Issue Description ### Problem or motivation The ACP remote lane currently needs to preserve safe resume behavior without repeatedly re-staging work that is already valid for the same session. The failure mode to avoid is letting one session reuse another session's staged workspace or credentials. ### Proposed solution Keep the session fingerprint as the scoping key and add a staged-runtime cache keyed to that fingerprint. When the fingerprint matches, reuse the already staged workspace and managed home. When the fingerprint changes, stage fresh. ### Alternatives considered - Always restage on resume: safest mechanically, but wastes work and breaks the compatible-resume optimization. - Reuse without fingerprint scoping: too risky because it could cross session boundaries. ### Roadmap alignment This is a narrow implementation change for the ACP remote lane and does not duplicate any broader roadmap item I could find in `ROADMAP.md`. ### Additional context The change preserves the existing session fingerprint contents and codex auth copy-back cadence while adding tests for compatible reuse, incompatible fresh staging, no cross-session credential reuse, and failed-turn eviction. ## What Changed - Added a staged-runtime cache in the ACP remote execution path keyed by the session fingerprint. - Reused the existing staged workspace and managed home for compatible resumes. - Kept incompatible fingerprints on the fresh staging path. - Preserved the existing session fingerprint contents and codex auth copy-back cadence. - Added tests for compatible reuse, incompatible fresh staging, no cross-session credential reuse, and failed-turn eviction. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts` (74/74 pass, including the active-turn lease regression) - `pnpm exec tsc -p packages/adapter-utils/tsconfig.json --noEmit` - Verified the PR touches only `packages/adapter-utils/src/acpx-engine/execute.ts` and `packages/adapter-utils/src/acpx-engine/execute.test.ts` ## Risks - A cache eviction bug could cause an unavailable or partially staged runtime to be reused. - If the fingerprint scoping regressed, a session could incorrectly reuse another session's state. - The change is isolated to the ACP remote lane, but it still affects resume behavior for that path. ## Model Used OpenAI Codex, GPT-5, reasoning-capable coding agent, tool-enabled session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.10 |
||
|
|
a7186dce4b |
fix(plugin-sdk): thread companyId through configChanged + fail-closed cross-tenant guard (LOOA-687) (#10096)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - First-party **plugins** run as isolated workers spawned by the host
`plugin-loader`, reading company-scoped config through a governed
`ctx.config.get(companyId)` channel.
> - The host→worker `configChanged` RPC carries `{ config, companyId }`,
but the SDK dispatch dropped the scope — `onConfigChanged(newConfig)`
was companyId-blind by design — so a **proactive** worker kept a single
worker-global config.
> - #10092 added a startup replay that fans out **every** stored
company's config through `configChanged`. With no deterministic
ordering, a plugin configured for more than one distinct company ends up
running as whichever DB row was delivered last.
> - That is a latent cross-tenant identity/secret confusion bug: one
company's bot token could be applied to another company's traffic.
> - This pull request threads `companyId` through `onConfigChanged` and
adds a fail-closed cross-tenant guard at the SDK layer, so a
single-tenant worker can never silently collapse to a second company's
config.
> - The benefit is that the config-delivery class is fixed at the SDK
boundary — before any genuinely multi-company proactive plugin ships —
without changing today's single-tenant behavior.
## Linked Issues or Issue Description
No public GitHub issue — describing in-PR (hardening / latent security):
**Latent cross-tenant config collapse.** The worker-side `configChanged`
dispatch forwarded only `config` and dropped `companyId`, so a proactive
plugin kept a single worker-global config. #10092's startup replay
delivers every configured company's config sequentially with no `ORDER
BY`, so a plugin with configs for more than one distinct company would
apply a nondeterministic last-write-wins global config (one tenant's
credential applied to another's traffic).
- Builds on and must merge after #10092.
- Not exploitable today: the only proactive consumer (the chat gateway)
has single-tenant config rows, so last-write-wins is a no-op. This is a
hardening pre-condition before any multi-company proactive plugin ships.
## What Changed
- **Thread scope through:** `onConfigChanged(newConfig, context)` with a
new exported `PluginConfigChangeContext { companyId }`. Backward
compatible — the second arg is optional; existing single-arg
implementations are unaffected.
- **Fail-closed cross-tenant guard** (`worker-rpc-host.ts`): a
single-tenant plugin that receives `configChanged` for a second,
distinct company with a *different* config is rejected with the new
`PLUGIN_RPC_ERROR_CODES.CROSS_TENANT_CONFIG` instead of silently
overwriting the applied tenant's config. Idempotent replays of the
*same* config under a different scope row remain allowed.
- **Opt-in `multiCompanyConfig: true`** on the plugin definition for
plugins that genuinely serve multiple companies from one worker (keying
per-company state on `context.companyId`); the guard is bypassed for
those.
- **Deterministic `ORDER BY companyId`** on `registry.listConfigs`, so
the startup replay binds a single-tenant worker to a stable company
across restarts.
- **Loader visibility:** a `CROSS_TENANT_CONFIG` rejection is logged at
`warn` (was best-effort `debug`) so the misconfiguration is surfaced.
- **Regression test**
(`packages/plugins/sdk/tests/worker-rpc-host.test.ts`): two distinct
companies delivered via the startup-replay path fail closed and stay
bound to the first company; an idempotent same-config replay under a
different scope row is allowed; a `multiCompanyConfig` plugin receives
per-company config with the correct `context.companyId`.
## Verification
- SDK `tsc --noEmit`: clean.
- SDK vitest `worker-rpc-host.test.ts`: 7/7 pass (incl. 3 new). The
two-distinct-company case **fails against pre-fix code** and passes
after the fix.
- #10092 embedded-postgres `plugin-config-startup-delivery.test.ts`: 3/3
pass (unaffected by the new `ORDER BY`).
- Full server `tsc --noEmit` against this SDK: clean.
## Risks
- **Low functional risk.** The second `onConfigChanged` arg is optional
and existing implementations are unchanged. Today's single-tenant
gateway keeps working — idempotent same-config replays are explicitly
allowed, so the go-live is preserved.
- **Behavioral shift on misconfig:** a genuinely multi-company plugin
that has NOT opted into `multiCompanyConfig` now fails closed
(`CROSS_TENANT_CONFIG`) rather than silently collapsing to one tenant.
This is the intended safer default; opt in with `multiCompanyConfig:
true` to serve multiple companies from one worker.
- **Not in scope (residual).** Per-company workers/connections for a
genuinely multi-company gateway increase resource use and are tracked
separately (ties into the #10092 fan-out/timeout follow-up). This PR
fixes the class and fails closed; it does not build multi-tenant
connection management.
## Model Used
Claude — Anthropic `claude-opus-4-8` (Opus 4.8), extended thinking, with
tool use / code execution via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change and contains no internal
Paperclip ticket id — branch predates this rule; not renaming an open PR
mid-review
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
doc surface; internal SDK/host behavior only
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: anicca <annica@Michaels-Mac-Studio.local>
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.723.0-canary.9
|
||
|
|
1ebf5254b6 |
fix(plugin-loader): deliver stored config to freshly-started plugin workers (#10092)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - One capability is first-party **plugins** that run as isolated workers spawned by the host `plugin-loader`, reading company-scoped config through a governed `ctx.config.get(companyId)` channel. > - A **proactive** plugin (e.g. a chat gateway that opens a Slack Socket Mode connection at startup) does its company work from `setup()`, where there is **no company-scoped invocation** — so `ctx.config.get()` is rejected with `company context is required`. > - The worker swallows that error and falls back to its default (feature-off) config, so the plugin comes up **inert** even though correct config exists in the database. > - This is a regression from #9557 ("governed access contracts"), which changed `plugin-loader.ts` `activatePlugin` from loading stored config into the worker bootstrap to `const config = {}`. > - This pull request replays each configured company's stored config to the freshly-started worker over the **same `configChanged` host→worker path an operator config-save already uses**. > - The benefit is that proactive plugins receive their config on worker start (both server boot and operator enable) without weakening the governed-access surface. ## Linked Issues or Issue Description No public GitHub issue — describing in-PR (bug): **Bug.** After a proactive plugin's worker spawns, it never receives its stored config. Governed access (`packages/plugins/sdk/src/host-client-factory.ts`) only resolves `config.get` inside a company-scoped invocation (event/action/tool, or explicit `params.companyId`). Proactive plugins operate from `setup()` where no such scope exists, so `config.get()` fails with `company context is required`, the worker falls back to defaults, and the feature stays disabled despite valid DB config. - Regression introduced by #9557. - Related follow-up (latent multi-company hardening): #10096. ## What Changed - `plugin-registry.ts`: add read-only `listConfigs(pluginId)` returning all stored company config rows for a plugin (scoped `where eq(pluginConfig.pluginId, pluginId)`). - `plugin-loader.ts`: after the worker starts in `activatePlugin`, replay each company's stored config through the existing `configChanged` host→worker RPC — one `{ config, companyId }` per row, the same payload shape as the operator config-save path in `routes/plugins.ts`. Best-effort and idempotent; covers both server-boot `loadAll` and operator enable. - test: DB-backed `plugin-config-startup-delivery.test.ts` covering `registry.listConfigs` completeness and cross-plugin isolation. ## Verification - `tsc --noEmit` on `@paperclipai/server` — clean. - New `plugin-config-startup-delivery.test.ts` (embedded-postgres, 3 cases) — pass. - Full PR CI green: typecheck, all server/e2e/serialized test shards, build, canary dry-run, verify, and the security scanners (Snyk, Socket, Superagent, Greptile). ## Risks - **Low functional risk.** Adds an outbound host→worker push that mirrors the already-shipped operator-save path. A worker without an `onConfigChanged` handler (or momentarily unavailable) simply keeps the runtime `ctx.config.get(companyId)` model. - **Startup fan-out.** One `configChanged` per configured company at activation (sequential, default RPC timeout). `plugin_config` rows are writable only by instance-admins, so fan-out size is operator-controlled — not a remote surface. - **No secret-handling change.** `configJson` is delivered as-is, exactly as `config.get`/operator-save already deliver it. No new secret sink; catch-blocks log only ids + `err.message` at debug, never `configJson`. - **Latent multi-company behavior (pre-existing, not introduced here).** The worker-side `configChanged` dispatch forwards only `config` (drops `companyId`), and `listConfigs` has no `ORDER BY`, so a plugin configured for **more than one** company would apply a nondeterministic last-write-wins global config. This is existing SDK behavior — operator-save already pushes into the same handler — and is **not reachable by the single-company consumer this fix targets**. Greptile flagged this shape (4/5). It is tracked and fixed as a separate, non-blocking hardening PR (#10096): thread `companyId` through `onConfigChanged`, deterministic ordering, bounded fan-out. ## Model Used Claude — Anthropic `claude-opus-4-8` (Opus 4.8), extended thinking, with tool use / code execution via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes` / `Refs` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change and contains no internal Paperclip ticket id — branch predates this rule; not renaming an open PR mid-review - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — no doc surface; internal SDK/host behavior only - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — 4/5; two latent multi-company items triaged as non-blocking and fixed in follow-up #10096 (see Risks) - [x] I will address all Greptile and reviewer comments before requesting merge — addressed: triaged as non-blocking follow-up in #10096 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.8 |
||
|
|
e1cc63328a |
feat(acpx): per-adapter managed-home seed (x3) + Codex auth copy-back for remote ACP lane (#10073)
## Thinking Path > - Paperclip is a control plane that orchestrates AI agents and adapter execution for human operators. > - Agents run across local and remote execution contexts, and reliability in remote sessions depends on consistent adapter bootstrapping. > - The ACP path must prepare per-adapter runtime homes so managed credentials and config are available in sandboxed runner environments. > - Before this change, the new remote ACP lane did not yet consistently stage managed-home paths for all affected adapters or restore Codex auth state on teardown. > - We added a shared per-adapter seam in the ACP engine, then wired Codex, Claude, and Gemini remote lanes to seed managed homes and remap to in-sandbox locations. > - Codex additionally reuses the existing atomic auth restore flow to copy auth back on teardown, matching CLI behavior. > - This improves remote runner parity with existing CLI behavior and avoids credential drift in shared-code-path executions. ## Linked Issues or Issue Description - This change continues the ACP remote managed-home work by completing per-adapter remote bootstrapping and Codex auth restore behavior for the remote ACP lane. - It specifically covers: `acpx-engine`, `codex-local`, `claude-local`, and `gemini-local`. - Related prior work in this repo: PR #10070. ## What Changed - Add a per-adapter remote managed-home seam (`prepareRemoteManagedHome`) in `acpx-engine` and thread it through ACP execution options. - Implement Codex ACP preparation to: - stage `CODEX_HOME` (auth/config/skills) into a sandboxed remote home, - repoint `CODEX_HOME` to the in-sandbox path, - and wire teardown copy-back through the existing managed-auth restore path. - Implement Claude ACP preparation to seed a sanitized `config-seed` into `CLAUDE_CONFIG_DIR` and remap that directory to the sandbox root. - Implement Gemini ACP preparation to seed `~/.gemini/skills`, set `HOME` to managed runtime root, and preselect API-key auth in `settings.json`. - Keep local and runner-less ACP→CLI behavior unchanged by only invoking the remote managed-home seam when running in remote ACP mode. - Preserve existing authorization and activity boundaries in the shared engine and adapter layers. ## Verification - `git log --oneline origin/master..origin/feat/acp-remote-managed-home-seed` confirms only the expected 4 commits. - `tsc --noEmit` is clean in `adapter-utils`, `codex-local`, `claude-local`, and `gemini-local`. - Vitest selection used during validation passed (120 tests across the ACP-related suites). ## Risks - If remote sandbox teardown occurs after token rotation but before restore timing, Codex credentials can become stale and require re-auth on next startup. - Partial provisioning of managed-home assets would cause adapter bootstrap failures in runner-backed ACP sessions. - This change is scoped to execution-path behavior; it should not affect CLI behavior. ## Model Used None — human-authored. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.7 |
||
|
|
8a3cc86531 |
feat(acpx): stage workspace + route in-sandbox cwd for remote ACP lane (#10070)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The shared ACP engine
(`packages/adapter-utils/src/acpx-engine/execute.ts`) is responsible for
launching local and remote agent processes via the ACP protocol
> - On runner-backed remote sandbox (Daytona) targets, `buildRuntime`
never crossed the CLI's staging seam: it never called
`prepareAdapterExecutionTargetRuntime`, left `runtimeRootDir: null` in
both the paperclip and process-session bridges, and handed the agent the
**HOST filesystem path** as the `session/new` cwd — meaning
Claude/Gemini silently operated on a path that does not exist inside the
sandbox (Codex additionally crashes on its HOST home path, addressed in
a follow-up PR)
> - The fix must cross the staging seam for remote sandboxes, thread the
real `runtimeRootDir` through both bridges, and bind the in-sandbox
workspace path as the session cwd — without touching local ACP runs or
the runner-less ACP→CLI fallback
> - This pull request introduces a `stageAcpRemoteRuntime` helper that
calls `prepareAdapterExecutionTargetRuntime` for runner-backed remote
runs, captures `{ workspaceRemoteDir, runtimeRootDir, assetDirs,
restoreWorkspace }`, and reuses the in-sandbox `workspaceRemoteDir` as
the single `sessionCwd` across `session/new`, the fingerprint,
compatibility check, persistence, `ensureSession`, the process-session
bridge cwd, and the error path
> - The benefit is that remote ACP runs now operate in the correct
in-sandbox cwd and receive a non-null `runtimeRootDir` in both bridges —
fixing silent wrong-cwd degradation for Claude/Gemini on Daytona
targets; this is PR 1 of 3 and seeds no credential material
## Linked Issues or Issue Description
No public GitHub issue exists for this change. Inline description
follows the feature request template:
### Subsystem affected
packages/adapters — agent adapter implementations
### Problem or motivation
The shared ACP engine
(`packages/adapter-utils/src/acpx-engine/execute.ts`) never crossed the
CLI's staging seam on runner-backed remote sandbox targets. It never
called `prepareAdapterExecutionTargetRuntime`, always passed
`runtimeRootDir: null` to both the paperclip and process-session
bridges, and handed the agent the **HOST filesystem path** as the ACP
`session/new` cwd. As a result, Claude and Gemini silently operated on a
cwd that does not exist inside the sandbox; Codex crashed with a fatal
error on the HOST `CODEX_HOME` path (that crash is in a follow-up PR).
### Proposed solution
Gate on `usesRunnerBackedSandbox` (`kind === "remote" && transport ===
"sandbox" && runner`). For runs that pass the gate, call
`prepareAdapterExecutionTargetRuntime` via a new `stageAcpRemoteRuntime`
helper that ships the workspace into the sandbox and captures `{
workspaceRemoteDir, runtimeRootDir, assetDirs, restoreWorkspace }`.
Thread the real `runtimeRootDir` into both bridges. Bind a single
`sessionCwd` (the in-sandbox `workspaceRemoteDir`) and use it at every
cwd-keyed session site (`session/new`, fingerprint, compat, persist,
`ensureSession`, process-session bridge, error path) so a warm/resumable
session is reused rather than invalidated. For local runs and the
runner-less ACP→CLI fallback, `sessionCwd` resolves to the HOST cwd —
byte-identical to the previous behavior.
### Alternatives considered
Patching each per-adapter bridge individually — rejected because the bug
is in the shared engine layer and the fix belongs there so all three
adapters (Codex, Claude, Gemini) benefit without per-adapter
duplication.
### Roadmap alignment
Internal correctness fix enabling remote ACP to work as designed; no new
user-facing features. This is PR 1 of 3 in a sequential chain: PR 1
(this PR) stages the workspace and routes the cwd; PR 2 adds per-adapter
home seeding and copy-back; PR 3 wires session-lifecycle restore.
## What Changed
- **`packages/adapter-utils/src/acpx-engine/execute.ts`** — added
`stageAcpRemoteRuntime` helper that calls
`prepareAdapterExecutionTargetRuntime` for runner-backed remote
sandboxes and returns `{ sessionCwd, runtimeRootDir, stagedRuntime }`;
`buildRuntime` now uses this helper to derive `sessionCwd` (in-sandbox
`workspaceRemoteDir` for remote; HOST cwd unchanged for
local/runner-less) and threads the real `runtimeRootDir` to both the
paperclip bridge and process-session bridge; `stagedRuntime` is stashed
for the follow-up credential PR
- **`packages/adapter-utils/src/acpx-engine/execute.test.ts`** — new
engine-level unit tests: staging seam crossed with no credential asset,
non-null `runtimeRootDir` in both bridges, in-sandbox `session/new` cwd,
warm-handle reuse after the cwd change, local-unchanged; 60 tests green
- **`packages/adapters/codex-local/src/server/acp.test.ts`** — new
per-adapter test: runner-backed remote asserts `ensureSession` cwd ==
`workspaceRemoteDir`; runner-less sandbox falls back to CLI
- **`packages/adapters/claude-local/src/server/acp.test.ts`** — same
per-adapter coverage for Claude
- **`packages/adapters/gemini-local/src/server/acp.test.ts`** — same
per-adapter coverage for Gemini
## Verification
- `adapter-utils` typecheck clean; `codex/claude/gemini-local` typecheck
clean
- Engine units (`acpx-engine/execute.test.ts`): 60/60 green — staging
seam crossed with no credential asset, non-null `runtimeRootDir` to both
bridges, in-sandbox `session/new` cwd, warm-handle reuse after cwd
change, local unchanged
- Per-adapter ACP test suites (`codex/claude/gemini-local`
`acp.test.ts`): 103 tests green — runner-backed remote asserts
`ensureSession` cwd == `workspaceRemoteDir`; runner-less sandbox falls
back to CLI
- CI green on PR (in progress)
## Risks
This is PR 1 of 3 in a strictly sequential chain; it seeds **no
credential material** (no `assets`, no `installCommand`). The
per-adapter home seeding is deferred to PR 2, which consumes the
`stagedRuntime` object stashed here. The `restoreWorkspace` callback is
carried on `stagedRuntime` for PR 3's session-lifecycle wiring (see the
`stageAcpRemoteRuntime` function comment). Local ACP runs and the
runner-less ACP→CLI fallback are untouched — `sessionCwd` resolves to
the HOST cwd for those paths, preserving existing behavior. The
`stageAcpRemoteRuntime` helper is gated on `usesRunnerBackedSandbox`, so
there is no regression risk for local or CLI-lane runs.
## Model Used
Claude Sonnet 4.6 (`claude-sonnet-4-6`) via Paperclip ACPX engine —
extended context, tool use enabled, co-authored with Paperclip agent
orchestration.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`,
`feat/...`) and contains no internal Paperclip ticket id or
instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Harold Kim <harold@paperclip.ing>
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.723.0-canary.6
|
||
|
|
5a00040c4c |
feat(plugin-sdk): interactions/approvals respond + attachment read capabilities for the chat gateway (#10066)
Adds the remaining 5 plugin capabilities + 7 worker→host RPC methods (interactions read/respond, approvals read/respond, attachment read) needed by the Slack chat gateway plugin (v0.5.0) to pass manifest capability validation and load. - Security review: PASS (LOOA-642) after the viewer-role privilege-escalation blocker (LOOA-648) was fixed on this branch (requireActiveHumanMember now rejects viewer on impersonation write-paths, matching assertCompanyAccess). - CI: Build, Typecheck, all server suites (3/3 + serialized 4/4), workspaces, e2e shard 2/2, and all security scanners (Snyk/Socket/Superagent/Greptile/security-review) green. - One e2e flake (signoff-policy 'non-participant cannot advance stage') is unrelated: it exercises execution-policy stage advancement (routes/issues.ts, untouched by this PR) and failed on a heartbeat_run_events FK race + 409 checkout conflict. Unblocks LOOA-629 (Slack gateway go-live) and the interview-ask feature. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.5 |
||
|
|
0ef3b320c7 |
Harden POST /plugins/install: canonicalize localPath for all instances; bundled-only floor for managed instances (#10067)
**Builds on** #10058 — managed detection keys off the *presence* of the `PAPERCLIP_MANAGED_CONFIG` env var that PR introduces, deliberately never its parsed body. **Summary.** Two layered hardenings of the plugin install route. (1) For **all** instances: `localPath` installs previously skipped the package-name validation entirely; the path is now null-byte-checked, resolved absolute, `realpath`'d (collapsing `..` traversal and symlinks), and required to be an existing directory before the loader ever sees it. (2) For instances running under a managed hosting control plane (detected by the *presence* of `PAPERCLIP_MANAGED_CONFIG` — deliberately never its body, so a corrupted document cannot widen the surface): registry/npm installs return 403, and `localPath` installs must canonicalize to inside the bundled plugin catalog root (`packages/plugins`) — a positive allowlist enforced in code at the route, independent of any flag value. Self-hosted behavior is otherwise unchanged. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The plugin system lets instance admins install plugins from a registry or from a local filesystem path, and plugin installation is code execution on the host > - The `localPath` branch of `POST /plugins/install` skips the validation applied to registry installs; the raw path reaches the plugin loader without canonicalization > - Separately, instances operated by a managed hosting control plane must constrain installs to the bundled plugin catalog, because there the host belongs to the operator, not the tenant > - This pull request canonicalizes and validates `localPath` for all instances, and adds a bundled-only install floor for managed instances > - The benefit is a smaller install-route attack surface everywhere, and a positive code-enforced allowlist where the operator owns the machine ## Linked Issues or Issue Description No public issue exists; `bug_report` template fields for the validation gap this PR fixes: - **What happened:** `POST /plugins/install` with `localPath` set bypasses the package-name validation entirely; the un-canonicalized path (relative segments, symlinks, no existence check) is handed straight to the plugin loader. - **Expected behavior:** path installs are validated like registry installs — null-byte-checked, resolved absolute, `realpath`'d, and required to be an existing directory before the loader sees them. - **Steps to reproduce:** as an instance admin, call `POST /plugins/install` with a `localPath` containing `..` traversal or a symlink pointing outside any plugin directory; observe the loader receives the raw path. Exploitability is bounded (the route already requires instance admin), so this is hardening of an admin-only surface rather than an open exploit. - **Version:** current `master`. The managed-instance bundled-only floor layered on top is new behavior (motivation: on managed hosting, arbitrary plugin install is arbitrary code execution on operator infrastructure), aligned with the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed - New `server/src/services/plugin-install-guard.ts` — three pure primitives: managed detection (presence-based), path canonicalization (null-byte check → absolute resolve → `realpath` → must be an existing directory), and segment-based containment in the bundled plugin catalog root. - Route enforcement in `server/src/routes/plugins.ts`: npm/registry installs return 403 on managed instances; `localPath` installs are canonicalized on every instance and, on managed instances, must land inside the bundled catalog root. - The plugin loader now receives the canonical path instead of the raw request string. ## Verification - 15 guard unit tests (`server/src/__tests__/plugin-install-guard.test.ts`): traversal, symlink escape, null byte, file-vs-directory, string-prefix sibling root. - 13 route security tests (`server/src/__tests__/plugin-install-route-security.test.ts`): 403 matrix on managed instances + self-hosted happy paths. - 36 existing plugin route authz tests green (`server/src/__tests__/plugin-routes-authz.test.ts`). - Server `tsc --noEmit` clean. ```bash cd server pnpm vitest run src/__tests__/plugin-install-guard.test.ts src/__tests__/plugin-install-route-security.test.ts src/__tests__/plugin-routes-authz.test.ts pnpm exec tsc --noEmit ``` ## Risks - Managed instances: npm/registry installs and out-of-catalog `localPath` installs now return 403 — intended new behavior, enforced in code rather than configuration. - All instances: `localPath` installs that previously pointed at nonexistent paths or non-directories now fail with 400 before reaching the loader (previously the loader failed later, less safely). Symlinked deployment layouts are handled by canonicalizing both sides of the containment check. - Self-hosted npm install path is unchanged. Low residual risk. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.4 |
||
|
|
04e070bf45 |
docs(skill): require agents to claim only monitors they actually scheduled (#10064)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work, and the shared `skills/paperclip/SKILL.md` is the behavioral
contract every managed agent follows each heartbeat.
> - Issue continuation between heartbeats depends on real, persisted
state: an issue only auto-resumes when it has a scheduled **issue
monitor** (`monitorNextCheckAt` + an execution-policy `monitor` block)
that the server's `tickDueIssueMonitors` scheduler polls and re-wakes
via `issue_monitor_due`.
> - A run/heartbeat is an ephemeral execution window — nothing keeps
"watching" after it exits — but the skill never said this, so agents
narrated a "watcher in this run" as if a live subscription existed.
> - That gap produced a concrete user-facing failure: an agent claimed
"a watcher in this run wakes me when CI + Greptile complete," then the
run ended with no monitor scheduled and nothing ever resumed, leaving
the user unsure whether a watcher existed at all.
> - This PR closes the gap by documenting what a monitor actually is and
adding hard rules so agents only claim a watcher they have actually
scheduled, describe it in checkable terms, and never imply a live
watcher on a task they mark `done`.
> - The benefit is that agent narration stays consistent with the
disposition guard and recovery classifier that already enforce these
paths in state, so users get accurate expectations about whether and
when a task will resume.
## Linked Issues or Issue Description
This is a documentation-only change to a shared agent skill, so no code
issue is required. The underlying problem it addresses:
**Problem or motivation** — Agents were telling users that a "watcher in
this run" would wake them when external checks (CI, Greptile) finished,
when no persisted issue monitor had been scheduled. Because a heartbeat
is ephemeral, no such watcher exists after the run exits, so the task
silently never resumed and the user was left confused about what would
happen next.
**Proposed solution** — Document, in the shared skill, exactly what an
issue monitor is (durable `monitorNextCheckAt` + execution-policy
`monitor` block, polled by `tickDueIssueMonitors`, re-woken via
`issue_monitor_due`) and add rules that agents may only claim a
watcher/monitor after actually scheduling one, must describe it in
checkable terms (kind / next check / timeout / attempts), and must never
imply a live watcher on a task being marked `done`.
**Alternatives considered** — Enforcing purely in server state (the
disposition guard and recovery classifier already reject
`in_review`/parked issues without a real wake path). That enforcement
exists but does not stop an agent from *narrating* a non-existent
watcher in a comment; aligning the skill guidance with the existing
state enforcement is the missing piece.
## What Changed
- Added a **"Monitors and Watchers (say only what you actually
scheduled)"** subsection to `skills/paperclip/SKILL.md` explaining that
a watcher does not live inside a run, and that only a persisted issue
monitor can auto-resume an issue (with the concrete fields and the
`tickDueIssueMonitors` / `issue_monitor_due` polling path).
- Added three behavioral rules: only claim a monitor after scheduling
one (and how to schedule/confirm it via `PATCH /api/issues/{id}` and
`monitor/check-now`); describe monitors in checkable terms; never imply
a live watcher on a task marked `done`.
- Cross-referenced the rule from the **Critical Rules** list.
- Tightened the final-disposition checklist so `in_review` /
`in_progress` continuation requires a real, non-null
`monitorNextCheckAt` rather than a merely described one.
## Verification
- Docs-only change to `skills/paperclip/SKILL.md`; no code paths are
affected.
- Confirmed every identifier referenced in the new text is real in the
codebase: `monitorNextCheckAt`, `monitorScheduledBy`,
`executionPolicy.monitor`, `tickDueIssueMonitors`, and the
`issue_monitor_due` wake reason.
- Rendered the Markdown to confirm the new subsection and the Critical
Rules bullet display correctly and links resolve within the document.
- `git diff` confirms the change is limited to the single skill file (14
insertions, 2 deletions).
## Risks
Low risk. This is guidance text in a shared agent skill with no runtime
or schema impact. Worst case is stylistic wording that can be refined in
a follow-up; it cannot break builds, migrations, or behavior. It
strengthens (never loosens) the existing disposition guarantees.
## Model Used
Claude Opus 4.8 (model id `claude-opus-4-8`, 1M-context variant) running
in an agent harness with extended thinking and tool use (file edit,
shell, git, GitHub CLI).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [ ] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
62e367c1b2 |
fix(ui): rewrite recovery and blocked-notice copy in plain language (#10065)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The issue-detail UI shows recovery cards and blocked/parked notices when a task loses its next step — a run finished with no disposition, a task is stranded, work is blocked behind other tasks, or an assigned item sits in the backlog > - That copy was written in the scheduler's internal vocabulary — "Corrective wake queued", "Graph Liveness", "lost a live action path", "the responsible" — which describes Paperclip's internals rather than the user's situation > - Users seeing these cards report having no idea what the card means or what they are supposed to do > - This pull request rewrites the user-facing copy in plain language and adds explicit calls to action that match the options in the card's Resolve menu > - The benefit is that a non-expert operator can read a recovery or blocked notice and immediately understand what happened and which action to take next ## Linked Issues or Issue Description No public GitHub issue exists for this; describing the problem here (bug-report format): - **What happened:** Recovery action cards and blocked notices render internal jargon, e.g. the headline "Paperclip detected this task lost a live action path. A recovery owner needs to act.", the chip "Corrective wake queued", the kind label "Graph Liveness", and phrases like "Comments still wake the responsible". Status values also appear as raw code literals (`in_progress`, `todo`). - **Expected behavior:** These notices should tell a normal user, in plain language, what happened and what to do next (retry the task, mark it done, send it for review, or record a blocker). - **Impact:** Operators stall on tasks that only need a simple disposition because the UI doesn't tell them that's what is being asked. Related prior work: #9417 (merged) made the reopen-suppressed blocked message explicit; this PR extends the same plain-language treatment to the rest of the recovery and blocked-notice copy. ## What Changed - Recovery card headlines for `missing_disposition`, `stranded_assigned_issue`, and `issue_graph_liveness` now say what Paperclip found and name the concrete next steps ("try the task again, mark it done, or send it for review") matching the card's Resolve menu. - The `issue_graph_liveness` kind label "Graph Liveness" is now "Task Needs Next Step", and the "Wake" metadata row is now "Follow-up". - Wake-policy chips describe actual behavior: "An agent will be asked to choose the next step" (was "Corrective wake queued"), "Board will decide", "Manual follow-up needed", "Repair needed before retry", "Check scheduled". - Blocked/waiting/parked notices say "the assignee" instead of "the responsible" / "responsible agent", and "notify" instead of "wake". - The still-needs-a-next-step notice drops raw `in_progress` code literals and keeps a plain-language option list (mark done or cancelled, send for review, record what is blocking it, delegate follow-up). - Parked-backlog notice renders "To do / In progress" as plain labels instead of code literals. - Component tests updated to pin the new copy and the successful-run example options. ## Verification - `cd ui && npx vitest run src/components/IssueRecoveryActionCard.test.tsx src/components/IssueBlockedNotice.test.tsx src/components/IssueAssignedBacklogNotice.test.tsx src/components/IssueChatThread.test.tsx` — 4 files, 126 tests, all passing. - Copy-only review: the diff touches display strings, one label map entry, and test assertions; no control flow, props, or identifiers change. ## Risks - Low risk — user-facing strings and test updates only. No behavior, API, or schema changes. The only functional surface is that anything keying off the displayed text (e.g. screenshots, external docs) will show the new wording. ## Model Used - Claude (Anthropic) — Claude Fable 5, model ID `claude-fable-5`, extended thinking enabled, running in Claude Code with agentic tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no docs reference this copy) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0f898f95af |
Render managed experimental settings as locked, with a 'Managed by Paperclip Cloud' badge (#10061)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The experimental settings page renders one interactive toggle per feature from the settings API > - On managed instances some values are enforced by the hosting control plane, and the API now reports those keys as managed > - Rendering enforced values as live toggles misleads users: the click appears to work, and the value silently snaps back > - This pull request renders managed keys as locked toggles with a "Managed by Paperclip Cloud" badge and guards the handlers so no PATCH can be emitted > - The benefit is UI honesty on managed instances, with self-hosted responses rendering exactly as before ## Linked Issues or Issue Description Builds on #10058 — renders the per-key `managedKeys` metadata #10058 adds to settings responses (typing shared from #10058 at rebase). No public issue exists; `feature_request` template fields: - **Problem or motivation:** on managed instances users see fully interactive toggles for settings the control plane enforces; changes appear to apply and never do, with no explanation. - **Proposed solution:** disabled toggle + badge + guarded handler driven by the settings response's managed-key metadata; the ~17 uniform setting cards are extracted into one shared component with copy, aria-labels, and patch payloads preserved verbatim. - **Alternatives considered:** hiding managed settings entirely (users lose sight of the effective value and why it is fixed); tooltip-only hints on still-active toggles (doomed PATCHes are still emitted and stripped server-side). - **Roadmap alignment:** supports the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed - When the settings API reports a feature key as managed (`managedKeys` from the managed-config overlay), the experimental settings page renders that toggle disabled with a badge and a guarded handler, so a click can never emit a PATCH. Previously, managed-instance users saw fully interactive toggles they could never actually change. Self-hosted responses (no `managedKeys`) render exactly as before. - The ~17 copy-pasted uniform setting cards are extracted into one `ExperimentalToggleCard` component with titles, descriptions, footnotes, aria-labels, and patch payloads preserved verbatim; the two bespoke cards get inline managed handling (the managed auto-recovery toggle also cannot open its preview dialog). - `ui/src/api/instanceSettings.ts` response typing now uses the shared `InstanceExperimentalSettingsWithManaged` / `ManagedSettingMetadata` types from #10058; `ui/src/pages/InstanceExperimentalSettings.tsx` locked rendering + card extraction; tests. ## Verification - 24 page tests (20 existing unmodified + 4 new: locked badge with no PATCH while unmanaged keys stay editable; managed auto-recovery opens no dialog; an open recovery preview closes with no PATCH when a refresh marks auto-recovery managed; self-hosted unaffected): `pnpm --filter @paperclipai/ui exec vitest run src/pages/InstanceExperimentalSettings.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` clean ## Risks - Low risk. UI-only change; no server or API behavior changes. Self-hosted responses carry no `managedKeys`, so the page renders exactly as before there. The card extraction preserves copy, aria-labels, and patch payloads verbatim, covered by the 20 pre-existing page tests passing unmodified. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c4cdcc4826 |
Generalize bundled plugin provisioning: ensureBundledKubernetesPlugin → ensureBundledPlugins (#10063)
**Builds on** #10058 — reads `plugins.autoInstall` from the parsed managed-config contract #10058 introduces (the interim `readManagedPluginAutoInstall` shim is retired at rebase). **Summary.** Boot-time bundled-plugin provisioning becomes catalog-driven. A new bundled-plugin catalog lists the sandbox providers shipped in-tree (keys like `kubernetes`, `daytona` → plugin key + path under the catalog root). Managed instances read `plugins.autoInstall` from `PAPERCLIP_MANAGED_CONFIG`; unknown keys or paths escaping the catalog root (symlinks resolved) **throw before listen** — a managed instance refuses to start rather than boot half-provisioned. Installation keeps today's mechanism: an in-process, fail-safe `loader.installPlugin({ localPath })` under a system actor — no HTTP route, no user, no role widening. Self-hosted boot is unchanged (kubernetes bundle only, existing env override honored, install failures still log-and-continue). **Semantics.** A plugin already present in any non-uninstalled state is skipped, so an operator-disabled plugin is never silently re-enabled; managed mode reinstalls soft-uninstalled bundles (the control plane owns provisioning); removal from the autoInstall list never auto-uninstalls. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox-provider plugins ship in-tree, but boot-time provisioning is hard-coded to exactly one of them (Kubernetes) via a bespoke function > - On managed hosting, tenant users have no install privileges, so any bundled plugin that is not provisioned at boot is unusable > - Widening install routes or granting roles to fix that would trade a provisioning gap for a security regression > - This pull request generalizes the existing boot installer into a catalog-driven `ensureBundledPlugins`, fed by `plugins.autoInstall` from `PAPERCLIP_MANAGED_CONFIG` > - The benefit is that managed tenants get working bundled plugins out of the box, through the same in-process, role-free mechanism the codebase already trusts, while self-hosted boot is unchanged ## Linked Issues or Issue Description No public issue exists; `feature_request` template fields: - **Problem or motivation:** on managed instances tenant users cannot install plugins (by design they never hold instance admin), so even plugins shipped with the product are unusable; boot provisioning currently knows only the Kubernetes bundle. - **Proposed solution:** a bundled-plugin catalog plus `ensureBundledPlugins(keys)` driven by the managed config; same in-process `loader.installPlugin({ localPath })` under a system actor; unknown keys or catalog-escaping paths fail startup; already-present plugins are skipped so operator-disabled plugins are never silently re-enabled. - **Alternatives considered:** granting tenant users install privileges (widens secrets/adapters/settings access to solve a one-button problem); a separate non-admin install route for bundled plugins (new authz surface; provisioning removes the need for any install action at all). - **Roadmap alignment:** supports the in-progress "Cloud deployments" milestone and builds on the shipped sandbox-provider milestone in `ROADMAP.md`. Refs #10058. ## What Changed - New `server/src/services/bundled-plugins.ts`: the bundled-plugin catalog, the fail-to-start resolver (`resolveBundledPluginInstalls`, positive allowlist + catalog-root containment with symlinks resolved), and the fail-safe installer (`ensureBundledPlugins`). - `server/src/app.ts`: replaces the hard-coded `ensureBundledKubernetesPlugin` boot hook with resolver + installer wiring, with test hooks (`managedPluginAutoInstall`, `bundledPluginCatalogRoot` options). - `server/src/index.ts`: passes `plugins.autoInstall` from the single fail-closed `PAPERCLIP_MANAGED_CONFIG` startup parse (#10058) into `createApp`; absent env means self-hosted and changes nothing. ## Verification - 24 new tests in `server/src/__tests__/bundled-plugins.test.ts` (catalog resolution, containment incl. symlink and `..` escapes, skip/reinstall matrix, self-hosted invariants, installer error paths) — all green. - 85 adjacent startup/plugin-route/auto-build/managed-config tests green (`managed-config`, `instance-settings-managed-overlay`, `plugin-install-autobuild`, `plugin-routes-authz`, `server-startup-feedback-export`). - Server `tsc --noEmit` clean. ```bash cd server npx vitest run src/__tests__/bundled-plugins.test.ts npx vitest run src/__tests__/managed-config.test.ts src/__tests__/instance-settings-managed-overlay.test.ts src/__tests__/plugin-install-autobuild.test.ts src/__tests__/plugin-routes-authz.test.ts src/__tests__/server-startup-feedback-export.test.ts npx tsc --noEmit ``` ## Risks - Managed instances with a malformed or unknown `plugins.autoInstall` entry now **refuse to start** (fail closed, by design) instead of booting half-provisioned; harness misconfiguration surfaces as a precise startup error. - Self-hosted behavior is unchanged (kubernetes bundle only, `PAPERCLIP_KUBERNETES_PLUGIN_PATH` honored without containment, install failures log-and-continue), so the default deployment path carries low risk. - No uninstall path exists in this module; removal from the autoInstall list can leave a previously provisioned plugin installed (intentional v1 semantics, documented in code). ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.3 |
||
|
|
216d3d2680 |
Managed-instance config: fail-closed PAPERCLIP_MANAGED_CONFIG parsing and read-time settings overlay (#10058)
**Builds on.** #10055 — the `catalogVersion` this config document pins is the feature-catalog artifact #10055 emits. **Summary.** Instances operated by a managed hosting control plane can now receive instance configuration through a single environment variable, `PAPERCLIP_MANAGED_CONFIG` (versioned JSON: `mode`, `catalogVersion`, `features`, `plugins.autoInstall`). When the variable is absent the instance is self-hosted and nothing changes. When present, parsing is strict and **fail-closed**: blank value, malformed JSON, unknown feature key, a feature key this build's feature catalog does not mark tier `managed`, missing required section, or unsupported version refuses startup with a precise error — a typo that silently does nothing is how a security control quietly fails. Managed feature values are overlaid **at read time** inside the instance settings service (never persisted), so a DB restore or manual row edit cannot resurrect a disabled capability; responses expose per-key `managedKeys` metadata (`managed: true`, `managedBy`) so clients can render locked state. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs both self-hosted and under managed hosting, where an operator's control plane owns instance configuration > - Today instance feature settings live only in the tenant database; a hosting control plane has no way to enforce a configuration that tenant-side writes or restores cannot undo > - Managed configuration will carry security posture, so delivery must be atomic and parsing must fail closed — a typo that silently does nothing is how a security control quietly fails > - This pull request adds strict parsing of one `PAPERCLIP_MANAGED_CONFIG` env var and overlays its feature values at read time inside the settings service, never persisting them > - The benefit is a minimal, auditable managed-hosting contract: absent var ⇒ self-hosted instances are byte-for-byte unchanged; present ⇒ deterministic, locked configuration surfaced to clients via per-key managed metadata ## Linked Issues or Issue Description Refs #966 — this PR delivers that issue's "managed config injection" hook, via a strict env-var contract rather than the config-file path it sketches; the issue's other hooks (identity header, health, usage webhook, lifecycle, external secrets, IAM auth) are out of scope, so the PR refs rather than closes it. *Mechanism differs from #966's proposal, so the `feature_request` fields are also filled in:* - **Problem or motivation:** managed hosting deployments need to centrally enable/disable instance features; DB-stored settings can be edited, restored, or migrated back to permissive values, and nothing marks a value as operator-enforced. - **Proposed solution:** one versioned JSON env var; fail-closed parse at startup; read-time overlay in the settings service (precedence: managed value over stored value over schema default); `managedKeys` metadata in settings responses so clients can render locked state. - **Alternatives considered:** per-feature env vars (non-atomic across a half-updated env set, unbounded env surface); seeding the DB at boot (persisted values can be edited or restored over, and cannot express "forced"); lenient warn-and-drop parsing (fails open — unacceptable for a security-bearing control). - **Roadmap alignment:** supports the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed - New `server/src/services/managed-config.ts` (pure parser over the env record) - Startup parse ordered before the first `instanceSettingsService` construction in `server/src/index.ts` - Read-time merge + `managedKeys` in the settings service - Shared validator updates ## Verification - 29 parser/overlay tests (fail-closed matrix incl. blank/whitespace env, missing sections, catalog-tier mismatch, empty-section happy path): `pnpm vitest run src/__tests__/managed-config.test.ts src/__tests__/instance-settings-managed-overlay.test.ts` (from `server/`) - 40 existing settings route/service tests green: `pnpm vitest run src/__tests__/instance-settings-routes.test.ts src/__tests__/instance-settings-service.test.ts` (from `server/`) - 15 shared validator tests: `pnpm vitest run src/validators/instance.test.ts` (from `packages/shared/`) - Server `tsc --noEmit` clean: `pnpm typecheck` (from `server/`) ## Risks - Self-hosted instances (no `PAPERCLIP_MANAGED_CONFIG` set) are byte-for-byte unchanged — the parser only runs when the variable is present. - For managed instances, a malformed document now refuses startup by design (fail-closed). This is an intentional behavioral guarantee, not a regression: the control plane owns the variable and a precise startup error is the contract. - Overlay values are never persisted, so no migration or data-shape risk. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.2 |
||
|
|
9edde68373 |
feat(kubernetes): native file-sync lifecycle hooks over pod exec (#10053)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - AI agents run in sandboxed execution environments (Kubernetes pods, Daytona workspaces, etc.) and need to sync files between the host and those environments — for workspace setup, asset delivery, and output retrieval > - The existing sync path for Kubernetes uses a base64-over-exec chunk loop: each ~4 MB chunk requires its own `execInPod` round-trip, so large syncs balloon into many exec calls with corresponding overhead > - `execInPod` supports piped stdin/stdout, meaning the full transfer can be done as a single exec that streams a raw `tar` archive over the data channel — one round-trip regardless of file size, with nothing base64-encoded and nothing buffered whole in memory on either side > - PR-1 (#10013, merged) added the `onEnvironmentSyncIn`/`onEnvironmentSyncOut` opt-in hook API to the sandbox provider interface and documented the protocol; PR-2 (#10028, merged) implemented these hooks for the Daytona provider > - This pull request implements the same two lifecycle hooks in the Kubernetes sandbox provider, so workspace/asset file sync streams through one `execInPod` per operation instead of the chunk loop > - The benefit is significantly fewer exec round-trips for large syncs and flat memory use on both host and pod, with security properties preserved: atomic replace, secret-mode enforcement, path confinement, TOCTOU-safe snapshot, and member-confinement on host-assembled archives from sandbox-authored tar output ## Linked Issues or Issue Description This is the third and final PR in a sequential series: - Refs #10013 — PR-1: opt-in sync hook API + provider docs (merged) - Refs #10028 — PR-2: native file-sync lifecycle hooks for Daytona provider (merged) **Feature:** Native single-exec file-sync lifecycle hooks for the Kubernetes sandbox provider. *Motivation:* The existing Kubernetes sync path encodes files as base64 and loops over `execInPod` one chunk at a time (~4 MB per exec). For large workspaces or asset sets this is slow and resource-intensive. The Kubernetes `execInPod` API supports piped stdin/stdout, enabling a raw-`tar` streaming transfer that needs only one exec regardless of file count or size and never buffers the whole payload in memory. *Proposed solution:* Implement `onEnvironmentSyncIn` and `onEnvironmentSyncOut` in the Kubernetes provider using a streaming `execInPod` with a tar pipeline — for syncIn the host builds the archive on disk and streams its raw bytes into the pod's stdin (`head -c <exact-size> | tar -x`, no base64); for syncOut in-pod `tar` writes to the exec's stdout and the host streams those bytes straight to a file. Path confinement, atomic replace, secret-mode enforcement, TOCTOU protection, and a streamed-bytes fail-closed guard are all enforced. ## What Changed - **New `src/file-sync.ts`** in `packages/plugins/sandbox-providers/kubernetes/` — `performSyncIn` and `performSyncOut` over an injected pod-exec closure, keeping transfer logic hermetically unit-testable - **New `execInPodStreaming` in `src/pod-exec.ts`** — a streaming exec primitive that binds a caller-supplied stdin readable and a stdout writable to the exec WebSocket data channel, added alongside the existing `execInPod` (which is unchanged). This lets a transfer stream raw bytes to/from disk instead of buffering the payload as a single string - **Updated `src/plugin.ts`** — registers `onEnvironmentSyncIn`/`onEnvironmentSyncOut`; resolves the `sandbox-cr` pod exactly like `onEnvironmentExecute` and delegates; `job` backend rejects file-sync calls explicitly (out of scope) - **syncIn path:** host builds the tarball to a temp file → streams its raw bytes over exec stdin, bounded in-pod by `head -c <exact-archive-size> | tar -x` (no base64 anywhere) → extract into a `/proc/self/fd`-pinned reserved `0700` staging dir → `chmod`-before-`mv -f` atomic replace per file (directory mappings use `followSymlinks`→`-h`) - **syncOut path:** in-pod validate + realpath-snapshot each source (closes the validation→copy TOCTOU window) → single-exec `tar -c` streamed over exec stdout → host streams that stdout straight to a temp file through a byte-counting transform → member-confined extraction of the sandbox-authored archive - **Security properties:** secret files land at requested mode with no widened window; every interpolated path is shell-quoted and confined lexically plus via in-pod `realpath`; the outbound stream is bounded by a **streamed-bytes disk guard** (`MAX_SYNC_OUTPUT_BYTES`, 8 GiB default, per-call overridable) that fails the transfer closed — writing no target file — if an untrusted pod emits more bytes than allowed. Neither host nor pod buffers the whole payload, so there is no in-memory size cap on the transfer - **No changes** to `execInPod`, `wrapCommandWithEnv`, or `FastUploadInterceptor` (the `environmentExecute` path is untouched) - **No dependency or lockfile changes** - **New tests** in `test/unit/file-sync.test.ts` (atomic-replace, `0600` secret mode, symlink preserve/deref, dir-mapping, exclude, path-confinement rejection, streamed-output guard fail-closed) and `test/unit/pod-exec.test.ts` (streaming stdin/stdout, caller-sink error fail-closed), plus extended `test/unit/plugin.test.ts` ## Follow-up: Legacy Job-Lease Base64 Fallback Fix Addresses the Greptile 4/5 blocking finding ("Handle existing job leases", `server/src/services/environment-runtime.ts`). Job leases provisioned before the `nativeFileSyncUnsupported` metadata flag existed carry `backend: "job"` but no flag, so `supportsSync()` treated them as native-capable and routed their sync to the pod-exec hook — which the job backend rejects (it has no exec channel) instead of using the byte-identical base64 fallback. The fix adds a belt-and-suspenders gate on the persisted `backend === "job"` field alongside the existing `nativeFileSyncUnsupported` flag check, so pre-existing job leases continue syncing via the base64 fallback after deployment. No behaviour change for `sandbox-cr` leases. ## Verification - `pnpm --filter @paperclipai/sandbox-provider-kubernetes test` — 19 files / 182 tests green, including the existing `upload-interceptor` and `pod-exec` suites - `tsc --noEmit` in the kubernetes package — 0 errors - The sync hooks are opt-in; existing `environmentExecute` behaviour is unaffected and tested by the unchanged existing suites ## Risks - **Opt-in only:** `onEnvironmentSyncIn`/`onEnvironmentSyncOut` are registered conditionally; providers that do not register them fall back to the existing chunk loop. No regression risk on the existing path. - **Shell-injection surface:** all path interpolation uses shell-quoting; paths are additionally confined lexically and via in-pod `realpath` before use. - **TOCTOU on syncOut:** the in-pod snapshot validates and records file metadata before the tar call, closing the window between validation and copy. - **Archive member confinement:** host-side reassembly rejects any tar member whose resolved path escapes the target directory, preventing a malicious in-pod tar from writing outside the intended destination. - **Untrusted-output volume:** an over-large outbound stream trips the streamed-bytes disk guard and fails closed (no target written and the temp sink is swept) rather than filling host disk or memory; the guard bounds disk unconditionally and bounds memory insofar as WebSocket write-backpressure holds. ## Model Used Anthropic Claude Sonnet 4.6 (`claude-sonnet-4-6`) — produced by a Claude-based AI agent using agentic tool use and multi-step code generation. 200K context window, extended reasoning, code execution and verification capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.1 |
||
|
|
ad74fb5450 |
Add a feature catalog build artifact derived from the experimental settings schema (#10055)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instances expose ~23 experimental feature settings, all declared in one shared zod schema and toggled per instance > - Deployment tooling and hosting control planes have no machine-readable list of those feature keys for a given release — the schema is only reachable from code that imports the package > - Any external system that references feature keys therefore does so as free text, and typos drift silently > - This pull request derives a versioned `feature-catalog.json` build artifact from the schema, with a compiler-checked metadata map so the schema stays the single source of truth > - The benefit is a stable contract external tooling can validate feature-key references against, with zero runtime behavior change ## Linked Issues or Issue Description No public issue exists; `feature_request` template fields: **Problem or motivation:** External deployment tooling cannot enumerate or validate an instance's feature keys per release; free-text references fail silently when keys are renamed or removed. **Proposed solution:** A metadata map keyed by the settings schema's own keys (compiler flags drift) plus a build step emitting `feature-catalog.json` (keys, tiers, defaults, `catalogVersion`) as a release artifact. **Alternatives considered:** A hand-maintained catalog file (drifts from the schema); serving the schema from a runtime API (requires a running instance at validation time — a build artifact works offline and pins to a release). **Roadmap alignment:** Supports the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed Adds a metadata map (title, description, tier, cloud/self-hosted defaults) keyed by the keys of `instanceExperimentalSettingsSchema`, so the schema stays the single source of truth and the compiler flags any drift. A new build step (`build:feature-catalog --version <v>`) emits `feature-catalog.json` — all 23 feature keys, their tiers, and a `catalogVersion` — as a release artifact that managed-hosting control planes can validate feature-flag writes against. No runtime behavior changes. - New `packages/shared/src/feature-catalog.ts`: per-flag metadata map keyed by a type derived from the settings schema (adding/removing/renaming a flag without updating the map is a compile error), plus `featureCatalogArtifactSchema` and `buildFeatureCatalogArtifact`/`renderFeatureCatalogArtifact` for the artifact - New `scripts/generate-feature-catalog.ts` wired as `pnpm build:feature-catalog --version <v>` - `scripts/create-github-release.sh` generates the artifact and uploads it as a GitHub Release asset (with a dry-run preview line) - Tests in `packages/shared/src/feature-catalog.test.ts` ## Verification - `vitest run packages/shared/src/feature-catalog.test.ts` — 9 tests: schema-key coverage, drift detection, artifact shape - `pnpm --filter @paperclipai/shared typecheck` - Artifact generation run end-to-end: `pnpm build:feature-catalog --version 0.0.0-test` emits 23 keys with `catalogVersion` ## Risks Low risk — no runtime behavior changes; the change is metadata, a build script, and a release-artifact emission step only. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.723.0-canary.0 |
||
|
|
e55d702916 |
feat(plugin-sdk): human-attributed issue comments for chat gateway plugins (#10050)
Adds the `issue.comments.create_human_attributed` capability and `ctx.issues.createComment` `actorUserId` option, with host-side active-human-member verification. LOOA-627. Co-Authored-By: Paperclip <noreply@paperclip.ing>v2026.722.0 canary/v2026.722.1-canary.0 |
||
|
|
1b8738da5c |
feat(cli): add --force to company export for non-interactive runs (#10054)
## Thinking Path
- `company export` writes into a target directory but aborts
non-interactively when that directory is non-empty ("already contains
files. Re-run interactively or choose an empty directory"), and there is
no override flag. Any automated caller that exports into a pre-existing
directory (for example a git clone used as a backup target) is therefore
stuck.
- The interactive confirmation is the right default for humans, but
automation needs an explicit opt-out rather than being forced to export
into a throwaway empty directory and copy the tree afterward.
- A `--force` flag that skips only the non-empty-directory confirmation
is the minimal change: it does not alter what gets written (existing
files are overwritten in place, nothing is bulk-deleted), so callers
keep full control over cleanup via their own VCS.
## What Changed
- Added a `--force` option to `company export` that skips the non-empty
output-directory confirmation for non-interactive/automated runs.
- Threaded the flag into `confirmOverwriteExportDirectory(outDir, {
force })`; behavior is unchanged when the flag is absent.
- Updated the non-interactive error message to also mention `--force`.
- Added focused unit tests covering: missing dir (resolves), empty dir
(resolves), non-empty dir without force (throws), non-empty dir with
force (resolves), and a path that exists but is a file (throws).
## Verification
- `pnpm --filter @paperclipai/cli exec vitest run
src/__tests__/company-export-force.test.ts` → 5/5 pass.
- `tsc --noEmit` over the CLI sources: no new type errors (the only
errors are pre-existing `@paperclipai/plugin-sdk` module-not-found in
`server/` from an unbuilt plugin sdk in the sandbox, unrelated to this
change).
- End-to-end against a local server: exporting into a non-empty
directory fails without `--force` and succeeds with it; a `.git`
directory and a sentinel file in the target were preserved; 128 files
written.
## Risks
- Low. The flag is opt-in and defaults to false; interactive and
empty-directory behavior is untouched. `--force` overwrites matching
files in place but never deletes unrelated files, so it cannot silently
wipe a directory.
## Model Used
Claude Opus 4.8 (claude-opus-4-8)
---
**Problem or motivation**
`company export` aborts when the `--out` directory is non-empty and
stdin/stdout are not a TTY, and there is no override flag. This makes it
impossible to run `company export` unattended into a pre-existing
directory such as a git clone.
**Proposed solution**
Add a `--force` flag that skips the non-empty-directory confirmation for
non-interactive callers. Files are still written on top of existing
content with no bulk delete, so unrelated files such as `.git` are
preserved.
**Alternatives considered**
Exporting into a fresh temp directory and copying the tree into the real
target afterward works but is clumsy and error-prone for automation;
broadening or removing the guard entirely would remove a useful safety
net for interactive users.
**Roadmap alignment**
Hardens the automated/unattended export path used by scheduled
company-backup routines.
---
- [x] I searched the repository and open pull requests for similar or
duplicate PRs and found none.
Co-authored-by: anicca <annica@MichaelacStudio.localdomain>
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.722.0-canary.4
|
||
|
|
b431d4eca1 |
fix(a11y): add ARIA progressbar to QuotaBar component (#1878)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Operators run those agents against paid model providers, so the web UI has a Costs surface that reports spend and quota utilisation > - Those figures are drawn as horizontal bars by the shared `QuotaBar` component (`ui/src/components/QuotaBar.tsx`), consumed by `BillerSpendCard` and `ProviderQuotaCard` > - `QuotaBar` renders its fill as a plain `<div>` whose CSS width is the only encoding of the percentage — no `role`, no value attributes, no accessible name > - A screen reader therefore announces nothing at all for these bars, so the spend and quota numbers they convey are unavailable to assistive-technology users (WCAG 2.1 SC 4.1.2, Name/Role/Value) > - The same gap is being closed for the other bars in this area by #1805 (BudgetPolicyCard) and #1869 (ProviderQuotaCard's inline bars); `QuotaBar` is the remaining shared component with no ARIA semantics > - This pull request adds the standard ARIA progressbar attributes to `QuotaBar`'s fill element, reusing the `label` prop the component already takes > - The benefit is that every progress bar on the Costs screen exposes its name and current value to assistive technology, with no visual or behavioural change for sighted users ## Linked Issues or Issue Description No existing GitHub issue — the problem is described in-PR below, following [`bug_report.yml`](.github/ISSUE_TEMPLATE/bug_report.yml). Related PRs from the same accessibility sweep (each covers a *different* component, so these are companions rather than duplicates — all three are currently open): - Refs #1805 — ARIA attributes for the BudgetPolicyCard progress bar - Refs #1869 — ARIA attributes for the ProviderQuotaCard inline bars **What happened?** On the Costs screen, the spend/quota bars rendered by `QuotaBar` (via `BillerSpendCard` and `ProviderQuotaCard`) are non-semantic `<div>` elements. Screen readers skip them entirely: no role, no value, no label is announced, so the percentage information is available only visually. **Expected behavior** Each bar should be exposed as a progress bar with an accessible name and its current value — e.g. announced as "Weekly spend: 45%, progress bar". **Steps to reproduce** 1. Run the app and open the Costs page. 2. Expand any provider or biller card so a quota/spend bar is visible. 3. Navigate to the bar with a screen reader (VoiceOver, NVDA, or Chrome DevTools → Accessibility pane). 4. Observe that the fill element has no role, no value, and no accessible name. **Paperclip version or commit** Reproduces on `master`; `ui/src/components/QuotaBar.tsx` has carried no ARIA attributes since the component was introduced. **Deployment mode** Not deployment-specific — the missing markup is in the shipped component. Verified in local dev (`pnpm dev`). **Agent adapter(s) involved** Not adapter-specific (core UI). ## What Changed - `ui/src/components/QuotaBar.tsx`: added `role="progressbar"` to the fill `<div>`. - Added `aria-valuenow={Math.round(clampedPct)}` with `aria-valuemin={0}` / `aria-valuemax={100}`, using the already-clamped percentage so the reported value can never fall outside 0–100. - Added an `aria-label` of the form `<label>: <pct>%`, reusing the existing `label` prop for the accessible name. - No changes to props, styling, layout, or rendering logic: 1 file, 5 added lines, 0 deleted. ## Verification - Manual: open Costs → expand a provider/biller card, inspect the bar in Chrome DevTools → Accessibility pane. The fill node now reports role `progressbar`, value `45`, min `0`, max `100`, and name "Weekly spend: 45%". - Manual: with VoiceOver/NVDA, the bar announces "Weekly spend: 45%, progress bar" instead of being skipped. - Visual regression check: the bar is unchanged for sighted users — only ARIA attributes were added, no class or style changes. - CI (lint, typecheck, build, tests) is green on this branch. - No unit test is added: the change is a set of static ARIA attributes on one element, and `QuotaBar` currently has no test file. Happy to add one if maintainers would like coverage here. ## Risks Low risk. Presentation-only accessibility metadata on a single element; no props, state, or styling change, and no other component is touched. The one debatable point is that the percentage appears in both `aria-label` and `aria-valuenow`, so some screen readers may announce it twice; both forms are valid, and the label is kept because it carries the bar's name alongside the value. Happy to drop the percentage from the label if reviewers prefer the terser announcement. ## Model Used <!-- @bluzername: please replace this line with the provider + exact model ID (and context window / reasoning mode if relevant), or "None — human-authored". Required by CONTRIBUTING.md. --> ## Checklist - [x] I have included a thinking path that traces from project context to this change - [ ] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (no docs cover this component's markup) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.722.0-canary.3 |
||
|
|
204c416478 |
fix(release): publish bundled packages with trusted npm staging (#10047)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip publishes its CLI, server, adapters, and shared packages through automated canary and stable release workflows. > - `@paperclipai/adapter-utils` bundles the patched `acpx` runtime, so it must use npm 11 for OIDC trusted publishing. > - The prior staging directory contained pnpm's `.pnpm` symlink forest, which crashes npm 11's directory-pack step on GitHub runners and produces a consumer-broken bundled dependency tree. > - This pull request rebuilds staged production dependencies as a physical npm tree, reapplies repository patches, and publishes that clean directory directly with npm 11 trusted publishing. > - The benefit is a release path that retains GitHub Actions OIDC trusted publishing while shipping a working patched acpx runtime to consumers. ## Linked Issues or Issue Description Refs: #9980, #10030, #10041 No public GitHub issue exists for this release failure. ### What happened? Canary and stable publishing began routing `@paperclipai/adapter-utils` through npm after it declared `bundleDependencies: ["acpx"]`. Publishing the pnpm-deployed directory with npm 11 crashes during npm's directory-pack phase on GitHub runners with `Exit handler never called!`. The same staged shape also produces a broken consumer artifact because acpx cannot resolve transitive runtime dependencies after installation. ### Expected behavior Bundled packages publish directly from a self-contained staging directory through npm 11 OIDC trusted publishing, and consumers receive a working patched acpx runtime with its transitive dependencies. ### Steps to reproduce 1. Stage `packages/adapter-utils` using the old `pnpm deploy`-only shape. 2. Publish that directory with npm 11 on a GitHub runner. 3. npm crashes before registry/OIDC activity while walking the `.pnpm` symlink forest. 4. Install an artifact packed from that old shape into a fresh npm project and run acpx; its runtime dependency resolution fails. ### Deployment mode GitHub Actions canary/stable release workflow. ### Relevant logs or output ```text npm error Exit handler never called! ``` ## What Changed - After `pnpm deploy`, remove the staged pnpm `node_modules` tree and run `npm install --omit=dev --ignore-scripts --no-audit --no-fund` to create a physical hoisted production tree. - Apply every root `pnpm.patchedDependencies` patch whose package is declared in the staged package's bundled dependencies, failing staging if any patch cannot apply. - Assert the staged acpx runtime contains the required `onAgentStderr` patch marker. - Publish the clean staging directory directly with pinned npm 11.18.0, retaining GitHub Actions OIDC trusted publishing, verbose diagnostics, and the duplicate-transparency-log retry without provenance. - Keep pinned npm 10.9.7 packing only for local/dry-run payload verification; registry publishing does not use a tarball argument. - Add focused coverage for npm-tree staging, patch application, direct directory publish arguments, and bundled tlog retries. ## Verification - `bash -n scripts/release-lib.sh scripts/release.sh` - `node --test scripts/release-lib.test.mjs scripts/acpx-patch-packaging.test.mjs` — 12/12 passed. - `pnpm test:release-registry` — 68/68 passed. - Real staging smoke: `node scripts/prepare-bundled-package.mjs packages/adapter-utils <stage>` produced a real `node_modules/acpx` directory, no `.pnpm` directory, and the `onAgentStderr` patch marker. - Real npm 11 directory-publish smoke: `npx --yes npm@11.18.0 publish --dry-run --tag canary --access public --loglevel verbose` packed 26 bundled dependencies and reached the expected existing-version registry rejection without `Exit handler never called!`. - The merge-triggered `publish_canary` workflow remains the live OIDC trusted-publishing verification. ## Risks - The live GitHub Actions trusted-publishing path can only be fully proven by the merge-triggered canary run; npm debug-log upload remains available if it fails. - Bundling acpx continues to freeze platform-specific transitive artifacts such as esbuild binaries from the Linux release runner. This is a pre-existing consequence of the bundling decision in #9980 and is not expanded here. - Rebuilding dependencies with npm depends on the exact bundled dependency versions in the staged manifest; staging fails hard if repository patches no longer apply. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.6-sol`, high reasoning mode, with repository tool use and code execution. The harness did not expose a model context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.722.0-canary.2 |
||
|
|
8f7509c28b |
fix(a11y): add aria-label to mobile tab selector in PageTabBar (#1871)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - People increasingly drive Paperclip from a phone, so the UI ships a mobile layout alongside the desktop one > - `PageTabBar` is the shared component behind the tab strip on nearly every detail page — AgentDetail, ProjectDetail, RoutineDetail, IssueDetail, Inbox, Costs > - On desktop it renders a Radix `TabsList`, whose `TabsTrigger`s carry their own accessible names; on mobile it swaps to a native `<select>` > - That `<select>` had no accessible name at all, so screen readers announced it only as "popup button" — a user could not tell what the control switches between > - Because the component is shared, the gap reproduced on every mobile page that uses tabs rather than on one screen > - This pull request adds `aria-label="Page section"` to the mobile `<select>` > - The benefit is that mobile screen-reader users get the same orientation desktop users already get from the tab triggers, from a one-line change with no visual or behavioral impact ## Linked Issues or Issue Description No existing public issue covers this, so the problem is described in-PR following `.github/ISSUE_TEMPLATE/bug_report.yml`: - **What happened** — On a mobile viewport, the `PageTabBar` `<select>` had no `aria-label`, no `<label>` association, and no visible text of its own. VoiceOver/TalkBack announce it as an unlabeled "popup button". - **Expected behavior** — The control announces what it switches between, matching the accessible naming the desktop `TabsTrigger`s already provide. - **Steps to reproduce** — 1. Open any detail page with tabs (agent, project, routine, issue). 2. Narrow the viewport to mobile width so the tab strip collapses to a `<select>`. 3. Focus the `<select>` with a screen reader. 4. Observe that no purpose is announced. - **Version / commit** — head `ef92d1c`, branch `fix/page-tab-bar-mobile-a11y`. - **Deployment mode** — Any. The change is UI-only and client-side. **Related prior PR:** #1532 (closed unmerged on 2026-03-23) made this same one-line change to `ui/src/components/PageTabBar.tsx` as part of a ~100-file batch. This PR is the focused standalone version of that fix. ## What Changed - Added `aria-label="Page section"` to the mobile `<select>` in `ui/src/components/PageTabBar.tsx`. One line added; no other files touched. ## Verification - **Automated:** `pnpm -C ui test` and the repo CI gates (lint, typecheck, build) — CI is currently green on `ef92d1c`. - **Manual:** Open any tabbed detail page, narrow the viewport until the tab strip becomes a `<select>`, and focus it with VoiceOver (macOS/iOS) or TalkBack (Android). It now announces "Page section, popup button" instead of an unlabeled "popup button". - **Inspector check:** In devtools, the `<select>` node's computed accessible name is "Page section" (previously empty). ## Risks Low risk. `aria-label` on a `<select>` is a presentation-free attribute: it changes nothing about layout, styling, DOM structure, event handling, or the desktop code path, which is untouched. No migration, no API change, no new dependency. The only debatable point is wording — "Page section" is a generic name shared by every call-site (see the note below). ## Model Used Not specified by the original author, and not recoverable from the commit metadata (no `Co-Authored-By` or model trailer on `ef92d1c`). @bluzername — please replace this line with the provider, model ID/version, and any relevant capability details, or "None — human-authored". ## Checklist - [x] I have included a thinking path that traces from project context to this change - [ ] I have specified the model used (with version and capability details) — see above; needs the author - [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`fix/page-tab-bar-mobile-a11y`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable — no test added; the change is a static attribute with no branching behavior - [ ] I have updated relevant documentation to reflect my changes — not applicable - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — currently 4/5, see below - [ ] I will address all Greptile and reviewer comments before requesting merge --- ### Maintainer note on the open Greptile comment This description was restructured to the repository PR template by a maintainer; the code and the author's intent are unchanged. Unchecked boxes above are ones only @bluzername can attest to. Greptile's one remaining comment asks for a `selectAriaLabel` prop so call-sites could override the label. We think the hardcoded label is correct here and match existing practice: shared components whose meaning is fixed own their label internally (`ThemeToggle.tsx`), while components whose label depends on the data they render take it as a prop (`CopyText.tsx`'s `ariaLabel`). This `<select>` always means "which page section", at every call-site, so a prop no caller would set would be unused API surface. |
||
|
|
f215444919 | feat(daytona): native file-sync lifecycle hooks over batch transfer (#10028) | ||
|
|
39b94d0e7a | fix(test): gate interaction-continuation retry barrier on terminal write (#10049) | ||
|
|
4e8cd757eb |
fix(release): add npm publish crash diagnostics (#10041)
## Thinking Path > - Paperclip publishes a coordinated set of packages through its release workflows > - Bundled packages use a pinned npm CLI so trusted publishing works consistently > - The canary publisher now crashes deterministically inside npm before useful output reaches the workflow log > - The npm debug log that contains the underlying failure disappears with the hosted runner > - This pull request upgrades the pinned publish CLI and preserves both verbose HTTP activity and npm debug logs on failure > - The benefit is that the plausible HTTP-layer fix ships immediately, while any remaining CI-only failure becomes diagnosable ## Linked Issues or Issue Description - **Problem:** The canary release workflow fails on the first bundled package with `npm error Exit handler never called!` and no preceding diagnostic output. - **Expected behavior:** Bundled packages publish through trusted publishing, or the workflow retains enough npm diagnostics to identify the actual failure. - **Reproduction:** Run the canary release workflow in GitHub Actions; the failure reproduced on both attempts of run 29948506814. - **Version/commit:** Current `master` after #10024 and #10030. - **Deployment mode:** GitHub-hosted release workflow using Node.js 24 and npm trusted publishing. - Related: #10024, #10030. ## What Changed - Bumped the bundled publish CLI from npm 11.16.0 to npm 11.18.0. - Added `--loglevel verbose` to bundled npm publish invocations. - Dumped the last 300 lines of every npm debug log after failed canary or stable publishes, with common registry credential forms redacted. - Updated release assertions to pin npm 11.18.0 and verify verbose logging. ## Verification - `pnpm test:release-registry` — 66 tests passed. - `bash -n scripts/release-lib.sh`. - Parsed `.github/workflows/release.yml` with Python/PyYAML. - Smoke-tested npm log redaction with representative Authorization, `_authToken`, and token environment values. - Smoke-tested npm debug-log redaction against Authorization, `_authToken`, and `npm_token` examples. - `git diff --check origin/master...HEAD`. - The merge-triggered canary workflow remains the live trusted-publishing verification. ## Risks - Low code risk: changes are isolated to the release publisher and its workflow diagnostics. - npm 11.18.0 could expose a different registry/runtime regression; failure-time debug log dumping makes that actionable. - Verbose npm output increases release log volume but does not change package contents or dist-tags; common credential forms are redacted before debug logs are printed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.3 Codex, tool-enabled coding agent with repository and shell execution; context window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8af70b9fae |
fix(test): drain in-flight heartbeat runs before liveness teardown (#10040)
## Thinking Path > - Paperclip runs AI agent heartbeats to manage work; each heartbeat dispatches `executeRun` fire-and-forget, which is intentional for concurrency > - The server escalation test suite (`heartbeat-issue-liveness-escalation.test.ts`) exercises `reconcileIssueGraphLiveness`, which heals a resolved-dependency wake by enqueuing an on-demand heartbeat run > - `enqueueWakeup` → `startNextQueuedRunForAgent` dispatches the run fire-and-forget (`void executeRun(...)`), so the background run outlives the awaited reconcile call > - The test's `afterEach` polled `heartbeat_runs.status` to wait for idle, but that flips to `completed` while `executeRun`'s finally block is still flushing events — the escaping `heartbeat_run_events` insert could land between the events delete and the runs delete, tripping the FK constraint > - This PR fixes the race deterministically by tracking in-flight `executeRun` promises and exposing `heartbeatService.drainActiveRunExecutions()`, which the suite awaits before clearing tables > - The benefit is a permanently reliable escalation test suite with no sleeps, no retry bumps, and no production behavior change ## Linked Issues or Issue Description **What happened?** The `heartbeat-issue-liveness-escalation.test.ts` suite intermittently failed in CI with: ``` delete on table "heartbeat_runs" violates foreign key constraint "heartbeat_run_events_run_id_heartbeat_runs_id_fk" ``` **Expected behavior** `afterEach` cleanup should complete without FK violations. **Steps to reproduce** The race is timing-dependent but surfaces reliably when the teardown window is artificially widened. `reconcileIssueGraphLiveness()` heals resolved-dependency wakes by dispatching a heartbeat run fire-and-forget (`void executeRun(...)`). The old `afterEach` polled `heartbeat_runs.status` — but that flips to `completed` while `executeRun`'s finally block still has pending `heartbeat_run_events` row writes. The escaping insert can land between the events delete and the runs delete. **Paperclip version or commit** Reproducible on current `master` (commit `b57aa9950c707a024156c34b79326a82b2dcca31`) ## What Changed - **`server/src/services/heartbeat.ts`** — tracks all in-flight `executeRun` promises in a module-level `Set`; exposes `heartbeatService(db).drainActiveRunExecutions()`, which loops until the set drains (a completing run can enqueue the next queued run in its finally, so a single `await` is not enough) - **`server/src/server-suites/heartbeat-issue-liveness-escalation.test.ts`** — replaces the poll-on-`heartbeat_runs.status` teardown with `await heartbeatService(db).drainActiveRunExecutions()` before clearing tables; removes the now-unnecessary `waitForHeartbeatRunToComplete` helper ## Verification ```bash # Full file (22 tests) npx vitest run server/src/server-suites/heartbeat-issue-liveness-escalation.test.ts # 12x stress loop (264 test-runs, 0 failures) for i in $(seq 1 12); do npx vitest run server/src/server-suites/heartbeat-issue-liveness-escalation.test.ts || break done # Type check the changed files npx tsc --noEmit ``` - 22/22 tests green locally - 12/12 full-file loop iterations: 264 test-runs / 264 afterEach cycles, 0 failures - Widened-teardown stress variant (failed deterministically before the fix) now passes with the drain ## Risks Low risk. The drain mechanism is additive — it only affects test teardown and could also be wired into graceful shutdown. The fire-and-forget dispatch in production is unchanged. The `Set`-based tracking adds negligible overhead per run dispatch (insert on dispatch, delete on completion). ## Model Used - **Provider:** Anthropic - **Model:** Claude Sonnet 4.6 (`claude-sonnet-4-6`) - **Context window:** 200K tokens - **Mode:** Tool use, code execution, extended reasoning ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b57aa9950c |
fix(test): stop flaky server-suite afterAll hook timeouts (#10024)
## Thinking Path > - Paperclip is an open-source AI agent management platform; its test suite spans a `server` package that mounts real embedded Postgres databases in `beforeAll`/`afterAll` hooks > - The `server` package CI shard runs all ~93 suites serially (`maxWorkers=1`) on a loaded CI host; each suite boots and tears down its own embedded Postgres in hook callbacks > - vitest's default `hookTimeout` is 10 seconds; under load, graceful embedded-Postgres shutdown occasionally crosses that threshold > - This produces intermittent `Error: Hook timed out in 10000ms` failures in `afterAll` hooks — not test assertion failures — and the suites pass on re-run, making them textbook flaky tests > - Inspecting `embedded-postgres@18.1.0-beta.16` shows that `stop()` takes no argument (no fast-shutdown mode), SIGINTs postgres (already PostgreSQL "fast shutdown"), and resolves only on the child's `exit` event with no internal time bound > - Two targeted fixes: (1) raise `hookTimeout` and `teardownTimeout` to 30 s in `server/vitest.config.ts` — one config change that eliminates the flake for all ~93 suites at once; (2) wrap `stop()` in a 5 s bounded `Promise.race` in the test helper so a slow shutdown can never hang the hook regardless of OS scheduling variance > - This PR changes only test-infra and test-config; no production-code behavior changes ## Linked Issues or Issue Description No public GitHub issue exists for this flake. Inline bug description (bug report template): **What happened?** The `General tests (server (N/3))` CI shards intermittently fail with `Error: Hook timed out in 10000ms` in `afterAll` hooks and pass on re-run. Every test assertion passes; only the teardown hook exceeds vitest's default timeout. **Expected behavior** CI passes reliably. Teardown timeouts should not be a source of flake. **Steps to reproduce** Run the server test suite repeatedly on a loaded host or in CI with `maxWorkers=1` — the shard occasionally crosses 10 s in `afterAll` during embedded-Postgres shutdown. **Paperclip version** `master`, any build that includes `server/vitest.config.ts` without an explicit `hookTimeout`. **Deployment mode** Self-hosted (CI). ## What Changed - **`server/vitest.config.ts`** — added `hookTimeout: 30000` and `teardownTimeout: 30000`. Removes flake across all ~93 server suites at once. 30 s gives generous headroom over observed worst-case teardown while still catching a genuinely hung hook. - **`packages/db/src/test-embedded-postgres.ts`** — added `stopEmbeddedPostgresBounded()`, a 5 s `Promise.race` wrapper around `stop()`. Applied at all three call sites inside `cleanup()`. Data dir is still removed unconditionally; errors are still swallowed; the null-instance guard is preserved. Existing behavior unchanged except the shutdown can no longer block indefinitely. ## Verification - `tsc --noEmit` clean on `packages/db` (built against worktree-local `shared`) - `packages/db` `client.test.ts` passes 14/14 — boots embedded Postgres and exercises the bounded teardown via `cleanup()` in `afterEach` - Standalone bounded-race semantics verified: hang resolves at the 5 s bound; late or immediate `stop()` rejection swallowed; no unhandled rejection; null-instance path safe - CI: all 3 server shards + split-verify lane (Async-Verification Gate) expected green after this PR ```bash # Reproduce the teardown test locally: cd packages/db && npx vitest run src/client.test.ts # Type-check packages/db: npx tsc --noEmit -p packages/db/tsconfig.json ``` ## Risks Low risk. No product-code changes — test-infra and test-config only. The vitest timeout increase is additive (raises the ceiling; never lowers it). The bounded race wrapper preserves prior teardown behavior exactly: data dir always removed, errors always swallowed, stop is still attempted. A worst-case outcome is that a genuinely hung `stop()` now surfaces as a test timeout at 30 s instead of 10 s — still caught, just later. ## Model Used - **Provider:** Anthropic - **Model:** Claude Sonnet 4.6 (`claude-sonnet-4-6`) - **Context window:** 200 k tokens - **Capabilities:** tool use, code execution, extended reasoning - **Mode:** Paperclip agent heartbeat (autonomous execution with human board oversight) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
47c38777d8 |
fix(release): use trusted publishing npm for bundled packages (#10030)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip publishes canary and stable packages through a shared release script > - GitHub Actions authenticates those publishes through npm trusted publishing and an OIDC identity token > - Bundled-dependency packages recently moved from pnpm publish to a pinned npm CLI to preserve their bundled files > - That pin selected npm 10, which cannot use trusted publishing, so the first bundled package failed with `ENEEDAUTH` > - This pull request keeps bundled-package packing on npm 10 while routing actual publishing through the trusted-publishing-capable npm 11 version > - The benefit is bundled packages keep their required npm packaging behavior while canary and stable releases authenticate successfully ## Linked Issues or Issue Description **What happened** The canary release job failed while publishing `@paperclipai/adapter-utils` with `ENEEDAUTH`. The package has bundled dependencies, so the release helper selected pinned `npm@10.9.7`; the workflow provides OIDC trusted publishing rather than an npm token, and npm 10 cannot use that authentication path. Because this is the first package attempted, the release exited before trying the remaining packages. **Expected behavior** Bundled-dependency packages should publish with an npm CLI that both preserves bundled dependencies and supports GitHub Actions trusted publishing. **Steps to reproduce** Run the canary release workflow from master after PR #9980. The `publish_canary` job reaches `@paperclipai/adapter-utils`, invokes `npx npm@10.9.7 publish`, and fails with `ENEEDAUTH`. **Deployment mode** GitHub Actions canary and stable npm release workflows. Refs #9980. ## What Changed - Kept bundled-package dry-run packing on npm `10.9.7`, which successfully produces the staged tarball. - Routed bundled-package publishing through npm `11.16.0`, which supports GitHub Actions trusted publishing. - Split the pack and publish helpers so future npm changes cannot silently couple the two compatibility requirements. - Updated focused release and ACPX packaging tests to enforce both versions and call paths. ## Verification - `pnpm test:release-registry` — 67 passed locally. - Initial all-npm-11 PR head: Canary Dry Run reproduced an npm-internal crash during bundled `pack`. - Current head `5b3961ed13dd26ed2d6b1096ea23fd91b32e4353`: Canary Dry Run passed with split npm pack/publish helpers. - All PR checks passed, including build, typecheck + release registry, server/workspace suites, both e2e shards, security gates, and Greptile. - Greptile reviewed the current head at 5/5 confidence with no blocking issues. ## Risks - Low risk: the change only separates the npm CLI used for bundled-package packing from the CLI used for publishing. - The versions remain explicitly pinned because npm 11.16.0 currently crashes on the bundled pack payload, while npm 10.9.7 cannot perform trusted publishing. - Focused tests assert both pins and both helper call paths, and the full Canary Dry Run passes on the current head. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.6-sol`, high reasoning mode, with repository, shell, GitHub CLI, and code-execution tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e1e35881af |
feat(cost-events): propagate issue.billing_code at heartbeat record time (#6821)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runs report progress through the heartbeat service, which writes the cost ledger (`cost_events`) as usage accrues > - `cost_events` already has a `billing_code` column, but nothing populates it — the heartbeat writes `issueId`/`projectId` and leaves `billing_code` NULL > - Issues carry a `billing_code`, so the attribution data sits one join away but never reaches the ledger rows > - Reporting therefore has to reconstruct attribution by joining back to `issues` at query time, which reflects the issue's *current* billing code rather than the one in effect when the cost was incurred > - This pull request threads `billingCode` through `resolveLedgerScopeForRun` so the heartbeat stamps it onto each `cost_events` row at record time > - The benefit is that attribution is captured at write time and stays correct if an issue's billing code later changes ## Linked Issues or Issue Description No existing public GitHub issue. Describing the problem in-PR: **Problem.** `cost_events` has a `billing_code` column that is never written. The heartbeat's cost-ledger insert records `issueId` and `projectId` but not the billing code of the issue the run belongs to, so every row lands with `billing_code` NULL. **Impact.** Cost-per-billing-code reporting has to derive attribution by joining `cost_events` back to `issues` at query time. That join returns the issue's billing code *as of the query*, not as of when the cost was incurred, so historical cost reports shift retroactively whenever an issue is re-coded. **Desired behaviour.** The billing code in effect at record time is stored on the `cost_events` row itself. **Related PRs.** #6820 — same change to the same file by the same author, opened separately. These are duplicates; only one should land. ## What Changed - `resolveLedgerScopeForRun` now selects `issues.billingCode` alongside `id` and `projectId`. - The scope object it returns gained a `billingCode` field, populated with `issue?.billingCode ?? null`. - The early-return path for runs with no issue in context returns `billingCode: null`. - The `costs.createEvent` call in `heartbeatService` passes `billingCode: ledgerScope.billingCode` alongside `issueId`/`projectId`. No schema migration: `cost_events.billing_code` already exists. ## Verification **No automated test accompanies this change.** There is currently no test asserting that a `cost_events` row carries the issue's billing code when an issue is in scope, or `null` when there is not. A reviewer should treat the checks below as manual verification only. Manual verification against a running instance: ```sql -- Non-NULL billing_code for recent runs on billed issues SELECT billing_code, COUNT(*) FROM cost_events WHERE created_at > NOW() - INTERVAL '1 hour' GROUP BY billing_code; -- Cost attribution query this change is intended to enable SELECT billing_code, SUM(cost_cents) FROM cost_events GROUP BY billing_code; ``` Expected: rows for runs attached to an issue with a billing code now carry that code; runs with no issue in context remain NULL. ## Risks Low risk in blast radius, with two things worth a reviewer's attention: - **Behavioural shift for consumers.** `cost_events.billing_code` was uniformly NULL and now starts arriving populated. Anything downstream that groups, filters, or dedupes on that column will see new values and new cardinality. Existing rows are not backfilled, so the column is mixed NULL/non-NULL across the historical boundary. - **No test coverage.** The null-fallback behaviour on both paths is asserted only by reading the code, not by a test. - **Migration safety:** not applicable — no schema change; the column already exists. - **Failure mode:** if `billingCode` were absent from the `issues` selection the value would silently be `undefined` rather than erroring, so the field is worth confirming in review. ## Model Used **TODO (author):** this section is required and cannot be completed on your behalf. Please state the provider and model name, the exact model ID/version, and the reasoning/thinking mode used — or "None — human-authored" if no AI model was involved. Per the template, the "Generated with Claude Code" footer is not a substitute for this section. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [ ] I have specified the model used (with version and capability details) — **pending author input, see above** - [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above — #6820 is a duplicate of this PR - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable — **no test added for the new field** - [x] I have updated relevant documentation to reflect my changes — not applicable, no user-facing or documented behaviour changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green — **`e2e` did not complete on `5ca5fde` (Playwright install timed out at 30m and the run was cancelled); all other checks pass** - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — **currently 4/5, sole finding being this description** - [ ] I will address all Greptile and reviewer comments before requesting merge --- <sub>This description was reformatted to `.github/PULL_REQUEST_TEMPLATE.md` by the Paperclip PR triage bot. The code was not modified. Checklist boxes reflect the PR's verifiable state at commit `5ca5fde`; unchecked items are genuinely outstanding, not oversights. The **Model Used** section requires input from the author. The previous description's `LEG-` reference was removed as an internal, instance-local identifier that the template prohibits.</sub> --------- Co-authored-by: Lead Backend Engineer Agent <backend1@legacykeeper.io> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |