mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 14:10:50 +02:00
e6ad2ea3d9fa59f222007b5cba1a94f2e89c56f4
3966
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e6ad2ea3d9 |
ci: activate stacked pull request optimization (#12509)
## Thinking Path > - Paperclip uses GitHub Actions to verify pull requests before merge. > - The pull request workflow calls a trusted reusable workflow at an immutable commit SHA. > - Pull request #12507 added stack-aware CI scope to that reusable workflow. > - The caller still points to the prior workflow commit, so the new behavior is not active. > - This pull request advances the caller to the merged commit from pull request #12507. > - The benefit is that native stack middle layers stop starting redundant full CI matrices. ## Linked Issues or Issue Description Refs #12507 **What existing behavior does this improve?** The pull request workflow still calls the trusted CI definition that predates stack-aware scope selection. **Subsystem affected** GitHub Actions pull request verification. **Current behavior** Every pull request layer in a native stack starts the complete test, build, canary, and E2E matrix. **Proposed behavior** Call the trusted workflow from merged master commit `39b8ee2960541d14b380f95365deecba6723d9bd`. That workflow runs full CI for ordinary, top, and lowest-unmerged pull requests. It keeps policy and required aggregate checks on middle layers. **Reason and benefit** This completes the two-step immutable workflow rollout from pull request #12507. Large stacks will use fewer runners and will spend less time waiting for duplicate jobs. **Breaking changes** Middle native stack layers no longer run the complete CI matrix. Required aggregate checks remain present and fail closed if the stack scope is missing or invalid. ## What Changed - Pin `.github/workflows/pr.yml` to merged master commit `39b8ee2960541d14b380f95365deecba6723d9bd`. - Activate the stack-aware trusted workflow that merged in pull request #12507. ## Verification - `node --test scripts/__tests__/e2e-shard.test.mjs` — 11 tests passed. - `actionlint .github/workflows/pr.yml .github/workflows/pr-trusted.yml` — passed. - `git diff --check origin/master...HEAD` — passed. - Confirm that the caller SHA equals the merge commit for pull request #12507. ## Risks - The caller is immutable and points to a commit that exists on `master`. - Ordinary pull requests and merge-relevant stack layers still run full CI. - A rollback can restore the prior immutable SHA in one line. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment suffix and context window are not exposed. The model used reasoning, repository tools, code execution, Git, and GitHub API access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
39b8ee2960 |
ci: optimize checks for stacked pull requests (#12507)
## Thinking Path > - Paperclip uses GitHub Actions to protect changes before they enter `master`. > - GitHub evaluates every pull request in a native stack against the stack base. > - The current workflow therefore starts the complete CI matrix for every layer in a stack. > - A large stack can queue many copies of the same integrated verification and delay every pull request. > - GitHub provides stack position and base metadata so workflows can select merge-relevant layers. > - This pull request keeps policy and required check names on every layer, but runs full CI only for ordinary pull requests, the top layer, and the lowest unmerged layer. > - The benefit is much lower CI load without weakening the required-check contract. ## Linked Issues or Issue Description **What existing behavior does this improve?** The trusted pull request workflow currently runs every test, build, canary, and E2E lane for every pull request in a native stack. **Subsystem affected** GitHub Actions pull request verification. **Current behavior** A stack with 61 pull requests can start 61 complete CI matrices after a cascading rebase. **Proposed behavior** Run the always-on policy job and stable required-check aggregators for every layer. Run the complete verification matrix only for ordinary pull requests, the top stack layer, and the lowest unmerged stack layer. **Reason and benefit** The top layer verifies the integrated stack. The lowest unmerged layer verifies the current merge candidate. Middle layers keep branch-protection checks without consuming the complete runner matrix. **Breaking changes** Middle stack layers no longer run the complete CI matrix. Their `ci / verify` and `ci / e2e` checks still require the policy job to pass and require every expensive lane to be intentionally skipped. ## What Changed - Add a fail-safe stack scope decision to the trusted PR runner gate. - Run typecheck, general tests, build, serialized tests, canary, and E2E shards only for ordinary, top, and lowest-unmerged pull requests. - Preserve the required `ci / verify` and `ci / e2e` names on every layer. - Make the required aggregators distinguish valid middle-layer skips from failures or missing scope decisions. - Add regression coverage for ordinary, top, bottom, middle, and malformed stack metadata. ## Verification - `node --test scripts/__tests__/e2e-shard.test.mjs` — 11 tests passed. - `actionlint .github/workflows/pr-trusted.yml .github/workflows/pr.yml` — passed. - `git diff --check origin/master...HEAD` — passed. - The caller remains pinned to the current trusted workflow. A separate activation change must advance the immutable SHA after this pull request lands. ## Risks - Incorrect stack classification could skip important jobs. Missing or malformed stack metadata defaults to full CI. - Middle-layer required checks depend on the policy job and verify that all expensive jobs have the `skipped` result. - The reusable workflow change does not become active until the immutable caller SHA advances in a separate change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment suffix and context window are not exposed. The model used reasoning, repository tools, code execution, Git, and GitHub API access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d7ff895a9a |
chore: add @forgottendev to CODEOWNERS (#12501)
## Thinking Path > - Paperclip uses GitHub ownership rules to protect critical repository files. > - These rules cover release infrastructure, GitHub configuration, skills, and dependency files. > - The current rules do not include `@forgottendev`. > - The new maintainer needs the same review scope as the existing code owners. > - This pull request adds `@forgottendev` to every existing CODEOWNERS rule. > - The benefit is consistent review ownership without a change to the protected path set. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change improves the GitHub review ownership for critical repository files. **Subsystem affected** Repository governance and GitHub configuration. **Current behavior** The 13 existing CODEOWNERS rules list `@cryppadotta`, `@devinfoley`, and `@nickyleach`. They do not list `@forgottendev`. **Proposed behavior** Every existing CODEOWNERS rule also lists `@forgottendev`. **Reason and benefit** This gives `@forgottendev` the same review ownership scope as the existing maintainers. It keeps ownership consistent across all protected paths. **Breaking changes** None. This change does not remove an owner or change a path pattern. ## What Changed - Added `@forgottendev` to all 13 existing entries in `.github/CODEOWNERS`. - Kept all existing owners and path patterns unchanged. ## Verification - Ran `git diff --check origin/master..HEAD`. - Confirmed that all 13 active CODEOWNERS rules contain `@forgottendev`. - Confirmed that the commit changes only `.github/CODEOWNERS`. - Confirmed that the GitHub account `forgottendev` exists. - Did not run the application test suite because this change only updates GitHub ownership metadata. ## Risks - Low risk. This is an additive ownership change. - GitHub can request review from `@forgottendev` for future pull requests that change a covered path. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5, with reasoning, shell access, and GitHub CLI tool use. The hosted context-window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.829.0-canary.2 |
||
|
|
35fca95626 |
chore(lockfile): refresh pnpm-lock.yaml (#12484)
## Thinking Path > - Paperclip uses workspace manifests to declare package dependencies. > - The lockfile records each workspace link for repeatable installs. > - PR #12469 added `@paperclipai/shared` to the Grok local adapter manifest. > - The `master` lockfile did not include this new workspace link. > - A frozen install requires the manifest and lockfile to match. > - This pull request regenerates the lockfile with the repository workflow. > - The benefit is a repeatable frozen install from `master`. ## Linked Issues or Issue Description Refs #12469 ## What Changed - Added the `@paperclipai/shared` workspace link to the Grok local adapter importer in `pnpm-lock.yaml`. ## Verification - Ran `CI=true pnpm install --frozen-lockfile` successfully. - Ran the full pull request CI suite under Node.js 24. All 23 jobs passed. ## Risks - Low risk. This pull request changes only one lockfile importer. - The package manifest already declares the dependency. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5 (`gpt-5`). The runtime does not expose the context limit. Agentic reasoning, repository tools, API tools, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
a20a4944ec |
feat: add Grok device login to the sandbox login panel (#12469)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses adapters to connect agents and model providers to its control plane > - The sandbox login panel supports displayed-code login for selected adapters > - Grok users need the same login path and a private credential home for later runs > - This pull request adds Grok support to the shared device-login path and preserves the existing Codex path > - The benefit is one secure login flow for both adapters with company-scoped credential storage ## Linked Issues or Issue Description **Agent or provider** Grok Local needs displayed-code login support in the sandbox login panel. **Why this adapter is useful** This change lets users sign in to Grok from the sandbox login panel. It also gives later Grok runs access to the stored credential. **How the agent is invoked** The Grok local adapter uses its login command through the shared displayed-code login flow. Later runs receive the managed home through `GROK_HOME`. **Additional context** The change uses adapter-scoped login lifecycle handling. It stores the credential in a company-scoped directory with mode `0700`, and it stores the credential file with mode `0600`. ## What Changed - Rename the shared device-login modules to adapter-neutral names. - Scope the shared login lifecycle to a closed adapter set. - Return the device-login URL that the provider prints. - Add the Grok prompt parser, login command, capability, and login panel entry. - Store the Grok credential in a private, company-scoped home directory. - Pass `GROK_HOME` to later Grok runs. - Add tests for the Grok adapter, the Daytona sandbox provider, the server login path, and the user interface. ## Verification - Run `pnpm vitest run packages/adapters/grok-local/src/server/adapter-auth-promotion.test.ts`. - Run the Grok adapter package suite. - Run the Daytona sandbox provider suite. - Run the server device-login suites. - Run the user interface suite. - Confirm the full CI suite passes. ## Risks The change extends shared login lifecycle code to another adapter. A regression could affect Codex login. The credential path uses explicit `chmod` calls to keep the directory at mode `0700` and the file at mode `0600`. ## Model Used OpenAI Codex, GPT-5. The runtime used tool calls and code review support. The runtime did not provide a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.829.0-canary.1 nightly/v2026.829.0-nightly.0 |
||
|
|
cec675ffca |
test(server): load the agent-permissions route module graph once per file (#12471)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server test suites verify agent permissions and route behavior > - The agent-permissions route suite rebuilt its full module graph for every test > - CPU load made that repeated work exceed the test timeout and caused intermittent failures > - This pull request loads the route module graph once for the file and resets each mock before every test > - The benefit is faster, stable test execution with the same test coverage and isolation ## Linked Issues or Issue Description **What happened?** `server/src/__tests__/agent-permissions-routes.test.ts` failed intermittently in continuous integration. The suite reset modules and re-imported the route module graph for every test. Under CPU load, one import took seconds instead of milliseconds and caused an expected response to become an HTTP 500. **Expected behavior** The suite should run all 54 cases without intermittent timeout failures. Each test should keep isolated mock state. **Steps to reproduce** 1. Run `npx vitest run server/src/__tests__/agent-permissions-routes.test.ts`. 2. Repeat the file run under high CPU load. 3. Compare the failure rate and run time before and after this change. **Paperclip version or commit** `c4d1af4216f174a92823ca3a20e0717c54371dd5` **Deployment mode** Built from source. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Not adapter-specific. This change affects a server test suite. **Database mode** Not database-related. ## What Changed - Load the route module graph one time for the describe block with the existing `hoistModuleGraph` helper. - Make `createApp` synchronous and read the hoisted graph. - Remove per-test `vi.resetModules()` and the 26 `vi.doUnmock(...)` calls. - Keep stable mock objects and reset each route-facing mock before every test. - Keep all 44 `it` blocks and 54 parameterized cases. ## Verification - `npx vitest run server/src/__tests__/agent-permissions-routes.test.ts` passes and reports 54 tests. - `npx tsc --noEmit -p server/tsconfig.json` passes with 0 errors. - Under 32 concurrent CPU-bound loops, the file passed 15 of 15 runs after this change, with 54 of 54 cases on each run. - The same test failed 1 of 15 runs before this change. - Per-run wall-clock time changed from about 29–34 seconds to about 8–10 seconds. ## Risks - Low risk. The change affects test setup only. - The hoisted mock objects keep stable identity, and `beforeEach` resets every route-facing mock. - The registration step arms no mock implementations, and the test adapter still unregisters in a `finally` block. ## Model Used OpenAI Codex, GPT-5. The runtime provided tool use and code execution. The runtime did not provide a context window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.829.0-canary.0 |
||
|
|
40d8cbc41a |
ci: activate stacked lockfile regeneration (#12464)
## Thinking Path > - Paperclip uses trusted GitHub Actions workflows to verify each pull request. > - The caller pins the reusable workflow to an immutable merged commit. > - Native stacked pull requests can inherit a stale lockfile from a parent layer. > - Pull request #12461 added safe merge-tree lockfile regeneration to the trusted workflow. > - The caller must now select that merged workflow version. > - This pull request advances the immutable pin to the merge commit from #12461. > - The benefit is reliable verification for native stacked pull requests. ## Linked Issues or Issue Description Refs #12461 ## What Changed - Pin the trusted pull request workflow to merged commit `1da6b37fc56dacf5e7ffbd31da35756a1cba41f8`. - Activate stale lockfile regeneration for native stacked pull requests. ## Verification - `actionlint .github/workflows/pr.yml` - `node --test scripts/__tests__/e2e-shard.test.mjs` ## Risks - Low risk. This changes only the immutable reusable-workflow pin. - The target commit is merged into `master` and passed all required checks. - The target workflow adds one lockfile-only install to the policy job. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with model `gpt-5`, reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ad474abece |
fix(server): load the mocked module graph once in the closed-workspace issue route suite (#12470)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server tests cover issue routes and execution workspace state > - The closed-workspace route suite reloads a mocked module graph before each test > - CPU contention can bind one test to the real service and hide a 500 response > - This pull request loads the mocked graph once and checks exact success statuses > - The benefit is a stable suite that detects route failures instead of accepting them ## Linked Issues or Issue Description **What happened?** The closed-workspace issue route suite reloaded and unmocked the module graph before each test. Under CPU contention, a route could bind to the real execution-workspaces service. The request then returned `500`, while a weak assertion accepted the result. **Expected behavior** The suite must use the configured service mocks for every test. Each success case must assert its exact expected HTTP status. **Steps to reproduce** 1. Run the closed-workspace route suite many times in parallel. 2. Use CPU contention during the run. 3. Observe intermittent failures or weak assertions that accept `500` responses. **Paperclip version or commit** Commit `0d5686b942f1d1366112cb922f95e5c8622bb9a6`. **Deployment mode** Built from source with the server test runner. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific. **Database mode** Not database-related. **Additional context** This pull request relates to [#11472](https://github.com/paperclipai/paperclip/pull/11472). It does not change product code. ## What Changed - Load the mocked service graph once per file with `hoistModuleGraph`. - Remove the per-test module reset and unmock cycle. - Assert exact success statuses for the three affected responses. - Add the missing `refreshReopenPendingConsumption` mock. - Use one named timeout constant for every `vi.waitFor` call. - Replace `setImmediate` barriers with fake-timer advances. ## Verification - `npx vitest run --project @paperclipai/server server/src/__tests__/issue-closed-workspace-routes.test.ts` passes 12 of 12 tests locally. - A 200-sample sweep with 20 parallel copies passed with zero failures. - An independent 100-sample sweep with 20 parallel copies passed with zero failures. - `npx tsc --noEmit -p server` reports the same 61 pre-existing errors with and without this change. - GitHub Actions must run the full server suite for final verification. ## Risks Low risk. This pull request changes one test file and does not change product code. The stricter assertions can expose a real route failure that the old suite hid. ## Model Used Codex, GPT-5, tool use and code review assistance. The exact context window and reasoning mode are not available in this handoff. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.828.0-canary.16 |
||
|
|
64b7dce0ad |
refactor(adapter-utils): replace the process-wide byte ledger with route-local byte bounds (#12465)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The adapter layer carries sandbox requests to host processes. > - The HTTP/2 bridge used one process-wide byte ledger for all routes. > - One busy route could exhaust that shared budget and move another route to file transport. > - This pull request gives each host retention site a fixed byte bound and limits concurrent HTTP/2 streams. > - The benefit is local protection: one route cannot consume the byte budget of another route. ## Linked Issues or Issue Description **What happened?** The HTTP/2 bridge used one aggregate byte ledger for retained bytes across all routes. A busy route could exhaust the shared budget and force an unrelated route to use file transport. **Expected behavior** Each route should protect its own retained bytes. A reset on one HTTP/2 stream should cancel only that stream's host forward. **Steps to reproduce** 1. Start the HTTP/2 bridge with multiple sandbox routes. 2. Send enough retained data through one route to reach the aggregate byte limit. 3. Send a request through a sibling route. 4. Observe that the sibling route can fall back to file transport because the first route used the shared ledger. **Paperclip version or commit** `47639e227e78e3c5e0dd1a3c0e2d792fe86895a3` **Deployment mode** Built from source with the adapter-utils and server test suites. ## What Changed - Bound each host retention site with a fixed local byte limit. - Limited concurrent live HTTP/2 streams with one built-in stream limit. - Bound each host forward and response-body read to its own HTTP/2 stream lifetime. - Removed the process-wide byte ledger, its environment override, its metrics, and its file-transport fallbacks. - Added tests for the stream limit, host body budget, and sibling-stream cancellation. ## Verification - Run `pnpm vitest run --project adapter-utils`. - Confirm that 996 adapter-utils tests pass. - Confirm that `test_live_forward_work_never_passes_the_stream_limit` passes. - Confirm that `test_the_host_body_budget_matches_the_stream_limit` passes. - Confirm that the sibling-stream cancellation test passes. - Run `pnpm tsc --noEmit`. - Confirm that all pull request checks pass. ## Risks The bridge no longer uses a process-wide byte ledger. A local bound or stream limit that is too low can reject or delay valid work. The tests cover the new limits and stream cancellation behavior. ## Model Used OpenAI GPT-5 Codex. Runtime model ID: GPT-5. The model used code execution and repository tools. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.828.0-canary.15 |
||
|
|
a3b78f9d34 |
fix(cli): make embedded-Postgres tests survive runner contention (#12466)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The CLI test suite uses embedded Postgres for worktree checks. > - Several tests start a cluster and run migrations on each test. > - Runner contention can make this cost class exceed small hand-written time budgets. > - A timed-out test can also leave the cluster alive while cleanup removes its data directory. > - This pull request gives the cost class one measured timeout and registers cleanup for the timeout path. > - The benefit is stable tests and fewer orphaned embedded Postgres processes. ## Linked Issues or Issue Description No public GitHub issue tracks this defect. Related PRs #10987 and #10961 address wider embedded-Postgres test setup issues. This pull request addresses the CLI timeout and cleanup path described below. **What happened?** Five CLI tests used different hand-written time budgets while they started embedded Postgres and ran migrations. Runner contention made one test exceed its budget. A timeout also skipped the `finally` cleanup path and left a Postgres process alive. **Expected behavior** Each embedded-Postgres test gets a budget that covers measured runner contention. Cleanup stops the cluster before the test removes its data directory, including after a timeout. **Steps to reproduce** 1. Run `npx vitest run cli/src/__tests__/worktree.test.ts` on a contended runner. 2. Observe the embedded-Postgres tests take longer than their small hand-written budgets. 3. Inspect the process list after a timeout and observe an orphaned Postgres process. **Paperclip version or commit** The defect reproduces on `master` before this pull request. **Deployment mode** Built from source with Vitest. **Installation method** Built from source with pnpm. ## What Changed - Export `EMBEDDED_POSTGRES_TEST_TIMEOUT_MS` from the embedded-Postgres test helper. - Apply the shared timeout to every test defined by `itEmbeddedPostgres`. - Replace the five hand-written timeout values in `worktree.test.ts`. - Register temporary-directory removal and cluster stop with `onTestFinished` in the correct order. - Restore the working directory during timeout cleanup. ## Verification - `npx vitest run cli/src/__tests__/worktree.test.ts` passes 63 tests. - `npx vitest run cli/src` passes 424 tests across 59 files. - `tsc --noEmit` in `cli/` adds no new error in the touched files. - CI must pass on the pull request head. - Greptile must report 5/5 with no unresolved blocking thread. ## Risks - Low risk. The change affects test helpers and test cleanup only. - The shared budget can lengthen a failing test before Vitest reports the failure. - The cleanup order depends on Vitest callback order, which the tests now use explicitly. ## Model Used OpenAI Codex, GPT-5. The model used tool calls and code review support. The context window size and internal reasoning mode were not exposed in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.828.0-canary.14 |
||
|
|
1da6b37fc5 |
fix(ci): regenerate stale stacked lockfiles (#12461)
## Thinking Path > - Paperclip uses trusted GitHub Actions workflows to verify every pull request > - Native stacked pull requests use another pull request branch as their base > - A parent layer can change a package manifest without committing `pnpm-lock.yaml` > - A child layer can inherit that manifest change without changing a manifest itself > - The current policy skips lockfile regeneration for that child and downstream frozen installs fail > - This pull request validates the complete merge tree and shares a regenerated lockfile only when needed > - The benefit is reliable stacked pull request verification without weakening the trusted workflow boundary ## Linked Issues or Issue Description **What happened?** A stacked child pull request inherited a package manifest change from its parent. The child did not change a manifest itself. The policy job skipped lockfile regeneration. Downstream jobs tried to restore an artifact that did not exist and then failed during frozen dependency installation. **Expected behavior** The policy job must validate the complete pull request merge tree. It must upload a regenerated lockfile when the checked-in lockfile is stale, including on a stacked child layer. **Steps to reproduce** 1. Create a parent pull request that changes `package.json` without committing `pnpm-lock.yaml`. 2. Create a child pull request on that branch without another manifest change. 3. Run the trusted pull request workflow for the child. 4. Observe that frozen dependency installation fails because no `pr-lockfile` artifact exists. **Paperclip version or commit** `f173ee09fa5c2ced7806bba47b54c3df853ab4df` **Deployment mode** GitHub Actions trusted pull request workflow. **Agent adapter(s) involved** Not adapter-specific. This is a core CI workflow bug. ## What Changed - Regenerate the lockfile from every checked-out merge tree. - Compare the generated lockfile with the checked-in copy before upload. - Download the artifact only when the policy job reports that it uploaded one. - Fail closed when a reported artifact is missing. - Add a workflow contract test for stacked lockfile handling. - Keep the caller pinned to the last merged trusted SHA; after this implementation merges, a separate activation PR will advance the immutable pin to its merge commit. ## Verification - `actionlint .github/workflows/pr-trusted.yml` - `node --test scripts/__tests__/e2e-shard.test.mjs` ## Risks - The policy job runs one lockfile-only install for every pull request. This can add a small amount of CI time. - A missing artifact now fails immediately when the policy job reports an upload. This is intentional because it exposes workflow corruption. - No runtime or product behavior changes. - The implementation/activation split is intentional: unmerged PR-authored workflow code must never execute on trusted runners. ## Model Used OpenAI Codex with model `gpt-5`, reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.828.0-canary.13 |
||
|
|
e127faa14c |
fix(adapter-utils): bound the ACP startup handshake and fence the abandoned session promise (#12454)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter utilities start and control agent sessions. > - The ACP startup handshake can stay pending when the sandbox transport closes. > - A pending handshake keeps the run active and prevents a clear operator result. > - This pull request bounds the handshake and fences its abandoned promise. > - The result gives each startup failure a terminal state and a safe host-authored diagnostic. ## Linked Issues or Issue Description No matching public issue or pull request appeared in the GitHub search for this failure. The issue details follow. **What happened?** The adapter engine awaited `runtime.ensureSession()` without a startup bound. A lost sandbox transport could leave the await pending. **Expected behavior** The engine must end the run when the startup deadline expires or the duplex transport closes. A late session result must not reopen the settled run. **Steps to reproduce** 1. Start an ACP-backed agent run. 2. Keep the ACP initialization call pending. 3. Let the startup deadline expire or close the duplex transport. 4. Confirm that the run reaches a terminal state and that a late session result does not reopen it. **Paperclip version or commit** `66e1c0df8b23cb8354b36dd446d9548dc4389191` merge base. **Deployment mode** Local dev (`pnpm dev`). **Installation method** Built from source (`pnpm dev`). **Agent adapter(s) involved** Custom / external plugin adapter. **Database mode** Not database-related. **Relevant logs or output** The new tests use fixed host-authored diagnostics for handshake guard failures and late close failures. **Additional context** The change updates the execution semantics document and adds regression coverage. The three existing failures in `execute.test.ts` also occur at the merge base. ## What Changed - Bound `runtime.ensureSession()` with a startup deadline and a duplex transport loss check. - Added terminal error codes for handshake timeout and transport loss. - Fenced late session resolution and rejection so the settled run has one owner. - Suppressed sandbox-controlled diagnostic values on the guard-failure and late-close paths. - Added regression tests for timeout, transport loss, late resolution, and late close rejection. - Documented the startup live-path contract. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` exits 0. - The engine test suite runs from the repository root. - The new regression cases pass. - The three known failures remain the only failures and also fail at the merge base. The board approved this pre-existing test exception. - Cold start and session resume cases pass. - All required GitHub checks pass. - Greptile reports 5/5 with no open P2 findings, recommendations, or follow-ups. ## Risks The startup guard changes only the ACP startup path. A slow but valid startup can now end at the configured deadline. The fence closes a late handle once and records fixed host-authored diagnostics. ## Model Used OpenAI Codex, GPT-5, current model version, tool use and code execution, with the full task context. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally. Three pre-existing failures remain and have an approved exception. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f173ee09fa |
ci: activate stale base handling (#12463)
## Thinking Path
> - Paperclip uses pull request checks to protect changes.
> - The trusted workflow selects GitHub or AWS runners.
> - Pull request #12462 removed an unreliable stale API equality.
> - The public caller must pin an immutable trusted workflow commit.
> - This pull request changes only that pin.
> - The benefit is correct AWS routing for trusted stacked pull
requests.
## Linked Issues or Issue Description
Refs #12339
Refs #12462
## What Changed
- Pin the thin pull request caller to
canary/v2026.828.0-canary.11
|
||
|
|
f9c32513b2 |
ci: ignore stale PR API base snapshots (#12462)
## Thinking Path > - Paperclip uses pull request checks to protect changes. > - The trusted CI gate selects GitHub or AWS runners. > - GitHub can return an old pull request base SHA after the live base advances. > - The signed event, live Git ref, ancestry, and merge parents provide the required proof. > - The stale API field rejects a safe run even when those proofs pass. > - This pull request removes that unreliable equality. > - The benefit is correct AWS routing for trusted stacked pull requests. ## Linked Issues or Issue Description Refs #12339 Refs #12459 ## What Changed - Stop treating pull request base.sha as a current-state signal. - Keep the signed event base SHA and live ref descendant check. - Keep the live base or synthetic base merge-parent proof. - Keep all numeric identity, repository, head SHA, and triggering actor checks. ## Verification - actionlint .github/workflows/pr-trusted.yml .github/workflows/pr.yml - github-runners/tests/test-workflow.sh .github/workflows/pr-trusted.yml .github/workflows/pr.yml - The corrected gate selected the Fleet label with the exact live event data from PR #12339 run 33201610330. ## Risks - The pull request API base SHA can be stale and is no longer compared. - Replaced ancestry, a changed head, a changed merge parent, and a changed merge tree still fail closed. ## Model Used - OpenAI Codex, GPT-5.6, with reasoning and terminal tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked existing public pull requests - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket ID - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation in the operations repository - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open findings - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
30ac5e116e |
ci: activate live stacked base validation (#12460)
## Thinking Path
> - Paperclip uses pull request checks to protect changes.
> - The trusted CI workflow selects GitHub or AWS runners.
> - Pull request #12459 fixed validation when a stacked base advances.
> - The public caller must use an immutable trusted workflow SHA.
> - This pull request changes only that SHA.
> - The benefit is automatic AWS routing for the affected trusted stack
runs.
## Linked Issues or Issue Description
Refs #12339
Refs #12459
## What Changed
- Pin the thin pull request caller to
|
||
|
|
f929355fb9 |
ci: allow validated stacked base advances (#12459)
## Thinking Path > - Paperclip uses pull request checks to protect changes. > - The trusted CI workflow selects GitHub or AWS runners. > - Stacked pull requests can advance their base branch while a gate waits. > - The gate already proves that the live base descends from the event base. > - An earlier exact base check rejects that safe state before the ancestry check runs. > - This pull request removes the conflicting check and verifies live API consistency. > - The benefit is automatic AWS routing for trusted stacked pull requests without weaker identity checks. ## Linked Issues or Issue Description Refs #12339 Refs #12457 ## What Changed - Allow the live pull request base SHA to advance from the signed event base snapshot. - Require the pull request API base SHA to match the live Git ref during validation. - Keep the numeric author, sender, and triggering actor checks. - Keep descendant ancestry and synthetic merge validation. ## Verification - actionlint .github/workflows/pr-trusted.yml .github/workflows/pr.yml - github-runners/tests/test-workflow.sh .github/workflows/pr-trusted.yml .github/workflows/pr.yml - The routing suite covers a live stacked base advance and a base change during validation. ## Risks - A trusted stacked run can use a newer descendant base than its signed event snapshot. - Replaced ancestry still fails closed. - A live base change during gate validation still fails closed. ## Model Used - OpenAI Codex, GPT-5.6, with reasoning and terminal tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in this pull request - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket ID - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation in the operations repository - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open findings - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c594811f3c |
ci: activate stacked merge validation (#12458)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Pull request checks protect the application and its contributors.
> - The trusted CI gate must validate both direct and stacked GitHub
merge shapes.
> - Pull request #12457 added that validation at an immutable master
SHA.
> - The active caller still pins the prior workflow version.
> - This pull request pins the caller to the newly authorized SHA.
> - The benefit is safe automatic AWS routing for trusted stacked pull
requests.
## Linked Issues or Issue Description
Refs #12457
**What existing behavior does this improve?**
The active caller uses a gate that fails closed on GitHub synthetic
stacked merge parents.
**Subsystem affected**
GitHub Actions pull request routing.
**Current behavior**
Trusted stacked pull requests run on GitHub-hosted runners after the
merge-parent check rejects the synthetic base merge.
**Proposed behavior**
The caller uses the authorized workflow SHA
|
||
|
|
7b199fcafa |
fix(ci): validate stacked PR merge refs (#12457)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request checks protect the application and its contributors. > - The trusted CI gate verifies signed event data against live GitHub state. > - GitHub gives a stacked pull request a synthetic merge commit for its base stack. > - The current validator expects the raw base SHA as the first merge parent. > - That assumption sends safe stacked pull requests to GitHub-hosted runners. > - This pull request validates the synthetic base merge and its ancestry explicitly. > - The benefit is safe automatic AWS routing for approved stacked pull requests. ## Linked Issues or Issue Description Refs #12455 Refs #12456 **What existing behavior does this improve?** The trusted runner gate validates merge commits for master-based pull requests but rejects GitHub synthetic base merges for stacked pull requests. **Subsystem affected** GitHub Actions pull request identity and merge-state validation. **Current behavior** A trusted stacked pull request passes all numeric identity checks. The gate fails closed because the first event merge parent is a GitHub synthetic base merge instead of the raw base snapshot SHA. **Proposed behavior** The gate verifies the current base-ref tip, base-snapshot ancestry, identical event and live merge parents, the child head parent, the synthetic base merge parents, and identical event and live merge trees. **Reason and benefit** The change preserves fail-closed live-state validation while allowing approved stacked pull requests to use the isolated AWS Fleet. **Breaking changes** None for untrusted contributors. Trusted stacked pull requests can select AWS after the caller pins this workflow version. **Additional context** GitHub builds a stacked test merge in two steps. It first merges the stack base into its own current base. It then uses that synthetic commit as the first parent of the child test merge. ## What Changed - Fetch and validate the current base branch ref. - Require the event base snapshot to remain an ancestor of that ref. - Validate direct and synthetic base merge parent shapes. - Preserve the existing child-head, live-state, merge-tree, repository, and actor checks. ## Verification - Ran actionlint on the trusted workflow. - Ran the external routing suite. - Added positive coverage for the GitHub stacked merge shape. - Added fail-closed coverage for replaced base ancestry and a synthetic merge that omits the current base ref. - Replayed the checks against the live merge shape for pull request #12340. ## Risks The gate has more GitHub API reads. Its five-minute timeout and fail-closed behavior limit the effect of API errors. The accepted synthetic commit must be the same in the event and live merge, must include the current base branch as a direct parent, and must produce the same merge tree. > For core feature work, check ROADMAP.md first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See CONTRIBUTING.md. ## Model Used OpenAI Codex on GPT-5. The exact deployment ID and context-window size are not exposed. The model used reasoning, tool use, GitHub API access, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked existing public items and described the issue in-PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
e48e0bd3c2 |
ci: activate trusted stacked PR routing (#12456)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Pull request checks protect the application and its contributors.
> - The trusted CI workflow sends approved contributors to isolated AWS
runners.
> - Pull request #12455 added safe support for non-master base branches.
> - The active caller still filters for master and uses the previous
immutable workflow SHA.
> - This pull request removes the base-branch trigger filter and pins
the caller to the authorized stack-aware SHA.
> - The benefit is that trusted stacked pull requests can use the AWS
fleet automatically.
## Linked Issues or Issue Description
Refs #12455
**What existing behavior does this improve?**
The active pull request caller only triggers for master and selects the
master-only trusted workflow version.
**Subsystem affected**
GitHub Actions pull request routing.
**Current behavior**
Trusted stacked pull requests either do not trigger the default caller
or use a branch-local older caller. Their heavy jobs remain in the
GitHub-hosted queue.
**Proposed behavior**
The caller triggers for every pull request base branch. It uses the
authorized stack-aware workflow at full SHA
|
||
|
|
d6b33d6c16 |
fix(ci): route trusted stacked PRs to AWS (#12455)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request checks protect the application and its contributors. > - The trusted CI workflow sends approved contributors to isolated AWS runners. > - The workflow currently limits AWS routing to pull requests that target master. > - Stacked pull requests target another branch and stay in the GitHub queue. > - This pull request keeps the identity checks and accepts a live nonempty base branch in the Paperclip repository. > - The benefit is that trusted stacked pull requests can use the AWS fleet automatically. ## Linked Issues or Issue Description **What existing behavior does this improve?** The trusted pull request runner gate currently sends only master-based pull requests to AWS. **Subsystem affected** GitHub Actions pull request routing. **Current behavior** A trusted contributor can pass all numeric identity checks. The gate still selects GitHub-hosted runners when the pull request targets another branch in a stack. **Proposed behavior** The gate accepts any nonempty base branch in the Paperclip base repository. It still verifies the live base ref, base SHA, head SHA, merge SHA, author ID, sender ID, and triggering actor ID. **Reason and benefit** Large pull request stacks currently add all heavy jobs to the limited GitHub-hosted queue. This change lets approved contributors use the isolated AWS fleet for those jobs. **Breaking changes** Trusted pull requests that target a non-master branch now use AWS instead of GitHub-hosted runners. Untrusted pull requests keep the current GitHub-hosted route. **Additional context** Pull request #12339 is the master-based first item in a stack. Its child pull requests show this queue pattern. ## What Changed - Replace the master-only route check with a nonempty base-ref check. Keep the existing base-repository and live-state checks. ## Verification - Ran actionlint on .github/workflows/pr-trusted.yml. - Ran the external routing test suite. It passed the trusted stacked-base case and all fail-closed identity cases. ## Risks The AWS trust boundary now includes code from a non-master base branch when the current pull request author, sender, and triggering actor are all approved numeric GitHub IDs. An approved account can already submit arbitrary head code. The runner remains isolated and has no production access or repository secrets. > For core feature work, check ROADMAP.md first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See CONTRIBUTING.md. ## Model Used OpenAI Codex on GPT-5. The exact deployment ID and context-window size are not exposed. The model used reasoning, tool use, GitHub API access, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8da49d6ea7 |
fix(ci): make runner image tests deterministic (#12452)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Pull request checks protect the runtime and authorization boundaries > - The same checks must produce the same result on GitHub and RunsOn Ubuntu images > - One runtime test assumed that a shell PID always owns the listening socket > - One watchdog test denied issue reads while it tried to test assignment denial > - These assumptions caused image-sensitive failures during the AWS runner canary > - This pull request tests the production contracts directly > - The benefit is a reliable CI result across both runner images ## Linked Issues or Issue Description Refs #12350 ## What Changed - Verify stale service ownership through the existing process-group ownership helper. - Allow normal issue reads in the watchdog reassignment fixture. - Assert that the watchdog reassignment reaches and denies the `tasks:assign` guard. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts -t "does not reuse a stopped auto-port service port while another process owns it"` - `pnpm exec vitest run server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts -t "still enforces normal assignment guards for watchdog reassignment"` - Ran the complete issue agent mutation ownership suite six times. All 522 test executions passed. - Ran the runtime regression case eight times. All eight test executions passed. ## Risks - Low risk. This pull request changes test fixtures and assertions only. - The process-group assertion matches the ownership rule that the runtime already uses. - The watchdog fixture still denies `issue:mutate` and `tasks:assign`. ## Model Used - OpenAI Codex with GPT-5.6 (`gpt-5.6-sol`). The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes, Closes, or Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run relevant tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes. No documentation change is required for this test-only fix. - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.828.0-canary.9 |
||
|
|
b2752b21d5 |
fix(adapter-utils): allow process sessions without birthtime (#12451)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Remote agent adapters use a process-session wrapper inside a sandbox. > - Some supported sandbox filesystems do not report inode creation time. > - The wrapper rejects a zero creation time before it launches the agent. > - This pull request accepts that filesystem shape and keeps the existing change-time probe. > - The benefit is that valid sandbox runs can start without a birth-time field. ## Linked Issues or Issue Description **What happened?** The remote process-session wrapper exited before it launched the agent when the sandbox filesystem reported `birthtimeMs` as zero. **Expected behavior** The wrapper must start on a filesystem that does not report inode creation time. **Steps to reproduce** 1. Start the remote process-session wrapper. 2. Make `lstat()` report a zero `birthtimeMs` for its session directory. 3. Observe that the pre-fix wrapper terminates before the child process starts. **Paperclip version or commit** Reproduced from `66e1c0df8b23cb8354b36dd446d9548dc4389191`. **Deployment mode** Self-hosted server with a remote sandbox runtime. ## What Changed - Allow a zero reported creation time for process-session directories. - Keep the probe that rejects a creation time copied from change time. - Add a regression test that launches and stops a session with zero birth time. ## Verification - `npx vitest run packages/adapter-utils/src/execution-target-stdin-race.test.ts` - `pnpm --filter @paperclipai/adapter-utils typecheck` ## Risks A filesystem without creation time can reduce the precision of sandbox-local path-swap detection. This change does not change host-file or Paperclip API authority. A follow-up will review that larger security posture alignment. > I checked `ROADMAP.md`. This is a focused compatibility bug fix for the existing sandbox-agent roadmap area. ## Model Used OpenAI Codex — GPT-5.6. The exact deployment suffix and context-window size are not exposed. The model used reasoning, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d9449e636e |
feat(onboarding): sign in to an agent provider during onboarding (#12440)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - New organizations create their first agent through the onboarding wizard > - The wizard does not show provider sign-in when a host credential is absent or unknown > - The create step also gives unclear feedback when the provider needs authentication > - This pull request adds a safe auth signal and a provider sign-in step for sandbox drivers > - The benefit is a clearer onboarding path with no token or account data in the signal ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (server API, shared types, and UI) **Problem or motivation** The onboarding wizard can fail when the selected provider needs authentication. It does not tell the person how to complete sign-in. **Proposed solution** Add a status-only provider auth signal. Show the sign-in panel for sandbox drivers when the signal says `absent` or `unknown`. Apply a stored Claude login to the new agent and block creation when the adapter test reports missing authentication. **Alternatives considered** The wizard could hide the sign-in panel when the signal read fails. This would hide a needed action, so this pull request shows the panel when the signal is unknown. **Roadmap alignment** The change supports the roadmap goal for scoped and audited credential bindings. **Additional context** The auth signal returns only `present`, `absent`, or `unknown`. It never returns a token, identifier, or account name. ## What Changed - Add `GET /api/companies/:companyId/adapters/:type/auth-signal` with company and permission checks. - Add shared auth-signal types and the UI query path. - Apply a stored Claude login by reference without reading its token. - Show the provider sign-in panel only for sandbox drivers with interactive terminal support. - Block agent creation when the provider test reports missing authentication. - Add route, wizard, and end-to-end test coverage. ## Verification - `pnpm --filter @paperclipai/server test adapter-auth-signal-routes` passes 50 tests. - `pnpm --filter @paperclipai/ui test OnboardingWizard` passes 69 tests. - `pnpm --filter @paperclipai/ui exec tsc --noEmit` exits with code 0. - The `e2e_shards` lane runs `tests/e2e/onboarding.spec.ts`. ## Risks The route reads a host-local readiness signal. It returns `unknown` on read errors and never exposes credential data. The UI may add a sign-in step when the signal is unavailable. ## Model Used OpenAI Codex, GPT-5, extended reasoning, tool use, and code execution. The exact context window was not provided. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
66e1c0df8b |
ci: pin PR base snapshot validation (#12450)
## Summary - rotate the thin caller from `5ac66b3fdd5dc22c0c4e5fdb234ac063cf1d9ff8` to `c119c4bee6ebb9c81791d7a6994f1be06d7cc22b` - keep both exact SHAs authorized during rotation - keep AWS routing disabled until verification completes ## Validation - actionlint passes - workflow-contract tests pass - internal trust-routing harness passescanary/v2026.828.0-canary.8 |
||
|
|
c119c4bee6 |
fix(ci): validate the PR base snapshot (#12449)
## Summary - validate the event base ref/SHA against the live PR state instead of requiring the moving `master` branch tip to remain unchanged while a hosted gate queues - retain exact event/live merge parent and tree validation, plus author/sender/rerun checks ## Canary finding A seven-minute hosted-gate queue allowed `master` to advance. Requiring the live branch tip to equal the event base snapshot would route otherwise valid trusted runs back to GitHub-hosted indefinitely on a busy repository. ## Validation - actionlint and workflow-contract tests pass - internal routing harness passes - replaced PR base snapshot, stale head, changed merge parent/tree, and untrusted actors all remain fail-closed - AWS routing remains disabled during rotation |
||
|
|
5a9c06ab66 |
ci: pin workflow merge ref validation (#12448)
## Summary
- rotate the thin caller from `b88fadb0390d2113933d5d155f16b07cbd5dafee`
to `5ac66b3fdd5dc22c0c4e5fdb234ac063cf1d9ff8`
- keep both exact SHAs authorized during rotation
- keep AWS routing disabled until verification completes
## Validation
- actionlint passes
- e2e workflow-contract tests pass
- internal routing harness validates a single output and `${{ github.sha
}}` merge-ref source
|
||
|
|
5ac66b3fdd |
fix(ci): validate the workflow merge ref (#12447)
## Summary
- use `${{ github.sha }}` as the event merge commit validated by the
trusted gate
- retain exact base/head parent and identical-tree comparison against
the current live merge ref
## Canary finding
GitHub leaves `pull_request.merge_commit_sha` empty on some `opened`
payloads even though the workflow runs against a valid merge ref. The
gate safely fell back to GitHub-hosted and no EC2 instance launched.
## Validation
- `actionlint .github/workflows/pr-trusted.yml .github/workflows/pr.yml`
- `node --test ./scripts/__tests__/e2e-shard.test.mjs`
- internal routing harness passes and asserts `EVENT_MERGE_SHA` is
sourced from `${{ github.sha }}`
- AWS routing remains disabled during rotation
|
||
|
|
c75507fd3f |
ci: pin single runner route output (#12445)
## Summary - rotate the thin PR caller from `3b295b05dc8c8dd82c12e4a9c6f721446c5cb2e8` to `b88fadb0390d2113933d5d155f16b07cbd5dafee` - keep both exact SHAs authorized during the rotation - keep AWS routing disabled until the caller and boundary verify ## Validation - `actionlint .github/workflows/pr.yml .github/workflows/pr-trusted.yml` - `node --test ./scripts/__tests__/e2e-shard.test.mjs` - internal routing harness requires exactly one runner output per gate execution |
||
|
|
b88fadb039 |
fix(ci): emit one runner route (#12444)
## Summary - emit exactly one `runner` job output from the trusted gate - write `ubuntu-latest` only inside fail-closed paths - write the Fleet label only after every identity, PR-state, merge-equivalence, and rerun-actor check passes ## Canary finding The live gate reached the trusted success notice, but GitHub retained the first of two duplicate `runner=` outputs, so policy still requested `ubuntu-latest`. No EC2 instance launched. Routing was disabled immediately. ## Validation - `actionlint .github/workflows/pr-trusted.yml .github/workflows/pr.yml` - `node --test ./scripts/__tests__/e2e-shard.test.mjs` - internal routing harness passes and now requires exactly one runner output in all cases - AWS routing remains disabled during rotation |
||
|
|
47bd4d4803 |
docs: point section 11 at the renamed implementation plan (#11677)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The `doc/` tree is how a contributor learns the system
> - Section 11 of `doc/DEPLOYMENT-MODES.md` is a map to the other
documents
> - One entry on that map points at a file that is not there
> - Commit `9c7d9ded1` moved every plan into `doc/plans/` with a date
prefix, and this link kept the old name
> - This pull request updates the link to the current file name
> - The benefit is that a reader who follows the map arrives at the plan
## Linked Issues or Issue Description
No issue exists. The problem is below.
**Issue type**
Docs. A link points at a file name that changed.
**Where is the issue?**
`doc/DEPLOYMENT-MODES.md`, line 176, in section 11 "Relationship to
Other Docs".
**What's wrong?**
The line reads:
```
- implementation plan: `doc/plans/deployment-auth-mode-consolidation.md`
```
That file is not in the repository. Commit `9c7d9ded1` ("docs: organize
plans into doc/plans
with date prefixes") renamed it to
`doc/plans/2026-02-23-deployment-auth-mode-consolidation.md`.
The renamed file is there today. The link text was not updated.
The other four entries in section 11 are correct. I checked each one:
`doc/SPEC-implementation.md`, `doc/DEVELOPING.md`, `doc/CLI.md` and
`doc/spec/invite-flow.md`
all exist.
**Suggested fix**
Use the new file name in the link. That is what this pull request does.
## What Changed
- `doc/DEPLOYMENT-MODES.md` line 176 now names
`doc/plans/2026-02-23-deployment-auth-mode-consolidation.md`
## Verification
Run this command. It prints the file, so the new name is correct:
```
ls doc/plans/2026-02-23-deployment-auth-mode-consolidation.md
```
Run this command. It prints nothing, so the old name is wrong:
```
ls doc/plans/deployment-auth-mode-consolidation.md
```
## Risks
Low risk. The change is one file name in one line of documentation. No
code changes.
## Model Used
Claude (Anthropic), model ID `claude-opus-5`, 1M context window,
extended thinking, with tool
use and code execution.
The path was found by [docproof](https://github.com/melbinjp/docproof),
an open source
checker that compares what documentation claims against what the
repository contains. A
person then read the document, checked the rename in `git log`, and
confirmed the other four
links in the same section.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
One box stays unticked and I would rather say so than tick it. I did not
run the test suite
locally, because the change is one line of prose and no test reads it.
The other two were unticked when this opened, because CI and Greptile
had not run yet. Both
have now. **29 checks pass and 2 report `skipping`** (Contributor trust,
Storybook visual
regression); none fail. Greptile returned **5/5** with no P2s: *"The
documentation-only change
appears safe to merge. The new path exists ... and replaces an obsolete
path that no longer
exists."*
|
||
|
|
de00d78854 |
ci: pin equivalent merge validation (#12442)
## Summary - rotate the thin PR caller from `d9fc93d8383ece6fba721881a7aba638867f4996` to `3b295b05dc8c8dd82c12e4a9c6f721446c5cb2e8` - keep both exact workflow SHAs authorized in the runner group during the rotation - keep `AWS_CI_ENABLED=false` until this caller is merged and verified ## Validation - `actionlint .github/workflows/pr.yml .github/workflows/pr-trusted.yml` - `node --test ./scripts/__tests__/e2e-shard.test.mjs` - full AWS/GitHub boundary verification passes with zero active runners |
||
|
|
3b295b05dc |
fix(ci): validate equivalent PR merge refs (#12441)
## Summary - validate the live master ref and event base SHA before AWS routing - accept GitHub synthetic merge commits only when the event and live commits have the exact expected base/head parents and identical tree - preserve fail-closed routing for malformed, stale, replaced, or untrusted events ## Canary finding A trusted reopened PR produced two synthetic merge SHAs with different timestamps but identical current base/head parents and tree. The former exact-SHA comparison safely fell back to GitHub-hosted runners, but could not route a valid event to AWS. ## Validation - `actionlint .github/workflows/pr-trusted.yml .github/workflows/pr.yml` - `node --test ./scripts/__tests__/e2e-shard.test.mjs` - real-event gate simulation selects the Fleet label for equivalent merge commits - negative simulations keep an untrusted sender and replaced head on `ubuntu-latest` - AWS routing remains disabled during rotation |
||
|
|
c916af0cc0 |
ci: call trusted PR workflow (#12439)
## Thinking Path > - Paperclip uses pull request CI to validate each proposed change > - The existing workflow defines every heavy job in a PR-controlled file > - A trusted reusable workflow now contains the synchronized CI definition > - The caller must use an immutable default-branch SHA > - This pull request replaces the duplicate job list with that pinned caller > - The benefit is automatic secure runner selection without workflow drift ## Linked Issues or Issue Description Refs #12436 Refs #12438 **What existing behavior does this improve?** This improves how the pull request workflow selects trusted CI capacity. **Subsystem affected** Cross-cutting CI automation. **Current behavior** The active workflow contains a duplicate list of all heavy jobs. It cannot use the administrator-controlled runner gate. **Proposed behavior** The active workflow calls the synchronized trusted workflow at an immutable SHA. The trusted workflow selects GitHub-hosted or isolated AWS capacity from the validated contributor identity. **Reason and benefit** The thin caller prevents pull request changes from replacing the external-runner security gate. It also keeps runner selection automatic. **Breaking changes** The check names gain the reusable workflow job prefix. AWS routing remains disabled until the canary starts. ## What Changed - Replaced the duplicated heavy CI job list with one reusable-workflow call. - Pinned the call to the reviewed default-branch commit. - Limited the caller token to actions, contents, and pull request read access. ## Verification - actionlint on both workflow files - Trusted-routing tests - Confirmed the pinned SHA contains the workflow and is an ancestor of master - Full AWS and GitHub runner-boundary verification with routing disabled ## Risks The check context names change when GitHub expands the reusable workflow. The rollout verifies the new aggregate contexts before branch rules change. The repository kill switch remains off during this pull request. ## Model Used OpenAI Codex with GPT-5, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked a related public PR or described the issue with the matching template fields - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d9fc93d838 |
ci: synchronize trusted PR policy (#12438)
## Thinking Path > - Paperclip uses pull request CI to protect changes before merge > - The trusted reusable workflow will select isolated AWS capacity > - The active workflow changed while the reusable workflow waited for merge > - The reusable policy must contain every current CI policy step before activation > - This pull request synchronizes the migration-order check and shell validation > - The benefit is one reviewed workflow version with verified job parity ## Linked Issues or Issue Description Refs #12436 **What existing behavior does this improve?** This improves the pull request CI workflow synchronization before AWS runner activation. **Subsystem affected** Cross-cutting CI automation. **Current behavior** The active workflow validates migration order. The new reusable workflow does not yet contain that check. **Proposed behavior** Both workflow definitions contain the same heavy jobs and policy steps before the active workflow becomes a thin caller. **Reason and benefit** The synchronization prevents policy drift during the two-step secure rollout. **Breaking changes** None. AWS routing remains disabled. ## What Changed - Added the current migration-order validation to the trusted workflow. - Added the existing shellcheck intent annotation to the active workflow. - Verified normalized heavy-job parity between both definitions. ## Verification - actionlint on both workflow files - Local trusted-routing and normalized workflow-parity tests - git diff --check ## Risks Low risk. The migration check already runs in active CI. This change copies it into the inactive trusted definition. AWS routing stays disabled. ## Model Used OpenAI Codex with GPT-5, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked a related public PR or described the issue with the matching template fields - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
07b6816829 |
ci: add trusted reusable PR workflow (#12436)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Pull request checks protect the quality of the control plane > - The current checks depend only on the shared GitHub-hosted runner limit > - Busy periods leave many pull request jobs queued even when external capacity is available > - Public pull request code must not select or directly access private runner infrastructure > - This pull request adds an inactive reusable workflow with a fail-closed identity gate > - A later pull request can pin this workflow by its full master commit SHA > - The benefit is automatic, controlled access to isolated runner capacity without changing current CI during bootstrap ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the pull request CI workflow. It prepares the existing checks to use an administrator-controlled runner selection. **Subsystem affected** Cross-cutting GitHub Actions CI configuration. **Current behavior** Every pull request job uses `ubuntu-latest`. Jobs wait when the GitHub-hosted concurrency limit is full. **Proposed behavior** Add a reusable copy of the current PR workflow. A GitHub-hosted gate validates durable numeric user IDs and current GitHub API state. The gate emits one runner label. The default and every validation failure use `ubuntu-latest`. The AWS label is possible only when an administrator enables it and every identity check passes. This bootstrap pull request does not change the active `.github/workflows/pr.yml` caller. A follow-up change will call this workflow by the full master commit SHA. **Reason and benefit** The split bootstrap creates an immutable trust boundary before external runners are reachable. It also keeps CI automatic for contributors. Contributors do not select a runner. **Breaking changes** None in this bootstrap pull request. The active PR workflow does not change. ## What Changed - Added an inactive `workflow_call` copy of the current PR checks. - Added a GitHub-hosted routing gate that checks the repository ID, pull request author ID, event sender ID, rerun actor ID, base branch, head SHA, merge SHA, and current pull request state. - Made every validation failure select `ubuntu-latest`. - Pinned every third-party action to a full commit SHA. - Disabled persistent checkout credentials for all jobs. - Limited the workflow token to Actions read, contents read, and pull request read access. ## Verification - `actionlint .github/workflows/pr-trusted.yml` - Ran the dedicated workflow routing test harness against `.github/workflows/pr-trusted.yml`. - Compared the job keys with `.github/workflows/pr.yml`. The new workflow contains every existing job plus the gate. - Verified each pinned action commit against its current GitHub major-version tag. ## Risks The gate could route a trusted pull request to the wrong runner if an identity check is incomplete. The gate checks durable numeric IDs from the event and current GitHub API state. It checks the rerun actor separately. It defaults to GitHub-hosted capacity before any validation runs. This file is inactive in this pull request. The follow-up caller and runner-group restriction must use the exact commit that reaches `master`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact serving model ID and context-window size are not exposed in this environment. The model used high-reasoning, terminal, GitHub API, browser, and web-research capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5f1e25c112 |
test(server): load the company-skills route module graph once per file (#12426)
## Thinking Path > - Paperclip is an open source app that helps people manage AI agents for work. > - The server provides company-scoped routes for company skills and their test runs. > - The authorization tests for these routes fail intermittently in continuous integration. > - The failure returns HTTP 500 instead of the expected HTTP 403. > - The test file resets and rebuilds its module graph before each test. > - This rebuild imports an unmocked issue service and raises a TypeError. > - This pull request loads the module graph once per describe block. > - The change keeps the test result stable and preserves all 53 tests. ## Linked Issues or Issue Description **What happened?** The company-skill test-run authorization tests failed intermittently in continuous integration. One test returned HTTP 500 instead of HTTP 403. **Expected behavior** Each unauthorized request must return HTTP 403. The test file must keep all 53 tests and skip none. **Steps to reproduce** 1. Run `npx vitest run server/src/__tests__/company-skills-routes.test.ts` before this change. 2. Repeat the run in continuous integration. 3. Observe the intermittent HTTP 500 result in an authorization case. **Paperclip version or commit** Commit `651d26a96f6e24811d336759d3e67ff3abb5ec29`. **Deployment mode** Built from source. Continuous integration runs the test suite. **Agent adapter(s) involved** Not adapter-specific. This issue affects server test module setup. **Database mode** Not database-related. ## What Changed - Load the mocked route module graph once for each describe block. - Use the existing `hoistModuleGraph` helper, as the cost service test does. - Remove twelve per-test `vi.doUnmock` calls and the redundant module rebuild. - Keep the change in `server/src/__tests__/company-skills-routes.test.ts` only. ## Verification - Run `npx vitest run server/src/__tests__/company-skills-routes.test.ts`. - Confirm that 53 tests pass and 0 tests skip. - Confirm that the diff changes only `server/src/__tests__/company-skills-routes.test.ts`. - Confirm that all Paperclip continuous integration checks pass. ## Risks Low risk. The change affects test setup only. It does not change production code, route behavior, or assertions. ## Model Used OpenAI GPT-5 (Codex), exact model ID `gpt-5`, tool use and code execution enabled. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dbf052577d |
Follow the current onboarding arc in the release smoke (#12423)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - A release is gated by the release smoke: it installs the published
`paperclipai` artifact into a Docker container and drives the sign-in →
onboarding → first-agent path with Playwright
> - That suite runs only from the release pipeline, never on a pull
request, so it sees the UI only after the UI has already changed
> - The onboarding wizard was rebuilt into the agent arc. The "Name your
organization" step, the "Start Onboarding" launcher, and the agent role
picker are all gone
> - The spec still waited for those, so it failed on its first assertion
and blocked every nightly and beta release
> - The failure was also hard to read. The workflow uploaded no
container logs, because it learned the container's name only after the
harness succeeded, and the harness ran the container with `--rm` and
deleted it before anything read it
> - This pull request rewrites the spec to follow the current arc, and
repairs the log capture at both ends
> - The benefit is that nightly and beta releases are unblocked, and the
next failure arrives with the logs attached
## Linked Issues or Issue Description
No existing issue. Describing it inline, following
`.github/ISSUE_TEMPLATE/bug_report.yml`.
Refs #12274 (removed the company-naming step from the wizard).
Refs #12135 (the previous alignment of this spec, before #12274).
Refs #12316 (open; also edits `scripts/docker-onboard-smoke.sh`, in the
bootstrap helpers rather than the container lifecycle, so the two
changes do
not overlap. Whichever lands second should rebase and re-run).
**What happened?**
The release smoke fails.
`tests/release-smoke/docker-auth-onboarding.spec.ts`
never gets past its first wait:
```
✘ tests/release-smoke/docker-auth-onboarding.spec.ts:43:3 › Docker authenticated onboarding smoke › logs in, completes onboarding, and hires the lead agent
Error: expect(locator).toBeVisible() failed — element(s) not found (timeout 20000ms)
> 33 | await expect(wizardHeading.or(startButton)).toBeVisible({ timeout: 20_000 });
```
The spec waits for an `h3` reading "Name your organization" or a
"Start Onboarding" button. Neither exists. #12274 removed the
company-naming
step; the string now survives only in a code comment and in
`ui/src/components/OnboardingWizard.step.test.tsx`, which asserts it is
*absent*. The steps after the first wait are stale too: the CTA on step
1 is
"Continue" and not "Next", the organization input's placeholder changed,
and
the agent step's `#onboarding-agent-role` picker is gone, so every
onboarding
hire is filed under the neutral `general` role.
The suite runs only from the release pipeline, so nothing on a pull
request
saw the drift. Both `smoke_nightly` and `smoke_beta` call the same
reusable
workflow, so every nightly and every beta was blocked.
The failure also arrived without diagnostics. The job's "Capture Docker
logs"
step is `if: always()`, but it is guarded on `SMOKE_CONTAINER_NAME`,
which the
"Launch Docker smoke harness" step writes to `$GITHUB_ENV` only *after*
the
harness returns. On any failure before that the guard is false, the step
does
nothing, and the upload reports "No files were found". Below that,
`scripts/docker-onboard-smoke.sh` starts the container with
`docker run -d --rm`, so the `docker stop` in its EXIT trap deletes the
container and its logs together — and a container that crashes on its
own is
removed the instant its process exits.
**Expected behavior**
The spec walks the onboarding arc the app actually presents, and proves
the
company is created, the lead agent is hired, and the first task is
seeded and
dispatched. When the smoke fails, the run's artifact carries the
container's
logs.
**Steps to reproduce**
1. Run the Release Smoke workflow against a published artifact that
carries
#12274, or run it locally:
`PAPERCLIPAI_VERSION=2026.828.0-canary.3 SMOKE_DETACH=true
./scripts/docker-onboard-smoke.sh`
2. Run `pnpm run test:release-smoke` against that container.
3. The single spec fails at `openOnboarding()` after 20 seconds.
4. In CI, open the run's `release-smoke` artifact. It has no
`docker-onboard-smoke.log`.
**Paperclip version or commit**
`2026.828.0-canary.3` (commit
nightly/v2026.828.0-nightly.0
beta/v2026.828.0-beta.0
v2026.831.0
canary/v2026.828.0-canary.5
|
||
|
|
7b73b08250 |
ci: enforce migration order against PR target (#12433)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - Paperclip applies database migrations in numeric order. > - Two branches can create the same migration number before either branch merges. > - The existing repository check can find duplicates only after both histories are present in one checkout. > - A pull request must compare its new migrations with the target branch before merge. > - This pull request adds that comparison to the existing PR policy job. > - The benefit is an early failure with exact renumbering instructions. ## Linked Issues or Issue Description **What existing behavior does this improve?** The PR policy check for files in `packages/db/src/migrations`. **Current behavior** A stale branch can add the same migration number as the target branch. The existing check does not compare PR additions with the target branch migration tip. **Proposed behavior** The policy job fails when a new PR migration number is not greater than every migration on the target branch. The error names the conflict, the next safe number, and the related files to update. **Reason and benefit** This prevents duplicate or out-of-order migration numbers from reaching `master`. It also gives contributors and agents a direct repair procedure. **Breaking changes** None. The change rejects migration numbering that is already unsafe. ## What Changed - Added a dependency-free check that compares new PR migration files with the target branch tip. - Added the check to the existing PR policy job. - Added tests for no-op, valid, duplicate, and lower-number cases. ## Verification - `node --test '.github/scripts/tests/*.test.mjs'` passed 133 tests. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm test:run` passed 4,889 tests and failed 24 unrelated macOS path and wildcard-listener tests that also affect the current `master` checkout. ## Risks - Low risk. The check reads Git history and does not modify migrations. - The check permits gaps. It only requires each new migration number to follow the target branch tip. - The existing migration check continues to validate duplicate numbers, snapshots, and journal entries inside the PR. ## Model Used OpenAI Codex, GPT-5. The exact serving model ID and context-window size are not exposed in this session. Reasoning, tool use, web access, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.828.0-canary.4 |
||
|
|
8316ceb0b9 |
Add the Better Auth issuer column so signup and sign-in work (#12396)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A self-hosted install in `authenticated` mode signs users in with Better Auth, mounted at `/api/auth` over a hand-written Drizzle `account` table in `packages/db` > - Better Auth 1.7.0 added a required `issuer` field to that `account` model, plus a unique index on `(issuer, accountId)` > - The dependency bump in #11886 changed only `server/package.json` and the lockfile, so the Drizzle table never grew the column > - The Drizzle adapter checks the model against the schema on every write, so `linkAccount` throws and sign-up answers 500 with an empty body; a fresh install cannot create its first user, and an upgraded install locks out every existing user > - This pull request adds the `issuer` column and its unique index, and migrates the column in with a backfill that covers every existing row > - The benefit is that sign-up and sign-in work again, on a new install and after an upgrade ## Linked Issues or Issue Description No existing issue. Describing it inline, following `.github/ISSUE_TEMPLATE/bug_report.yml`. Refs #11886 (the dependency bump that introduced the required field). Refs #12269 (an earlier attempt at this fix; its backfill covers only `provider_id = 'credential'`). **What happened?** Sign-up fails on a self-hosted install. `POST /api/auth/sign-up/email` answers HTTP 500 with a zero-byte body. The server log carries: ``` [Better Auth]: The field "issuer" does not exist in the "account" Drizzle schema. # SERVER_ERROR: [BetterAuthError: The field "issuer" does not exist in the "account" Drizzle schema.] ``` The request writes the `user` row and then fails on the `account` row. The address is stuck after that: a second sign-up answers 422 `USER_ALREADY_EXISTS`, sign-in answers 401, and password reset answers 400 `RESET_PASSWORD_DISABLED` because the account that would hold the password does not exist. An upgraded install is worse. `sign-in/email` matches the credential account on `account.issuer === 'local:credential'`. Rows written before the upgrade have no issuer, so every existing user is locked out. **Expected behavior** `POST /api/auth/sign-up/email` answers 2xx and writes both the `user` row and its credential `account` row. `POST /api/auth/sign-in/email` then answers 2xx and sets a session cookie. An install that upgrades keeps its existing users. **Steps to reproduce** 1. Start a server from `master` with `PAPERCLIP_DEPLOYMENT_MODE=authenticated` against an empty database. 2. `curl -X POST http://127.0.0.1:<port>/api/auth/sign-up/email -H 'Content-Type: application/json' -H 'Origin: http://127.0.0.1:<port>' --data '{"name":"A","email":"a@example.com","password":"a-long-password"}'` 3. The response is HTTP 500 with an empty body. **Paperclip version or commit** `master` atcanary/v2026.828.0-canary.3 |
||
|
|
dc7a1a020a |
fix(adapter-utils): skip the remote session close when the duplex channel is already lost (#12394)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The adapter runtime settles each run through a duplex control
channel
> - A lost channel can leave the remote session-close call without a
usable peer
> - The call has no deadline, so run teardown can wait for the full
adapter timeout
> - This pull request skips that remote call after the runtime latches
channel loss
> - The benefit is faster run finalization while the local cleanup
effects remain
## Linked Issues or Issue Description
**What happened?**
The run teardown placed a remote session-close call over a duplex
control channel that the runtime had already latched as lost. The call
blocked until the adapter execution timeout released it.
**Expected behavior**
Run teardown should release the local warm handle and continue when the
duplex control channel has already failed.
**Steps to reproduce**
1. Start an adapter run with the duplex control channel.
2. Latch a channel-loss state before settlement.
3. Use a runtime whose close call never resolves.
4. Confirm that teardown returns without a remote close call.
**Paperclip version or commit**
Commit
canary/v2026.828.0-canary.2
|
||
|
|
7895f7f2b0 |
Install the declared Sentry server package into the hosted image (#12330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip supports opt-in Sentry error monitoring for server and browser errors. > - The hosted image must include the server package when an operator sets SENTRY_DSN. > - The server package is an optional peer in the source tree, so the image did not include it. > - This pull request installs the declared server package in the hosted image and checks the result. > - The benefit is a hosted tenant can send server errors without a manual package install. ## Linked Issues or Issue Description No public issue exists for this change. **What happened?** The hosted image did not include the declared @sentry/node server package. A hosted tenant could set SENTRY_DSN, but the server could not load the package from the image. **Expected behavior** The hosted image must include the exact @sentry/node version from server/package.json. The self-hosted image must remain without this optional package. **Steps to reproduce** 1. Build or pull the hosted image. 2. Resolve @sentry/node from the server package path. 3. Compare its version with server/package.json. 4. Confirm that the tsx loader path still resolves. **Paperclip version or commit** Commit b6ff556a33ebdbe764b7f495951cd59009776608. **Deployment mode** Docker hosted image. ## What Changed - Add a cloud-server-deps Docker stage that installs the declared @sentry/node version in isolation. - Copy the isolated package into the cloud image without changing the production image. - Add a probe that checks the tsx loader and the resolved Sentry version. - Run the probe after the hosted image push in the Docker workflow. - Add server tests and update the observability documentation. ## Verification - Run `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/cloud-image-sentry.test.ts`. - Confirm that the changed test passes in CI. - Confirm that all pull request checks pass. - Note that the Docker workflow does not run for pull requests. It runs after a push to master, for configured tags, or after manual dispatch. ## Risks - Low risk. The production image body stays unchanged. - The cloud image adds the declared Sentry package and a small dependency tree. - The workflow probe fails if the image loses the tsx loader or resolves a different Sentry version. ## Model Used OpenAI GPT-5; exact model version supplied by the execution service; tool use and code execution; context window not specified. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b7fd6c59b6 |
fix(server): load the costs-service route module graph once per file (#12375)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server test suite checks budget and cost routes > - The costs-service test rebuilt the full route module graph before every test > - Synchronous graph rebuilds caused long stalls under CPU load > - This pull request loads the mocked graph once for each describe block and keeps per-test mock setup > - The benefit is a stable 17-test file without production code changes ## Linked Issues or Issue Description **What happened?** The costs-service route test rebuilt its full mocked module graph before every test. Under CPU load, a rebuild sometimes stalled a test past the 15-second timeout. **Expected behavior** The test file should load its mocked route graph once for each describe block while each test keeps isolated mock behavior. **Steps to reproduce** 1. Run the costs-service route test under synthetic CPU load. 2. Repeat the file test 30 times. 3. Observe intermittent test timeouts before this change. **Paperclip version or commit** The test used the current master branch at the time of this change. **Deployment mode** Built from source. ## What Changed - Add `hoistModuleGraph` to load the mocked route graph once for each describe block. - Keep per-test mock setup in `beforeEach` so test isolation stays unchanged. - Keep all 17 tests and their assertions. - Remove the module graph rebuild from the per-test path. ## Verification - Run `npx vitest run src/__tests__/costs-service.test.ts` from `server/`. - Confirm that the file reports 17 tests and zero skipped tests. - Confirm that 30 runs under the same synthetic CPU load report 0.0% failure after the change, compared with 10.0% before the change. - Confirm that mutation checks still fail when each authorization guard is broken. ## Risks This change affects test setup only. The main risk is weaker test isolation if a mock keeps state between tests. Each test still re-arms its mock behavior in `beforeEach`, and the full assertion set remains. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This bug fix does not add a core feature. ## Model Used OpenAI GPT-5. The model used tool calls and code execution. The exact context window and reasoning configuration were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (the submitting engineer ran the file before handoff) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.828.0-canary.1 |
||
|
|
4436cf00a2 |
Render the onboarding agent arc's real steps in Storybook (#12369)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - New tenants meet it through an onboarding wizard: create your agent, connect a model, review > - Those three screens only render for a signed-in account that owns a provisioned stack, so the only way to look at one was to walk a real signup > - Round 4 redesigned all three and shipped them without anyone seeing them render; when the connect step then failed on a live stack, the review step behind it could not be reached at all > - Storybook already mounts the app's provider stack and stubs `/api`, and already has an `Onboarding/Agent arc` story — but that story previews `AgentCapsule`, which the wizard stopped using in round 4 > - This pull request mounts the real wizard in Storybook at each step, and replaces the stale capsule stories with the component the arc actually renders > - The benefit is that these screens can be reviewed, and regressions seen, without provisioning anything ## Linked Issues or Issue Description No existing issue. Describing it inline, following `.github/ISSUE_TEMPLATE/enhancement.yml`. **What existing behavior does this improve?** Reviewing the tenant onboarding wizard. Today it can only be seen by signing up for a real account and provisioning a real stack. **Subsystem affected** `ui` — the onboarding wizard and its Storybook coverage. **Current behavior** `OnboardingWizard` renders only for a signed-in account that owns a company. There is no route, harness, or story that mounts it, so no screen in the agent arc can be looked at in isolation. The existing `Onboarding/Agent arc` story previews `AgentCapsule` in `slot`/`configured`/`online` and describes it as what "the wizard holds in one tree slot" — round 4 replaced that with `PillGuy` and `dormant`/`alive`, so the story documents a component the arc no longer renders. **Proposed behavior** Stories that mount the real `OnboardingWizard` at each of the three steps against the existing Storybook API fixtures, plus stories for `PillGuy` and its transition. **Reason and benefit** Round 4 shipped three redesigned screens that nobody could see. The connect step then failed on staging, which made the review step unreachable even with an account — reviewing it required hand-editing `localStorage`. Stories remove that whole class of problem. **Breaking changes** None. Storybook-only; no product code is touched. ## What Changed - `CreateYourAgent`, `ConnectAModel`, `Review` — the real wizard, per step. - `PillStates` and `PillMorph` — the two states, and the transition on a loop with a toggle. The morph is the arc's payoff and the hardest thing to judge from a still. - Replaces the `AgentCapsule` stories in this file with `PillGuy`. `AgentCapsule` is still used by `DesignGuide` and keeps its coverage there. - Four routes added to the Storybook fetch fixtures: `/api/instance/settings`, `…/environments`, `…/adapters/:type/models`. The empty environment list is the cloud-tenant shape, and also the state that produces the "no managed sandbox environment is available" notice — worth being able to look at rather than only meeting it on a live stack. ## Three properties of the wizard the stories had to respect Each of these cost a debugging cycle, so they are documented at the call site: 1. **The draft is seeded during render, not in an effect.** Roughly twenty `useState(saved?.x ?? default)` initializers read the restored blob exactly once, so a draft written after mount arrives too late. 2. **Nothing mounts until the companies list settles.** The wizard's own mount gate waits on `isFetching`, but that query is *disabled* until the account settles, and a disabled query is not fetching. Mounting straight away gets an inner wizard that reads a null draft, falls back to `initialStep`, and then persists that back over the seed. A real session never hits this because the dashboard has already loaded the list. 3. **The review step opens with no `initialStep`.** An explicit option takes precedence over saved state by design, so passing one clamps 5 to 4 and lands on Connect. ## Verification - `npx vitest run src/components/OnboardingWizard.test.tsx src/components/OnboardingWizard.step.test.tsx` — 44 tests, all passing. - `npx tsc -p tsconfig.json --noEmit` — no new errors. (`src/lib/sentry.*` reports pre-existing missing-type errors for `@sentry/browser` on master.) - Each of the three step stories loaded in a running Storybook and read back: - `create-your-agent` → "STEP 1 OF 3 / Create your first agent / Name / Next" - `connect-a-model` → "STEP 2 OF 3 / Connect a model / Paperclip works with your existing subscription or API keys. / Claude Code / Codex / Advanced settings / Connect" - `review` → "STEP 3 OF 3 / Let's get started... / Darnold is ready to work! / Get started" One difference from a real walk, worth knowing before treating a story as ground truth: the review story has no Back button, because `entryStep` is 5 there and back-navigation is bounded by where the run entered. ## Risks Low. Storybook-only — no product code, routes, or bundles change. The four added fetch fixtures are inside the Storybook mock and cannot affect the app. The one thing to watch: these stories mount the real wizard, so a future change to how it restores drafts or decides its initial step can break them. That is arguably the point — the stories would be the first place it shows — but it does mean they are coupled to internals rather than to a prop surface, and the three notes above are what a maintainer needs. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use: file editing, shell, and browser automation for loading and reading back each story. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>canary/v2026.828.0-canary.0 |
||
|
|
17ebcc65b7 |
refactor(daytona): simplify the login session-home create (#12348)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers let agents run in remote environments > - The Daytona login flow creates a home directory for each login session > - The create path ran owner, mode, and link-type checks inside the sandbox > - These checks cannot protect the host because sandbox code can change the checked state > - This pull request uses one `mkdir -p` command and removes the unused helper scripts > - The benefit is a simpler login path with the host-side credential checks unchanged ## Linked Issues or Issue Description **What happened?** The Daytona login flow used helper scripts and inside-sandbox checks for the session-home directory. The standalone package build also copied a scripts directory that no longer existed after the helper scripts were removed. **Expected behavior** The login flow must create the session home with one `mkdir -p` command. The package build must complete without copying a removed directory. **Steps to reproduce** 1. Build the Daytona plugin package. 2. Start a Daytona device login. 3. Inspect the session-home create command and the package output. **Paperclip version or commit** Commit `dfdf5914ba37caa1e3bc380236844a7d76237e12`. **Deployment mode** Built from source. ## What Changed - Replace the session-home helper checks with one `mkdir -p` command. - Remove the two unused session-home helper scripts. - Remove the dead build copy steps for the deleted scripts directory. - Add coverage for a failed session-home create command. - Keep the host-side credential reader unchanged. - Keep the Kubernetes provider package unchanged. ## Verification - Package unit tests pass: 220 passed, 6 skipped. - Package typecheck passes. - The package build passes and emits 56 files in `dist`. - The roadmap check confirms that this change stays within the planned sandbox-provider work. - GitHub search found no open duplicate or related pull request. ## Risks The login flow no longer reports owner, mode, or link-type errors from inside the sandbox. Those checks did not protect the host. The host-side credential reader still uses no-follow path opens and accepts only a regular file with owner and exact mode `0600`. Risk is low because this change removes checks that cannot enforce the host security boundary. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This pull request updates an existing sandbox-provider path, not a new core feature. ## Model Used OpenAI Codex, GPT-5. The runtime provides tool use and code execution. The runtime does not expose the context window size or a more specific model identifier. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.827.0-canary.9 |
||
|
|
3d9be6f7fe |
refactor(daytona): simplify the inbound file-sync path (#12329)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers run agent work in isolated environments > - The Daytona inbound file-sync path writes selected sandbox data to host paths > - Its in-sandbox lexical and realpath guards do not protect host-write or Paperclip API authority > - This pull request removes those guards and simplifies the inbound file-sync commands > - The benefit is a smaller path with the same outbound controls and host extraction validation ## Linked Issues or Issue Description No public issue exists for this change. This pull request describes the enhancement. **What existing behavior does this improve?** It improves inbound file and directory synchronization for the Daytona sandbox provider. **Subsystem affected** `packages/plugins` — the Daytona sandbox provider. **Current behavior** The inbound path checks lexical and realpath confinement inside the sandbox. It uses file-descriptor-pinned commands for file promotion, archive extraction, and decompression. These checks do not protect host-write or Paperclip API authority. **Proposed behavior** Remove the inbound lexical and realpath checks. Use `mv -f` for file mappings, `tar -xf` for directory mappings, and `zstd -d -o` for decompression. Keep outbound source checks, atomic snapshot downloads, and tarball member validation. **Reason and benefit** The sandbox boundary protects the relevant authorities. The removed guards run inside that boundary and only produce early errors. The simpler commands reduce code and preserve the controls that protect the host boundary. **Breaking changes** The inbound path no longer rejects mappings because of sandbox-side lexical or realpath confinement. Outbound source validation and tarball member validation remain unchanged. ## What Changed - Remove lexical and realpath confinement guards from inbound file mappings, inbound directory mappings, and post-upload command working directories. - Replace file-descriptor-pinned promotion with one `mv -f` command per file mapping. - Replace file-descriptor-pinned extraction with one `tar -xf` command per directory mapping. - Replace retained-descriptor decompression with `zstd -d -o`. - Keep outbound source guards, atomic snapshot-and-download, and tarball member validation. - Update Daytona tests for the new command shapes and remove tests for the removed rejections. ## Verification - The author ran the Daytona package test suite: 232 tests passed and 6 tests skipped. - The author ran the Daytona package typecheck successfully. - Review the diff and confirm it changes only the three Daytona files named in this description. - Confirm the Storybook check may report `SKIPPED` as an expected repository state. ## Risks The inbound path now trusts the sandbox boundary for host-write protection. A later change that gives sandbox code host-write or Paperclip API authority could require new guards. Outbound source checks and archive member validation remain in place. No Kubernetes or core runtime file changes exist. ## Model Used OpenAI GPT-5; exact model version and context window were not provided; tool use and code review assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.827.0-canary.8 |
||
|
|
036600d922 |
fix(db): close test database clients before the embedded Postgres cluster stops (#12335)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses database clients and embedded PostgreSQL test fixtures > - A fixture stopped its embedded PostgreSQL cluster while clients still held connections > - The postgres.js driver then scheduled a write on a stopped connection > - That write escaped the timer callback and caused a test process to exit with an error > - This pull request closes registered clients before the fixture stops its cluster > - The benefit is stable test teardown and clear failure reporting in continuous integration ## Linked Issues or Issue Description Refs: #10869 **What happened?** An embedded PostgreSQL test fixture stopped its cluster while database clients still held open connections. The postgres.js driver then scheduled a deferred write on a dead connection. The write caused an unhandled error after the test shard reported success. **Expected behavior** The fixture closes all live clients for its cluster before it stops the embedded PostgreSQL cluster. Tests then finish without a deferred write on a dead connection. **Steps to reproduce** 1. Run the database regression test with the embedded PostgreSQL fixture. 2. Stop the fixture while its database client still has an open connection. 3. Observe the deferred write and the process exit status. **Paperclip version or commit** Branch base: |
||
|
|
666f5a6e69 |
fix(server): make the wake-claim lease test deterministic (#12331)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server test suite checks question response delivery and wake claims > - One test used wall-clock time and could start a second delivery under load > - The second delivery reused one promise resolver and could hang for 15 seconds > - This pull request uses the injected clock and one resolver for each wakeup > - The benefit is a deterministic test that fails at once if a second wakeup occurs ## Linked Issues or Issue Description **What happened?** The wake-claim lease test slept for 70 milliseconds before it ran the pending sweep. Under load, the lease could look stale during that interval. The sweep then started a second delivery. The second delivery reused one promise resolver, so the test hung until the 15-second suite timeout. **Expected behavior** The test must control the time used by the service. One wakeup must use one resolver. An unexpected second wakeup must fail at once. **Steps to reproduce** 1. Run the question response delivery test under CPU load. 2. Let the test sleep before the pending sweep. 3. Observe that a second delivery can start and the test can reach the 15-second timeout. **Paperclip version or commit** Commit `0dd735e53a5cc9f6d3395834b826dfb0b1da2ea9`. **Deployment mode** Built from source with the server test suite. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific (core test issue). **Database mode** Not database-related. **Additional context** The change affects one test file. It does not change production source code. ## What Changed - Drive the test with the service's injected clock. - Give each wakeup call its own promise resolver. - Assert that lease renewal advances the last attempt time. - Assert that the wakeup runs one time and the attempt count stays at 1. ## Verification - Run `pnpm exec vitest run server/src/services/__tests__/question-response-delivery.test.ts`. - The changed file reports 29 passing tests. - Run the changed file 25 times, including 5 runs under CPU load. ## Risks Low risk. The change affects one test file and test setup only. It does not change production behavior. ## Model Used OpenAI GPT-5, exact runtime model ID supplied by the Paperclip agent environment, tool use and code review assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.827.0-canary.7 |
||
|
|
bdd8f1bedb |
Repair the drizzle snapshot so generate emits no spurious migration (#12333)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip keeps its state in PostgreSQL, and `packages/db` owns that schema through Drizzle > - `drizzle-kit generate` writes a new migration by diffing `packages/db/src/schema/` against the newest snapshot in `packages/db/src/migrations/meta/`, so the snapshot must describe the schema that the migrations produce > - Snapshot `0228` recorded the new `error_count` column on the wrong table, and snapshot `0229` inherited the error, so the newest snapshot no longer matched the schema > - Because of that, `generate` on `master` folded the drift into any new migration: it emitted an `ADD COLUMN` for a column that migration `0228` already creates, which fails on a fresh database, plus an out-of-scope `DROP COLUMN` > - This pull request moves the column entry to the correct table in both snapshots and adds a test that repeats the diff `generate` performs > - The benefit is that the next person who generates a migration gets only their own change, and CI fails if the snapshot drifts again ## Linked Issues or Issue Description No existing issue. The description below follows `.github/ISSUE_TEMPLATE/bug_report.yml`. **What happened?** `drizzle-kit generate` on `master` emits a wrong migration. The newest snapshot, `packages/db/src/migrations/meta/0229_snapshot.json`, disagrees with the schema in two places. It omits `issue_question_response_deliveries.error_count`, which `0228_nasty_grim_reaper.sql` creates. It also carries `decision_archive_notification_outbox.error_count`, which no migration ever creates and the Drizzle schema never declared. Snapshot `0228` introduced both halves of the error: it added the new `error_count` column to `decision_archive_notification_outbox` instead of the table that the same migration creates. Snapshot `0229` copied it forward. Any new migration therefore starts with two statements that do not belong to it: ```sql ALTER TABLE "issue_question_response_deliveries" ADD COLUMN "error_count" integer DEFAULT 0 NOT NULL; ALTER TABLE "decision_archive_notification_outbox" DROP COLUMN "error_count"; ``` The `ADD COLUMN` fails on a fresh database, because migration `0228` already creates that column. The `DROP COLUMN` targets a column that does not exist on any deployment. **Expected behavior** `drizzle-kit generate` reports "No schema changes, nothing to migrate" on a clean checkout of `master`, and a new migration contains only the author's own schema change. **Steps to reproduce** 1. Check out `master` at commit `bc1a21564`. 2. Run `pnpm install`. 3. Run `pnpm --filter @paperclipai/db generate`. 4. Read the emitted `packages/db/src/migrations/0230_*.sql`. It contains the two statements above, and no schema file was changed. **Paperclip version or commit** `master` at `bc1a21564`. The drift entered in #12307 (snapshot `0228`) and was carried forward by #12291 (snapshot `0229`), which worked around it by building its snapshot by hand. **Deployment mode** Not deployment specific. It affects anyone who generates a migration, and it affects any fresh database that would later run the bad migration. **Database mode** All PostgreSQL modes: embedded, local Docker, and hosted. ## What Changed - Moved the `error_count` column entry from `decision_archive_notification_outbox` to `issue_question_response_deliveries` in `packages/db/src/migrations/meta/0228_snapshot.json` and `packages/db/src/migrations/meta/0229_snapshot.json`. Both files keep their `id` and `prevId`, so the snapshot chain is unchanged. - Added `packages/db/src/migration-snapshot-drift.test.ts`. It reads the newest snapshot named by `_journal.json`, serializes the schema modules with `generateDrizzleJson`, and asserts that `generateMigration` returns no statements. This is the same diff that `generate` performs. - Documented the snapshot rule in `doc/DATABASE.md` under a new "Migration snapshots" section. No migration SQL was added, renumbered, or edited. No schema file changed. The database is correct as it is; only the snapshot was wrong. Why both snapshots and not only the newest one: `0229` is the file that `generate` reads, so repairing it is what fixes the bug. `0228` holds the same error, and `drizzle-kit drop` removes the last migration and its snapshot, which would promote `0228` back to newest and bring the drift back. Repairing both removes that trap. Snapshots are never applied to a database, so neither edit changes any deployment. ## Verification Commands run from the repository root. - `pnpm --filter @paperclipai/db generate` — "No schema changes, nothing to migrate 😴". It writes no SQL file, no snapshot, and no journal entry. `git status` stays clean. Before the fix, the same command wrote `0230_fast_caretaker.sql` with the two spurious statements. - The repaired `0229_snapshot.json` is byte-identical to the snapshot that a real `generate` run produced, except for the `id` and `prevId` that keep the chain intact. - Chain check: the repaired `0228` and `0229` snapshots now differ by exactly the two columns that `0229_drop_company_brand_color_and_attachment_max_bytes.sql` drops, `companies.brand_color` and `companies.attachment_max_bytes`, and by nothing else. - Database check: applied all 229 migrations in order to an embedded PostgreSQL, then compared the live schema with the repaired snapshot. 179 tables and 2687 columns match, with no missing column, no extra column, and no nullability difference. The same comparison against the pre-fix snapshot reports exactly two problems: `column only in snapshot: decision_archive_notification_outbox.error_count` and `column only in database: issue_question_response_deliveries.error_count`. This harness was a scratch script and is not part of the pull request. - `pnpm --filter @paperclipai/db typecheck` — pass. It runs `check:migrations`, which is `check-migration-numbering` and `check-migration-safety`. - `npx vitest run --root packages/db` — 28 files, 102 tests, all pass. This includes the new test. - New test, negative case: with the pre-fix `0229_snapshot.json` restored, `migration-snapshot-drift.test.ts` fails and prints exactly the two spurious statements, plus the instruction to run `generate`. It passes on the repaired snapshot. It takes about 1.2 seconds and needs no database. - `node scripts/check-forbidden-tokens.mjs` and `node scripts/check-no-git-push.mjs` — pass. ## Risks Low risk. A Drizzle snapshot is a build-time record for `drizzle-kit generate`. It is never applied to a database, so this change cannot alter any deployment, and no operator action is needed. Databases that already ran migrations `0228` and `0229` are correct today and stay correct. The proof is a clean `generate`: the command that produced the wrong migration now reports "No schema changes, nothing to migrate" and writes nothing. Two smaller notes: - The new test depends on `drizzle-kit/api`, which is already a dev dependency of `packages/db`. If a future `drizzle-kit` upgrade changes that surface, the test fails loudly at import rather than passing silently. - The test imports every module in `packages/db/src/schema/`, which is the same set that `drizzle.config.ts` points the CLI at. It deduplicates by object identity, because the barrel re-exports the same table objects and `drizzle-kit` rejects a table it sees twice. ## Model Used Claude (Anthropic), Claude Opus, agentic tool use via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.827.0-canary.6 |
||
|
|
bc1a21564f |
Remove the company brand color and per-company attachment limit (#12291)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - A company is the top-level container, and the company General page
holds its settings
> - Two of those settings did almost nothing: the brand color only
tinted the generated company icon, and the attachment size limit sat
under the deployment-level `PAPERCLIP_ATTACHMENT_MAX_BYTES` cap that
already bounded every upload
> - A setting that changes one icon hue, and a setting that can only
lower a limit the operator already set, are not worth the page space or
the code that carries them
> - This pull request deletes both settings from the UI, the validators,
the API contract, the server, and the database
> - With the deployment cap as the only limit left, the message a person
sees when an upload is rejected has to name that limit in terms they can
act on, so the raw byte count becomes a human-readable size
> - The benefit is a shorter company General page for every deployment,
one attachment limit instead of two, and less code between an upload and
its ceiling
## Linked Issues or Issue Description
No existing issue. The description below follows
`.github/ISSUE_TEMPLATE/enhancement.yml`.
**What existing behavior does this improve?**
The company General page (`/company/settings`), the `PATCH
/api/companies/{companyId}` and `PATCH
/api/companies/{companyId}/branding` request contracts, and the
attachment upload limit on task, case, and company-import uploads.
**Subsystem affected**
Cross-cutting: `ui/`, `server/`, `packages/shared`, `packages/db`.
**Current behavior**
The company General page shows an "Appearance" section with three
controls: Logo, Brand color, and Attachment size limit. The brand color
is a hex value that feeds one thing — the hue of the generated company
pattern icon. Companies that never set one already get a hue derived
from the company name. The attachment size limit is a per-company byte
count stored on `companies.attachment_max_bytes`. Every upload path
clamps it against the deployment-level `PAPERCLIP_ATTACHMENT_MAX_BYTES`
cap, so the per-company value can only lower a limit the operator
already chose.
**Proposed behavior**
The Appearance section keeps the Logo control only. The company pattern
icon always derives its hue from the company name. Every attachment path
reads the deployment cap directly, so `PAPERCLIP_ATTACHMENT_MAX_BYTES`
is the single limit. An upload rejected by that limit says so in human
units — "File is larger than the 10 MB limit" rather than a raw byte
count. The `companies.brand_color` and `companies.attachment_max_bytes`
columns are dropped, and both fields leave the company API contract.
**Reason and benefit**
Both settings ask an operator to make a decision that changes almost
nothing. The brand color moves one icon hue on a page that also lets you
upload a real logo, which overrides the icon entirely. The attachment
limit reads as a real control but cannot raise anything, so it is a
second place to look when an upload is rejected. Removing both shortens
the page every deployment sees, removes a company-scoped read from the
task attachment upload path, and leaves one attachment limit to reason
about instead of two.
**Breaking changes**
The company API responses no longer include `brandColor` or
`attachmentMaxBytes`, and `GET /api/invites/{token}` no longer includes
`companyBrandColor`. `PATCH /api/companies/{companyId}/branding` is
strict, so a request that sends `brandColor` now returns 400; the
non-strict `PATCH /api/companies/{companyId}` schema strips it. Company
packages exported by older versions still import: the portability
company manifest schema is non-strict, so the retired keys are stripped
and ignored rather than rejected. Companies that stored a brand color
lose it — their icon reverts to the name-derived hue that every company
without a color already used.
## What Changed
- Removed the "Brand color" and "Attachment size limit" fields from the
company General page, along with their state, dirty checks, save
payload, and Save-button gating.
- Removed `brandColor` and `attachmentMaxBytes` from
`createCompanySchema`, `updateCompanySchema`, and
`updateCompanyBrandingSchema`, and deleted the now-orphaned
`DEFAULT_COMPANY_ATTACHMENT_MAX_BYTES` and
`MAX_COMPANY_ATTACHMENT_MAX_BYTES` constants.
- Removed both fields from the `Company` type, the portability manifest
type and schema, and the `companiesApi.update` payload allowlist.
- Dropped `brandColor` from `CompanyPatternIcon` and its callers, so the
icon hue always comes from the company name. Deleted the now-unused
`hexToHue` helper and the now-unused `pickTextColorForSolidBg` export.
- Stopped emitting `brandColor` from the company service selection and
from the invite-summary and invite-branding payloads in
`server/src/routes/access.ts`.
- Replaced `normalizeIssueAttachmentMaxBytes` with the deployment cap:
task attachments, case attachments, and company import now use
`MAX_ATTACHMENT_BYTES` directly. The helper is deleted.
- Added `formatAttachmentSize()` next to `MAX_ATTACHMENT_BYTES` and
routed every over-limit message through it, so a rejected upload names
the limit in human units instead of raw bytes: `Image exceeds 10485760
bytes` becomes `Image is larger than the 10 MB limit`. Enforcement is
unchanged — the same single cap, the same multer limits, the same status
codes and response shapes.
- Added migration
`0229_drop_company_brand_color_and_attachment_max_bytes.sql` and removed
both columns from the Drizzle `companies` schema.
- Kept legacy imports working: the portability company manifest schema
is non-strict, so older packages carrying the retired keys still import
with the keys ignored.
- Updated the skill API reference and the implementation spec, and
pruned the token-extraction allowlist entries that the removed code made
stale.
## Verification
Commands run from the repository root:
- `pnpm --filter @paperclipai/shared typecheck` — pass
- `pnpm --filter @paperclipai/db typecheck` — pass (includes
`check:migrations`, which validates the new migration number and journal
entry)
- `pnpm --filter @paperclipai/ui typecheck` — pass
- server typecheck via `node_modules/.bin/tsc --noEmit` in `server/` —
pass. `pnpm --filter @paperclipai/server typecheck` could not run
locally because it builds the Rust runner first and `cargo` is not
installed on this machine; the TypeScript step it wraps is the command
above.
- `npx vitest run packages/shared/src/validators/company.test.ts` — 6
passed
- `npx vitest run server/src/__tests__/company-portability.test.ts` — 90
passed
- `npx vitest run server/src/__tests__/attachment-types.test.ts
server/src/__tests__/assets.test.ts
server/src/__tests__/issue-attachment-routes.test.ts
server/src/__tests__/company-portability.test.ts
server/src/__tests__/cases-routes.test.ts` — 165 passed (the
human-readable limit messages)
- `npx vitest run server/src/__tests__/company-branding-route.test.ts
server/src/__tests__/issue-attachment-routes.test.ts
server/src/__tests__/invite-summary-route.test.ts
server/src/__tests__/openclaw-invite-prompt-route.test.ts
server/src/__tests__/companies-route-cross-company-authz.test.ts` — all
passed
- `npx vitest run cli/src/__tests__/company.test.ts
cli/src/__tests__/company-delete.test.ts` — 27 passed
- `npx vitest run` in `ui/` — 4425 passed, 1 pre-existing failure
unrelated to this change (`OnboardingWizard.test.tsx` "renders instead
of throwing when the browser denies storage access", which also fails on
`master`)
- `npx vitest run` in `server/` — see the note below
- `node scripts/check-token-gates.mjs` — no new violations; the only
reported violations are the pre-existing `PillGuy.tsx` ones present on
`master`
New tests added:
- `packages/shared/src/validators/company.test.ts` — the create and
update schemas strip the retired keys, the strict branding schema
rejects `brandColor`, and the portability manifest schema accepts a
legacy entry carrying both keys and drops them.
- `server/src/__tests__/company-branding-route.test.ts` — `PATCH
/api/companies/{companyId}/branding` returns 400 for `brandColor` and
does not call the company service.
- `server/src/__tests__/company-portability.test.ts` — a legacy package
that declares `brandColor` and `attachmentMaxBytes` imports
successfully, and neither key reaches `companies.create`.
- `server/src/__tests__/issue-attachment-routes.test.ts` — the effective
task attachment limit is the deployment cap, and the route no longer
loads the company to size an upload.
- `server/src/__tests__/attachment-types.test.ts` —
`formatAttachmentSize()` renders the default cap as `10 MB`, keeps one
decimal place for fractional sizes and drops a trailing `.0`, falls back
to KB and bytes for small caps, steps up to GB, and never emits `NaN`
for a degenerate input.
- `server/src/__tests__/assets.test.ts` — the asset-image and
company-logo routes both return the human-readable limit message on an
over-cap upload.
## Merge with master
`master` moved while this was open, and the merge needed two
resolutions:
- **`ui/src/pages/CompanySettings.tsx`.** #12243 reworded the
user-facing
copy from "company" to "organization", and that rewording landed inside
the "Brand color" and "Attachment size limit" hints — the two fields
this change deletes. Both fields are removed, so the conflicted block is
dropped whole. The Logo field and every other copy change from #12243
are
kept.
- **Migration renumbered 0228 -> 0229.** #12307 landed
`0228_nasty_grim_reaper`, so this migration is now
`0229_drop_company_brand_color_and_attachment_max_bytes`. Its snapshot
is
rebuilt from master's `0228_snapshot.json` with only the two `companies`
columns removed, and `meta/_journal.json` is master's journal plus a
single `idx: 229` entry. `pnpm --filter @paperclipai/db
check:migrations`
passes.
The snapshot was rebuilt by hand rather than taken from `drizzle-kit
generate`, because master's `0228_snapshot.json` has drifted from
master's
own schema: `issue_question_response_deliveries.error_count` is created
by
master's 0228 SQL but missing from its snapshot, and the snapshot still
carries `decision_archive_notification_outbox.error_count`. Regenerating
folds both into this migration, and the resulting `ADD COLUMN
error_count`
would fail on a fresh database where master's 0228 already created that
column. Rebuilding from master's snapshot leaves that drift exactly
where
it is and keeps this migration to the two column drops. The drift is
pre-existing on master and is not addressed here.
## Risks
- **The migration is a destructive column drop.**
`0229_drop_company_brand_color_and_attachment_max_bytes.sql` removes
`companies.brand_color` and `companies.attachment_max_bytes`. It is safe
because both features are removed in the same change and nothing reads
either column after it. The statements use `DROP COLUMN IF EXISTS`,
matching the convention of the recent drop migrations in this
repository. The drop is not reversible: a downgrade after this migration
loses any stored values.
- **Stored brand colors are lost.** A company that had set a color now
renders the name-derived icon hue that every company without a color
already used. No other surface changes, and an uploaded logo still
overrides the icon.
- **API response shape narrows.** `brandColor` and `attachmentMaxBytes`
leave the company payloads, and `companyBrandColor` leaves the invite
summary payload. A client reading those fields now sees `undefined`. The
bundled UI and CLI are updated in this change.
- **Legacy imports are covered.** Packages exported by older versions
still carry both keys. The manifest schema is non-strict, so the keys
are stripped rather than rejected, and a test locks that in.
- **The over-limit message strings changed.** Anything matching on the
old `... exceeds N bytes` text — a test, a script, or a client that
string-matches `body.error` — needs updating. The status codes (422) and
response shapes are unchanged, so structured clients are unaffected.
- **Attachment limits can only widen.** A deployment that had lowered a
company below the deployment cap now allows uploads up to the cap for
that company. Lower `PAPERCLIP_ATTACHMENT_MAX_BYTES` if a smaller
ceiling is needed.
- **Storybook visual baselines shift** for the `CompanyPatternIcon`
matrix story, because those fixtures had brand colors. That workflow
runs only on a PR labeled `storybook-visual`, so it does not gate this
PR; regenerate the baselines if the label is added.
## Model Used
Claude (Anthropic), Claude Opus, agentic tool use via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
canary/v2026.827.0-canary.5
|
||
|
|
7b91fe9ea7 |
Hide host-path and execution-engine surfaces in managed-sandbox-only mode (#12293)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - An instance can turn on the `enableManagedSandboxOnly` feature, which hides the local environment and runs every agent in the platform-managed environment > - That feature already gated the environment pickers, the onboarding wizard, and the server-side run selection, but many other screens still showed absolute paths on the execution host and still let the user pick an execution engine > - On such an instance those controls name a filesystem the user cannot reach; a path written there is stored and then ignored, which reads as a broken control > - This pull request hides the remaining host-path and execution-engine surfaces behind the same feature, adds a server rule that refuses a project-workspace path write while the feature is on, and closes a related route gap in the isolated-workspace pages > - The benefit is that a managed instance shows no host path and no folder picker anywhere, and a write that carries a path now fails with a clear message instead of being silently discarded ## Linked Issues or Issue Description No public issue exists. The description below follows `.github/ISSUE_TEMPLATE/enhancement.yml`. **What existing behavior does this improve?** The `enableManagedSandboxOnly` instance feature, and the UI surfaces that show a host filesystem path: project properties, the new-project dialog, the project workspace and execution workspace detail pages, the workspace and task cards, plugin local folders, and the agent configuration form with its per-adapter fields. It also improves route gating for `enableIsolatedWorkspaces`. **Subsystem affected** Cross-cutting (`ui/` and `server/`). **Current behavior** When `enableManagedSandboxOnly` is on, the local environment disappears from the environment pickers and the server refuses to run an agent on the local host. Everything else stays visible. A user still sees: - the project "Local folder" row, its absolute path, and the Set/Change/Clear buttons - the "Local folder" field and its "Choose" folder picker in the new-project dialog - the "Local path" field and fact row on a project workspace - the "Paths" and "Lifecycle commands" groups on an execution workspace - the working directory on workspace cards, task properties, and runtime service rows - the plugin "Local folders" section - "Working directory (deprecated)", "Command", "Execution engine", "ACP server command", "ACP state directory", and "Agent instructions file" in the agent configuration form A path typed into any of these names a filesystem no agent on the instance uses. The project workspace API also accepts a `cwd` write and stores it. Separately, `/workspaces`, `/execution-workspaces/*`, and `/projects/:projectId/workspaces/:workspaceId` render for anyone who types or bookmarks the URL, even with `enableIsolatedWorkspaces` off. Only the sidebar entry reads that flag. **Proposed behavior** With `enableManagedSandboxOnly` on, none of those surfaces render. A project whose codebase came from a managed checkout keeps its one-line "Paperclip-managed folder." label and shows no path. The non-path controls stay: repo URL, branch, service URL, port, command output, ACP session mode, ACP non-interactive permissions, Codex fast mode, and the sandbox toggles. The project-workspace create and patch routes, and the nested workspace on project create, answer `422` with "This instance runs agents only in the platform-managed environment; local folders are not configurable." when the payload carries a non-null `cwd`. A `cwd: null` write still passes, so an instance that just turned the feature on can clear a stale path. With `enableIsolatedWorkspaces` off, the three workspace route groups redirect to the dashboard. **Reason and benefit** A control that cannot do anything is worse than a missing control: the user fills it in, saves, and gets no error and no effect. The server rule turns that silent no-op into a clear refusal. The route gate stops a feature that an instance has turned off from staying reachable by URL, which is the same standard the Cases, Pipelines, and hidden-settings pages already meet. **Breaking changes** None for a default instance: both flags are off by default for self-hosted and managed instances, so nothing changes unless an operator turns them on. Stored `adapterConfig` values are never cleared, so turning the feature off restores every previous value. ## What Changed - Add `ui/src/hooks/useManagedSandboxOnly.ts`, modelled on `useAppsEnabled`, for components that do not already read the experimental settings. It exposes `hideHostPaths`, which fails closed while the settings query is in flight, so a cold cache never flashes a host path before the policy resolves. Components that keep their own settings read compute the same gate from `isFetched`. - Add `managedSandboxOnly` to `AdapterConfigFieldsProps` and populate it where `AgentConfigForm` builds the adapter field props. Resolve the effective instructions-file gate once as `hideInstructionsFile || hideHostPaths`, so every adapter hides that path field with no per-adapter edit. - Hide under the flag: the project "Local folder" block and its absolute-path edit panel (a managed checkout keeps its label, without the path); the new-project "Local folder" field; the project-workspace "Local path" field and fact row; the execution-workspace "Paths" and "Lifecycle commands" groups; the working directory on the workspace summary card, the task workspace card, the task properties "Folder" row, and the runtime service rows; the plugin "Local folders" section; "Working directory (deprecated)" and "Command" in the agent form; and the per-adapter "Execution engine", "ACP server command", and "ACP state directory" for `claude_local`, `codex_local`, and `gemini_local`. - Drop two working-directory fallbacks that had no gate to read: the close-workspace dialog now falls back to "No additional details", and the reuse-existing workspace label and picker subtitle fall back to a neutral phrase. - Refuse a non-null `cwd` with `422` on `POST /projects/:id/workspaces`, `PATCH /projects/:id/workspaces/:workspaceId`, and the nested workspace on `POST /companies/:companyId/projects`, following the `assertNoAgentHostWorkspaceCommandMutation` precedent on those routes. - Add `IsolatedWorkspacesRouteGate` and wrap the `/workspaces`, `/execution-workspaces/*`, and `/projects/:projectId/workspaces/:workspaceId` routes with it. - Leave the SSH "Remote workspace path" and the workspace file browser alone, with a comment explaining why. ## Verification Automated: - `pnpm --filter @paperclipai/ui exec vitest run` — 479 of 480 files pass (4455 of 4456 tests). The one failure is `OnboardingWizard.test.tsx > renders instead of throwing when the browser denies storage access`, which also fails on `origin/master` and is unrelated to this change. - `pnpm --filter @paperclipai/server exec vitest run project workspace instance-settings` — 45 of 50 files pass. Four files fail on macOS for reasons unrelated to this change: `workspace-instance-cleanup`, `workspace-runtime`, `execution-workspace-runtime-control-conflict`, and `workspace-runtime-exposure` compare `/var/...` against the resolved `/private/var/...` or bind real ports. The same files fail on a clean `master` checkout on the same machine. - `pnpm --filter @paperclipai/ui typecheck` - `tsc --noEmit` in `server/` (after `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps`). The package `typecheck` script also builds the Rust runner, which needs `cargo`; it is not installed on the machine that ran this. New and extended tests: - `ui/src/adapters/managed-sandbox-only-config-fields.test.tsx` — the three adapters drop the execution engine, the ACP paths, the instructions-file path, and every "Choose" button when the flag is on, and keep the non-path controls. - `ui/src/components/AgentConfigForm.render.test.tsx` — flag-on and flag-off renders for the working directory, the command, the engine, the ACP paths, and the resolved adapter field props. - `ui/src/components/ProjectProperties.managed-sandbox.test.tsx`, `ui/src/components/NewProjectDialog.managed-sandbox.test.tsx`, `ui/src/pages/ProjectWorkspaceDetail.test.tsx`, `ui/src/components/ProjectWorkspaceSummaryCard.test.tsx`, `ui/src/components/WorkspaceRuntimeControls.test.tsx`. - `ui/src/components/IsolatedWorkspacesRouteGate.test.tsx` — redirect when off, render when on, and render nothing while the flag query is in flight. - "Still loading" cases for the project properties, the new-project dialog, the workspace summary card, the runtime service rows, and the agent configuration form, each asserting that no host path renders before the policy resolves. - `server/src/__tests__/project-workspace-managed-sandbox-routes.test.ts` — the `422` on all three write paths, the `cwd: null` pass-through, and the flag-off pass-through. Manual check to reproduce: turn on Managed Environment Only in instance experimental settings, then open a project, the new-project dialog, an agent's configuration, and a workspace page. No path, folder icon, or "Choose" button appears. Turn the setting off and each control returns with its stored value. No documentation change was needed. The operator-facing text for both settings lives in the feature catalog entry, which already states the contract this pull request now enforces across the UI. ## Risks - Low. Both flags default to off, so a default instance is unchanged. - The hidden fields are presentation only. No stored `adapterConfig` value is cleared, because an import carries adapter configuration written on another instance and clearing it would break that flow. Turning the setting off shows every previous value again. - The `422` is the one behavior change for an API caller, and only while the setting is on. `cwd: null` still passes so a stale path can be cleared. - The route gate renders nothing until the flag query settles, so an instance with isolated workspaces on never flashes a redirect. An instance with the feature off now redirects a bookmarked workspace URL to the dashboard. - Every host-path guard fails closed while the settings query is in flight, so a default instance shows those controls a moment later than before on a cold load. That is the safe direction: the alternative flashes a path a managed instance must never show. - Two path surfaces stay on purpose, each with a comment: the SSH "Remote workspace path" is a path on the user's own remote host, and the workspace file browser shows workspace-relative paths. The instance Adapters page also keeps its "Local path" install option, since that page is an instance-admin surface the hosting operator can already hide through the hidden-settings mechanism. ## Model Used Claude (Anthropic), Claude Opus, 1M context window, extended thinking, agentic tool use through Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.827.0-canary.4 |