mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
ff5fd62d077fa4a75bb14fd01b85041ff1c85e7e
3567
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ff5fd62d07 |
Resolve the environment secret companyId context on first save (#11291)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments give agent runs an execution target, and their env vars can carry secrets, so writes need a company scope for secret bindings > - `PATCH /environments/:id` resolves that scope from a query param, existing bindings, or a single-membership actor — and fails closed otherwise > - A fresh environment has no bindings yet, so an admin whose memberships cannot pin one company gets a 422 on the very first env var save and can never create the first binding > - #11200 made this reachable from the UI: it opened the envVars-only edit on the managed sandbox row, but the UI never sends the company it already knows > - This pull request sends the page's selected company on environment updates and adds a server fallback to the instance's only company > - The benefit is that the first env var save works on fresh environments, while multi-company ambiguity still fails closed ## Linked Issues or Issue Description Refs #11200 **What happened?** On a fresh platform-managed sandbox environment, the first "Save environment variables" in the UI fails with HTTP 422: "Environment secret management requires a companyId context during the instance-scoped transition." The same failure hits any environment update that touches `envVars` or `config` when the environment has no secret bindings yet and the actor's memberships do not name exactly one company (for example an instance admin provisioned without a membership row, or a user in two companies). **Expected behavior** The save succeeds. The client tells the server which company scopes the secret bindings, and on a single-company instance the server can resolve the only possible scope by itself. **Steps to reproduce** 1. Provision a platform-managed sandbox environment (managed-config `environments` entry) so the row exists with no secret bindings. 2. Sign in as an instance admin whose memberships do not resolve to exactly one company. 3. Open Environments, edit the managed row, add an env var, and save. 4. The PATCH returns 422 with the companyId-context error. ## What Changed - `ui/src/api/environments.ts`: `environmentsApi.update` accepts an optional `companyId` and sends it as the `companyId` query param the server already reads. Renamed the `customImageCompanyQuery` helper to `companyIdQuery` since it now serves plain updates and probes too. - `ui/src/pages/CompanyEnvironments.tsx`: both the managed envVars-only save and the general environment save pass `selectedCompanyId`. - `server/src/routes/environments.ts`: `resolveEnvironmentSecretContextCompanyId` falls back to the instance's only company when exactly one exists, mirroring the existing fallback in `resolveCustomImageCompanyId`. Multi-company instances still fail closed. - Tests: three new route tests (explicit query context for a multi-company actor, single-company-instance fallback for a membership-less admin, fail-closed regression on a multi-company instance), and updated the SSH-probe and probe-ambiguity tests for the new fallback. ## Verification - `cd server && pnpm vitest run src/__tests__/environment-routes.test.ts` — 78 pass. - `cd ui && pnpm vitest run src/pages/CompanyEnvironments.test.tsx` — 22 pass. - Full `server` and `ui` vitest suites and `tsc --noEmit` on both packages pass locally. - Manual: on a managed instance, edit the managed sandbox environment, add an env var, save. Before: 422. After: the save persists and the PATCH carries `?companyId=`. ## Risks - Low risk. The explicit query param path already existed on the server; the UI now uses it. - Behavioral shift: on single-company instances, environment secret-context resolution (including `POST /environments/:id/probe`) now resolves the only company instead of returning no context. That lets probes resolve secret-backed config where they previously returned a 422 asking for an explicit companyId. Multi-company instances keep the fail-closed behavior, covered by a regression test. - Self-hosted single-company instances gain the same first-save fix; no schema or config changes. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f1931d0e14 |
test(server): fix flaky workspace-busy retry-row read race (#11293)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server heartbeat system records each agent run and its retry state > - A workspace-busy deferral cancels one run before it inserts the scheduled retry row > - The test helper can return after the cancel write and before the retry-row insert > - A direct read can then return no row and fail a valid retry assertion > - This pull request makes presence reads wait for the retry row > - The benefit is stable test coverage without a production behavior change ## Linked Issues or Issue Description Refs: #10806 **What happened?** The workspace-busy test read the retry row after the first deferral write. The helper returned before the scheduled-retry insert completed. The read then returned no row and failed the retry assertions. **Expected behavior** The test must wait until the scheduled-retry row exists before it checks retry-row fields. The production write order must stay unchanged. **Steps to reproduce** 1. Add a 300 ms delay between the deferral writes. 2. Run `server/src/__tests__/heartbeat-workspace-busy.test.ts`. 3. Observe failures at retry-row presence checks. 4. Add the bounded polling helper. 5. Run the test file again and observe that all presence checks pass. **Paperclip version or commit** Commit `d9b6e8a6e62b9b56919fc9c52d294e8ac569f70f`. **Deployment mode** Local test run from source. ## What Changed - Add `waitForRetryRun`, which polls for the retry row with a 10 second timeout and a 50 millisecond interval. - Use the helper at every test site that reads a retry row after deferral. - Keep direct reads at absence assertions. - Keep production code unchanged. ## Verification - Injected a temporary 300 millisecond delay between the two production writes and reproduced the five presence-site failures. - Applied the helper with the delay and passed the test file 15 out of 15 times. - Removed the temporary production delay. - Ran the changed test file 25 consecutive times with 0 failures. - Ran TypeScript checks for the changed test file with no errors. ## Risks Low risk. This pull request changes test code only. The helper has a bounded timeout. Production behavior and retry-row assertions remain unchanged. ## Model Used OpenAI Codex, GPT-5, reasoning mode, tool use, and code execution. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c57c0f7498 |
fix(sandbox-providers): accept bsdtar listings in the syncOut tarball confinement check (#11289)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers (Daytona, Kubernetes) sync run results back to the host with a sandbox-authored tarball > - Before extraction, a confinement check parses the host `tar -tvf` listing and fails closed on unparseable lines > - The parser only understands the GNU tar listing dialect; macOS ships bsdtar, whose ls-style listing never matches > - Every sandbox syncOut on a macOS host therefore aborts with "refusing tarball with an unparseable entry listing", and the run fails at copy-back > - This pull request teaches the parser both dialects while keeping the fail-closed and traversal guarantees > - The benefit is that Daytona and Kubernetes sandbox runs work on macOS hosts, with no behavior change on Linux ## Linked Issues or Issue Description No existing issue. Bug description: **What happened?** On a macOS host, every Daytona sandbox run fails at syncOut. The adapter reports: `Daytona syncOut refusing tarball with an unparseable entry listing: -rw-r--r-- 0 daytona daytona 7560 Aug 11 21:43 AGENTS.md`. The Kubernetes provider has the same parser and fails the same way. **Expected behavior** The confinement check accepts a well-formed listing from the host tar, whichever dialect the host tar emits. It still rejects members that escape the extraction directory, and it still fails closed on lines it cannot parse. **Steps to reproduce** 1. Run Paperclip on macOS (system tar is bsdtar). 2. Configure an agent with the Daytona sandbox provider. 3. Trigger any run that syncs files back from the sandbox. 4. The run fails at syncOut with the unparseable-entry-listing error, because bsdtar prints `<perms> <links> <user> <group> <size> <Mon> <day> <time|year> <name>` while the parser expects the GNU `<perms> <owner>/<group> <size> <date> <time> <name>` shape. **Operating system** macOS (bsdtar 3.5.3). Linux hosts with GNU tar are unaffected. ## What Changed - Extracted the listing-line parse in both providers' `file-sync.ts` into an exported `parseTarVerboseListingLine` that accepts the GNU/busybox dialect and the bsdtar (libarchive) dialect. - The GNU shape now requires the slash-joined `<owner>/<group>` field. This keeps the two shapes mutually exclusive. Without it, a bsdtar line with numeric uid/gid satisfies the loose GNU pattern shifted by one field, which would hide a leading `../` from the traversal check. - Unparseable lines still fail closed. This includes device-node entries, whose size column is `major,minor` in both dialects. - Made the path-traversal fixture in the Daytona suite portable: GNU spells member renaming `--transform`, bsdtar spells it `-s`. - Added a Daytona test that refuses a sandbox-authored tarball carrying a symlink whose target escapes the extraction dir. - Added parser unit tests for both dialects (file, dir, symlink, hardlink, numeric owner, year-form dates, fail-closed lines) to both providers' suites. ## Verification - `pnpm test` in `packages/plugins/sandbox-providers/daytona`: 136/136 pass on a macOS host. On unpatched `master` the round-trip test fails there with the unparseable-entry-listing error. - `pnpm test` in `packages/plugins/sandbox-providers/kubernetes`: the new parser tests pass; no new failures against the `master` baseline on the same host. - `pnpm typecheck` passes in both packages. - CI runs the same suites on Linux/GNU tar and proves the GNU path is unchanged. ## Risks - Low risk. The GNU pattern is one token stricter (`<owner>/<group>` must contain `/`). GNU and busybox tar always print the slash-joined owner field, so accepted GNU listings are unchanged. - The bsdtar branch only widens acceptance on hosts that were failing 100% of syncOuts before, so no working deployment changes behavior. - The check still fails closed on anything neither pattern matches. ## Model Used Claude Fable 5 (`claude-fable-5`, Claude Code CLI, extended thinking + tool use). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
01112c350c |
fix(ui): always give respondWith a real Response in the sw fetch fallback (#11292)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The UI registers a service worker (`ui/public/sw.js`) with a
network-first fetch handler whose cache is an offline fallback
> - The fallback hands `event.respondWith` the result of
`caches.match(...)`, which resolves `undefined` on a cache miss — and in
the navigation branch, `caches.match("/") || offlineResponse` never uses
the fallback because `caches.match` returns a promise, which is always
truthy
> - When the network fetch rejects (server restart, deploy, brief
outage) and the cache misses, the browser fails the request with
`Uncaught (in promise) TypeError: Failed to convert value to
'Response'`, so navigation breaks outright instead of degrading to the
offline page
> - This pull request awaits the cache lookups and guarantees a real
`Response` on every path
> - The benefit is that brief server unavailability degrades to the
offline fallback instead of a dead navigation
## Linked Issues or Issue Description
No existing issue. Description follows the bug template:
**What happened?**
Navigating while the server was briefly unavailable (mid-restart)
produced `The FetchEvent for "…" resulted in a network error response:
the promise was rejected.` and `sw.js:1 Uncaught (in promise) TypeError:
Failed to convert value to 'Response'.` The navigation failed instead of
showing the offline fallback.
**Expected behavior**
A failed navigation serves the cached app shell when present, otherwise
the "Offline" 503 response. A failed asset fetch serves its cache entry
when present, otherwise a proper network-error response. `respondWith`
always receives a real `Response`.
**Steps to reproduce**
1. Load the app so `sw.js` is active; ensure `/` is not in the service
worker cache (fresh cache version).
2. Restart or stop the backend.
3. Navigate to any page: the fetch rejects, `caches.match` misses, and
the browser logs the conversion TypeError with a failed navigation.
## What Changed
- `ui/public/sw.js`: the fetch fallback awaits `caches.match(...)` and
returns the "Offline" 503 for navigations and `Response.error()` for
assets when the cache misses.
- `ui/src/lib/sw-offline-fallback.test.ts`: evaluates the real `sw.js`
in a sandboxed scope and covers the three fallback paths; the two
miss-path tests fail against the previous code.
## Verification
- `pnpm vitest run src/lib/sw-offline-fallback.test.ts` in `ui/` — 3
tests pass.
- Verified both miss-path tests fail against the unmodified `sw.js`.
## Risks
Low risk. The change only affects the fetch-rejection path; successful
fetches and cache hits behave exactly as before. `Response.error()`
mirrors what the browser would produce for an unhandled failed no-cors
fetch.
## Model Used
- Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code
CLI with tool use (code search, edit, test execution).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
2c53437fc9 |
fix(server): authenticate cloud-proxied browsers on the live-events websocket (#11290)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The UI receives live run/issue events over a websocket at `/api/companies/:id/events/ws`; the server authorizes upgrades with a bearer token or a Better Auth session > - On a cloud-managed deployment, browsers authenticate through trusted `x-paperclip-cloud-*` headers injected by the managing front door — they never hold a local Better Auth session, and the Express middleware lane that understands those headers is not consulted for websocket upgrades > - Every browser websocket upgrade behind the front door therefore resolves no identity and is rejected 403: the live-events socket has never connected on a managed instance, leaving permanent reconnect churn and console failure noise while the UI silently degrades to polling > - This pull request adds a cloud-actor lane to the upgrade authorization, reusing the same trusted-header resolver the HTTP middleware uses > - The benefit is working realtime updates on managed instances, an end to the reconnect churn, and unchanged self-hosted behavior ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** On a cloud-managed instance, the browser console shows `WebSocket connection to 'wss://…/api/companies/<id>/events/ws' failed:` repeating indefinitely for every company, on a healthy instance. The server rejects each upgrade with 403 because `authorizeUpgrade` in `server/src/realtime/live-events-ws.ts` only knows bearer tokens and Better Auth sessions, while cloud-proxied browsers authenticate via `x-paperclip-cloud-*` trusted headers (handled only by the Express `actorMiddleware` lane in `server/src/middleware/auth.ts`). **Expected behavior** A browser that authenticates through the trusted cloud headers can open the live-events websocket for any company in its membership scope, exactly as it can call the HTTP API for those companies. **Steps to reproduce** 1. Run Paperclip in `authenticated` mode behind a proxy that injects the `x-paperclip-cloud-*` headers with a valid `PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN`. 2. Load any company page in a browser (no local Better Auth session). 3. HTTP API calls succeed; every `/events/ws` upgrade is rejected 403 and the UI retries forever. ## What Changed - `server/src/middleware/auth.ts`: `resolveCloudTenantActor` now accepts a minimal `CloudActorHeaderSource` (`header(name)`) instead of an Express `Request` — `Request` satisfies it unchanged — plus `cloudActorHeaderSourceFromHeaders` to adapt raw `IncomingMessage.headers`. - `server/src/realtime/live-events-ws.ts`: `authorizeUpgrade` gains an injected `resolveCloudActor` lane, tried before the Better Auth session fallback in `authenticated` mode. A resolved cloud actor is authoritative: the upgrade is authorized only for a company in the actor's membership scope (`companyIds`, the same scope the HTTP lane grants). Absent/unresolvable cloud headers fall through to the session path. - `server/src/index.ts`: wires `resolveCloudActor` through `resolveCloudTenantActor` + the header shim. The resolver self-gates: without `PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN` and a matching trust token it returns null, so self-hosted deployments never take this path. - Tests: upgrade authorized for an in-scope company (session resolver not consulted), rejected for an out-of-scope company, fall-through to session auth when no cloud actor resolves; header-shim resolution from a raw lowercased header map including `string[]` values. ## Verification - `pnpm vitest run server/src/__tests__/live-events-ws.test.ts server/src/middleware/cloud-tenant-actor.test.ts` — 25 tests pass. - `pnpm typecheck` in `server/` — clean. - Not verified live end-to-end: that requires a managed instance running this build; the direct probe evidence (HTTP authenticated fine, every WS upgrade 403) matches the code path exactly. ## Risks Low risk. The new lane only activates when the deployment configures the cloud trust token and the request presents it; both checks already protect the HTTP lane. Authorization scope is the same `companyIds` set the HTTP middleware computes (primary stack company plus the user's real membership rows). The cloud resolver's user/company materialization writes are debounced (existing behavior shared with the HTTP lane), so websocket reconnect storms do not amplify database writes. Self-hosted instances see no behavioral change, covered by the fall-through test. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution; diagnosis included live websocket handshake probes against a managed instance). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.812.0-canary.9 |
||
|
|
6a5b293240 |
test(server): fix onboarding first-task teardown foreign-key race (#11284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server coordinates issue work and heartbeat runs. > - The onboarding first-task route sends an assignment wake in the background. > - The route test removes related database rows during teardown. > - A late heartbeat run can keep foreign-key child rows alive during teardown. > - This pull request drains the wake and deletes run rows in foreign-key order. > - The benefit is a stable test that keeps the onboarding behavior unchanged. ## Linked Issues or Issue Description **What happened?** The onboarding first-task route sent a background assignment wake. The test teardown removed parent rows before the wake-created heartbeat rows finished. **Expected behavior** The test teardown should wait for the background wake and remove heartbeat rows before it removes their parent rows. **Steps to reproduce** 1. Run the onboarding first-task route test. 2. Repeat the test many times. 3. Observe an intermittent foreign-key error during teardown. **Paperclip version or commit** Commit `c30fe965920eeb7e7fb88e17574a65bed8fc01a4`. **Deployment mode** Local dev (pnpm dev). **Installation method** Built from source (pnpm dev / pnpm build). **Agent adapter(s) involved** Not adapter-specific (core bug). **Database mode** Embedded PGlite (default — DATABASE_URL unset). ## What Changed - Stub the server adapter in the route test so the dispatched run finishes at once. - Drain heartbeat runs to quiescence before teardown. - Delete heartbeat runs and child rows before their parent rows. - Delete runtime state and company skill rows in foreign-key order. - Keep the route behavior and all three test assertions unchanged. ## Verification - Run `pnpm exec vitest run src/__tests__/issue-onboarding-first-task-routes.test.ts` from the `server` package. - The author ran the suite 25 times with 25 passes. - The suite reproduced the teardown foreign-key error before this change. ## Risks Low risk. This change affects one test file and does not change product code or route behavior. ## Model Used OpenAI GPT-5. The model used tool calls and code review assistance. The exact context window and reasoning mode were not exposed in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.812.0-canary.8 |
||
|
|
f9bd0438e1 |
fix(server): stop terminal workspace reaper starving on oldest candidates (#11238)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server runs a scheduled reaper that archives terminal workspaces after it checks their state. > - The reaper reads candidates in `updatedAt` order and skips candidates that do not qualify for archive. > - The fixed page kept the same skipped candidates at the front, so the reaper did not inspect later eligible workspaces. > - This pull request adds a keyset cursor and a throttled log for sweeps that archive no workspace. > - The benefit is that the reaper inspects all candidates over time and reports an inert sweep. ## Linked Issues or Issue Description **What happened?** The terminal workspace reaper inspected a fixed page of old candidates. Ineligible candidates stayed in that page, so the reaper skipped later eligible workspaces on every sweep. **Expected behavior** The reaper must inspect each candidate over time and archive every eligible terminal workspace. **Steps to reproduce** 1. Create more than 50 terminal workspace candidates. 2. Keep the oldest page ineligible for archive. 3. Place an eligible workspace after that page. 4. Run repeated reaper sweeps. 5. Observe that the later eligible workspace remains unarchived. **Paperclip version or commit** Commit `3efdf555e6e14a46747c796c3c554438bfc03261`. **Deployment mode** Built from source with the server test suite. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific (core bug). **Database mode** Not database-related. ## What Changed - Add a keyset cursor that uses `(updatedAt, id)` order across reaper pages. - Reset the cursor at the end of the candidate set so the next sweep starts at the beginning. - Add a throttled log when a sweep inspects candidates but archives none. - Add regression tests for archive delivery and starvation. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts` — 45 tests pass. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/server-startup-feedback-export.test.ts` — 16 tests pass. - `pnpm --filter @paperclipai/server typecheck` — clean. ## Risks Low risk. The change affects only candidate paging and the related reaper log. The cursor resets after the candidate set, so the sweep remains periodic. ## Model Used Codex, OpenAI GPT-5, extended reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (no exact duplicate found; related scheduler PR #10911 is distinct) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no documentation applies; this is an internal reaper behavior change) - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.812.0-canary.7 |
||
|
|
1a377424db |
fix(ui): add mobile blocker actions (#11282)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators use task properties to inspect and change task relationships. > - A blocked-by chip linked directly to the blocking task. > - Its remove control appeared only on hover, so touch users could not reach it. > - This pull request opens a small action menu when a user taps the chip on mobile. > - The menu lets the user visit the task or start the existing blocker-removal confirmation. > - The benefit is that touch users can manage blockers without changing the fast desktop flow. ## Linked Issues or Issue Description **What happened?** On a phone-width layout, a tap on a blocked-by chip opened the blocking task immediately. The remove control appeared only on hover, so a touch user could not remove the blocker. **Expected behavior** A tap on a blocked-by chip on mobile opens a menu. The menu offers `Visit task` and `Remove blocker` actions. **Steps to reproduce** 1. Open a task that has a blocker. 2. Use a viewport below the mobile breakpoint. 3. Open the task properties. 4. Tap the blocked-by chip. **Paperclip version or commit** Reproduced on `e5a7fd7038` from `master`. **Deployment mode** Built from source with the local development workflow. ## What Changed - Added a mobile-only action menu to blocked-by chips. - Kept the direct task link and hover/focus remove control on desktop. - Reused the existing removal confirmation before the relation update. - Added focused regression coverage for the mobile visit and remove choices. - Added a phone-width Storybook state with the action menu open. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/IssueProperties.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm build-storybook` - Opened the new Storybook state in Playwright Chromium with a Pixel 5 viewport. Confirmed that both actions are visible and fit in the viewport. ## Risks - Low risk. The behavior change is limited to the existing mobile breakpoint. - Desktop navigation and blocker removal keep their current behavior. - The menu uses the shared dropdown and dialog primitives. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex `gpt-5.6-sol`, xhigh reasoning. The Codex CLI managed the context window for this run. The model used repository tools, code execution, tests, and browser automation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.812.0-canary.6 |
||
|
|
e5a7fd7038 |
Add sandbox device-login for the Codex adapter (#11237)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Agent adapters connect Paperclip to tools such as the Codex command line tool. > - A sandboxed Codex agent may start without a credential. > - The operator needs a safe sign-in flow that does not expose credentials to the shared package or the sandbox. > - This pull request adds a company-scoped device-login flow with a temporary Daytona sandbox. > - The flow promotes the credential only after readiness checks pass and removes the temporary sandbox after use. > - The result lets an operator sign in to a sandboxed Codex agent from the agent form. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** A Codex adapter that runs in a sandbox cannot authenticate when the company has no pre-provisioned Codex credential. **Proposed solution** Add a company-scoped device-login session. Start a temporary sandbox, run `codex login --device-auth`, stream the code and URL, verify readiness, promote the credential, and delete the sandbox. **Alternatives considered** Pre-provisioning a credential does not support first-time sandbox login. Keeping the credential in the login sandbox does not provide a durable company credential. **Roadmap alignment** This supports the roadmap item for cloud and sandbox agents. **Additional context** The flow uses a five-minute cleanup reaper, compare-and-set status changes, and a PostgreSQL advisory lock to protect promotion and cleanup. ## What Changed - Add the adapter login-session contract, database table, and migration. - Add company-scoped server routes and a service for sandbox device login. - Add credential promotion, readiness checks, and cleanup after login. - Add restart-safe cleanup for abandoned login sandboxes. - Add sandbox login controls to the agent creation and edit forms. - Keep device-login and vendor identifiers out of public shared and adapter UI symbols. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local exec vitest run` passed with 310 tests at the submitted commit. - The server login route, service, and reaper tests passed with 45 tests at the submitted commit. - The agent form render tests passed with 26 tests at the submitted commit. - The public-symbol leak check passed at the submitted commit. - A live Daytona sign-in flow still requires confirmation by a user with a live sandbox. ## Risks The migration adds a new company-scoped table. A promotion or cleanup race could remove a credential or leave a sandbox active, so the service uses claims, compare-and-set transitions, and an advisory lock. The live Daytona flow needs operator confirmation because local tests do not provide a real browser sign-in. ## Model Used OpenAI Codex, GPT-5, tool use and code execution, extended reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.812.0-canary.5 |
||
|
|
9c941169a6 |
fix(ui): keep new task dialog visible above mobile keyboard (#11281)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators create tasks in a dialog that includes the assignee and project fields. > - Mobile browsers reduce and offset the visual viewport when the on-screen keyboard opens. > - The dialog used layout viewport units, so its upper fields could move off-screen while the user typed. > - This pull request makes the dialog follow the live visual viewport and keeps the focused editor visible. > - The benefit is that operators can see the task context and the field they edit on mobile devices. ## Linked Issues or Issue Description **What happened?** On mobile browsers, opening the keyboard in the new-task dialog could move the assignee and project fields above the visible screen. The active editor could also become difficult to see. **Expected behavior** The full dialog must stay inside the visible browser area. The active editor and task controls must remain reachable while the on-screen keyboard is open. **Steps to reproduce** 1. Open Paperclip on a mobile browser. 2. Open the new-task dialog. 3. Focus the title or description editor to open the on-screen keyboard. 4. Observe that the upper fields can move outside the visible viewport. **Paperclip version or commit** Reproduced before commit `838cdbb325` on `master`. **Deployment mode** Local dev (`pnpm dev`) in a mobile browser viewport. ## What Changed - Read `window.visualViewport` while the dialog is open. - Apply token-based dialog geometry when the visual viewport is constrained. - Keep the focused editor visible after viewport resize and scroll events. - Add unit coverage for visual viewport updates and focus scrolling. - Add Playwright coverage for mobile, tablet, desktop keyboard, and unconstrained desktop layouts. ## Verification - `pnpm exec vitest run ui/src/components/NewIssueDialog.test.tsx` — 27 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — passed with all gates clean. - `pnpm --filter @paperclipai/ui build-storybook` — passed. - `pnpm exec playwright test tests/storybook-visual/new-issue-dialog-viewport.spec.ts --config tests/storybook-visual/playwright.config.ts` — 4 tests passed. ## Risks - Low risk. The custom geometry only activates when `visualViewport.height` is less than `window.innerHeight`. - Browsers without the Visual Viewport API keep the existing dialog primitive behavior. - The browser test checks hit targets and visible bounds at mobile, tablet, and desktop widths. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The session used reasoning, repository tools, shell execution, and browser automation. The service did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.812.0-canary.4 |
||
|
|
67001ec6eb |
chore(db): treat Drizzle migration snapshots as binary in diffs (#11254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip stores its state in a Postgres database, managed in `packages/db`. > - The schema uses Drizzle. `drizzle-kit` writes a full-schema snapshot to `packages/db/src/migrations/meta/` for every migration. > - Each snapshot is a large generated JSON file. One snapshot is about 39k lines. > - PR #11240 marked these files `linguist-generated=true`. That collapses the file view and drops the files from language stats. > - But `linguist-generated` does not change the PR line-count badge. Git still counts every snapshot line, so a PR that adds one migration shows a ~39k-line badge (two migrations show ~78k). PR #11237 shows +84,665 for this reason. > - This pull request adds `-diff` to the same files, so git treats them as binary and their lines leave the +/- count. > - The benefit is a pull request badge that reflects the real code change, not generated snapshot noise. ## Linked Issues or Issue Description No existing issue. This is a small follow-up to merged PR #11240. Description follows the enhancement template: **Problem / Motivation** PR #11240 marked `packages/db/src/migrations/meta/**` as `linguist-generated=true`. That attribute collapses the diff in the Files-changed view and removes the files from language stats, but it does not remove their lines from the PR additions/deletions badge. Each Drizzle snapshot is a full copy of the schema (~39k lines), so any PR that adds a migration still shows a huge line count. PR #11237 shows +84,665, of which ~78k are two generated snapshots. **Proposed Solution** Add `-diff` to the same glob. Git then treats the snapshots as binary. `git diff --numstat` reports `-` for these files, so their lines leave the +/- badge and GitHub shows "Binary file not shown" in place of the full JSON. **Alternatives Considered** - Keep only `linguist-generated`: leaves the misleading ~39k/78k badge on every migration PR. - Use the `binary` macro (`-diff -merge -text`): also disables EOL normalization. This repo has open CRLF/LF work, so `-text` is left unset on purpose. ## What Changed - Added `-diff` to `src/migrations/meta/**` in `packages/db/.gitattributes`. - Kept `linguist-generated=true` (language stats) and `-merge` (no auto-merge of generated snapshots). - Left `-text` unset on purpose, so end-of-line normalization stays intact. - Left the `.sql` migration files untouched, so their diffs stay visible for review. ## Verification Check the attributes and confirm the snapshot is now treated as binary: ``` git check-attr linguist-generated diff merge -- \ packages/db/src/migrations/meta/0031_snapshot.json \ packages/db/src/migrations/0009_fast_jackal.sql git diff --numstat origin/master...origin/feat/adapter-sandbox-login -- \ packages/db/src/migrations/meta/0214_snapshot.json ``` Expected: ``` packages/db/src/migrations/meta/0031_snapshot.json: linguist-generated: true packages/db/src/migrations/meta/0031_snapshot.json: diff: unset packages/db/src/migrations/meta/0031_snapshot.json: merge: unset packages/db/src/migrations/0009_fast_jackal.sql: linguist-generated: unspecified packages/db/src/migrations/0009_fast_jackal.sql: diff: unspecified packages/db/src/migrations/0009_fast_jackal.sql: merge: unspecified - - packages/db/src/migrations/meta/0214_snapshot.json ``` The snapshot reports `-` in numstat (binary, not counted). The `.sql` migration keeps normal diff behavior. GitHub reads the rule from the PR tree, so the badge drops on the next PR that touches these files. ## Risks Low risk. The change only affects how git and GitHub render and count generated files. It does not touch application code, the schema, or any migration. - With `-diff`, GitHub and local `git diff` no longer show a text diff for a snapshot. This is intended; the files are generated and are not reviewed by hand. The raw file is still viewable. - `-text` is left unset, so this change does not affect the CRLF/LF handling that other PRs (for example #8922) address. ## Model Used Claude Opus 4.8 (Anthropic), model ID `claude-opus-4-8`, used through Claude Code with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — not applicable; this change touches no code path, only git/GitHub file handling - [x] I have added or updated tests where applicable — not applicable; `.gitattributes` behavior is verified with `git check-attr` and `git diff --numstat` (see Verification) - [x] I have updated relevant documentation to reflect my changes — not applicable; the `.gitattributes` file documents its own rules inline - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.812.0-canary.3 |
||
|
|
f16071a290 |
fix(ui): use a fictional tailnet hostname in the vite proxy test (#11248)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The UI dev server proxies `/api` to the backend and injects `x-forwarded-host`; a unit test asserts that injection with a sample Host header > - The sample Host header is a contributor's real machine and tailnet hostname, committed to the public repository > - Real personal hostnames do not belong in a public codebase, and this one also contains the contributor's OS username, so `scripts/check-forbidden-tokens.mjs` (which forbids the local username) blocks `npm` publishing from that contributor's machine > - This pull request replaces the fixture with a fictional tailnet-style hostname > - The benefit is no personal identifiers in the test fixtures and a passing forbidden-token check for every contributor ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** `ui/src/lib/vite-api-proxy.test.ts` uses a real contributor dev-machine hostname as its `Host` header fixture. `node scripts/check-forbidden-tokens.mjs` fails on that contributor's machine because the hostname contains their OS username, blocking the publish flow. Introduced in #10718. **Expected behavior** Test fixtures use fictional hostnames. The forbidden-token check passes on every contributor machine. **Steps to reproduce** 1. On a machine whose OS username appears in the fixture hostname, run `node scripts/check-forbidden-tokens.mjs`. 2. The check reports the two lines in `ui/src/lib/vite-api-proxy.test.ts` and blocks with exit code 1. ## What Changed - `ui/src/lib/vite-api-proxy.test.ts`: the `Host` fixture is now `dev-box.tail1234.ts.net:3101` (fictional). The test only asserts that whatever host arrives is injected as `x-forwarded-host`, so the value is arbitrary. ## Verification - `pnpm vitest run src/lib/vite-api-proxy.test.ts` in `ui/` — 5 tests pass. - `node scripts/check-forbidden-tokens.mjs` — "No forbidden tokens found" on the previously affected machine. - Note: git history retains the old value; this removes it from the current tree only. ## Risks None. A test fixture string with no behavioral coupling. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.812.0-canary.2 |
||
|
|
d90f4d488e |
fix(ui): survive first load against a cold backend without a blank page (#11246)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The web UI holds live websocket connections for run events, coordinates cross-tab polling through a leader-election store, and renders app chrome (sidebar, providers) around a routed outlet > - When the backend is still cold-starting (managed hosting wake, server restart, reverse proxy up before the app), the event websockets refuse connections and the first SPA load mounts against a dead backend > - In that state the mount cascade can exceed React's nested update limit (minified error #185); the crash originates in shell hooks outside the routed error boundary, so React unmounts the entire root to a blank page, and the dead page keeps retrying the websocket on a flat 1.5s timer until the user hard-refreshes > - This pull request removes the wasted nested commits from the shared-polling subscription path, adds exponential backoff to the transcript websocket reconnect, and adds a last-resort app-shell error boundary > - The benefit is that a cold or briefly unreachable backend degrades to a recoverable state instead of a blank page that hammers the server ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** On the first load against a backend that was still starting, the app showed its loading animation and then a blank page. The console showed repeated `WebSocket connection to 'wss://…/api/companies/<id>/events/ws' failed` lines and `Uncaught Error: Minified React error #185` with a stack through the shared-polling coordinator's `subscribe`. The websocket retries continued indefinitely on the dead page. A manual refresh fixed it. **Expected behavior** A backend that is briefly unreachable degrades gracefully: websocket reconnects back off, the UI keeps rendering from cache, and even a worst-case crash shows a reload prompt instead of a blank page. **Steps to reproduce** 1. Serve the UI while the backend API is still starting (websocket upgrades and API calls refused). 2. Load any company page with several shared-polling consumers mounted (dashboard with sidebar). 3. Observe repeated websocket failures; on affected loads the page goes blank with React error #185. ## What Changed - `ui/src/hooks/useSharedPolling.ts`: coordinator snapshot notifications now keep the previous state object when leadership did not change, so React bails out instead of scheduling a nested re-render. `subscribe` invokes its listener synchronously from inside the mount effect with a fresh object each time; before this change every mount and notify burned nested-update budget even with no value change — the crash frame in the field report was exactly this `subscribe → setState` call. - `ui/src/components/transcript/useLiveRunTranscripts.ts`: the live event websocket reconnect backs off exponentially (1.5s → 15s cap, reset on successful open), mirroring `LiveUpdatesProvider`, instead of a flat 1.5s retry. - `ui/src/components/AppErrorBoundary.tsx` (+ wiring in `ui/src/main.tsx`): a dependency-free boundary above the router and providers. `RouteErrorBoundary` only guards the routed `<Outlet />`; a crash in the shell around it had no boundary, so React unmounted the root to a blank page. The boundary renders a reload prompt with the error message. - Tests: `useSharedPollingSnapshot.test.tsx` (mount costs no extra commit — fails against the previous code; a real leadership change re-renders exactly once and ticks stay quiet), a backoff test in `useLiveRunTranscripts.test.tsx` (delays grow 1.5s → 3s → 6s and reset after a successful open), and `AppErrorBoundary.test.tsx` (render throw, effect throw, healthy pass-through). ## Verification - `pnpm vitest run` in `ui/` over the touched suites (shared polling, cross-tab poll, transcripts, boundary): 34 tests pass. - `pnpm typecheck` in `ui/` — clean. - The snapshot regression test was verified to fail against the pre-change hook (extra commit per mount). - Not reproduced end-to-end: the exact 50-update cascade from the field crash needs a live cold backend; the change removes the identified per-mount/per-notify nested commits at the reported crash frame, bounds the reconnect load, and guarantees the shell can no longer blank the page. ## Risks Low risk. The snapshot change only suppresses re-renders whose state is value-identical; leadership changes propagate exactly as before. The backoff only lengthens retry delays after consecutive failures and resets on success. The new boundary renders children untouched unless an error reaches it; behavior on healthy loads is unchanged. Self-hosted deployments see the same code paths — the cold-backend window simply rarely occurs there. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution; diagnosis included mapping the production minified stack to source). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9adeb4a9d0 |
fix(ui): fetch the web app manifest with credentials (#11245)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The web UI is a PWA-capable SPA; `ui/index.html` links `/site.webmanifest` so browsers can read app metadata > - Browsers fetch `<link rel="manifest">` in "omit credentials" mode unless the link opts in with `crossorigin="use-credentials"` > - Self-hosted this is harmless, but when Paperclip runs behind an authenticating reverse proxy (a managed hosting front door), the cookie-less manifest request is rejected with 401 on every page load and logs a console error pair on each navigation > - This pull request adds `crossorigin="use-credentials"` to the manifest link so the request carries the same session cookies as every other same-origin asset request > - The benefit is a clean console and a servable manifest in proxied deployments, with self-hosted behavior unchanged ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** On every page load behind an authenticating reverse proxy, the browser logs `Failed to load resource: the server responded with a status of 401` for `/site.webmanifest`, plus `Manifest fetch from … failed, code 401`. The proxy rejects the request because the browser sends the manifest fetch without cookies. **Expected behavior** The manifest request carries the same session credentials as every other same-origin asset request, so the proxy can authenticate and serve it. No console errors. **Steps to reproduce** 1. Serve Paperclip behind a reverse proxy that requires a session cookie for all app routes. 2. Sign in and load any page. 3. Open the browser console: the manifest fetch fails with 401 while all other assets load. ## What Changed - `ui/index.html`: the manifest link now carries `crossorigin="use-credentials"`. - `ui/src/lib/pwa-install-mode.test.ts`: a regression test asserts the attribute stays on the link. ## Verification - `pnpm vitest run src/lib/pwa-install-mode.test.ts` in `ui/` — 2 tests pass. - Manual check of the rendered link tag in `ui/index.html`. ## Risks Low risk. The manifest is same-origin, so `use-credentials` only switches the fetch from "omit" to the include behavior all other same-origin requests already have. Self-hosted deployments see no change. Cross-origin manifest hosting is not used in this project. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d5bb396518 |
fix: pass sandbox provider credential env vars to plugin workers; hide Local default under managed-sandbox-only (#11244)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments give each agent run an execution target, and sandbox providers (Daytona, E2B, Novita, exe.dev) run as plugin workers > - A managed deployment provisions one platform-managed sandbox row with no credential in config; the provider is documented to fall back to its process env var (for example `DAYTONA_API_KEY`) > - Plugin workers spawn with a scrubbed environment, so that fallback never sees the host env var — probe and lease acquisition fail with "require an API key in config or DAYTONA_API_KEY" even when the deployment sets the var > - Separately, the managed-sandbox-only mode hides local rows from every list, but the instance Default picker renders a hardcoded synthetic "Local" option that no filter touches > - This pull request forwards each bundled provider's documented credential env var to its own plugin worker, and gates the synthetic Local option on the flag > - The benefit is that the documented host-env credential fallback works for plugin-backed providers, and managed-sandbox-only instances no longer offer Local anywhere ## Linked Issues or Issue Description **Subsystem affected** Plugin worker environment construction (`server/src/services/plugin-loader.ts`) and the environments UI (instance Default picker, agent form inherited-environment label). **Problem or motivation** Two follow-ups to the managed-sandbox-only mode (#11200), both found on a live managed deployment: 1. The deployment sets `DAYTONA_API_KEY` as a server env var and the managed sandbox row omits `config.apiKey` by contract. "Test Connection" fails with `Sandbox environment probe failed for provider "daytona". Daytona sandbox environments require an API key in config or DAYTONA_API_KEY.` A real agent run fails the same way at lease acquisition. The cause: sandbox providers run as plugin workers, and `buildPluginWorkerEnv` passes only model-provider keys and in-cluster Kubernetes vars. The provider's own documented credential env var never reaches the worker, so the in-plugin `process.env` fallback reads nothing. The self-hosted path has the same gap: the Daytona plugin README documents `DAYTONA_API_KEY` as a host-level fallback, and it does not work today. 2. With `enableManagedSandboxOnly` on, the instance Default environment picker still shows "Local". The server filters local *rows* out of the list, and the client filter mirrors that for cached lists, but this option is a hardcoded `<option value="">Local</option>` — not a list row — so no filter removes it. Selecting it writes a null default, which run selection then rejects fail-closed. **Proposed solution** Forward each bundled sandbox provider's documented credential env var into its plugin worker, keyed by the manifest's declared `environmentDrivers[].driverKey` so a worker only receives its own provider's credential (daytona → `DAYTONA_API_KEY`, e2b → `E2B_API_KEY`, exe-dev → `EXE_API_KEY`, novita → `NOVITA_API_KEY`). Keep the existing gate: only plugins that declare `environment.drivers.register` receive any passthrough. In the UI, render the synthetic Local option only when managed-sandbox-only is off; under the flag show a disabled "Select environment" placeholder only while no default is stamped yet, and stop the agent form's inherited label from reading "Local". **Alternatives considered** Adding `DAYTONA_API_KEY` to the existing `ADAPTER_ENV_PASSTHROUGH` list was rejected: that list goes to every environment-driver plugin, so each provider would receive every other provider's credential. A manifest schema field for declared credential env vars was rejected as heavier than needed: the bundled providers are known, and the mapping lives next to the two existing passthrough lists. ## What Changed - `server/src/services/plugin-loader.ts`: new `SANDBOX_PROVIDER_CREDENTIAL_ENV_PASSTHROUGH` map (driverKey → documented credential env vars). `buildPluginWorkerEnv` reads the manifest's `environmentDrivers` and forwards only the matching vars, after the existing `environment.drivers.register` gate. Blank values stay excluded. - `server/src/__tests__/plugin-database.test.ts`: the daytona worker receives `DAYTONA_API_KEY` and not another provider's key; a plugin whose drivers have no mapping (kubernetes) receives no credential var. - `ui/src/pages/CompanyEnvironments.tsx`: the Default picker's synthetic Local option renders only when managed-sandbox-only is off. Under the flag, a disabled "Select environment" placeholder renders only while the default is unset. - `ui/src/pages/CompanyEnvironments.test.tsx`: the Local option is present by default and absent under the flag; saved non-local environments stay selectable. - `ui/src/components/AgentConfigForm.tsx`: the inherited-environment label falls back to "Managed sandbox" instead of "Local" under the flag. ## Verification - `server`: `npx vitest run src/__tests__/plugin-database.test.ts -t buildPluginWorkerEnv` — 5 passed (3 existing, 2 new). - `ui`: `npx vitest run src/pages/CompanyEnvironments.test.tsx` — 22 passed (2 new); `npx vitest run src/components/AgentConfigForm.render.test.tsx` — 10 passed. - `tsc --noEmit` clean in `server` and `ui`. - Live managed deployment: confirmed the tenant service env carries `DAYTONA_API_KEY` while the probe fails with the exact message above, which pins the root cause to the worker env, not delivery. ## Risks - The worker env grows by exactly one var per matching bundled provider, only when the deployment sets it and only for plugins that declare a matching environment driver. Plugins without a mapping see no change. - Self-hosted behavioral shift is the fix itself: a host-level `DAYTONA_API_KEY` (or E2B/EXE/NOVITA equivalent) now reaches the provider as its README documents. Deployments that set the var but expected it to stay inert had no working configuration to preserve — the provider errored on every keyless probe and run. - UI change is inert unless `enableManagedSandboxOnly` is on (default false everywhere). ## Model Used Claude Fable 5 (`claude-fable-5`) via Claude Code — extended thinking, tool use, parallel read-only subagents for the two root-cause traces. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.812.0-canary.1 |
||
|
|
0aa743fc30 |
build(db): clean dist before drizzle generate (#11241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Database migrations are generated by drizzle-kit, which reads the schema from the db package's built `dist/schema/*.js` > - `tsc` never deletes stale outputs, so a long-lived checkout keeps compiled schema files whose sources were deleted long ago > - A `generate` run in such a checkout sees those ghost tables and sweeps phantom `CREATE TABLE` statements into an unrelated migration > - This pull request makes `generate` clean `dist` before building, so drizzle always diffs against exactly the current schema sources > - The benefit is that no contributor can accidentally resurrect deleted tables inside a new migration ## Linked Issues or Issue Description **What happened?** Running `pnpm --filter @paperclipai/db generate` in a months-old checkout produced a migration re-creating `cloud_upstream_connections`, `cloud_upstream_runs`, and `company_secret_pools` — tables whose schema sources were deleted in #10507. The compiled copies were still in `dist/schema/`, and `drizzle.config.ts` reads the schema from `dist`, so drizzle treated them as new tables missing from the snapshot. **Expected behavior** `generate` diffs the current schema sources only; deleted tables can never reappear in a generated migration. **Steps to reproduce** 1. Build the db package, then delete a schema source file without cleaning `dist`. 2. Run `pnpm --filter @paperclipai/db generate`. 3. The generated migration re-creates the deleted table. ## What Changed - `packages/db/package.json`: the `generate` script runs `pnpm run clean` before `tsc`, so the drizzle-kit input is always a fresh build of the current sources. ## Verification - In a checkout carrying stale `dist/schema/cloud_upstreams.js` / `company_secret_pools.js` artifacts, `generate` produced a phantom migration before this change; after a clean build it reports "No schema changes, nothing to migrate". With this change the clean happens inside `generate` itself. ## Risks - Low risk: `generate` is a developer-only script; the change only adds the existing `clean` step ahead of the existing build, at the cost of a full rebuild per generate. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.812.0-canary.0 |
||
|
|
b5ebda1dca |
fix(grok-local): report real token usage and cost instead of hardcoded zeros (#10433)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Cost/usage tracking is core to that: the dashboard shows per-agent spend so a team can see what their AI workforce is costing them > - The `grok_local` adapter (xAI's Grok Build CLI) is a newer adapter than `claude_local`/`codex_local`, and its usage/cost wiring was left incomplete > - Every `grok_local` run persists `usage.inputTokens/outputTokens/cachedInputTokens = 0` and `costUsd = null` in `heartbeat_runs`, unconditionally, even though the underlying `grok` CLI reports real, non-zero token counts and cost per turn in its own JSON stream > - This pull request wires the parser to actually read `usage`/`total_cost_usd` from the CLI's terminal `end` event, threads those values into the adapter's execution result, and marks them `usageBasis: "per_run"` so the heartbeat service doesn't incorrectly delta them against a prior run on a resumed session (matching how `claude_local`/`codex_local` already do this) > - The benefit is accurate cost/usage visibility for any self-hosted Paperclip instance running Grok Build agents, instead of a dashboard that always reads zero ## Linked Issues or Issue Description Fixes: #10432 ## What Changed - `packages/adapters/grok-local/src/server/parse.ts`: `parseGrokJsonl()` now reads `usage.input_tokens` / `usage.output_tokens` / `usage.cache_read_input_tokens` / `total_cost_usd` from the terminal `end` event and returns them on `ParsedGrokJsonl` (previously discarded entirely). - `packages/adapters/grok-local/src/server/execute.ts`: `toResult()` now populates `usage.inputTokens/outputTokens/cachedInputTokens` from the parsed values instead of hardcoded `0`, sets `usageBasis: "per_run"` (each `--single` invocation reports usage for just that process, not a running session total), and surfaces `costUsd` only when `billingType === "api"` (metered) — subscription/OAuth billing has no marginal dollar cost, so it stays `null` there, but token counts are populated for both billing types since usage visibility is useful regardless of billing model. - `packages/adapters/grok-local/src/server/parse.test.ts`: added a test asserting usage/cost extraction from a representative `end` event payload, and updated the existing exact-equality test for the new fields. - `packages/adapters/grok-local/src/server/execute.test.ts`: added a test covering both subscription billing (tokens populated, `costUsd: null`) and API-key billing (tokens populated, real `costUsd`), and asserting `usageBasis: "per_run"` in both cases. ## Verification - `pnpm vitest run packages/adapters/grok-local/src/server/parse.test.ts packages/adapters/grok-local/src/server/execute.test.ts` — 9/9 passed - `tsc --noEmit` on the `grok-local` package — clean - Verified against a real self-hosted Paperclip instance running `grok` CLI `0.2.112` with SuperGrok subscription (OAuth) auth: confirmed the raw CLI stream reports real `usage`/`total_cost_usd` (e.g. `"usage":{"input_tokens":21560,...},"total_cost_usd":0.0564448`) that was previously discarded before ever reaching `heartbeat_runs.usage_json`, which always showed all-zero tokens regardless of real usage. ## Risks - Low risk, additive change scoped entirely to the `grok_local` adapter's usage/cost reporting path — no change to control flow, session handling, or process execution. - `usageBasis: "per_run"` mirrors the existing, already-tested pattern in `claude_local`/`codex_local` execute paths, so the heartbeat service's per-run vs. session-cumulative delta logic is exercised the same way. - `costUsd` is intentionally left `null` for subscription/OAuth billing (no behavior change there beyond now-populated token counts) to avoid implying a dollar cost that doesn't exist for flat-rate billing. ## Model Used Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code, no extended thinking. Root cause was found by comparing real `grok` CLI JSON stream output (captured directly from a live invocation) against the persisted `heartbeat_runs.usage_json` row for the same run on a self-hosted instance, then reading `parse.ts`/`execute.ts` source to confirm the hardcoded zero values. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`fix/grok-local-usage-cost-tracking`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (no user-facing docs reference this internal usage-reporting behavior) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending at time of writing) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (addressed the one P1 raised — `usageBasis: "per_run"`) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c0bdf26633 |
chore(db): collapse Drizzle migration snapshot diffs and block auto-merge (#11240)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip stores its state in a Postgres database, managed in `packages/db`. > - The schema uses Drizzle. `drizzle-kit` writes a full-schema snapshot to `packages/db/src/migrations/meta/` for every migration. > - Each snapshot is a large generated JSON file. One snapshot is over 200 KB. > - GitHub shows these files as 30k+ line diffs in a pull request. The diffs add no review value, because a human never edits the files. > - The files also merge badly. `drizzle-kit` computes the `id`/`prevId` chain and the full schema state, so a line-level merge of two snapshots produces a file that no real `generate` run creates. > - This pull request adds a `.gitattributes` file that marks the snapshot directory as generated and blocks its auto-merge. > - The benefit is clean pull request diffs and a loud conflict that forces the correct fix when two branches add a migration. ## Linked Issues or Issue Description No existing issue. This is a small repository-hygiene change. Description follows the enhancement template: **Problem / Motivation** Every migration adds a full-schema snapshot JSON under `packages/db/src/migrations/meta/`. These files are large and generated. GitHub renders them as 30k+ line diffs in pull requests, which buries the real change (the `.sql` migration) in noise. The files also have no meaningful line-level merge: `drizzle-kit` computes each snapshot's `id`/`prevId` chain and full schema state. **Proposed Solution** Add `packages/db/.gitattributes`: - `linguist-generated=true` on `src/migrations/meta/**` — GitHub collapses the diff and drops the files from language stats. - `-merge` on the same glob — git refuses the line-level merge and raises a conflict instead of fabricating an invalid snapshot. **Alternatives Considered** - `-diff` / `binary`: hides the diff completely and blocks text merge, but also blocks any local `git diff` and gives a worse conflict experience. `linguist-generated` keeps the file expandable and text-based, so it is the lighter option. - Do nothing: leaves the noisy diffs and the risk of a silent bad merge. ## What Changed - Added `packages/db/.gitattributes`. - Marked `src/migrations/meta/**` as `linguist-generated=true` to collapse the snapshot and journal diffs on GitHub. - Set `-merge` on the same files so git raises a conflict instead of auto-merging generated snapshots. - Left the `.sql` migration files untouched, so their diffs stay visible for review. ## Verification Run `git check-attr` against the affected files and a control `.sql` file: ``` git check-attr linguist-generated merge -- \ packages/db/src/migrations/meta/0031_snapshot.json \ packages/db/src/migrations/meta/_journal.json \ packages/db/src/migrations/0009_fast_jackal.sql ``` Expected output: ``` packages/db/src/migrations/meta/0031_snapshot.json: linguist-generated: true packages/db/src/migrations/meta/0031_snapshot.json: merge: unset packages/db/src/migrations/meta/_journal.json: linguist-generated: true packages/db/src/migrations/meta/_journal.json: merge: unset packages/db/src/migrations/0009_fast_jackal.sql: linguist-generated: unspecified packages/db/src/migrations/0009_fast_jackal.sql: merge: unspecified ``` The snapshot and journal files carry both attributes. The `.sql` migration keeps its normal diff and merge behavior. GitHub applies the rule from the pull request tree, so the collapse shows on the next pull request that touches these files. ## Risks Low risk. The change only affects git and GitHub display and merge behavior for generated files. It does not touch application code, the schema, or any migration. - `-merge` leaves the current-branch version in the working tree on conflict and marks the file conflicted. It does not insert conflict markers into the JSON. The correct resolution stays "renumber the later migration and regenerate", then commit. - Related open pull requests #879 and #8922 also add `.gitattributes` rules for migration files, but for CRLF/LF hash mismatches on Windows. If either lands, a follow-up can merge the rules into one file. There is no functional overlap with this change. ## Model Used Claude Opus 4.8 (Anthropic), model ID `claude-opus-4-8`, used through Claude Code with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — not applicable; this change touches no code path, only git/GitHub file handling - [x] I have added or updated tests where applicable — not applicable; `.gitattributes` behavior is verified with `git check-attr` (see Verification) - [x] I have updated relevant documentation to reflect my changes — not applicable; the `.gitattributes` file documents its own rules inline - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3b74ff4813 |
fix(ui): use issuePrefix instead of name-derived prefix in create dialog badges (#8550)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents at work > - Paperclip stores an `issuePrefix` on each company (e.g. "OPS") used for issue identifiers (`OPS-1`) and company-prefixed routes (`/OPS/dashboard`) > - The create-dialog badges in NewIssueDialog, NewProjectDialog, and NewGoalDialog were derived from the company display name using `company.name.slice(0, 3).toUpperCase()` — so "Acme Labs" showed "ACM" > - This is misleading because the badge visually represents the issue prefix, but actually shows an unrelated 3-letter slice of the display name > - When a company has `issuePrefix = "OPS"` but `name = "Acme Labs"`, the badge showed "ACM" while issues use "OPS-1" > - This pull request replaces `name.slice(0, 3).toUpperCase()` with `company.issuePrefix` in all three dialog badge components > - The benefit is that the badge now matches the actual prefix used for issues and routes, eliminating confusion ## Linked Issues or Issue Description Fixes: #8501 ## What Changed - `ui/src/components/NewIssueDialog.tsx` (line ~1339): Replaced `company.name.slice(0, 3).toUpperCase()` with `company.issuePrefix` in the selected-company header badge - `ui/src/components/NewIssueDialog.tsx`: Replaced `company.name.slice(0, 3).toUpperCase()` with `company.issuePrefix` in the company picker list badge - `ui/src/components/NewProjectDialog.tsx`: Replaced `selectedCompany.name.slice(0, 3).toUpperCase()` with `selectedCompany.issuePrefix` in the selected-company header badge - `ui/src/components/NewGoalDialog.tsx`: Replaced `selectedCompany.name.slice(0, 3).toUpperCase()` with `selectedCompany.issuePrefix` in the selected-company header badge ## Verification 1. Create or configure a company whose `issuePrefix` differs from the first 3 letters of its display name (e.g. name = "Acme Labs", issuePrefix = "OPS") 2. Open the New Task dialog — the selected-company header badge should show "OPS", not "ACM" 3. Open the company picker dropdown inside the New Task dialog — each company list badge should show the actual `issuePrefix` 4. Open the New Project dialog — the selected-company header badge should show "OPS" 5. Open the New Goal dialog — the selected-company header badge should show "OPS" 6. Verify that companies whose prefix matches the first 3 letters (e.g. name="Ops Team", prefix="OPS") still display correctly **Before/After:** - Before: Company "Acme Labs" with `issuePrefix = "OPS"` showed badge "ACM" - After: Same company shows badge "OPS" (Screenshots require running the UI locally against a test instance with the relevant company configuration.) ## Risks Low risk — this is a purely visual change to 3 React component badge labels. No API changes, no schema changes, no behavioral changes to issue creation or routing. The `issuePrefix` field is already loaded on the company objects used by these components. ## Model Used - **Provider:** OpenCode - **Model:** MiMo v2.5 Free - **Reasoning:** N/A (standard code generation) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched GitHub for duplicate or related PRs and found none targeting the same badge code - [x] I have linked the existing issue with Fixes: #8501 - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change (`fix/issue-8501`) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>canary/v2026.811.0-canary.18 |
||
|
|
0a95ada1be |
feat(server): chunked import preview endpoint and resumable upload in the Import page and CLI (#11224)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The previous pull request added server-side chunked resumable import transfers; without clients, large imports still ride the single fragile upload > - The Import page and the CLI need to slice large packages, upload parts with retry and progress, resume after interruptions, and preview before applying > - Preview is the missing server piece: the browser flow is preview-then-import, so a completed spool must be previewable without re-uploading > - This pull request adds the transfer preview endpoint, switches the Import page to the chunked path for zips over 48 MB, and teaches the CLI the same for oversized local imports > - The benefit is that large imports get progress, per-part retry, and resume in both clients, while small imports keep the exact single-shot path they have today ## Linked Issues or Issue Description **What happened?** With only the server transfer routes in place, users still upload large company packages as one request from the Import page and the CLI: no progress indication, no retry below the whole file, and no resume after a dropped connection or refresh. The preview-then-import flow also cannot run against an uploaded transfer, forcing a second full upload. **Expected behavior** A large package uploads once as verified parts with visible progress; preview and import both run against the uploaded spool; an interrupted upload resumes with only the missing parts re-sent; packages at or below 48 MB behave exactly as before. **Steps to reproduce** 1. Select a 500 MB zip on the Import page over an unreliable connection. 2. Watch the single upload fail near the end and restart from zero, twice — once for preview, once for import. 3. Same story headless via the CLI. ## What Changed - Server: `POST /import/transfers/:id/preview` runs the existing preview logic against the completed spool (shared assembly + whole-file verification helper with apply); preview neither completes the run nor deletes the spool, so the subsequent apply reuses it. Missing parts respond with the missing list. - UI: zips over 48 MB take the chunked path in both preview and import — the file is sliced into 32 MB parts hashed with WebCrypto (single ArrayBuffer, no second copy), the transfer is created or resumed (the create response's missing-parts list drives what uploads), parts upload sequentially with three attempts each and visible progress, then transfer preview/apply replace the multipart calls. The existing preview pane, collision handling, adapter overrides, and async job polling are unchanged; ≤ 48 MB keeps the single-shot path. - CLI: oversized local `.zip` or folder imports zip/slice/hash with node crypto, upload with resume and per-part retry and progress lines, and use transfer preview/apply. Small packages keep the inline path byte-identical. - Failure honesty: adapters/API errors fail open to existing behavior; a part failing all attempts surfaces a durable error panel with resume intact. ## Verification - Server: preview-then-apply on one spool (run stays open, spool intact, then apply completes), preview with missing parts rejected — added to the transfer route suite (embedded Postgres). - UI suite: large file takes the chunked path (manifest shape, part uploads, progress, apply on a resumed transfer, single-shot endpoints never called), small file stays single-shot, part failure after three attempts surfaces the error panel without running preview, resume re-uploads only the missing part. - CLI: manifest slicing/hashing, threshold behavior for zip and folder sources, folder-zip round-trip through the real zip reader, upload resume/retry/exhaustion/already-completed, full-command chunked and small-zip inline flows. - Server, ui, cli typechecks clean. Exact counts in the PR checks. ## Risks - The 48 MB threshold only routes between two verified paths; behavior below it is untouched. - Chunked CLI imports use the board-scoped transfer routes, so oversized CLI imports need board credentials (agent tokens keep the agent-safe small-file path). No privilege change — board actors already had the generic routes — but the two size regimes differ semantically; called out for review. - A CLI dry-run over the threshold uploads parts before previewing; the spool persists (24 h sweep) and a later apply resumes without re-upload — inherent to preview-against-spool. Stacked on #11223 — merge that first; this PR then shows only the preview endpoint and client changes. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.17 |
||
|
|
8f478242f1 |
fix(adapter-utils): close sandbox stdin file race with atomic write and fault-tolerant poller (#11235)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters use execution targets to exchange input and output with sandbox processes > - The sandbox input path can expose partial files, and the poller can delete or stop on invalid input > - These timing windows can lose agent input without a clear error > - This pull request makes host writes atomic and makes the poller retry invalid files before it drops them > - The benefit is reliable sandbox input delivery with visible failure after bounded retries ## Linked Issues or Issue Description Closes #10874 ## What Changed - Decode host input into a temporary file, then rename it onto the final JSON path. - Apply the same atomic write pattern to the filesystem client. - Parse each input file before deletion. - Retry parse failures and drop a file after the bounded retry limit with an error event. - Add regression tests for empty, partial, and permanently malformed input files. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/execution-target-stdin-race.test.ts` passes 5/5 tests. - Related sandbox callback, execution target, and sandbox execution suites pass 71/71 tests. - TypeScript checks pass for the changed files. - The regression suite fails on the old code and passes on this change. ## Risks - Low risk. The change affects sandbox input file handling and adds bounded retry behavior. - A permanently malformed file now creates an error event after the retry limit. > This bug fix does not add a core feature, so a roadmap change is not needed. ## Model Used Codex, OpenAI GPT-5, current agent runtime, large context window, tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
23a1b025c2 |
feat(server): chunked resumable company import transfers (#11223)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import moves large packages into an instance, and since the upload cap rose to 1 GB, the transport is the weak point: one HTTP request, buffered fully in memory, with no resume > - A dropped connection at 90% of an 800 MB upload starts the whole transfer over, and a server restart loses all progress > - This pull request adds the server side of chunked resumable import transfers: a durable run ledger and routes that accept the same import zip as verified ~32 MB parts spooled to disk > - An interrupted transfer resumes from the parts already uploaded — across dropped connections, page refreshes, and server restarts — and peak upload memory drops from the whole package to one part > - The benefit is that large imports become reliable on real-world connections instead of all-or-nothing ## Linked Issues or Issue Description **What happened?** Large company imports travel as a single HTTP upload. On a slow or flaky connection, any interruption discards all progress and the upload restarts from zero. The server buffers the entire compressed package in memory during upload. A server restart mid-upload loses the transfer entirely. With the upload cap now at 1 GB, these failure modes govern exactly the imports the cap was raised for. **Expected behavior** A large import upload survives interruptions: already-transferred data is kept and verified, only the missing remainder is re-sent, and the server's memory use during upload is bounded by a part, not the package. **Steps to reproduce** 1. Import a multi-hundred-MB company package over a connection that drops mid-upload. 2. The upload fails; retrying starts from byte zero. 3. Repeat on an unstable connection and the import may never complete. ## What Changed - New `company_transfer_runs` table (drizzle schema + migration) and `companyTransferRunService`: one row per transfer with a content-derived idempotency key, per-part completion recorded atomically and idempotently, resume scoped to actor and direction, completed runs short-circuiting retries of identical content. - New transfer routes beside the existing import routes, same authorization: declare a sliced zip (`POST /import/transfers` — validates cap, 64 MB part ceiling, contiguity, size sums, sha256 format), upload parts (`PUT .../parts/:n` — raw body, hash-and-size verified before an atomic write to a disk spool under the instance root; re-uploads are no-op successes), poll resume state (`GET .../:id` — missing parts recomputed from disk), and apply (`POST .../:id/apply` — requires all parts, re-verifies the assembled zip against the whole-file hash fail-closed, then feeds the existing import pipeline through factored helpers rather than duplicated logic). - Hourly sweep fails and cleans spools idle for 24 h; a swept transfer honestly reports all parts missing on resume. - Strict UUID gating on run ids before any filesystem path construction. - The existing single-shot upload path is untouched; clients arrive in the follow-up PR. ## Verification - Transfer route suite (embedded Postgres): create/upload/status/apply round-trip with a real imported company, out-of-order parts, wrong-hash part rejected and unrecorded, re-upload no-op, apply-with-missing-parts rejection, resume after failure with prior progress intact, assembled-hash mismatch failing closed with spool deletion, actor scoping 404s, async-job apply, sweep followed by honest resume. - Ledger suite (embedded Postgres): part idempotency, actor/direction scoping, completed-run short-circuit, cancelled runs staying cancelled. - Existing portability route suite unchanged and green; server + db typechecks clean. Exact counts in the PR checks. ## Risks - New routes are additive; the existing import path is untouched. The transfer routes carry the same board authorization as the import routes they sit beside. - Disk spool: bounded by the existing upload cap per transfer, cleaned on success, failure, hash mismatch, and by the 24 h sweep. Spool paths are strict-UUID-gated. - The apply step still materializes the assembled zip in memory once (same profile as today's single-shot import at apply time); upload-time memory drops to one part. - Known limitation, deliberate: transfers are keyed on content alone, so identical package content cannot be imported twice without re-exporting (surfaced explicitly to the caller). Acceptable for v1; noted for review. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.16 |
||
|
|
2494a2a0fe |
perf: add repeatable issue-detail baseline rig (#10409)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The issue detail page is a core operator surface where perceived latency directly affects task navigation > - Performance work needs repeatable evidence so later optimizations can be compared against the same scenarios > - The page did not expose stable user-timing marks for its header or first useful content > - There was also no isolated seeded browser rig that measured warm navigation, cold deep links, waterfalls, or server time > - This pull request adds the instrumentation and a one-command Playwright baseline harness > - The benefit is that issue-page performance changes can be validated with reproducible median measurements instead of anecdotes ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: `ui/`, `server/`, and browser performance tooling. **Problem or motivation** The issue detail page performs a large client bootstrap and request fan-out, but the repository lacks stable user-timing boundaries and a repeatable benchmark. That makes performance changes difficult to compare and allows regressions to be judged from anecdotes instead of consistent evidence. **Proposed solution** Add stable header/content paint measures, development/QA-only lifecycle vital reporting, aggregate server timing for the issue endpoint, and a seeded Playwright command that runs warm/cold scenarios under throttled and unthrottled profiles with N≥5 median reporting. **Alternatives considered** Ad hoc DevTools recordings were rejected because they are not repeatable or reviewable. Production telemetry was rejected because this baseline should not change production data collection. A unit-only harness was rejected because it cannot capture browser bootstrap, rendering, and network waterfall costs. **Roadmap alignment** The roadmap calls for agent performance to be measurable over time. This change applies that evidence-first principle to a core operator page and does not duplicate a listed roadmap deliverable. **Additional context** The generated report includes warm and cold medians, TTFB/FCP/LCP where applicable, request and byte totals before first useful content, JavaScript bytes, and issue endpoint server timing. ## What Changed - Added `issue-detail:navigate→header-paint` and `issue-detail:navigate→content-paint` user-timing measures to the issue detail page. - Added development/QA-only TTFB, LCP, and INP console reporting without production telemetry delivery. - Added `Server-Timing` for `GET /api/issues/:id`. - Added `pnpm exec playwright test --config tests/perf/issue-detail/playwright.config.ts`, which seeds an isolated instance and runs N≥5 warm/cold samples under unthrottled and Fast 4G/4x CPU profiles. - Added Markdown, raw JSON, and Chrome-trace outputs with median baseline tables and waterfall data. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm check:token-gates` - `npx playwright test --config tests/perf/issue-detail/playwright.config.ts --list` - `pnpm exec playwright test --config tests/perf/issue-detail/playwright.config.ts` — passed 20 samples in 9.4 minutes (5 runs × 2 scenarios × 2 profiles) for the baseline; post-review integrity reruns also exercised the corrected paths, while this shared runner intermittently killed Chromium processes, so the rig now performs one bounded browser-crash retry per sample. - Baseline medians: warm unthrottled 278/447 ms header/content; cold unthrottled 646/646 ms; warm throttled 1240/2060 ms; cold throttled 3932/3933 ms. ## Risks - Low product risk: the new browser measurements are development/QA tooling and the UI timing work does not change visible layout. - `Server-Timing` exposes only aggregate handler duration, not query contents or private identifiers. - Native INP reporting uses supported browser event timing entries and silently no-ops where unsupported. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.4, tool-assisted coding and browser execution with reasoning enabled; context-window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dev Agent <dev@paperclip.ing>canary/v2026.811.0-canary.15 |
||
|
|
5bb2490b86 |
feat(ui): surface all issue documents and agent artifacts in chat-style sidebar (#11226)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The task detail page uses a chat-style thread with a right sidebar; the sidebar has a Plan tab and an Artifacts tab (#11101 made this UI the default) > - The Plan tab only showed the one issue document named `plan`, and the Artifacts tab only listed formal work products; other agent-authored documents (for example a `synthesis` doc) and agent-attached files were invisible in the sidebar > - Users could see an agent mention a document in the thread but had no way to find that document in the sidebar, which breaks trust in the task view as the record of the work > - This pull request surfaces every non-system issue document in the Plan tab, composes the Artifacts tab from work products, documents, and agent-created attachments, and gives thread images a full-screen lightbox with download > - The benefit is that anything an agent produces on a task is now reachable from the sidebar, while user uploads stay with their comments in the thread ## Linked Issues or Issue Description Refs #11101 (chat-style task UI default — this PR extends its sidebar). **Subsystem affected** Task detail UI (chat-style thread sidebar): Plan tab, Artifacts tab, and thread attachment rendering in `ui/src`. **Current behavior** The Plan tab renders only the issue document literally named `plan`. The Artifacts tab renders only formal work products. Agent-authored documents with any other name, and files agents attach to comments, do not appear anywhere in the sidebar. Thread images open as bare links. **Proposed behavior** The Plan tab lists every non-system issue document, with the `plan` document first and the others rendered inline below it. The Artifacts tab composes three sources — work products, issue documents, and agent-created comment attachments — deduplicated against attachment-backed work products via `metadata.attachmentId`, and shows whenever any source is non-empty. Work-product rows without a resolvable attachment or document fall back to links found in their metadata so they stay clickable. Images in the thread open a shared full-screen lightbox with a download action. Files uploaded by users stay thread-only and are not mixed into the Artifacts tab. **Reason and benefit** Agents routinely produce documents that are not named `plan` and attach files to their comments. Users reading the thread must be able to find every one of those outputs from the sidebar. Redundant surfacing is acceptable; an unfindable document is not. **Breaking changes** None. This is additive rendering; no schema or API changes. ## What Changed - `IssuePropertiesPlansTab.tsx`: renders all non-system issue documents, `plan` primary, others inline below via `MarkdownBody` - `IssuePropertiesArtifactsTab.tsx`: composes work products + documents + agent-created attachments with dedupe; rows without an attachment/document target fall back to `metadata` links - `IssueProperties.tsx`: Artifacts tab visibility now derives from the composed source set - New `ui/src/lib/issue-artifacts.ts`: pure composition/dedupe logic, unit-tested - New `ui/src/components/task-chat/task-chat-attachments.ts`: splits agent vs user comment attachments, unit-tested - `TaskChatBubble.tsx`: thread images open the shared full-screen lightbox with download - `useIssueDocuments.ts`: hook now exposes the full issue-document list ## Verification - `pnpm typecheck` — passes across the workspace - `pnpm check:token-gates` — 3/3 CLEAN - `cd ui && pnpm vitest run src/lib/issue-artifacts.test.ts src/components/task-chat/task-chat-attachments.test.ts src/pages/IssueDetail.test.tsx` — 74 tests pass - Manual: open a task whose agent created a document not named `plan` (for example `synthesis`); confirm it appears in the Plan tab below the plan and in the Artifacts tab; confirm an image the agent attached appears under Artifacts; confirm a user-uploaded image stays only in the thread and opens full screen with a download button Snapshot baselines are intentionally not updated for this visual change, per the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks - Low risk: rendering-only change scoped to the task sidebar and thread bubbles; composition logic is pure and unit-tested - Dedupe relies on `metadata.attachmentId` linkage; a work product with malformed metadata would render as a duplicate row (cosmetic only) ## Model Used - Claude (Anthropic), model id `claude-fable-5`, extended thinking enabled, agentic tool use via Claude Agent SDK (Claude Code harness) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0044fa8904 |
Let tenants edit env vars on managed sandbox environments; add managed-sandbox-only mode (#11200)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments give each agent run an execution target: the local host, SSH, or a sandbox provider > - A managed deployment can provision one platform-managed sandbox environment through the `PAPERCLIP_MANAGED_CONFIG` `environments` section > - That row is fully locked today. A tenant cannot add environment variables for their agents. There is also no way to hide local execution — run selection falls back to the local row > - A platform that manages the sandbox for its tenants needs both: the tenant adds env vars (and nothing else), and local execution is neither visible nor reachable > - This pull request opens exactly one tenant edit (env vars) on the managed sandbox row, and adds an `enableManagedSandboxOnly` mode that hides local and makes run selection fail closed > - The benefit is a complete managed-sandbox experience with no change for self-hosted instances ## Linked Issues or Issue Description **Subsystem affected** Environments (managed sandbox provisioning, environment routes, run environment selection) and the environments UI. **Problem or motivation** Platform-provisioned sandbox environments (`metadata.managedByPaperclip`) reject every write on cloud-managed instances. Agents often need environment variables inside their sandbox. The tenant has no way to set them on the managed row. Separately, an operator cannot remove local execution: the environment list always shows the local row, and run selection falls back to it when no default is set. **Proposed solution** Allow an envVars-only PATCH on the managed sandbox row, and echo those env vars back for editing. Add a managed-tier feature (`enableManagedSandboxOnly`) that hides the local environment from all read surfaces and redirects local-landing run selection to the managed sandbox environment, failing closed when it is unavailable. **Alternatives considered** UI-only hiding of the local row. This was rejected: it does not stop a run from resolving to local, so it is presentation without enforcement. Full unlock of the managed row was also rejected: name, driver, and config stay platform-owned so boot reconciliation cannot fight tenant edits. ## What Changed - `server/src/routes/environments.ts`: the platform-provisioned write floor admits an envVars-only PATCH on the generalized managed sandbox row (sandbox driver, `managedByPaperclip`, not legacy kubernetes-marker rows). Name, driver, config, status, metadata, and DELETE stay rejected. The read floor stops blanking env vars on that row; credential-shaped config keys stay redacted for every actor. Legacy kubernetes-marker rows keep the full floor. - Same file: under `enableManagedSandboxOnly`, the environments list and the by-id read omit the local row for every actor, including instance admins. - `server/src/services/execution-workspace-policy.ts`: `resolveExecutionWorkspaceEnvironmentId` gains the managed-sandbox-only inputs. A selection that lands on the local environment is redirected to the managed sandbox environment. With no active managed row it throws `ManagedSandboxUnavailableError` — never local. Non-local selections (ssh, user-created sandboxes) are untouched. - `server/src/services/heartbeat.ts`: the run path reads the flag, looks up the managed row (`findManagedSandboxEnvironment`, new read-only finder in `environments.ts`), and passes both to the resolver. Mirrors the forced-kubernetes precedent, which keeps precedence when both regimes are on. - `server/src/services/managed-environments.ts`: after a successful reconcile, the instance default environment moves to the managed sandbox row when the current default is unset, local, or dangling. A tenant-chosen custom environment is never overridden. - `packages/shared`: new `enableManagedSandboxOnly` key (schema default false, catalog tier `managed`, cloudDefault false, selfHostedDefault false) and the matching interface field. - UI: managed rows show a "Managed by Paperclip" lock badge; editing one opens a dedicated env-vars-only editor that sends the one PATCH shape the server admits (the old full form failed with a 403 on save). New `ui/src/lib/managed-sandbox-environment.ts` mirrors the local filter for cached lists (applied in the project picker; the agent picker already excluded local). The experimental settings page gains the toggle at its alphabetical card position. `environmentsApi.update` now declares the `envVars` field it already sent. - Tests: environment route floor coverage (envVars-only accepted, mixed bodies rejected, legacy rows still blanked and locked, local hidden and 404 under the flag, self-hosted unchanged), an embedded-postgres service test pinning that boot reconciliation never touches tenant env vars, resolver redirect/fail-closed cases, managed-environments default-stamping cases, and UI lib/settings tests. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/environment-routes.test.ts src/__tests__/environment-service.test.ts src/__tests__/execution-workspace-policy.test.ts src/services/managed-environments.test.ts` — all pass. - `pnpm --filter @paperclipai/shared exec vitest run` — 425 pass (catalog/schema default parity is pinned by an existing test). - `pnpm --filter @paperclipai/ui exec tsc --noEmit` and the affected UI suites (CompanyEnvironments, InstanceExperimentalSettings incl. card-order test, new lib test) — all pass. - Full workspace `pnpm test`: 3,414 passed. 17 files report failures on this machine; the identical 17 fail on a clean `origin/master` worktree in the same environment (git-worktree/skills/embedded-postgres environment dependencies and plugin-SDK zero-test collections). One additional file (`issue-monitor-scheduler.test.ts`) failed one timing-sensitive test in one of two full-suite runs and passes 7/7 in isolation on this branch — a flake in a domain this diff does not touch. The branch introduces no new failures. - Self-hosted zero-delta: every new behavior is gated on the cloud-managed instance check or the new flag, which defaults to false in schema and catalog; pinned by the "does not floor platform-marked rows on self-hosted instances" and flag-off tests. ## Risks - Behavior is opt-in twice over: the write-floor exception applies only to rows the managed-config provisioner stamps, and the hiding/forcing applies only when `enableManagedSandboxOnly` is on (default false everywhere). Self-hosted instances see no change. - The env-vars echo is scoped to the generalized managed sandbox row; legacy kubernetes-marker rows keep the blanket floor because pre-generalization builds may have written platform values there. - Fail-closed run selection means a managed instance with the flag on and an archived managed row (provider plugin down) refuses runs with a precise error instead of running locally. That is the intended posture. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.14 |
||
|
|
b847e8b6f6 |
perf(server): reduce issue detail request overhead (#10414)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent work > - Opening an issue fans out into several authenticated issue-detail reads, so repeated work on that path directly affects perceived latency > - Those reads repeated issue and authorization lookups, returned full private JSON even when unchanged, and performed non-critical bookkeeping writes on the request path > - Interaction reads also performed lifecycle writes even though `GET` must be read-only > - This pull request adds request-scoped reuse, private conditional responses, read-only interaction access, and bounded write debouncing without crossing actor, request, or company boundaries > - The result is less database, serialization, logging, and response-body work while preserving authorization and interaction lifecycle invariants ## Linked Issues or Issue Description This is the server-only latency phase. Related work is tracked separately in #10415 (aggregate view), #10416 (warm navigation, merged into the base), and #10463 (bundle split). This pull request intentionally excludes those scopes. **What happened?** Opening an issue detail view caused avoidable server costs: repeated issue and authorization reads within one request, full private JSON responses when a representation was unchanged, writes during interaction-list reads, production debug transport setup, and immediate bookkeeping writes for cloud tenant activity and board-key usage. **Expected behavior** All successful JSON `GET /api/issues/:id/*` responses should support strong private ETags and `304 Not Modified`. Repeated work may be reused only within the current request. `GET /interactions` must not modify stored interactions. Non-critical activity timestamps may be debounced without weakening authentication or stale instance-admin cleanup. **Steps to reproduce** 1. Start Paperclip in local development or self-hosted server mode. 2. Open one issue and request its detail subresources with the same authenticated actor. 3. Repeat a successful JSON request with its `ETag` in `If-None-Match`. 4. Observe `304 Not Modified`, no interaction writes from `GET /interactions`, and unchanged authorization boundaries. **Deployment mode / installation** - Local development or self-hosted server - Built from source - Core server behavior; not adapter-specific ## What Changed - Added strong ETags and `Cache-Control: private, must-revalidate` to successful JSON reads under `/api/issues/:id/*`, including standards-compliant `If-None-Match` handling. - Added request-scoped promise memoization for issue and authorization lookups; no authorization result survives the request. - Made `GET /interactions` read-only, moved supersession and terminal-state handling to mutation paths, and prevented plugin callers from accepting or rejecting interactions after an issue closes. - Removed the production debug-file logger transport while preserving development formatting. - Debounced cloud-tenant activity and board-key `lastUsedAt` persistence, while keeping stale instance-admin deletion unconditional and authentication checks per request. - Added focused tests for ETags, request isolation, authorization lifecycle behavior, interaction invariants, logger configuration, and retry-safe debounce behavior. ## Verification - `pnpm exec vitest run server/src/__tests__/private-json-etag.test.ts server/src/__tests__/issue-thread-interaction-routes.test.ts` — 2 files, 23 tests passed. - Focused Vitest run covering request memoization, authorization, interactions, plugin orchestration, logger, cloud tenant, board auth, and issue services — 9 files, 264 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check origin/master...HEAD` — passed. - Scope guardrails: 21 changed files under `server/src`; no lockfile, workflow, migration, UI, aggregate-view, or bundle-split changes. ## Risks - Strong ETags hash each successful serialized JSON response. This adds a small CPU cost but avoids transferring unchanged bodies. - Debounced bookkeeping timestamps can lag by the bounded debounce interval. They are non-critical usage metadata; authentication still runs per request, and stale instance-admin deletion remains unconditional. - Legacy pending interactions on terminal issues are projected as expired by reads and are finalized only by mutation paths. The stored record remains unchanged on `GET` by design. - No database schema or migration changes are included. > This is a focused performance correction and does not duplicate a planned core feature in `ROADMAP.md`. ## Model Used OpenAI Codex using `gpt-5.3-codex` for the initial implementation and `gpt-5.6-sol` for isolation, verification, and PR preparation, with reasoning, repository tool use, code execution, and GitHub CLI access. The runtimes did not expose authoritative context-window sizes. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
145d86911b |
Remove decision training UI (#11225)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Decisions desk shows work that needs an operator response. > - It also exposed decision-training actions and a separate training library. > - Paperclip does not plan to use these training surfaces now. > - Keeping inactive controls makes the Decisions workflow harder to scan. > - This pull request removes the training UI and keeps the backend snapshot contract unchanged. > - The benefit is a smaller and clearer Decisions workflow without a data migration. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Decisions desk currently exposes training controls, training state, and a separate training library route. **Subsystem affected** `ui/` — React and Vite board UI. **Current behavior** Operators can open a training library from the Decisions toolbar. They can also mark a decision for training from rows and inspect the result in a drawer. **Proposed behavior** Remove the training controls, badges, drawer, library pages, and routes from the Decisions UI. Keep the server APIs and stored training examples unchanged. **Reason and benefit** The product does not plan to use decision training now. Removing the unused surfaces reduces Decisions UI noise and avoids presenting a workflow that operators should not use. **Breaking changes** The `/decisions/training` UI routes are no longer registered. Existing server endpoints and stored decision-training data remain compatible. ## What Changed - Removed decision-training controls and state from Decisions toolbars, rows, queue pages, and shelves. - Removed the training drawer, library, inspector, API client, helpers, query keys, and routes. - Added route and row regressions that assert training UI does not return. - Updated the Decisions Storybook description to match the available controls. ## Verification - `pnpm exec vitest run ui/src/App.test.tsx ui/src/components/AttentionQueueRow.test.tsx` — 32 tests passed. - `pnpm check:token-gates` — all gates clean. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — server and UI partitions passed. The CLI partition had one environment-only failure because this agent runtime injects static AWS credentials. The exact CLI file passed all 8 tests when those credential variables were unset. - Searched `ui/src` and `ui/storybook` for the removed training routes, drawer, library, badges, and actions. Only negative regression assertions remain. ## Risks - Low implementation risk. This change deletes UI-only entry points and does not change the database or server APIs. - Saved training-page URLs no longer render a board route. This is the intended behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5.6-sol`, with `xhigh` reasoning. The runtime did not expose the context-window size. The agent used repository tools, shell execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.811.0-canary.13 |
||
|
|
45dfb183b6 |
fix(ui): add undo action to inbox archive toast (#11220)
## Thinking Path > - Paperclip helps operators supervise AI-agent work. > - The Mine inbox keeps tasks that need an operator's attention in one place. > - Operators can archive a task from its detail page after they finish triage. > - That action is easy to select accidentally and did not offer immediate recovery. > - This pull request adds Undo to the archive success toast and keeps inbox caches consistent. > - The benefit is fast recovery without searching for or reopening the task. ## Linked Issues or Issue Description **What happened?** Archiving a task from the Mine inbox removed it and showed a success toast with no recovery action. **Expected behavior** The success toast should offer Undo. Selecting Undo should restore the task through the existing unarchive API while preserving a consistent inbox view. **Steps to reproduce** 1. Open a task from the Mine inbox. 2. Select the archive action. 3. Observe that the task leaves the inbox and the success toast has no Undo action. **Paperclip version or commit** Reproduced on `master` before this change. **Deployment mode** Local dev, built from source. Related prior work: #9931 and #10668. ## What Changed - Add an Undo action to the successful inbox archive toast. - Optimistically restore the task in captured inbox query caches before the unarchive request completes. - Cancel in-flight inbox fetches and clear the local archive guard so stale responses cannot hide the restored task. - Reapply the archive guard and remove the cached task if the unarchive request fails. - Add regression tests for successful Undo, the in-flight cache race, and failed Undo rollback behavior. ## Verification - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx ui/src/lib/inboxArchiveCache.test.ts` — 51 tests passed on the final head. - `pnpm check:token-gates` — passed. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — the general server and UI groups passed. One CLI doctor assertion detected injected host AWS credentials and passed all 8 tests with those unrelated variables unset. A task-watchdog scheduler test also passed all 18 tests in isolation after one full-suite timing failure. ## Risks Low risk. The change uses the existing unarchive endpoint and inbox cache helpers. Undo failure returns the task to its archived state, shows an error toast, and invalidates the inbox queries for server reconciliation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, model ID GPT-5. The service manages the context window. Reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7734f4b32d |
fix(ui): keep slash autocomplete scrollable in dialogs (#11222)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip helps operators manage AI-agent companies. > - Operators create tasks and comments through shared rich-text editors. > - These editors show slash-command and mention matches in a floating menu. > - Modal dialogs treat that body-level menu as outside content and cancel its wheel and touch movement. > - This pull request keeps scroll events inside the floating menu and preserves native scrolling. > - The benefit is that operators can reach every match with a mouse wheel, a trackpad, or a touch screen. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. **What happened?** Slash-command and mention menus could contain more matches than their visible height. When an editor was inside a modal dialog, the modal scroll lock canceled wheel and touch movement on the body-level menu portal. Operators could not scroll to later matches. **Expected behavior** The autocomplete menu must scroll with a mouse wheel, a two-finger trackpad gesture, and a vertical touch gesture. Keyboard selection and normal editor behavior must stay unchanged. **Steps to reproduce** 1. Open a task or comment editor inside a modal dialog. 2. Enter a slash command or mention query that has more matches than the menu can show. 3. Try to scroll the menu with a wheel, trackpad, or touch gesture. **Paperclip version or commit** `7ea2068ef8` on `master`. **Deployment mode** Local development UI built from source. ## What Changed - Keep wheel and touch movement inside the shared autocomplete menu portal. - Add vertical overscroll containment while preserving native momentum scrolling. - Add a regression test that mounts the real dialog and verifies that wheel and touch movement stay uncanceled. ## Verification - `pnpm --dir ui exec vitest run src/components/MarkdownEditor.test.tsx` - `pnpm check:token-gates` - `pnpm -r typecheck` - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run` - `pnpm build` The two AWS variables are omitted from the full test command because this agent runtime injects static AWS credentials. One unrelated CLI doctor test correctly warns when those credentials are present. The CI environment does not inject them. ## Risks - Low risk. Event propagation stops only on the open autocomplete menu portal. - Ancestor listeners no longer receive wheel or touch movement from that menu. Native menu scrolling and option-level touch handling still receive the events. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with model ID `gpt-5`. The deployment suffix and context-window size are not exposed to the agent. The model used agentic reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.811.0-canary.12 |
||
|
|
3e1ea39ff3 |
fix(inbox): honor saved policy for explicit targets (#11221)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100): short sentences, one instruction per sentence, simple
approved vocabulary, and the active voice. -->
## Thinking Path
> - Paperclip is the open source control plane people use to manage
AI-agent companies and their work
> - Each user can let agents tidy that user's Mine inbox
> - The profile control saves either an open policy or an agent
allowlist
> - Explicit inbox archive requests checked only the separate
`inbox:manage` grant
> - This made the saved profile control ineffective for explicit user
targets
> - This pull request makes authorization honor the target user's saved
policy
> - The benefit is that the UI control and the API now enforce the same
user choice
## Linked Issues or Issue Description
**What happened?**
An agent received `403 inbox_cross_user_grant_required` when it archived
an issue with an explicit `userId`. The denial occurred even when that
user had enabled inbox management for the agent in Profile Settings. The
authorization service checked only `principal_permission_grants` for
explicit targets and ignored the saved user inbox policy.
**Expected behavior**
An explicit target is allowed when the target user saved an `open`
policy or an allowlist that contains the agent. An unsaved default-open
policy must remain limited to the responsible-user path. A scoped
`inbox:manage` grant must remain an administrative override.
**Steps to reproduce**
1. Save an inbox-agent allowlist for a user.
2. Include the acting agent in that allowlist.
3. Call `POST /api/issues/{issueId}/inbox-archive` with that user's
explicit `userId`.
4. Observe the incorrect `403 inbox_cross_user_grant_required` response
on the previous implementation.
**Paperclip version or commit**
Reproduced on `7ea2068ef8`.
**Deployment mode**
Self-hosted server.
**Installation method**
Built from source with pnpm.
**Agent adapter(s) involved**
Not adapter-specific. This is a core authorization bug.
**Database mode**
External Postgres in production. The regression tests use embedded
PostgreSQL.
**Access context**
Agent bearer authentication.
Related foundations: #9658 and #9724.
## What Changed
- Read the target user's saved inbox-agent policy before the
explicit-target decision.
- Allow saved `open` policies and matching allowlists for explicit
targets.
- Keep unsaved implicit-open policies responsible-user-only.
- Keep scoped `inbox:manage` grants as administrative overrides.
- Add service and route regressions for allow, deny, archive, unarchive,
and audit metadata.
- Update the implementation contract and agent-facing inbox API
guidance.
## Verification
- `pnpm exec vitest run
server/src/__tests__/authorization-service.test.ts
server/src/__tests__/inbox-archive-routes.test.ts` — 66 tests passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `git diff --check origin/master...HEAD` — passed.
## Risks
- Low risk. The change is limited to explicit inbox targets with a saved
policy.
- A missing policy row still denies explicit cross-user access.
- A non-matching allowlist and a disabled policy still deny access
unless a scoped administrative grant applies.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- OpenAI Codex based on GPT-5. The runtime did not expose the exact
model build or context-window size. The agent used reasoning, repository
tools, code execution, and focused test execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.811.0-canary.11
|
||
|
|
71e9d6bb0b |
feat(release): candidate-branch beta builds and the release checklist (#11209)
> Follow-up to #11208 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channels promote artifacts along canary → nightly → beta → stable, with the happy path being promotion of an existing build > - When one or two targeted fixes are needed before a beta or stable, the only options today are waiting for the next nightly or absorbing a whole day of unrelated master changes > - The channel model was designed with an escape hatch for exactly this: short-lived candidate branches carrying only cherry-picked fixes > - This pull request implements candidate-branch beta builds with full verification, documents the stable fix path through the soak-justification gate, and adds the release captain's checklist > - The benefit is that a surgical fix can ship forward without either delay or blast radius, with its provenance recorded ## Linked Issues or Issue Description Refs #11008 — completes the fix-path half of the channel model introduced there. **Subsystem affected** Release automation: `scripts/release.sh`, `.github/workflows/release.yml`, `doc/RELEASING.md`, new `doc/RELEASE-CHECKLIST.md`, tests. **Problem or motivation** Beta promotion only accepts commits that already shipped as a nightly, and stable promotion expects a soaked beta. There is no supported way to ship one or two cherry-picked fixes between lanes: an urgent fix must wait for the nightly cycle or pull in every unrelated master change from the day. The original channel design called for candidate branches to cover this, and they were deferred from the initial implementation. **Proposed solution** Candidate-branch beta builds: cut `candidate/beta-<target>` from a nightly's source commit, cherry-pick the fixes, and dispatch `channel: beta` with the new `candidate_branch` input. Selection enforces the naming convention, rejects heads that already shipped as a beta or predate the candidate tooling, and records the cherry-picked commits in the job summary. Because candidate heads never went through a canary or nightly, publication is gated on a full `release-verify` run (promoted nightlies keep skipping re-verification). The stable fix path (`candidate/release-<target>` as `source_ref`) works through the existing soak gate: the justification requirement is the deliberate, recorded trade-off for shipping unsoaked bits, and is now documented as such. ## What Changed - `scripts/release.sh`: `--from-candidate` flag (beta only) waives the shipped-a-nightly requirement while keeping the duplicate-beta guard - `.github/workflows/release.yml`: `candidate_branch` dispatch input; candidate mode in `select_beta` (naming validation, duplicate and tooling-era rejection, cherry-pick recording); new `verify_beta_candidate` job gating candidate publishes on full verification - `doc/RELEASING.md`: beta fix-path and stable fix-path sections - `doc/RELEASE-CHECKLIST.md` (new): the release captain's checklist for all four lanes as built - Tests: dry-run fixture coverage for `--from-candidate` (waives the nightly guard, keeps the duplicate guard, rejected outside beta) and wiring tests for candidate validation plus the verification gate ## Verification - `node --test` on the four affected suites: 42 pass in total (17 + 25 across the two runs), including the 5 new tests - `bash -n` on `release.sh`; YAML parse of the workflow - After merge: exercise the path end to end the first time a real cherry-picked beta is needed — dispatch with a `candidate/beta-*` branch and confirm the summary records the picks and verification runs ## Risks - Candidate builds bypass the smoke-tested-nightly provenance by design; the compensating controls are full verification before publish, the post-publish beta smoke, the human `npm-beta` gate, and recorded cherry-picks - The stable fix path rides the existing justification mechanism rather than adding a second bypass — one recorded escape hatch, not two ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.10 |
||
|
|
7ea2068ef8 |
fix(files): only highlight accessible workspace file links (#11090)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task comments can contain references to files in project and execution workspaces. > - Paperclip detected path-shaped inline code and showed it as an actionable file chip. > - The UI did not first confirm that the current board session could open the file. > - Missing, denied, ambiguous, remote, and unsupported files therefore looked actionable and failed after a click. > - This pull request adds an issue-scoped availability check and promotes only confirmed files to chips. > - The benefit is that the task thread shows a file action only when that action can succeed. ## Linked Issues or Issue Description **What happened?** Task comments promoted path-shaped inline code to file chips before Paperclip checked the file. A chip could point to a missing, denied, ambiguous, remote, or non-previewable file. The action then failed after the user selected it. **Expected behavior** Paperclip must show a file chip only after the server confirms that the current board session can open the exact file reference. All other path-shaped text must stay ordinary inline code. **Steps to reproduce** 1. Add a task comment that contains inline code with a missing or inaccessible workspace path. 2. Open the task thread as a board user. 3. Observe that the path looks like an actionable file chip. 4. Select the chip and observe that the file cannot open. **Paperclip version or commit** `19be4cf927` and earlier. **Deployment mode** Local dev and self-hosted server. **Access context** Board user. ## What Changed - Added shared request, response, and validation contracts for batched workspace-file availability checks. - Added an issue-scoped server endpoint that resolves file references with company, issue, workspace, and preview-access checks. - Added bounded batch concurrency and tests for missing, denied, ambiguous, remote, unsupported, and available files. - Added an issue-scoped UI availability registry that deduplicates, batches, caches, and invalidates file checks. - Changed task-comment markdown rendering so only confirmed files get chip styling and file-viewer behavior. - Bound each chip to the exact workspace target that passed the availability check. ## Verification - `pnpm exec vitest run packages/shared/src/workspace-file-resource.test.ts server/src/__tests__/file-resources.test.ts ui/src/components/MarkdownBody.test.tsx ui/src/components/WorkspaceFileMarkdownBody.availability.test.tsx ui/src/lib/remark-workspace-file-refs.test.ts ui/src/lib/workspace-file-availability.test.ts` — 93 passed, 35 skipped. - `pnpm check:token-gates` — clean. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — all server and UI groups passed. One unchanged CLI test saw the run-injected static AWS credentials and expected only its local `AWS_PROFILE`. The same test passed, 8 of 8, after removing only `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` from its process environment. ## Risks - File chips now appear after an asynchronous availability check, so path-shaped text can briefly render as inline code. - Availability results use the existing 30-second file-resource cache window. File-resource invalidation forces a new check. - The endpoint limits each request to 100 references and the client chunks larger sets. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The service did not expose a more specific model ID or context-window size. The agent used high-reasoning mode, repository tools, command execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.811.0-canary.9 |
||
|
|
815e49bb7c | feat: make chat-style tasks the default experience (#11101) | ||
|
|
2da6a248c3 |
fix(release): surface recovery commands when a lane tag push is rejected (#11208)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's promotion lanes publish to npm, then push a lane tag and dispatch the Docker image build at that tag > - The first nightly of the beta-tooling merge published to npm and then died at the tag push: GITHUB_TOKEN may not create refs pointing at workflow-modifying commits from dispatch or scheduled runs > - The failure was a bare `remote rejected` with no guidance, leaving the release half-finished (npm live, no tag, no images) until an operator reverse-engineered the recovery > - This pull request makes every lane's tag push degrade into exact recovery instructions in the job summary > - The benefit is that a rare platform-permission rejection becomes a two-minute runbook operation instead of a forensic exercise ## Linked Issues or Issue Description Refs #11008 — the incident occurred promoting that change's own merge commit, the first workflow-modifying commit to flow through the lanes it introduced. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31445344811 published `2026.811.0-nightly.0` to npm, then failed pushing `nightly/v2026.811.0-nightly.0`: `refusing to allow a GitHub App to create or update workflow .github/workflows/release.yml without workflows permission`. The tagged commit modifies workflow files, and GITHUB_TOKEN may not create refs pointing at such commits from dispatch or scheduled runs (push-event runs are exempt, which is why the canary tag on the same commit succeeded). The job failed with no explanation and the Docker dispatch never ran. **Proposed solution** Wrap the nightly, beta, and stable tag pushes: on rejection, write the exact recovery commands into the job summary — create and push the tag with maintainer credentials, dispatch `docker.yml` at the tag, and for stable also run `create-github-release.sh` — then fail the job. Document the cause and recovery in the failure playbooks and pin the three recovery blocks with a wiring test. ## What Changed - `.github/workflows/release.yml`: recovery-summary wrappers on the nightly, beta, and stable tag-push steps - `doc/RELEASING.md`: failure-playbook entry for the workflows-permission rejection - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting all three lanes carry the recovery summary ## Verification - Wiring tests: 6 pass; YAML parse of the workflow - The recovery commands are exactly the ones used to resolve the real incident (tag push + `docker.yml` dispatch for `nightly/v2026.811.0-nightly.0`) ## Risks - Low. The happy path is unchanged (a successful push skips the wrapper); the failure path trades a bare error for actionable output and still fails the job, since the release state is genuinely incomplete ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.8 |
||
|
|
1ea2f0e2d6 |
feat(cli): add 'paperclipai channels' to show release lanes and the current one (#11210)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channels (canary → nightly → beta → stable) select by install target, and users discover them today only through maintainer-oriented docs > - A user who wants to know "which lane am I on, and what else is there" has no self-serve answer > - The channel rollout planned a read-only CLI command for exactly this > - This pull request adds `paperclipai channels`: every lane with the version its dist-tag resolves to, the install command for it, and which lane the running install follows > - The benefit is self-serve lane discovery without reading release documentation ## Linked Issues or Issue Description Refs #11008 — the user-facing discovery surface for the channel model completed there. **Subsystem affected** CLI: `cli/src/commands/channels.ts` (new), `cli/src/index.ts`, `doc/CHANNELS.md`, tests. **Problem or motivation** Channel selection is install-based (`@latest` / `@beta` / `@nightly` / `@canary`), but nothing in the product tells a user which channel their install follows or what the other lanes currently resolve to. The information lives in `doc/CHANNELS.md` and the npm registry, neither of which a running install surfaces. **Proposed solution** A read-only `paperclipai channels` command: prints each channel with the version its dist-tag currently resolves to (per-lane registry lookups that degrade to `unavailable` individually), the install command for each, and the running install's lane parsed from its version suffix — source checkouts carry the repository's placeholder version and are reported as unmapped rather than guessed. `--json` emits the same data for scripting. ## What Changed - `cli/src/commands/channels.ts` (new): channel table, lane parsing, registry resolution, human and `--json` output - `cli/src/index.ts`: registers `channels` - `doc/CHANNELS.md`: "Seeing where you are" section - `cli/src/__tests__/channels.test.ts` (new): lane parsing including unknown versions, full resolution against a fake runner, per-lane degradation, table/dist-tag sync ## Verification - `vitest run cli/src/__tests__/channels.test.ts`: 5 pass - `pnpm typecheck` in `cli/` - Live run against the real registry shows all four lanes with their current versions (`2026.722.0` / `2026.811.0-beta.0` / `2026.811.0-nightly.0` / canary) and correctly reports a source checkout as unmapped ## Risks - Low. Read-only command reusing the existing `resolvePublishedVersion` registry helper; no state, no auth, no publish surface ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.7 |
||
|
|
c4abecb2c4 |
fix(skills): refresh project folders in place (#11066)
## Thinking Path > - Paperclip helps operators manage agent skills across a company. > - The installed skills view groups project-backed skills into folders. > - The view showed two folder creation controls and only offered a global project scan. > - Operators need one clear folder action and a refresh action for the selected project. > - This pull request keeps folder creation in the folder rail and adds a scoped project refresh. > - The benefit is a calmer skills view and faster, more precise project skill updates. ## Linked Issues or Issue Description No public GitHub issue exists for this focused UI bug. **What happened?** The installed skills view repeated the folder creation action in the toolbar. A selected project folder also had no way to refresh only its own project skills. **Expected behavior** The folder rail must own folder creation. A selected project-backed folder must offer a refresh action that scans only that project and refreshes the skill and folder queries. **Steps to reproduce** 1. Open the installed skills view for a company with project-backed skill folders. 2. Select a project folder. 3. Observe the duplicate folder action and the absence of a project-scoped refresh action. **Paperclip version or commit** Reproduced before this two-commit fix on `master`. **Deployment mode** Local development with `pnpm dev`. ## What Changed - Removed the duplicate toolbar folder creation button when the folder rail exists. - Preserved the toolbar folder action when no folder rail exists. - Added a refresh action beside the breadcrumb for a selected project-backed folder. - Passed the selected project ID to the project scan API. - Refreshed both the installed skill list and skill folder data after scans. - Added component tests for compact folder creation, the empty-folder fallback, and scoped project refresh. ## Verification - `pnpm exec vitest run ui/src/pages/CompanySkills.test.tsx` — 20 tests passed. - `pnpm check:token-gates` — passed with all three gates clean. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — the server and UI stages passed 7,475 tests. The CLI stage then found one environment-sensitive AWS doctor assertion because this agent runtime injects static AWS credentials. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts --project paperclipai` — all 8 tests passed. - GitHub CI — all latest-head checks passed. ## Risks - Low risk. The scoped refresh depends on the existing `project:<id>` folder system key. - The global scan path is unchanged. - There are no schema, migration, API contract, dependency, workflow, or documentation changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 family. The runtime did not expose a more specific model ID or context-window size. The agent used high-reasoning mode, repository tools, GitHub tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.811.0-canary.6 |
||
|
|
b58ce27a02 |
fix: isolate execution workspace summaries (#10790)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip gives operators a summary for each workspace. > - An execution workspace detail page used the parent project-workspace summary slot. > - Two execution workspaces under one project workspace could therefore show the same summary. > - This pull request gives each execution workspace its own summary scope. > - It also limits the summary snapshot and generated issue to that execution workspace. > - The benefit is that a new or parallel execution workspace cannot inherit unrelated status. ## Linked Issues or Issue Description **What happened?** An execution workspace detail page read and refreshed the summary slot for its parent project workspace. Parallel execution workspaces could show the same status and include issues from each other. **Expected behavior** Each execution workspace must have one isolated summary slot. Its generated snapshot must include only issues assigned to that execution workspace. **Steps to reproduce** 1. Create two execution workspaces under one project workspace. 2. Add different issues to each execution workspace. 3. Generate the summary in the first execution workspace. 4. Open the second execution workspace. 5. Observe that the old implementation could reuse the first summary. **Paperclip version or commit** The problem exists on `master` before this pull request. **Deployment mode** The issue affects both local trusted and authenticated deployments. ## What Changed - Added `execution_workspace` to the shared summary-slot scope contract. - Validated execution-workspace ownership and stored generated summary issues on the correct execution workspace. - Limited execution-workspace snapshots to issues with the matching execution workspace ID. - Updated the execution workspace page to use its own summary slot. - Updated Summarizer instructions, routine options, catalog metadata, documentation, and regression tests. ## Verification - `NODE_ENV=test pnpm exec vitest run packages/shared/src/summary-slot.test.ts server/src/__tests__/summary-slots.test.ts ui/src/pages/ExecutionWorkspaceDetail.test.tsx` — 30 focused tests passed; the embedded-Postgres server tests were run outside the process-restricted sandbox. - `pnpm check:token-gates` — passed. - `pnpm --filter @paperclipai/skills-catalog validate` — passed with 17 catalog skills. - [Latest-head GitHub Actions](https://github.com/paperclipai/paperclip/actions/runs/31491475405) — all 22 jobs passed on `beea14cbaf`, including typecheck, build, server/workspace tests, serialized suites, e2e, canary, and aggregate verification. One unrelated adapter cleanup test initially hit an `ENOTEMPTY` temp-directory race; its single permitted rerun passed. - Greptile — 5/5 confidence on `beea14cbaf`, 12 files reviewed, zero comments added, and zero unresolved threads. ## Risks - Low risk. The new scope is additive. - Existing project and project-workspace summary slots keep their current keys and behavior. - A summary generated for an execution workspace now excludes sibling workspace issues by design. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The deployment does not expose a more specific model ID or context-window value. It used agentic reasoning, repository tools, code execution, and GitHub tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.811.0-canary.5 |
||
|
|
9cdaa5416e |
fix(ui): remember folded inbox subtasks (#11069)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The inbox helps operators scan parent tasks and their sub-tasks > - Operators can fold a parent task to hide its sub-tasks > - The inbox previously forgot that fold state after a page refresh > - This pull request stores the fold state for each company and restores it when the inbox loads > - The benefit is that the inbox keeps the operator's chosen task layout across page refreshes ## Linked Issues or Issue Description **What happened?** The inbox reset every folded parent task after a page refresh. This made all nested sub-tasks visible again. **Expected behavior** The inbox must keep each folded or unfolded parent state after a page refresh. The state must remain separate for each company. **Steps to reproduce** 1. Open the inbox with parent and child tasks. 2. Fold one parent task. 3. Refresh the page. 4. Observe that the child task is visible again without this fix. **Paperclip version or commit** Current `master` before this pull request. **Deployment mode** Local dev and built-from-source deployments. ## What Changed - Added company-scoped local storage helpers for collapsed inbox parent IDs. - Restored the stored parent fold state when the inbox mounts or the selected company changes. - Saved both direct toggle changes and explicit collapse changes. - Added helper tests and an inbox remount regression test for both folded and unfolded states. ## Verification - `pnpm exec vitest run ui/src/lib/inbox.test.ts ui/src/pages/Inbox.test.tsx` — 77 tests passed. - `pnpm check:token-gates` — all gates passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/ui build` — passed. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — 3,521 tests passed and four skipped. One unrelated server test on the current base fails because it reads `heartbeat.scheduling_suppressed` instead of `issue_commented`; the same test fails alone and this pull request changes only inbox UI files. ## Risks - Low risk. The state is local to the browser and scoped by company ID. - Old parent IDs can remain in local storage after tasks are deleted, but they do not affect visible tasks. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.6-sol`, with reasoning, tool use, and code execution. The runtime does not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.811.0-canary.4 |
||
|
|
66575fe519 |
fix(paperclip-page): scope uploader credentials to the publish helper (#10894)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents publish static pages with the paperclip-page skill and its `publish.sh` helper > - The skill docs told operators to bind the page-uploader IAM keys as the global `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` > - Static env keys have precedence over `AWS_PROFILE` in the AWS CLI and in all AWS SDKs > - Because of this, each agent run lost the host role identity and lost access to Secrets Manager and other AWS services > - This pull request adds namespaced credential variables that apply only to the helper's own `aws` calls > - The benefit is a stable host AWS identity in agent runs, with no change to page publishing ## Linked Issues or Issue Description No public GitHub issue exists. Description of the problem: **What happened?** Agent runs on a host with `AWS_PROFILE` set lost access to AWS Secrets Manager. The failures looked intermittent. The cause is deterministic: the page-uploader keys were bound as global `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` in agent run environments. These static keys shadow `AWS_PROFILE`. Each process in the agent run then used the S3-upload-only uploader identity. **Expected behavior** The page-uploader credentials apply only to the page publish helper. All other processes keep the host identity from `AWS_PROFILE`. **Steps to reproduce** 1. Set `AWS_PROFILE` to a role with Secrets Manager access. 2. Export `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` for an IAM user without that access. 3. Run `aws sts get-caller-identity`. The identity is the IAM user, not the role. 4. Run `aws secretsmanager list-secrets`. The call fails with `AccessDeniedException`. ## What Changed - `publish.sh` reads `PAPERCLIP_PAGE_AWS_ACCESS_KEY_ID` and `PAPERCLIP_PAGE_AWS_SECRET_ACCESS_KEY`, with optional `PAPERCLIP_PAGE_AWS_SESSION_TOKEN`. - The helper applies these values only to its own `aws` invocations. It clears ambient `AWS_PROFILE` and `AWS_SESSION_TOKEN` for those calls. - Credential precedence is: namespaced key pair, then `PAPERCLIP_PAGE_AWS_PROFILE`, then the ambient credential chain. Existing global-name bindings continue to work during migration. - Validation: the key pair must be set together. The pair plus `PAPERCLIP_PAGE_AWS_PROFILE` is an error. A session token without the pair is an error. - `SKILL.md` and `README.md` now instruct operators to bind the secrets under the namespaced names and explain the shadowing hazard. ## Verification - Run `node --test .agents/skills/paperclip-page/scripts/publish.test.mjs`. All 11 tests pass. - New tests cover: the incomplete key pair, the pair-plus-profile conflict, the token-without-pair error, and a fake-`aws` environment capture that proves the helper's calls see the page keys while `AWS_PROFILE` and `AWS_SESSION_TOKEN` stay unset. - Run `bash -n .agents/skills/paperclip-page/scripts/publish.sh` for a syntax check. ## Risks - Low risk. The change is contained in one skill helper and its documents. - The ambient credential chain remains the fallback, so current deployments do not break before operators rebind the secrets. - Operators must rebind the two page secrets to the namespaced names to get the benefit. The README documents this. ## Model Used Claude Fable 5 (`claude-fable-5`), Anthropic. Context window: 1,000,000 tokens (128K max output). Agentic coding session with extended thinking and tool use (Claude Code harness). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d648becb90 |
refactor(ci): split workspaces-a into two Vitest native shards
Split the slow workspaces-a CI lane into two Vitest native shards and keep release verification in parity. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.811.0-canary.3 nightly/v2026.811.0-nightly.1 |
||
|
|
6601014898 |
fix(release): reject promotion sources that predate their channel tooling (#11197)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem promotes builds along canary → nightly → beta → stable, and each publish job checks out the promotion's source commit and runs that tree's release tooling > - The first beta dispatch failed with `unexpected argument: beta`: the selected nightly's source predated the beta channel, so its `release.sh` did not know the argument > - The failure was clean (argument parsing, nothing published) but cryptic, and the same trap waits for any promotion of a source older than its target channel's tooling > - This pull request makes the selection jobs reject such sources with an actionable error and documents the property > - The benefit is that a bootstrapping or old-source promotion fails in seconds with instructions, instead of mid-publish with a parser error ## Linked Issues or Issue Description Refs #11008 — the guard hardens the beta promotion flow introduced there, after its first dispatch surfaced the gap described below. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31444045044 (first beta dispatch) failed in `publish_beta` with `unexpected argument: beta`. Promotions deliberately build from the pinned source commit, which means they also run that commit's `scripts/release.sh` — and a source that predates the target channel's introduction cannot publish it. Nothing guards this today; the error surfaces deep in the publish job with no explanation. **Proposed solution** Guard at selection time: `select_nightly` requires the source canary's `release.sh` to know the nightly channel, and `select_beta` requires the source nightly's `release.sh` to know the beta channel. Each guard literally matches the channel case arm and fails closed with a clear message naming the remedy (promote a newer source). Document the tooling-era property in `RELEASING.md` and pin the guards with a wiring test. ## What Changed - `.github/workflows/release.yml`: tooling-era guards in `select_nightly` and `select_beta` - `doc/RELEASING.md`: documents that promotions run the source commit's release tooling - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test pinning both guards ## Verification - Wiring tests: 5 pass - Guard expressions exercised against real commits: accepts the beta-capable merge commit of the beta-channel change, rejects a pre-beta commit - YAML parse of the workflow - After merge: the next beta dispatch selects a beta-capable nightly and passes the guard ## Risks - Low. Selection-time check only; the guards match the channel case arm literally and fail closed (with the same actionable message) if that line is ever reformatted ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.2 |
||
|
|
35aaaa0bd0 |
feat(server): preserve task timestamps and hierarchy through company import/export (#11193)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company export/import moves a whole company — agents, tasks, comments — between instances as a portable bundle > - The bundle never carried task timestamps or parent links: the export writes neither, the importer lets database defaults stamp "now", and sub-tasks arrive flattened > - Boards sort by recency, so every imported task showing "created just now" collapses the task list into import order, and the task hierarchy the user built is gone > - This pull request adds created/updated/started/completed/cancelled timestamps and a parent link to the bundle (schema v7), preserves them end to end on import, and keeps comment imports from clobbering a preserved updated time > - The benefit is that an imported company reads like the company the user left: same recency order, same task tree ## Linked Issues or Issue Description **What happened?** After a company import, every task showed as created at import time. Recency sorting collapsed to import order, and parent/child task nesting disappeared. The user called out losing "the meaningful task hierarchy and recency sorting". Cause: the export bundle has no fields for task timestamps or parent links, the importer lets `defaultNow()` win on insert, and the comment importer bumps every touched task's `updatedAt` to now. **Expected behavior** An imported company preserves each task's creation/update/start/completion times and its position in the task tree, so sorting and nesting on the destination match the source. **Steps to reproduce** 1. On a source instance, create tasks over several days, including sub-tasks nested under parents. 2. Export the company and import it into another instance. 3. Every task shows the import moment as its creation/update time and all tasks are top-level. ## What Changed - Export writes `createdAt`/`updatedAt`/`startedAt`/`completedAt`/`cancelledAt` (ISO, only when set) and `parent: <taskSlug>` into each task's bundle extension; a parent outside the export selection drops the edge with an aggregate warning, mirroring the existing blocker-edge warning (`server/src/services/company-portability.ts`). - Bundle schema version 6 → 7. All new fields are optional: v5/v6 bundles import unchanged with a version-aware downlevel warning; bundles newer than the board still fail closed. - Manifest parsing validates the new timestamps like comment timestamps (invalid → warn and ignore, never a hard failure); shared types and the zod validator carry the new optional fields. - Import resolves parent slugs to pre-generated destination ids, drops self-references and cycles from tampered bundles with warnings, and orders rows parents-first because the self-referencing FK is checked per insert chunk. - `importIssues` writes the preserved timestamps (falling back to insert time when absent; `startedAt` stays null unless bundle-carried, per #11191's semantics) and `parentId`. - `addImportedComments` no longer blanket-bumps `updatedAt = now()`; it takes `GREATEST(updated_at, newest imported comment createdAt)`, so a preserved update time never regresses while unpreserved rows keep the old behavior. ## Verification - `pnpm vitest run server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts server/src/__tests__/productivity-review-service.test.ts` — 102 passed, 1 pre-existing opt-in benchmark skip. Includes: full round-trip with exact timestamp equality and a 3-deep parent chain against embedded Postgres; v6 back-compat (defaults + warning); forward-compat rejection (v8); cycle/self-reference/invalid-timestamp tampered-bundle handling; comment-bump preserve-awareness in both directions. - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/shared typecheck` — clean. ## Risks - **Rollout ordering**: a board on the previous build (max schema v6) refuses bundles exported by this build (stamped v7) — the existing newer-than-supported rejection, working as designed. Cross-instance moves need the importing board upgraded first. Called out here so operators aren't surprised during the transition window. - Parent edges from tampered bundles are dropped with warnings rather than failing the import; blocker relations already behave this way. - Timestamps are data-only; no destination schema migration. Stacked on #11191 (its commit is included here) — merge #11191 first; this PR then shows only the v7 changes. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4c062a0eb2 |
fix(ui): stop coercing imported agents to the destination CEO adapter (#11192)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import lets an operator bring a package of agents into an instance, and each agent declares which adapter runs it (Claude Code, Codex, and so on) > - The export and the server importer preserve each agent's adapter faithfully, but the Import page seeds an adapter override for every agent with the destination CEO's adapter before the user touches anything > - Every imported agent therefore arrives as the CEO's adapter (usually Claude Code) even when the source package holds a mix, and the picker shows the coerced value as if it were the source's, so nothing looks wrong > - This pull request makes the manifest adapter the default, sends overrides only for agents the user actually changed, and replaces the silent coercion with an explicit per-agent fallback warning when the destination truly lacks the source adapter > - The benefit is that a mixed Claude/Codex team imports as a mixed Claude/Codex team, and any real adapter gap is visible instead of silent ## Linked Issues or Issue Description **What happened?** A user imported a company package whose agents were a mix of Claude Code and Codex on the source instance. After the import, every agent was configured as Claude Code. The import preview showed no sign that anything had been changed. Cause: the Import page initializes its adapter-override map by assigning every agent the destination CEO's adapter type and sends that override for every agent, overriding the manifest's per-agent adapter server-side. For imports into a new company, the "CEO adapter" is read from whichever unrelated company is currently selected. **Expected behavior** Imported agents keep the adapter declared in the package. An override is sent only when the operator explicitly picks a different adapter, or when the source adapter is not installed on the destination — and in that case the page must say so per agent, not silently substitute. **Steps to reproduce** 1. On a source instance, create a company with one Claude Code agent and one Codex agent, and export it. 2. Import the package on another instance whose CEO uses Claude Code, changing nothing in the import dialog. 3. Both agents arrive configured as Claude Code; the Codex identity is gone. ## What Changed - The preview no longer seeds adapter overrides; the override map starts empty, and the picker displays each agent's manifest adapter (`ui/src/pages/CompanyImport.tsx`). - `buildFinalAdapterOverrides` sends an entry only when the effective adapter differs from the manifest or the agent's adapter config was edited — untouched agents flow through with no override. - The page fetches the destination's installed adapters (existing `adaptersApi.list()` client). When a manifest adapter is missing or disabled on the destination, only that agent defaults to the CEO's adapter, with a visible amber warning naming both adapters. If the adapters request fails, the page fails open: manifest adapters are kept and no coercion happens. - Tests: untouched mixed-adapter import sends no overrides; a user-changed agent sends exactly one; a missing destination adapter produces the fallback plus rendered warning for that agent only; an adapters-endpoint failure produces no coercion. ## Verification - `npx vitest run ui/src/pages/CompanyImport.test.tsx` — 19 passed (15 pre-existing + 4 new). - `pnpm --filter ./ui typecheck` (`tsc -b`) — clean. ## Risks - Behavior change: users who previously relied on the silent conversion (importing packages that reference adapters they don't have) now get an explicit per-agent fallback with a warning — same outcome, visible. The server's hard rejection of unknown adapter types remains the backstop for API callers. - UI-only change; no server or schema impact. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d816eb8095 |
fix(server): keep imported tasks quiescent under the productivity review sweep (#11191)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import brings a full company package — agents, tasks, routines — into an instance, with `pauseAutomations` promising a quiet landing > - The pause covers the imported entities, but the destination's own productivity-review sweep does not know the difference between imported rows and live work > - The importer stamps every imported in-progress task with `startedAt = now()`, so six hours later the sweep's long-active check fires on every one of them and floods the board with review tasks and agent wakeups > - This pull request stops fabricating `startedAt` on import and makes the sweep skip tasks whose assignee agent is paused > - The benefit is that an import lands quietly: no surprise review-task storm, and paused teams stay paused until the operator activates them ## Linked Issues or Issue Description **What happened?** After importing a company package with automations paused, a batch of "productivity review" tasks appeared roughly six hours later — one for every imported in-progress task — each with an owner-agent wakeup. The user described it as jarring and wasteful. Cause: `importIssues` fabricates `startedAt = now()` for imported in-progress rows, and `reconcileProductivityReviews` considers any assigned in-progress task without checking whether the assignee agent is paused, so its long-active-duration evidence (6 h threshold) trips on the fabricated timestamp. **Expected behavior** An import with paused automations must be quiescent: no destination sweep should generate work from imported rows until the operator unpauses the imported team. A paused agent must not accumulate review tasks it cannot act on. **Steps to reproduce** 1. Import a company package containing tasks with status `in_progress` assigned to agents, with "pause automations" enabled. 2. Wait for the productivity-review reconcile (runs at startup and on the heartbeat scheduler tick) more than six hours after the import. 3. Observe one new review task plus an owner wakeup per imported in-progress task. ## What Changed - `importIssues` no longer fabricates `startedAt` for imported `in_progress` rows; it inserts null (`server/src/services/issues.ts`). Audited every consumer of `issues.startedAt` — all are null-tolerant, and normal checkout/status-transition paths set the value when work really starts. - `reconcileProductivityReviews` skips candidates whose assignee agent is `paused`, counting them as skipped (`server/src/services/productivity-review.ts`). This is a general rule, not import-specific: a paused agent cannot act on a review. - Tests: paused-assignee candidate with an old `startedAt` creates no review, and creates one after unpausing; imported in-progress issue lands with null `startedAt` (embedded-Postgres import test); the pre-existing long-active regression test still passes. ## Verification - `pnpm vitest run server/src/__tests__/productivity-review-service.test.ts server/src/__tests__/company-portability-import-batching.test.ts` — 20 passed, 1 pre-existing opt-in benchmark skip. - `pnpm vitest run server/src/__tests__/company-portability.test.ts` — 78 passed. - `pnpm --filter @paperclipai/server typecheck` — clean. ## Risks - Behavior change beyond imports: tasks assigned to paused agents no longer receive productivity reviews anywhere. This is intended — the review would target an agent that cannot respond — and reviews resume on the first reconcile after unpausing. - Imported in-progress tasks now carry no `startedAt` until real work starts on the destination. The one sweep that read the fabricated value is the one this PR quiets; all other consumers fall back safely (audit in the commit body). - Low risk otherwise: no schema change, no API shape change. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8f7b8b3fda |
feat(release): add human-gated beta channel with stable soak enforcement (#11008)
> Follow-up to #11006 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem now publishes canary (every master push), nightly (scheduled, smoke-gated, added in #11006), and stable (manual) > - There is still no human-approved release-candidate lane between nightly and stable, and nothing enforces that a stable actually soaked anywhere before shipping > - Betas need a real approval gate, and stables need a soak policy that is data, not prose > - This pull request adds the beta channel: a manual promotion of a chosen nightly behind the `npm-beta` environment gate, re-smoked after publish, plus a stable preflight that enforces a 3-day beta soak with a written-justification bypass > - The benefit is a complete canary → nightly → beta → stable train where every stable shipped as a beta first, and emergencies leave a written trace ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** After #11006 the project has canary and nightly prerelease lanes, but no release-candidate lane. Stable promotion has no enforced soak: any ref can ship as stable directly. There is no approval boundary for a broader-audience prerelease, and no structured way to record why an emergency release skipped validation. **Proposed solution** Add a `beta` channel: a manual dispatch that promotes a chosen nightly's source commit, publishes behind the `npm-beta` GitHub environment (required reviewers are the gate), re-smokes the published beta, and tags `beta/vX`. Enforce in the stable path that the source commit shipped as a beta at least 3 days earlier (measured from the beta's npm publish time), with a `skip_soak_justification` input as the recorded emergency bypass. **Alternatives considered** Codifying the soak policy in docs only. Rejected: an unenforced policy decays; the preflight makes the policy executable while the justification input keeps the emergency path usable and auditable. ## What Changed - `scripts/release.sh` + `scripts/release-lib.sh`: `beta` channel — requires HEAD to carry a `nightly/v*` tag, publishes the package set as `YYYY.MDD.P-beta.N` under dist-tag `beta`, tags `beta/vYYYY.MDD.P-beta.N` - `.github/workflows/release.yml`: - `channel: beta` dispatch path: `select_beta` resolves the newest (or an explicit `source_version`) nightly and fails loudly on selection problems; `publish_beta` runs behind the `npm-beta` environment, pushes the tag, and dispatches `docker.yml`; `smoke_beta` re-runs the release smoke suite against the exact published beta version - stable path: new `preflight_stable` job enforces the 3-day beta soak from the beta's npm publish time; `skip_soak_justification` bypasses with the reason echoed into the job summary; dry runs report without blocking - `.github/workflows/docker.yml`: `beta/v*` tags publish `:beta` on both images, with exact version stamping - `.github/workflows/release-smoke.yml`: `beta` added to the dispatch choice list - Docs: `CHANNELS.md` beta entries; `RELEASING.md` beta lane, soak gate, and failure playbook; `RELEASE-AUTOMATION-SETUP.md` `npm-beta` environment setup, including the warning to create the environment before the first beta dispatch (GitHub auto-creates unprotected environments on first reference) - Tests: beta version-counting coverage in `scripts/release-registry-versions.test.mjs`; beta identity and nightly-tag guard coverage in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the two touched suites: 17 pass, including the 3 new beta tests - `bash -n` on both shell scripts and YAML parse of all three workflows - After merge, in order: create the `npm-beta` environment, dispatch `channel: beta` with `dry_run: true` to preview, then a real promotion of a published nightly through the approval gate, then a stable dry-run against a young beta to see the soak gate report ## Risks - If the `npm-beta` environment does not exist when the first beta dispatch runs, GitHub creates it with no protection rules and the beta publishes without approval. Mitigated by documentation and by creating the environment before merge (operator step) - Until the first beta exists, every stable dispatch requires `skip_soak_justification`. This is deliberate — the first beta ships immediately after this merges — but it is a behavior change to the stable dispatch - The soak clock reads the beta's npm publish time from the registry; a registry outage makes the preflight fall back to requiring justification (fail-closed) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.0 beta/v2026.811.0-beta.0 nightly/v2026.811.0-nightly.0 |
||
|
|
459469638e |
feat(adapter-codex-local): add secure device-login building blocks (#11097)
## Thinking Path > - Paperclip connects AI agents to local and remote runtimes. > - The Codex local adapter needs a safe device-login flow. > - A future sandbox integration needs strict prompt validation, secret protection, cleanup, and private credential storage. > - This pull request adds tested building blocks for that flow. > - The result gives a later Daytona integration a clear security boundary. ## Linked Issues or Issue Description No public issue covers this change. **Problem or motivation** The Codex local adapter has no safe, reusable flow to prove device login inside an isolated sandbox. **Proposed solution** Add parser, runner, credential export, and proof helpers. Validate the prompt, protect login data, store credentials in a private run-scoped home, and dispose all sandbox resources. **Alternatives considered** Do not connect a production Daytona driver in this change. Use an injected sandbox driver and focused tests first. This keeps the security controls testable before live provider integration. **Roadmap alignment** The change extends the Codex local adapter. It does not add a core Paperclip route or duplicate a planned core feature. **Additional context** The flow keeps the login URL, code, and token out of logs, results, and errors. The proof home uses a company-scoped root and a run-scoped private directory. ## What Changed - Add a pure parser for the exact Codex device-login URL and one-time code shape. - Add a sandbox runner with prompt handling, timeout, cancellation, and disposal. - Add a credential export step with company scoping, path checks, payload checks, private modes, locking, and cleanup. - Add redacted device-login fixtures and focused tests for parsing, secret redaction, runner outcomes, credential export, and cleanup. ## Verification - Run `pnpm --filter @paperclipai/adapter-codex-local exec vitest run`. - Run `pnpm --filter @paperclipai/adapter-codex-local exec tsc --noEmit`. - Review tests for strict URL and code validation, timeout, cancellation, disposal, secret redaction, path safety, payload safety, file modes, and cleanup. ## Risks - This change provides building blocks, not a live Daytona proof. - A later integration must connect the runner to a concrete sandbox driver. - Credential export depends on existing Codex authentication cache helpers. - Incorrect path or payload assumptions can reject valid credentials. ## Model Used Codex, GPT-5, tool use, code execution, and repository review. The Paperclip runtime controls the exact context window and reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.810.0-canary.6 |
||
|
|
5a0985f80a |
test(release-smoke): update onboarding spec for the mission-first wizard (#11190)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane gates every nightly publish on the release smoke suite, which drives real onboarding in a browser against the published artifact > - With the harness fixed (#11187, #11189), the gate reached the Playwright suite for the first time in CI — and the spec still walks the old onboarding wizard, so it fails at "Create your first agent" on every current build > - The wizard was redesigned to a mission-first five-step flow, and the spec rotted silently because the suite never ran in CI before > - This pull request rewrites the spec to drive the current wizard end to end > - The benefit is a smoke gate that actually tests today's product, verified against a real published canary ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `tests/release-smoke/docker-auth-onboarding.spec.ts`. **Problem or motivation** Nightly run 31431273139 failed in the smoke Playwright suite: the spec expects the old wizard step "Create your first agent", but current builds show the redesigned mission-first flow (front door → company → mission → team lead → connect model → review). The page snapshot in the run artifact shows the "Define your mission" step where the spec expected the agent step. Both retries failed identically — this is deterministic spec drift, not flake. **Proposed solution** Rewrite the spec for the current flow: fill the company name, define the mission directly (confirming creates the company), name the team lead, hire it through the adapter step — the adapter environment probe reports unhealthy in the CLI-less smoke container by design and must not block the hire — then launch to the dashboard. Assert the company, the ceo-role agent, and the company goal through the API. The first-task and assignment-run assertions are removed together with the wizard flow that created them. ## What Changed - `tests/release-smoke/docker-auth-onboarding.spec.ts`: rewritten for the mission-first wizard; sign-in and wizard-opening helpers and the company-name step are unchanged ## Verification - Full local run against the real nightly candidate: launched the smoke container for `paperclipai@2026.810.0-canary.3` via `scripts/docker-onboard-smoke.sh` (with the #11189 bind fix), then ran `pnpm run test:release-smoke` against it — 1 passed (4.5s) - After merge: dispatch `release.yml` with `channel: nightly` to run the full gate in CI ## Risks - Low. Test-only change. The spec now asserts less about first-task creation because the wizard no longer creates a first task; if a first-run trigger returns to onboarding, the spec should grow that assertion back ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (CI artifact forensics, UI source tracing, local Docker + Playwright reproduction and verification). All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.810.0-canary.5 |
||
|
|
f94f6003c6 |
fix(release-smoke): pin the smoke container to the lan bind preset (#11189)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane gates every nightly on the release smoke suite, which boots the published artifact in a Docker container and drives real onboarding > - The gate kept failing even after the readiness budget fix (#11187), and the new container-log dump revealed the server was healthy but listening on 127.0.0.1 inside the container, unreachable through Docker's port mapping > - `onboard --yes` without an explicit `--bind` prefers trusted-local quickstart defaults: it writes a loopback bind into the instance config and ignores the deployment env vars the harness passes, and that config outranks `HOST` at runtime > - This pull request pins the smoke container to the `lan` bind preset and adds a wiring test for it > - The benefit is a working nightly gate, verified end to end against a real published canary ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `docker/Dockerfile.onboard-smoke`, `scripts/__tests__/release-verify-workflow.test.mjs`. **Problem or motivation** Nightly run 31428558684 failed in smoke with the server unreachable at the mapped port for the full 420 second budget. The container logs (captured thanks to #11187) show a fully booted server with `Bind loopback (127.0.0.1)`. The harness sets `HOST=0.0.0.0` and the deployment env vars, but `onboard --yes` without `--bind` deliberately prefers trusted-local defaults, writes `bind: loopback` into the instance config, and the config outranks `HOST` at runtime. A loopback listener inside a container is invisible to the port mapping, so the health check can never pass. This behavior predates the current stable, so the harness was silently broken against every recent version — it only surfaced now because the nightly lane is the suite's first CI consumer. **Proposed solution** Pass `--bind lan` in the smoke container command (the flag is supported by `latest` and canary alike; it selects the all-interfaces preset and keeps the env-driven authenticated deployment), and pin the flag with a wiring test so it cannot regress silently. ## What Changed - `docker/Dockerfile.onboard-smoke`: the onboard command is now `onboard --yes --bind lan --data-dir ...`, with a comment explaining why the flag is load-bearing - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting the smoke Dockerfile pins a non-loopback bind preset ## Verification - Full local harness run against the real nightly candidate `2026.810.0-canary.1`: container healthy, bind banner shows `lan (0.0.0.0)`, authenticated bootstrap completed (admin created, bootstrap invite accepted, board session verified), `/api/health` returns `bootstrapStatus: ready` - `node --test scripts/__tests__/release-verify-workflow.test.mjs`: 4 pass - After merge: dispatch `release.yml` with `channel: nightly` to run the gate end to end in CI ## Risks - Low. The change only affects the smoke container. `--bind lan` inside a container exposes the port to the container network only; reachability from outside still goes through Docker's explicit port mapping ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (CI log forensics, upstream source tracing, local Docker reproduction and verification). All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.810.0-canary.4 nightly/v2026.810.0-nightly.0 |
||
|
|
30f6999cbe |
fix(release-smoke): configurable readiness timeout and diagnostics for slow containers (#11187)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane (#11006) gates every nightly publish on the release smoke suite, which boots the published artifact in a Docker container > - The suite's first CI execution failed at the health readiness check: the harness hard-codes a 90 second budget, but a CI container cold-installs paperclipai from npm and initializes embedded postgres with no warm caches > - When the timeout expired with the container still running, the harness printed no container logs, so the failure gave no diagnostics > - This pull request makes the readiness budget configurable, raises it for CI, and dumps container logs on timeout > - The benefit is that the nightly gate measures the artifact, not the runner's cold caches, and a red smoke run is diagnosable from its logs ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `scripts/docker-onboard-smoke.sh`, `.github/workflows/release-smoke.yml`. **Problem or motivation** Run 31426044332 (first forced nightly after #11006) failed in `smoke_nightly` with `server did not become ready at http://localhost:3232/api/health` after exactly 90 seconds. The harness's readiness window is hard-coded to 90 attempts at 1 second. Locally that works because the npm cache is warm; in CI the container downloads the full package set and embedded postgres first. The timeout path also printed no container logs when the container was still running, so there was no way to see how far boot had progressed. **Proposed solution** Make the readiness budget an environment variable (`SMOKE_READY_TIMEOUT_SECONDS`, default unchanged at 90 for local use), set it to 420 in the CI workflow, and dump the last 150 container log lines when the readiness check times out on a still-running container. ## What Changed - `scripts/docker-onboard-smoke.sh`: `SMOKE_READY_TIMEOUT_SECONDS` env var (default 90) replaces the hard-coded readiness budget; timeout with a still-running container now prints the tail of `docker logs` - `.github/workflows/release-smoke.yml`: sets `SMOKE_READY_TIMEOUT_SECONDS=420` for CI runs ## Verification - `bash -n` on the harness and YAML parse of the workflow - The real proof is the next `channel: nightly` dispatch of `release.yml`, which re-runs this suite in CI with the new budget ## Risks - Low. The local default is unchanged; CI runs simply wait longer before declaring failure, and a genuinely broken artifact still fails (with logs now) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. Diagnosis from CI run logs; patch model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.810.0-canary.3 |
||
|
|
5ca752dc81 |
fix(server): raise company import zip upload limit to 1 GB and make it operator-configurable (#11184)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import/export lets an operator move a full company package between instances, with the Import page uploading the package as one compressed `.zip` > - The server caps that upload at 128 MB, and real company packages with attachments now exceed it — imports fail at the preview step > - The failure message tells the user to use the CLI folder import, but that path posts inline JSON capped at 64 MB, so the advice is a dead end for exactly these packages > - This pull request raises the zip upload cap to a 1 GB default, makes it operator-configurable through an environment variable, scales the decompression-bomb guards from the cap in effect, and replaces the misleading hint > - The benefit is that large real-world company packages import successfully, and operators with unusual needs can tune the cap without a code change ## Linked Issues or Issue Description **What happened?** A company import fails at the preview step with `Preview failed: Import package exceeds 134217728 bytes`. The package is a valid Paperclip export. Its compressed size is larger than the 128 MB server cap (one reported package is 257 MB). The error panel suggests the CLI folder import, but that path sends the package as one inline JSON body capped at 64 MB, so it also fails. **Expected behavior** A valid company package of realistic size imports successfully through the Import page. If a package is too large, the error must state the limit clearly and suggest a step that can work. **Steps to reproduce** 1. Export a company with enough attachments to make the compressed package larger than 128 MB. 2. Open the Import page and upload the `.zip`. 3. Click "Preview import". 4. The preview fails with `Import package exceeds 134217728 bytes`. **Deployment mode** Reported from a managed deployment; the limit applies to all deployment modes. ## What Changed - Raise `PORTABLE_ZIP_UPLOAD_LIMIT_BYTES` from 128 MB to a 1 GB default (`server/src/http/body-limits.ts`). - Add the `PAPERCLIP_IMPORT_ZIP_MAX_BYTES` environment override. Invalid or non-positive values fall back to the default. - Scale the zip decompression-bomb guard from the configured cap at the import route: the aggregate inflated ceiling is 4x the cap. The per-entry ceiling stays at 512 MB because V8's string length limit applies to an entry regardless (`server/src/routes/companies.ts`, `packages/shared/src/portability-zip.ts`). - Report the 422 limit error in MB instead of raw bytes. - Replace the "use the CLI folder import for very large packages" hint on preview failure with advice that works: re-export the package without large attachments (`ui/src/pages/CompanyImport.tsx`). - Update the stale comment in `ui/src/lib/import-preflight.ts` that made the same CLI claim. - Add tests for the new default, the env override, and the invalid-override fallback. ## Verification - `pnpm vitest run server/src/__tests__/body-limits.test.ts packages/shared/src/portability-zip.test.ts server/src/__tests__/company-portability-routes.test.ts server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts` — all pass. - `pnpm vitest run ui/src/pages/CompanyImport.test.tsx` — passes, including the updated failure-panel copy assertion. - `pnpm typecheck` — clean across the workspace. - Manual: upload a `.zip` larger than the configured cap; the preview fails with `Import package exceeds the 1024 MB upload limit` and the new hint. A package between 128 MB and 1 GB now previews and imports. ## Risks - Peak per-import memory rises with the cap: the upload is buffered in memory and unzipped in one pass. A 1 GB compressed package can use several GB transiently. Imports are instance-admin actions, so the exposure is a deliberate operator action, not anonymous traffic. Operators on small hosts can lower the cap with `PAPERCLIP_IMPORT_ZIP_MAX_BYTES`. - The aggregate bomb guard moves from a fixed 512 MB to 4x the configured cap. It still bounds expansion far below what a decompression bomb needs. - No migration and no API shape change. The 422 message text changes; no code matches on the old text. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (file edits, local test runs, live-instance inspection). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.810.0-canary.2 |