mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
codex/plugin-task-execution
2085
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0a95ada1be |
feat(server): chunked import preview endpoint and resumable upload in the Import page and CLI (#11224)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The previous pull request added server-side chunked resumable import transfers; without clients, large imports still ride the single fragile upload > - The Import page and the CLI need to slice large packages, upload parts with retry and progress, resume after interruptions, and preview before applying > - Preview is the missing server piece: the browser flow is preview-then-import, so a completed spool must be previewable without re-uploading > - This pull request adds the transfer preview endpoint, switches the Import page to the chunked path for zips over 48 MB, and teaches the CLI the same for oversized local imports > - The benefit is that large imports get progress, per-part retry, and resume in both clients, while small imports keep the exact single-shot path they have today ## Linked Issues or Issue Description **What happened?** With only the server transfer routes in place, users still upload large company packages as one request from the Import page and the CLI: no progress indication, no retry below the whole file, and no resume after a dropped connection or refresh. The preview-then-import flow also cannot run against an uploaded transfer, forcing a second full upload. **Expected behavior** A large package uploads once as verified parts with visible progress; preview and import both run against the uploaded spool; an interrupted upload resumes with only the missing parts re-sent; packages at or below 48 MB behave exactly as before. **Steps to reproduce** 1. Select a 500 MB zip on the Import page over an unreliable connection. 2. Watch the single upload fail near the end and restart from zero, twice — once for preview, once for import. 3. Same story headless via the CLI. ## What Changed - Server: `POST /import/transfers/:id/preview` runs the existing preview logic against the completed spool (shared assembly + whole-file verification helper with apply); preview neither completes the run nor deletes the spool, so the subsequent apply reuses it. Missing parts respond with the missing list. - UI: zips over 48 MB take the chunked path in both preview and import — the file is sliced into 32 MB parts hashed with WebCrypto (single ArrayBuffer, no second copy), the transfer is created or resumed (the create response's missing-parts list drives what uploads), parts upload sequentially with three attempts each and visible progress, then transfer preview/apply replace the multipart calls. The existing preview pane, collision handling, adapter overrides, and async job polling are unchanged; ≤ 48 MB keeps the single-shot path. - CLI: oversized local `.zip` or folder imports zip/slice/hash with node crypto, upload with resume and per-part retry and progress lines, and use transfer preview/apply. Small packages keep the inline path byte-identical. - Failure honesty: adapters/API errors fail open to existing behavior; a part failing all attempts surfaces a durable error panel with resume intact. ## Verification - Server: preview-then-apply on one spool (run stays open, spool intact, then apply completes), preview with missing parts rejected — added to the transfer route suite (embedded Postgres). - UI suite: large file takes the chunked path (manifest shape, part uploads, progress, apply on a resumed transfer, single-shot endpoints never called), small file stays single-shot, part failure after three attempts surfaces the error panel without running preview, resume re-uploads only the missing part. - CLI: manifest slicing/hashing, threshold behavior for zip and folder sources, folder-zip round-trip through the real zip reader, upload resume/retry/exhaustion/already-completed, full-command chunked and small-zip inline flows. - Server, ui, cli typechecks clean. Exact counts in the PR checks. ## Risks - The 48 MB threshold only routes between two verified paths; behavior below it is untouched. - Chunked CLI imports use the board-scoped transfer routes, so oversized CLI imports need board credentials (agent tokens keep the agent-safe small-file path). No privilege change — board actors already had the generic routes — but the two size regimes differ semantically; called out for review. - A CLI dry-run over the threshold uploads parts before previewing; the spool persists (24 h sweep) and a later apply resumes without re-upload — inherent to preview-against-spool. Stacked on #11223 — merge that first; this PR then shows only the preview endpoint and client changes. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2494a2a0fe |
perf: add repeatable issue-detail baseline rig (#10409)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The issue detail page is a core operator surface where perceived latency directly affects task navigation > - Performance work needs repeatable evidence so later optimizations can be compared against the same scenarios > - The page did not expose stable user-timing marks for its header or first useful content > - There was also no isolated seeded browser rig that measured warm navigation, cold deep links, waterfalls, or server time > - This pull request adds the instrumentation and a one-command Playwright baseline harness > - The benefit is that issue-page performance changes can be validated with reproducible median measurements instead of anecdotes ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: `ui/`, `server/`, and browser performance tooling. **Problem or motivation** The issue detail page performs a large client bootstrap and request fan-out, but the repository lacks stable user-timing boundaries and a repeatable benchmark. That makes performance changes difficult to compare and allows regressions to be judged from anecdotes instead of consistent evidence. **Proposed solution** Add stable header/content paint measures, development/QA-only lifecycle vital reporting, aggregate server timing for the issue endpoint, and a seeded Playwright command that runs warm/cold scenarios under throttled and unthrottled profiles with N≥5 median reporting. **Alternatives considered** Ad hoc DevTools recordings were rejected because they are not repeatable or reviewable. Production telemetry was rejected because this baseline should not change production data collection. A unit-only harness was rejected because it cannot capture browser bootstrap, rendering, and network waterfall costs. **Roadmap alignment** The roadmap calls for agent performance to be measurable over time. This change applies that evidence-first principle to a core operator page and does not duplicate a listed roadmap deliverable. **Additional context** The generated report includes warm and cold medians, TTFB/FCP/LCP where applicable, request and byte totals before first useful content, JavaScript bytes, and issue endpoint server timing. ## What Changed - Added `issue-detail:navigate→header-paint` and `issue-detail:navigate→content-paint` user-timing measures to the issue detail page. - Added development/QA-only TTFB, LCP, and INP console reporting without production telemetry delivery. - Added `Server-Timing` for `GET /api/issues/:id`. - Added `pnpm exec playwright test --config tests/perf/issue-detail/playwright.config.ts`, which seeds an isolated instance and runs N≥5 warm/cold samples under unthrottled and Fast 4G/4x CPU profiles. - Added Markdown, raw JSON, and Chrome-trace outputs with median baseline tables and waterfall data. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm check:token-gates` - `npx playwright test --config tests/perf/issue-detail/playwright.config.ts --list` - `pnpm exec playwright test --config tests/perf/issue-detail/playwright.config.ts` — passed 20 samples in 9.4 minutes (5 runs × 2 scenarios × 2 profiles) for the baseline; post-review integrity reruns also exercised the corrected paths, while this shared runner intermittently killed Chromium processes, so the rig now performs one bounded browser-crash retry per sample. - Baseline medians: warm unthrottled 278/447 ms header/content; cold unthrottled 646/646 ms; warm throttled 1240/2060 ms; cold throttled 3932/3933 ms. ## Risks - Low product risk: the new browser measurements are development/QA tooling and the UI timing work does not change visible layout. - `Server-Timing` exposes only aggregate handler duration, not query contents or private identifiers. - Native INP reporting uses supported browser event timing entries and silently no-ops where unsupported. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.4, tool-assisted coding and browser execution with reasoning enabled; context-window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
5bb2490b86 |
feat(ui): surface all issue documents and agent artifacts in chat-style sidebar (#11226)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The task detail page uses a chat-style thread with a right sidebar; the sidebar has a Plan tab and an Artifacts tab (#11101 made this UI the default) > - The Plan tab only showed the one issue document named `plan`, and the Artifacts tab only listed formal work products; other agent-authored documents (for example a `synthesis` doc) and agent-attached files were invisible in the sidebar > - Users could see an agent mention a document in the thread but had no way to find that document in the sidebar, which breaks trust in the task view as the record of the work > - This pull request surfaces every non-system issue document in the Plan tab, composes the Artifacts tab from work products, documents, and agent-created attachments, and gives thread images a full-screen lightbox with download > - The benefit is that anything an agent produces on a task is now reachable from the sidebar, while user uploads stay with their comments in the thread ## Linked Issues or Issue Description Refs #11101 (chat-style task UI default — this PR extends its sidebar). **Subsystem affected** Task detail UI (chat-style thread sidebar): Plan tab, Artifacts tab, and thread attachment rendering in `ui/src`. **Current behavior** The Plan tab renders only the issue document literally named `plan`. The Artifacts tab renders only formal work products. Agent-authored documents with any other name, and files agents attach to comments, do not appear anywhere in the sidebar. Thread images open as bare links. **Proposed behavior** The Plan tab lists every non-system issue document, with the `plan` document first and the others rendered inline below it. The Artifacts tab composes three sources — work products, issue documents, and agent-created comment attachments — deduplicated against attachment-backed work products via `metadata.attachmentId`, and shows whenever any source is non-empty. Work-product rows without a resolvable attachment or document fall back to links found in their metadata so they stay clickable. Images in the thread open a shared full-screen lightbox with a download action. Files uploaded by users stay thread-only and are not mixed into the Artifacts tab. **Reason and benefit** Agents routinely produce documents that are not named `plan` and attach files to their comments. Users reading the thread must be able to find every one of those outputs from the sidebar. Redundant surfacing is acceptable; an unfindable document is not. **Breaking changes** None. This is additive rendering; no schema or API changes. ## What Changed - `IssuePropertiesPlansTab.tsx`: renders all non-system issue documents, `plan` primary, others inline below via `MarkdownBody` - `IssuePropertiesArtifactsTab.tsx`: composes work products + documents + agent-created attachments with dedupe; rows without an attachment/document target fall back to `metadata` links - `IssueProperties.tsx`: Artifacts tab visibility now derives from the composed source set - New `ui/src/lib/issue-artifacts.ts`: pure composition/dedupe logic, unit-tested - New `ui/src/components/task-chat/task-chat-attachments.ts`: splits agent vs user comment attachments, unit-tested - `TaskChatBubble.tsx`: thread images open the shared full-screen lightbox with download - `useIssueDocuments.ts`: hook now exposes the full issue-document list ## Verification - `pnpm typecheck` — passes across the workspace - `pnpm check:token-gates` — 3/3 CLEAN - `cd ui && pnpm vitest run src/lib/issue-artifacts.test.ts src/components/task-chat/task-chat-attachments.test.ts src/pages/IssueDetail.test.tsx` — 74 tests pass - Manual: open a task whose agent created a document not named `plan` (for example `synthesis`); confirm it appears in the Plan tab below the plan and in the Artifacts tab; confirm an image the agent attached appears under Artifacts; confirm a user-uploaded image stays only in the thread and opens full screen with a download button Snapshot baselines are intentionally not updated for this visual change, per the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks - Low risk: rendering-only change scoped to the task sidebar and thread bubbles; composition logic is pure and unit-tested - Dedupe relies on `metadata.attachmentId` linkage; a work product with malformed metadata would render as a duplicate row (cosmetic only) ## Model Used - Claude (Anthropic), model id `claude-fable-5`, extended thinking enabled, agentic tool use via Claude Agent SDK (Claude Code harness) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0044fa8904 |
Let tenants edit env vars on managed sandbox environments; add managed-sandbox-only mode (#11200)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments give each agent run an execution target: the local host, SSH, or a sandbox provider > - A managed deployment can provision one platform-managed sandbox environment through the `PAPERCLIP_MANAGED_CONFIG` `environments` section > - That row is fully locked today. A tenant cannot add environment variables for their agents. There is also no way to hide local execution — run selection falls back to the local row > - A platform that manages the sandbox for its tenants needs both: the tenant adds env vars (and nothing else), and local execution is neither visible nor reachable > - This pull request opens exactly one tenant edit (env vars) on the managed sandbox row, and adds an `enableManagedSandboxOnly` mode that hides local and makes run selection fail closed > - The benefit is a complete managed-sandbox experience with no change for self-hosted instances ## Linked Issues or Issue Description **Subsystem affected** Environments (managed sandbox provisioning, environment routes, run environment selection) and the environments UI. **Problem or motivation** Platform-provisioned sandbox environments (`metadata.managedByPaperclip`) reject every write on cloud-managed instances. Agents often need environment variables inside their sandbox. The tenant has no way to set them on the managed row. Separately, an operator cannot remove local execution: the environment list always shows the local row, and run selection falls back to it when no default is set. **Proposed solution** Allow an envVars-only PATCH on the managed sandbox row, and echo those env vars back for editing. Add a managed-tier feature (`enableManagedSandboxOnly`) that hides the local environment from all read surfaces and redirects local-landing run selection to the managed sandbox environment, failing closed when it is unavailable. **Alternatives considered** UI-only hiding of the local row. This was rejected: it does not stop a run from resolving to local, so it is presentation without enforcement. Full unlock of the managed row was also rejected: name, driver, and config stay platform-owned so boot reconciliation cannot fight tenant edits. ## What Changed - `server/src/routes/environments.ts`: the platform-provisioned write floor admits an envVars-only PATCH on the generalized managed sandbox row (sandbox driver, `managedByPaperclip`, not legacy kubernetes-marker rows). Name, driver, config, status, metadata, and DELETE stay rejected. The read floor stops blanking env vars on that row; credential-shaped config keys stay redacted for every actor. Legacy kubernetes-marker rows keep the full floor. - Same file: under `enableManagedSandboxOnly`, the environments list and the by-id read omit the local row for every actor, including instance admins. - `server/src/services/execution-workspace-policy.ts`: `resolveExecutionWorkspaceEnvironmentId` gains the managed-sandbox-only inputs. A selection that lands on the local environment is redirected to the managed sandbox environment. With no active managed row it throws `ManagedSandboxUnavailableError` — never local. Non-local selections (ssh, user-created sandboxes) are untouched. - `server/src/services/heartbeat.ts`: the run path reads the flag, looks up the managed row (`findManagedSandboxEnvironment`, new read-only finder in `environments.ts`), and passes both to the resolver. Mirrors the forced-kubernetes precedent, which keeps precedence when both regimes are on. - `server/src/services/managed-environments.ts`: after a successful reconcile, the instance default environment moves to the managed sandbox row when the current default is unset, local, or dangling. A tenant-chosen custom environment is never overridden. - `packages/shared`: new `enableManagedSandboxOnly` key (schema default false, catalog tier `managed`, cloudDefault false, selfHostedDefault false) and the matching interface field. - UI: managed rows show a "Managed by Paperclip" lock badge; editing one opens a dedicated env-vars-only editor that sends the one PATCH shape the server admits (the old full form failed with a 403 on save). New `ui/src/lib/managed-sandbox-environment.ts` mirrors the local filter for cached lists (applied in the project picker; the agent picker already excluded local). The experimental settings page gains the toggle at its alphabetical card position. `environmentsApi.update` now declares the `envVars` field it already sent. - Tests: environment route floor coverage (envVars-only accepted, mixed bodies rejected, legacy rows still blanked and locked, local hidden and 404 under the flag, self-hosted unchanged), an embedded-postgres service test pinning that boot reconciliation never touches tenant env vars, resolver redirect/fail-closed cases, managed-environments default-stamping cases, and UI lib/settings tests. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/environment-routes.test.ts src/__tests__/environment-service.test.ts src/__tests__/execution-workspace-policy.test.ts src/services/managed-environments.test.ts` — all pass. - `pnpm --filter @paperclipai/shared exec vitest run` — 425 pass (catalog/schema default parity is pinned by an existing test). - `pnpm --filter @paperclipai/ui exec tsc --noEmit` and the affected UI suites (CompanyEnvironments, InstanceExperimentalSettings incl. card-order test, new lib test) — all pass. - Full workspace `pnpm test`: 3,414 passed. 17 files report failures on this machine; the identical 17 fail on a clean `origin/master` worktree in the same environment (git-worktree/skills/embedded-postgres environment dependencies and plugin-SDK zero-test collections). One additional file (`issue-monitor-scheduler.test.ts`) failed one timing-sensitive test in one of two full-suite runs and passes 7/7 in isolation on this branch — a flake in a domain this diff does not touch. The branch introduces no new failures. - Self-hosted zero-delta: every new behavior is gated on the cloud-managed instance check or the new flag, which defaults to false in schema and catalog; pinned by the "does not floor platform-marked rows on self-hosted instances" and flag-off tests. ## Risks - Behavior is opt-in twice over: the write-floor exception applies only to rows the managed-config provisioner stamps, and the hiding/forcing applies only when `enableManagedSandboxOnly` is on (default false everywhere). Self-hosted instances see no change. - The env-vars echo is scoped to the generalized managed sandbox row; legacy kubernetes-marker rows keep the blanket floor because pre-generalization builds may have written platform values there. - Fail-closed run selection means a managed instance with the flag on and an archived managed row (provider plugin down) refuses runs with a precise error instead of running locally. That is the intended posture. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
145d86911b |
Remove decision training UI (#11225)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Decisions desk shows work that needs an operator response. > - It also exposed decision-training actions and a separate training library. > - Paperclip does not plan to use these training surfaces now. > - Keeping inactive controls makes the Decisions workflow harder to scan. > - This pull request removes the training UI and keeps the backend snapshot contract unchanged. > - The benefit is a smaller and clearer Decisions workflow without a data migration. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Decisions desk currently exposes training controls, training state, and a separate training library route. **Subsystem affected** `ui/` — React and Vite board UI. **Current behavior** Operators can open a training library from the Decisions toolbar. They can also mark a decision for training from rows and inspect the result in a drawer. **Proposed behavior** Remove the training controls, badges, drawer, library pages, and routes from the Decisions UI. Keep the server APIs and stored training examples unchanged. **Reason and benefit** The product does not plan to use decision training now. Removing the unused surfaces reduces Decisions UI noise and avoids presenting a workflow that operators should not use. **Breaking changes** The `/decisions/training` UI routes are no longer registered. Existing server endpoints and stored decision-training data remain compatible. ## What Changed - Removed decision-training controls and state from Decisions toolbars, rows, queue pages, and shelves. - Removed the training drawer, library, inspector, API client, helpers, query keys, and routes. - Added route and row regressions that assert training UI does not return. - Updated the Decisions Storybook description to match the available controls. ## Verification - `pnpm exec vitest run ui/src/App.test.tsx ui/src/components/AttentionQueueRow.test.tsx` — 32 tests passed. - `pnpm check:token-gates` — all gates clean. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — server and UI partitions passed. The CLI partition had one environment-only failure because this agent runtime injects static AWS credentials. The exact CLI file passed all 8 tests when those credential variables were unset. - Searched `ui/src` and `ui/storybook` for the removed training routes, drawer, library, badges, and actions. Only negative regression assertions remain. ## Risks - Low implementation risk. This change deletes UI-only entry points and does not change the database or server APIs. - Saved training-page URLs no longer render a board route. This is the intended behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5.6-sol`, with `xhigh` reasoning. The runtime did not expose the context-window size. The agent used repository tools, shell execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
45dfb183b6 |
fix(ui): add undo action to inbox archive toast (#11220)
## Thinking Path > - Paperclip helps operators supervise AI-agent work. > - The Mine inbox keeps tasks that need an operator's attention in one place. > - Operators can archive a task from its detail page after they finish triage. > - That action is easy to select accidentally and did not offer immediate recovery. > - This pull request adds Undo to the archive success toast and keeps inbox caches consistent. > - The benefit is fast recovery without searching for or reopening the task. ## Linked Issues or Issue Description **What happened?** Archiving a task from the Mine inbox removed it and showed a success toast with no recovery action. **Expected behavior** The success toast should offer Undo. Selecting Undo should restore the task through the existing unarchive API while preserving a consistent inbox view. **Steps to reproduce** 1. Open a task from the Mine inbox. 2. Select the archive action. 3. Observe that the task leaves the inbox and the success toast has no Undo action. **Paperclip version or commit** Reproduced on `master` before this change. **Deployment mode** Local dev, built from source. Related prior work: #9931 and #10668. ## What Changed - Add an Undo action to the successful inbox archive toast. - Optimistically restore the task in captured inbox query caches before the unarchive request completes. - Cancel in-flight inbox fetches and clear the local archive guard so stale responses cannot hide the restored task. - Reapply the archive guard and remove the cached task if the unarchive request fails. - Add regression tests for successful Undo, the in-flight cache race, and failed Undo rollback behavior. ## Verification - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx ui/src/lib/inboxArchiveCache.test.ts` — 51 tests passed on the final head. - `pnpm check:token-gates` — passed. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — the general server and UI groups passed. One CLI doctor assertion detected injected host AWS credentials and passed all 8 tests with those unrelated variables unset. A task-watchdog scheduler test also passed all 18 tests in isolation after one full-suite timing failure. ## Risks Low risk. The change uses the existing unarchive endpoint and inbox cache helpers. Undo failure returns the task to its archived state, shows an error toast, and invalidates the inbox queries for server reconciliation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, model ID GPT-5. The service manages the context window. Reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7734f4b32d |
fix(ui): keep slash autocomplete scrollable in dialogs (#11222)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip helps operators manage AI-agent companies. > - Operators create tasks and comments through shared rich-text editors. > - These editors show slash-command and mention matches in a floating menu. > - Modal dialogs treat that body-level menu as outside content and cancel its wheel and touch movement. > - This pull request keeps scroll events inside the floating menu and preserves native scrolling. > - The benefit is that operators can reach every match with a mouse wheel, a trackpad, or a touch screen. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. **What happened?** Slash-command and mention menus could contain more matches than their visible height. When an editor was inside a modal dialog, the modal scroll lock canceled wheel and touch movement on the body-level menu portal. Operators could not scroll to later matches. **Expected behavior** The autocomplete menu must scroll with a mouse wheel, a two-finger trackpad gesture, and a vertical touch gesture. Keyboard selection and normal editor behavior must stay unchanged. **Steps to reproduce** 1. Open a task or comment editor inside a modal dialog. 2. Enter a slash command or mention query that has more matches than the menu can show. 3. Try to scroll the menu with a wheel, trackpad, or touch gesture. **Paperclip version or commit** `7ea2068ef8` on `master`. **Deployment mode** Local development UI built from source. ## What Changed - Keep wheel and touch movement inside the shared autocomplete menu portal. - Add vertical overscroll containment while preserving native momentum scrolling. - Add a regression test that mounts the real dialog and verifies that wheel and touch movement stay uncanceled. ## Verification - `pnpm --dir ui exec vitest run src/components/MarkdownEditor.test.tsx` - `pnpm check:token-gates` - `pnpm -r typecheck` - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run` - `pnpm build` The two AWS variables are omitted from the full test command because this agent runtime injects static AWS credentials. One unrelated CLI doctor test correctly warns when those credentials are present. The CI environment does not inject them. ## Risks - Low risk. Event propagation stops only on the open autocomplete menu portal. - Ancestor listeners no longer receive wheel or touch movement from that menu. Native menu scrolling and option-level touch handling still receive the events. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with model ID `gpt-5`. The deployment suffix and context-window size are not exposed to the agent. The model used agentic reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7ea2068ef8 |
fix(files): only highlight accessible workspace file links (#11090)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task comments can contain references to files in project and execution workspaces. > - Paperclip detected path-shaped inline code and showed it as an actionable file chip. > - The UI did not first confirm that the current board session could open the file. > - Missing, denied, ambiguous, remote, and unsupported files therefore looked actionable and failed after a click. > - This pull request adds an issue-scoped availability check and promotes only confirmed files to chips. > - The benefit is that the task thread shows a file action only when that action can succeed. ## Linked Issues or Issue Description **What happened?** Task comments promoted path-shaped inline code to file chips before Paperclip checked the file. A chip could point to a missing, denied, ambiguous, remote, or non-previewable file. The action then failed after the user selected it. **Expected behavior** Paperclip must show a file chip only after the server confirms that the current board session can open the exact file reference. All other path-shaped text must stay ordinary inline code. **Steps to reproduce** 1. Add a task comment that contains inline code with a missing or inaccessible workspace path. 2. Open the task thread as a board user. 3. Observe that the path looks like an actionable file chip. 4. Select the chip and observe that the file cannot open. **Paperclip version or commit** `19be4cf927` and earlier. **Deployment mode** Local dev and self-hosted server. **Access context** Board user. ## What Changed - Added shared request, response, and validation contracts for batched workspace-file availability checks. - Added an issue-scoped server endpoint that resolves file references with company, issue, workspace, and preview-access checks. - Added bounded batch concurrency and tests for missing, denied, ambiguous, remote, unsupported, and available files. - Added an issue-scoped UI availability registry that deduplicates, batches, caches, and invalidates file checks. - Changed task-comment markdown rendering so only confirmed files get chip styling and file-viewer behavior. - Bound each chip to the exact workspace target that passed the availability check. ## Verification - `pnpm exec vitest run packages/shared/src/workspace-file-resource.test.ts server/src/__tests__/file-resources.test.ts ui/src/components/MarkdownBody.test.tsx ui/src/components/WorkspaceFileMarkdownBody.availability.test.tsx ui/src/lib/remark-workspace-file-refs.test.ts ui/src/lib/workspace-file-availability.test.ts` — 93 passed, 35 skipped. - `pnpm check:token-gates` — clean. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — all server and UI groups passed. One unchanged CLI test saw the run-injected static AWS credentials and expected only its local `AWS_PROFILE`. The same test passed, 8 of 8, after removing only `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` from its process environment. ## Risks - File chips now appear after an asynchronous availability check, so path-shaped text can briefly render as inline code. - Availability results use the existing 30-second file-resource cache window. File-resource invalidation forces a new check. - The endpoint limits each request to 100 references and the client chunks larger sets. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The service did not expose a more specific model ID or context-window size. The agent used high-reasoning mode, repository tools, command execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
815e49bb7c | feat: make chat-style tasks the default experience (#11101) | ||
|
|
c4abecb2c4 |
fix(skills): refresh project folders in place (#11066)
## Thinking Path > - Paperclip helps operators manage agent skills across a company. > - The installed skills view groups project-backed skills into folders. > - The view showed two folder creation controls and only offered a global project scan. > - Operators need one clear folder action and a refresh action for the selected project. > - This pull request keeps folder creation in the folder rail and adds a scoped project refresh. > - The benefit is a calmer skills view and faster, more precise project skill updates. ## Linked Issues or Issue Description No public GitHub issue exists for this focused UI bug. **What happened?** The installed skills view repeated the folder creation action in the toolbar. A selected project folder also had no way to refresh only its own project skills. **Expected behavior** The folder rail must own folder creation. A selected project-backed folder must offer a refresh action that scans only that project and refreshes the skill and folder queries. **Steps to reproduce** 1. Open the installed skills view for a company with project-backed skill folders. 2. Select a project folder. 3. Observe the duplicate folder action and the absence of a project-scoped refresh action. **Paperclip version or commit** Reproduced before this two-commit fix on `master`. **Deployment mode** Local development with `pnpm dev`. ## What Changed - Removed the duplicate toolbar folder creation button when the folder rail exists. - Preserved the toolbar folder action when no folder rail exists. - Added a refresh action beside the breadcrumb for a selected project-backed folder. - Passed the selected project ID to the project scan API. - Refreshed both the installed skill list and skill folder data after scans. - Added component tests for compact folder creation, the empty-folder fallback, and scoped project refresh. ## Verification - `pnpm exec vitest run ui/src/pages/CompanySkills.test.tsx` — 20 tests passed. - `pnpm check:token-gates` — passed with all three gates clean. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — the server and UI stages passed 7,475 tests. The CLI stage then found one environment-sensitive AWS doctor assertion because this agent runtime injects static AWS credentials. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts --project paperclipai` — all 8 tests passed. - GitHub CI — all latest-head checks passed. ## Risks - Low risk. The scoped refresh depends on the existing `project:<id>` folder system key. - The global scan path is unchanged. - There are no schema, migration, API contract, dependency, workflow, or documentation changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 family. The runtime did not expose a more specific model ID or context-window size. The agent used high-reasoning mode, repository tools, GitHub tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b58ce27a02 |
fix: isolate execution workspace summaries (#10790)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip gives operators a summary for each workspace. > - An execution workspace detail page used the parent project-workspace summary slot. > - Two execution workspaces under one project workspace could therefore show the same summary. > - This pull request gives each execution workspace its own summary scope. > - It also limits the summary snapshot and generated issue to that execution workspace. > - The benefit is that a new or parallel execution workspace cannot inherit unrelated status. ## Linked Issues or Issue Description **What happened?** An execution workspace detail page read and refreshed the summary slot for its parent project workspace. Parallel execution workspaces could show the same status and include issues from each other. **Expected behavior** Each execution workspace must have one isolated summary slot. Its generated snapshot must include only issues assigned to that execution workspace. **Steps to reproduce** 1. Create two execution workspaces under one project workspace. 2. Add different issues to each execution workspace. 3. Generate the summary in the first execution workspace. 4. Open the second execution workspace. 5. Observe that the old implementation could reuse the first summary. **Paperclip version or commit** The problem exists on `master` before this pull request. **Deployment mode** The issue affects both local trusted and authenticated deployments. ## What Changed - Added `execution_workspace` to the shared summary-slot scope contract. - Validated execution-workspace ownership and stored generated summary issues on the correct execution workspace. - Limited execution-workspace snapshots to issues with the matching execution workspace ID. - Updated the execution workspace page to use its own summary slot. - Updated Summarizer instructions, routine options, catalog metadata, documentation, and regression tests. ## Verification - `NODE_ENV=test pnpm exec vitest run packages/shared/src/summary-slot.test.ts server/src/__tests__/summary-slots.test.ts ui/src/pages/ExecutionWorkspaceDetail.test.tsx` — 30 focused tests passed; the embedded-Postgres server tests were run outside the process-restricted sandbox. - `pnpm check:token-gates` — passed. - `pnpm --filter @paperclipai/skills-catalog validate` — passed with 17 catalog skills. - [Latest-head GitHub Actions](https://github.com/paperclipai/paperclip/actions/runs/31491475405) — all 22 jobs passed on `beea14cbaf`, including typecheck, build, server/workspace tests, serialized suites, e2e, canary, and aggregate verification. One unrelated adapter cleanup test initially hit an `ENOTEMPTY` temp-directory race; its single permitted rerun passed. - Greptile — 5/5 confidence on `beea14cbaf`, 12 files reviewed, zero comments added, and zero unresolved threads. ## Risks - Low risk. The new scope is additive. - Existing project and project-workspace summary slots keep their current keys and behavior. - A summary generated for an execution workspace now excludes sibling workspace issues by design. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The deployment does not expose a more specific model ID or context-window value. It used agentic reasoning, repository tools, code execution, and GitHub tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9cdaa5416e |
fix(ui): remember folded inbox subtasks (#11069)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The inbox helps operators scan parent tasks and their sub-tasks > - Operators can fold a parent task to hide its sub-tasks > - The inbox previously forgot that fold state after a page refresh > - This pull request stores the fold state for each company and restores it when the inbox loads > - The benefit is that the inbox keeps the operator's chosen task layout across page refreshes ## Linked Issues or Issue Description **What happened?** The inbox reset every folded parent task after a page refresh. This made all nested sub-tasks visible again. **Expected behavior** The inbox must keep each folded or unfolded parent state after a page refresh. The state must remain separate for each company. **Steps to reproduce** 1. Open the inbox with parent and child tasks. 2. Fold one parent task. 3. Refresh the page. 4. Observe that the child task is visible again without this fix. **Paperclip version or commit** Current `master` before this pull request. **Deployment mode** Local dev and built-from-source deployments. ## What Changed - Added company-scoped local storage helpers for collapsed inbox parent IDs. - Restored the stored parent fold state when the inbox mounts or the selected company changes. - Saved both direct toggle changes and explicit collapse changes. - Added helper tests and an inbox remount regression test for both folded and unfolded states. ## Verification - `pnpm exec vitest run ui/src/lib/inbox.test.ts ui/src/pages/Inbox.test.tsx` — 77 tests passed. - `pnpm check:token-gates` — all gates passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/ui build` — passed. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — 3,521 tests passed and four skipped. One unrelated server test on the current base fails because it reads `heartbeat.scheduling_suppressed` instead of `issue_commented`; the same test fails alone and this pull request changes only inbox UI files. ## Risks - Low risk. The state is local to the browser and scoped by company ID. - Old parent IDs can remain in local storage after tasks are deleted, but they do not affect visible tasks. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.6-sol`, with reasoning, tool use, and code execution. The runtime does not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4c062a0eb2 |
fix(ui): stop coercing imported agents to the destination CEO adapter (#11192)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import lets an operator bring a package of agents into an instance, and each agent declares which adapter runs it (Claude Code, Codex, and so on) > - The export and the server importer preserve each agent's adapter faithfully, but the Import page seeds an adapter override for every agent with the destination CEO's adapter before the user touches anything > - Every imported agent therefore arrives as the CEO's adapter (usually Claude Code) even when the source package holds a mix, and the picker shows the coerced value as if it were the source's, so nothing looks wrong > - This pull request makes the manifest adapter the default, sends overrides only for agents the user actually changed, and replaces the silent coercion with an explicit per-agent fallback warning when the destination truly lacks the source adapter > - The benefit is that a mixed Claude/Codex team imports as a mixed Claude/Codex team, and any real adapter gap is visible instead of silent ## Linked Issues or Issue Description **What happened?** A user imported a company package whose agents were a mix of Claude Code and Codex on the source instance. After the import, every agent was configured as Claude Code. The import preview showed no sign that anything had been changed. Cause: the Import page initializes its adapter-override map by assigning every agent the destination CEO's adapter type and sends that override for every agent, overriding the manifest's per-agent adapter server-side. For imports into a new company, the "CEO adapter" is read from whichever unrelated company is currently selected. **Expected behavior** Imported agents keep the adapter declared in the package. An override is sent only when the operator explicitly picks a different adapter, or when the source adapter is not installed on the destination — and in that case the page must say so per agent, not silently substitute. **Steps to reproduce** 1. On a source instance, create a company with one Claude Code agent and one Codex agent, and export it. 2. Import the package on another instance whose CEO uses Claude Code, changing nothing in the import dialog. 3. Both agents arrive configured as Claude Code; the Codex identity is gone. ## What Changed - The preview no longer seeds adapter overrides; the override map starts empty, and the picker displays each agent's manifest adapter (`ui/src/pages/CompanyImport.tsx`). - `buildFinalAdapterOverrides` sends an entry only when the effective adapter differs from the manifest or the agent's adapter config was edited — untouched agents flow through with no override. - The page fetches the destination's installed adapters (existing `adaptersApi.list()` client). When a manifest adapter is missing or disabled on the destination, only that agent defaults to the CEO's adapter, with a visible amber warning naming both adapters. If the adapters request fails, the page fails open: manifest adapters are kept and no coercion happens. - Tests: untouched mixed-adapter import sends no overrides; a user-changed agent sends exactly one; a missing destination adapter produces the fallback plus rendered warning for that agent only; an adapters-endpoint failure produces no coercion. ## Verification - `npx vitest run ui/src/pages/CompanyImport.test.tsx` — 19 passed (15 pre-existing + 4 new). - `pnpm --filter ./ui typecheck` (`tsc -b`) — clean. ## Risks - Behavior change: users who previously relied on the silent conversion (importing packages that reference adapters they don't have) now get an explicit per-agent fallback with a warning — same outcome, visible. The server's hard rejection of unknown adapter types remains the backstop for API callers. - UI-only change; no server or schema impact. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5ca752dc81 |
fix(server): raise company import zip upload limit to 1 GB and make it operator-configurable (#11184)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import/export lets an operator move a full company package between instances, with the Import page uploading the package as one compressed `.zip` > - The server caps that upload at 128 MB, and real company packages with attachments now exceed it — imports fail at the preview step > - The failure message tells the user to use the CLI folder import, but that path posts inline JSON capped at 64 MB, so the advice is a dead end for exactly these packages > - This pull request raises the zip upload cap to a 1 GB default, makes it operator-configurable through an environment variable, scales the decompression-bomb guards from the cap in effect, and replaces the misleading hint > - The benefit is that large real-world company packages import successfully, and operators with unusual needs can tune the cap without a code change ## Linked Issues or Issue Description **What happened?** A company import fails at the preview step with `Preview failed: Import package exceeds 134217728 bytes`. The package is a valid Paperclip export. Its compressed size is larger than the 128 MB server cap (one reported package is 257 MB). The error panel suggests the CLI folder import, but that path sends the package as one inline JSON body capped at 64 MB, so it also fails. **Expected behavior** A valid company package of realistic size imports successfully through the Import page. If a package is too large, the error must state the limit clearly and suggest a step that can work. **Steps to reproduce** 1. Export a company with enough attachments to make the compressed package larger than 128 MB. 2. Open the Import page and upload the `.zip`. 3. Click "Preview import". 4. The preview fails with `Import package exceeds 134217728 bytes`. **Deployment mode** Reported from a managed deployment; the limit applies to all deployment modes. ## What Changed - Raise `PORTABLE_ZIP_UPLOAD_LIMIT_BYTES` from 128 MB to a 1 GB default (`server/src/http/body-limits.ts`). - Add the `PAPERCLIP_IMPORT_ZIP_MAX_BYTES` environment override. Invalid or non-positive values fall back to the default. - Scale the zip decompression-bomb guard from the configured cap at the import route: the aggregate inflated ceiling is 4x the cap. The per-entry ceiling stays at 512 MB because V8's string length limit applies to an entry regardless (`server/src/routes/companies.ts`, `packages/shared/src/portability-zip.ts`). - Report the 422 limit error in MB instead of raw bytes. - Replace the "use the CLI folder import for very large packages" hint on preview failure with advice that works: re-export the package without large attachments (`ui/src/pages/CompanyImport.tsx`). - Update the stale comment in `ui/src/lib/import-preflight.ts` that made the same CLI claim. - Add tests for the new default, the env override, and the invalid-override fallback. ## Verification - `pnpm vitest run server/src/__tests__/body-limits.test.ts packages/shared/src/portability-zip.test.ts server/src/__tests__/company-portability-routes.test.ts server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts` — all pass. - `pnpm vitest run ui/src/pages/CompanyImport.test.tsx` — passes, including the updated failure-panel copy assertion. - `pnpm typecheck` — clean across the workspace. - Manual: upload a `.zip` larger than the configured cap; the preview fails with `Import package exceeds the 1024 MB upload limit` and the new hint. A package between 128 MB and 1 GB now previews and imports. ## Risks - Peak per-import memory rises with the cap: the upload is buffered in memory and unzipped in one pass. A 1 GB compressed package can use several GB transiently. Imports are instance-admin actions, so the exposure is a deliberate operator action, not anonymous traffic. Operators on small hosts can lower the cap with `PAPERCLIP_IMPORT_ZIP_MAX_BYTES`. - The aggregate bomb guard moves from a fixed 512 MB to 4x the configured cap. It still bounds expansion far below what a decompression bomb needs. - No migration and no API shape change. The 422 message text changes; no code matches on the old text. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (file edits, local test runs, live-instance inspection). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
19be4cf927 |
refactor(a11y): add scope=col to the agent costs table headers (#1789)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The agent detail page has a Costs section. That section renders a data table of per-run spend. > - The table has a header row, but its `<th>` elements carry no `scope` attribute. > - A screen reader uses `scope="col"` to bind each data cell to its column header. Without it, the reader announces a number without telling the user which column it belongs to. > - A table of costs is exactly the case where that hurts. Every cell is a bare figure. > - This pull request adds `scope="col"` to the five header cells in that table. > - The benefit is that assistive technology announces the cost table correctly. The change is markup only, so sighted users see no difference. ## Linked Issues or Issue Description No public GitHub issue covers this. The problem is described in-PR, following the enhancement template. **What existing behavior does this improve?** The Costs table rendered by `CostsSection` in `ui/src/pages/AgentDetail.tsx`. **Subsystem affected** ui/ — React + Vite board UI **Current behavior** The table renders five header cells: Date, Run, Input, Output, and Cost. None of them set `scope`. A screen reader must guess the header-to-cell relationship, so a user hears a value with no column name attached to it. **Proposed behavior** Each header cell sets `scope="col"`. A screen reader then announces the column name together with each cell, so a cost figure is read as part of the Cost column. **Reason and benefit** `scope` is the standard way to associate header cells with data cells in an HTML table. The attribute has no visual effect, so the fix carries no design cost and makes the table usable with a screen reader. **Breaking changes** None. `scope` is a presentational-neutral HTML attribute. No component API, no styling, and no test changes. **Related pull requests** - #2215 proposed the same attribute for the Routines table. It is closed, because that table no longer exists on master. - #1524 and #1522 applied `scope="col"` to other tables. Both are closed. ## What Changed - Added `scope="col"` to the five `<th>` elements in the `CostsSection` table in `ui/src/pages/AgentDetail.tsx`. - Rebased the branch onto current master. - Dropped the original `ui/src/pages/Routines.tsx` hunks. Master rebuilt the Routines page around folder-grouped rows, so the table those hunks targeted no longer exists. - Dropped the original `HintIcon` opacity change. It altered a visible colour, which is out of scope for a markup-only accessibility fix. ## Verification - Run `pnpm --filter @paperclipai/ui typecheck` and `pnpm --filter @paperclipai/ui build`. This is a markup-only change, so a clean type-check and build is the relevant automated signal. - Open an agent detail page and go to the Costs section. Inspect the header row. Each `<th>` now carries `scope="col"`. - Navigate the same table with a screen reader, cell by cell. Each cell is announced with its column name. - Compare the rendered page before and after. It is unchanged, because `scope` has no styling effect. ## Risks Low risk. The change adds one standard HTML attribute to five header cells in a single table. It introduces no code path, changes no component API, and has no visual effect. The worst case is that the attribute is redundant for a reader that already infers the column, which is harmless. ## Model Used Anthropic Claude Opus 5, exact model ID `claude-opus-5`. It ran with extended thinking and repository read/write tools, inside a maintainer-operated triage agent. The model rebased the branch, dropped the two out-of-scope hunks, and wrote this description. The original change was authored by @bluzername, and the model used for that work is not recorded here. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` references) - [x] I have considered and documented any risks above - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run tests locally and they pass — not run. This is a markup-only change and the package has no test covering this table. - [ ] I have added or updated tests where applicable — no test added, which is why this PR is titled `refactor:`. - [ ] I have updated relevant documentation to reflect my changes — no documentation describes this markup. - [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work — not checked by the maintainer who rebased this. - [ ] All Paperclip CI gates are green — CI re-runs on this push. - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — Greptile re-reviews on this push. --------- Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
cc35c3c395 |
feat: structure and humanize recovery notices (#11075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip posts system comments when automatic run recovery cannot continue > - These comments currently mix the main event with recovery identifiers and routing details > - The task chat shell also renders these comments as large raw text blocks > - Operators need a short explanation first and inspectable evidence on demand > - This pull request emits structured recovery notices and renders them as compact humanized rows > - The benefit is a quieter task thread that keeps the full recovery evidence available ## Linked Issues or Issue Description Related prior extraction source: #11070. This pull request replaces only its structured recovery notice slice with a focused branch based on current master. **What existing behavior does this improve?** Paperclip recovery escalations and the experimental task chat system-comment renderer. **Current behavior** Recovery escalation comments put action identifiers, owner details, run details, and failure codes into the visible markdown body. The task chat shell renders the complete system comment as a large text block. **Proposed behavior** The server emits a short system notice with typed metadata sections. The task chat shell classifies known recovery families and renders one compact row. An operator can expand the row to inspect the full body and metadata. **Reason and benefit** The main thread stays readable during repeated recovery activity. Typed links and evidence remain available without exposing raw failure text in the default view. **Breaking changes** The visible recovery comment body is shorter. Recovery action deduplication now reads the structured metadata and still recognizes legacy body markers. No API schema or database migration changes. ## What Changed - Emit stranded recovery escalations with `system_notice` presentation and typed recovery, owner, run, and failure-code metadata. - Share bounded metadata row builders across recovery notice producers and preserve legacy deduplication compatibility. - Humanize known recovery notice families and render compact expandable task-chat rows. - Route system-authored comments ahead of derived agent authorship so recovery notices do not appear as agent bubbles. - Add focused server and UI regression coverage. ## Verification - `pnpm check:token-gates` — 3/3 clean. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/shared exec vitest run src/validators/issue.test.ts` — 32 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/services/recovery/stranded-notice.test.ts src/__tests__/issue-recovery-actions.test.ts` — 57 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/services/recovery/successful-run-handoff.test.ts src/services/recovery/stranded-notice.test.ts` — 39 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts -t 'escalates an exhausted failed successful-run handoff without using generic continuation recovery first|escalates an exhausted successful handoff run that still leaves no disposition'` — 2 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts -t 'blocks assigned todo work after the one automatic dispatch recovery was already used'` — passed. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/system-notice-humanizer.test.ts src/components/task-chat/TaskChatSystemNotice.test.tsx src/components/task-chat/task-chat-adapter.test.ts` — 15 tests passed. - Storybook visual baselines were not updated because this chat-shell path has no affected snapshot baseline. Focused rendering tests and token gates cover this change. ## Risks - Consumers that parse recovery action identifiers from comment markdown must move to structured metadata. Server deduplication remains backward compatible with legacy comments. - The humanizer uses stable recovery-family phrases. Unknown notices use a generic truncated first-sentence fallback. - The UI changes only the experimental task chat presentation. The stored comment body and expanded metadata remain available. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5, reasoning mode, repository tools, shell execution, and GitHub integration. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4e9a78db58 |
feat(ui): persist task chat composer drafts (#11076)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People send task instructions through the board task chat. > - A page refresh or task switch can discard an unfinished message in the redesigned composer. > - The existing task chat already supplies a task-specific draft key. > - The redesigned composer must use that key without changing attachment or send behavior. > - This pull request restores, saves, and clears text drafts in the redesigned composer. > - The benefit is that users can return to unfinished task messages without losing their text. ## Linked Issues or Issue Description Related prior work: #11070. This pull request extracts only the final composer draft behavior from that larger draft. **Subsystem affected** ui/ — React + Vite board UI. **Problem or motivation** The redesigned task chat composer does not use the draft key that the task thread already provides. A refresh, navigation, or unmount can lose an unfinished message. **Proposed solution** Persist text drafts by task key in local storage. Restore a draft when the composer mounts. Save changes after a short delay and flush pending text during unload or unmount. Clear the draft only after a successful send. **Alternatives considered** The composer could save on every keystroke. A short delay avoids unnecessary synchronous storage writes. The feature could also stay in the larger predecessor PR, but a focused PR is easier to review and verify. **Roadmap alignment** This is a focused usability improvement for the task conversation surface. It does not add or duplicate a roadmap capability. ## What Changed - Added safe draft storage helpers for load, save, and clear operations. - Connected the task-specific draft key to the redesigned task chat composer. - Preserved drafts across debounce windows, unmounts, page unloads, failed sends, and React Strict Mode probes. - Cleared drafts after successful sends without changing current attachment safeguards. - Added focused composer and thread integration tests. ## Verification - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/ui exec vitest run src/components/task-chat/TaskChatComposer.test.tsx src/components/TaskChatThread.test.tsx` ## Risks - Local storage can be unavailable or full. The helpers catch storage errors and keep the composer usable. - Only text is persisted. Attachments, work mode, and assignee selections remain session state. - Draft keys remain task-scoped, so text does not cross task boundaries. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with model `gpt-5`. The context-window size is not exposed in this environment. The model used agentic reasoning, tool use, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1791979512 |
refactor(a11y): add ARIA progressbar attributes to BudgetPolicyCard (#1805)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Humans oversee those agents in teams, so each team needs its own spend controls > - The BudgetPolicyCard component shows how much of a budget is consumed > - The utilization bar in that card is a styled div with no ARIA role, value, or label > - A screen reader user therefore cannot hear how much budget is used > - This pull request adds role="progressbar" and the matching ARIA value attributes > - The benefit is that assistive technology announces budget utilization the same way sighted users see it ## Linked Issues or Issue Description No existing issue. Related pull requests: #1869 and #1878 add the same progressbar semantics to ProviderQuotaCard and QuotaBar. They touch different files. The problem follows the bug report template: **What happened?** The budget utilization bar in `BudgetPolicyCard` renders as a plain `div`. It has no `role`, no `aria-valuenow`, and no accessible name. A screen reader announces nothing for it. The user can read the "Remaining" amount, but not the utilization percentage. **Expected behavior** The bar is announced as a progress bar. It reports the current utilization percentage, with its minimum and maximum. **Steps to reproduce** 1. Open a project or agent page that shows the budget card. 2. Start VoiceOver (Cmd+F5 on macOS). 3. Move the cursor to the budget utilization bar. 4. VoiceOver announces nothing. **Paperclip version or commit** master, `ui/src/components/BudgetPolicyCard.tsx`. **Deployment mode** Local dev (pnpm dev). ## What Changed - Added `role="progressbar"` to the inner bar element. - Added `aria-valuenow` with the rounded utilization percentage. - Added `aria-valuemin={0}` and `aria-valuemax={100}`. - Added `aria-label` in the form `Budget utilization: 73% used`. `aria-valuenow` and `aria-label` use the same `progress` value. That value is already capped at 100 by `Math.min(100, summary.utilizationPercent)`, so the reported value stays inside the min/max range when a scope is over budget. An earlier revision also changed the budget amount `Input` to `type="number"`. That change is removed. It changed input behavior and was not an accessibility fix. See "Risks". ## Verification 1. Open a page that shows the budget card. 2. Start VoiceOver (Cmd+F5 on macOS) and move to the utilization bar. 3. VoiceOver announces "Budget utilization: X% used, progress indicator". 4. Inspect the element. `aria-valuenow` equals the displayed percentage, and it stays at 100 when utilization is above 100%. The change adds attributes only. There is no visual change. ## Risks Low risk. The change adds ARIA attributes to one element in one component. It changes no logic, no styling, and no layout. The `type="number"` change is removed on purpose. With `type="number"`, a browser reports an empty string for text it cannot parse. `parseDollarInput("")` then returns `0` instead of `null`, so the existing "Enter a valid non-negative dollar amount." error never appears and unparseable input silently becomes a $0.00 budget. The `inputMode="decimal"` input keeps that validation path working. ## Model Used Original implementation by @bluzername. The author did not state a model. Scope reduction, branch update, and this description: Claude Opus 5 (1M context, extended thinking, tool use), prepared under maintainer review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change - [x] I have considered and documented any risks above - [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
384e5f6178 |
revert(ui): back out onboarding port (#10786) (#11067)
This reverts commit |
||
|
|
0a511ed1b0 |
feat(apps): support multiple provider connections (#11060)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps subsystem connects company tools through governed provider connections. > - A company can need more than one account for the same provider. > - The current database constraint and Apps flow assume one named connection per company. > - New quarantined actions also need an explicit review decision before activation. > - This pull request supports multiple provider connections and complete action review decisions. > - The benefit is safer access control and a clear multi-account Apps workflow. ## Linked Issues or Issue Description Refs: #11040 **Subsystem affected** Cross-cutting. This change affects the Apps UI, the tool access API, the shared request contract, and the database schema. **Problem or motivation** The connection name constraint prevents a company from keeping more than one connection for a provider. The Apps UI also reuses an existing OAuth connection when a user asks to connect another account. Action review can enable selected entries without recording a decision for every quarantined action. **Proposed solution** Remove the company and connection name uniqueness constraint. Let users open, count, edit, and create multiple provider connections. Require the finish request to cover every quarantined action exactly once before the server activates reviewed entries. **Alternatives considered** The UI could generate unique internal names and keep the database constraint. This would preserve a one-connection assumption in the data model and would make display names part of identity. The server could also infer review decisions from enabled actions. This would not distinguish a reviewed disabled action from an action that the user did not review. **Roadmap alignment** This change extends the completed MCP Tool Gateway and Apps milestone. It also supports the Connected Apps roadmap item. It follows the navigation and connection management work in #11040. ## What Changed - Remove the company-scoped connection name uniqueness index with an ordered and idempotent migration. - Add a reviewed action list to the finish-app contract and reject incomplete or duplicate review decisions. - Activate reviewed entries and keep unreviewed quarantined entries blocked. - Enable a completed connection and preserve the company and connection scope in all updates. - Show provider connection counts and open the provider setup page from Browse. - Let users edit existing connections or connect another account without reusing an active OAuth connection. - Update focused server and UI coverage for multiple connections and action review. ## Verification - Ran the focused Apps UI suite. All 116 tests passed in 11 files. - Ran the focused server and CLI suite. All 276 tests passed in 3 files. - Ran `pnpm --filter @paperclipai/db check:migrations`. The migration safety check passed. - Ran `pnpm -r typecheck`. All projects passed. - Ran `pnpm build`. All projects built successfully. - Ran `pnpm test:run`. It passed 3,735 tests and skipped 4 tests. One worktree-safety assertion failed because the execution workspace reloads its worktree marker. The same test passed with an isolated non-worktree marker. - Ran `pnpm check:token-gates`. It reports 12 existing violations in the unchanged `PaperclipOrbit3D.tsx` file from the target branch. - Started the six affected Playwright specifications. Chromium could not start because the host does not provide `libatk-1.0.so.0`. The GitHub e2e jobs will verify these specifications. - GitHub Actions passed every final-head CI gate, including all three e2e shards and the aggregate `e2e` and `verify` jobs. - Greptile reviewed final commit `9af9200426` at 5/5 with zero review threads. ## Risks - Removing the name uniqueness index permits duplicate display names. Stable connection IDs and UIDs remain unique within a company. - The finish-app endpoint accepts the new review field as optional for backward compatibility. When clients send it, the server requires a complete decision for all quarantined actions. - Multiple OAuth connections depend on the explicit new-connection route flag. Focused tests cover active and draft connection reuse. - The migration is ordered after migration 0210. Its `DROP INDEX IF EXISTS` statement is safe to repeat. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with the `gpt-5.6-sol` model assisted this change. The agent used repository tools, code execution, test execution, and agentic reasoning. The Codex runtime manages the context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b18b0fc39b |
feat: refine app connections and legacy worktree startup (#11040)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps UI manages app discovery and app connections. > - The managed worktree runtime starts agent work in repository worktrees. > - The Apps routes do not match the main discovery flow, and the connections view lacks a delete action. > - Legacy managed worktrees can also start before their pending seed operation runs. > - This pull request makes app discovery the main Apps route and makes connection management explicit. > - It also seeds legacy managed worktrees before runtime startup and makes the CLI read the repository-local config. > - The benefit is a clearer Apps workflow and a safer managed-worktree startup path. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the Apps navigation, app connection management, managed git-worktree startup, and CLI worktree selection. **Subsystem affected** Cross-cutting. The change affects `ui/`, `server/`, `cli/`, and development documentation. **Current behavior** The `/apps` route opens the connections list while discovery uses a nested route. The connections list has no delete action. Some legacy managed worktrees can start runtime work before their pending seed operation runs. The CLI can also read an ambient Paperclip config instead of the repository-local config. **Proposed behavior** The `/apps` route opens Browse, and `/apps/connections` opens the connection list. Users can delete a connection after confirmation. Runtime startup seeds legacy managed worktrees when required. The CLI resolves the current worktree from the repository-local `.paperclip/config.json` file. **Reason and benefit** Users can discover apps from the canonical Apps route and can manage existing connections from a dedicated route. Legacy worktrees receive their required repository content before agent runtime starts. CLI worktree selection stays scoped to the current repository. **Breaking changes** The `/apps` and `/apps/browse` route behavior changes. Old Browse links redirect to `/apps`. The change does not modify an API schema or database schema. ## What Changed - Make Browse the canonical `/apps` page and move the connection list to `/apps/connections`. - Align Apps navigation, redirects, attention links, empty states, and connection actions with the new routes. - Add connection deletion with confirmation and clear failure feedback. - Seed legacy managed git worktrees before runtime startup when their seed status is pending. - Read the CLI worktree selection from the repository-local Paperclip config. - Update focused UI, server, CLI, and development documentation coverage. ## Verification - Ran 202 focused UI, server, and CLI tests. All tests passed. - Ran `pnpm -r typecheck`. All projects passed. - Ran `pnpm build`. All projects built successfully. - Ran `pnpm test:run`. The server and UI stages passed 7,168 tests. The CLI stage found one environment-sensitive secrets test because this workspace injects static AWS credentials. The isolated CLI file passed all 8 tests after those injected variables were unset. - Ran `pnpm check:token-gates`. It reports 12 existing color-token violations in the unchanged `PaperclipOrbit3D.tsx` file from the target branch. This pull request does not modify that file. - Ran focused regression coverage for repository-root CLI config resolution and connection deletion state. All tests and affected package typechecks passed. - Collected all 27 tests in the six changed Playwright specifications successfully. - GitHub Actions passed every latest-head CI gate, including all three e2e shards and the aggregate `e2e` and `verify` jobs. - Greptile reviewed the final commit at 5/5 with zero unresolved threads. ## Risks - Existing bookmarks for `/apps/browse` redirect to `/apps`. - Connection deletion changes visible connection state and requires user confirmation. - The legacy seed path runs only for managed git worktrees with pending seed state. Tests cover the startup condition. - The rebase preserves the target branch's direct OAuth policy for the Notion connection flow. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with the GPT-5 model family assisted this change. The agent used reasoning, repository tools, code execution, and test execution. The runtime does not expose the exact model snapshot or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
01b51dc0e5 |
fix(runtime): stop Live badge and Working shimmer after task teardown (#10985)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task list and the chat views show a Live badge and a Working shimmer for an issue that has an active run. > - A finished task kept the Live badge and the Working shimmer after the run ended and the sandbox stopped. > - The user interface reads run liveness from the `heartbeat_runs.status` row. The run finalizer writes the terminal status in a step that is separate from the agent `status=done` update. When the sandbox or the run process stops between the two steps, `heartbeat_runs.status` stays `running` forever. > - A run row that stays `running` makes a finished task look perpetually Live, and the user interface has no guard for an issue that already reached a terminal status. > - This pull request closes the invariant "environment lease released implies the run is terminal" on the server, and adds a user interface guard that suppresses live state for a terminal issue. > - The benefit is that a finished task stops showing Live and Working, both at the source (the run row) and at the surface (the badge and the shimmer). ## Linked Issues or Issue Description **Bug description** - A completed task kept the Live badge and the Working shimmer after its run ended and the sandbox was torn down. **Steps to reproduce** - Run an agent task to completion. Let the sandbox tear down while the run finalizer is between the `status=done` update and the terminal run-status write. - Open the task list or the chat view for the finished task. **Expected behavior** - A finished task shows no Live badge and no Working shimmer. **Actual behavior (before this change)** - The finished task showed the Live badge and the Working shimmer because its `heartbeat_runs.status` row stayed `running`. This pull request supersedes the two separate pull requests #10954 (frontend) and #10955 (backend). It carries all of their changes for the same race. ## What Changed Server: - Run teardown terminalizes a still-running or still-queued run before it releases the environment lease. It writes `succeeded` when the issue already reached `done`, `cancelled` when the issue is `cancelled`, and `interrupted` otherwise. It never overwrites a status that another path already made terminal. - The recovery stale-lock sweep terminalizes an orphaned running run to `interrupted` after it confirms the process and the sandbox are both gone. It requires recorded process metadata, so it never terminalizes a live run, a queued run, or a scheduled retry. - Each terminal transition writes a run event. - The stale-lock sweep continues and clears the lock when the audit write fails. It logs the failure loudly. - New server tests cover both invariants. User interface: - A shared guard suppresses the Live badge and the Working shimmer when the issue status is terminal. - The guard keeps non-terminal `queued` and `running` issues live. - The guard prefers the newest issue live-status snapshot. - New user interface tests cover the guard and the snapshot preference. ## Verification Server: - `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && tsc --noEmit` in `server/` — 0 errors. - `pnpm --filter server test heartbeat-run-lease-release-terminalization.test.ts recovery-stale-issue-lock-sweep.test.ts` — 12 tests pass. User interface: - `pnpm --filter @paperclipai/ui typecheck` — 0 errors. - `pnpm exec vitest run ui/src/lib/liveIssueIds.test.ts ui/src/lib/issue-chat-messages.test.ts` — 40 tests pass. ## Risks - Low risk. The server change only forces a still-live run row to a terminal status when the lease releases or when the recovery sweep confirms the process is dead. It never overwrites an existing terminal status, and it guards the recovery path with process metadata to avoid terminalizing a live run. - The user interface change is additive. The guard only suppresses live state for a terminal issue and keeps queued and running issues live. - No database migration. No change to any external endpoint. ## Model Used - Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
11e56654f8 |
feat(ui): port onboarding flow from prototype; add cloud + local variants (#10786)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - First-run onboarding is the subsystem that turns a brand-new install
into a working company: it creates the company, its goal, a lead agent,
and that agent's first task
> - The existing `OnboardingWizard` carried all of that wiring
correctly, but its UI had drifted from the current design direction, and
a separate design prototype (`paperclip-onboard`) existed as a
standalone visual mock with no backend
> - Porting the prototype's *logic* would have thrown away working,
well-tested backend orchestration; leaving the two apart meant the
design never shipped
> - Separately, cloud and local (self-hosted) installs need meaningfully
different first runs — local has no sign-in and must let the user pick a
locally-installed CLI adapter — so a single linear wizard could not
serve both
> - This pull request rebuilds the presentational layer from the
prototype on top of the existing backend orchestration, and splits it
into two thin flow containers over a shared core
> - The benefit is that the shipped onboarding matches the intended
design, cloud and local can diverge without duplicating logic, and each
can later ship to a different app version while sharing one set of step
components
## Linked Issues or Issue Description
No existing issue — describing inline (feature request).
**What problem does this solve?**
Onboarding is the first thing a new user sees, and the shipped wizard
had drifted from the current design. In parallel, cloud and local
installs need different first-run paths: local has no hosted sign-in,
and its agent runs on a CLI adapter installed on the user's machine,
which the cloud path never has to ask about. There was no way to express
that difference without either forking the whole wizard or bolting
conditionals onto a single linear flow.
**Proposed solution**
Extract the onboarding step views and shell into a shared core, then
compose two thin flow containers (cloud and local) over it. Keep all
backend orchestration in the existing `useOnboardingFlow` hook so no
working logic is rewritten.
**Alternatives considered**
- *Single flow with a `variant` prop* — most DRY, but the two flows are
intended to ship on different app versions, and a shared file would have
to be split later anyway.
- *Two fully independent copies* — simplest per-flow, but every shared
refinement (spacing, motion, copy) would have to be made twice and would
drift.
## What Changed
- **Shared core** under `ui/src/components/onboarding/`:
`OnboardingScaffold` owns the full-screen shell and the single
`AnimatePresence` step crossfade, so both flows transition identically;
step views (Start / Company / Agent / Task), `FooterNav`, `AgentPreview`
and the motion constants are extracted for reuse.
- **`CloudOnboardingFlow`** — `start → company → agent → task`; mounted
in the real app via `OnboardingWizardVariant`. Behaviour matches the
retired wizard, including `previewMock` and the existing-company ("add
an agent") entry point.
- **`LocalOnboardingFlow`** — skips sign-in and adds an optional email
ask (with a privacy assurance), a local model/adapter step that hires
with `requireEnvProbe: true`, and a "star us on GitHub" interstitial
before completing. **Harness-only for now** — the real app still mounts
the cloud flow.
- **Deleted `OnboardingWizard.tsx`** (1,786 lines); updated its
Storybook stories and the `OnboardingWizardVariant` test to the new
components.
- **Orbiting 3D paperclip backdrop** behind the auth and welcome screens
(`three`), code-split so it only downloads on those screens; honours
`prefers-reduced-motion` and disposes its GL context on unmount.
- **`motion`** added for step transitions and the agent-capsule
choreography.
- Visual values routed through design tokens per `DESIGN.md`; `Stepper`
generalized to take a step total (backward compatible); `/design-guide`
page and the component index updated.
- **Standalone preview harness** (`ui/onboarding-preview.html`) with
`?flow=` and `?step=` for backend-free review, wired as a second Vite
rollup input.
- **Adapter env probe bound to the adapter it ran against.**
`hireLeadAgent` reused `adapterEnvResult` for any adapter, so when a
hire failed and the user picked a *different* local adapter and retried,
the previous adapter's verdict satisfied the `requireEnvProbe` guard
while the hire posted the new adapter's config — hiring it unprobed. The
cache is now keyed on the adapter type plus the exact config posted to
the test endpoint, the config is built once and shared by probe and
hire, a failed probe clears the cache, and `clearAdapterEnvResult()`
(called on adapter change) stops the step displaying a stale verdict.
Cloud is unaffected — it hires with `requireEnvProbe: false`. Reported
by Greptile.
- **E2E specs re-pointed at the new flow.** Four specs still drove the
deleted wizard (`onboarding`, `conference-room-typing-intro`,
`planning-mode-visual-verification`, `nux-phase4-screenshots`) and
failed with `element(s) not found` on `"Name your company"` /
`input[placeholder="Acme Corp"]`. Rather than repeat the new drive
sequence four times, `tests/e2e/onboarding-flow.ts` adds one driver per
step (`startCloudOnboarding`, `completeCompanyStep`,
`completeAgentStep`, `completeTaskStep`, `completeCloudOnboarding`) and
the specs import it, so the next flow change touches a single file. Two
now-dead `**/test-environment` route stubs went with it — the cloud flow
hires with `requireEnvProbe: false`, so that probe never fires.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `npx vitest run` over the onboarding suites
(`OnboardingWizardVariant`, `AgentCapsule`, `onboarding-launch`,
`onboarding-goal`, `onboarding-route`, `onboarding-adapter-config`) — 33
tests pass.
- `pnpm --filter @paperclipai/ui build` — succeeds; the three.js chunk
splits out separately (522 kB raw / 133 kB gzip) rather than entering
the main bundle.
- Both flows driven end-to-end in the preview harness in `previewMock`
(no database writes), plus the cloud flow rendered in the real
authenticated app at `/onboarding` to confirm the mount swap.
- The four re-pointed e2e specs pass locally against the new flow.
- New `ui/src/hooks/useOnboardingFlow.test.tsx` — 4 cases pinning the
adapter-probe cache (switch-adapter retry, cold path, explicit clear,
and the cloud flow's `requireEnvProbe: false`). Verified non-vacuous:
the switch-adapter case fails against the pre-fix code.
- Rebased onto current `master`; `pnpm-lock.yaml` is deliberately
**not** committed — `.github/workflows/pr.yml` regenerates it when a
manifest changes and shares it with downstream jobs as the `pr-lockfile`
artifact.
## Risks
- **Deleting `OnboardingWizard.tsx` is the one change that alters
existing app behaviour.** The cloud flow is intended to be
behaviour-equivalent, and its entry points are covered by the updated
`OnboardingWizardVariant` test, but this is the area to review most
closely.
- **Conflict risk with open PRs that touch the old wizard**: #9900,
#9501, #8982 and #6636 all modify
`ui/src/components/OnboardingWizard.tsx`, which this PR removes.
Whichever lands second will need its change re-applied to the new step
components. Flagging so ordering can be decided deliberately.
- **New dependencies**: `motion` and `three` (+ `@types/three`). `three`
is large, so it is lazily imported and code-split — it does not affect
the main bundle. Both are MIT.
- The **local flow is not reachable in the app** yet (harness/canary
only), so it carries no runtime risk today; wiring it up is a follow-up.
- The auth screens remain **presentational only** — they are not wired
to real auth, unchanged from before this PR.
- **Pre-existing, not introduced here:** `OnboardingWizardVariant`
renders outside `<Routes>` in `App.tsx`, so its `useParams()` never
resolves `:companyPrefix` and `/{prefix}/onboarding` opens the welcome
screen instead of jumping to the agent step. `master` has the identical
structure, so this PR faithfully ports existing behaviour; the working
"add an agent" entry is the launcher card behind the overlay, which is
what the screenshot spec drives. Worth a separate fix.
## Model Used
Claude Opus 5 (`claude-opus-5`) via Claude Code, with extended thinking
and tool use (repo search/edit, local test + build execution, and
browser-driven visual verification of the rendered flows). Portions of
the session also ran on `claude-opus-4-8` and `claude-fable-5`.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
b67c512f82 |
feat(ui): live run label → run detail, running row → task detail (#11034)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The agent detail page has a Dashboard tab. The Dashboard tab shows a "Live Run" section for the agent's current heartbeat. > - The "Live Run" section has two clickable pieces: the section heading and the running row. Both pieces linked to the same run detail page. > - Two controls that go to the same place waste a navigation affordance and hide the task the agent runs. > - This pull request splits the two destinations. The heading goes to the run. The running row goes to the task. > - The benefit is that a user reaches the run internals from the label and the work item from the row, in one click each. ## Linked Issues or Issue Description No public GitHub issue exists for this change. The description follows the enhancement issue template. **What existing behavior does this improve?** The "Live Run" section on the agent detail page, Dashboard tab (the `LatestRunCard` component in `ui/src/pages/AgentDetail.tsx`). **Current behavior** The "Live Run" heading and the running row both link to the run detail page (`/agents/:agentId/runs/:runId`). The row shows the run code and an invocation-source chip. There is a separate "View details →" link that also goes to the run detail page. A user cannot reach the task the run works on from this section. **Proposed behavior** The heading becomes a link to the run detail page and appends the short run code, shown as `Live Run · <run code>`. The redundant "View details →" link is removed. The running row links to the task detail page when the run's context snapshot resolves to a known issue, and the row then shows the task status glyph, the task slug, and the task title. A pure timer heartbeat with no resolvable task keeps the previous behavior: run code plus source chip, linking to the run detail page. **Reason and benefit** The heading and the row now go to distinct, intuitive destinations. A user reaches the run internals from the label and the work item from the row, each in a single click. The fallback keeps heartbeats with no task readable and avoids a blank or broken row. ## What Changed - Made the "Live Run" / "Latest Run" heading a `Link` to the run detail page and appended the short run code (`run.id.slice(0, 8)`) in a mono span, formatted `Live Run · <run code>`. Kept the pulsing live dot. - Removed the redundant "View details →" link. - Changed the running row `Link` target to the task detail page (`/issues/:identifier`) when a task resolves, falling back to the run detail page otherwise. - Resolved the task from the run context snapshot (`contextSnapshot.issueId`, falling back to `contextSnapshot.taskId`) against a `Map` of the agent's assigned issues threaded in from `AgentOverview`. - When a task resolves, replaced the run code and source chip in the row with the task status glyph (`StatusGlyph`), the task slug, and the task title. Kept the running spinner, the run status badge, and the timestamp. ## Verification - `pnpm check:token-gates` → 3/3 gates clean. - `pnpm --filter ui typecheck` → passes. - `pnpm --filter ui exec vitest run src/pages/AgentDetail.progress.test.ts src/pages/AgentDetail.instructions.test.tsx` → 10/10 pass. - Manual (needs a reviewer with a browser): open an agent detail page → Dashboard tab. - For a live issue-execution run: the heading reads `Live Run · <run code>` and opens the run detail page; the row shows the task status icon, slug, and title and opens the task detail page. - For a pure timer heartbeat with no task: the row falls back to run code + source chip and opens the run detail page. No blank row. - Confirm both states in light and dark mode. ## Risks Low risk. The change is presentational and scoped to one component. The task lookup is defensive: it reads the context snapshot with a fallback key and only renders the task row when the issue is present in the already-loaded assigned-issue set, so an unknown or missing issue degrades to the previous run-detail behavior rather than breaking. ## Model Used - Provider: Anthropic (Claude). - Model: claude-opus-4-8 (Opus 4.8). - Context window: 200K. - Reasoning mode: extended thinking, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5da382fd59 |
feat(skills): require explicit merge modes (#10978)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can select company skills and synchronize them to adapter runtimes > - The skill sync API replaced the complete selection without an explicit destructive choice > - Company package import also replaced conflicting skills by default > - These defaults could remove operator edits during setup and import reruns > - This pull request adds explicit assignment merge modes and safe package conflict handling > - The benefit is that reruns preserve operator work unless the caller explicitly requests replacement ## Linked Issues or Issue Description **What existing behavior does this improve?** This change improves agent skill synchronization and company package import. **Subsystem affected** This is a cross-cutting change across the shared contracts, server, CLI, and UI. **Current behavior** Agent skill synchronization replaces the full desired skill set from a modeless request. Package import replaces a conflicting skill when the caller does not select a conflict mode. **Proposed behavior** Agent skill synchronization requires `add`, `remove`, or `replace`. Package import skips conflicts by default. Each imported skill reports whether it was created, renamed, replaced, or skipped. **Reason and benefit** Setup and import reruns must preserve operator edits by default. Explicit destructive modes make data loss less likely and make each outcome inspectable. **Breaking changes** Callers of the agent skill sync API must now send `mode`. Callers that need the former behavior must send `replace`. Package import now uses `skip` when `onConflict` is absent. ## What Changed - Added required `add`, `remove`, and `replace` modes to the shared agent skill sync contract. - Added actionable `422` validation for missing or invalid modes. - Updated first-party UI and CLI callers with explicit modes. - Changed package skill conflict handling to use `skip` by default. - Kept plugin-owned and built-in stock skill imports on explicit `replace`. - Added created, renamed, replaced, and skipped results to company imports. - Added regression coverage for merge modes and package conflict outcomes. ## Verification - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run:serialized` (128 suites passed) - `pnpm --filter @paperclipai/skills-catalog test` (20 tests passed) - Focused agent skill route, company skill service, portability, CLI, and UI tests passed. - GitHub CI passed build, typecheck, canary, all general and serialized test shards, all browser shards, policy, security, and final verification on commit `2cfbb3e4c5`. - Greptile reviewed the latest commit at 5/5 with zero unresolved threads. ## Risks - This change intentionally rejects modeless agent skill sync requests. - The safe package default can leave an existing skill unchanged where the old default overwrote it. - All first-party callers now select a mode. Regression tests cover each outcome. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.6-sol` through Codex. The runtime used agentic reasoning, tool use, code execution, and repository editing. The runtime did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
03cfad7ceb |
feat(apps): connect Notion through MCP OAuth (#11009)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents governed access to external tools. > - The Apps gallery lists Notion, but the server required manually configured OAuth credentials. > - Notion's hosted MCP server supports OAuth discovery and dynamic client registration. > - Notion also requires HTTPS or a loopback HTTP redirect URI. > - This pull request adds a direct Notion MCP OAuth path with PKCE and reusable dynamic clients. > - It also adds the current Apps UI states for connect and reauthorization. > - The benefit is a secure Notion connection with no manual client credential setup. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Apps gallery, Apps connect route, OAuth token lifecycle, and managed MCP gateway. **Subsystem affected** `server/`, `packages/shared/`, `scripts/`, and `ui/`. **Current behavior** The Notion gallery cards are disabled. The server uses the classic Notion OAuth endpoints and requires operator-supplied client credentials. It does not register an OAuth client from provider metadata. Concurrent refreshes can also replay a rotating refresh token. **Proposed behavior** Enable the Notion Apps flow. Discover OAuth metadata from `https://mcp.notion.com/mcp`. Register and reuse a public RFC 7591 client with PKCE. Require HTTPS or loopback HTTP callbacks. Serialize refreshes, store each rotated refresh token before the new access token can be used, and show a reconnect state for `invalid_grant`. **Reason and benefit** Operators can connect the built-in Notion MCP app without creating or copying OAuth credentials. Paperclip keeps dynamic clients and rotating tokens in the company secret store. **Breaking changes** None. Explicit environment client credentials still take priority. Existing Slack and Linear OAuth endpoint hints remain unchanged. Other OAuth apps remain disabled unless they are allowlisted. **Additional context** PR #10910 is a related, broader Connections v3 wizard replacement. This PR is the focused current Apps flow. The MCP Tool Gateway and Connected Apps items in `ROADMAP.md` cover this planned capability. ## What Changed - Classify all 20 reviewed Notion MCP tools with provider-scoped read and write defaults. - Require approval for selected Notion mutations, including move, duplicate, and convert actions that generic verb matching missed. - Preserve company-scoped connection and catalog resolution for Notion profiles and policies. - Add RFC 7591 dynamic client registration with `token_endpoint_auth_method=none` and mandatory PKCE. - Store the dynamic client ID on the connection and store any returned client secret in the company secret store. - Reuse the registered client for later connects and keep explicit environment credentials as the first choice. - Discover protected-resource and authorization-server metadata from the Notion MCP endpoint. - Add `redirectConstraints: "https-or-loopback-http"` to the generated Notion app definition and shared contract. - Reject non-loopback plain HTTP callbacks before network access with a TLS setup error. - Serialize client registration and token refresh operations within the server process. - Store a rotated refresh token before publishing the refreshed access token. - Treat `invalid_grant` as terminal and move the connection to a clear reauthorization state. - Add focused coverage for registration reuse, callback constraints, refresh rotation, and terminal grants. - Enable the Notion Apps route and add connect, redirect, success, error, and reconnect UI states. - Keep non-allowlisted OAuth apps blocked and cover the UI policy with regression tests. ## Verification - The focused Notion policy integration test passed with embedded PostgreSQL. - The focused 20-tool classification test passed. - The server typecheck passed on the governance head. - `pnpm -r typecheck` passed on the rebased head. - `pnpm --filter @paperclipai/server typecheck` passed after the security follow-up. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts -t \u0027DCR|refresh tokens|invalid_grant|abandoned lease\u0027` passed 10 focused security tests. - `pnpm build` passed on the rebased head. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts -t 'OAuth|oauth'` passed 14 tests. - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts` passed 5 tests. - The complete server group passed 3,686 tests with 4 skipped. - The complete UI group passed 3,656 tests. - The full local runner found one environment-only CLI failure because this agent runtime injects static AWS credentials into a test that expects `AWS_PROFILE` only. `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` passed all 8 tests. - The prior UI verification passed 55 focused tests, `pnpm check:token-gates`, the Storybook build, and review of six 1440 x 1000 screenshots. - OAuth request sequence: protected-resource metadata `GET https://mcp.notion.com/.well-known/oauth-protected-resource/mcp`; authorization metadata `GET https://mcp.notion.com/.well-known/oauth-authorization-server`; dynamic registration `POST https://mcp.notion.com/register`; authorization `GET https://mcp.notion.com/authorize`; token exchange and refresh `POST https://mcp.notion.com/token`; MCP traffic `POST https://mcp.notion.com/mcp`. - The live metadata and registration probe confirmed that Notion accepts HTTPS and loopback HTTP redirects. It rejects a plain HTTP private hostname. - A later QA task owns the full browser consent and managed gateway tool-list dry run against a configured HTTPS deployment. ## Risks - Notion can add tools. Unrecognized names use the generic classifier, and new or changed risky tools stay quarantined after connection activation. - A deployment that uses a private non-loopback hostname must configure HTTPS before it can connect Notion. - Dynamic registration creates a provider-side client. Paperclip reuses it because registration does not provide a standard delete operation. - Refresh coordination uses a database CAS lease across service instances. An unclean crash leaves an uncertain lease and requires reconnect instead of risking refresh-token replay. - The current Apps surface overlaps with PR #10910. Merge order can require a small conflict resolution if that PR lands first. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex on a GPT-5 runtime. The exact deployment ID and context window are not exposed. The runtime used reasoning, repository tools, code execution, and network tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ea83c5c822 |
feat(task-chat): bring back copy/👍/👎 actions on the agent bubble footer (#11025)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task view shows a threaded conversation between the user and the agent, with agent replies rendered as bubbles that carry a "✓ Worked · N tools" summary line. > - The conference-room chat already offers per-message copy and thumbs-up / thumbs-down feedback, but the redesigned task thread dropped these controls from the agent bubble footer. > - Users lose a quick way to copy an agent reply or send feedback on it, and the redesign silently ignored the feedback-vote props it was already given. > - This pull request prepends a copy · thumbs-up · thumbs-down cluster to the bubble summary line and wires the existing feedback-vote props through. > - The benefit is a consistent feedback surface across both chat views, with no new API. ## Linked Issues or Issue Description No public GitHub issue exists for this change; the underlying issue is described inline below following the feature template (`.github/ISSUE_TEMPLATE/feature_request.yml`). #### Problem or motivation The redesigned task thread renders each agent reply with a "✓ Worked · N tools · <timestamp>" summary line, but it dropped the copy and thumbs-up / thumbs-down controls that the conference-room chat still shows. Users can no longer copy an agent reply or vote feedback from the task thread. The redesign component already received `feedbackVotes` and `onVote` props but ignored them. #### Proposed solution Prepend a copy · 👍 · 👎 cluster to the summary line, leading the always-visible timestamp, reusing the shared `IssueChatFeedbackButtons` so both chat views speak the same feedback language. Anchor the cluster to the turn's summary row (a sibling of the expandable tool-history fold) so it stays on the summary line whether the tool history is collapsed or expanded. #### Alternatives considered Placing the cluster inside the expandable fold — rejected because expanding the tool history then re-centered the actions to the middle of the tall fold. #### Roadmap alignment UI polish to the task thread; no core-roadmap overlap. ## What Changed - Add `TaskChatBubbleActions`: a copy · thumbs-up · thumbs-down cluster built on the shared `IssueChatFeedbackButtons`. - Render the cluster on the agent bubble's "✓ Worked · …" summary line, leading the timestamp; runless agent replies get the same cluster with the timestamp trailing. Human and system bubbles are unchanged. - Add a `leading` slot to `TaskChatTurn` so the actions sit on the summary row, a sibling of the tool-history fold, and stay anchored when the fold expands. - Wire the redesign to the `feedbackVotes` / `onVote` props it already received. - Add a demo binding in the `TaskChatLab` dev harness. ## Verification - `pnpm check:token-gates` — 3/3 gates CLEAN. - `pnpm typecheck` — clean across all packages. - `cd ui && pnpm vitest run src/components/task-chat/TaskChatBubble.test.tsx src/components/task-chat/TaskChatTurn.test.tsx` — 31/31 pass. - Manual: open a task thread, confirm the copy / 👍 / 👎 cluster shows on the agent bubble summary line before the timestamp, copy works, votes toggle, and the cluster stays on the summary line when the tool history is expanded. Visual change: snapshot baselines are intentionally not updated, per the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks Low risk. UI-only change scoped to the redesigned task-chat bubble footer. It reuses an existing shared feedback component and existing vote props; no API, schema, or server change. Human and system bubbles are untouched. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended thinking enabled, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d84c5eae7a |
feat(ui): hide task priority from the UI (keep data model) (#11024)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks and issues carry a `priority` field that renders across many product surfaces: the detail header, the Triage properties panel, Kanban and thread cards, the New Task composer, list Sort/Group/Filter menus, search filters, and the dashboard chart. > - Product feedback found the priority level adds visual noise and decision cost without clear value in day-to-day task flow. > - We want to remove priority from the interface, but keep the data model, API, validation, and search DSL fully intact so the choice is reversible with no migration. > - This pull request hides every priority indicator and control behind one compile-time flag, `SHOW_TASK_PRIORITY_UI`, set to `false`. > - The benefit is a calmer, simpler UI now, with a single-boolean revive path and zero data loss. ## Linked Issues or Issue Description <!-- No public GitHub issue exists. Describing in-PR per the feature template. --> **Subsystem affected** The web UI (`ui/src`): issue detail, properties panel, Kanban/thread cards, New Task dialog, issues list Sort/Group/Filter menus, search filter bar/sheet, dashboard charts, and the design-guide showcase. **Problem or motivation** The task/issue priority level appears across many surfaces and adds visual clutter and decision overhead without pulling its weight in normal task flow. We want it gone from the interface without discarding the underlying data or breaking anything that depends on it. **Proposed solution** Add a single compile-time UI flag, `SHOW_TASK_PRIORITY_UI` (default `false`), and gate every priority indicator and control behind it. Leave the data model, API params, Zod validation (including the `"medium"` default), and the search filter DSL untouched. Reviving priority is a one-line flip of the flag back to `true`. **Alternatives considered** Deleting the priority code and schema outright. Rejected: it is irreversible, needs a data migration, and throws away a field the API and search still support. A gated flag keeps the change reversible and low risk. **Roadmap alignment** UI simplification. This is a presentation-only change; it does not alter core agent or data behavior. ## What Changed - Added `ui/src/lib/ui-flags.ts` exporting `SHOW_TASK_PRIORITY_UI: boolean = false` (typed `boolean` so gated branches are not flagged as dead code). - Gated the priority row in the Triage properties panel and the editable priority control in the issue detail header (plus its skeleton seed). - Gated the per-card priority icon in `KanbanBoard` and in `IssueThreadInteractionCard`. - Hid the priority chip and the mobile "more" menu priority section in the New Task dialog. The submit path still sends the `"medium"` default. - Removed the Priority options from the issues list Sort and Group-by menus; the comparator and grouping logic stay dormant. - Hid the Priority sections in the issue filters popover and in the search filter bar and sheet. The `priority:` search DSL and filter state stay functional at the data layer. - Suppressed active-filter priority pills for consistency. - Gated the "Tasks by Priority" dashboard chart and the design-guide priority showcase subsection. - Left activity-feed "changed priority" history text intact as a historical record. - Updated call-site tests to assert priority UI is absent while the flag is off, added focused hidden-surface tests, and added a test that proves creating a task still persists `priority: "medium"`. ## Verification - `pnpm check:token-gates` — all 3 gates clean. - `pnpm --filter @paperclipai/ui typecheck` — clean. - `pnpm --filter @paperclipai/ui exec vitest run` on the touched surfaces (IssueProperties, IssueFiltersPopover, IssuesList, NewIssueDialog, IssueDetail, PriorityIcon and its interaction test) — all green under `TZ=UTC`. - Manual: with the flag off, priority does not appear in the detail header, Triage panel, New Task composer, Sort/Group/Filter menus, or the dashboard chart. Creating a task still persists `priority: "medium"`, and the `priority:` search token still filters at the data layer. ## Risks Low risk. The change is presentation-only and additive: no data model, API, validation, or search-DSL changes. The priority code paths remain compiled and tested; flipping `SHOW_TASK_PRIORITY_UI` to `true` restores the full UI. Visual snapshot baselines are intentionally not updated per the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended thinking enabled, with tool use / code execution in an agentic coding harness. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f258b34bbd |
fix(ui): unify cloud-managed sign-out (#10994)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip can run as a self-hosted app or as a Cloud-managed tenant. > - These modes need different sign-out sequences because Cloud owns three sessions. > - Several visible controls implemented sign-out separately and could choose different paths. > - This pull request adds one Cloud-aware sign-out action and moves every visible control to it. > - The benefit is one safe sign-out path in Cloud and unchanged local sign-out in self-hosted deployments. ## Linked Issues or Issue Description Refs #2073. That older PR adds a separate company-settings sign-out surface. This change centralizes the existing account, company, and instance-settings surfaces and preserves the self-hosted behavior described there. **What happened?** Visible sign-out controls used separate implementations. A Cloud-managed control could call the app-local sign-out endpoint and open the local auth page. That path did not enter the Cloud-owned logout sequence for the tenant, Cloud, and identity sessions. **Expected behavior** Every visible sign-out control must use one action. Cloud-managed instances must navigate the top-level window to the same-origin `/cloud/logout` route without a local sign-out call first. Authenticated self-hosted instances must keep the local sign-out API and cache invalidation behavior. **Steps to reproduce** 1. Open a Cloud-managed tenant. 2. Use the account menu, company menu, or instance-settings sign-out control. 3. Observe that independently implemented controls can enter different sign-out paths. **Paperclip version or commit** The problem reproduces at `656ecfa585b31938e2685ffab3db22e794474803`. **Deployment mode** Cloud-managed tenant built from source. The regression tests also cover authenticated self-hosted mode. ## What Changed - Added `useSignOut` as the shared Cloud-aware sign-out action. - Navigated Cloud-managed sessions to `/cloud/logout` exactly once without calling local auth first. - Preserved local API sign-out and cache invalidation for authenticated self-hosted sessions. - Migrated the account menu, company menu, and instance general settings to the shared action. - Added focused tests for mode selection, menu closure, pending state, failure state, and settings behavior. ## Verification - `pnpm exec vitest run ui/src/hooks/useSignOut.test.tsx ui/src/components/SidebarAccountMenu.test.tsx ui/src/components/SidebarCompanyMenu.test.tsx ui/src/pages/InstanceGeneralSettings.test.tsx` — 26 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — passed. - `git diff --check origin/master..HEAD` — passed. ## Risks - Low risk. The change centralizes existing behavior and adds no schema, API, telemetry, or style-token changes. - Cloud mode depends on the existing health decision. Tests pin both mode branches. - The change does not alter Fetch Metadata, CSRF, cookie, prefetch, or return-URL protections owned by the Cloud logout route. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The runtime does not expose a more specific model ID or context-window size. The agent used high-reasoning mode, shell tools, and API tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
656ecfa585 |
fix(server): keep Date fields intact through secret redaction; harden chat notice timestamps (#10984)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The task chat thread renders issue comments, system notices, and run
transcripts
> - The server routes comment payloads through the run-secret redaction
walker before it sends them
> - The walker rebuilds each object with `Object.entries`, and this
collapses `Date` instances to `{}`
> - The chat renderer then calls `.toISOString()` on an invalid date and
throws, and the thread falls back to the error banner
> - This pull request keeps `Date` instances intact in redacted
responses and makes the renderer safe against bad timestamps
> - The benefit is that task threads with system notices render
correctly again
## Linked Issues or Issue Description
**What happened**
Task threads that contain a system notice showed the banner "Chat
renderer hit an internal state error." in place of the conversation.
This occurred on many tasks.
**Expected behavior**
The thread renders all comments and system notices with correct
timestamps.
**Steps to reproduce**
1. Open a task that has at least one system notice comment (for example
a "Workspace ready" notice).
2. `GET /api/issues/{id}/comments` returns `createdAt: {}` for every
comment because the secret-redaction walker collapses `Date` objects.
3. The system-notice row calls `new Date({}).toISOString()`. This throws
`RangeError: Invalid time value` and trips the thread error boundary.
**Version / deployment**
Regression from #9934 (`e43f187ca`). It applies to all deployments that
include that commit.
## What Changed
- `server/src/services/run-secret-redaction.ts`:
`redactRegisteredSecretValues` now returns `Date` instances as-is. Dates
hold no redactable text, and the `Object.entries` rebuild turned them
into `{}`.
- `ui/src/components/IssueChatThread.tsx`: the system-notice row formats
its timestamp with a new `toValidIsoString` helper. A value that does
not parse as a date now degrades to "no timestamp" instead of a render
crash.
- Regression tests at three layers:
- Walker unit tests: `Date` values survive with and without registered
secret values.
- Route test: `GET /issues/:id/comments` serializes `createdAt` /
`updatedAt` as ISO strings.
- Render test: a system notice with a malformed `createdAt` renders
without the error boundary.
## Verification
- `npx vitest run --root server
src/__tests__/run-secret-redaction.test.ts` — 5 passed.
- `npx vitest run --root server
src/__tests__/issue-comment-redaction.test.ts` — 4 passed (embedded
Postgres route test).
- `cd ui && npx vitest run src/components/IssueChatThread.test.tsx
src/components/IssueChatThreadSystemNotice.test.tsx
src/lib/issue-chat-messages.test.ts` — 121 passed.
- Each new test was run against the unfixed code and failed there, which
confirms it guards the regression.
- A local sweep rendered 47 real issue threads through
`IssueChatThread`: 7 tripped the boundary before the fix, 0 after.
## Risks
- Low risk. The server change only preserves `Date` objects that the
walker destroyed before. String redaction behavior does not change, and
the registry-key stripping does not change.
- The UI change only affects the timestamp of system-notice rows and
omits it when the value is invalid.
## Model Used
- Claude Fable 5 (`claude-fable-5`), Anthropic. Agentic coding session
with tool use (file edits, shell, Vitest). No extended-context or
special reasoning mode.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
2ea22d6cea |
fix(ui): white text on light-mode user chat bubbles (#10952)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task detail view shows the conversation as chat bubbles. The requester's own messages sit in a solid accent-colored bubble. > - The bubble container sets `text-white`, but the message body renders through `MarkdownBody`. Tailwind prose tokens (`--tw-prose-body`) win over the inherited container color. > - `prose-invert` only lightens the prose text in dark mode. In light mode the prose body stayed its default dark color, so the text read as near-black on the blue bubble and was hard to read. > - This pull request maps the human bubble's prose tokens to the inherited text color in both themes. > - The benefit is that the requester's chat text is readable white-on-blue in light mode, and dark mode stays exactly as it was. ## Linked Issues or Issue Description <!-- No public GitHub issue exists; described in-PR per the bug report template. --> **What happened?** In light mode, the text inside the user's own chat bubbles in the task detail view rendered as dark (near-black) on the solid blue accent background. This made the requester's messages hard to read. **Expected behavior** The text inside the user's accent-colored chat bubbles should be white in light mode, matching the bubble's `text-white` intent. Dark mode already rendered correctly and should not change. **Steps to reproduce** 1. Open a chat-style task detail view in light mode. 2. Post a message as the requester (human) so it renders in the solid blue accent bubble. 3. Observe the body text renders dark on blue instead of white. **Paperclip version or commit** Reproduces on `master` (branched from `814cb3367`). **Agent adapter(s) involved** Not adapter-specific (core UI bug). ## What Changed - Add the existing `paperclip-markdown-on-accent` class to the human-branch `MarkdownBody` in `TaskChatBubble.tsx`. This class (already used by `IssueChatThread` for the same accent bubble) maps prose body/heading tokens to `currentColor`, so the text follows the bubble's `text-white` in both themes. - Apply the same class to the human-branch `MarkdownBody` in `TaskChatDescriptionBubble.tsx` (the description-as-first-bubble surface) for consistency. - Add unit tests covering that the human accent bubble carries the on-accent class and the agent/neutral bubbles do not. ## Verification - `pnpm check:token-gates` → 3/3 CLEAN. - `pnpm --filter ./ui vitest run src/components/task-chat/TaskChatBubble.test.tsx` → 9/9 passing. - Manual: in light mode, the requester's chat bubble text renders white on blue; agent/neutral bubbles unchanged; dark mode unchanged. This is a visual change. Snapshot baselines are intentionally not updated, per `doc/design/DECISION-SHEET.md` → "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks Low risk. The change is scoped to the human-branch `MarkdownBody` className on two chat-bubble components and only remaps prose color tokens to the inherited text color. Agent and neutral bubbles are untouched, and dark mode behavior is unchanged. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, ~200K context window, extended thinking mode, with tool use / code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f554d67377 |
fix(server): add explicit review verdict policies (#10931)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Issues use `in_review` to request a final decision from an authorized writer > - The server rejected an assignee agent that tried to close its own review, even when the issue had no independent-review rule > - This rejection stopped the default agent workflow and did not represent the configured execution-stage rules > - Paperclip needs an open default and explicit issue-level constraints for teams that require an independent or human verdict > - This pull request removes the unconditional rejection and adds `anyone`, `not_creator`, and `human_only` review policies > - The benefit is a working default path with opt-in, authenticated verdict controls ## Linked Issues or Issue Description Refs #10635, #4429, and #10671. The related public work covers execution-stage independence, self-approval fallback behavior, and durable review paths. This change is distinct. It controls who can resolve an issue review verdict. It keeps configured execution stages active. ## What Changed - Added a nullable `review_policy` issue column. Null has the same meaning as `anyone`. The migration does not backfill existing issues. - Added shared create, update, response, and compact issue contracts for `anyone`, `not_creator`, and `human_only`. - Removed the unconditional agent self-approval rejection for `in_review` issues. - Added one reusable verdict-actor check for terminal status changes and pending interaction accept or reject actions. - Used the authenticated principal type for `human_only`. Agent keys and run tokens remain agent principals. - Used the latest transition into `in_review` to identify the requester for `not_creator`. - Added actionable 403 responses that name the policy, the allowed actor, and the next step. - Kept the configured execution-stage transition and signoff behavior. - Added focused contract, helper, status-route, interaction-route, and execution-stage regression tests. - Updated the implementation specification for the new issue field. ## Verification - `pnpm exec vitest run packages/shared/src/validators/issue.test.ts server/src/__tests__/issue-review-policy.test.ts server/src/__tests__/issue-stalled-review-decision-routes.test.ts --reporter=dot` passed: 42 tests. - `pnpm --filter @paperclipai/shared typecheck` passed. - `pnpm --filter @paperclipai/db typecheck` passed, including migration numbering and safety checks. - `pnpm --filter @paperclipai/server typecheck` passed. - `pnpm run typecheck:build-gaps` passed across server, CLI, plugin SDK/examples, plugin wiki, and UI. - `git diff --check origin/master...HEAD` passed. - SecurityEngineer review approved the authenticated-principal checks and accepted policy-relaxation tradeoff with no required changes. - Greptile reviewed the latest head at 5/5 with zero inline comments or follow-ups. - The latest-head GitHub rollup passed build, typecheck, server/workspace tests, serialized suites, canary, e2e, and external security checks. ## Risks - The migration adds one nullable text column. It has no default and no backfill. - `not_creator` reads the latest recorded transition into `in_review`. It denies the verdict when it cannot identify the requester. - Agents can change or relax `reviewPolicy` when they have issue write access. This is intentional for this issue-level control. - Null and `anyone` do not add a database query to the verdict path. - Configured execution-stage checks still run after the issue-level policy check. > This work aligns with the completed "Agent Reviews and Approvals" and "Enforced Outcomes" roadmap items. It does not add a new roadmap capability. ## Model Used - OpenAI Codex, GPT-5. The exact deployment ID and context-window size are not exposed to the agent. The run used reasoning, repository tools, code execution, and GitHub CLI access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5b62a3883f |
feat(settings): add experimental Simplified English Interactions flag (#10934)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents ask humans for decisions through interaction blocks: plan confirmations, structured questions, suggested tasks, and checkbox prompts > - Each agent writes these decision prompts in its own style, so operators can get long or unclear text at the exact moment they must decide > - There is no instance-level control that makes agents use a controlled language for these decision points only > - This pull request adds an experimental setting that tells agents to write all user-interaction content in ASD-STE100 Simplified Technical English, with the context the user needs and the effect of each choice > - The benefit is faster, clearer human decisions, with agent thinking and normal responses unchanged ## Linked Issues or Issue Description Refs #10410 (optional `/simplified-english` skill in the skills catalog; this PR adds the instance-level toggle for interactions). **Problem or motivation** Agent-posted user interactions (plan confirmations, structured questions, suggested-task proposals, checkbox prompts) are written in each agent's default style. Operators who want fast, unambiguous decisions have no way to ask agents to use a controlled language for exactly those decision points. **Proposed solution** Add an experimental instance setting, `enableSimplifiedEnglishInteractions` ("Simplified English Interactions"). When it is on, the server sets `simplifiedEnglishInteractions: true` in the heartbeat wake payload. The shared wake-prompt renderer, used by every adapter, then emits a directive: write all user-interaction content in ASD-STE100 Simplified Technical English, state what information the user needs to decide, and state what happens for each choice. The directive applies to interaction content only. Thinking, comments, documents, and other responses keep their usual style. **Alternatives considered** Per-agent instructions work today, but someone must maintain them on every agent. An instance-level toggle applies uniformly and turns off in one place. Server-side rewriting of interaction payloads was rejected: post-hoc translation is lossy and cannot add the decision context that only the agent has. **Roadmap alignment** Extends the experimental settings surface with another opt-in agent-behavior refinement, consistent with existing prompt-side flags. ## What Changed - Added `enableSimplifiedEnglishInteractions` to the experimental instance-settings zod schema, mirror type, and feature catalog (default off, tier preference) in `packages/shared`. - Server `instance-settings.ts` normalizes the flag on both read branches; `heartbeat.ts` reads it once and passes `simplifiedEnglishInteractions` into `buildPaperclipWakePayload`. - Shared adapter renderer (`packages/adapter-utils/src/server-utils.ts`): added the field to `PaperclipWakePayload`, normalization, and an `- interaction language (experimental): ...` directive emitted in both fresh and resume prompt lanes, so one injection point covers all adapters. - UI: new experimental settings card "Simplified English Interactions" in `ui/src/pages/InstanceExperimentalSettings.tsx`, in alphabetical card order; fixtures updated. - Tests: renderer coverage for flag on/off in both lanes, plus schema/catalog/UI fixture updates. ## Verification - From the repo root: `node_modules/.bin/vitest run packages/adapter-utils` (88/88), `packages/shared` validators (25/25), server instance-settings + heartbeat suites (47/47 and 70/70 consumer tests), `ui` settings tests (32/32). - Typecheck is clean in all four touched packages. - Manual check: turn the flag on in Settings → Experimental, wake an agent, and confirm the wake prompt contains the interaction-language directive; turn it off and confirm the directive is absent. ## Risks - Low risk: the flag defaults to off, and the only behavior change is one extra directive line in the wake prompt when an operator turns it on. - The directive is advisory to the agent; models can still deviate from STE. No data or API shape changes; no migration. ## Model Used - Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended thinking, agentic tool use via Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e43f187cad |
feat(secrets): add human-approved secret proposals (#9934)
## Thinking Path > - Paperclip is the control plane people use to manage AI-agent companies. > - Agents can encounter credentials during work. > - Directly creating live secrets or bindings would bypass human governance. > - Proposal records must remain inert and separate from live secret resolution until an authorized human approves them. > - Approval must reuse the existing secret-create and protected agent-config write paths. > - This pull request adds the propose, review, approve, and reject lifecycle. > - The benefit is that agents can safely hand credentials into Paperclip without exposing plaintext or gaining authority to activate them. ## Linked Issues or Issue Description Follow-on to #9921, which established run-bound agent secret access. **Problem / motivation:** Agents can receive credentials during work. There is no governed way for them to propose a credential or binding without exposing plaintext in work artifacts or immediately creating live access. **Proposed solution:** Store agent-authored proposals outside live secret tables. Encrypt each proposed value and register exact-value redaction when Paperclip receives it. Require an authorized human to approve or reject each proposal. Approval executes through the normal write paths as the human approver. Binding proposals can target only the proposer or its downward reporting chain under the restrictive V1 policy. **Alternatives considered:** We rejected live secrets with a `proposed` status. That design would put untrusted rows in resolver, list, and sync paths. It would also allow uniqueness squatting. We rejected direct agent binding writes because a binding is an agent-config write and must keep the existing human permission gate. **Roadmap alignment:** This change extends the run-bound agent secret-access foundation in #9921 with a governed proposal workflow. ## Security Verdict Q0 SecEng verdict: **PASS-with-required-changes**. The review accepted the separate proposal-table design and required the implementation to: - fail closed unless both encryption and exact-value run redaction registration succeed; - scrub ciphertext idempotently on reject, withdraw, and expiry, with audit-visible state; - treat agent justification as hostile input and foreground action, target, provenance, and approver permissions; - snapshot and re-check the target agent plus reports-to chain at approval to prevent org-chart laundering; - make cascade approval atomic and fail closed if either secret creation or binding authorization fails; - deny low-trust, `skill_test`, `task_bridge`, and non-run-bound sources consistently; and - execute approval through the normal human secret/config write paths, including protected-change gates. Those requirements are implemented and covered by focused service, route, and UI tests. Residual V1 risk remains the accepted 14-day encrypted retention window. Proposal-time redaction also cannot clean a value that leaked before the propose call. ## What Changed - Added `company_secret_proposals`, migration `0207`, shared proposal contracts, and a state-machine service for create, approve, reject, withdraw, cascade, expiry, and ciphertext scrubbing. - Added run-bound agent proposal routes and board review routes. The routes derive provenance from authentication and enforce source restrictions, company isolation, chain-of-command checks, approval-as-approver, wake-on-resolution, and dual audit trails. - Added durable per-run exact-value redaction registration so proposal values remain redacted on later read surfaces. - Added the Secrets **Proposals** tab and agent configuration **Proposed access** rows. The UI shows fingerprint and length only. It also frames agent justification as untrusted input, runs permission preflight, supports approve and reject actions, and confirms cascades. - Updated OpenAPI, agent skill guidance, API reference documentation, and focused server and UI regression coverage. - Rebased the branch onto current `master` and renumbered the proposal migration after `0206`. ## QA Acceptance Results Q5 QA verdict: **PASS — 9/9 acceptance criteria met**, with one Minor non-blocking follow-up. - **AC1:** proposed values never echo, never appear in live lists/resolvers, and expose only fingerprint + length to board reviewers. - **AC2:** restrictive `self_and_reports` matrix passes: self/downward allowed; upward/lateral denied. - **AC3:** secret approval uses the normal create path, honors rename overrides, records proposer/approver provenance, and scrubs ciphertext. - **AC4:** approved bindings materialize and resolve through the target agent's runtime list/fetch routes. - **AC5:** pending-secret bindings require cascade; cascade succeeds atomically and permission failures leave nothing applied. - **AC6:** reject, withdraw, dependent rejection, and expiry paths scrub ciphertext and preserve reasons/audit state. - **AC7:** token/source and approver denial matrix passes through live checks plus focused route tests. - **AC8:** proposal lifecycle events and reused `secret.created`/config-write events form the required dual audit trail; origin-issue notification and wake are queued. - **AC9:** both review surfaces render and execute correctly; UI approval materializes the binding. QA also confirmed zero plaintext occurrences for all exercised proposal values in server logs. The single finding is that the company-level `bindingTargetPolicy` toggle is not wired yet. V1 is hardcoded to the restrictive `self_and_reports` policy. The matrix is correct and the follow-up is tracked separately, so QA classified it as non-blocking. ## Verification - Focused server proposal and redaction suite: 83 tests pass. - Focused proposal review UI suite: 54 tests pass. - Embedded-Postgres migration reapply test: 1 test passes with the documented 30-second timeout. - `pnpm --filter @paperclipai/db typecheck` passes, including migration numbering and safety checks. - `pnpm --filter @paperclipai/shared typecheck` passes. - `pnpm --filter @paperclipai/ui typecheck` passes. - `pnpm check:token-gates` passes with all gates clean. - Q5 exercised the complete propose, review, approve, bind, and runtime-resolve flow over real HTTP, JWT, and database paths. It verified 9/9 acceptance criteria. ## Risks - Proposal ciphertext is retained encrypted for up to 14 days while pending. Terminal-state and expiry scrub paths reduce but do not remove server-compromise risk during that window. - The V1 target policy is restrictive but not yet company-configurable. A separate follow-up owns that change. - A new migration can require another renumber if another migration lands before maintainers merge this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent. The exact runtime model ID and context-window size are not exposed. The agent used reasoning, repository editing, terminal execution, Paperclip API, and GitHub CLI capabilities. Q3 UI work also records Claude Opus 4.8 assistance in its commit trailers. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f950952de7 |
fix: reliably show plans in the Plan pane and restore sticky plan confirmation CTAs (#10930)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Chat-style tasks show an agent's plan in a dedicated "Plan" pane, and a plan confirmation lets the user accept or request changes to that plan > - When an agent asked for confirmation but never actually published the plan document (it only wrote the plan in a comment or a question), the Plan pane rendered empty, and the confirmation call-to-action that used to sit pinned at the bottom of the pane had disappeared > - A user asked to confirm a plan they cannot see, with no visible CTA, is stuck — the feature silently fails > - This pull request closes the gap on both sides: it prevents plan confirmations that don't point at a real, latest plan revision, it teaches agents to publish the plan document before confirming, and it restores the sticky confirmation action bar and an explanatory empty state so the pane never goes silently blank > - The benefit is that when a plan is expected, it reliably shows up in the right pane with reachable accept/revise actions ## Linked Issues or Issue Description <!-- No public GitHub issue exists; describing in-PR per the bug template. --> **Bug report** - **What happened:** A task in planning mode could present a plan confirmation while the Plan pane stayed empty (no plan document rendered), and the plan-card confirmation CTAs that were previously pinned to the bottom of the Plan pane no longer appeared. - **Expected behavior:** When a plan is expected, the plan document appears in the Plan pane; when a plan is genuinely missing, the pane explains why rather than showing nothing; and the accept/request-changes CTAs stay visible and reachable while the plan scrolls. - **Steps to reproduce:** Put a task in planning mode with the chat-style task view enabled, have an agent create a plan confirmation without first publishing the `plan` document, and open the Plan tab — the pane is blank and the confirmation actions are missing. - **Deployment mode:** Local dev and self-hosted; UI + server. Related PR (not a duplicate): #9609 "Pin pending confirmations by composer" pins confirmations in a different surface (the composer); this PR restores the Plans-pane action bar and the server/agent guarantees behind it. ## What Changed - **Server:** Reject a `request_confirmation` whose target is a plan document unless a plan document exists and the target points at its *latest* revision, so a confirmation can never reference a plan the pane cannot render (`readPlanTarget` is now exported for reuse). - **Agent instructions:** The CEO and default agent instruction bundles now spell out a plan-publish contract — publish the `plan` document, re-`GET` it and capture `latestRevisionId`, then create the confirmation targeting that revision; never present a plan only in a thread comment or via `ask_user_questions`. - **UI — sticky CTAs:** Restore the plan confirmation action bar pinned to the bottom of the Plans tab so accept/revise stay reachable while the plan scrolls. - **UI — diagnostics:** Keep the Plan tab visible whenever an issue is in planning mode (even before a plan document exists) and show an empty state explaining why the pane is empty instead of rendering nothing. - **UI — annotations:** Add a `panelPlacement="inline"` mode so the plan-document annotation panel renders in document flow instead of as a floating side panel when hosted in the narrow task properties pane. ## Verification - `pnpm check:token-gates` → 3/3 CLEAN - `pnpm typecheck` → clean (all packages) - UI: `pnpm --filter @paperclipai/ui exec vitest run src/components/issue-properties/IssuePlanConfirmationActionBar.test.tsx src/components/IssueProperties.test.tsx src/components/IssueDocumentAnnotations.test.tsx` → 69 passed - Server: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/issue-thread-interaction-routes.test.ts src/__tests__/agent-skills-routes.test.ts` → 60 passed - Manual: with a planning-mode task, the Plan tab stays visible, shows the plan document (or a diagnostic empty state), and the confirmation CTAs stay pinned at the bottom. Visual note: snapshot baselines are intentionally not updated — per `doc/design/DECISION-SHEET.md` "Per-change snapshot verification demoted to dormant (Jul 13 2026)". The `storybook-visual` label is intentionally not added. ## Risks Low-to-moderate. The server change adds a validation gate on plan-document confirmations: an interaction that targets a stale or nonexistent plan revision is now rejected with a 422 instead of being created. This is the intended guarantee, but any caller that relied on creating such confirmations will now need to publish the plan document first (which the updated agent instructions cover). UI changes are additive to the Plans tab and gated by the existing chat-style-task experimental flag. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended thinking enabled, with tool use (file editing, shell, test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dc71fef6bf |
feat(ui): show task identifier in task-detail breadcrumb header (#10933)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task detail page shows a breadcrumb header with the task status glyph and the task title. > - The breadcrumb did not show the task identifier, so a reader could not name the task without opening extra context. > - Agents and people refer to tasks by identifier, so the identifier belongs next to the title. > - This pull request renders the task identifier in the breadcrumb header, between the status glyph and the title. > - The benefit is faster reference: a reader sees the task key and the title together at the top of the page. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task detail breadcrumb header. It shows the status glyph and the task title, but not the task identifier. **Subsystem affected** Web UI — the breadcrumb bar on the task detail page (`ui/src/components/BreadcrumbBar.tsx`, `ui/src/context/BreadcrumbContext.tsx`, `ui/src/pages/IssueDetail.tsx`). **Current behavior** The breadcrumb header renders the status glyph and then the task title. The task identifier does not appear in the header. **Proposed behavior** The breadcrumb header renders the task identifier between the status glyph and the title. The identifier uses gray monospace styling from design tokens (`font-mono text-muted-foreground`). **Reason and benefit** A reader can name and reference the task from the header without opening more context. The identifier and the title appear together. **Breaking changes** None. The identifier field is optional. Crumbs without an identifier render as before. ## What Changed - Add an optional `identifier` field to the `Breadcrumb` type and include it in the `breadcrumbsEqual` comparison so an identifier change triggers a fresh render. - Add a `CrumbIdentifier` helper in `BreadcrumbBar` that renders the identifier in gray monospace (`font-mono text-muted-foreground`), placed after the leading status glyph in each crumb variant. - Wire the issue identifier onto the task crumb in `IssueDetail`. - Add unit tests that cover the identifier field in `breadcrumbsEqual` (fresh render on change, no-op on identical value). ## Verification - `pnpm check:token-gates` → 3/3 gates CLEAN (color literals, arbitrary bracket values, raw font-size). - `pnpm --filter @paperclipai/ui exec vitest run src/context/BreadcrumbContext.test.tsx` → 4/4 tests pass. - `pnpm typecheck` → the four changed files typecheck clean. - Manual: open a task detail page. The breadcrumb header shows the status glyph, then the task identifier in gray monospace, then the title. Visual change. Snapshot baselines are intentionally not updated, per `doc/design/DECISION-SHEET.md` → "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks Low risk. The change is additive and the identifier field is optional. It touches only the breadcrumb header rendering and the equality check. No data model or API change. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended thinking enabled, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ac3b2e1d7a |
fix(ui): use cloud logout for managed sign-out (#10937)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip supports self-hosted and Cloud-managed authenticated deployments > - A Cloud-managed tenant uses the Cloud harness to own the full browser session > - The account menu treated every authenticated deployment as Cloud and called the local sign-out API first > - This pull request uses the existing Cloud health metadata as the mode gate > - Cloud-managed sign-out now starts the harness logout round trip with a top-level navigation > - Self-hosted authenticated sign-out keeps the existing local API flow > - The benefit is a complete Cloud logout without changing self-hosted behavior ## Linked Issues or Issue Description Related prior work: Refs #10802. **What happened?** The account menu called the app-local sign-out endpoint before it moved an authenticated browser to the Cloud logout route. It also used authenticated deployment mode as the Cloud test. This test included self-hosted authenticated instances. **Expected behavior** A Cloud-managed tenant must navigate the top-level browser directly to `/cloud/logout`. A self-hosted authenticated instance must keep the app-local sign-out flow. **Steps to reproduce** 1. Open a Cloud-managed tenant. 2. Open the account menu. 3. Select **Sign out**. 4. Observe that the browser returns through the tenant auth route instead of completing the Cloud logout round trip. **Paperclip version or commit** Reproduced on `master` after `76f442040c`. **Deployment mode** Paperclip Cloud-managed authenticated deployment. ## What Changed - Read the existing Cloud instance metadata in the account menu. - Navigate directly to `/cloud/logout` for Cloud-managed instances without calling the local sign-out API. - Keep the local sign-out API and cache refresh for self-hosted authenticated instances. - Add regression coverage for both sides of the mode gate. ## Verification - `pnpm exec vitest run ui/src/components/SidebarAccountMenu.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` ## Risks - Low risk. The change is limited to the account-menu action. - The Cloud branch depends on the existing `health.cloud` metadata that already gates other Cloud UI behavior. - The self-hosted regression test verifies that authenticated mode alone does not select the Cloud route. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The runtime provided agentic reasoning, repository tools, shell execution, and test execution. The exact internal model ID and context window are not exposed to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f5e9ca3e89 |
fix(adapters): keep user-scoped env bindings on the agent Test action (#10926)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The agent Test action builds adapter config from the form.
> - The build-config parser kept plain and secret_ref bindings.
> - It dropped user_secret_ref bindings on the create path.
> - This PR shares one parser that keeps every binding shape.
> - The test path now sends the same env binding set that a real run
sees.
> - The benefit is one fix across every adapter build-config path.
## Linked Issues or Issue Description
**What happened?**
The agent Test action dropped a user-scoped env binding in create mode.
The same agent config worked in a real run. Related public PRs: #10115,
#9321, #9921, #8825.
**Expected behavior**
The Test action should keep user-scoped env bindings and resolve them
like a real run.
**Steps to reproduce**
1. Set a user-scoped env binding on an agent config form.
2. Run Test in create mode.
3. The probe runs without the variable.
**Paperclip version or commit**
|
||
|
|
b1b7a9dff6 |
feat(settings): alphabetize experimental cards and drop the Experimental chip (#10924)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Instance Settings area exposes an Experimental page that lists opt-in feature toggles as cards. > - New experimental features were appended to the list over time, so the cards sat in insertion order with no predictable arrangement. > - An unordered list is hard to scan when you are looking for one specific feature. > - Each card also carried a small "Experimental" secondary badge, which is redundant on a page that is itself titled Experimental. > - This pull request sorts every card alphabetically by its title and removes that redundant badge. > - The benefit is a list that is faster to scan and headings that are less cluttered. ## Linked Issues or Issue Description No public GitHub issue exists for this change, so the enhancement is described inline following `.github/ISSUE_TEMPLATE/enhancement.yml`: **What existing behavior does this improve?** The Instance Settings → Experimental page, which lists opt-in feature toggles as a stack of cards. **Subsystem affected** UI — the Instance Experimental settings page (`ui/src/pages/InstanceExperimentalSettings.tsx`). **Current behavior** Cards render in insertion order (the order features happened to be added), so finding a specific feature means scanning the whole list. Several headings also carry a redundant "Experimental" secondary badge. **Proposed behavior** Cards render top-to-bottom in A→Z order by title, and no card shows an "Experimental" secondary badge. Toggle logic, footnotes, conditional visibility, and the "Managed by Paperclip Cloud" badge are unchanged. **Reason and benefit** Alphabetical order makes the list predictable and quick to scan for a specific feature. The "Experimental" badge repeats information already conveyed by the page title, so removing it declutters the headings. **Breaking changes** None. This touches card render order and the removal of a decorative badge only — no state, persistence, toggle, or visibility logic changes. ## What Changed - Sorted every card on the Instance Experimental settings page alphabetically by its heading title. - Removed the redundant "Experimental" secondary badge from the card headings (previously on Apps, Cases, and Chat-Style Tasks). - Added tests asserting the cards render in case-insensitive alphabetical order and that no card renders an "Experimental" secondary badge. - No behavior change: toggle handlers, footnotes, managed-key handling, and conditional cards (Conference Room Chat, worktree-scoped run) are untouched and now sort into their alphabetical slots. ## Verification - `pnpm check:token-gates` → all 3 gates CLEAN. - `npx vitest run ui/src/pages/InstanceExperimentalSettings.test.tsx` → 32/32 tests pass (the suite renders the real component and now covers ordering + badge removal). - `pnpm --filter @paperclipai/ui typecheck` (`tsc -b`) → clean. - Manual: open Instance Settings → Experimental. The cards read A→Z and no card shows an "Experimental" chip. ## Risks Low risk. The change is limited to one page component: card render order and the removal of a decorative badge, plus new tests. No state, persistence, toggle, or visibility logic is modified. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended thinking enabled, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
72b509c895 |
Recognize delivered workspaces and reap terminal worktrees (#10908)
## Thinking Path > - Paperclip separates workspace provisioning lifecycle from whether the work was actually delivered. > - Git ancestry alone cannot recognize squash merges or deliveries into a branch other than the workspace base. > - A merged pull request linked from a terminal issue is stronger delivery evidence for those cases. > - The read contract should expose that evidence without changing persisted workspace schema. > - Cleanup must remain conservative: terminal descendants, delivered work, and no active run checkout are all required. > - Reusing the existing cleanup primitives keeps service shutdown, lease cleanup, activity logging, and archival behavior consistent. > - Focused regression coverage locks in both the honest read signal and the fail-closed reaper guards. ## Linked Issues or Issue Description **What existing behavior does this improve?** Execution workspace close-readiness payloads and terminal workspace cleanup. **Current behavior** Delivered squash-merged or cross-branch workspaces can remain `active` and report a permanent “not merged” warning because git ancestry does not contain their original commits. **Proposed behavior** Read payloads distinguish PR-confirmed delivery, ancestry delivery, unmerged work, and unknown state. Fully terminal delivered workspace trees are archived only when no active run holds the checkout. **Reason and benefit** Operators and automation receive an honest delivery signal, while shipped worktrees stop looking active forever and genuinely unmerged work retains its warning. **Breaking changes** The workspace payload gains a derived field. Existing fields and persistence remain unchanged; no database migration is required. **What happened?** A delivered workspace can remain `active` and warn that it is not merged forever after its issue ships through a squash or cross-branch pull request. **Expected behavior** Pull-request delivery should be represented honestly, and a fully terminal delivered workspace should become cleanup-eligible when no run holds its checkout. **Steps to reproduce** 1. Create an issue workspace with commits ahead of its configured base. 2. Deliver those commits with a squash merge or into a different target branch. 3. Mark the source issue and descendants done, then read workspace close readiness. Before this change, the workspace remains active with a “not merged” warning indefinitely. ## What Changed - Added the derived `deliveryState` workspace contract: `merged_via_pr`, `merged_by_ancestry`, `unmerged`, or `unknown`. - Extracted a shared GitHub pull-request merge classifier and reused it for merge confirmations and workspace delivery checks. - Suppressed false ancestry warnings when a terminal issue has ground-truth merged-PR evidence. - Added an idempotent terminality reaper with descendant-terminal, active-run, and delivered-work guards. - Restricted PR delivery evidence to the source issue, then required live merged state plus matching GitHub repository, head branch, and current workspace HEAD; persisted status, stale PRs, lexical mentions, inbound references, and descendant PRs cannot authorize cleanup. - Preserved workspaces with modified or untracked files even when their committed HEAD was delivered. - Bounded both long-lived pull-request state caches to 1,000 entries with oldest-entry eviction. - Routed eligible workspaces through existing runtime shutdown, lease cleanup, activity logging, and archival machinery with exclusive Git index, HEAD, and branch-ref locks plus non-forced removal. - Added regression coverage for delivery derivation, warning behavior, reaper guards, scheduler wiring, and squash/cross-branch delivery. ## Verification - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts src/__tests__/merged-pr-confirmation-sweep.test.ts src/__tests__/server-startup-feedback-export.test.ts --reporter=verbose` — 63 passed - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts --reporter=verbose` after review hardening — 43 passed - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/execution-workspaces-service.test.ts src/__tests__/merged-pr-confirmation-sweep.test.ts src/__tests__/external-objects-service.test.ts --reporter=dot` on the final local head — 73 passed - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-workspace-busy.test.ts --reporter=verbose` — 15 passed - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run` — server 3,662 passed (4 skipped), UI 3,599 passed, CLI 327 passed, shared 415 passed, and skills catalog 20 passed; the aggregate DB stage ran both source and built copies of one unrelated embedded-Postgres migration test and both reached its 5-second timeout - `pnpm --filter @paperclipai/db exec vitest run src/status-card-migrations.test.ts --reporter=verbose` — isolated aggregate-timeout verification passed in 3.99 seconds - `NODE_ENV=production pnpm build` - `pnpm check:token-gates` ## Risks The reaper intentionally fails closed when issue terminality, pull-request state, git ancestry, or checkout ownership cannot be proven. GitHub lookups can delay classification and cleanup but cannot cause an unproven workspace to be archived. Automated terminal archival holds exclusive Git index, HEAD, and branch-ref locks across validation and removal, skips configured destructive hooks, and uses non-forced removal so dirty writes fail closed. Reopening a source issue does not restore an archived workspace; it emits an audit event so a human or agent can re-provision explicitly. ## Model Used OpenAI Codex, GPT-5. The runtime did not expose a more specific model ID or context-window size. Reasoning, tool use, repository editing, test execution, and GitHub CLI access were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c2b41bb7cd |
fix(issues): quiet missing-disposition warnings while a live continuation is running (#10899)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - When an agent run ends without recording a disposition, Paperclip
raises a "missing disposition" handoff so the work does not silently
stall
> - The server already tracks whether such an issue has a live
continuation (a running or queued run, or a queued wake) in
`successfulRunHandoff.hasLiveContinuation`
> - But no UI surface read that flag, so an issue that an agent was
actively working on still showed the "This task still needs a next step"
banner, a loud thread warning, and "Needs next step" badges
> - This pull request makes every missing-disposition complaint respect
liveness: warn only when no live agent is on the issue and it is really
stuck
> - The benefit is that users see the warning only when action is
needed, and the noise disappears while an agent is already handling the
issue
## Linked Issues or Issue Description
No public GitHub issue exists for this bug. Description follows the
bug-report template:
**What happened?**
An issue that a live agent run was actively working on showed the
"missing disposition" warning banner, a loud thread notice, and "Needs
next step" badges at the same time. The API payload for that issue
showed `successfulRunHandoff.required: true` together with
`hasLiveContinuation: true` and a `liveRunId`, but the UI ignored the
liveness fields.
**Expected behavior**
The missing-disposition warning appears only when the issue has no live
run or queued wake. A live agent records a disposition when its run
ends. Paperclip complains only if the run ends and no disposition
exists.
**Steps to reproduce**
1. Let a run finish on an in-progress issue without a disposition.
Paperclip raises the handoff and queues a corrective wake.
2. Open the issue page while the corrective run (or any new run) is
live.
3. See the banner, the badges, and the loud thread notice — all visible
while the agent works.
**Paperclip version or commit**
Current `master` (reproduced at commit
|
||
|
|
427509e6e0 |
fix(ui): show the synced company logo on the Cloud org switcher trigger (#10917)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - On Paperclip Cloud a tenant instance holds exactly one company, and the cloud control plane pushes the stack's uploaded workspace icon into that company's branding. > - The cloud-mode organization switcher trigger always rendered the deterministic monogram, so an uploaded organization logo never appeared in the app chrome. > - This pull request renders the trigger through the tenant company's logo and brand color, with the monogram as the fallback. > - The benefit is that the logo a customer uploads for their organization actually shows up inside their Paperclip app. ## Linked Issues or Issue Description No existing public issue found (searched open/closed PRs and issues for "organization logo", "switcher logo", "company logo cloud" — closest related PR is #10850, which introduced the cloud-mode switcher). Describing in-PR: **Subsystem affected** Board UI: the sidebar organization switcher in Paperclip Cloud mode (`SidebarCompanyMenu`). **Problem or motivation** A Cloud customer uploads an organization logo when creating their workspace; the control plane syncs it into the tenant company's branding (`company.logoUrl`). But the cloud branch of the switcher trigger rendered `StackIcon` — monogram-only by design for stack rows — for the trigger too, ignoring `selectedCompany.logoUrl`. Result: the uploaded logo never appears in the app chrome; users see a letter tile instead. **Proposed solution** Add a `CurrentStackIcon` for the trigger that passes the selected company's `logoUrl`/`brandColor` into `CompanyPatternIcon`, seeded by the stack display name. Falls back to the exact previous monogram when no logo is set. Stack rows are unchanged: the portfolio payload deliberately carries no hot-linkable icon URL for other stacks. **Alternatives considered** Fetching per-stack icons for the rows was rejected: the cloud portfolio payload carries no icon URLs (embedding signed, expiring control-plane URLs would be wrong), and the defect is the current organization's chrome, which the already-synced company logo covers. ## What Changed - `ui/src/components/SidebarCompanyMenu.tsx`: cloud-mode trigger renders the tenant company logo (fallback: monogram); stack rows untouched; self-hosted path untouched. - `ui/src/components/SidebarCompanyMenu.test.tsx`: new regression test that the trigger carries the company logo while stack rows keep monograms; the `CompanyPatternIcon` mock now exposes `logoUrl`. ## Verification - `pnpm --dir ui exec vitest run src/components/SidebarCompanyMenu.test.tsx`: 12/12 pass (11 existing + 1 new). - `pnpm --dir ui exec tsc --noEmit`: clean. ## Risks - Cloud-only rendering branch; self-hosted trigger rendering is untouched. - If the branding sync has not run yet, the trigger shows the same monogram as before — no regression, and it upgrades in place once `logoUrl` arrives. ## Model Used - Claude Fable 5 (`claude-fable-5`), Anthropic. The run used extended reasoning, repository tools, shell execution, and GitHub integration. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1fa36be353 |
fix(ui): use HTTP-safe clipboard copy everywhere (#10875)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators often open self-hosted Paperclip over plain HTTP on a LAN or private network. > - Browser Clipboard API writes are not reliable in that insecure context. > - Paperclip already has one shared helper with a legacy copy fallback, but many current copy actions bypass it. > - This pull request routes every core UI copy action and the first-party workspace-diff plugin through the shared helper. > - The benefit is consistent copy behavior on HTTPS, localhost, and plain-HTTP private deployments. ## Linked Issues or Issue Description Refs #3529. This change supersedes the stale prior attempt in #3531. Current master has more copy surfaces and a first-party plugin UI bridge that the prior branch does not cover. ## What Changed - Replaced direct Clipboard API writes and duplicate fallback implementations across the current core UI with `copyTextToClipboard`. - Added an HTTP-safe clipboard function to the plugin UI SDK and wired the host bridge to the same implementation. - Migrated the first-party workspace-diff plugin to the plugin SDK clipboard function. - Added unit coverage for native rejection fallback and plugin host delegation. - Added a source-level regression test that rejects new direct clipboard writes outside the shared implementation. - Documented the plugin UI clipboard function. ## Verification - `NODE_ENV=test pnpm exec vitest run ...` for 14 affected suites: 164 tests passed. - `pnpm exec vitest run tests/ui-clipboard.test.ts` in `packages/plugins/sdk`: 1 test passed. - `NODE_ENV=test pnpm -r typecheck`: passed for 31 workspace projects. - `NODE_ENV=test pnpm test:run`: passed. - `NODE_ENV=production pnpm build`: passed. - `pnpm check:token-gates`: passed with all gates clean. ## Risks Low risk. Secure contexts still use the modern Clipboard API. Plain HTTP and rejected modern writes use the existing `execCommand("copy")` fallback. That API is deprecated, but it is the compatibility path required for insecure contexts. The change has no schema, API, or visual design effect. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, `gpt-5.6-sol`. The runtime did not expose a context-window size. Reasoning, tool use, repository editing, test execution, and GitHub CLI access were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
14d755824c |
Remove decision and review summaries from issue headers (#10891)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Operators use issue pages to read task state and control task work > - The issue header showed separate summaries for open decisions and review paths > - These summaries repeated state that belongs in the Decisions view > - The extra sections added noise before the issue description and thread > - This pull request removes both header summaries and keeps decision actions in the Decisions view > - The benefit is a simpler issue header with one place for decision work ## Linked Issues or Issue Description **What existing behavior does this improve?** The issue detail header shows separate pending-decision and review-path sections. **Subsystem affected** `ui/` — React and Vite board UI. **Current behavior** An issue header can show a decision strip and a larger review panel before the issue content. **Proposed behavior** The issue header does not show either decision section. Operators continue to manage decisions and stalled reviews in the Decisions view. **Reason and benefit** This removes duplicate decision state from the issue header and reduces visual noise. **Breaking changes** The issue page no longer provides these summaries or shortcuts. Decision data, review state, and the Decisions view do not change. ## What Changed - Removed the pending-decision strip and review-path panel from the issue detail header. - Deleted the two unused header components and the panel-specific test. - Kept stalled-review actions and their Storybook examples in the Decisions queue. - Added an issue-detail regression test that covers both removed sections. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/pages/IssueDetail.test.tsx` (46 tests passed) - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm build-storybook` - `git diff --check` ## Risks - Low risk. This change removes two issue-header surfaces. It does not change decision APIs or data. - Users must open the Decisions view to find pending decisions and stalled-review actions. > This change does not duplicate planned core work in `ROADMAP.md`. GitHub searches found no related open issue or pull request. ## Model Used - OpenAI Codex, GPT-5. The exact deployment ID and context window are not exposed. Tool use and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ef33c1d9ed |
fix(decisions): retire completed-target decisions and link targets from the card (#10892)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The decisions desk shows pending decisions that need an operator response > - A strict decision cannot apply its effects after its target task changes > - A decision still remained pending when every target task finished after proposal > - The card also linked only the origin task, even when the decision acted on another task > - This pull request expires those moot decisions and links their target tasks > - The benefit is an accurate queue and a clear path to the work that each decision affects ## Linked Issues or Issue Description Related PR: #10801 removes the issue-page decision strip, which makes clear queue provenance more important. **What happened?** A strict decision stayed pending until its time-to-live limit after every target task reached `done`. The decision card linked only the origin task. The origin task is where the agent proposed the decision, and it can differ from the task that the decision affects. An operator could therefore open a finished task with no visible decision and no explanation of the real target. **Expected behavior** Paperclip must expire a strict decision when all of its targets finish after the decision is proposed. The card must show and link every target task that differs from the origin task. **Steps to reproduce** 1. Create a strict decision that targets an active task from a different origin task. 2. Move the target task to `done` without resolving the decision. 3. Run the decision expiry sweep. 4. Observe that the old code keeps the decision open until its time-to-live limit. 5. Observe that the old card links only the origin task. **Paperclip version or commit** The bug reproduces on upstream `master` before this pull request. **Deployment mode** Local dev and self-hosted server modes are affected because the behavior is in the shared decision service and board UI. ## What Changed - Expire an open strict decision with reason `target_completed` when every strict target reached `done` after proposal. - Keep decisions that intentionally target an already-finished task. - Keep lenient-only decisions open. - Keep continuation delivery consistent with other expiry reasons. - Add target-task links to the decision card provenance line. - Use one shared target-ID helper across signing, execution, expiry, card provenance, and resolver preloading. - Add service and UI regression tests for primary, secondary, and target-completed cases. ## Verification - `pnpm exec vitest run ui/src/components/DecisionCard.test.tsx server/src/__tests__/decisions-service.test.ts` — 51 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — all gates clean. - `git diff --check origin/master...HEAD` — passed. ## Risks - Low migration risk. This change does not alter the database schema. - The expiry sweep performs the existing strict-target query and adds a snapshot comparison before expiry. - A decision remains open if any strict target is active or if a target was already `done` at proposal time. > The roadmap lists work queues as planned. This pull request fixes the existing decisions desk. It does not add a new queue subsystem. ## Model Used - Implementation: Anthropic Claude through Claude Code. The runtime did not expose the exact model snapshot or context-window size. The model used reasoning, repository tools, code execution, and test execution. - PR preparation: OpenAI Codex with GPT-5. The runtime did not expose a dated model snapshot or context-window size. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8142e54150 |
feat(activity): merge the audit page into one rich Activity page (#10838)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The board has two different pages for change history: a basic Activity list and a rich Audit feed > - The two pages show the same kind of information, so an operator must guess which page to open > - The basic list also caps at 200 rows and has no filters, so it hides older changes > - This pull request merges both pages into one Activity page that is built on the rich audit feed > - The page adds a scope toggle for all actors or agent actions only, and it hides privileged controls from members who do not have the audit permission > - The benefit is one obvious place to answer "who changed what", for every member, with filters and full history ## Linked Issues or Issue Description Related pull requests in this stack (open both before this one): - Refs #10830 — adds the company prefix to the board audit route. - Refs #10831 — adds the two-tier all-actors scope to the audit endpoint. This pull request calls that scope. This branch is stacked on those two pull requests. The diff therefore shows their commits until they merge. After they merge, this pull request contains only the last two commits: the page merge and the actor-label fix. **Problem or motivation** The board has two overlapping history pages. `/:company/activity` renders a plain list that is capped at 200 rows and has no filters. The audit page renders a filtered, paginated feed of agent actions, but it is a separate sidebar item and it was reachable only by members with the audit permission. A member who wants to know who changed an issue must know which of the two pages answers the question. **Proposed solution** Keep one sidebar item, "Activity", and build it on the rich feed. Add a scope toggle: "All activity" reads every actor kind, and "Agent actions" keeps the earlier audit behavior. Put the scope in the `mode` query parameter so a person can link to it. Show the responsible-user filter and the CSV export only to callers that the server answers at the privileged tier. Redirect the earlier audit paths to the merged page with the agent scope preset, so old links continue to work. Delete the plain list page. **Alternatives considered** Keeping both pages and adding filters to the plain list. That duplicates the feed logic and keeps the "which page?" problem. Deleting the audit page instead was also rejected, because the audit feed has the pagination, filters, and export that the plain list does not. **Roadmap alignment** The roadmap marks the activity log and action attribution as shipped. This change improves that shipped capability. It does not add a new subsystem. ## What Changed - Added a scope toggle to `AuditFeed`. "All activity" requests `actorScope=all`, and "Agent actions" keeps the earlier agent-only request. Cursor pagination works in both scopes. - Stored the scope in the `mode` query parameter, so a person can bookmark or share a scope. - Made the page chrome permission-aware. The toggle, the responsible-user filter, and the CSV export appear only when the server answers at the privileged tier. A basic member sees the shared feed and no upsell wall. - Replaced the sidebar "Audit" item. The sidebar now has one "Activity" item. - Redirected `/:company/audit` and the unprefixed `/audit` to `/:company/activity?mode=agents`. - Deleted the earlier `ui/src/pages/Activity.tsx` list page and the `CompanyAudit` page wrapper. Added `CompanyActivity` as the single route target. - Fixed the actor label for stripped rows. The basic tier removes the agent id but keeps the actor kind, so every agent row rendered as "System". Rows now fall back to the actor kind: "Agent", "User", "Plugin", or "System". - Widened the responsible-user filter control, which truncated its own label. - Resolved agent names on the basic tier. The basic tier removes the privileged `agentId` but keeps the acting principal `actorId`, and the company agent directory this page already reads is authorization-filtered. The feed therefore resolves an agent actor from `agentId` first and from an agent-typed `actorId` second. Hiding the name only in the UI gave no confidentiality benefit, because any reader could join the retained id against the readable directory. Agents that the directory filters out still fall back to the generic kind label. No server payload or permission was widened. - Fixed a stuck state in the access-downgrade recovery. A downgrade between cursor requests leaves full-tier and basic-tier pages in one cache, which starts a single recovery refetch. If that refetch did not clear the mix, the cached pages kept the condition true, the "Refreshing audit access…" banner rendered permanently, and it hid the error state together with its "Try again" button. The banner is now tied to an outstanding attempt. The refetch effect also depended on the whole query object, which changes identity every render, so it repeated the request on each render; the attempt is now tracked in state and runs once per downgrade. - Kept the agent detail "Audit" tab unchanged. That tab passes a locked agent id, which keeps the earlier privileged scope and hides the toggle. The `GET /companies/:id/activity` endpoint stays. The dashboard still reads it. This pull request does not change that endpoint. ## Verification - `pnpm exec vitest run ui/src/pages/audit/AuditFeed.test.tsx ui/src/App.activity-routing.test.tsx ui/src/lib/company-routes.test.ts ui/src/components/Sidebar.test.tsx server/src/__tests__/activity-routes.test.ts server/src/__tests__/agent-action-audit-routes.test.ts` — all tests pass. - New `ui/src/App.activity-routing.test.tsx` drives the real route table. It asserts that the company activity path resolves, and that both the company audit path and the unprefixed audit path reach the activity path with the agent scope preset. - New `AuditFeed` tests cover the scope toggle, the basic tier without privileged chrome, the locked-agent case, the actor-kind fallback label, basic-tier name resolution, and both downgrade-recovery paths (the refetch errors, and the refetch returns a still-mixed pair). - Mutation-checked the three new guards: disabling each one fails the test that covers it, so none of them pass vacuously. - `pnpm -r typecheck` is clean. Both design token gates are clean. - Rendered every state in a browser at 1440x900 and at 390x844: both scopes, the basic member view, the loading state, the error state, the filtered-empty state, and the true-empty state. A designer reviewed the renders and approved them. ## Risks - The default company page now reads the all-actors scope, which returns more rows than the earlier agent-only query. Cursor pagination and the existing page limit bound each request. - The page is now visible to every company member. The server decides what each member sees. The UI only hides controls that the caller cannot use. Refs #10831 for the server rules and tests. - The basic tier now shows agent names that the previous revision withheld. The name was already recoverable from the retained `actorId` through the readable agent directory, so this closes an inconsistency rather than widening access. A security reviewer chose this outcome over stripping `actorId`. - Old audit links now redirect. The redirect keeps the agent scope, so a person who bookmarked the audit page sees the same rows. - Low migration risk. There is no database change. > The roadmap marks activity log and action attribution as shipped. This change improves that existing capability. ## Model Used Claude Opus 5 (`claude-opus-5`, 1M context) with extended thinking and tool use, run through Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
68ddd6a7a0 |
feat(activity): add two-tier all-actors audit feed (#10831)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Operators need one activity feed for human, agent, plugin, and system changes > - The existing audit endpoint returns only rows that have agent attribution > - The full audit view also requires a dedicated permission > - This pull request adds an explicit all-actors scope with basic and privileged access tiers > - The benefit is that company members can inspect the shared activity history while sensitive attribution and export controls stay protected ## Linked Issues or Issue Description **What existing behavior does this improve?** The company audit activity endpoint and the board audit route. **Subsystem affected** Server REST API and board UI routing/API contracts. **Current behavior** The agent-action audit endpoint excludes activity without an agent ID. It also rejects company members who do not have the full audit permission. **Proposed behavior** Callers can opt into `actorScope=all`. A company member receives all actor kinds with sensitive attribution fields removed. A permitted board user receives complete rows and can use attribution filters. The default scope and CSV permission remain unchanged. **Reason and benefit** The board needs one chronological activity source for user, agent, plugin, and system actions. A two-tier response keeps the feed useful without widening access to detailed attribution or export capabilities. **Breaking changes** None. The endpoint keeps the existing agent-only scope and permission behavior by default. ## What Changed - Added `actorScope=all` to the unified audit query and included activity from every actor type. - Added a company-readable basic tier that removes run, responsible-user, agent, and details attribution. - Kept attribution filters and CSV export behind `audit:view_agent_actions`. - Added route and integration coverage for basic readers, permitted readers, pagination, filter denial, and all actor kinds. - Added the missing unprefixed `/audit` redirect and company route classification. ## Verification - `pnpm exec vitest run server/src/__tests__/activity-routes.test.ts server/src/__tests__/agent-action-audit-routes.test.ts ui/src/lib/company-routes.test.ts --reporter=verbose` (35 tests passed) - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` ## Risks - The all-actors query can return more rows than the legacy agent-only query. Cursor pagination and existing limits bound each request. - The basic tier intentionally exposes action and actor-kind context. It removes detailed run, agent, responsible-user, and details attribution. - The legacy endpoint behavior remains the default, which reduces compatibility risk. > The roadmap marks activity log and action attribution as shipped. This change improves that existing capability and does not introduce a separate workflow system. ## Model Used - OpenAI Codex, `gpt-5.6-sol`, 114K context, agentic reasoning with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f0ed524ffa |
fix(ui): prefix audit board route (#10830)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board UI keeps company work under a company-prefixed route. > - The Audit sidebar link used a bare `/audit` path. > - The route helper treated `audit` as a company prefix because the board-route list did not include it. > - The router also had no redirect for a bare `/audit` deep link. > - This pull request registers Audit in both places and adds regression coverage. > - The benefit is that the Audit sidebar link and old bare deep links open the active company's audit feed. ## Linked Issues or Issue Description Related PR: #9744 **What happened?** The Audit sidebar link opened `/audit`. The router interpreted `AUDIT` as a company prefix and showed the invalid-company page. **Expected behavior** The Audit sidebar link must open `/<company-prefix>/audit`. A bare `/audit` deep link must redirect to the active company. **Steps to reproduce** 1. Open a company board. 2. Select Audit in the sidebar. 3. Observe that the app opens `/audit` and shows an invalid-company error. **Paperclip version or commit** Reproduced on master after #9744. **Deployment mode** Board UI in local or self-hosted deployments. ## What Changed - Added `audit` to the board-route root list. - Added the unprefixed `/audit` redirect route. - Added regression tests for Audit prefixing, prefix extraction, and relative-path conversion. ## Verification - `pnpm exec vitest run ui/src/lib/company-routes.test.ts` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - Manual check: select Audit in the sidebar and confirm the URL is `/<company-prefix>/audit` and the audit feed renders. ## Risks - Low risk. This change only reserves one existing board route and adds one redirect. - A company cannot use `AUDIT` as an issue prefix after this change. That prefix already conflicts with the existing Audit board page. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The runtime does not expose a more specific model ID or context-window size. The agent used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e8ae5286eb |
feat(issues): explain cross-task agent writes with attribution, audit receipts, and actionable denials (#10843)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents write to tasks they do not own. They comment, they change
fields, and the control plane now permits this by default for
standard-trust agents on any task they can read
> - This makes a task thread ambiguous. A reader sees a comment from an
agent that is not the assignee, but no surface says whose authority that
write rode
> - The same gap applies to field edits. The activity stream named the
verb, but it did not show the before value, the after value, or the
reason the write was permitted
> - The remaining refusals are also opaque. An agent that hits a wall
receives a 403 with no boundary name, no actor who can act, and no
sanctioned path. One real incident spent a full detour to find the
workaround
> - This pull request adds the three surfaces that make open cross-task
writes legible: an attribution chip, a field-level audit receipt, and an
actionable denial contract shared by the API and the UI
> - The benefit is that a reader can answer "who did this, on whose
authority, and was it allowed?" on the task itself, and a blocked writer
is told what to do next
## Linked Issues or Issue Description
No public issue exists for this work, so the enhancement is described
here.
**What existing behavior does this improve?**
Cross-task agent writes are permitted, but they are not explained. A
task thread can hold comments from agents that are not the assignee, and
the activity stream can hold field changes made by those agents. Neither
surface names the responsible user behind the write. When a write is
refused, the error text does not name the boundary or the way forward.
**Subsystem affected**
Issue detail UI (comment thread and activity stream), the issue write
authorization responses in the server, and the shared copy contract that
both consume.
**Current behavior**
- An agent comment on a task the agent does not own looks the same as an
assignee comment.
- An `issue.updated` activity row states the verb only. It does not show
the field-level before and after values, the responsible user, or the
authorization reason.
- A refused write returns a short message such as an ownership error.
The message does not state which rule fired, who is able to perform the
action, or which alternative path is sanctioned.
**Proposed behavior**
- An agent comment on a task the agent does not own carries a chip that
reads "for {user}". The chip names the responsible user. Its tooltip
states that the author is not the assignee and cannot exceed that user's
permissions.
- Each `issue.updated` row shows a receipt: the changed fields with
before and after values, the responsible user, and the authorization
reason. This applies to board edits as well as agent edits.
- Each refusal states three things: the boundary that fired, who is able
to act, and the sanctioned path. The API error body and the in-app
notice use the same words, because both read one shared contract.
Related pull requests, found by searching this repository:
- Refs #10837 — merged. It added the default-open cross-task write rule,
the comment attribution data, and the per-run containment cap that this
pull request makes visible.
- Refs #10114 — open. It proposes a narrower authorization change in the
same area.
- Refs #7998 — open. It proposes append-only cross-assignee comments as
an alternative to opening writes.
## What Changed
- Adds `packages/shared/src/issue-write-denial.ts`. This is one copy
contract for eight ways an issue write can be refused: not visible,
responsible-user ceiling, responsible user unavailable, excluded actor
class, assignee run lock, per-run cross-task cap, missing run context,
and rejected attribution. Each entry names the boundary, who can act,
and the sanctioned path.
- Maps server authorization decisions onto that contract in
`server/src/routes/issues.ts` and
`server/src/services/cross-issue-influence-limit.ts`. The flattened
`error` string carries all three obligations, and `details.code` lets
the UI render the same words. The two cap codes keep the names they
already ship under.
- Adds `CommentAttributionChip`. It renders "for {user}" beside the
author name on agent comments where the author is not the assignee. It
renders nothing when no responsible user is recorded, so older rows stay
clean. It is wired into both `IssueChatThread` and the flagged
`TaskChatThread` redesign.
- Adds `IssueFieldChangeReceipt`. It renders the change receipt under
`issue.updated` rows in the activity stream. Ids resolve to agent and
user names where the directory is loaded. Server-truncated text is
labelled as a preview, so the receipt never implies that it shows a
whole value.
- Adds `IssueWriteDenialNotice`. It renders the shared copy in the app,
keyed off the denial events the server logs on a task.
- Adds a public `/ux-lab/cross-issue-collaboration` page. It renders all
three surfaces and their edge cases for review without a seeded thread.
This follows the existing `ux-lab` pages.
## Verification
Automated, all green:
```
pnpm --filter @paperclipai/shared exec vitest run src/issue-write-denial.test.ts # 17 tests
pnpm --filter @paperclipai/ui exec vitest run src/components/IssueWriteDenialNotice.test.tsx \
src/components/IssueFieldChangeReceipt.test.tsx src/components/CommentAttributionChip.test.tsx \
src/lib/issue-change-receipt.test.ts src/lib/comment-attribution.test.ts # 46 tests
pnpm --filter @paperclipai/server exec vitest run src/__tests__/cross-issue-influence-limit.test.ts \
src/__tests__/issue-comment-attribution-audit-routes.test.ts \
src/__tests__/issue-agent-mutation-ownership-routes.test.ts \
src/__tests__/low-trust-red-team-routes.test.ts # 98 tests
```
`tsc --noEmit` passes for the shared, ui, and server packages.
Manual, in a browser:
1. Start the UI only: `pnpm --filter @paperclipai/ui exec vite`.
2. Open `/ux-lab/cross-issue-collaboration`. No session is needed,
because `ux-lab` routes are public.
3. All three surfaces were captured at 1440x900 in light mode and dark
mode, and at 390x844. The page reported no errors.
4. The chip tooltip was opened by a hover and by a keyboard focus.
Rendering the page found defects that the tests had missed. Three copy
and contrast defects were fixed, and two of them are now pinned by a
test. A design review then found three layout defects, which are also
fixed: the denial notice orphaned its label when a value wrapped, the
receipt icon wrapped onto its own line at narrow widths, and the chip
tooltip was reachable by hover only.
## Risks
Low risk, and additive.
- Every new surface renders nothing when its data is absent. Comments
without a recorded responsible user show no chip, and activity events
without a receipt show no receipt, so existing rows do not change.
- No migration is included. The data these surfaces read already ships.
- The wire values of the two per-run cap denial codes are unchanged.
Only the human-readable text changes, plus six codes that had no
`details.code` before.
- The denial copy is read by agents as well as people. If wording must
change later, one shared module is the only place to change it.
- Roadmap check: this extends the completed "Activity log & action
attribution" area rather than duplicating planned core work.
## Model Used
Claude Opus 5 (Anthropic), model id `claude-opus-5[1m]`, 1M context
window, extended thinking, with tool use and code execution. It ran as
an agent in Claude Code and drove a real browser to capture the review
screenshots.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
5858ccb981 |
feat: make in-app features cloud-aware (#10850)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators use the same board application in self-hosted and Paperclip Cloud deployments. > - A Cloud tenant contains one company, so an in-app company switch does not change the active Cloud stack. > - Cloud operators need the sidebar and company surfaces to use the signed-in user's stack portfolio. > - The server must derive Cloud identity and links from trusted instance context instead of client input. > - This pull request adds canonical Cloud context, a trusted stack portfolio proxy, and Cloud-aware navigation. > - The benefit is consistent stack switching on Cloud while self-hosted company behavior stays unchanged. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: server REST routes and the React board UI. **Problem or motivation** A Cloud-managed instance contains one company. The existing company switcher could only switch records inside that tenant. It could not move the operator to another Cloud stack. The existing header also gave long organization names too little width. **Proposed solution** Expose a canonical public Cloud context in health data. Add a trusted server proxy for the current user's stack portfolio. Use that data in the board UI to switch stacks with top-level navigation. Keep the existing company behavior on self-hosted instances. Move search into the navigation and keep long organization names inside the sidebar panel. **Alternatives considered** An in-app `/stacks` route was rejected because Cloud tenant hosts reserve that path and stack selection must wake or authenticate another tenant. Client-supplied user identity was rejected because the server can derive the trusted Cloud actor. **Roadmap alignment** This change advances the Cloud deployments milestone. It keeps the product local-first and Cloud-ready without changing the self-hosted mental model. ## What Changed - Added canonical Cloud instance context and public health metadata. - Added a Cloud-only stack portfolio proxy with trusted actor forwarding and per-user caching. - Prevented normal company creation on Cloud-managed instances. - Switched the sidebar and Companies page from company actions to stack actions on Cloud. - Added full-page stack navigation and Cloud create-stack links. - Moved search into the sidebar navigation so the organization name keeps more width. - Added truncation and hover recovery for long organization and stack names. - Added server and UI regression coverage for Cloud and self-hosted behavior. - Updated the implementation specification for the Cloud contracts. ## Verification - `node scripts/check-token-gates.mjs` passed. All three token gates are clean. - `pnpm --dir server exec vitest run src/__tests__/health.test.ts src/__tests__/cloud-instance.test.ts src/__tests__/cloud-routes.test.ts src/__tests__/company-cloud-floor.test.ts src/__tests__/company-portability-routes.test.ts` passed: 5 files and 66 tests. - `pnpm --dir ui exec vitest run src/components/SidebarCompanyMenu.test.tsx` passed: 1 file and 11 tests. - Pre-PR QA report `7da87ca7` passed all 8 acceptance criteria with real HTTP route factories and real Chromium screenshots in Cloud and self-hosted modes. - Security reviews passed for the canonical Cloud context and stack portfolio proxy. ## Risks - Cloud stack switching depends on the configured Cloud application and tenant portfolio URLs. - The new health `cloud` block is public by design, but it contains only canonical public instance metadata. - The stack proxy fails closed on self-hosted instances and derives the user identity from the trusted actor. - Self-hosted navigation and company creation retain their existing paths and behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5`. The run used reasoning, repository tools, shell execution, and GitHub integration. The deployment did not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |