mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
efcce9cc8e17e2b30f6f2edf30e00a77b0bf19bb
3057
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
efcce9cc8e |
fix(adapters): record unpriced CLI usage (#9505)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Budgets and spend telemetry are control-plane safety features, not just reporting > - Local Codex and Claude adapters can execute through either ACP or their native CLI engines > - The ACP lane records usage and reported cost, but CLI JSON output often reports tokens without a price > - The CLI lane was either losing per-run usage semantics or coercing missing cost to zero, making real usage indistinguishable from a genuinely free run > - This pull request preserves CLI usage as per-run totals and records token-bearing runs without a reported price as explicitly unpriced ledger events > - The benefit is accurate usage accounting and a visible pricing gap instead of silently misleading zero-cost telemetry ## Linked Issues or Issue Description Refs #9471 Refs #9230 **Bug description** A `codex_local` run using the CLI engine can emit a final `turn.completed` event with millions of input tokens and tens of thousands of output tokens while the agent's spend ledger remains indistinguishable from a true zero-usage, zero-cost run. Claude CLI output has the same missing-price edge case. **Expected behavior** Token-bearing CLI runs should persist their usage. If the adapter reports a price, the ledger should record it as reported; if the CLI reports usage but no price, the ledger should explicitly mark the event as unpriced rather than silently treating missing price data as a reported `$0` cost. **Reproduction shape** 1. Configure `codex_local` with `engine: cli`. 2. Run a task that produces a `turn.completed` usage payload. 3. Observe token usage in the run stream. 4. Before this change, missing price data is represented as ordinary zero-cost spend and the CLI usage basis is not consistently propagated. ## What Changed - Mark Codex and Claude native CLI usage totals as `per_run` and propagate that basis through success and failure results. - Stop coercing missing Claude CLI cost to `0`. - Add `cost_status` to cost events with `reported` and `unpriced` values, including an idempotent migration and shared validation/types. - Persist token-bearing runs without a reported price as `unpriced` ledger events while retaining zero cents until an authoritative price exists. - Add parser, execute-path, heartbeat-accounting, and cost-service regression coverage for both local CLI adapters. - Document the cost-status invariant and CLI accounting behavior. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/parse.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/claude-local-execute.test.ts server/src/__tests__/heartbeat-cost-accounting.test.ts server/src/__tests__/costs-service.test.ts` — 6 files / 102 tests passed. - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` — includes migration numbering and safety checks. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Existing cost rows default to `reported`, preserving current interpretation; only new token-bearing events with absent cost are marked `unpriced`. - This change does not invent model pricing. Budget hard stops still cannot charge an unknown amount, but operators and evals can now distinguish missing pricing from a genuinely reported zero cost. - Consumers that enumerate cost-event fields should tolerate the additive `costStatus` field. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model `gpt-5.3-codex`, with repository tool use and code execution; default reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ce7dedf33d |
perf(ci): balance general-server test shards by recorded suite duration (#9516)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its PR CI runs the general-server vitest lane pinned to `maxWorkers=1` and sharded across 3 runners (introduced in #8360) > - Suites were assigned to shards round-robin by sorted file index, so shard test time was unbalanced: a recent PR run split 73s / 153s / 115s, and the heaviest shard made "General tests (server 2/3)" the slowest check in the whole workflow at 314s wall > - The slowest shard sets the lane's wall time, so unbalanced partitions waste the other two runners and stretch the PR critical path > - This pull request replaces the round-robin assignment with a deterministic longest-processing-time partition weighted by a checked-in per-suite duration manifest > - The benefit is near-even shard weights (projected 113s / 113s / 113s with the current manifest), taking roughly 40s off the PR critical path with no reduction in coverage ## Linked Issues or Issue Description - Refs #8360 (introduced the 3-way general-server sharding this PR rebalances) - No public issue exists. Problem: the general-server test lane's round-robin shard assignment ignores per-suite duration, so one shard can carry multiple 30s+ suites while another finishes in half the time; the slowest shard alone determines the check's wall time. ## What Changed - `scripts/general-server-shard.mjs` (new): manifest loader and deterministic LPT (longest-processing-time) partitioner; suites missing from the manifest get the median recorded weight, and a missing or malformed manifest degrades to uniform weights so the lane never fails on stale data - `scripts/general-server-shard-durations.json` (new): per-suite duration manifest sampled from a real PR run (240 suites); the `$comment` field documents how to regenerate it - `scripts/run-vitest-stable.mjs`: both shard-selection sites (run and `--dry-run`) now use the balanced partition instead of index round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: 6 new tests covering skew-balance vs round-robin, determinism, median fallback for unlisted suites, malformed-manifest degradation, manifest coverage of the current suite set, and real-partition balance - `server/src/__tests__/heartbeat-issue-rewake-throttle.test.ts`: hardened the `afterEach` sweep — post-run bookkeeping (run-event records, follow-up wake scheduling) can still insert rows briefly after a run reaches a terminal status, and a late insert landing between the `agent_wakeup_requests` and `agents` deletes failed teardown with a foreign-key violation on the first CI attempt of this PR; the sweep now retries so a late background write cannot take down the shard - `release-verify.yml` shares the same runner script and inherits the balancing with no workflow change ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass (run against current master) - `npx vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6/6 pass against embedded Postgres with the hardened teardown - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass - `node scripts/run-vitest-stable.mjs --dry-run` with each shard flag shows every suite assigned exactly once across the 3 shards, with projected weights ~113s each ## Risks - Low risk: partition changes which runner executes which suite, not what runs; a completeness test asserts every suite is assigned to exactly one shard - The duration manifest will drift as suites are added/changed; unlisted suites get the median weight and a coverage test flags when the manifest covers less than half the suite set, so drift degrades balance gracefully rather than breaking the lane > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking enabled, agentic tool use (file edits, shell, test execution) via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude (Paperclip SWE) <noreply@paperclip.ing> |
||
|
|
b49d178c46 |
fix(ui): experiments auto-recovery dialog leaves UI dimmed and locked after enabling (#9513)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Operators tune instance behavior through Settings → Experiments,
where experimental features are toggled on and off
> - The task graph liveness auto-recovery experiment shows a
confirmation dialog (preview of what would be recovered) before it is
enabled
> - After confirming with "Enable only" or "Enable and run", the
dialog's Radix overlay and the `pointer-events: none` body lock were
left behind, dimming the page and blocking all interaction until a
refresh
> - The dialog was unconditionally mounted and only closed inside the
mutation's `onSuccess`, so the overlay teardown depended on the mutation
outcome and could race or never happen
> - This pull request closes the dialog before the mutation fires in
both confirm flows, clears the pending preview alongside the open flag,
and mounts the dialog conditionally so its overlay fully unmounts
> - The benefit is that enabling an experiment behaves like every other
settings change: the dialog goes away, the page stays interactive, and
errors surface in the page-level error banner instead of a dead UI
## Linked Issues or Issue Description
No public GitHub issue exists for this bug; description follows the bug
report template. Refs #4587 (the PR that introduced the configurable
liveness auto-recovery controls this dialog belongs to).
**What happened?** In Settings → Experiments, toggling on "Task graph
liveness auto-recovery" and confirming via "Enable only" left the whole
UI dimmed and unclickable. The dialog content disappeared, but the modal
overlay and the `pointer-events: none` lock on `<body>` remained until a
full page refresh.
**Expected behavior:** Confirming (or dismissing) the auto-recovery
dialog should close it completely and return the page to a fully
interactive state, with the toggle reflecting the new setting.
**Steps to reproduce:**
1. Open Settings → Experiments.
2. Toggle on "Task graph liveness auto-recovery"; the confirmation
dialog with the recovery preview appears.
3. Click "Enable only".
4. The dialog content disappears but the page stays dimmed and nothing
is clickable; refreshing restores the UI and shows the setting was
applied.
**Paperclip version or commit:** master @
|
||
|
|
0e21a27301 |
fix: forward onSpawn to hermes and process adapters for PID persistence (#8722)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter layer (hermes-local, process adapters) delegates agent execution to child processes via `runChildProcess()` > - `runChildProcess()` accepts an `onSpawn` callback to report child PID and process group info, but the hermes and process adapters were not forwarding `ctx.onSpawn` to this call > - Without PID persistence, the orphan reaper cannot distinguish live runs from abandoned processes, causing false-positive reaps and 5-minute timeout errors for active runs > - This pull request adds `onSpawn: ctx.onSpawn` to both adapter call sites and declares the option in the `runChildProcess` wrapper type > - The benefit is that the orphan reaper can now correctly track live child processes, eliminating false-positive reaps ## Linked Issues or Issue Description Fixes #8723 Fixes false-positive orphan reaps in hermes-local and process adapters by forwarding the `onSpawn` callback to `runChildProcess()`. All other adapters (claude-local, codex-local, cursor-local, gemini-local, grok-local, opencode-local, pi-local) already forward `ctx.onSpawn` — these two were the only ones missing it. ## What Changed - `server/src/adapters/utils.ts`: Added `onSpawn?` to the `runChildProcess()` options type so callers can forward the callback - `server/src/adapters/process/execute.ts`: Forward `ctx.onSpawn` to `runChildProcess()` - `packages/adapters/hermes/src/server/execute.ts`: Forward `ctx.onSpawn` to `runChildProcess()` ## Verification - `pnpm -r typecheck` passes across all packages - Confirmed all other adapters already forward `ctx.onSpawn` (12 grep matches across 9 adapter files) - The 3-line diff is additive only — no existing behavior is changed, only a previously-ignored callback is now forwarded ## Risks Low risk. This is a 3-line additive change. The `onSpawn` parameter is optional (`?`) so existing callers are unaffected. The callback is already well-established across all other adapters. ## Model Used Hermes Agent (by Nous Research) — xiaomi/mimo-v2.5-pro via OpenRouter, with tool use (file editing, git, GitHub API). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass (typecheck passes) - [x] I have added or updated tests where applicable (N/A — type-level fix only, no behavioral change) - [x] I have updated relevant documentation to reflect my changes (N/A — internal fix) - [x] I have considered and documented any risks above --------- Co-authored-by: Zephyr <zephyr@motoyuki.dev>canary/v2026.713.0-canary.2 |
||
|
|
634ae1298f |
fix(ui): consolidate live/running blues; stop inbox unread badge indenting the row (#9383)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The web UI leans on a shared design-token + component system so surfaces stay visually consistent as they grow > - Two small inconsistencies had crept in: several distinct blues were used to signal "live/running" agent state across the sidebar, task header, and chat thread; and in the Inbox an unread task's mark-read dot was pushing that row's status icon and title one column right of read rows > - Both read as "not quite aligned" in daily use and undercut the polish of the lists work that just landed > - This pull request consolidates the live/running blues onto one shared recipe and stops the unread dot from indenting the row > - The benefit is one consistent "live" blue everywhere and Inbox rows that line up whether read or unread ## Linked Issues or Issue Description No public GitHub issue exists for this work; describing inline per the bug-report template. - **Problem**: (1) the same concept — an agent actively working — rendered in three visibly different blues: the sidebar `N live` dot, the task-detail "Live" badge, and the chat-thread "RUNNING" badge each used a different token/recipe. (2) In the Inbox, unread rows carry a leading mark-read dot that occupies the chevron column, but a per-row spacer was still rendering in that same column — so on unread rows the status icon + title were shifted one column (~24px) further right than read rows. Most visible when grouped by workspace. - **Steps to reproduce**: open the Inbox with a mix of read and unread tasks (group by workspace). The unread rows' status icons sit further right than the read rows'. Separately, compare the blue of the sidebar `N live` dot, a task's "Live" header badge, and a chat "RUNNING" badge — they don't match. - **Expected behavior**: unread and read rows align on the same status column, with the unread dot centered on the workspace group chevron; and all three "live/running" affordances share one blue. ## What Changed - Added a shared `liveBlueBadge` recipe in `ui/src/lib/status-colors.ts` and pointed the task-detail **Live** badge (`IssueDetail.tsx`) and the chat-thread **RUNNING** badge (`IssueChatThread.tsx`) at it; removed the now-redundant `brandChipBadge` usage from the chat thread and a stray `🔵` breadcrumb prefix. - Changed the sidebar **`N live`** dot (`SidebarNavItem.tsx`) to the same `blue-600 / dark:blue-400` as its adjacent label text. - **Inbox** (`Inbox.tsx`): skip the per-row leading spacer when the unread mark-read dot is present, so the dot alone fills the chevron column. Unread rows' status icon + title now sit in the same column as read rows, and the dot centers on the workspace group chevron. - **Test** (`Inbox.test.tsx`): added a regression test asserting an unread leaf row renders the mark-read dot and drops the spacer, while a read row keeps the spacer. ## Verification - `pnpm typecheck` — clean (all packages) - `pnpm check:token-gates` — 3/3 CLEAN - `cd ui && pnpm vitest run src/pages/Inbox.test.tsx` — 14/14 (includes the new regression test) - Full Storybook visual suite (514 stories, both themes) — green locally (CI cannot run this suite yet — the baseline-manifest archive is unpublished, a pre-existing condition from #9134) - Manual (workspace-grouped Inbox, 2× dark): measured the unread badge center at the same x as the workspace chevron (276 = 276) and the unread-row status icon at the same x as read-row status icons (292 = 292). Before/after screenshots in a PR comment below. ## Risks Low risk — presentation only. No data, routing, or state changes. The blue consolidation is a token/class swap; the Inbox change removes a redundant spacer element on unread rows only (read rows and non-grouped/mobile views are unaffected). The unread-row behavior is covered by the new unit test. ## Model Used Claude (Anthropic), Opus 4.8 — model id `claude-opus-4-8`; extended thinking + tool use, driving local verification (typecheck, token gates, vitest, Playwright visual suite + pixel measurements). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8775bde4ce |
Add runtime asset build-gap guard
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server package ships runtime asset trees used by built-in agents
and onboarding templates.
> - A prior server build omitted those source asset trees from `dist`,
allowing a built artifact to differ from runtime expectations.
> - The copy step is now present, but the existing build-gap gate only
checked TypeScript coverage for packages whose build skips `tsc`.
> - This pull request extends that standing gate so source asset files
under the server runtime asset trees must exist at the matching `dist`
paths after build.
> - The benefit is that future server asset additions fail loudly in CI
instead of silently shipping an incomplete `dist`.
## Linked Issues or Issue Description
### What happened?
After a server build, runtime asset files under
`server/src/built-ins/**` and `server/src/onboarding-assets/**` could be
missing from `dist/` with no build failure. The existing build-gap gate
only checked TypeScript coverage for packages that skip `tsc`; it did
not verify that non-TypeScript source assets were copied to `dist`. A
server build that forgot the `cp -R` step, or that added a new asset
tree without updating the copy command, would produce an incomplete
`dist` without any CI signal.
### Expected behavior
After `pnpm --filter @paperclipai/server build`, every
non-TypeScript/non-JavaScript source file under `server/src/built-ins/`
and `server/src/onboarding-assets/` must exist at the matching path
under `server/dist/`. If any file is missing, the build-gap gate must
exit non-zero with a diagnostic listing the missing files and the
command to fix them.
### Steps to reproduce
1. Remove a copied runtime asset: `rm
server/dist/built-ins/agents/reflection-coach/AGENTS.md`
2. Run the guard: `node scripts/run-typecheck-build-gaps.mjs
--runtime-assets-only`
3. Before this fix: the command exits 0 and the missing file goes
undetected.
### Paperclip version or commit
Reproduced on `master` at `c36f1a4af` (`@paperclipai/server` 0.3.1).
### Deployment mode
Not deployment-specific — the build-gap check runs in CI on any
checkout.
## What Changed
- Extended `scripts/run-typecheck-build-gaps.mjs` with a source-derived
server runtime asset parity check for non-`.ts`/non-`.js` files under
`server/src/built-ins/**` and `server/src/onboarding-assets/**`.
- Added a guard-only mode, `--runtime-assets-only`, for focused
pass/fail verification after a server build.
- Wired `pnpm run typecheck:build-gaps` to prepare plugin SDK build
deps, build the server package, then run the existing build-gap gate
plus the new asset check.
## Verification
Pass path:
```text
$ pnpm --filter @paperclipai/plugin-sdk ensure-build-deps
> @paperclipai/plugin-sdk@1.0.0 ensure-build-deps .../packages/plugins/sdk
> node ../../../scripts/ensure-plugin-build-deps.mjs
$ pnpm --filter @paperclipai/server build
> @paperclipai/server@0.3.1 build .../server
> tsc && mkdir -p dist/onboarding-assets dist/built-ins && cp -R src/onboarding-assets/. dist/onboarding-assets/ && cp -R src/built-ins/. dist/built-ins/
$ node scripts/run-typecheck-build-gaps.mjs --runtime-assets-only
[typecheck:build-gaps] server runtime assets present in dist: 7 file(s)
```
Regression simulation (guard catches the missing file):
```text
$ rm server/dist/built-ins/agents/reflection-coach/AGENTS.md
$ node scripts/run-typecheck-build-gaps.mjs --runtime-assets-only
[typecheck:build-gaps] Missing server runtime asset(s) in dist:
- source: server/src/built-ins/agents/reflection-coach/AGENTS.md
expected dist: server/dist/built-ins/agents/reflection-coach/AGENTS.md
Run pnpm --filter @paperclipai/server build and ensure source runtime asset trees are copied into dist.
```
Standing gate (full end-to-end):
```text
$ pnpm run typecheck:build-gaps
[typecheck:build-gaps] typechecking 4 workspace(s): paperclipai, @paperclipai/plugin-authoring-smoke-example, @paperclipai/plugin-llm-wiki, @paperclipai/ui
[typecheck:build-gaps] server runtime assets present in dist: 7 file(s)
```
## Risks
Low risk. The check only reads source and dist files during the
build-gap gate. The main tradeoff is that the gate now builds
`@paperclipai/server` so a clean checkout has generated `dist` content
to validate.
## Model Used
OpenAI Codex, GPT-5 based coding agent with repository tool use and
shell execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.713.0-canary.1
|
||
|
|
c36f1a4afd |
fix(ui): make mobile decision rows readable (#9472)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies > - Human operators use the Decisions attention queue to review and resolve work that needs them > - Decision rows were composed as a fixed content column plus a right-side controls column > - At phone widths, timestamps, actions, menus, and evidence thumbnails compressed the decision headline until it was barely readable > - This pull request makes each row respond to its own container width and stacks metadata, content, evidence, and actions on narrow surfaces while preserving the dense desktop layout > - The benefit is a useful, thumb-reachable Decisions workflow on phones and narrow side panels without regressing wide-screen density or scrolling performance ## Linked Issues or Issue Description No public GitHub issue exactly matches this bug, so it is described here using the bug-report fields. **What happened** Decision rows used a fixed two-column layout. On narrow screens, the right-hand timestamp, overflow menu, decision buttons, and optional thumbnails squeezed the headline into a truncated sliver. **Expected behavior** Decision headlines should remain readable on mobile, supporting context should flow below the headline, and primary actions should remain easy to tap. Wide rows should retain the compact desktop presentation. **Steps to reproduce** 1. Open the Decisions / What needs me surface with populated attention items. 2. Reduce the row container to a phone-width layout (approximately 390px). 3. Observe rows with multiple actions or evidence thumbnails. **Paperclip version / deployment mode** Current `master`, board UI in local or hosted deployments. **Related public work found during dedup search** - Refs: #9311 — original What needs me attention queue work. - Refs: #9468 — recent Decisions scrolling performance work preserved by this change. ## What Changed - Reworked `AttentionQueueRow` into a container-query-driven vertical stack on narrow surfaces, with the existing compact layout restored at wide row widths. - Made decision titles wrap to two lines, moved project/evidence context below the headline, and promoted actions to full-width mobile tap targets. - Preserved upstream row memoization and `content-visibility` scrolling optimizations while rebasing onto current `master`. - Added three 390px Storybook scenarios covering populated rows, type/detail variants, and snoozed/dismissed curtains. - Updated the focused row test to assert the new thumbnail/context alignment. ## Verification - `pnpm exec vitest run ui/src/components/AttentionQueueRow.test.tsx` — 1 file passed, 16 tests passed. - `pnpm check:token-gates` — all token gates clean. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/ui build-storybook` — completed successfully. - `git diff --check public/master...HEAD` — passed. ## Risks - Low risk: the behavior is isolated to the Decisions row presentation and its Storybook coverage. - Container-query breakpoints could need future visual tuning for unusual embedded widths, but the wide layout remains available at the row-level breakpoint. - The mobile layout increases row height by design in exchange for readable content and usable actions. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex CLI coding agent. The exact model ID and context-window size are not exposed to this runtime; reasoning, repository editing, shell execution, and test execution capabilities were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8a0db228a6 |
perf(ui): improve Decisions scrolling performance (#9468)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Operators use the Decisions page to review an uncapped attention feed across active, snoozed, and dismissed items > - Large feeds mounted every row eagerly, and routine interactions re-rendered the full queue > - That made initial paint and scrolling progressively slower as decision history accumulated > - This pull request bounds rendering, stabilizes row props, and lets off-screen rows skip layout and paint work > - The benefit is a responsive Decisions page even for companies with large attention histories ## Linked Issues or Issue Description ### What happened? Opening `/decisions` for a company with a large attention history eagerly mounted every visible-feed row. Expanding, selecting, dismissing, snoozing, or restoring an item could also re-render the entire queue. ### Expected behavior The page should render a bounded initial window, progressively reveal more rows near the scroll boundary, and avoid re-rendering unaffected rows during interactions. ### Steps to reproduce 1. Populate a company with hundreds of attention items. 2. Open `/decisions`. 3. Scroll and interact with individual rows. 4. Observe increasing initial render, layout, paint, and interaction cost on the previous implementation. ### Paperclip version or commit Reproduced on `master` before this PR. ### Deployment and installation Local development, built from source. This is a core UI issue, not adapter- or database-specific. ### Additional context Searched open public issues and PRs; no duplicate was found. ## What Changed - Added a pure `planAttentionRenderRows` helper that allocates one render budget across active groups and open snoozed/dismissed curtains in document order. - Render 50 rows initially and add 100 more when the Decisions page approaches the scroll boundary. - Memoized `AttentionQueueRow`, stabilized parent callbacks and inbox dismissal actions, and passed row items through a shared expand callback. - Added `content-visibility: auto` and intrinsic containment so accumulated off-screen rows avoid unnecessary layout and paint work. - Added render-plan coverage and a regression test proving identical row props do not re-render after a parent update. ## Verification - `pnpm -C ui typecheck` - `pnpm -C ui exec vitest run src/lib/attention.test.ts src/components/AttentionQueueRow.test.tsx src/components/Sidebar.test.tsx src/pages/Inbox.test.tsx` — 96 tests passed - `pnpm check:token-gates` — all gates clean ## Risks - Low risk: the change is UI-only and does not alter API or database contracts. - The main behavioral risk is incorrect row-budget accounting across collapsed groups or open curtains; the pure planner has focused tests for ordering, truncation, and collapsed/closed sections. - Progressive rendering means rows beyond the current budget are intentionally absent until scrolling nears the boundary, matching the existing Issues list pattern. > This is a targeted performance fix and does not overlap planned core feature work in `ROADMAP.md`. ## Model Used - Anthropic Claude Fable 5 assisted with the implementation using repository tools and code execution. - OpenAI Codex `gpt-5.6-sol` prepared and verified the PR with high reasoning effort, repository tools, shell execution, and GitHub/Paperclip API access. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no documentation changes were required for this UI-only behavior) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c8253e3641 |
fix(adapters): inject execution contract once per fresh heartbeat (#9469)
## Thinking Path > - Paperclip coordinates AI-agent work through repeated heartbeat runs. > - Adapter prompts combine a default heartbeat template with scoped wake context. > - Fresh heartbeats received the same execution contract from both layers, wasting prompt tokens and obscuring which layer owns the contract. > - Resume deltas and template-less adapters do not share that composition path, so removing the wake-payload copy unconditionally would drop required guidance. > - Empty comment batches also emitted instructions and metadata that only matter when comments exist. > - This pull request makes execution-contract inclusion explicit by prompt path, preserves OpenClaw gateway behavior, and suppresses no-op comment boilerplate. > - The benefit is one contract per heartbeat path and roughly 300 fewer prompt tokens on a fresh zero-comment wake. ## Linked Issues or Issue Description - Fixes #9221 - Refs #9200 - Refs #7634 ## What Changed - Stop emitting the execution-contract paragraph from fresh scoped wake payloads because the default heartbeat template already contains the full contract. - Keep the contract in resume deltas, and add `includeExecutionContract` for adapters that do not render the default heartbeat template. - Opt `openclaw-gateway` into wake-payload contract rendering so template-less gateway runs retain the guidance. - Omit comment-batch acknowledgement/fetch guidance and empty `pending comments` / `latest comment id` metadata when a fresh wake has no pending comments. - Add regression and acceptance coverage proving composed fresh prompts contain `Execution contract` exactly once while resume and template-less paths retain it. Measured effect: the fresh zero-comment wake block drops from 1,840 to 855 characters (about 300 tokens saved per fresh heartbeat; about 220 on comment wakes), and the composed fresh prompt contains `Execution contract` once instead of twice. ## Verification - `npx vitest run packages/adapter-utils/src/server-utils.test.ts` — 63 passed - `npx vitest run server/src/__tests__/codex-local-execute.test.ts` — 13 passed - `npx vitest run server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/openclaw-gateway-adapter.test.ts server/src/__tests__/low-trust-red-team-routes.test.ts` — 27 passed - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed - `pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck` — passed ## Risks - Low risk: prompt text and adapter composition only; no database or API migration. - The main compatibility risk is a template-less adapter losing the contract. The explicit option and OpenClaw gateway regression coverage protect the known template-less path. - External adapters that call `renderPaperclipWakePrompt` directly can opt into `includeExecutionContract: true` when they do not render the default template. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.4, reasoning mode with tool use and code execution; context-window size is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5dff52631d |
fix(heartbeat): throttle redundant issue re-wakes (#9470)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents and their work > - Heartbeat admission decides when an agent should start another adapter session for an issue > - After process-loss recovery, assignment pollers and reconcilers can repeatedly request another wake while the issue remains `in_progress` > - When the preceding runs succeeded without issue-visible progress, those event-free wakes provide no new information but still pay the full cost of an adapter session > - Existing liveness evidence is too broad for this case because workspace tool calls can make a run look active without moving the issue > - This pull request adds an issue-scoped admission throttle for consecutive no-progress re-wakes while preserving every wake that carries new information or recovery intent > - The benefit is bounded recovery cost without delaying comments, operator actions, failures, or other meaningful events ## Linked Issues or Issue Description No public GitHub issue exists for this bug. **What happened?** After a process died, external wake drivers could re-wake the same agent for the same `in_progress` issue every few seconds. Each succeeded run that produced no issue-visible progress could be followed by another full adapter session despite no new issue input. In the observed recovery smoke, one recovery consumed 25 sessions and 2.4× the direct-run cost. **Expected behavior** Repeated event-free re-wakes should back off after consecutive successful runs produce no issue-visible progress. Any new information, explicit operator intent, or failed-run recovery should continue immediately. **Steps to reproduce** 1. Start an issue heartbeat and simulate process loss while the issue remains `in_progress`. 2. Allow assignment/reconciliation drivers to request repeated event-free wakes for the same agent and issue. 3. Complete each follow-up run successfully without adding a comment, issue mutation, document, work product, interaction, or continuation. 4. Observe repeated adapter sessions starting every few seconds without new issue input. **Environment** - Version: reproduced on `master` before this change - Deployment: local development, built from source - Adapter scope: core bug; not adapter-specific - Database: reproduced and tested with embedded Postgres ## What Changed - Add a pure issue re-wake throttle that detects consecutive succeeded runs without issue-visible progress and applies a 120-second exponential cooldown capped at 30 minutes. - Gate event-free `enqueueWakeup` requests and return the explicit skip reason `issue_rewake_throttled` while the cooldown is active. - Always bypass throttling for comment wakes, new issue activity, explicit resumes, `forceFreshSession`, event-shaped reasons, and post-failure recovery. - Add focused pure unit coverage and database-backed heartbeat admission coverage for throttle and bypass behavior. ## Verification - `cd server && pnpm vitest run src/__tests__/issue-rewake-throttle.test.ts` — 12 passed. - `cd server && pnpm vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6 passed with embedded Postgres. - `cd server && pnpm run typecheck` — passed. - Neighbor suites previously verified: `heartbeat-dependency-scheduling`, `heartbeat-process-recovery`, `run-continuations`, `heartbeat-issue-liveness-escalation`, `recovery-stale-issue-lock-sweep`, and `heartbeat-comment-wake-batching` — 131 tests passed. ## Risks - A progress classifier that is too narrow could defer a legitimate event-free poll; the cooldown is bounded and new issue activity bypasses it immediately. - A progress classifier that is too broad could allow the original heartbeat storm; tests intentionally distinguish issue-visible mutations from workspace-only activity. - Low compatibility risk: no schema, API contract, or migration changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent. The runtime does not expose the exact underlying model ID or context-window size; reasoning, terminal tool use, code inspection, GitHub CLI access, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public PR branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9e7e84e3fe |
fix(adapters): propagate ACP-lane usage and cost into spend telemetry (#9471)
## Thinking Path > - Paperclip is the open source control plane for running and governing AI-agent companies. > - Adapter executions feed token usage, billing identity, and run cost into the control plane's spend telemetry. > - The default ACP execution lane for local Claude and Codex adapters did not propagate per-turn usage or cost, so paid runs could be recorded with zero spend and no tokens. > - Claude CLI result events could also undercount output tokens by reading only the main-loop usage block instead of the complete per-model ledger. > - The shared executor needs to distinguish per-run usage from session-cumulative usage so the server does not apply the wrong delta heuristic. > - This pull request captures ACP usage and cumulative-cost deltas, resolves adapter billing identity, uses Claude's complete model-usage ledger, and preserves per-run usage in server normalization. > - The benefit is accurate token and cost accounting across the default paid Claude and Codex execution paths. ## Linked Issues or Issue Description ### What happened? Paid `claude_local` and `codex_local` runs using the default ACP engine can complete successfully while the control plane records zero or null cost and missing token usage. Claude CLI result parsing can additionally undercount output tokens when subagent or sidechain usage is present. ### Steps to reproduce 1. Run a paid Claude or Codex local adapter through the ACP engine. 2. Complete a turn that reports usage and cumulative cost through ACP status/events. 3. Inspect the execution result and normalized run telemetry. ### Expected behavior The execution result contains per-turn token usage, a per-run USD cost delta, and the correct billing identity. Server normalization records those per-run values without applying a session-cumulative delta a second time. ### Actual behavior before this change ACP execution results returned no usage and `costUsd: null` with unknown billing. The server therefore recorded zero spend and no tokens for paid runs. Claude CLI parsing could use an incomplete usage block. ## What Changed - Capture ACP usage from runtime status and `usage_update` events, reporting it as `usageBasis: per_run`. - Convert agent-reported cumulative ACP cost into a per-turn delta, including counter-reset and no-report safeguards. - Add a shared billing-identity resolver and map Claude and Codex authentication/provider modes to control-plane billing types. - Prefer Claude result-event `modelUsage` totals so subagent and sidechain tokens are included. - Skip the server's session-cumulative usage delta when an adapter explicitly reports per-run usage. - Add regression coverage for usage capture, event fallback, cost resets, stale reports, billing identities, model-usage totals, and server spend normalization. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts packages/adapters/claude-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/acp.test.ts server/src/__tests__/costs-service.test.ts server/src/__tests__/monthly-spend-service.test.ts` — 6 files, 126 tests passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - `pnpm --filter @paperclipai/adapter-claude-local typecheck` — passed. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - A broader Claude-local suite has a pre-existing rate-limit classification failure in `test.probe.test.ts`; it also fails on clean `master` and is unrelated to this change. ## Risks - Cost reporting depends on the agent's cumulative counter semantics; reset handling falls back to the post-turn amount and is covered by regression tests. - Incorrect billing-mode inference could misclassify spend; provider/auth mappings mirror each adapter's existing CLI behavior and have focused tests. - The new `usageBasis` contract changes server normalization only when adapters explicitly opt into `per_run`; existing adapters retain prior behavior. - No database migration, workflow, lockfile, or UI changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Implementation commit: Anthropic Claude Fable 5, tool-enabled coding workflow (exact context window and runtime configuration were not recorded in the commit metadata). - PR preparation and verification: OpenAI Codex, tool-enabled coding agent (runtime model ID and context window are not exposed to this session). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.713.0-canary.0 |
||
|
|
4a40c0cb13 |
feat(routines): gate scheduled runs on external activity (#9436)
## Thinking Path > - Paperclip is the open source app people use to manage AI-agent companies and their recurring work > - Scheduled routines provide native cron-driven execution for recurring agent tasks > - Watcher-style routines currently dispatch a model run even when the control plane has been quiet since their last useful run > - Existing pause, catch-up, and concurrency policies do not distinguish external work from a routine's own bookkeeping > - This pull request adds a generic activity gate that checks company-scoped activity provenance before scheduled dispatch > - The benefit is backward-compatible zero-token quiet skips while real human, agent, or delegated-child activity still wakes the routine ## Linked Issues or Issue Description - Refs #8534 ## What Changed - Added `activity_gate_policy` and `activity_gate_scope` routine columns with backward-compatible `always` / `company` defaults. - Added a company-bounded `evaluateActivityGate()` predicate that uses the last dispatched run as its open window, excludes the routine's own execution runs and scheduler bookkeeping, ignores pure-read actions, and supports company/project scope. - Integrated the predicate into scheduled ticks after pause/worktree eligibility checks; quiet ticks create visible skipped run-history rows with reason `no_external_activity` and gate-window diagnostics without advancing the activity window. - Kept webhook, manual, and API dispatch paths ungated; catch-up schedules evaluate the gate once per scheduler tick. - Added migration-default, provenance predicate, project-scope, quiet-window, scheduler, and webhook-bypass coverage. ## Verification - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts` — 51 tests passed - Embedded Postgres `EXPLAIN` for the company-scope gate scan: ```text Limit (cost=24.56..24.58 rows=1 width=24) -> Incremental Sort (cost=24.56..24.60 rows=2 width=24) Sort Key: activity.created_at, activity.id Presorted Key: activity.created_at -> Nested Loop Anti Join (cost=0.44..24.55 rows=1 width=24) Join Filter: (own_run.id = activity.run_id) -> Index Scan using activity_log_company_created_idx on activity_log activity (cost=0.15..8.19 rows=1 width=40) Index Cond: ((company_id = '00000000-0000-0000-0000-000000000001'::uuid) AND (created_at > (now() - '01:00:00'::interval)) AND (created_at <= now())) ``` ## Risks - The migration adds two non-null text columns, but constant defaults preserve all existing routine behavior and avoid a backfill step. - Project scope resolves activity through issue/run/routine provenance; tests cover in-project and cross-project issue activity, while every top-level and correlated query remains company-bounded. - This is the scheduler/schema foundation. Public API validation and documentation for configuring the new fields are intentionally handled in the next scoped follow-up. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.4` with medium reasoning, repository/tool access, terminal code execution, and test execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR extends the existing Scheduled Routines roadmap item - [x] I have searched GitHub for duplicate or related PRs and linked the related efficiency request above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing configuration is exposed in this scoped foundation PR) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.712.0-canary.0 |
||
|
|
e4e12bfb89 |
fix(workspaces): persist readiness state and validate ports (#9408)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution workspaces can run managed services that must report reliable lifecycle and readiness state > - Service startup previously waited for readiness before committing the starting row, making concurrent control actions see stale state > - Fixed service ports also needed clearer configuration and ownership diagnostics to avoid cross-workspace collisions > - This pull request persists startup state before readiness, validates port ownership, and exposes configurable service ports in the workspace UI > - The benefit is dependable service controls and actionable diagnostics when workspace runtimes start slowly or compete for ports ## Linked Issues or Issue Description ### What happened? Slow-starting workspace services could remain invisible to concurrent stop/restart controls until readiness completed, and fixed-port conflicts lacked enough ownership context for safe repair. ### Expected behavior A starting service is persisted immediately, control operations can observe it, configured ports are editable, and conflicts identify the owning process/workspace. ### Steps to reproduce 1. Configure a workspace service that delays binding its HTTP port. 2. Start the service and immediately request another control action. 3. Observe stale persisted state before this change. 4. Configure two workspaces for the same fixed port and observe limited conflict diagnostics. ### Paperclip version or commit `origin/master` at `02e2dd271` ### Deployment mode Local dev; built from source; not adapter-specific; database-backed workspace runtime state. ## What Changed - Commit the `starting` runtime-service row before waiting for readiness and transition it after the probe completes. - Add port-owner inspection and cross-workspace conflict details to local service supervision. - Preserve configurable runtime service ports through workspace configuration updates. - Surface service-port editing and validation in the execution workspace details UI. - Add server and UI regression coverage for slow readiness, concurrent controls, port persistence, and conflict diagnostics. ## Verification - `vitest --project @paperclipai/server src/__tests__/workspace-runtime.test.ts src/__tests__/execution-workspaces-service.test.ts` — 118 tests passed. - `vitest --project @paperclipai/ui src/pages/ExecutionWorkspaceDetail.service-ports.test.ts` — 4 tests passed. - `node scripts/check-token-gates.mjs` — all token gates clean. ## Risks - Moderate risk: changes touch workspace service lifecycle persistence and local process/port inspection. - No schema migration is required; tests exercise slow readiness, concurrent control, persisted ports, and cross-workspace conflicts. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and code execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.711.0-canary.6 |
||
|
|
a739a8dfba |
docs(skill): clarify workspace port-conflict recovery (#9407)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Isolated development workspaces need a repeatable run, verify, and repair procedure > - Port conflicts can be caused by another live Paperclip run that keeps respawning and reclaiming a configured port > - Restarting the target service alone does not resolve that class of conflict > - This pull request teaches the workspace-repair skill to identify the owner, guard the master checkout, and stop the conflicting run before repair > - The benefit is safer recovery guidance that addresses the actual port owner instead of creating restart loops ## Linked Issues or Issue Description ### What happened? The workspace repair procedure could recommend restarting a managed service while a separate live run still owned and reclaimed the configured port. ### Expected behavior The procedure identifies the owning process/run, protects the live master checkout, and stops the conflicting owner before restarting the intended service. ### Steps to reproduce 1. Start two managed workspace runs configured for the same fixed port. 2. Restart only the target workspace service. 3. Observe the sibling run reclaiming the port and the repair failing to hold. ### Paperclip version or commit `origin/master` at `02e2dd271` ### Deployment mode Local dev; built from source; not adapter-specific; not database-related. ## What Changed - Expand the port-conflict diagnosis to distinguish dead owners from live respawning runs. - Add master-checkout safety checks before killing or restarting processes. - Document owner-first recovery and explicit verification of final port ownership. - Tighten the success checklist so a repaired workspace must prove health and correct ownership. ## Verification - Reviewed the rendered Markdown diff and command sequence for consistent owner-first recovery. - No executable code changes are included in this documentation-only PR. ## Risks - Low risk: documentation and agent procedure only. - Process termination guidance remains intentionally guarded by owner identification and master-worktree checks. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and code execution; context-window size was not exposed by the runtime. The original change also credits Claude Fable 5 in the commit trailer. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
07e528256e |
fix(ui): bound shared polling cache (#9406)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The board coordinates repeated API polling across tabs to reduce redundant requests > - The shared polling coordinator retained cached result and publication entries after the last subscriber left > - Dynamic polling keys could therefore grow those maps for the lifetime of the page > - This pull request evicts inactive keys while preserving useful short-lived handoff state and request deduplication > - The benefit is bounded client memory without regressing cross-tab polling behavior ## Linked Issues or Issue Description ### What happened? Shared polling cached result/publication entries indefinitely after a polling key no longer had subscribers. ### Expected behavior Inactive keys are eventually removed, while recently published values remain available long enough for normal subscriber handoff. ### Steps to reproduce 1. Create and unsubscribe many distinct shared polling keys in one page lifetime. 2. Inspect the coordinator's cached results and publication timestamps. 3. Observe that the old maps retain every historical key. ### Paperclip version or commit `origin/master` at `02e2dd271` ### Deployment mode Local dev; built from source; not adapter-specific; not database-related. ## What Changed - Track inactive polling keys and schedule bounded cache eviction. - Preserve cached data while a key is active or inside its retention window. - Cancel stale cleanup timers when polling resumes and clear coordinator caches during disposal. - Add focused fake-timer coverage for retention, resubscription, and disposal behavior. ## Verification - `vitest --project @paperclipai/ui src/lib/cross-tab-poll.test.ts` — 11 tests passed. ## Risks - Low-to-moderate risk: eviction timing affects client polling coordination. - Tests cover the retention boundary, resumed subscriptions, and coordinator cleanup to reduce regression risk. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and code execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5618ea91f6 |
fix(ui): wrap company skill source paths (#9405)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company skills expose their source metadata in a narrow details sidebar > - Long filesystem paths and repository locators were truncated, hiding the part operators often need to distinguish sources > - The sidebar can preserve the complete value by wrapping at arbitrary path boundaries instead of ellipsizing it > - This pull request renders full source paths and repository labels without widening the layout > - The benefit is that operators can inspect and copy the actual skill source from the UI ## Linked Issues or Issue Description ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I can reproduce this on `master`. - [x] I have confirmed the behavior originates in Paperclip itself, not an agent adapter, API provider, or local configuration. ### What happened? Long company-skill source paths and repository locators were truncated in the skill details sidebar. ### Expected behavior The complete source value remains visible and wraps within the available sidebar width. ### Steps to reproduce 1. Open a company skill whose source path is longer than the details sidebar. 2. View the Source field. 3. Observe that the old UI replaces the middle or end of the value with an ellipsis. ### Paperclip version or commit `origin/master` at `02e2dd271`. ### Deployment mode Local dev (`pnpm dev`). ### Installation method Built from source (`pnpm dev` / `pnpm build`). ### Agent adapter(s) involved None; this is a company-skills UI layout issue. ### Logs, configuration, or screenshots Not applicable; the behavior is directly visible in the Source field. ### Additional context The narrow sidebar should remain width-constrained. Wrapping intentionally trades vertical space for full source inspectability. ## What Changed - Replace source-path truncation with width-constrained arbitrary wrapping. - Apply the same wrapping behavior to linked repository/source labels. - Add a regression test proving the full long path is rendered without ellipsis. ## Verification - `vitest --project @paperclipai/ui src/pages/CompanySkills.test.tsx` — 11 tests passed. - `node scripts/check-token-gates.mjs` — all token gates clean. ## Risks - Low risk: the change is limited to text layout in the company skill details view. - Very long unbroken values may make the Source section taller, intentionally trading vertical space for inspectability. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and code execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.711.0-canary.5 |
||
|
|
02e2dd271b |
feat(ui): add selection debug instrumentation (#9397)
Adds debug-gated selection instrumentation and focused tests for issue document annotations. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.711.0-canary.4 |
||
|
|
7fd321d622 |
feat(ui): use Lucide icons for task status glyphs (#9395)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Operators scan task state constantly, so the task **status** vocabulary (backlog / todo / in progress / in review / done / blocked / cancelled) has to read instantly > - Those statuses render through one shared component, `StatusGlyph`, whose icons were hand-rolled SVG geometry lifted from an internal spec > - Hand-rolled glyphs are harder to reason about, drift from the rest of the UI (which uses Lucide everywhere else), and mix fill/stroke styles across statuses > - This pull request swaps the hand-rolled geometry for named Lucide icons — one clean, consistent icon family — with no change to colours, sizing, or accessibility > - The benefit is a status icon set that is consistent with the rest of the app's iconography, trivially adjustable (change a mapping, not SVG path math), and simpler to maintain ## Linked Issues or Issue Description No existing public GitHub issue. Describing the change in-PR (feature/polish): **Problem / motivation.** The task status icons in `StatusGlyph` were bespoke inline SVGs (a half-filled disc for *in progress*, a filled disc + knockout check for *done*, ring+bar for *blocked*, ring+slash for *cancelled*, etc.). The rest of the UI uses [Lucide](https://lucide.dev) icons, so the status set was the odd one out — and its mixed fill/stroke shapes were harder to scan and to tweak. **Proposed solution.** Map each status to a Lucide icon and render that instead: | Status | Lucide icon | | --- | --- | | backlog | `circle-dashed` | | todo | `circle` | | in_progress | `rotate-cw` | | in_review | `circle-dot` | | done | `circle-check` | | blocked | `circle-minus` | | cancelled | `ban` | | in_queue (covered-blocked) | `circle-minus`, recoloured blue | Colours (the `--status-task-icon-*` tokens), the `sm/md/lg` size scale, `currentColor` recolouring, and the `role="img"` / `aria-label` behaviour are all unchanged — only the shapes change. **Alternatives considered.** Keeping the bespoke geometry (rejected: inconsistent with the app and harder to maintain). **Related PRs** (linked for reviewer context, not dependencies): - Refs #8580 — the merged PR that established the current hand-rolled status glyphs this PR restyles. - Refs #8838 — open PR forwarding Radix trigger props through `StatusGlyph`; touches the same component (no overlap with this change). - Refs #1760 — open proposal to redesign the *cancelled* status icon specifically; this PR moves cancelled to Lucide `ban`. ## What Changed - `ui/src/components/StatusGlyph.tsx`: replaced the per-status hand-rolled SVG `glyphBody()` geometry with a `status → Lucide icon` map (`circle-dashed`, `circle`, `rotate-cw`, `circle-dot`, `circle-check`, `circle-minus`, `ban`). Kept the token-driven colour wiring, size scale, `currentColor` recolouring, a11y label handling, and the `in_queue` = blocked-icon-recoloured-blue behaviour. - `ui/src/components/StatusGlyph.test.tsx`: updated to lock the new icon mapping (per-status Lucide class, size scale, colour var, `in_queue`, a11y) instead of the old geometry. Net: two files, +74 / −138 (the component got smaller). Because every status surface (list, board, detail header, status picker, sub-task/blocked-by pills, chips) routes through `StatusGlyph`, this single-component edit covers them all. ## Verification - `pnpm check:token-gates` → **3/3 clean** (no hardcoded colour/spacing/font values introduced). - `pnpm typecheck` → clean across all packages. - `cd ui && pnpm vitest run` → **2509/2509 passing**, including the updated `StatusGlyph` test. - Manual: ran the worktree dev server and confirmed the new icons render everywhere (task list, task detail, related-task chips, and the status picker showing all seven). **Storybook visual-regression note:** this is an intentional visual change, so the status-icon stories will diff against the published baseline. The baseline snapshots need to be regenerated and republished by a maintainer (`pnpm test:storybook-visual:update` from a trusted environment) as part of accepting this change — the visual-regression CI check is expected to be red until then. No baseline is published in the environment this PR was authored in, so that step is left to a maintainer. ## Risks - **Low risk / cosmetic.** No logic, data, or API changes — only the rendered icon shapes. Colours, sizes, and accessibility labels are unchanged. - The most noticeable shifts are *in progress* (half-disc → rotating arrow), *done* (solid disc+check → outline circle+check), and *cancelled* (ring+slash → ban). These are deliberate. - The only CI check expected to fail is the Storybook visual-regression job, pending a maintainer baseline update (see Verification). ## Model Used Claude Opus 4.8 (`claude-opus-4-8`), run in Claude Code with extended thinking and tool use (file edits, local test runs, browser-driven visual verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.711.0-canary.3 |
||
|
|
e0f1905222 |
[codex] Quiet packaged version fallback diagnostics (#9207)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server startup path reports the product version from package metadata and, in source checkouts, Git metadata. > - Packaged installs can run from `node_modules`, where Git metadata is normally unavailable and that absence is expected. > - The fallback path was still attempting Git metadata probing in packaged contexts, which could print scary diagnostic noise during onboarding even though the package version fallback was working. > - This pull request makes the packaged path skip Git probing only when the package does not look like a source checkout, and keeps fallback diagnostics opt-in. > - The benefit is a quieter first-run experience without weakening source-checkout version detection or debug diagnostics. ## Linked Issues or Issue Description No public GitHub issue exists. ### Bug Report #### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest released version of Paperclip or can reproduce on `master`. - [x] I have confirmed the error originates in Paperclip itself, not in an agent adapter, API provider, or local configuration. #### What happened? When Paperclip starts from a packaged install, server version resolution can fall back from Git metadata to package metadata. That expected fallback path could emit scary Git diagnostic noise during onboarding even though startup could continue normally. #### Expected behavior Packaged Paperclip startup should use package metadata quietly when Git metadata is unavailable. Source checkouts should still use Git-derived versions, and operators who explicitly opt into version-resolution diagnostics should still receive useful Git failure details. #### Steps to reproduce 1. Run Paperclip from a packaged install where the server package is under `node_modules` and does not include package-local Git metadata. 2. Start the server in an environment where `git describe` cannot resolve repository metadata for that package. 3. Observe that version fallback can produce Git diagnostic noise during startup even though the package version fallback is expected. #### Paperclip version or commit Reproduced against the pre-fix server version resolution behavior on `master`-derived builds. #### Deployment mode Self-hosted server / packaged local install. #### Installation method npm / pnpm package install. #### Agent adapter(s) involved Not adapter-specific; this is core server startup/version behavior. #### Database mode Not database-related. #### Access context Unclear / not applicable. #### Relevant logs or output Git fallback diagnostics from `git describe` could appear during packaged startup. The exact path and Git output depend on the operator environment. #### Additional context The fix keeps diagnostics available behind `PAPERCLIP_DEBUG_VERSION_RESOLUTION=1` and preserves source-checkout Git version detection, including source paths that happen to contain a `node_modules` segment. #### Privacy checklist - [x] I have reviewed all pasted output for PII and redacted where necessary. ## What Changed - Skip Git metadata probing for packaged installs under `node_modules` only when no package-local Git metadata is present. - Preserve Git-derived version detection for source or linked workspace checkouts, even when their path contains a `node_modules` segment. - Keep fallback diagnostics behind the existing debug/diagnostic opt-in path. - Include useful Git failure details such as stderr/stdout/stack/cause when diagnostics are enabled. - Add version tests covering packaged fallback behavior, source-checkout detection, richer diagnostics, and quiet default output. ## Verification - `pnpm vitest run server/src/__tests__/version.test.ts` passed after the Greptile follow-up changes. - `pnpm --filter @paperclipai/server typecheck` passed. - `git diff --check` passed. - Greptile completed with confidence score 5/5 and no blocking issues on the latest reviewed commit. ## Risks Low risk. The change is scoped to version fallback behavior. Source-checkout Git version detection remains covered, while packaged `node_modules` contexts intentionally rely on package metadata instead of Git probing. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent with shell/tool use in the Paperclip workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.711.0-canary.2 |
||
|
|
49d1abc458 |
fix(ui): detail the env unsaved-changes banner and guard against draft loss (#9391)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instance settings include an Environments section where operators configure execution environments, each with an environment-variables editor for run-time bindings > - The editor showed a bare "Unsaved changes" banner that never said which variables changed, sometimes appeared the moment a saved config was opened (a lossy round-trip through the editor's emit rules made clean values look dirty), and the environment form let you navigate away without any confirmation, silently dropping the draft > - Operators could not tell what was unsaved, distrusted the phantom banner, and lost half-finished environment edits to a stray click — the agent configuration page already confirms before discarding, so environments behaved inconsistently > - This pull request lists the new/edited/removed variable names under the banner, normalizes both sides of the dirty comparison so saved values no longer look dirty on open, and confirms before cancel, in-app navigation, or tab unload while the form has unsaved changes > - The benefit is that the banner is trustworthy and specific, and unsaved environment edits can no longer be lost without an explicit confirmation ## Linked Issues or Issue Description Related (not fixed by this PR): #8930 introduced the current environment-variables editor and its unsaved-changes banner; #9386 moved environment create/edit from a modal to routed pages, which this PR's navigation guard builds on. No existing public issue for the defects themselves; described per the bug report template: **What happened?** The environment-variables editor in Environments settings showed a bare "Unsaved changes" banner with no indication of which variables changed. For some saved configurations (names with surrounding whitespace, incomplete secret references, duplicate names differing only by whitespace) the banner appeared immediately on opening the edit form, before any user input. Navigating away from the environment form — cancel, an in-app link, or closing the tab — silently discarded the draft with no confirmation. **Expected behavior** The banner should say which variables are new, edited, or removed; a freshly opened saved configuration should show no banner; and leaving the form with unsaved changes should require an explicit confirmation, consistent with the agent configuration page. **Steps to reproduce** 1. Open Settings → Instance settings → Environments and edit an environment whose saved config round-trips lossily (e.g. an env var name stored with trailing whitespace) — the "Unsaved changes" banner appears with no user edits. 2. Add or edit a variable — the banner gives no hint of what is unsaved. 3. With a dirty draft, click any in-app link or Cancel — the draft is dropped with no confirmation. **Deployment mode** Self-hosted (local development instance), reproducible on `master`. ## What Changed - The unsaved-changes banner in `EnvironmentVariablesEditor` now renders a change summary line — `New: … · Edited: … · Removed: …` — showing up to three names per group with a `+N more` overflow and the full list in a `title` tooltip. A rename shows as one addition plus one removal. - Dirty detection normalizes both the committed value and the draft through the same rules the editor uses when emitting values (trimmed names, incomplete secret refs dropped, last-writer-wins on trimmed duplicates), so a saved config that round-trips lossily no longer shows a phantom banner on first open. - The editor exposes an `onDirtyChange` callback and warns via `beforeunload` while its local draft is dirty. - The environment create/edit page (`CompanyEnvironments`) tracks a payload-level baseline fingerprint of the form as initialized and treats the page as having unsaved changes when the current form differs from it or the editor draft is dirty. While dirty it confirms ("Discard unsaved environment changes?") on Cancel, intercepts same-origin in-app link clicks, and warns on tab unload. ## Verification - `node_modules/.bin/vitest run ui/src/pages/CompanyEnvironments.test.tsx ui/src/components/environment-variables-editor/EnvironmentVariablesEditor.test.tsx` — 48 tests pass, including new coverage for: the change-summary banner text, no phantom banner for lossy round-trip values, beforeunload only while dirty, cancel confirmation on the edit page, and unload/link-click warnings after edits are staged into the form. - `tsc -b` in `ui/` passes. - Manual: edit an environment, add/edit/remove variables, observe the summary line; click Cancel or an in-app link and observe the confirmation; save and observe navigation proceeds without prompting. ## Risks - Low risk, UI-only. The click interceptor is scoped to same-origin anchor navigation while the environment form page has unsaved changes and is removed on cleanup; modified-key/middle-button clicks and external links are left alone. - The dirty-normalization intentionally ignores differences the editor could never persist (incomplete secret refs, untrimmed duplicate names); those were previously reported as unsaved changes that could not be saved away. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code, with extended thinking and tool use (code editing, test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.711.0-canary.1 |
||
|
|
9cde4e128c |
feat: run ACP sessions in sandbox execution targets (#9390)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters (Claude, Codex, Gemini) default to the ACP engine lane, which needs a live bidirectional stdio session with the agent process > - Sandbox execution targets only exposed one-shot command execution, so every ACP-capable adapter refused remote targets and fell back to the CLI lane with a "supports only the local Paperclip host" warning > - Running agents in sandboxes is a core deployment mode, and losing ACP there means losing streaming updates, structured events, and default-lane parity with local runs > - This pull request adds a provider-agnostic process-session bridge that relays the ACP stdio session into the sandbox over the existing sandbox runner contract, and updates the adapters to use it > - The benefit is that the default ACP lane now behaves the same on the local host and in any sandbox provider, with CLI fallback reserved for targets that genuinely cannot host a bidirectional session ## Linked Issues or Issue Description No existing public issue covers this; inline description following the feature request template: **Problem or motivation** Configuring an ACP-capable adapter (e.g. Claude) with a sandbox environment made every run fall back to the CLI lane with the warning "Claude ACP currently supports only the local Paperclip host, but this run targets a remote environment." The ACP engine only knew how to spawn a local subprocess, while sandbox providers only expose one-shot command execution — so there was no way to hold the bidirectional stdio session ACP requires. **Proposed solution** Add a process-session bridge in `adapter-utils`: a local ACPX-spawnable proxy script connects to a token-authenticated loopback TCP server, which relays JSON-framed stdin/stdout/stderr events to and from a small relay script executed inside the sandbox via the provider's ordinary runner. Claude/Codex/Gemini adapters now treat sandbox targets with a runner as ACP-capable, resolve agent commands against the remote target, and fall back to CLI only when the sandbox exposes no bidirectional path. The sandbox callback bridge injects a run-scoped API endpoint and bridge token so the agent inside the sandbox can reach Paperclip (including work-product handoffs) without ever receiving the host run JWT. **Alternatives considered** A provider-specific lane was prototyped first: Daytona minting SSH access metadata at lease time, converted into an SSH execution target. It was dropped because it only worked for providers able to advertise SSH, added per-provider surface area, and left every other sandbox provider on the CLI fallback. The merged design rides the one-shot runner contract all providers already implement; a regression test pins that sandbox targets stay on the bridge lane even when lease metadata advertises SSH access. **Roadmap alignment** Directly advances the "Cloud / Sandbox agents" roadmap item — agents running in remote and sandboxed environments keep the same control-plane behavior as local ones. No overlap with other planned core work. ## What Changed - `packages/adapter-utils/src/execution-target.ts`: new `startAdapterExecutionTargetProcessSessionBridge()` plus helpers — writes a token-authenticated local proxy script (spawnable by ACPX) and a remote relay script synced into the sandbox, with a loopback TCP server streaming JSON-framed stdio between them; events emitted before the ACP client attaches are buffered so none are lost. - `packages/adapter-utils/src/acpx-engine/execute.ts`: the ACP engine can execute against remote sandbox targets through the bridge instead of requiring a local subprocess, including remote cwd/env shaping. - `packages/adapter-utils/src/sandbox-callback-bridge.ts`: sandbox-scoped API bridging extended to allow work-product handoffs; the sandbox payload env carries a bridge token, never the host run JWT. - `packages/adapters/claude-local`, `codex-local`, `gemini-local` (`src/server/acp.ts`): default-lane selection no longer rejects all remote targets; command resolution is remote-aware (`ensureAdapterExecutionTargetCommandResolvable`, `resolveAdapterExecutionTargetCwd`); the fallback reason is now scoped to sandboxes that expose only one-shot execution. - `server/src/__tests__/environment-execution-target.test.ts`: pins that sandbox targets resolve to the bridge lane, including when lease metadata advertises SSH access. - Non-sandbox remote targets (e.g. SSH) keep the CLI lane: the ACP engine's remote transport is sandbox-only, so default-lane selection falls back for those targets across all three adapters, and tests covering CLI-specific remote behavior pin `engine: "cli"` explicitly. - The bridge authenticates loopback connections before they can own the session or receive buffered output (token required, idle unauthenticated peers dropped), and remote event writes are serialized so the exit event always lands after stdout/stderr have drained. - Daytona plugin: formatting-only residue from the earlier iteration; no functional change. ## Verification - `vitest run` over the touched suites — `packages/adapter-utils/src/acpx-engine/execute.test.ts`, `packages/adapter-utils/src/execution-target-sandbox.test.ts`, `packages/adapter-utils/src/sandbox-callback-bridge.test.ts`, the three adapter `acp.test.ts` files, and `server/src/__tests__/environment-execution-target.test.ts` — 102 tests pass. - End to end: with a Claude agent configured on a Daytona sandbox environment, the primary-model test now selects the default ACP lane (no fallback warning), and the full round trip (wake → sandbox execution → API bridge → comment post) was exercised twice from inside a live sandbox. ## Risks - Behavioral shift: adapters that previously always fell back to CLI on sandbox targets now default to ACP there; `engine=cli` still pins the CLI lane explicitly. - The bridge relays stdio as JSON lines over loopback TCP guarded by a per-session random token; the remote relay runs inside the sandbox under the provider's runner. Providers with slow one-shot execution will see higher session startup latency — the CLI fallback remains for genuinely incapable targets. - No schema or migration changes. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) — extended thinking enabled, agentic tool use via the Claude Agent SDK harness; implementation iterated with local Vitest verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no shipped docs describe the old local-only ACP limitation) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.local>canary/v2026.711.0-canary.0 |
||
|
|
0f9b1d399c |
fix(ui): make environment edit a routed page (#9386)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instance settings include an Environments section where operators configure sandbox/SSH/local execution environments, including interactive custom-image setup sessions with a browser terminal > - The environment create/edit form was rendered inside a modal dialog, so pressing Escape anywhere — including inside the embedded SSH terminal while capturing a snapshot — closed the whole modal and destroyed the in-progress session > - Environment editing is a heavyweight, long-lived flow; losing it to a reflexive Escape keypress is destructive and surprising > - This pull request converts environment create/edit from a modal into routed standalone pages, so Escape no longer dismisses the form > - The benefit is that terminal sessions and half-completed edits survive Escape, and the flow gets shareable URLs and normal back/forward navigation ## Linked Issues or Issue Description No existing public issue; described per the bug report template: **What happened?** While editing an environment's sandbox snapshot in the embedded SSH terminal, pressing Escape (e.g. to exit a mode inside the terminal) closed the entire environment edit modal, discarding the setup session and any unsaved form state. **Expected behavior** Escape inside the terminal or form should not dismiss the environment editor. A heavyweight flow like environment configuration should be a standalone page where Escape behaves as expected within the focused widget. **Steps to reproduce** 1. Open Instance settings → Environments and edit a sandbox environment 2. Start a custom image setup session and focus the browser terminal 3. Press Escape 4. The modal closes and the session context is lost ## What Changed - Converted the environment create/edit dialog in `CompanyEnvironments.tsx` into routed pages at `/company/settings/instance/environments/new` and `/company/settings/instance/environments/:environmentId/edit` - Registered the new routes in `App.tsx` and wired breadcrumbs for the list/create/edit states - Form state now initializes from the route (create vs edit) instead of dialog open/close state, and successful saves navigate back to the environments list - Updated `CompanyEnvironments.test.tsx` and `CompanySettings.test.tsx` to render through a router with the new routes and assert against the routed form page instead of a dialog ## Verification - `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx ui/src/pages/CompanySettings.test.tsx` — 22/22 passing - `tsc --noEmit` on the `ui` package — clean - Behavioral coverage: the updated tests exercise the routed create/edit pages end to end (open edit via the list, interact with the setup-session controls on the form page, save navigates back to the list); with the form no longer in a dialog there is no Escape-close handler to trigger ## Risks - Low risk; UI-only routing change. Deep links into the old modal state do not exist (the modal had no URL), so no redirects are needed - The edit page resolves the environment from the route param; a stale/unknown id falls back to the environments list ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Fable 5), extended thinking enabled, agentic tool use via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Cody <cody@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.17 |
||
|
|
b15115e05b |
fix(ui): keep rendered markdown list markers visible (#9359)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The board UI renders agent/user-authored Markdown in task descriptions, comments, and other work-thread surfaces > - Those Markdown surfaces often live inside cards and containers that constrain overflow > - Ordered-list markers are painted outside the list content box, so too little inline padding can clip multi-digit markers at the left edge > - This pull request keeps the shared Markdown list gutter compact while giving ordered lists enough marker space for two- and three-digit counters > - The benefit is that long numbered lists in board-facing Markdown render correctly without widening unordered-list gutters or changing API, data, or editor behavior ## Linked Issues or Issue Description No public GitHub issue found. Bug description: - What happened: rendered Markdown ordered lists with multi-digit items could show clipped marker digits when the list was flush against an overflow-constrained container. - Expected behavior: ordered-list markers such as `10.` and `100.` should render fully in task descriptions and comments. - Steps to reproduce: render a `.paperclip-markdown` ordered list with at least 100 items inside a container that clips overflow and has no extra left gutter. - Paperclip version/commit: current `master` before this PR. - Deployment mode: board UI, deployment-mode independent. Related search result: - Refs #2049 because it also touches rendered Markdown list presentation, but it styles GFM task-list checkboxes and does not address ordered-list marker clipping. ## What Changed - Set the shared `.paperclip-markdown` list padding to a compact `1.5rem` baseline for bullets and lists. - Added an ordered-list-only `2.5rem` padding override so outside-positioned multi-digit ordered-list markers have enough inline-start room. - Added a focused stylesheet regression test that verifies unordered-list gutters stay compact while ordered lists keep the larger marker gutter. - Restored the exact maintainer-skill marker phrase expected by the existing server skill utility contract test, fixing an unrelated latest-head CI failure from current `master`. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/MarkdownListStyles.test.ts` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui build` - `pnpm exec vitest run server/src/__tests__/paperclip-skill-utils.test.ts` ## Risks - Low risk: ordered lists in rendered Markdown get a larger left gutter; unordered lists keep a smaller shared gutter. - Low risk: the skill-doc marker change is text-only and matches the existing server test contract. - No database, API, migration, auth, adapter, or telemetry changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-enabled software-engineering session. Context window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.16 |
||
|
|
36ec79c196 |
feat: add attention queue and Decisions surface (#9380)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies, where operators need a reliable way to find and act on work awaiting their input. > - The attention and issue-thread interaction subsystems expose those decision points across server APIs and the board UI. > - The previous navigation and interaction presentation left these actions fragmented and did not offer a controlled rollout for the Decisions surface. > - This branch adds the attention feed, richer interaction cards, grouping, dismiss/snooze behavior, and a gated Decisions sidebar entry. > - It also keeps experimental settings and API contracts synchronized, with an idempotent migration for the new dismissal state. > - This pull request delivers the complete, tested attention/Decisions experience as one reviewable unit. ## Linked Issues or Issue Description - Adds an operator-focused attention queue and Decisions experience: grouped decision cards, semantic interaction actions, dismiss/snooze handling, resilient interaction states, and an experimental flag to control the Decisions navigation entry. ## Feature Context ### Problem or Motivation Operators currently have to hunt across approvals, interactions, failed runs, and budget alerts to find decisions that need their action. ### Proposed Solution Provide a gated Decisions attention queue that groups actionable items, supports direct resolution, and preserves operator control through dismiss and snooze actions. ### Alternatives Considered Keep separate, source-specific views only; this leaves cross-cutting operator decisions fragmented and harder to prioritize. ### Roadmap Alignment This improves the V1 control-plane operator workflow by making pending governed actions discoverable in one company-scoped surface. ## What Changed - Added server attention-feed services, routes, interaction handling, dismiss/snooze support, and an idempotent `0145` inbox-dismissal migration. - Added shared attention, inbox-dismissal, and experimental-settings contracts. - Added Decisions/attention UI, interaction-card states, sidebar badge/navigation integration, grouping, keyboard support, and Storybook coverage. - Added tests for attention behavior, thread interactions, settings normalization, dismissals, and API behavior. - Removed generated screenshots from the final PR diff and rebased the branch onto current `master`. ## Verification - `pnpm check:token-gates` — passed. - `pnpm exec vitest run packages/shared/src/issue-thread-interactions.test.ts server/src/__tests__/attention-service.test.ts server/src/__tests__/inbox-dismissals.test.ts server/src/__tests__/issue-thread-interactions-service.test.ts server/src/__tests__/issue-thread-interaction-routes.test.ts ui/src/lib/attention.test.ts ui/src/components/AttentionQueueRow.test.tsx ui/src/components/IssueThreadInteractionCard.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx` — passed: 158 tests across 9 focused files. - GitHub Actions for `ad636f560`: build and typecheck/release-registry have passed; remaining general-server and Greptile checks are in progress. ## Risks - Moderate: this is a cross-layer attention/interaction feature with a new migration and navigation behavior. - The `enableDecisions` experimental setting defaults to off, limiting rollout impact. - Existing dismissal data is backfilled to `dismiss`; the migration is idempotent and uses guarded constraint creation. > ROADMAP.md was checked; no duplicate planned core feature was identified. Related open pull requests were searched before opening this PR. ## Model Used - OpenAI GPT-5.5 via Codex CLI, with tool use and local code execution. Context-window size unavailable in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public PR branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally; focused tests pass and the remaining unrelated AWS test failure is documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.710.0-canary.15 |
||
|
|
ac66fd65cb |
Fix Cody default model adapter test config (#9365)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent configuration includes adapter-specific model settings and a
built-in adapter test action so operators can verify runtime
configuration before saving changes.
> - Cody/Codex-style local adapters can use an adapter default model
when the user clears the explicit model field.
> - The adapter test path still passed an object containing `model:
undefined` in some create/edit flows, which is different from omitting
the model and can break default-model behavior.
> - The previous fix was reverted because it also included an unrelated
skill documentation edit.
> - This pull request reapplies only the UI default-model test-config
fix, with no doc or skill changes.
> - The benefit is that testing Cody/Codex adapter settings with the
default model follows the same contract as saving default model
settings: no explicit model key is sent.
## Linked Issues or Issue Description
Bug report:
- Summary: Testing a Cody/Codex local agent after selecting the default
model could send an adapter config with an undefined model value instead
of omitting the model key.
- Expected behavior: Clearing the model to use the adapter default
should test with `adapterConfig: {}` unless another model is explicitly
selected.
- Actual behavior: The UI test-config path could preserve `model:
undefined`, causing the adapter test to fail instead of exercising the
default model.
- Related PRs: Reapplies the UI-only portion of #9361 after #9363
reverted the original PR.
## What Changed
- Exported and reused `omitUndefinedEntries` so adapter test config
payloads drop undefined adapter config entries before calling the test
endpoint.
- Hardened the current model display value so create-mode values that
are nullish or non-string do not crash the model selector/test flow.
- Added render coverage for editing a Codex agent back to the default
model and for testing a create form with the default model.
## Verification
- `pnpm exec vitest run
ui/src/components/AgentConfigForm.render.test.tsx`
- `pnpm check:token-gates`
- Confirmed `git diff origin/master --name-only` contains only:
- `ui/src/components/AgentConfigForm.render.test.tsx`
- `ui/src/components/AgentConfigForm.tsx`
- `ui/src/lib/agent-config-patch.ts`
## Risks
Low risk. The change only removes `undefined` adapter config entries
from the UI adapter-test payload and adds focused render coverage.
Explicit model values and other adapter config fields are preserved.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
OpenAI Codex, GPT-5 coding agent, tool-use enabled. Context window size
not exposed in this runtime.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.710.0-canary.14
|
||
|
|
d1f6a6850a |
Fix agent detail URL after agent rename (#9340)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The agent detail page uses route references that can be based on an agent's URL key. > - Renaming an agent can change that URL key while the browser is still on the old route. > - After save or rollback, refetching the stale route reference can render an "Agent not found" state even though the agent still exists. > - This pull request redirects the detail page to the updated canonical route when the saved agent's route reference changes. > - The benefit is that agent renames keep users on the same configuration workflow without landing on a stale URL. ## Linked Issues or Issue Description - Refs #1848 - Related public search performed for agent rename/not-found issues and PRs; no closer in-flight PR was found. - Bug context: after saving a renamed agent or rolling back to a revision with a different name-derived URL key, the agent detail page could continue using the old URL and show "Agent not found". ## What Changed - Added a small route-sync helper that compares the previous and updated agent route refs after mutations. - Redirects the agent detail page with `replace: true` when a save or rollback changes the canonical route ref. - Removes the stale detail-query cache entry so the old route reference is not refetched after a rename. ## Verification - Local outgoing patch scan for common secrets, private paths/emails, and internal issue/link references: no matches. - `corepack pnpm install --frozen-lockfile` - `corepack pnpm --dir ui run typecheck` - `corepack pnpm --dir ui exec vitest run src/pages/AgentDetail.progress.test.ts src/App.test.tsx` - `corepack pnpm check:token-gates` ## Risks - Low risk: the redirect only runs when the updated agent resolves to a different route ref than the current agent. - If a future mutation response omits both URL key and name, the existing route-ref fallback behavior still applies. ## Model Used OpenAI Codex, GPT-5 coding agent via the local Codex adapter, with tool-assisted repository inspection, shell execution, and GitHub API use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
17dde9d3f2 |
fix(sandbox): keep custom-image snapshots applied to config tests, probes, and saves (#9385)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox environments can capture reusable custom images (provider snapshots) so agents boot with pre-installed tools and CLI logins > - The custom-image runtime fingerprint check included provider secret-ref paths (e.g. the Daytona `apiKey`), while capture-time fingerprinting excluded them, so any config carrying a credential never matched its captured snapshot > - As a result, agent config tests and environment probes silently booted the provider base image instead of the snapshot, test sandboxes were deleted before operators could inspect them, and any environment save orphaned the snapshot without warning > - The UI compounded the confusion by displaying an internal template id that matches nothing in the provider dashboard > - This pull request aligns runtime fingerprints with capture-time exclusions, re-stamps fingerprints on saves that cannot affect the snapshot (warning when they can), archives test/probe sandboxes instead of deleting them, and surfaces the provider snapshot ref in the UI > - The benefit is that custom images actually apply to config tests and probes, survive unrelated config edits, and are debuggable against the provider dashboard ## Linked Issues or Issue Description No public GitHub issue exists for this; describing it in-PR per the bug template. Related: Refs #9329 (saved-environment probe company context — this branch carries an equivalent fix), Refs #8794 (introduced reusable sandbox custom images). **What happened?** With a Daytona environment whose provider config stores the API key as a secret reference and an active captured custom-image snapshot: - Agent config tests and environment probes booted the provider base image (`daytonaio/sandbox:0.8.0`) instead of the captured snapshot, so CLI upgrades/logins baked into the snapshot were missing and the probe reported "login required" and an outdated CLI. - The environment card showed an internal template id (e.g. `b5be03e1-ca5…`) that does not correspond to any snapshot name in the provider dashboard, making the active image impossible to correlate. - Test/probe sandboxes were deleted immediately after the run, so the sandbox a test used could not be inspected afterwards. - Saving the environment config (even fields unrelated to the image) changed the stored fingerprint, silently detaching the snapshot with no warning. **Expected behavior** Config tests and probes boot the captured snapshot when one is active; the UI shows the provider-facing snapshot/template ref; test sandboxes stay inspectable for a short window; unrelated config edits keep the snapshot linked, and edits that genuinely invalidate it produce an explicit warning. **Steps to reproduce** 1. Configure a sandbox environment on Daytona with the API key stored as a company secret reference. 2. Capture a custom image snapshot from the environment page and mark it active (e.g. after installing/logging into a CLI in the setup sandbox). 3. Run the agent config test or an environment probe: the sandbox boots the base image, not the snapshot, and the sandbox is deleted immediately after the test. 4. Save the environment config with an unrelated field change: the snapshot silently stops applying. **Paperclip version or commit** `master` at the merge-base of this branch. **Deployment mode** Self-hosted local instance (macOS, pnpm dev server) with the Daytona sandbox provider plugin. ## What Changed - Runtime custom-image fingerprint checks now exclude provider secret-ref paths, matching capture-time exclusions, so configs carrying credentials match their captured snapshots (`environment-custom-image-runtime.ts`). - Agent config tests and saved-environment probes force fresh, non-reused sandboxes and pass company context so lease-backed probes can resolve company secrets and boot the real snapshot (`environment-probe.ts`, `routes/agents.ts`, `routes/environments.ts`). - Test/probe sandboxes are released by archiving (stop + 60-minute provider-side auto-delete) instead of immediate deletion, so operators can inspect the exact sandbox a test used (Daytona plugin). - On environment PATCH save, changes that cannot affect the captured snapshot re-stamp the template's source fingerprint so the snapshot stays linked; boot-source or provider-identity changes (new manifest field `templateIdentityPaths`) mark the template detached and the save response reports it (`environment-custom-images.ts`, shared plugin types/validators). - The custom-image overview exposes `activeTemplateMatchesConfig`; the environments UI shows the provider snapshot/template ref (internal id moved to a tooltip), warns via toast when a save detaches the snapshot, and shows a persistent "Not in use" warning when the active template no longer matches the saved config (`CompanyEnvironments.tsx`, `api/environments.ts`). ## Verification - `pnpm vitest run server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/environment-probe.test.ts server/src/__tests__/environment-routes.test.ts server/src/__tests__/agent-test-environment-routes.test.ts` — server coverage for fingerprint exclusions, re-stamp/detach on save, probe company context, and fresh-sandbox test behavior. - `pnpm vitest run packages/plugins/sandbox-providers/daytona/src/plugin.test.ts` — archive-on-release and snapshot ref handling. - `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx` — snapshot ref display, detach toast, and "Not in use" warning. - Manually verified end-to-end on a live self-hosted instance against real Daytona: config test boots the captured snapshot (CLI login and version persist), the test sandbox remains visible in the provider dashboard as archived, and saving unrelated fields keeps the snapshot applied. ## Risks - Fingerprint exclusion widening: a provider credential rotation alone no longer detaches a captured snapshot; that is the intended behavior (the snapshot content does not depend on the credential), and provider-identity fields (e.g. Daytona `apiUrl`) still detach via `templateIdentityPaths`. - Archived test sandboxes consume provider-side resources for up to their auto-delete window instead of being freed immediately; bounded (60 minutes) and only for test/probe sandboxes. - New optional manifest field `templateIdentityPaths` is backward-compatible; providers that omit it keep current matching behavior. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.710.0-canary.13 |
||
|
|
70ce005bef |
Ensure worktree execution starts only after activation (#9374)
## Thinking Path > - Paperclip is the open-source control plane people use to manage AI agents and their work. > - Its scheduler, routines, and heartbeat services decide when agents automatically begin work. > - Experimental per-worktree execution is useful for isolated development, but enabling it previously allowed automatic services to consider an existing backlog. > - A worktree activation must therefore create a durable eligibility boundary rather than merely toggle execution on. > - This pull request records an activation cutoff and applies it consistently to automatic routine and heartbeat dispatch. > - The result is that an enabled worktree executes only work created after its own activation, while non-worktree behavior remains unchanged. ## Linked Issues or Issue Description **Problem type:** Bug / safety regression **Summary:** Enabling experimental run execution in an existing worktree could start automatic scheduler, routine, watchdog, and heartbeat activity for work created before that worktree was explicitly armed. **Expected behavior:** A worktree that has execution enabled only considers automatically dispatched work created on or after its activation timestamp. Ambiguous activation state fails closed. Non-worktree instances keep their existing behavior. **Related public work:** Refs #8275 (runtime worktree policy gating); this PR adds an activation-time boundary for automatic execution rather than changing the general runtime policy. ## What Changed - Persist a worktree execution activation timestamp and originating instance ID; stamp them only when the experimental toggle changes from disabled to enabled. - Resolve activation state fail-closed when the cutoff is missing, invalid, disabled, or belongs to another instance. - Gate automatic routine scheduling, webhooks, watchdog activity, and heartbeat selection at the activation cutoff; manual runs remain available. - Share the canonical worktree truthy-environment helper across routine dispatch and agent inbox filtering. - Add cutoff and truthy-runtime regression coverage, plus experimental-settings UI states that explain armed and suppressed execution. ## Verification - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts server/src/__tests__/instance-settings-service.test.ts` — passes: 2 files, 60 tests. - `pnpm --filter @paperclipai/server typecheck` — passes. - Existing CI completed successfully before the follow-up review fixes; this branch was rebased onto the latest `origin/master` before retesting. ## Risks - **Behavioral:** Automatic worktree execution is intentionally more restrictive; pre-existing work is suppressed until newly created after activation. - **Operational:** A malformed or cross-instance activation record fails closed, requiring an operator to disable and re-enable the experimental toggle on the intended worktree. - **Compatibility:** The worktree environment now accepts all canonical truthy values (`1`, `true`, `yes`, and `on`) consistently; non-worktree instances are unaffected. - **Branch metadata:** This existing execution-workspace branch predates the current naming rule and cannot be renamed under this task's workspace contract; the code and PR title do not include internal ticket references. > `ROADMAP.md` was checked; this targeted execution-safety fix does not duplicate planned core work. ## Model Used - Anthropic Claude Code — assisted with the original implementation; exact model identifier and context window were not recorded in the repository metadata. - OpenAI Codex CLI — assisted with PR preparation and review fixes; exact model identifier and context window are not exposed in this execution environment. Used with terminal tooling, code editing, and targeted test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.710.0-canary.12 |
||
|
|
23f34491e2 |
Fix apiCompression corrupting and dropping Better Auth responses for gzip clients (#9381)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its server fronts every API route — including Better Auth sign-in — with Express middleware, and #9190 added an `apiCompression` middleware that gzips JSON responses over 1KB > - That middleware buffers `res.write()` chunks with `String(chunk)`, but Better Auth (via better-call) streams `Uint8Array` chunks and commits headers with `writeHead()` before streaming > - `String(Uint8Array)` serializes the body to comma-separated decimal bytes (~3.4x inflation), and once the inflated body crossed the 1KB threshold, `setHeader()` threw `ERR_HTTP_HEADERS_SENT` and the catch handler destroyed the socket > - Every real browser sends `Accept-Encoding: gzip`, so sign-in returned zero bytes (`net::ERR_EMPTY_RESPONSE` / "Failed to fetch"), while curl without `Accept-Encoding` worked — making the bug easy to misdiagnose as a client or network issue > - This pull request makes the middleware byte-safe for `Uint8Array` chunks, passes through responses whose headers are already committed, and falls back to the uncompressed body instead of destroying the connection when compression fails > - The benefit is that browser sign-in (and any other streamed binary-chunk response) works again for gzip-accepting clients, with regression tests locking in all three behaviors ## Linked Issues or Issue Description Refs #9190 (introduced the `apiCompression` middleware). No public GitHub issue exists; bug description: - **What happened:** Sign-in from any real browser failed with `net::ERR_EMPTY_RESPONSE` / "Failed to fetch". The server logged `ERR_HTTP_HEADERS_SENT` from the compression middleware and destroyed the response socket, so zero bytes reached the client. - **Expected:** `/api/auth/*` responses are delivered intact regardless of the client's `Accept-Encoding`. - **Steps to reproduce:** Run the server with API compression active, open the web UI in a browser (which sends `Accept-Encoding: gzip`), and attempt email/password sign-in. The auth response body exceeds ~300 bytes, so after the ~3.4x stringification inflation it crosses the 1024-byte compression threshold and the response is destroyed. `curl` without `Accept-Encoding` succeeds against the same server. - **Scope:** Any route that streams `Uint8Array` chunks and/or commits headers via `writeHead()` before writing — in practice all Better Auth routes served through better-call. ## What Changed - `server/src/middleware/api-compression.ts`: - Buffer `res.write()` chunks with a `toBodyBuffer()` helper that converts `Uint8Array`/`ArrayBuffer` views via `Buffer.from()` instead of `String()`, so binary chunks are preserved byte-for-byte. - Pass responses through untouched once headers are already sent (`writeHead()`-style streaming), since compression headers can no longer be set at that point. - On any compression failure, write the original uncompressed body instead of calling `res.destroy()`, so clients get a valid (just uncompressed) response rather than a dropped connection. - `server/src/__tests__/api-compression.test.ts`: three new regression tests — small `writeHead`+`Uint8Array` responses are delivered byte-for-byte, large ones no longer drop the connection, and `Uint8Array` JSON bodies gzip without corruption (includes `/api/auth-bridge` and `/api/uint8-json` test routes mirroring better-call's streaming pattern). ## Verification - `cd server && pnpm vitest run src/__tests__/api-compression.test.ts` — 10/10 passing (7 pre-existing + 3 new regression tests). - Manual: with the fix, browser sign-in against a dev instance succeeds for gzip-accepting clients; before the fix the same request returned `net::ERR_EMPTY_RESPONSE`. ## Risks - Low risk. The middleware still compresses large text/JSON responses exactly as before; the changes only affect paths that previously produced corrupted or destroyed responses. - Behavioral shift: responses whose headers were already committed are now delivered uncompressed instead of being (incorrectly) buffered — this is strictly less surprising than the previous corrupted output. - Failure-path shift: a compression error now yields an uncompressed 200 response instead of a dropped connection. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking enabled, running via Claude Code / Paperclip agent harness with tool use (shell, file edit, test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>canary/v2026.710.0-canary.11 |
||
|
|
1e8ede4e1e |
fix(adapter-utils): runChildProcess escalates to SIGKILL on liveness, not child.killed (#8598)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Those agents run as child processes spawned by
`@paperclipai/adapter-utils`'s `runChildProcess`, which arms a
parent-side wall-clock timer at `timeoutSec` to bound a hung run
> - At the deadline `runChildProcess` sends SIGTERM, then after a grace
window escalates to SIGKILL — the SIGKILL backstop is what turns a
wedged child into a dead PID so the scheduler can reclaim and retry it
> - On the **direct-child fallback** path (`signalRunningProcess`, used
on win32 and whenever process-group signaling is unavailable or throws)
the escalation was gated on `!child.killed`
> - But Node sets `ChildProcess.killed` to `true` the instant a signal
is *successfully sent*, not when the process exits — so once the earlier
SIGTERM has been sent, `child.killed` is already `true`, the
`!child.killed` guard is `false`, and the SIGKILL escalation never runs
> - A child that ignores SIGTERM (e.g. a graceful-shutdown handler
wedged on a socket) is therefore never force-killed, outlives its
deadline, and for an unattended scheduler sits running forever with no
terminal state
> - This PR gates the fallback escalation on real liveness (`exitCode
=== null && signalCode === null`), so SIGKILL fires precisely while the
child is still alive
> - The benefit is the hard timeout actually guarantees termination
(except true uninterruptible D-state) on every platform/configuration,
not just where the process-group path is available
## Linked Issues or Issue Description
No existing public issue — describing the bug inline, following
`.github/ISSUE_TEMPLATE/bug_report.yml`:
### What happened?
When `@paperclipai/adapter-utils`'s `runChildProcess` reaches
`timeoutSec` and the spawned child ignores SIGTERM, the SIGKILL
escalation on the **direct-child fallback** path
(`signalRunningProcess`, taken on win32 or whenever `process.kill(-pgid,
…)` is unavailable or throws) never fires, so the child outlives its
deadline indefinitely. Root cause: the escalation is gated on
`!running.child.killed`, and `ChildProcess.killed` reflects only that a
signal was *successfully sent* (per the Node docs it "does not indicate
that the child process has been terminated"). After the deadline
SIGTERM, `child.killed` is already `true`, so `!child.killed` is `false`
and the follow-up SIGKILL is suppressed.
### Expected behavior
After the grace window, a child that is still alive is force-killed with
SIGKILL regardless of whether SIGTERM was already sent — the hard
timeout should guarantee termination (except true uninterruptible
D-state) on every platform/configuration.
### Steps to reproduce
1. Spawn a child that installs a no-op `SIGTERM` handler and never exits
(e.g. `process.on('SIGTERM', () => {}); setInterval(() => {}, 1000)`).
2. Drive it through the direct-child fallback, i.e.
`signalRunningProcess({ child, processGroupId: null }, …)` (the path
used on win32 / when group signaling is unavailable).
3. Send SIGTERM (the child swallows it; `child.killed` becomes `true`),
then send SIGKILL.
4. On the pre-fix `!child.killed` guard the SIGKILL call is a no-op and
the PID survives past its deadline. Covered by the new regression test
in this PR.
### Paperclip version or commit
Reproduces on `master` (the `signalRunningProcess` fallback). Also
present in published `@paperclipai/adapter-utils` (e.g. `2026.325.0`),
where the same `!child.killed` guard sits on the single direct-child
escalation path.
_Searched the open PR list for duplicates/related work on
`runChildProcess` / `signalRunningProcess` / SIGKILL escalation; found
none._
## What Changed
- `packages/adapter-utils/src/server-utils.ts`: in
`signalRunningProcess`, replace the direct-child fallback guard
`!running.child.killed` with `running.child.exitCode === null &&
running.child.signalCode === null` (real liveness). The process-group
path is unchanged.
- `packages/adapter-utils/src/server-utils.ts`: `export`
`signalRunningProcess` so the fallback branch can be unit-tested
directly.
- `packages/adapter-utils/src/server-utils.test.ts`: add a companion
regression test (POSIX-only, like the sibling timeout tests) that forces
the fallback (`processGroupId: null`) — sends SIGTERM (child swallows
it, `child.killed` becomes `true`), asserts the child is still alive,
then sends SIGKILL and asserts the PID dies. Also keeps the end-to-end
`runChildProcess` SIGTERM-ignoring test.
## Verification
```
npx vitest run packages/adapter-utils/src/server-utils.test.ts # 52 passed
npx tsc --noEmit # clean
```
- **Regression proof:** reverting the guard to `!running.child.killed`
makes the new fallback test fail (`waitForPidExit` → false; the child
survives); the liveness guard makes it pass. This addresses the prior
review note that the existing test only exercised the process-group path
(which already escalated correctly on POSIX) and never reached the
changed branch.
## Risks
Low. A one-line guard change scoped to the direct-child fallback; the
process-group path is untouched. SIGKILL is only sent when
`exitCode`/`signalCode` are both still `null`, i.e. the process is
provably alive, so the change cannot signal an already-reaped/recycled
PID. New tests are POSIX-only and `skipIf(win32)`, consistent with the
sibling timeout tests in this file.
## Model Used
Anthropic **Claude Opus 4.8**, driven via the Cursor agent (extended
reasoning + tool use, large context). Diff, tests, and the regression
proof above were produced and run by the agent; reviewed by a human
before pushing.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
|
||
|
|
1fe89eb8f8 |
Enforce durable external-wait liveness (#9373)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat/recovery subsystem decides whether an agent run has a durable continuation path after the process stops. > - External waits need stricter semantics than local background watchers: a killed local process is not durable, while a first-class blocker/monitor/scheduled wake is. > - Without that distinction, recovery can repeatedly treat adapter-failed continuations as live work and obscure the real reason a task stopped. > - This pull request adds explicit durable external-wait liveness handling and documents the expected execution semantics. > - It also improves operator-visible recovery evidence so invalid external-wait paths explain why they were rejected. > - The benefit is clearer recovery behavior, fewer duplicate continuation recoveries, and a safer contract for monitor-backed external waits. ## Linked Issues or Issue Description - Refs #5978 - Related PRs: #4988, #7495, #8502 ## What Changed - Added durable external-wait liveness classification so local/background watchers are not accepted as durable live paths after the owning process exits. - Preserved first-class blocker/monitor/scheduled wake paths as valid external-wait continuations. - Added backend regression coverage for killed watcher failure, monitor-backed durable wait resumption, normal completion, blocker behavior, and no duplicate recovery. - Added adapter utility coverage for terminal cleanup behavior used by local process adapters. - Surfaced invalid external-wait recovery evidence in the recovery action card and run ledger. - Updated execution semantics documentation and the V1 implementation contract. ## Verification - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed. - `node scripts/run-vitest-stable.mjs --mode general --group general-server` equivalent lane passed in CI-clean env: 238 files, 2164 tests passed, 1 skipped. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-a` passed in fully Paperclip-env-clean env: UI 305 files / 2430 tests; CLI 43 files / 230 tests. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-b` passed in fully Paperclip-env-clean env: shared/db/adapters/plugin packages all green. - `node scripts/run-vitest-stable.mjs --mode serialized` passed in fully Paperclip-env-clean env: 107 serialized server suites green, including 84/84 heartbeat-process-recovery tests. - `pnpm build` passed in fully Paperclip-env-clean env. Notes: running `pnpm test:run` directly inside the Paperclip heartbeat environment exposed local harness env contamination in existing tests (`PAPERCLIP_CONFIG`, `PAPERCLIP_DB_BACKUP_DIR`, and `PAPERCLIP_WORKTREE_START_POINT`). Re-running the same lanes with inherited `PAPERCLIP_*` and port env removed produced the CI-equivalent green results above. ## Risks - Medium behavioral risk: this changes recovery classification for stopped local external-wait processes, so adapters relying on unmanaged background watchers must use blockers, monitors, scheduled wakes, or explicit durable handoff instead. - Low UI risk: recovery-card copy changes are covered by component tests and Storybook screenshot QA. - No database migration is included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-enabled terminal/code execution. Exact context-window metadata was not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.710.0-canary.10 |
||
|
|
1f07690184 |
fix(ui): keep issue threads from jumping to latest comment (#9354)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue and item detail pages use a shared issue chat thread to show comments, runs, activity, and interactions. > - That thread still defaulted to landing on the latest comment when messages first loaded. > - On long issue/item pages, that default can yank the operator away from the top of the page before they choose to inspect the newest message. > - Deep links to comment hashes can create the same kind of initial viewport jump when they are used as generic navigation targets. > - This pull request makes initial latest-comment and initial thread-hash scrolling opt-in instead of default behavior. > - The benefit is stable initial page position across issue-thread surfaces while keeping the explicit Jump to latest control available. ## Linked Issues or Issue Description No exact public GitHub issue was found for this bug. Bug description: - What happened: opening a page with a shared issue conversation thread could automatically move the viewport toward the newest comment/thread target. - Expected behavior: ordinary page loads should keep the initial viewport stable unless the user explicitly clicks Jump to latest. - Steps to reproduce: open an issue or item detail page with a long conversation thread and observe whether the page jumps to the newest thread entry on initial load. - Paperclip version/commit: reproduced while working on the current `master` branch lineage. - Deployment mode: local trusted/dev UI. Related public thread/comment UX work: Refs #3916, Refs #7972, Refs #8800. ## What Changed - Changed `IssueChatThread` so initial latest-comment scrolling defaults to off. - Added a separate opt-in for initial thread-hash scrolling, also defaulting to off. - Preserved stale deleted-comment hash cleanup without scrolling the page. - Updated regression coverage so default initial load stays put, comment hashes do not scroll by default, and manual Jump to latest still scrolls. ## Verification - `pnpm --filter @paperclipai/ui typecheck` passed on the clean PR branch. - `pnpm --dir ui exec vitest run src/pages/IssueDetail.test.tsx -t "loads from the pending state into issue detail without changing hook order"` passed on the clean PR branch. - `pnpm --dir ui exec vitest run src/components/IssueChatThread.test.tsx` was attempted on the clean PR branch, but the file fails before changed assertions with the existing `TypeError: act is not a function` test-harness issue across 58 tests; 14 tests passed. - Static check: no `autoScrollToLatestOnInitialLoad={true}` or `autoScrollToHashOnInitialLoad={true}` call sites remain in `ui/src`. ## Risks Low risk. This only changes initial scroll defaults in the shared issue thread. The main behavioral shift is that direct comment/thread hashes no longer auto-scroll on first load unless a caller explicitly opts in; the Jump to latest button and post-submit scroll behavior are unchanged. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding-agent runtime; exact context window not exposed in this environment; tool-enabled repository inspection, editing, testing, git, GitHub CLI, and Paperclip API usage. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.9 |
||
|
|
be1fcb2b46 |
Fix agent sidebar liveness churn (#9358)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board UI includes an agents sidebar so operators can see which agents are currently active. > - That sidebar depends on live-run polling, heartbeat events, and cross-tab cache sharing to stay current without overloading the API. > - The sidebar was visually churning because agents could leave the live section immediately after a run ended, while progress events and cross-tab broadcasts kept forcing hot query updates. > - This pull request stabilizes live sidebar membership and makes shared polling broadcasts monotonic/deduplicated. > - The benefit is a calmer operator sidebar that still reflects real live state without flashing between stale and fresh snapshots. ## Linked Issues or Issue Description No public GitHub issue exists for this operator-facing bug. Bug summary: - What happened: the agents sidebar could flash or reshuffle around active agents while live-run and heartbeat data was updating. - Expected behavior: active and recently-active agents should remain visually stable, and cross-tab cache sharing should not overwrite fresher data with older snapshots. - Reproduction context: run Paperclip with multiple tabs or rapid live-run/progress updates and watch the agents sidebar while agents enter/leave live execution. - Deployment mode: local/operator board UI. Related PR: - Supersedes #9357, which carried the same fixes on a branch/title/body that were not suitable for public contribution hygiene. ## What Changed - Restored the maintainer-only warning wording in the developer skill guide so the existing server skill-utils CI gate passes on current master. - Added a 120-second linger window for streamlined sidebar agent rows so an agent does not immediately disappear from the live section as soon as its last run ends. - Deferred the recent-agent fallback until there are no live or lingering agents, while keeping the live badge tied only to actually-live runs. - Stopped broad live-runs/heartbeats/agents-list invalidation on every run progress event, while preserving targeted agent-detail invalidation. - Added producer timestamps to cross-tab shared polling result messages so older-or-equal snapshots are dropped before `setQueryData`. - Added per-resource broadcast dedupe/rate limiting so tabs do not rebroadcast equivalent cached data in a loop. - Added focused coverage for sidebar linger behavior, staggered multi-agent linger expiry, live update invalidation scope, shared polling timestamp handling, and cross-tab broadcast dedupe. ## Verification Run locally on the rebased PR branch: - `pnpm --filter @paperclipai/ui exec vitest run src/components/SidebarAgents.test.tsx src/context/LiveUpdatesProvider.test.ts` — 44 tests passed. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/cross-tab-poll.test.ts src/hooks/useSharedPolling.test.ts` — 10 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — exit code 0. ## Risks - Low migration risk: the sidebar/polling changes are UI/client cache behavior only, with no database or API contract changes. - Sidebar visibility now intentionally lingers for 120 seconds after the last live run; stale rows could remain briefly visible, but their live badge is removed when they are no longer actually live. - Cross-tab broadcasts are now more conservative; a missed publish should be corrected by the next normal poll or accepted newer timestamp. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex via the Paperclip Codex coding-agent runtime; exact API model identifier and context-window size are not exposed in this environment. The agent used terminal/tool execution for repository inspection, focused tests, branch preparation, and PR creation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.8 |
||
|
|
e84731af70 |
Revert "Fix default model adapter test config" (#9363)
Reverts paperclipai/paperclip#9361 |
||
|
|
ebd62ca5ae |
Fix default model adapter test config (#9361)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board UI lets operators create and edit agent adapter configuration, including a primary model field and an adapter test action. > - For Cody and similar adapter forms, selecting the default model means the model value is intentionally unset so the adapter can use its default. > - The adapter test path still allowed `model: undefined` to survive in the generated adapter config, which could send an invalid test payload instead of omitting the field. > - This pull request normalizes create/edit adapter test config so default-model selections omit `model` entirely. > - The benefit is that testing an agent configured to use the adapter default model exercises the same clean config shape that should be saved and run. ## Linked Issues or Issue Description No public GitHub issue was found for this local UI bug, so the problem is described inline. Bug description: - What happened: using the adapter test action after choosing the default model could include `model: undefined` in adapter config and surface a UI/runtime error instead of testing with the adapter default. - Expected behavior: choosing the default model should omit the `model` field from adapter config so the adapter default is used. - Steps to reproduce: edit a Codex/Cody-style agent with a concrete model, switch the model selector to Default, then run the adapter Test action. - Paperclip version/commit: current `master` before this PR. - Deployment mode: board UI, deployment-mode independent. ## What Changed - Sanitized adapter test config assembly so undefined adapter config entries are omitted before the test request is sent. - Made create-mode current model display resilient when the model is unset for adapter defaults. - Added regression coverage for editing an existing agent from a concrete model back to Default and testing it. - Added regression coverage for create-mode testing with an unset/default model. - Hardened the developer skill wording used by the existing server skill utility contract test. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/AgentConfigForm.render.test.tsx` - `pnpm exec vitest run server/src/__tests__/paperclip-skill-utils.test.ts` - GitHub PR checks on this branch are green, including Typecheck + Release Registry, Build, General tests, e2e, verify, security scans, and Greptile Review. ## Risks - Low risk: this only removes undefined values from adapter test config payloads, which aligns with the existing persisted patch behavior. - Low risk: default-model display now treats unset create-mode model values as an empty string. - No database, API schema, migration, auth, or adapter runtime contract changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected - check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-enabled software-engineering session. Exact context window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <cody@paperclip.ing>canary/v2026.710.0-canary.7 |
||
|
|
a4993a72a6 |
Fix live run streaming text readability (#9330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue thread UI renders live agent output from adapter run logs and transcript parsing. > - Some adapter streams emit many small or repeated token chunks, and live UI updates can expose partial words, duplicated slices, or transient markdown placeholders. > - That makes active run updates look like gibberish even when the underlying agent output is valid. > - The fix needs to preserve raw logs while making the live thread view stable, readable, and ordered. > - This pull request adds monotonic run-log sequencing, safer live transcript dedupe/order handling, markdown placeholder hiding, and readable live text stabilization. > - The benefit is a live issue thread that updates smoothly without showing confusing partial parser artifacts. ## Linked Issues or Issue Description No public GitHub issue exists yet, so this PR includes the bug details inline. ### What happened? Live run updates in the issue thread can show confusing repeated or partial text while an adapter is streaming. The visible text appears to lose parsing boundaries during active updates, especially with ACP-style token deltas, so the live output can briefly render duplicated chunks, incomplete words, or HTML-comment placeholders. ### Expected behavior Live text should remain readable while preserving the underlying run output for raw inspection. ### Steps to reproduce 1. Start a live agent run whose adapter emits small stdout token deltas. 2. Watch the issue thread while the run is still active. 3. Observe transient duplicated chunks, incomplete words, or markdown placeholder artifacts in the live rendered text. ### Paperclip version or commit Reproduced against current `master` before this PR branch. ### Deployment mode Local dev issue-thread UI with live local adapter runs. ### Additional context GitHub PR search for `live run streaming text markdown transcript` found one broad merged PR, `#252` (“Dotta updates - sorry it's so large”), but no targeted duplicate for this live streaming readability bug. ## What Changed - Added per-run monotonic sequence numbers to persisted and live run-log chunks. - Dedupe and order live transcript chunks by sequence before falling back to timestamp ordering. - Hide markdown HTML comment placeholder text from rendered markdown output. - Smooth live issue-thread text updates so partial additions reveal at readable word boundaries and sliding-window removals do not produce gibberish. - Added coverage for run-log ordering/deduping, markdown comment hiding, live issue-thread stabilization, and Greptile-reviewed edge cases where overlap rewrites could synthesize text or no-boundary additions could stay hidden. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts src/components/MarkdownBody.test.tsx src/components/transcript/useLiveRunTranscripts.test.tsx` passed before the review fix: 3 files, 86 tests. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts` passed after the review fix: 1 file, 30 tests. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts src/components/transcript/useLiveRunTranscripts.test.tsx` passed after the final Greptile overlap fix: 2 files, 40 tests. - `pnpm check:token-gates` passed. - Local PII/secret scan of touched files found only expected code/test words such as `secret`, `token`, and redaction-related strings; no literal credentials found. - `pnpm -r typecheck` passed after restoring declared dependencies with `CI=1 pnpm install --frozen-lockfile` and running with a short `TMPDIR` because `tsx` IPC sockets fail under the long sandbox temp path. - `pnpm build` passed with existing Vite CSS/font/chunk warnings. - GitHub PR checks passed on head `4c052dfe86aecb5feb73504e6b48843f68fce813`: build, typecheck/release registry, server and workspace test shards, serialized server suites, e2e, canary dry run, policy, review, Socket, Superagent, Snyk, and verify. - Greptile review passed on head `4c052dfe86aecb5feb73504e6b48843f68fce813` with confidence score 5/5 and no blocking issues found. - `pnpm test:run` failed in unrelated server workspace tests on this macOS local environment: - `server/src/__tests__/heartbeat-workspace-branch-containment.test.ts`: two assertions compare `/tmp/...` with `/private/tmp/...`. - `server/src/__tests__/heartbeat-worktree-suppression.test.ts`: expected one heartbeat run but observed two, followed by cleanup fallout in the full run. - Isolated rerun of those two server suites reproduced the same three failures. ## Risks - Low product risk for the UI changes: the readable smoothing only affects active live-run display stabilization, not stored comments or raw run logs. - Moderate verification risk: local full Vitest did not pass because of unrelated server workspace tests. Targeted tests for this change, typecheck, token gates, and build passed. - Run-log sequence fields are optional for compatibility with older log rows that do not include `seq`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-using local workspace execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
991279f52c |
Fix Skill Studio markdown dirty tracking (#9356)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Skill Studio is the UI surface for editing skill package files, including `SKILL.md`. > - Markdown files are split into frontmatter fields plus a rich markdown body editor. > - The dirty-state guard intentionally ignores markdown editor normalization during initial mount. > - That guard was too narrow: rich markdown body edits could happen before the file was marked as user-interacted, so edits did not reliably enable Save. > - This pull request broadens the user-interaction signals around the markdown body editor and adds a regression test for saving body edits. > - The benefit is that editing `SKILL.md` in Skill Studio now behaves like normal file editing: changes show as unsaved and Save persists the full markdown document. ## Linked Issues or Issue Description No public GitHub issue exists. Inline bug description follows. ### What happened? Editing a Skill Studio markdown body did not reliably mark the file dirty, so the Save action could remain unavailable or fail to persist the body edit. ### Expected behavior User edits in the markdown body editor should mark the file unsaved and allow saving the updated `SKILL.md` content. ### Steps to reproduce 1. Open a Skill Studio markdown file such as `SKILL.md`. 2. Edit the markdown body in the rich editor. 3. Observe whether the unsaved state appears and Save becomes enabled. 4. Save and reload the file. ### Paperclip version or commit Reproduced on `master` before this fix. ### Deployment mode Local dev (`pnpm dev`). ## What Changed - Mark markdown body interaction on capture-phase key, pointer, paste, drop, before-input, and input events around the rich editor. - Preserve the existing guard that prevents MDXEditor mount-time normalization from dirtying a clean file. - Add a Skill Studio regression test that edits the markdown body, observes the Unsaved state, enables Save, and verifies the saved `SKILL.md` includes both frontmatter and the edited body. ## Verification - `pnpm vitest run ui/src/pages/SkillStudio.test.tsx` Manual reviewer path: - Open a Skill Studio markdown file such as `SKILL.md`. - Edit the body text in the rich markdown editor. - Confirm the UI shows an unsaved state and the Save button is enabled. - Save and confirm the updated markdown body persists. ## Risks Low risk. The change only broadens interaction detection before applying existing dirty-state logic, and the guard still prevents initial editor normalization from marking an unopened file dirty. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex CLI using GPT-5, with repository file editing, shell execution, and GitHub CLI tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a02fe8d575 |
Update Codex adapter GPT-5.6 defaults (#9352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Codex local is the adapter subsystem that exposes OpenAI Codex CLI model choices to agents and issue overrides. > - OpenAI has GPT-5.6 Codex-capable models that should appear in Paperclip's built-in Codex model list and refresh behavior. > - Paperclip's server model listing falls back to the adapter metadata and merges OpenAI refresh results with known Codex defaults. > - This pull request updates the Codex default model metadata to include GPT-5.6 options and adds regression coverage for fallback and refresh paths. > - The benefit is that operators can select the new Codex models without relying on manual model IDs, and refresh behavior keeps known GPT-5.6 options visible. ## Linked Issues or Issue Description Refs #9322. Refs #9342. Refs #9346. ### Agent or provider Codex CLI (OpenAI). ### Why this adapter is useful OpenAI's GPT-5.6 Codex-capable models should be available in Paperclip's Codex adapter defaults and model refresh path. ### How the agent is invoked `codex` ## What Changed - Changed the `codex_local` default model metadata from `gpt-5.5` to `gpt-5.6`. - Added `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna` to the built-in Codex adapter model list. - Updated adapter and server model-listing tests to cover GPT-5.6 fallback and refresh behavior. - Aligned Codex Fast mode support and helper text with the new `gpt-5.6` default, while preserving GPT-5.5, GPT-5.4, and manual model ID support. ## Verification - `git diff --check origin/master...HEAD` - `pnpm exec vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts server/src/__tests__/adapter-models.test.ts server/src/__tests__/adapter-model-refresh-routes.test.ts` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` ## Risks Medium risk because changing `DEFAULT_CODEX_LOCAL_MODEL` from `gpt-5.5` to `gpt-5.6` changes the adapter's default model selection for new blank configurations. The model-list additions are otherwise low risk and covered by adapter/server metadata tests. This PR intentionally overlaps related PRs #9342 and #9346, so reviewers may prefer to close or fold it into one of those branches. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent based on GPT-5, with shell, git, GitHub CLI, and repository editing tool use. Exact served model ID and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
953b315dfb |
Shorten skill frontmatter descriptions (#9353)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip agents can load repository and catalog skills, and Codex renders skill names and frontmatter descriptions into startup context. > - Long descriptions consume the fixed skill metadata budget before Codex can use the progressively disclosed skill bodies. > - The repo `.agents/skills` descriptions and a few shipped catalog descriptions had grown into operational documentation instead of short trigger metadata. > - This pull request keeps the strongest trigger language in frontmatter while leaving detailed procedures in each skill body. > - The benefit is lower prompt overhead, more reliable skill triggering, and a regression guard that prevents description drift from returning. ## Linked Issues or Issue Description No public GitHub issue found for this maintenance item. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am working against `master`. - [x] I have confirmed the issue originates in Paperclip's shipped skill metadata, not in a local agent adapter or provider. ### What happened? Codex startup renders discovered skill names and frontmatter descriptions into a fixed skill metadata budget. Several repository skill descriptions and one shipped catalog description had grown into long-form operational guidance, which can force Codex to truncate descriptions before the model has enough trigger signal to select the right skill. ### Expected behavior Skill frontmatter descriptions should stay short trigger summaries: one capability sentence plus a “use when” clause. Detailed procedures should stay in the skill body and load only after the skill triggers. ### Steps to reproduce 1. Inspect `.agents/skills/*/SKILL.md` and `packages/skills-catalog/catalog/**/SKILL.md` frontmatter descriptions. 2. Measure folded YAML `description` values. 3. Observe descriptions above the intended short-trigger range, including descriptions above 300 characters. 4. Run the new shipped catalog test to verify future descriptions stay capped. ### Paperclip version or commit Reproduced on `master` at `cc81eefb6047d8eaf57faf785f421c03dc97073c`. ### Deployment mode Local dev / source checkout metadata inspection. This is not database-related. ### Installation method Built from source. ### Agent adapter(s) involved Codex, because Codex startup uses the skill metadata prompt budget. The metadata source itself is core repository/catalog content. ### Database mode Not database-related. ### Access context Not applicable; this is static repository metadata. ### Node.js version `v22.22.2` in the verification environment. ### Operating system Linux container environment. ### Relevant logs or output Final measurement after this PR: 29 source `SKILL.md` files, max description length 215 chars, 5,449 total description chars, estimated 1,363 description tokens at 4 chars/token. ### Relevant config None. ### Additional context The shipped catalog manifest was regenerated so the generated package metadata matches the edited catalog `SKILL.md` sources. ### Privacy checklist - [x] I have reviewed all pasted output for PII and redacted where necessary. ## What Changed - Shortened long `.agents/skills/*/SKILL.md` frontmatter descriptions to concise capability plus use-when trigger clauses. - Shortened the over-budget shipped skills catalog descriptions for wireframe, Paperclip capsules, and reflection coach. - Regenerated `packages/skills-catalog/generated/catalog.json` so shipped metadata matches source skill frontmatter. - Added a Vitest regression guard that caps repo skill source descriptions and generated catalog descriptions at 300 characters. ## Verification - `pnpm --filter @paperclipai/skills-catalog build:manifest` - `pnpm --filter @paperclipai/skills-catalog test` — 5 files passed, 19 tests passed - `pnpm --filter @paperclipai/skills-catalog typecheck` - Final measurement: 29 source `SKILL.md` files, max description length 215 chars, 5,449 total description chars, estimated 1,363 description tokens at 4 chars/token. Note: the clean PR worktree was created from `origin/master` and contains only this commit, but it does not have `node_modules`; running `pnpm --filter @paperclipai/skills-catalog test` there failed at tool/package resolution (`vitest`, `tsc`, `@paperclipai/shared`). The dependency-equipped workspace passed the commands above before the commit was cherry-picked onto the clean branch. ## Risks Low risk. This changes skill metadata and tests only. The main risk is over-trimming a useful trigger phrase, mitigated by keeping explicit “use when” clauses and leaving detailed guidance in the skill bodies. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent via Paperclip/Codex, with shell and file-edit tool use. Exact API model ID and context window were not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f5f461729 |
Avoid startup crash when Reflection Coach assets are missing (#9351)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server has built-in agent definitions that are loaded during startup and used to provision optional operational agents such as Reflection Coach > - Reflection Coach stores richer stock instructions, a routine description, and a bundled skill as markdown assets outside the TypeScript module body > - A deployed server can fail before it is healthy if one of those copied markdown assets is absent from `server/dist` > - A recent build fix preserves those assets during normal server builds, but runtime should still degrade gracefully if a packaged asset is missing or unreadable > - This pull request adds resilient loading for built-in Reflection Coach assets and keeps a minimal compiled fallback available > - The benefit is that a missing optional built-in agent file no longer turns into a process-wide startup crash ## Linked Issues or Issue Description Bug fix. No public GitHub issue found for this exact startup crash. - Related public PR: #9339 - What happened: the server could throw `ENOENT` while importing the built-in agent service if `server/dist/built-ins/agents/reflection-coach/AGENTS.md` was missing from a deployed build. - Expected behavior: the server should keep starting, log that the built-in asset was missing, and use safe fallback text for the optional built-in agent resource. - Steps to reproduce: build the server, remove the compiled Reflection Coach `AGENTS.md` asset from `server/dist`, then import/start the server path that loads built-in agent definitions. - Paperclip version/commit: reproduced against a deployed build containing the Reflection Coach built-in agent assets; fixed against current `master` after #9339. - Deployment mode: Node server deployment using compiled `server/dist` output. ## What Changed - Added built-in agent text loading that checks the compiled asset path first, then source/package fallback paths, then a minimal compiled-in fallback string. - Added fallback text for Reflection Coach instructions, routine description, and bundled skill content so startup does not depend on optional markdown assets being present. - Added regression coverage for readable candidate selection and missing-file fallback behavior. ## Verification - `pnpm -w exec vitest run server/src/__tests__/built-in-agents.test.ts` — 1 test file passed, 24 tests passed. - `pnpm --filter @paperclipai/server build` — server TypeScript build completed and copied `src/built-ins` into `dist/built-ins`. - Manual smoke: temporarily moved `server/dist/built-ins/agents/reflection-coach/AGENTS.md`, imported `server/dist/services/built-in-agents.js` through the repo-pinned `tsx` runtime, and confirmed Reflection Coach definitions still loaded with output `reflection-coach:3732`; the asset was restored afterward. ## Risks Low risk. The normal path still uses the full packaged markdown assets. The fallback path is only used when those files are missing or unreadable, and it logs a warning so packaging drift remains visible. ## Model Used OpenAI GPT-5 Codex coding agent, with repository tool access and shell-based verification. Exact context window was not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.6 |
||
|
|
cc81eefb60 |
Make plan-approval continuations durable after failed wakes (#9331)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents often work from reviewed plans that are approved through issue-thread interactions. > - Accepting a plan is not just a UI decision; it must reliably resume the assignee so approved work continues. > - A failed continuation wake could leave an approved plan stranded in review with no durable retry or visible recovery path. > - This pull request makes approved plan continuations retryable, recoverable, and visible when resume fails. > - The benefit is that operators can trust plan approval to either resume the agent or produce an explicit actionable failure instead of silent limbo. ## Linked Issues or Issue Description No public GitHub issue exists. Inline bug report follows the repository bug template. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest released version of Paperclip or can reproduce on `master`. - [x] I have confirmed the error originates in Paperclip itself, not in my agent adapter, API provider, or local configuration. ### What happened? When a plan-confirmation interaction was accepted, the assignee continuation wake could fail before useful agent execution. In that case the issue could remain in review even though the plan had been approved, because the failed wake was fire-and-forget and there was no durable retry or recovery path for accepted continuations. ### Expected behavior Accepted plan continuations should either wake the assignee successfully, retry bounded infrastructure failures, recover dropped wakes, or surface an explicit failure state that operators can act on. ### Steps to reproduce 1. Create an issue with an assignee and a plan confirmation that wakes the assignee on accept. 2. Accept the confirmation. 3. Simulate a pre-flight continuation failure, such as process loss before agent start or workspace validation failure. 4. Observe that the approved issue can remain in review without an active assignee wake or visible retry/failure state. ### Paperclip version or commit Reproducible on `master` before this PR's retry/recovery changes. ### Deployment mode Local dev (`pnpm dev`) and server-side recovery paths. ### Installation method Built from source (`pnpm install`, `pnpm dev`, test runner). ### Agent adapter(s) involved Not adapter-specific; this is a core continuation/recovery bug. The tests cover local-agent failure shapes without relying on a provider-specific API. ### Database mode Embedded Postgres test database for verification. The affected logic is database-backed and applies to normal Postgres deployments as well. ### Access context Board accepts the interaction; agent execution resumes through the assignee wake path. ### Relevant logs or output No sensitive logs are needed. The regression tests simulate the failed wake and recovery states directly. ### Relevant config No special config is required beyond an assignee with wake-on-demand enabled. ### Additional context This PR also prevents a stale workspace-validation payload from quarantining another issue's active workspace and prevents unrelated successful runs from masking a continuation that never resumed. ### Privacy checklist - [x] I have reviewed all pasted output for PII, usernames, file paths, API keys, tokens, and company names, and redacted where necessary. ## What Changed - Added bounded infrastructure retries for failed accepted-interaction continuation wakes. - Extended stranded issue recovery so dropped accepted-plan continuation wakes are requeued. - Recorded and rendered explicit resume-failure state on accepted confirmation cards. - Added clean-workspace fallback for workspace-validation failures while preventing cross-issue workspace quarantine. - Tightened recovery so unrelated successful runs do not mask an accepted continuation that never resumed. - Added focused server/UI coverage for retry scheduling, recovery, visible failure state, and interaction card rendering. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-retry-scheduling.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts` — 106 tests passed. - Earlier branch verification also covered the issue-thread interaction card tests for the visible resume-failure UI. - GitHub CI is green on the replacement PR head, and Greptile reports 5/5 with no blocking issues. ## Risks - Medium behavioral risk: this changes recovery behavior for accepted continuation interactions and workspace-validation retries. - Mitigation: retries are bounded, scoped to same-company issue context, and workspace quarantine now requires ownership by the issue being retried. - Existing stored confirmation results remain compatible because the new resume-failure field is optional. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-enabled terminal workflow. The runtime does not expose an exact context-window value to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.5 |
||
|
|
d166069bc4 |
Preserve built-in agent assets in server builds (#9339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server package ships compiled runtime code plus static runtime assets. > - Built-in agent definitions live under `server/src/built-ins` and runtime code resolves them relative to compiled server files. > - The server build already copied onboarding assets into `dist`, but it did not copy built-in agent assets alongside the compiled code. > - Packaged server builds could therefore miss built-in agent definitions even though source-based development runs worked. > - This pull request extends the server build copy step to preserve built-in agent assets in `dist/built-ins`. > - The benefit is packaged server builds keep the same built-in agent runtime assets available as source-based development runs. ## Linked Issues or Issue Description No public GitHub issue found. This PR describes the underlying bug inline using the bug report template fields. ### What happened? `@paperclipai/server` build output copied `server/src/onboarding-assets` into `server/dist/onboarding-assets`, but did not copy `server/src/built-ins` into `server/dist/built-ins`. Runtime code for built-in agents resolves those assets relative to the compiled server files, so packaged builds could omit built-in agent markdown assets that are present during source-based development. ### Expected behavior Packaged server builds should include built-in agent assets under `server/dist/built-ins`, matching the runtime location expected by the compiled server code. ### Steps to reproduce 1. Check out current `master` before this PR. 2. Run `pnpm --filter @paperclipai/server build`. 3. Check for `server/dist/built-ins/agents/reflection-coach/AGENTS.md`. 4. Observe that the built-in agent asset is missing from the server build output. ### Paperclip version or commit Reproduces on current `master` before this PR. The fix is verified on commit `2b89984ccb7857f06359bf65c48222f110c7aeff`. ### Deployment mode Build/package artifact behavior. This can affect any deployment mode that runs from the built server package rather than directly from source. ### Installation method Built from source with `pnpm --filter @paperclipai/server build`. ### Agent adapter(s) involved Not adapter-specific. This is a core server packaging bug for built-in agent assets. ### Database mode Not database-related. ### Access context Not applicable. This happens during package build output generation. Related search: - Searched public PRs/issues for `built-ins build copy repo:paperclipai/paperclip`. - Found no directly related open issue. One old closed Hermes adapter PR was not directly related. ## What Changed - Updated the `@paperclipai/server` build script to create `dist/built-ins`. - Added the copy step from `server/src/built-ins` into `server/dist/built-ins` alongside the existing onboarding asset copy. - Added a focused server package build-script test that asserts both onboarding and built-in static runtime asset directories are copied into `dist`. ## Verification - `pnpm exec vitest run server/src/__tests__/server-package-build-script.test.ts` - `pnpm --filter @paperclipai/server build` - `test -f server/dist/built-ins/agents/reflection-coach/AGENTS.md` ## Risks Low risk. This changes only the package build asset copy step and adds focused test coverage. The main risk is build-script portability, but it follows the existing `mkdir -p` and `cp -R` pattern already used for onboarding assets. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent in a tool-enabled Paperclip heartbeat. Exact model ID and context-window size are not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.4 |
||
|
|
05973b2073 |
Enable sandbox environments for Grok local adapter (#9338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local CLI adapters can run against Paperclip-managed execution environments instead of only the host filesystem. > - The environment picker and environment capability API derive sandbox support from shared adapter capability lists. > - The Grok Build adapter is implemented as a local CLI adapter, but it was missing from those shared environment capability lists. > - That made Grok agents look local-only even when sandbox environments were configured. > - This pull request registers `grok_local` in the shared adapter constants and remote-managed environment support path. > - The benefit is that Grok Build agents can select the same local, SSH, and sandbox environment overrides as other local CLI adapters. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Inline bug report follows. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on `master`. - [x] I have confirmed the error originates in Paperclip itself, not in the Grok adapter provider or local configuration. ### What happened? When configuring a Grok Build local agent in the board UI, Paperclip did not expose configured sandbox environments as selectable environment overrides. The shared environment capability helper treated `grok_local` as local-only because it was missing from the remote-managed local adapter allowlist. ### Expected behavior Grok Build should behave like other local CLI adapters: when environments are enabled and a runnable sandbox environment exists, the agent configuration form should show the environment override selector and allow the sandbox to be selected. ### Steps to reproduce 1. Enable environments in instance experimental settings. 2. Configure at least one runnable sandbox environment. 3. Open the agent configuration form for a Grok Build local agent. 4. Observe that the sandbox environment is not offered as an override before this fix. ### Paperclip version or commit `master` before this PR. ### Deployment mode Local dev (`pnpm dev`). ### Installation method Built from source (`pnpm dev` / `pnpm build`). ### Agent adapter(s) involved - Grok Build local adapter. - Core bug in shared environment capability logic. ### Database mode Not database-related. ### Access context Board human operator. ### Relevant logs or output No runtime error is emitted; the issue is a missing UI option caused by shared capability metadata. ### Relevant config No secret-bearing config required. Reproduction only needs environments enabled and a runnable sandbox environment configured. ### Privacy checklist - [x] I have reviewed all pasted output for PII, usernames, file paths, API keys, tokens, company names, and redacted where necessary. ## What Changed - Added `grok_local` to the shared built-in adapter type list. - Added `grok_local` to the remote-managed adapter set used by environment capability helpers. - Added shared regression coverage for Grok local sandbox provider and driver support. - Added a UI render regression test that confirms Grok Build agents show the environment override when a runnable sandbox exists. ## Verification - `git diff --check` - Changed-file secret scan with `rg` for common token/key patterns. - `pnpm --filter @paperclipai/shared exec vitest run src/environment-support.test.ts` - `pnpm --dir ui exec vitest run src/components/AgentConfigForm.render.test.tsx` ## Risks - Low risk. This expands environment support for an existing local adapter to match the local CLI adapter behavior already used by Claude, Codex, Gemini, OpenCode, Cursor, and Pi. - Operators still need at least one configured runnable sandbox environment before a Grok agent has a sandbox option to select. - This PR was created from a Paperclip execution workspace branch whose name is runtime-provided; the PR body intentionally avoids internal issue identifiers or instance-local links. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 coding agent, tool-enabled terminal/code execution environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5c85ae64a0 |
Cases: experimental first-class case object (#9198)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board currently uses issues for execution, but longer-lived content work needs a separate object that can survive beyond a single task thread. > - The Cases subsystem adds an experimental, company-scoped record for content artifacts and their supporting metadata. > - The backend needs durable storage, API routes, revision history, issue linkage, and company-boundary enforcement before the UI can depend on Cases. > - The UI needs an opt-in navigation surface, list/detail views, reference chips, and issue-page context so operators can inspect Cases without making them the default workflow. > - The agent-facing skills need a contract for creating and updating Cases so automated content workflows can dogfood the feature. > - This pull request ships that experimental end-to-end path behind the `enableCases` flag. > - The benefit is a first-class place to collect content work, references, attachments, revisions, and related execution threads without polluting the core issue model. ## Linked Issues or Issue Description No public GitHub issue exists for this experimental feature. Feature request fields: ### Problem Content-oriented work such as release notes, announcements, docs, and campaigns can span many execution issues, which makes the final artifact hard to find and reason about after the execution thread moves on. ### Proposed solution Add an experimental Cases object that is company-scoped, linked to issues, queryable through the API, inspectable in the board UI, and writable by agent workflows through documented conventions. ### Alternatives considered Continue encoding content artifacts directly in issues or documents only. That keeps the data model smaller, but it does not give operators a stable artifact-centric view or a clean way to link related execution history. ### Roadmap alignment Checked `ROADMAP.md`; this PR does not duplicate an existing planned core roadmap item. ## What Changed - Added the `cases` data model, migration, schema exports, and experimental `enableCases` instance setting. - Added company-scoped Cases API routes for list/detail/update, issue links, revisions, children, activity events, annotations, attachments, and idempotent agent-oriented upserts. - Scoped case and issue lookup helpers before access checks so inaccessible cross-company identifiers resolve as not found rather than leaking existence. - Fixed case PATCH timestamp handling so non-status updates cannot overwrite `completedAt` from a stale pre-transaction row snapshot. - Moved Cases list type/status/project filters into the server request before the server-side limit is applied, including multi-select filters and no-project filtering. - Added backend route coverage for creation, updates, idempotency, issue linking, attribution, company-boundary enforcement, OpenAPI registration, list filtering, timestamp patch behavior, and inaccessible lookup regressions. - Added the experimental Cases UI surface: sidebar entry, gated routes, list filters/grouping, detail overview, activity, revisions, children, attachments, and issue-page case rail. - Added case reference rendering and company-prefixed case href generation so case links resolve directly inside the active company route. - Added Paperclip skill documentation for agent workflows that create or update Cases. - Wired release-content skills to emit Cases for dogfooding. - Rebased onto current `master` and renumbered the Cases migrations to `0143`/`0144` after the latest upstream migration sequence. ## Verification - Current PR head: `ecc13be0d`. - Rebased on current `master` (`606aa4f266`) and pushed to the existing PR branch. - `git diff --check origin/master...HEAD` — passed before the first update push; subsequent committed diffs were also checked with `git diff --check` before commit. - Guardrails checked: no `pnpm-lock.yaml` changes, no `.github/workflows` changes, and changed-file count is below the Greptile 100-file limit. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/cases-routes.test.ts src/__tests__/instance-settings-service.test.ts src/__tests__/openapi-routes.test.ts` — passed, 3 files / 26 tests before review-fix commits. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/cases-routes.test.ts` — passed after each server-side Greptile fix, latest 1 file / 15 tests. - `pnpm --filter @paperclipai/server typecheck` — passed after the timestamp and lookup fixes. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Cases.test.tsx src/pages/CaseDetail.test.tsx src/pages/CompanySkills.test.tsx src/App.cases-routing.test.tsx` — passed, 4 files / 30 tests before review-fix commits. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Cases.test.tsx` — passed after the list-filter fix, 1 file / 12 tests. - `pnpm --filter @paperclipai/ui typecheck` — passed after the list-filter fix. - `pnpm check:token-gates` — passed after UI changes. - Remote PR checks on head `ecc13be0d` are green: Paperclip CI, build, typecheck, test matrix, e2e, Canary Dry Run, policy, commit review, Superagent Security Scan, Socket, Snyk, and Greptile passed; Storybook visual regression is skipped and security-review is neutral. - Greptile Review: 5/5 confidence, zero unresolved Greptile threads. ## Risks - Medium feature risk because this introduces a new experimental domain object across database, server, shared contracts, skills, and UI. - The feature is gated behind `enableCases`, which limits default operator exposure while the model is exercised. - Case links now prefer company-prefixed hrefs; the unprefixed redirect remains for externally entered URLs. - Cases list filtering now sends multi-select filters to the server before limiting; the UI still applies the same local filters as a second pass for ancestor/context rows. - Migrations were renumbered on top of current master; the SQL uses guarded `IF NOT EXISTS` / `ADD COLUMN IF NOT EXISTS` patterns where relevant for safer replay. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex in the Paperclip local coding environment was used for this PR curation, rebase verification, review-fix implementation, push, and PR description update. The runtime exposes tool use and shell execution; context-window size is not exposed by this Paperclip adapter. Several implementation commits also include AI co-author trailers recorded in git history. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.710.0-canary.3 |
||
|
|
9a1d4b7983 |
fix(ui): use prose editor for markdown agent instructions (#9332)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent setup depends on instruction files that are readable and editable from the board UI > - The instructions tab already receives server-side metadata describing whether each file is Markdown > - The UI was deciding Markdown editor usage primarily from the file extension, which makes extensionless Markdown instruction files feel like raw code > - This pull request makes the instructions editor trust server Markdown metadata for existing files and keep extension fallback only for new unsaved files > - The benefit is that AGENTS-style prose instructions render and edit like prose while explicitly non-Markdown files still use the raw textarea ## Linked Issues or Issue Description - Refs #8201 - Refs #5652 - Refs #3427 - Refs #2068 - Related PRs: #2468, #2620 ## What Changed - Use server `markdown` metadata from instruction file details/summaries to choose the prose Markdown editor for existing instruction files. - Keep extension-based Markdown detection only for pending new files before server metadata exists. - Remove the monospace content styling from the Markdown editor path so prose instructions read like normal text. - Add focused tests for extensionless Markdown files, new `.md` files, and `.md` files explicitly marked non-Markdown by the server. ## Verification - `pnpm check:token-gates` - `pnpm exec vitest run ui/src/pages/AgentDetail.instructions.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` ## Risks - Low risk. The editor selection now depends on server metadata for existing files, so incorrect server metadata would choose the wrong editor. The fallback still preserves extension-based behavior for newly created unsaved files. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 via Codex coding agent, tool-enabled terminal workflow. Context window details were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.2 |
||
|
|
606aa4f266 |
feat(search): filters, sorting, operators & command-palette parity (#9327)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company search is the primary way operators find issues, comments, documents, artifacts, agents, and projects across a busy company > - Search previously supported only a bare text query: no way to narrow by status/assignee/project/label/date, no sort control, no typed operators, and weak relevance/snippets meant hunting through noise > - As companies accumulate tens of thousands of items, unfiltered single-sort search stops scaling for day-to-day operator workflows > - This pull request adds a full filtering model (filter bar, chips, mobile sheet, URL state), sort modes, typed query operators (`status:`, `assignee:`, `type:`, …) with command-palette parity, relevance/snippet/deep-link improvements, zero-results recovery, and the supporting shared validators, backend service work, and DB indexes > - The benefit is that operators can go from a vague query to the exact item in a couple of keystrokes, on desktop and mobile, with shareable filtered-search URLs ## Linked Issues or Issue Description No existing public GitHub issue; describing the underlying feature request inline (per feature_request template): - **Problem:** Company search accepted only a plain text query. Users could not filter results by status, assignee, project, label, or recency; could not change result ordering; and got no guidance when filters emptied the result set. - **Desired solution:** Structured search filters (UI controls + typed query operators + URL parameters), selectable sort modes, better relevance and snippets with exact deep links, and parity between the search page and the command palette. - **Alternatives considered:** Client-side filtering of unfiltered results (does not scale past the fetch limit); a separate "advanced search" page (splits the surface and duplicates state handling). Related (not duplicate) PRs found while searching: #4848 (issue search query planning), #8235 (search rate limiting). ## What Changed - **Shared contract:** new search filter/sort/count/zero-results types and validators in `packages/shared` (`validators/search.ts`, types index). - **Backend:** `server/src/services/company-search.ts` supports issue filters, sort modes, per-filter option counts, snippets, artifact visibility, and zero-results loosen suggestions; single-statement match replaces per-scope scans and predicates are trigram-index compatible (~3.7s → ~350ms on a live 14.8k-hit corpus). - **DB:** migration `0142_company_search_sort_indexes.sql` adds the supporting indexes. - **Search page (`ui/src/pages/Search.tsx`):** filter bar, removable chips, mobile filter sheet with result-count preview, sort menu, URL round-tripping, zero-results recovery UI. - **Query operators (`ui/src/lib/search-query-parser.ts`):** typed operators parsed into filters, operator autocomplete, filter pills. - **Command palette:** operator-aware parsing and full-search handoff. - **Stale-operator fix (latest commit):** typed operator filters are no longer folded into persistent URL-filter state, so deleting a token (e.g. removing `status:blocked` from the input) actually removes the filter from subsequent requests; filter-control edits materialize control state and strip typed tokens so a removed chip cannot resurrect from the input. ## Verification - `cd ui && npx vitest run src/pages/Search.test.tsx` — 19 tests including two new red→green regressions for the stale-operator paths (both fail on the previous commit, pass now). - `cd ui && npx vitest run src/components/CommandPalette.test.tsx` and `cd server && npx vitest run src/services/company-search-service.test.ts` — operator parity and backend filter/sort/count coverage. - `cd ui && npx tsc --noEmit` — clean. - Manual: open `/search`, type `auth status:blocked`, confirm the status filter applies; delete `status:blocked`, confirm results are unfiltered again; drive the same filters from the filter bar/chips/mobile sheet and confirm the URL round-trips (reload/back/forward preserves state). - Full end-to-end QA pass (9/9 acceptance checks) against the wireframes on desktop (1280px) and mobile (390px) with a live API and browser automation. ## Risks - Additive migration (indexes only, no data rewrites) — safe to roll forward; index creation cost is paid once at migrate time. - Search request shape gains optional parameters only; old clients keep working. - Behavioral shift: filter-control edits now strip typed operator tokens from the query text (their values persist as filter state) — deliberate, so removed filters stay removed. - Ranking changes alter result ordering for existing queries; covered by service tests and the QA pass. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic, extended thinking + tool use) — stale-operator-filter fix, regression tests, PR preparation. - GPT-5 Codex (`codex_local` adapter) and Claude Opus 4.6 (`claude-opus-4-6`) — earlier implementation phases (backend contract, filter UI, operators, ranking) under agent orchestration. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details — pre-existing branch name retained to avoid closing/reopening the PR - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing docs affected) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending re-run on latest commit) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending re-review of the stale-filter fix) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.710.0-canary.1 |
||
|
|
cec0fc249a |
[codex] Parallelize release verify workflow (#9168)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Releases publish the same app and package set that operators install, so release verification should keep full release-strength coverage. > - The release workflow currently verifies stable and canary releases with one serial job that typechecks, runs all tests, and builds. > - The PR workflow already proves the test surface can be split into grouped general suites and serialized shards without changing coverage. > - This pull request extracts the release verify work into a reusable workflow and fans out the independent lanes. > - The benefit is faster stable and canary release verification while preserving the existing publish and preview gates. ## Linked Issues or Issue Description No public GitHub issue exists for this CI improvement. **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Release verification spends most of its wall time in a single serial test step even though the same stable test surface is already partitioned for PR CI. Stable dispatches and master-push canaries therefore wait on one long runner after setup, typecheck, tests, and build run sequentially. **Proposed solution** Add a reusable release verification workflow with parallel typecheck, grouped general tests, serialized test shards, and build lanes. Have both stable and canary release verification call it with the ref they need to verify. **Alternatives considered** Keeping the serial `pnpm test:run` job preserves the old shape but keeps stable and canary releases waiting on one long runner. Skipping verification when a source SHA already has green CI would be faster, but adds stale-check and lookup risk beyond this change. **Roadmap alignment** No overlapping item found in `ROADMAP.md`; this is release CI maintenance. **Additional context** The new workflow keeps the release-strength full `pnpm -r typecheck`, uses the existing stable test grouping/sharding entry points, and leaves publish/preview jobs unchanged. ## What Changed - Added `.github/workflows/release-verify.yml` as a `workflow_call` workflow accepting a `ref` input. - Split release verification into parallel `typecheck`, `general_tests`, `serialized_tests`, and `build` jobs with 20-minute lane timeouts. - Mirrored the PR workflow's stable test partition: `general-server` shards 1-3, `general-workspaces-a`, `general-workspaces-b`, and four serialized shards. - Replaced `release.yml` `verify_canary` and `verify_stable` job bodies with calls to the reusable workflow while leaving publish and preview jobs unchanged. - Added a Node test that guards the release workflow delegation and split verify surface. ## Verification - `actionlint 1.7.12 .github/workflows/release.yml .github/workflows/release-verify.yml` - `node ./scripts/release-package-map.mjs check` - `node --test ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check` ## Risks - Release verification now starts more jobs per release event, increasing total runner setup/install minutes. This matches the existing PR CI tradeoff and should reduce release wall time substantially. - The called workflow checks out the requested ref shallowly. That is intentional for verify lanes; publish and preview jobs still retain their existing full-history checkouts. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent in local tool-use mode with shell execution, repository editing, GitHub connector access, and medium reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d3e26a8d02 |
docs: point Paperclip docs links at docs.paperclip.ing (#9300)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The public README files are owned discovery surfaces for users who arrive from GitHub or npm. > - Some documentation links still used the old `https://paperclip.ing/docs` redirect path. > - Redirect hops are worse for users and for SEO because crawlers and readers do not land on the canonical docs host immediately. > - Package metadata should still point at the GitHub repository, because npm package homepages are expected to identify the source/project page. > - This pull request updates only explicit documentation links to the canonical docs subdomain. > - The benefit is a smaller, clearer link sweep with no package homepage metadata change. ## Linked Issues or Issue Description - No public GitHub issue exists for this small documentation maintenance change. - Problem: public README documentation links used a redirecting docs URL instead of the canonical docs host. - Expected behavior: README documentation links should point directly at `https://docs.paperclip.ing`. - Scope: root README, CLI README, and the Hermes adapter README docs reference. - Related search results reviewed: #793, #592, #675, and this PR. No open duplicate PR was found for this README-only canonical docs URL sweep. ## What Changed - Updated the root README Docs navigation link from `https://paperclip.ing/docs` to `https://docs.paperclip.ing`. - Updated the CLI README Docs navigation link from `https://paperclip.ing/docs` to `https://docs.paperclip.ing`. - Updated the Hermes adapter README Paperclip Docs link from `https://paperclip.ing/docs` to `https://docs.paperclip.ing`. - Kept all `package.json` homepage fields pointing at the Paperclip GitHub repository or package-specific GitHub README pages. ## Verification - `git diff --check origin/master...HEAD` - `rg 'https://paperclip\.ing/docs|https://docs\.paperclip\.ing' README.md cli/README.md packages/adapters/hermes/README.md` - `rg '"homepage": "https://docs\.paperclip\.ing"' -g 'package.json'` returned no matches. - Reviewed the final diff against `origin/master`; only the three README files changed. ## Risks - Low risk: this is a documentation-only URL update. - The main review risk is scope creep into npm metadata; that was explicitly avoided by keeping package homepages on GitHub. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5-based Codex coding agent with local shell, git, GitHub CLI, and Paperclip API tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3369c0dab7 |
fix(prompt): render exact branch name with backtick-safe fence in wake branch guard (#9326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run heartbeats inside execution workspaces pinned to a specific git branch; the wake prompt now carries a one-time "stay on this branch" guard (#9319) > - Greptile's final review round on #9319 landed after the PR merged: the guard sanitized the branch name by stripping backticks, which mutates the ref, so the prompt could tell the agent to stay on a branch name that does not exist > - A guard that names the wrong branch defeats its purpose and still leaves the workspace contract breakable > - This pull request keeps the pinned ref name exact and instead escapes it at render time with a backtick fence longer than any backtick run inside the name (standard Markdown inline-code escaping) > - The benefit is the guard always names the real branch while a hostile ref name still cannot close the code span or inject prompt text ## Linked Issues or Issue Description Refs #9319 — follow-up addressing the final Greptile review round that arrived after that PR merged. ## What Changed - `normalizePaperclipWakeExecutionWorkspace` no longer strips backticks from the branch name; it removes only control characters (illegal in git ref names, and the newline route into the prompt), trims, and caps length. - Added a `markdownInlineCode` helper that wraps a value in an inline-code span whose backtick fence is one longer than the longest backtick run in the value, and used it when rendering the branch guard line. - Updated the hostile-branch-name test to assert the exact ref is preserved and fenced, and control characters are removed. ## Verification - `pnpm vitest run packages/adapter-utils/src/server-utils.test.ts` — 58 tests pass, including the updated hostile-branch-name case. - `npx tsc --noEmit` in `packages/adapter-utils` — clean. - Manual: render a wake payload with `branchName: "evil` + backtick + `name"` and confirm the guard line reads ``` `` evil`name `` ``` and the span does not break. ## Risks - Low risk: prompt-rendering-only change; the normalized payload shape is unchanged. Branch names containing backticks (extremely rare) now render exactly instead of mutated. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) with extended thinking and tool use, via Claude Code / Paperclip agent harness. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.0 |
||
|
|
f3ca4d24bc |
fix: repair dirty/foreign-branch execution worktrees (#9297)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Execution workspaces are the bridge between Paperclip's control plane and a local agent's checked-out repository state. > - When a workspace is restored after a failed or interrupted run, the recorded branch can disagree with the branch currently checked out on disk. > - A clean branch mismatch can be reconciled safely, but a dirty mismatch needs a lossless path that does not discard uncommitted agent work. > - This pull request adds a quarantine-and-restore path that saves dirty work to a rescue branch, restores the recorded branch, and exposes the repair from the board UI and run page. > - The benefit is that operators can recover wedged execution workspaces without losing work or moving another live branch unexpectedly. ## Linked Issues or Issue Description No public GitHub issue exists for this workspace-recovery failure, so this PR includes the bug report inline. **What happened** A git worktree-backed execution workspace could become wedged when Paperclip expected one branch but found a different checked-out branch with dirty tracked or untracked files. The existing safe repair path refused the restore, leaving the source task blocked with no lossless one-click recovery path. **Expected behavior** Paperclip should preserve dirty work before restoring the recorded workspace branch. If another live workspace claims the checked-out branch, or an attached runtime service is active, the repair should refuse with clear operator-facing evidence instead of risking work loss or file contention. **Steps to reproduce** Create a git worktree execution workspace whose persisted branch name differs from the checked-out branch, add dirty tracked or untracked files in that worktree, then trigger workspace validation or use the branch reconcile endpoint. Before this change, the dirty mismatch remained blocked because Paperclip had no quarantine restore mode. **Paperclip version or commit** Observed on the pre-fix workspace-recovery implementation. Verified on this PR head after rebasing onto current `master`. **Deployment mode** Local trusted development/worktree deployments using git worktree execution workspaces and optional workspace runtime services. ## What Changed - Added dirty-worktree quarantine repair that creates a rescue branch, commits dirty tracked and untracked files there, restores the recorded branch, writes audit comments/activity, and preserves the live foreign branch ref. - Added `quarantine_restore` branch reconcile API support, recovery-action resolution, source-task wake behavior, execution-review preservation, claimant refusal, runtime-service refusal, and coverage for the non-transactional git ordering. - Added board UI controls for the repair action in the recovery card plus a compact failed-run workspace recovery surface that uses the same reconcile handlers. - Hardened Greptile follow-up cases by best-effort restoring the recorded branch after a mid-sequence rescue commit failure and by refusing quarantine restore while attached runtime services are active. ## Verification - `pnpm exec vitest run server/src/__tests__/execution-workspaces-service.test.ts -t "quarantine_restore"` - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts -t "workspace dirty quarantine branch repair"` - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts -t "repairs clean unrecorded branch drift|adopts unrecorded forward branch drift"` - `pnpm exec vitest run server/src/__tests__/workspace-runtime-routes-authz.test.ts` - `pnpm --filter @paperclipai/server typecheck` - Earlier PR verification covered the route, service, heartbeat, UI component, and run-page recovery surfaces; Cutter posted public preview screenshots for the repair popover and run-page panel at https://github.com/paperclipai/paperclip/pull/9297#issuecomment-4926934211. - GitHub PR checks are green on `dcac76b05f4cf6e1ee16544c2831d83c7857e475`. - Greptile is 5/5 with zero annotations and no unresolved review threads on `dcac76b05f4cf6e1ee16544c2831d83c7857e475`. ## Risks - Moderate risk because the change intentionally runs git commands against local worktrees; the implementation refuses dirty repair when another claimant or active runtime service is detected and records rescue refs for auditability. - Compatibility / release-note callout for self-hosted operators: existing instances that left `enableWorkspaceBranchReconcileForward` unset now get automatic forward branch reconciliation during heartbeat workspace recovery. Operators who want the previous advisory-only behavior can set `experimental.enableWorkspaceBranchReconcileForward` to `false`; dirty quarantine repair can likewise be disabled with `experimental.enableWorkspaceDirtyQuarantineRepair: false`. - If the git rescue succeeds but a later database write fails, the worktree may already be restored while the recovery action remains open; this ordering is documented in code because git side effects cannot participate in the database transaction. - UI risk is limited to the workspace recovery surfaces and covered by component tests plus the existing Cutter visual preview. ## Model Used OpenAI Codex coding agent based on GPT-5, with repository tool use, shell execution, and local test execution. Earlier preserved commits on this branch also show Claude Code / Claude Opus 4.8 assistance in their commit metadata. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] Branch naming exception documented: this PR preserves the existing worktree branch requested for publication while keeping the PR title and body public-facing - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |