mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 00:54:38 +02:00
70ce005bef50a59f4dec45386dcd884a6fac0ea7
3030
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
70ce005bef |
Ensure worktree execution starts only after activation (#9374)
## Thinking Path > - Paperclip is the open-source control plane people use to manage AI agents and their work. > - Its scheduler, routines, and heartbeat services decide when agents automatically begin work. > - Experimental per-worktree execution is useful for isolated development, but enabling it previously allowed automatic services to consider an existing backlog. > - A worktree activation must therefore create a durable eligibility boundary rather than merely toggle execution on. > - This pull request records an activation cutoff and applies it consistently to automatic routine and heartbeat dispatch. > - The result is that an enabled worktree executes only work created after its own activation, while non-worktree behavior remains unchanged. ## Linked Issues or Issue Description **Problem type:** Bug / safety regression **Summary:** Enabling experimental run execution in an existing worktree could start automatic scheduler, routine, watchdog, and heartbeat activity for work created before that worktree was explicitly armed. **Expected behavior:** A worktree that has execution enabled only considers automatically dispatched work created on or after its activation timestamp. Ambiguous activation state fails closed. Non-worktree instances keep their existing behavior. **Related public work:** Refs #8275 (runtime worktree policy gating); this PR adds an activation-time boundary for automatic execution rather than changing the general runtime policy. ## What Changed - Persist a worktree execution activation timestamp and originating instance ID; stamp them only when the experimental toggle changes from disabled to enabled. - Resolve activation state fail-closed when the cutoff is missing, invalid, disabled, or belongs to another instance. - Gate automatic routine scheduling, webhooks, watchdog activity, and heartbeat selection at the activation cutoff; manual runs remain available. - Share the canonical worktree truthy-environment helper across routine dispatch and agent inbox filtering. - Add cutoff and truthy-runtime regression coverage, plus experimental-settings UI states that explain armed and suppressed execution. ## Verification - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts server/src/__tests__/instance-settings-service.test.ts` — passes: 2 files, 60 tests. - `pnpm --filter @paperclipai/server typecheck` — passes. - Existing CI completed successfully before the follow-up review fixes; this branch was rebased onto the latest `origin/master` before retesting. ## Risks - **Behavioral:** Automatic worktree execution is intentionally more restrictive; pre-existing work is suppressed until newly created after activation. - **Operational:** A malformed or cross-instance activation record fails closed, requiring an operator to disable and re-enable the experimental toggle on the intended worktree. - **Compatibility:** The worktree environment now accepts all canonical truthy values (`1`, `true`, `yes`, and `on`) consistently; non-worktree instances are unaffected. - **Branch metadata:** This existing execution-workspace branch predates the current naming rule and cannot be renamed under this task's workspace contract; the code and PR title do not include internal ticket references. > `ROADMAP.md` was checked; this targeted execution-safety fix does not duplicate planned core work. ## Model Used - Anthropic Claude Code — assisted with the original implementation; exact model identifier and context window were not recorded in the repository metadata. - OpenAI Codex CLI — assisted with PR preparation and review fixes; exact model identifier and context window are not exposed in this execution environment. Used with terminal tooling, code editing, and targeted test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.710.0-canary.12 |
||
|
|
23f34491e2 |
Fix apiCompression corrupting and dropping Better Auth responses for gzip clients (#9381)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its server fronts every API route — including Better Auth sign-in — with Express middleware, and #9190 added an `apiCompression` middleware that gzips JSON responses over 1KB > - That middleware buffers `res.write()` chunks with `String(chunk)`, but Better Auth (via better-call) streams `Uint8Array` chunks and commits headers with `writeHead()` before streaming > - `String(Uint8Array)` serializes the body to comma-separated decimal bytes (~3.4x inflation), and once the inflated body crossed the 1KB threshold, `setHeader()` threw `ERR_HTTP_HEADERS_SENT` and the catch handler destroyed the socket > - Every real browser sends `Accept-Encoding: gzip`, so sign-in returned zero bytes (`net::ERR_EMPTY_RESPONSE` / "Failed to fetch"), while curl without `Accept-Encoding` worked — making the bug easy to misdiagnose as a client or network issue > - This pull request makes the middleware byte-safe for `Uint8Array` chunks, passes through responses whose headers are already committed, and falls back to the uncompressed body instead of destroying the connection when compression fails > - The benefit is that browser sign-in (and any other streamed binary-chunk response) works again for gzip-accepting clients, with regression tests locking in all three behaviors ## Linked Issues or Issue Description Refs #9190 (introduced the `apiCompression` middleware). No public GitHub issue exists; bug description: - **What happened:** Sign-in from any real browser failed with `net::ERR_EMPTY_RESPONSE` / "Failed to fetch". The server logged `ERR_HTTP_HEADERS_SENT` from the compression middleware and destroyed the response socket, so zero bytes reached the client. - **Expected:** `/api/auth/*` responses are delivered intact regardless of the client's `Accept-Encoding`. - **Steps to reproduce:** Run the server with API compression active, open the web UI in a browser (which sends `Accept-Encoding: gzip`), and attempt email/password sign-in. The auth response body exceeds ~300 bytes, so after the ~3.4x stringification inflation it crosses the 1024-byte compression threshold and the response is destroyed. `curl` without `Accept-Encoding` succeeds against the same server. - **Scope:** Any route that streams `Uint8Array` chunks and/or commits headers via `writeHead()` before writing — in practice all Better Auth routes served through better-call. ## What Changed - `server/src/middleware/api-compression.ts`: - Buffer `res.write()` chunks with a `toBodyBuffer()` helper that converts `Uint8Array`/`ArrayBuffer` views via `Buffer.from()` instead of `String()`, so binary chunks are preserved byte-for-byte. - Pass responses through untouched once headers are already sent (`writeHead()`-style streaming), since compression headers can no longer be set at that point. - On any compression failure, write the original uncompressed body instead of calling `res.destroy()`, so clients get a valid (just uncompressed) response rather than a dropped connection. - `server/src/__tests__/api-compression.test.ts`: three new regression tests — small `writeHead`+`Uint8Array` responses are delivered byte-for-byte, large ones no longer drop the connection, and `Uint8Array` JSON bodies gzip without corruption (includes `/api/auth-bridge` and `/api/uint8-json` test routes mirroring better-call's streaming pattern). ## Verification - `cd server && pnpm vitest run src/__tests__/api-compression.test.ts` — 10/10 passing (7 pre-existing + 3 new regression tests). - Manual: with the fix, browser sign-in against a dev instance succeeds for gzip-accepting clients; before the fix the same request returned `net::ERR_EMPTY_RESPONSE`. ## Risks - Low risk. The middleware still compresses large text/JSON responses exactly as before; the changes only affect paths that previously produced corrupted or destroyed responses. - Behavioral shift: responses whose headers were already committed are now delivered uncompressed instead of being (incorrectly) buffered — this is strictly less surprising than the previous corrupted output. - Failure-path shift: a compression error now yields an uncompressed 200 response instead of a dropped connection. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking enabled, running via Claude Code / Paperclip agent harness with tool use (shell, file edit, test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>canary/v2026.710.0-canary.11 |
||
|
|
1e8ede4e1e |
fix(adapter-utils): runChildProcess escalates to SIGKILL on liveness, not child.killed (#8598)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Those agents run as child processes spawned by
`@paperclipai/adapter-utils`'s `runChildProcess`, which arms a
parent-side wall-clock timer at `timeoutSec` to bound a hung run
> - At the deadline `runChildProcess` sends SIGTERM, then after a grace
window escalates to SIGKILL — the SIGKILL backstop is what turns a
wedged child into a dead PID so the scheduler can reclaim and retry it
> - On the **direct-child fallback** path (`signalRunningProcess`, used
on win32 and whenever process-group signaling is unavailable or throws)
the escalation was gated on `!child.killed`
> - But Node sets `ChildProcess.killed` to `true` the instant a signal
is *successfully sent*, not when the process exits — so once the earlier
SIGTERM has been sent, `child.killed` is already `true`, the
`!child.killed` guard is `false`, and the SIGKILL escalation never runs
> - A child that ignores SIGTERM (e.g. a graceful-shutdown handler
wedged on a socket) is therefore never force-killed, outlives its
deadline, and for an unattended scheduler sits running forever with no
terminal state
> - This PR gates the fallback escalation on real liveness (`exitCode
=== null && signalCode === null`), so SIGKILL fires precisely while the
child is still alive
> - The benefit is the hard timeout actually guarantees termination
(except true uninterruptible D-state) on every platform/configuration,
not just where the process-group path is available
## Linked Issues or Issue Description
No existing public issue — describing the bug inline, following
`.github/ISSUE_TEMPLATE/bug_report.yml`:
### What happened?
When `@paperclipai/adapter-utils`'s `runChildProcess` reaches
`timeoutSec` and the spawned child ignores SIGTERM, the SIGKILL
escalation on the **direct-child fallback** path
(`signalRunningProcess`, taken on win32 or whenever `process.kill(-pgid,
…)` is unavailable or throws) never fires, so the child outlives its
deadline indefinitely. Root cause: the escalation is gated on
`!running.child.killed`, and `ChildProcess.killed` reflects only that a
signal was *successfully sent* (per the Node docs it "does not indicate
that the child process has been terminated"). After the deadline
SIGTERM, `child.killed` is already `true`, so `!child.killed` is `false`
and the follow-up SIGKILL is suppressed.
### Expected behavior
After the grace window, a child that is still alive is force-killed with
SIGKILL regardless of whether SIGTERM was already sent — the hard
timeout should guarantee termination (except true uninterruptible
D-state) on every platform/configuration.
### Steps to reproduce
1. Spawn a child that installs a no-op `SIGTERM` handler and never exits
(e.g. `process.on('SIGTERM', () => {}); setInterval(() => {}, 1000)`).
2. Drive it through the direct-child fallback, i.e.
`signalRunningProcess({ child, processGroupId: null }, …)` (the path
used on win32 / when group signaling is unavailable).
3. Send SIGTERM (the child swallows it; `child.killed` becomes `true`),
then send SIGKILL.
4. On the pre-fix `!child.killed` guard the SIGKILL call is a no-op and
the PID survives past its deadline. Covered by the new regression test
in this PR.
### Paperclip version or commit
Reproduces on `master` (the `signalRunningProcess` fallback). Also
present in published `@paperclipai/adapter-utils` (e.g. `2026.325.0`),
where the same `!child.killed` guard sits on the single direct-child
escalation path.
_Searched the open PR list for duplicates/related work on
`runChildProcess` / `signalRunningProcess` / SIGKILL escalation; found
none._
## What Changed
- `packages/adapter-utils/src/server-utils.ts`: in
`signalRunningProcess`, replace the direct-child fallback guard
`!running.child.killed` with `running.child.exitCode === null &&
running.child.signalCode === null` (real liveness). The process-group
path is unchanged.
- `packages/adapter-utils/src/server-utils.ts`: `export`
`signalRunningProcess` so the fallback branch can be unit-tested
directly.
- `packages/adapter-utils/src/server-utils.test.ts`: add a companion
regression test (POSIX-only, like the sibling timeout tests) that forces
the fallback (`processGroupId: null`) — sends SIGTERM (child swallows
it, `child.killed` becomes `true`), asserts the child is still alive,
then sends SIGKILL and asserts the PID dies. Also keeps the end-to-end
`runChildProcess` SIGTERM-ignoring test.
## Verification
```
npx vitest run packages/adapter-utils/src/server-utils.test.ts # 52 passed
npx tsc --noEmit # clean
```
- **Regression proof:** reverting the guard to `!running.child.killed`
makes the new fallback test fail (`waitForPidExit` → false; the child
survives); the liveness guard makes it pass. This addresses the prior
review note that the existing test only exercised the process-group path
(which already escalated correctly on POSIX) and never reached the
changed branch.
## Risks
Low. A one-line guard change scoped to the direct-child fallback; the
process-group path is untouched. SIGKILL is only sent when
`exitCode`/`signalCode` are both still `null`, i.e. the process is
provably alive, so the change cannot signal an already-reaped/recycled
PID. New tests are POSIX-only and `skipIf(win32)`, consistent with the
sibling timeout tests in this file.
## Model Used
Anthropic **Claude Opus 4.8**, driven via the Cursor agent (extended
reasoning + tool use, large context). Diff, tests, and the regression
proof above were produced and run by the agent; reviewed by a human
before pushing.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
|
||
|
|
1fe89eb8f8 |
Enforce durable external-wait liveness (#9373)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat/recovery subsystem decides whether an agent run has a durable continuation path after the process stops. > - External waits need stricter semantics than local background watchers: a killed local process is not durable, while a first-class blocker/monitor/scheduled wake is. > - Without that distinction, recovery can repeatedly treat adapter-failed continuations as live work and obscure the real reason a task stopped. > - This pull request adds explicit durable external-wait liveness handling and documents the expected execution semantics. > - It also improves operator-visible recovery evidence so invalid external-wait paths explain why they were rejected. > - The benefit is clearer recovery behavior, fewer duplicate continuation recoveries, and a safer contract for monitor-backed external waits. ## Linked Issues or Issue Description - Refs #5978 - Related PRs: #4988, #7495, #8502 ## What Changed - Added durable external-wait liveness classification so local/background watchers are not accepted as durable live paths after the owning process exits. - Preserved first-class blocker/monitor/scheduled wake paths as valid external-wait continuations. - Added backend regression coverage for killed watcher failure, monitor-backed durable wait resumption, normal completion, blocker behavior, and no duplicate recovery. - Added adapter utility coverage for terminal cleanup behavior used by local process adapters. - Surfaced invalid external-wait recovery evidence in the recovery action card and run ledger. - Updated execution semantics documentation and the V1 implementation contract. ## Verification - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed. - `node scripts/run-vitest-stable.mjs --mode general --group general-server` equivalent lane passed in CI-clean env: 238 files, 2164 tests passed, 1 skipped. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-a` passed in fully Paperclip-env-clean env: UI 305 files / 2430 tests; CLI 43 files / 230 tests. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-b` passed in fully Paperclip-env-clean env: shared/db/adapters/plugin packages all green. - `node scripts/run-vitest-stable.mjs --mode serialized` passed in fully Paperclip-env-clean env: 107 serialized server suites green, including 84/84 heartbeat-process-recovery tests. - `pnpm build` passed in fully Paperclip-env-clean env. Notes: running `pnpm test:run` directly inside the Paperclip heartbeat environment exposed local harness env contamination in existing tests (`PAPERCLIP_CONFIG`, `PAPERCLIP_DB_BACKUP_DIR`, and `PAPERCLIP_WORKTREE_START_POINT`). Re-running the same lanes with inherited `PAPERCLIP_*` and port env removed produced the CI-equivalent green results above. ## Risks - Medium behavioral risk: this changes recovery classification for stopped local external-wait processes, so adapters relying on unmanaged background watchers must use blockers, monitors, scheduled wakes, or explicit durable handoff instead. - Low UI risk: recovery-card copy changes are covered by component tests and Storybook screenshot QA. - No database migration is included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-enabled terminal/code execution. Exact context-window metadata was not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.710.0-canary.10 |
||
|
|
1f07690184 |
fix(ui): keep issue threads from jumping to latest comment (#9354)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue and item detail pages use a shared issue chat thread to show comments, runs, activity, and interactions. > - That thread still defaulted to landing on the latest comment when messages first loaded. > - On long issue/item pages, that default can yank the operator away from the top of the page before they choose to inspect the newest message. > - Deep links to comment hashes can create the same kind of initial viewport jump when they are used as generic navigation targets. > - This pull request makes initial latest-comment and initial thread-hash scrolling opt-in instead of default behavior. > - The benefit is stable initial page position across issue-thread surfaces while keeping the explicit Jump to latest control available. ## Linked Issues or Issue Description No exact public GitHub issue was found for this bug. Bug description: - What happened: opening a page with a shared issue conversation thread could automatically move the viewport toward the newest comment/thread target. - Expected behavior: ordinary page loads should keep the initial viewport stable unless the user explicitly clicks Jump to latest. - Steps to reproduce: open an issue or item detail page with a long conversation thread and observe whether the page jumps to the newest thread entry on initial load. - Paperclip version/commit: reproduced while working on the current `master` branch lineage. - Deployment mode: local trusted/dev UI. Related public thread/comment UX work: Refs #3916, Refs #7972, Refs #8800. ## What Changed - Changed `IssueChatThread` so initial latest-comment scrolling defaults to off. - Added a separate opt-in for initial thread-hash scrolling, also defaulting to off. - Preserved stale deleted-comment hash cleanup without scrolling the page. - Updated regression coverage so default initial load stays put, comment hashes do not scroll by default, and manual Jump to latest still scrolls. ## Verification - `pnpm --filter @paperclipai/ui typecheck` passed on the clean PR branch. - `pnpm --dir ui exec vitest run src/pages/IssueDetail.test.tsx -t "loads from the pending state into issue detail without changing hook order"` passed on the clean PR branch. - `pnpm --dir ui exec vitest run src/components/IssueChatThread.test.tsx` was attempted on the clean PR branch, but the file fails before changed assertions with the existing `TypeError: act is not a function` test-harness issue across 58 tests; 14 tests passed. - Static check: no `autoScrollToLatestOnInitialLoad={true}` or `autoScrollToHashOnInitialLoad={true}` call sites remain in `ui/src`. ## Risks Low risk. This only changes initial scroll defaults in the shared issue thread. The main behavioral shift is that direct comment/thread hashes no longer auto-scroll on first load unless a caller explicitly opts in; the Jump to latest button and post-submit scroll behavior are unchanged. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding-agent runtime; exact context window not exposed in this environment; tool-enabled repository inspection, editing, testing, git, GitHub CLI, and Paperclip API usage. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.9 |
||
|
|
be1fcb2b46 |
Fix agent sidebar liveness churn (#9358)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board UI includes an agents sidebar so operators can see which agents are currently active. > - That sidebar depends on live-run polling, heartbeat events, and cross-tab cache sharing to stay current without overloading the API. > - The sidebar was visually churning because agents could leave the live section immediately after a run ended, while progress events and cross-tab broadcasts kept forcing hot query updates. > - This pull request stabilizes live sidebar membership and makes shared polling broadcasts monotonic/deduplicated. > - The benefit is a calmer operator sidebar that still reflects real live state without flashing between stale and fresh snapshots. ## Linked Issues or Issue Description No public GitHub issue exists for this operator-facing bug. Bug summary: - What happened: the agents sidebar could flash or reshuffle around active agents while live-run and heartbeat data was updating. - Expected behavior: active and recently-active agents should remain visually stable, and cross-tab cache sharing should not overwrite fresher data with older snapshots. - Reproduction context: run Paperclip with multiple tabs or rapid live-run/progress updates and watch the agents sidebar while agents enter/leave live execution. - Deployment mode: local/operator board UI. Related PR: - Supersedes #9357, which carried the same fixes on a branch/title/body that were not suitable for public contribution hygiene. ## What Changed - Restored the maintainer-only warning wording in the developer skill guide so the existing server skill-utils CI gate passes on current master. - Added a 120-second linger window for streamlined sidebar agent rows so an agent does not immediately disappear from the live section as soon as its last run ends. - Deferred the recent-agent fallback until there are no live or lingering agents, while keeping the live badge tied only to actually-live runs. - Stopped broad live-runs/heartbeats/agents-list invalidation on every run progress event, while preserving targeted agent-detail invalidation. - Added producer timestamps to cross-tab shared polling result messages so older-or-equal snapshots are dropped before `setQueryData`. - Added per-resource broadcast dedupe/rate limiting so tabs do not rebroadcast equivalent cached data in a loop. - Added focused coverage for sidebar linger behavior, staggered multi-agent linger expiry, live update invalidation scope, shared polling timestamp handling, and cross-tab broadcast dedupe. ## Verification Run locally on the rebased PR branch: - `pnpm --filter @paperclipai/ui exec vitest run src/components/SidebarAgents.test.tsx src/context/LiveUpdatesProvider.test.ts` — 44 tests passed. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/cross-tab-poll.test.ts src/hooks/useSharedPolling.test.ts` — 10 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — exit code 0. ## Risks - Low migration risk: the sidebar/polling changes are UI/client cache behavior only, with no database or API contract changes. - Sidebar visibility now intentionally lingers for 120 seconds after the last live run; stale rows could remain briefly visible, but their live badge is removed when they are no longer actually live. - Cross-tab broadcasts are now more conservative; a missed publish should be corrected by the next normal poll or accepted newer timestamp. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex via the Paperclip Codex coding-agent runtime; exact API model identifier and context-window size are not exposed in this environment. The agent used terminal/tool execution for repository inspection, focused tests, branch preparation, and PR creation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.8 |
||
|
|
e84731af70 |
Revert "Fix default model adapter test config" (#9363)
Reverts paperclipai/paperclip#9361 |
||
|
|
ebd62ca5ae |
Fix default model adapter test config (#9361)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board UI lets operators create and edit agent adapter configuration, including a primary model field and an adapter test action. > - For Cody and similar adapter forms, selecting the default model means the model value is intentionally unset so the adapter can use its default. > - The adapter test path still allowed `model: undefined` to survive in the generated adapter config, which could send an invalid test payload instead of omitting the field. > - This pull request normalizes create/edit adapter test config so default-model selections omit `model` entirely. > - The benefit is that testing an agent configured to use the adapter default model exercises the same clean config shape that should be saved and run. ## Linked Issues or Issue Description No public GitHub issue was found for this local UI bug, so the problem is described inline. Bug description: - What happened: using the adapter test action after choosing the default model could include `model: undefined` in adapter config and surface a UI/runtime error instead of testing with the adapter default. - Expected behavior: choosing the default model should omit the `model` field from adapter config so the adapter default is used. - Steps to reproduce: edit a Codex/Cody-style agent with a concrete model, switch the model selector to Default, then run the adapter Test action. - Paperclip version/commit: current `master` before this PR. - Deployment mode: board UI, deployment-mode independent. ## What Changed - Sanitized adapter test config assembly so undefined adapter config entries are omitted before the test request is sent. - Made create-mode current model display resilient when the model is unset for adapter defaults. - Added regression coverage for editing an existing agent from a concrete model back to Default and testing it. - Added regression coverage for create-mode testing with an unset/default model. - Hardened the developer skill wording used by the existing server skill utility contract test. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/AgentConfigForm.render.test.tsx` - `pnpm exec vitest run server/src/__tests__/paperclip-skill-utils.test.ts` - GitHub PR checks on this branch are green, including Typecheck + Release Registry, Build, General tests, e2e, verify, security scans, and Greptile Review. ## Risks - Low risk: this only removes undefined values from adapter test config payloads, which aligns with the existing persisted patch behavior. - Low risk: default-model display now treats unset create-mode model values as an empty string. - No database, API schema, migration, auth, or adapter runtime contract changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected - check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-enabled software-engineering session. Exact context window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <cody@paperclip.ing>canary/v2026.710.0-canary.7 |
||
|
|
a4993a72a6 |
Fix live run streaming text readability (#9330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue thread UI renders live agent output from adapter run logs and transcript parsing. > - Some adapter streams emit many small or repeated token chunks, and live UI updates can expose partial words, duplicated slices, or transient markdown placeholders. > - That makes active run updates look like gibberish even when the underlying agent output is valid. > - The fix needs to preserve raw logs while making the live thread view stable, readable, and ordered. > - This pull request adds monotonic run-log sequencing, safer live transcript dedupe/order handling, markdown placeholder hiding, and readable live text stabilization. > - The benefit is a live issue thread that updates smoothly without showing confusing partial parser artifacts. ## Linked Issues or Issue Description No public GitHub issue exists yet, so this PR includes the bug details inline. ### What happened? Live run updates in the issue thread can show confusing repeated or partial text while an adapter is streaming. The visible text appears to lose parsing boundaries during active updates, especially with ACP-style token deltas, so the live output can briefly render duplicated chunks, incomplete words, or HTML-comment placeholders. ### Expected behavior Live text should remain readable while preserving the underlying run output for raw inspection. ### Steps to reproduce 1. Start a live agent run whose adapter emits small stdout token deltas. 2. Watch the issue thread while the run is still active. 3. Observe transient duplicated chunks, incomplete words, or markdown placeholder artifacts in the live rendered text. ### Paperclip version or commit Reproduced against current `master` before this PR branch. ### Deployment mode Local dev issue-thread UI with live local adapter runs. ### Additional context GitHub PR search for `live run streaming text markdown transcript` found one broad merged PR, `#252` (“Dotta updates - sorry it's so large”), but no targeted duplicate for this live streaming readability bug. ## What Changed - Added per-run monotonic sequence numbers to persisted and live run-log chunks. - Dedupe and order live transcript chunks by sequence before falling back to timestamp ordering. - Hide markdown HTML comment placeholder text from rendered markdown output. - Smooth live issue-thread text updates so partial additions reveal at readable word boundaries and sliding-window removals do not produce gibberish. - Added coverage for run-log ordering/deduping, markdown comment hiding, live issue-thread stabilization, and Greptile-reviewed edge cases where overlap rewrites could synthesize text or no-boundary additions could stay hidden. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts src/components/MarkdownBody.test.tsx src/components/transcript/useLiveRunTranscripts.test.tsx` passed before the review fix: 3 files, 86 tests. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts` passed after the review fix: 1 file, 30 tests. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts src/components/transcript/useLiveRunTranscripts.test.tsx` passed after the final Greptile overlap fix: 2 files, 40 tests. - `pnpm check:token-gates` passed. - Local PII/secret scan of touched files found only expected code/test words such as `secret`, `token`, and redaction-related strings; no literal credentials found. - `pnpm -r typecheck` passed after restoring declared dependencies with `CI=1 pnpm install --frozen-lockfile` and running with a short `TMPDIR` because `tsx` IPC sockets fail under the long sandbox temp path. - `pnpm build` passed with existing Vite CSS/font/chunk warnings. - GitHub PR checks passed on head `4c052dfe86aecb5feb73504e6b48843f68fce813`: build, typecheck/release registry, server and workspace test shards, serialized server suites, e2e, canary dry run, policy, review, Socket, Superagent, Snyk, and verify. - Greptile review passed on head `4c052dfe86aecb5feb73504e6b48843f68fce813` with confidence score 5/5 and no blocking issues found. - `pnpm test:run` failed in unrelated server workspace tests on this macOS local environment: - `server/src/__tests__/heartbeat-workspace-branch-containment.test.ts`: two assertions compare `/tmp/...` with `/private/tmp/...`. - `server/src/__tests__/heartbeat-worktree-suppression.test.ts`: expected one heartbeat run but observed two, followed by cleanup fallout in the full run. - Isolated rerun of those two server suites reproduced the same three failures. ## Risks - Low product risk for the UI changes: the readable smoothing only affects active live-run display stabilization, not stored comments or raw run logs. - Moderate verification risk: local full Vitest did not pass because of unrelated server workspace tests. Targeted tests for this change, typecheck, token gates, and build passed. - Run-log sequence fields are optional for compatibility with older log rows that do not include `seq`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-using local workspace execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
991279f52c |
Fix Skill Studio markdown dirty tracking (#9356)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Skill Studio is the UI surface for editing skill package files, including `SKILL.md`. > - Markdown files are split into frontmatter fields plus a rich markdown body editor. > - The dirty-state guard intentionally ignores markdown editor normalization during initial mount. > - That guard was too narrow: rich markdown body edits could happen before the file was marked as user-interacted, so edits did not reliably enable Save. > - This pull request broadens the user-interaction signals around the markdown body editor and adds a regression test for saving body edits. > - The benefit is that editing `SKILL.md` in Skill Studio now behaves like normal file editing: changes show as unsaved and Save persists the full markdown document. ## Linked Issues or Issue Description No public GitHub issue exists. Inline bug description follows. ### What happened? Editing a Skill Studio markdown body did not reliably mark the file dirty, so the Save action could remain unavailable or fail to persist the body edit. ### Expected behavior User edits in the markdown body editor should mark the file unsaved and allow saving the updated `SKILL.md` content. ### Steps to reproduce 1. Open a Skill Studio markdown file such as `SKILL.md`. 2. Edit the markdown body in the rich editor. 3. Observe whether the unsaved state appears and Save becomes enabled. 4. Save and reload the file. ### Paperclip version or commit Reproduced on `master` before this fix. ### Deployment mode Local dev (`pnpm dev`). ## What Changed - Mark markdown body interaction on capture-phase key, pointer, paste, drop, before-input, and input events around the rich editor. - Preserve the existing guard that prevents MDXEditor mount-time normalization from dirtying a clean file. - Add a Skill Studio regression test that edits the markdown body, observes the Unsaved state, enables Save, and verifies the saved `SKILL.md` includes both frontmatter and the edited body. ## Verification - `pnpm vitest run ui/src/pages/SkillStudio.test.tsx` Manual reviewer path: - Open a Skill Studio markdown file such as `SKILL.md`. - Edit the body text in the rich markdown editor. - Confirm the UI shows an unsaved state and the Save button is enabled. - Save and confirm the updated markdown body persists. ## Risks Low risk. The change only broadens interaction detection before applying existing dirty-state logic, and the guard still prevents initial editor normalization from marking an unopened file dirty. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex CLI using GPT-5, with repository file editing, shell execution, and GitHub CLI tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a02fe8d575 |
Update Codex adapter GPT-5.6 defaults (#9352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Codex local is the adapter subsystem that exposes OpenAI Codex CLI model choices to agents and issue overrides. > - OpenAI has GPT-5.6 Codex-capable models that should appear in Paperclip's built-in Codex model list and refresh behavior. > - Paperclip's server model listing falls back to the adapter metadata and merges OpenAI refresh results with known Codex defaults. > - This pull request updates the Codex default model metadata to include GPT-5.6 options and adds regression coverage for fallback and refresh paths. > - The benefit is that operators can select the new Codex models without relying on manual model IDs, and refresh behavior keeps known GPT-5.6 options visible. ## Linked Issues or Issue Description Refs #9322. Refs #9342. Refs #9346. ### Agent or provider Codex CLI (OpenAI). ### Why this adapter is useful OpenAI's GPT-5.6 Codex-capable models should be available in Paperclip's Codex adapter defaults and model refresh path. ### How the agent is invoked `codex` ## What Changed - Changed the `codex_local` default model metadata from `gpt-5.5` to `gpt-5.6`. - Added `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna` to the built-in Codex adapter model list. - Updated adapter and server model-listing tests to cover GPT-5.6 fallback and refresh behavior. - Aligned Codex Fast mode support and helper text with the new `gpt-5.6` default, while preserving GPT-5.5, GPT-5.4, and manual model ID support. ## Verification - `git diff --check origin/master...HEAD` - `pnpm exec vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts server/src/__tests__/adapter-models.test.ts server/src/__tests__/adapter-model-refresh-routes.test.ts` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` ## Risks Medium risk because changing `DEFAULT_CODEX_LOCAL_MODEL` from `gpt-5.5` to `gpt-5.6` changes the adapter's default model selection for new blank configurations. The model-list additions are otherwise low risk and covered by adapter/server metadata tests. This PR intentionally overlaps related PRs #9342 and #9346, so reviewers may prefer to close or fold it into one of those branches. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent based on GPT-5, with shell, git, GitHub CLI, and repository editing tool use. Exact served model ID and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
953b315dfb |
Shorten skill frontmatter descriptions (#9353)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip agents can load repository and catalog skills, and Codex renders skill names and frontmatter descriptions into startup context. > - Long descriptions consume the fixed skill metadata budget before Codex can use the progressively disclosed skill bodies. > - The repo `.agents/skills` descriptions and a few shipped catalog descriptions had grown into operational documentation instead of short trigger metadata. > - This pull request keeps the strongest trigger language in frontmatter while leaving detailed procedures in each skill body. > - The benefit is lower prompt overhead, more reliable skill triggering, and a regression guard that prevents description drift from returning. ## Linked Issues or Issue Description No public GitHub issue found for this maintenance item. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am working against `master`. - [x] I have confirmed the issue originates in Paperclip's shipped skill metadata, not in a local agent adapter or provider. ### What happened? Codex startup renders discovered skill names and frontmatter descriptions into a fixed skill metadata budget. Several repository skill descriptions and one shipped catalog description had grown into long-form operational guidance, which can force Codex to truncate descriptions before the model has enough trigger signal to select the right skill. ### Expected behavior Skill frontmatter descriptions should stay short trigger summaries: one capability sentence plus a “use when” clause. Detailed procedures should stay in the skill body and load only after the skill triggers. ### Steps to reproduce 1. Inspect `.agents/skills/*/SKILL.md` and `packages/skills-catalog/catalog/**/SKILL.md` frontmatter descriptions. 2. Measure folded YAML `description` values. 3. Observe descriptions above the intended short-trigger range, including descriptions above 300 characters. 4. Run the new shipped catalog test to verify future descriptions stay capped. ### Paperclip version or commit Reproduced on `master` at `cc81eefb6047d8eaf57faf785f421c03dc97073c`. ### Deployment mode Local dev / source checkout metadata inspection. This is not database-related. ### Installation method Built from source. ### Agent adapter(s) involved Codex, because Codex startup uses the skill metadata prompt budget. The metadata source itself is core repository/catalog content. ### Database mode Not database-related. ### Access context Not applicable; this is static repository metadata. ### Node.js version `v22.22.2` in the verification environment. ### Operating system Linux container environment. ### Relevant logs or output Final measurement after this PR: 29 source `SKILL.md` files, max description length 215 chars, 5,449 total description chars, estimated 1,363 description tokens at 4 chars/token. ### Relevant config None. ### Additional context The shipped catalog manifest was regenerated so the generated package metadata matches the edited catalog `SKILL.md` sources. ### Privacy checklist - [x] I have reviewed all pasted output for PII and redacted where necessary. ## What Changed - Shortened long `.agents/skills/*/SKILL.md` frontmatter descriptions to concise capability plus use-when trigger clauses. - Shortened the over-budget shipped skills catalog descriptions for wireframe, Paperclip capsules, and reflection coach. - Regenerated `packages/skills-catalog/generated/catalog.json` so shipped metadata matches source skill frontmatter. - Added a Vitest regression guard that caps repo skill source descriptions and generated catalog descriptions at 300 characters. ## Verification - `pnpm --filter @paperclipai/skills-catalog build:manifest` - `pnpm --filter @paperclipai/skills-catalog test` — 5 files passed, 19 tests passed - `pnpm --filter @paperclipai/skills-catalog typecheck` - Final measurement: 29 source `SKILL.md` files, max description length 215 chars, 5,449 total description chars, estimated 1,363 description tokens at 4 chars/token. Note: the clean PR worktree was created from `origin/master` and contains only this commit, but it does not have `node_modules`; running `pnpm --filter @paperclipai/skills-catalog test` there failed at tool/package resolution (`vitest`, `tsc`, `@paperclipai/shared`). The dependency-equipped workspace passed the commands above before the commit was cherry-picked onto the clean branch. ## Risks Low risk. This changes skill metadata and tests only. The main risk is over-trimming a useful trigger phrase, mitigated by keeping explicit “use when” clauses and leaving detailed guidance in the skill bodies. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent via Paperclip/Codex, with shell and file-edit tool use. Exact API model ID and context window were not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f5f461729 |
Avoid startup crash when Reflection Coach assets are missing (#9351)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server has built-in agent definitions that are loaded during startup and used to provision optional operational agents such as Reflection Coach > - Reflection Coach stores richer stock instructions, a routine description, and a bundled skill as markdown assets outside the TypeScript module body > - A deployed server can fail before it is healthy if one of those copied markdown assets is absent from `server/dist` > - A recent build fix preserves those assets during normal server builds, but runtime should still degrade gracefully if a packaged asset is missing or unreadable > - This pull request adds resilient loading for built-in Reflection Coach assets and keeps a minimal compiled fallback available > - The benefit is that a missing optional built-in agent file no longer turns into a process-wide startup crash ## Linked Issues or Issue Description Bug fix. No public GitHub issue found for this exact startup crash. - Related public PR: #9339 - What happened: the server could throw `ENOENT` while importing the built-in agent service if `server/dist/built-ins/agents/reflection-coach/AGENTS.md` was missing from a deployed build. - Expected behavior: the server should keep starting, log that the built-in asset was missing, and use safe fallback text for the optional built-in agent resource. - Steps to reproduce: build the server, remove the compiled Reflection Coach `AGENTS.md` asset from `server/dist`, then import/start the server path that loads built-in agent definitions. - Paperclip version/commit: reproduced against a deployed build containing the Reflection Coach built-in agent assets; fixed against current `master` after #9339. - Deployment mode: Node server deployment using compiled `server/dist` output. ## What Changed - Added built-in agent text loading that checks the compiled asset path first, then source/package fallback paths, then a minimal compiled-in fallback string. - Added fallback text for Reflection Coach instructions, routine description, and bundled skill content so startup does not depend on optional markdown assets being present. - Added regression coverage for readable candidate selection and missing-file fallback behavior. ## Verification - `pnpm -w exec vitest run server/src/__tests__/built-in-agents.test.ts` — 1 test file passed, 24 tests passed. - `pnpm --filter @paperclipai/server build` — server TypeScript build completed and copied `src/built-ins` into `dist/built-ins`. - Manual smoke: temporarily moved `server/dist/built-ins/agents/reflection-coach/AGENTS.md`, imported `server/dist/services/built-in-agents.js` through the repo-pinned `tsx` runtime, and confirmed Reflection Coach definitions still loaded with output `reflection-coach:3732`; the asset was restored afterward. ## Risks Low risk. The normal path still uses the full packaged markdown assets. The fallback path is only used when those files are missing or unreadable, and it logs a warning so packaging drift remains visible. ## Model Used OpenAI GPT-5 Codex coding agent, with repository tool access and shell-based verification. Exact context window was not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.6 |
||
|
|
cc81eefb60 |
Make plan-approval continuations durable after failed wakes (#9331)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents often work from reviewed plans that are approved through issue-thread interactions. > - Accepting a plan is not just a UI decision; it must reliably resume the assignee so approved work continues. > - A failed continuation wake could leave an approved plan stranded in review with no durable retry or visible recovery path. > - This pull request makes approved plan continuations retryable, recoverable, and visible when resume fails. > - The benefit is that operators can trust plan approval to either resume the agent or produce an explicit actionable failure instead of silent limbo. ## Linked Issues or Issue Description No public GitHub issue exists. Inline bug report follows the repository bug template. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest released version of Paperclip or can reproduce on `master`. - [x] I have confirmed the error originates in Paperclip itself, not in my agent adapter, API provider, or local configuration. ### What happened? When a plan-confirmation interaction was accepted, the assignee continuation wake could fail before useful agent execution. In that case the issue could remain in review even though the plan had been approved, because the failed wake was fire-and-forget and there was no durable retry or recovery path for accepted continuations. ### Expected behavior Accepted plan continuations should either wake the assignee successfully, retry bounded infrastructure failures, recover dropped wakes, or surface an explicit failure state that operators can act on. ### Steps to reproduce 1. Create an issue with an assignee and a plan confirmation that wakes the assignee on accept. 2. Accept the confirmation. 3. Simulate a pre-flight continuation failure, such as process loss before agent start or workspace validation failure. 4. Observe that the approved issue can remain in review without an active assignee wake or visible retry/failure state. ### Paperclip version or commit Reproducible on `master` before this PR's retry/recovery changes. ### Deployment mode Local dev (`pnpm dev`) and server-side recovery paths. ### Installation method Built from source (`pnpm install`, `pnpm dev`, test runner). ### Agent adapter(s) involved Not adapter-specific; this is a core continuation/recovery bug. The tests cover local-agent failure shapes without relying on a provider-specific API. ### Database mode Embedded Postgres test database for verification. The affected logic is database-backed and applies to normal Postgres deployments as well. ### Access context Board accepts the interaction; agent execution resumes through the assignee wake path. ### Relevant logs or output No sensitive logs are needed. The regression tests simulate the failed wake and recovery states directly. ### Relevant config No special config is required beyond an assignee with wake-on-demand enabled. ### Additional context This PR also prevents a stale workspace-validation payload from quarantining another issue's active workspace and prevents unrelated successful runs from masking a continuation that never resumed. ### Privacy checklist - [x] I have reviewed all pasted output for PII, usernames, file paths, API keys, tokens, and company names, and redacted where necessary. ## What Changed - Added bounded infrastructure retries for failed accepted-interaction continuation wakes. - Extended stranded issue recovery so dropped accepted-plan continuation wakes are requeued. - Recorded and rendered explicit resume-failure state on accepted confirmation cards. - Added clean-workspace fallback for workspace-validation failures while preventing cross-issue workspace quarantine. - Tightened recovery so unrelated successful runs do not mask an accepted continuation that never resumed. - Added focused server/UI coverage for retry scheduling, recovery, visible failure state, and interaction card rendering. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-retry-scheduling.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts` — 106 tests passed. - Earlier branch verification also covered the issue-thread interaction card tests for the visible resume-failure UI. - GitHub CI is green on the replacement PR head, and Greptile reports 5/5 with no blocking issues. ## Risks - Medium behavioral risk: this changes recovery behavior for accepted continuation interactions and workspace-validation retries. - Mitigation: retries are bounded, scoped to same-company issue context, and workspace quarantine now requires ownership by the issue being retried. - Existing stored confirmation results remain compatible because the new resume-failure field is optional. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-enabled terminal workflow. The runtime does not expose an exact context-window value to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.5 |
||
|
|
d166069bc4 |
Preserve built-in agent assets in server builds (#9339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server package ships compiled runtime code plus static runtime assets. > - Built-in agent definitions live under `server/src/built-ins` and runtime code resolves them relative to compiled server files. > - The server build already copied onboarding assets into `dist`, but it did not copy built-in agent assets alongside the compiled code. > - Packaged server builds could therefore miss built-in agent definitions even though source-based development runs worked. > - This pull request extends the server build copy step to preserve built-in agent assets in `dist/built-ins`. > - The benefit is packaged server builds keep the same built-in agent runtime assets available as source-based development runs. ## Linked Issues or Issue Description No public GitHub issue found. This PR describes the underlying bug inline using the bug report template fields. ### What happened? `@paperclipai/server` build output copied `server/src/onboarding-assets` into `server/dist/onboarding-assets`, but did not copy `server/src/built-ins` into `server/dist/built-ins`. Runtime code for built-in agents resolves those assets relative to the compiled server files, so packaged builds could omit built-in agent markdown assets that are present during source-based development. ### Expected behavior Packaged server builds should include built-in agent assets under `server/dist/built-ins`, matching the runtime location expected by the compiled server code. ### Steps to reproduce 1. Check out current `master` before this PR. 2. Run `pnpm --filter @paperclipai/server build`. 3. Check for `server/dist/built-ins/agents/reflection-coach/AGENTS.md`. 4. Observe that the built-in agent asset is missing from the server build output. ### Paperclip version or commit Reproduces on current `master` before this PR. The fix is verified on commit `2b89984ccb7857f06359bf65c48222f110c7aeff`. ### Deployment mode Build/package artifact behavior. This can affect any deployment mode that runs from the built server package rather than directly from source. ### Installation method Built from source with `pnpm --filter @paperclipai/server build`. ### Agent adapter(s) involved Not adapter-specific. This is a core server packaging bug for built-in agent assets. ### Database mode Not database-related. ### Access context Not applicable. This happens during package build output generation. Related search: - Searched public PRs/issues for `built-ins build copy repo:paperclipai/paperclip`. - Found no directly related open issue. One old closed Hermes adapter PR was not directly related. ## What Changed - Updated the `@paperclipai/server` build script to create `dist/built-ins`. - Added the copy step from `server/src/built-ins` into `server/dist/built-ins` alongside the existing onboarding asset copy. - Added a focused server package build-script test that asserts both onboarding and built-in static runtime asset directories are copied into `dist`. ## Verification - `pnpm exec vitest run server/src/__tests__/server-package-build-script.test.ts` - `pnpm --filter @paperclipai/server build` - `test -f server/dist/built-ins/agents/reflection-coach/AGENTS.md` ## Risks Low risk. This changes only the package build asset copy step and adds focused test coverage. The main risk is build-script portability, but it follows the existing `mkdir -p` and `cp -R` pattern already used for onboarding assets. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent in a tool-enabled Paperclip heartbeat. Exact model ID and context-window size are not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.4 |
||
|
|
05973b2073 |
Enable sandbox environments for Grok local adapter (#9338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local CLI adapters can run against Paperclip-managed execution environments instead of only the host filesystem. > - The environment picker and environment capability API derive sandbox support from shared adapter capability lists. > - The Grok Build adapter is implemented as a local CLI adapter, but it was missing from those shared environment capability lists. > - That made Grok agents look local-only even when sandbox environments were configured. > - This pull request registers `grok_local` in the shared adapter constants and remote-managed environment support path. > - The benefit is that Grok Build agents can select the same local, SSH, and sandbox environment overrides as other local CLI adapters. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Inline bug report follows. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on `master`. - [x] I have confirmed the error originates in Paperclip itself, not in the Grok adapter provider or local configuration. ### What happened? When configuring a Grok Build local agent in the board UI, Paperclip did not expose configured sandbox environments as selectable environment overrides. The shared environment capability helper treated `grok_local` as local-only because it was missing from the remote-managed local adapter allowlist. ### Expected behavior Grok Build should behave like other local CLI adapters: when environments are enabled and a runnable sandbox environment exists, the agent configuration form should show the environment override selector and allow the sandbox to be selected. ### Steps to reproduce 1. Enable environments in instance experimental settings. 2. Configure at least one runnable sandbox environment. 3. Open the agent configuration form for a Grok Build local agent. 4. Observe that the sandbox environment is not offered as an override before this fix. ### Paperclip version or commit `master` before this PR. ### Deployment mode Local dev (`pnpm dev`). ### Installation method Built from source (`pnpm dev` / `pnpm build`). ### Agent adapter(s) involved - Grok Build local adapter. - Core bug in shared environment capability logic. ### Database mode Not database-related. ### Access context Board human operator. ### Relevant logs or output No runtime error is emitted; the issue is a missing UI option caused by shared capability metadata. ### Relevant config No secret-bearing config required. Reproduction only needs environments enabled and a runnable sandbox environment configured. ### Privacy checklist - [x] I have reviewed all pasted output for PII, usernames, file paths, API keys, tokens, company names, and redacted where necessary. ## What Changed - Added `grok_local` to the shared built-in adapter type list. - Added `grok_local` to the remote-managed adapter set used by environment capability helpers. - Added shared regression coverage for Grok local sandbox provider and driver support. - Added a UI render regression test that confirms Grok Build agents show the environment override when a runnable sandbox exists. ## Verification - `git diff --check` - Changed-file secret scan with `rg` for common token/key patterns. - `pnpm --filter @paperclipai/shared exec vitest run src/environment-support.test.ts` - `pnpm --dir ui exec vitest run src/components/AgentConfigForm.render.test.tsx` ## Risks - Low risk. This expands environment support for an existing local adapter to match the local CLI adapter behavior already used by Claude, Codex, Gemini, OpenCode, Cursor, and Pi. - Operators still need at least one configured runnable sandbox environment before a Grok agent has a sandbox option to select. - This PR was created from a Paperclip execution workspace branch whose name is runtime-provided; the PR body intentionally avoids internal issue identifiers or instance-local links. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 coding agent, tool-enabled terminal/code execution environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
5c85ae64a0 |
Cases: experimental first-class case object (#9198)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board currently uses issues for execution, but longer-lived content work needs a separate object that can survive beyond a single task thread. > - The Cases subsystem adds an experimental, company-scoped record for content artifacts and their supporting metadata. > - The backend needs durable storage, API routes, revision history, issue linkage, and company-boundary enforcement before the UI can depend on Cases. > - The UI needs an opt-in navigation surface, list/detail views, reference chips, and issue-page context so operators can inspect Cases without making them the default workflow. > - The agent-facing skills need a contract for creating and updating Cases so automated content workflows can dogfood the feature. > - This pull request ships that experimental end-to-end path behind the `enableCases` flag. > - The benefit is a first-class place to collect content work, references, attachments, revisions, and related execution threads without polluting the core issue model. ## Linked Issues or Issue Description No public GitHub issue exists for this experimental feature. Feature request fields: ### Problem Content-oriented work such as release notes, announcements, docs, and campaigns can span many execution issues, which makes the final artifact hard to find and reason about after the execution thread moves on. ### Proposed solution Add an experimental Cases object that is company-scoped, linked to issues, queryable through the API, inspectable in the board UI, and writable by agent workflows through documented conventions. ### Alternatives considered Continue encoding content artifacts directly in issues or documents only. That keeps the data model smaller, but it does not give operators a stable artifact-centric view or a clean way to link related execution history. ### Roadmap alignment Checked `ROADMAP.md`; this PR does not duplicate an existing planned core roadmap item. ## What Changed - Added the `cases` data model, migration, schema exports, and experimental `enableCases` instance setting. - Added company-scoped Cases API routes for list/detail/update, issue links, revisions, children, activity events, annotations, attachments, and idempotent agent-oriented upserts. - Scoped case and issue lookup helpers before access checks so inaccessible cross-company identifiers resolve as not found rather than leaking existence. - Fixed case PATCH timestamp handling so non-status updates cannot overwrite `completedAt` from a stale pre-transaction row snapshot. - Moved Cases list type/status/project filters into the server request before the server-side limit is applied, including multi-select filters and no-project filtering. - Added backend route coverage for creation, updates, idempotency, issue linking, attribution, company-boundary enforcement, OpenAPI registration, list filtering, timestamp patch behavior, and inaccessible lookup regressions. - Added the experimental Cases UI surface: sidebar entry, gated routes, list filters/grouping, detail overview, activity, revisions, children, attachments, and issue-page case rail. - Added case reference rendering and company-prefixed case href generation so case links resolve directly inside the active company route. - Added Paperclip skill documentation for agent workflows that create or update Cases. - Wired release-content skills to emit Cases for dogfooding. - Rebased onto current `master` and renumbered the Cases migrations to `0143`/`0144` after the latest upstream migration sequence. ## Verification - Current PR head: `ecc13be0d`. - Rebased on current `master` (`606aa4f266`) and pushed to the existing PR branch. - `git diff --check origin/master...HEAD` — passed before the first update push; subsequent committed diffs were also checked with `git diff --check` before commit. - Guardrails checked: no `pnpm-lock.yaml` changes, no `.github/workflows` changes, and changed-file count is below the Greptile 100-file limit. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/cases-routes.test.ts src/__tests__/instance-settings-service.test.ts src/__tests__/openapi-routes.test.ts` — passed, 3 files / 26 tests before review-fix commits. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/cases-routes.test.ts` — passed after each server-side Greptile fix, latest 1 file / 15 tests. - `pnpm --filter @paperclipai/server typecheck` — passed after the timestamp and lookup fixes. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Cases.test.tsx src/pages/CaseDetail.test.tsx src/pages/CompanySkills.test.tsx src/App.cases-routing.test.tsx` — passed, 4 files / 30 tests before review-fix commits. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Cases.test.tsx` — passed after the list-filter fix, 1 file / 12 tests. - `pnpm --filter @paperclipai/ui typecheck` — passed after the list-filter fix. - `pnpm check:token-gates` — passed after UI changes. - Remote PR checks on head `ecc13be0d` are green: Paperclip CI, build, typecheck, test matrix, e2e, Canary Dry Run, policy, commit review, Superagent Security Scan, Socket, Snyk, and Greptile passed; Storybook visual regression is skipped and security-review is neutral. - Greptile Review: 5/5 confidence, zero unresolved Greptile threads. ## Risks - Medium feature risk because this introduces a new experimental domain object across database, server, shared contracts, skills, and UI. - The feature is gated behind `enableCases`, which limits default operator exposure while the model is exercised. - Case links now prefer company-prefixed hrefs; the unprefixed redirect remains for externally entered URLs. - Cases list filtering now sends multi-select filters to the server before limiting; the UI still applies the same local filters as a second pass for ancestor/context rows. - Migrations were renumbered on top of current master; the SQL uses guarded `IF NOT EXISTS` / `ADD COLUMN IF NOT EXISTS` patterns where relevant for safer replay. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex in the Paperclip local coding environment was used for this PR curation, rebase verification, review-fix implementation, push, and PR description update. The runtime exposes tool use and shell execution; context-window size is not exposed by this Paperclip adapter. Several implementation commits also include AI co-author trailers recorded in git history. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.710.0-canary.3 |
||
|
|
9a1d4b7983 |
fix(ui): use prose editor for markdown agent instructions (#9332)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent setup depends on instruction files that are readable and editable from the board UI > - The instructions tab already receives server-side metadata describing whether each file is Markdown > - The UI was deciding Markdown editor usage primarily from the file extension, which makes extensionless Markdown instruction files feel like raw code > - This pull request makes the instructions editor trust server Markdown metadata for existing files and keep extension fallback only for new unsaved files > - The benefit is that AGENTS-style prose instructions render and edit like prose while explicitly non-Markdown files still use the raw textarea ## Linked Issues or Issue Description - Refs #8201 - Refs #5652 - Refs #3427 - Refs #2068 - Related PRs: #2468, #2620 ## What Changed - Use server `markdown` metadata from instruction file details/summaries to choose the prose Markdown editor for existing instruction files. - Keep extension-based Markdown detection only for pending new files before server metadata exists. - Remove the monospace content styling from the Markdown editor path so prose instructions read like normal text. - Add focused tests for extensionless Markdown files, new `.md` files, and `.md` files explicitly marked non-Markdown by the server. ## Verification - `pnpm check:token-gates` - `pnpm exec vitest run ui/src/pages/AgentDetail.instructions.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` ## Risks - Low risk. The editor selection now depends on server metadata for existing files, so incorrect server metadata would choose the wrong editor. The fallback still preserves extension-based behavior for newly created unsaved files. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 via Codex coding agent, tool-enabled terminal workflow. Context window details were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.2 |
||
|
|
606aa4f266 |
feat(search): filters, sorting, operators & command-palette parity (#9327)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company search is the primary way operators find issues, comments, documents, artifacts, agents, and projects across a busy company > - Search previously supported only a bare text query: no way to narrow by status/assignee/project/label/date, no sort control, no typed operators, and weak relevance/snippets meant hunting through noise > - As companies accumulate tens of thousands of items, unfiltered single-sort search stops scaling for day-to-day operator workflows > - This pull request adds a full filtering model (filter bar, chips, mobile sheet, URL state), sort modes, typed query operators (`status:`, `assignee:`, `type:`, …) with command-palette parity, relevance/snippet/deep-link improvements, zero-results recovery, and the supporting shared validators, backend service work, and DB indexes > - The benefit is that operators can go from a vague query to the exact item in a couple of keystrokes, on desktop and mobile, with shareable filtered-search URLs ## Linked Issues or Issue Description No existing public GitHub issue; describing the underlying feature request inline (per feature_request template): - **Problem:** Company search accepted only a plain text query. Users could not filter results by status, assignee, project, label, or recency; could not change result ordering; and got no guidance when filters emptied the result set. - **Desired solution:** Structured search filters (UI controls + typed query operators + URL parameters), selectable sort modes, better relevance and snippets with exact deep links, and parity between the search page and the command palette. - **Alternatives considered:** Client-side filtering of unfiltered results (does not scale past the fetch limit); a separate "advanced search" page (splits the surface and duplicates state handling). Related (not duplicate) PRs found while searching: #4848 (issue search query planning), #8235 (search rate limiting). ## What Changed - **Shared contract:** new search filter/sort/count/zero-results types and validators in `packages/shared` (`validators/search.ts`, types index). - **Backend:** `server/src/services/company-search.ts` supports issue filters, sort modes, per-filter option counts, snippets, artifact visibility, and zero-results loosen suggestions; single-statement match replaces per-scope scans and predicates are trigram-index compatible (~3.7s → ~350ms on a live 14.8k-hit corpus). - **DB:** migration `0142_company_search_sort_indexes.sql` adds the supporting indexes. - **Search page (`ui/src/pages/Search.tsx`):** filter bar, removable chips, mobile filter sheet with result-count preview, sort menu, URL round-tripping, zero-results recovery UI. - **Query operators (`ui/src/lib/search-query-parser.ts`):** typed operators parsed into filters, operator autocomplete, filter pills. - **Command palette:** operator-aware parsing and full-search handoff. - **Stale-operator fix (latest commit):** typed operator filters are no longer folded into persistent URL-filter state, so deleting a token (e.g. removing `status:blocked` from the input) actually removes the filter from subsequent requests; filter-control edits materialize control state and strip typed tokens so a removed chip cannot resurrect from the input. ## Verification - `cd ui && npx vitest run src/pages/Search.test.tsx` — 19 tests including two new red→green regressions for the stale-operator paths (both fail on the previous commit, pass now). - `cd ui && npx vitest run src/components/CommandPalette.test.tsx` and `cd server && npx vitest run src/services/company-search-service.test.ts` — operator parity and backend filter/sort/count coverage. - `cd ui && npx tsc --noEmit` — clean. - Manual: open `/search`, type `auth status:blocked`, confirm the status filter applies; delete `status:blocked`, confirm results are unfiltered again; drive the same filters from the filter bar/chips/mobile sheet and confirm the URL round-trips (reload/back/forward preserves state). - Full end-to-end QA pass (9/9 acceptance checks) against the wireframes on desktop (1280px) and mobile (390px) with a live API and browser automation. ## Risks - Additive migration (indexes only, no data rewrites) — safe to roll forward; index creation cost is paid once at migrate time. - Search request shape gains optional parameters only; old clients keep working. - Behavioral shift: filter-control edits now strip typed operator tokens from the query text (their values persist as filter state) — deliberate, so removed filters stay removed. - Ranking changes alter result ordering for existing queries; covered by service tests and the QA pass. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic, extended thinking + tool use) — stale-operator-filter fix, regression tests, PR preparation. - GPT-5 Codex (`codex_local` adapter) and Claude Opus 4.6 (`claude-opus-4-6`) — earlier implementation phases (backend contract, filter UI, operators, ranking) under agent orchestration. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details — pre-existing branch name retained to avoid closing/reopening the PR - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing docs affected) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending re-run on latest commit) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending re-review of the stale-filter fix) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.710.0-canary.1 |
||
|
|
cec0fc249a |
[codex] Parallelize release verify workflow (#9168)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Releases publish the same app and package set that operators install, so release verification should keep full release-strength coverage. > - The release workflow currently verifies stable and canary releases with one serial job that typechecks, runs all tests, and builds. > - The PR workflow already proves the test surface can be split into grouped general suites and serialized shards without changing coverage. > - This pull request extracts the release verify work into a reusable workflow and fans out the independent lanes. > - The benefit is faster stable and canary release verification while preserving the existing publish and preview gates. ## Linked Issues or Issue Description No public GitHub issue exists for this CI improvement. **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Release verification spends most of its wall time in a single serial test step even though the same stable test surface is already partitioned for PR CI. Stable dispatches and master-push canaries therefore wait on one long runner after setup, typecheck, tests, and build run sequentially. **Proposed solution** Add a reusable release verification workflow with parallel typecheck, grouped general tests, serialized test shards, and build lanes. Have both stable and canary release verification call it with the ref they need to verify. **Alternatives considered** Keeping the serial `pnpm test:run` job preserves the old shape but keeps stable and canary releases waiting on one long runner. Skipping verification when a source SHA already has green CI would be faster, but adds stale-check and lookup risk beyond this change. **Roadmap alignment** No overlapping item found in `ROADMAP.md`; this is release CI maintenance. **Additional context** The new workflow keeps the release-strength full `pnpm -r typecheck`, uses the existing stable test grouping/sharding entry points, and leaves publish/preview jobs unchanged. ## What Changed - Added `.github/workflows/release-verify.yml` as a `workflow_call` workflow accepting a `ref` input. - Split release verification into parallel `typecheck`, `general_tests`, `serialized_tests`, and `build` jobs with 20-minute lane timeouts. - Mirrored the PR workflow's stable test partition: `general-server` shards 1-3, `general-workspaces-a`, `general-workspaces-b`, and four serialized shards. - Replaced `release.yml` `verify_canary` and `verify_stable` job bodies with calls to the reusable workflow while leaving publish and preview jobs unchanged. - Added a Node test that guards the release workflow delegation and split verify surface. ## Verification - `actionlint 1.7.12 .github/workflows/release.yml .github/workflows/release-verify.yml` - `node ./scripts/release-package-map.mjs check` - `node --test ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check` ## Risks - Release verification now starts more jobs per release event, increasing total runner setup/install minutes. This matches the existing PR CI tradeoff and should reduce release wall time substantially. - The called workflow checks out the requested ref shallowly. That is intentional for verify lanes; publish and preview jobs still retain their existing full-history checkouts. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent in local tool-use mode with shell execution, repository editing, GitHub connector access, and medium reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d3e26a8d02 |
docs: point Paperclip docs links at docs.paperclip.ing (#9300)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The public README files are owned discovery surfaces for users who arrive from GitHub or npm. > - Some documentation links still used the old `https://paperclip.ing/docs` redirect path. > - Redirect hops are worse for users and for SEO because crawlers and readers do not land on the canonical docs host immediately. > - Package metadata should still point at the GitHub repository, because npm package homepages are expected to identify the source/project page. > - This pull request updates only explicit documentation links to the canonical docs subdomain. > - The benefit is a smaller, clearer link sweep with no package homepage metadata change. ## Linked Issues or Issue Description - No public GitHub issue exists for this small documentation maintenance change. - Problem: public README documentation links used a redirecting docs URL instead of the canonical docs host. - Expected behavior: README documentation links should point directly at `https://docs.paperclip.ing`. - Scope: root README, CLI README, and the Hermes adapter README docs reference. - Related search results reviewed: #793, #592, #675, and this PR. No open duplicate PR was found for this README-only canonical docs URL sweep. ## What Changed - Updated the root README Docs navigation link from `https://paperclip.ing/docs` to `https://docs.paperclip.ing`. - Updated the CLI README Docs navigation link from `https://paperclip.ing/docs` to `https://docs.paperclip.ing`. - Updated the Hermes adapter README Paperclip Docs link from `https://paperclip.ing/docs` to `https://docs.paperclip.ing`. - Kept all `package.json` homepage fields pointing at the Paperclip GitHub repository or package-specific GitHub README pages. ## Verification - `git diff --check origin/master...HEAD` - `rg 'https://paperclip\.ing/docs|https://docs\.paperclip\.ing' README.md cli/README.md packages/adapters/hermes/README.md` - `rg '"homepage": "https://docs\.paperclip\.ing"' -g 'package.json'` returned no matches. - Reviewed the final diff against `origin/master`; only the three README files changed. ## Risks - Low risk: this is a documentation-only URL update. - The main review risk is scope creep into npm metadata; that was explicitly avoided by keeping package homepages on GitHub. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5-based Codex coding agent with local shell, git, GitHub CLI, and Paperclip API tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3369c0dab7 |
fix(prompt): render exact branch name with backtick-safe fence in wake branch guard (#9326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run heartbeats inside execution workspaces pinned to a specific git branch; the wake prompt now carries a one-time "stay on this branch" guard (#9319) > - Greptile's final review round on #9319 landed after the PR merged: the guard sanitized the branch name by stripping backticks, which mutates the ref, so the prompt could tell the agent to stay on a branch name that does not exist > - A guard that names the wrong branch defeats its purpose and still leaves the workspace contract breakable > - This pull request keeps the pinned ref name exact and instead escapes it at render time with a backtick fence longer than any backtick run inside the name (standard Markdown inline-code escaping) > - The benefit is the guard always names the real branch while a hostile ref name still cannot close the code span or inject prompt text ## Linked Issues or Issue Description Refs #9319 — follow-up addressing the final Greptile review round that arrived after that PR merged. ## What Changed - `normalizePaperclipWakeExecutionWorkspace` no longer strips backticks from the branch name; it removes only control characters (illegal in git ref names, and the newline route into the prompt), trims, and caps length. - Added a `markdownInlineCode` helper that wraps a value in an inline-code span whose backtick fence is one longer than the longest backtick run in the value, and used it when rendering the branch guard line. - Updated the hostile-branch-name test to assert the exact ref is preserved and fenced, and control characters are removed. ## Verification - `pnpm vitest run packages/adapter-utils/src/server-utils.test.ts` — 58 tests pass, including the updated hostile-branch-name case. - `npx tsc --noEmit` in `packages/adapter-utils` — clean. - Manual: render a wake payload with `branchName: "evil` + backtick + `name"` and confirm the guard line reads ``` `` evil`name `` ``` and the span does not break. ## Risks - Low risk: prompt-rendering-only change; the normalized payload shape is unchanged. Branch names containing backticks (extremely rare) now render exactly instead of mutated. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) with extended thinking and tool use, via Claude Code / Paperclip agent harness. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.710.0-canary.0 |
||
|
|
f3ca4d24bc |
fix: repair dirty/foreign-branch execution worktrees (#9297)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Execution workspaces are the bridge between Paperclip's control plane and a local agent's checked-out repository state. > - When a workspace is restored after a failed or interrupted run, the recorded branch can disagree with the branch currently checked out on disk. > - A clean branch mismatch can be reconciled safely, but a dirty mismatch needs a lossless path that does not discard uncommitted agent work. > - This pull request adds a quarantine-and-restore path that saves dirty work to a rescue branch, restores the recorded branch, and exposes the repair from the board UI and run page. > - The benefit is that operators can recover wedged execution workspaces without losing work or moving another live branch unexpectedly. ## Linked Issues or Issue Description No public GitHub issue exists for this workspace-recovery failure, so this PR includes the bug report inline. **What happened** A git worktree-backed execution workspace could become wedged when Paperclip expected one branch but found a different checked-out branch with dirty tracked or untracked files. The existing safe repair path refused the restore, leaving the source task blocked with no lossless one-click recovery path. **Expected behavior** Paperclip should preserve dirty work before restoring the recorded workspace branch. If another live workspace claims the checked-out branch, or an attached runtime service is active, the repair should refuse with clear operator-facing evidence instead of risking work loss or file contention. **Steps to reproduce** Create a git worktree execution workspace whose persisted branch name differs from the checked-out branch, add dirty tracked or untracked files in that worktree, then trigger workspace validation or use the branch reconcile endpoint. Before this change, the dirty mismatch remained blocked because Paperclip had no quarantine restore mode. **Paperclip version or commit** Observed on the pre-fix workspace-recovery implementation. Verified on this PR head after rebasing onto current `master`. **Deployment mode** Local trusted development/worktree deployments using git worktree execution workspaces and optional workspace runtime services. ## What Changed - Added dirty-worktree quarantine repair that creates a rescue branch, commits dirty tracked and untracked files there, restores the recorded branch, writes audit comments/activity, and preserves the live foreign branch ref. - Added `quarantine_restore` branch reconcile API support, recovery-action resolution, source-task wake behavior, execution-review preservation, claimant refusal, runtime-service refusal, and coverage for the non-transactional git ordering. - Added board UI controls for the repair action in the recovery card plus a compact failed-run workspace recovery surface that uses the same reconcile handlers. - Hardened Greptile follow-up cases by best-effort restoring the recorded branch after a mid-sequence rescue commit failure and by refusing quarantine restore while attached runtime services are active. ## Verification - `pnpm exec vitest run server/src/__tests__/execution-workspaces-service.test.ts -t "quarantine_restore"` - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts -t "workspace dirty quarantine branch repair"` - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts -t "repairs clean unrecorded branch drift|adopts unrecorded forward branch drift"` - `pnpm exec vitest run server/src/__tests__/workspace-runtime-routes-authz.test.ts` - `pnpm --filter @paperclipai/server typecheck` - Earlier PR verification covered the route, service, heartbeat, UI component, and run-page recovery surfaces; Cutter posted public preview screenshots for the repair popover and run-page panel at https://github.com/paperclipai/paperclip/pull/9297#issuecomment-4926934211. - GitHub PR checks are green on `dcac76b05f4cf6e1ee16544c2831d83c7857e475`. - Greptile is 5/5 with zero annotations and no unresolved review threads on `dcac76b05f4cf6e1ee16544c2831d83c7857e475`. ## Risks - Moderate risk because the change intentionally runs git commands against local worktrees; the implementation refuses dirty repair when another claimant or active runtime service is detected and records rescue refs for auditability. - Compatibility / release-note callout for self-hosted operators: existing instances that left `enableWorkspaceBranchReconcileForward` unset now get automatic forward branch reconciliation during heartbeat workspace recovery. Operators who want the previous advisory-only behavior can set `experimental.enableWorkspaceBranchReconcileForward` to `false`; dirty quarantine repair can likewise be disabled with `experimental.enableWorkspaceDirtyQuarantineRepair: false`. - If the git rescue succeeds but a later database write fails, the worktree may already be restored while the recovery action remains open; this ordering is documented in code because git side effects cannot participate in the database transaction. - UI risk is limited to the workspace recovery surfaces and covered by component tests plus the existing Cutter visual preview. ## Model Used OpenAI Codex coding agent based on GPT-5, with repository tool use, shell execution, and local test execution. Earlier preserved commits on this branch also show Claude Code / Claude Opus 4.8 assistance in their commit metadata. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] Branch naming exception documented: this PR preserves the existing worktree branch requested for publication while keeping the PR title and body public-facing - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4d898aa7af |
feat(prompt): one-time execution-workspace branch guard in wake prompt (PAP-13326) (#9319)
## Summary
When a task runs in a branch-pinned execution workspace, agents
sometimes switch or rename the workspace branch, which breaks the
worktree contract. This adds a short, one-time prompt hint telling the
agent to stay on the pinned branch.
- **heartbeat.ts**: after the execution workspace is resolved, attach
`executionWorkspace: { branchName }` to the wake payload (only when a
branch pin exists — agent-home runs without a branch are untouched).
- **server-utils.ts (adapter-utils)**: normalize the new payload field
and render one bullet in `renderPaperclipWakePrompt`:
> `- execution workspace branch: you are running in an execution
workspace on branch \`<name>\`. Do not switch, rename, or re-point this
branch; keep all commits on it.`
- The hint renders **only on non-resumed sessions** — resume-delta
prompts skip it, so it appears the first time an issue's session starts,
not on every turn, and it never pollutes the issue thread. One renderer
change covers every adapter (claude, codex, cursor, gemini, grok,
opencode, pi, hermes, acpx engine) with zero per-adapter edits.
## Tests
- `server-utils.test.ts`: branch guard renders on first prompt, absent
on resumed-session prompts, absent when no branch is pinned; payload
round-trips through `stringifyPaperclipWakePayload`.
- `heartbeat-workspace-branch-containment.test.ts`: the finalize-path
adapter mock now asserts the wake payload the adapter receives carries
the branch pin matching `context.paperclipWorkspace.branchName`
(end-to-end heartbeat wiring, embedded postgres). All 6 pass.
- Full `server-utils` (56) and acpx-engine execute (34) suites green;
adapter-utils typechecks clean; no new server tsc errors.
PAP-13326
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.709.0-canary.15
|
||
|
|
176645187c |
Fix request storm polling and issue-list coalescing (#9190)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The board UI keeps issue, agent, activity, and run state fresh through polling across several pages and sidebar surfaces > - When multiple components or browser tabs poll the same company data at the same time, the API can receive bursts of duplicate issue-list requests > - Those duplicate requests increase database and server load without returning meaningfully different data > - This pull request adds server-side compression/coalescing plus client-side visibility-aware and cross-tab shared polling > - The benefit is lower request volume during normal board usage while preserving fresh UI data for active users ## Linked Issues or Issue Description No exact public GitHub issue was found for this request. Problem: - The board can issue redundant polling requests for the same issue-list data from multiple UI surfaces and tabs. - In busy operator sessions, those bursts can trigger request-storm behavior and unnecessary issue-list load. - Expected behavior is to reuse identical in-flight work server-side and reduce hidden-tab or duplicate-tab polling client-side while preserving normal refresh behavior. Related public context found during duplicate search: - #8206 covers a different board UI 404-storm scope. - #5165 covers separate issue-list behavior around page-size truncation. ## What Changed - Added API compression middleware foundation and a company/created-at index for heartbeat run access. - Added server-side issue-list request storm detection and identical in-flight request coalescing. - Added UI fetch metadata, visibility-aware polling, and request deduplication for issue/activity/client calls. - Added cross-tab shared polling primitives and wired them into the sidebar, inbox, dashboard, issue, project, routine, and agent surfaces. - Resolved the latest `master` migration collision by keeping upstream `0140_built_in_managed_resources.sql` and renumbering this branch's heartbeat-run index migration to `0141_heartbeat_runs_company_created_at_index.sql`; the SQL uses `CREATE INDEX IF NOT EXISTS` for idempotency. - Stabilized server heartbeat cleanup tests exposed by the PR check matrix. - Fixed the Greptile compression follow-up by weakening strong ETags on encoded JSON responses and bypassing compression for streamed/download responses. ## Verification - `pnpm exec vitest run ui/src/components/IssuesList.test.tsx ui/src/pages/Inbox.test.tsx` — passed after resolving the latest `master` conflict in `IssuesList.tsx` and updating the 200-result cap expectations. - `pnpm check:token-gates` — passed after the UI conflict resolution. - `jq -e '.entries | length as $n | (map(.idx) | unique | length == $n) and (map(.tag) | unique | length == $n)' packages/db/src/migrations/meta/_journal.json` — passed after renumbering the migration to `0141`. - `pnpm exec vitest run server/src/__tests__/api-compression.test.ts` - passed after the compression follow-up. - `pnpm exec vitest run server/src/__tests__/issue-list-assignee-filter-routes.test.ts` - passed after the compression follow-up. - `pnpm --filter @paperclipai/server typecheck` - passed after the compression follow-up. - Greptile Review for head `8dbddac41ec273fda404100b4981ddb912fad57b` - passed after the latest conflict/migration fix; all Greptile review threads are resolved. - `pnpm exec vitest run server/src/__tests__/api-compression.test.ts server/src/__tests__/issue-list-assignee-filter-routes.test.ts ui/src/api/client.test.ts ui/src/api/issues.test.ts ui/src/lib/polling.test.ts ui/src/lib/cross-tab-poll.test.ts ui/src/pages/Inbox.test.tsx ui/src/components/SidebarProjects.test.tsx ui/src/components/SidebarAccountMenu.test.tsx ui/src/lib/issueDetailCache.test.ts` — passed. - `pnpm exec vitest run ui/src/api/client.test.ts` — passed. - `pnpm exec vitest run ui/src/pages/Inbox.test.tsx ui/src/components/SidebarProjects.test.tsx ui/src/components/SidebarAccountMenu.test.tsx ui/src/lib/issueDetailCache.test.ts` — passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-worktree-suppression.test.ts` — passed. - `pnpm exec vitest run server/src/__tests__/low-trust-red-team-routes.test.ts` — passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - GitHub PR checks for head `8dbddac41ec273fda404100b4981ddb912fad57b`: all GitHub Actions/status checks passed; Greptile, Superagent, Socket, Snyk, build, typecheck/release registry, general tests, serialized server suites, e2e, canary dry run, policy, and commitperclip review are green; Storybook visual regression and security-review were skipped/neutral by policy. - Confirmed this branch does not include `pnpm-lock.yaml` or `.github/workflows` changes. ## Risks - Medium risk: issue-list coalescing changes request timing and cache semantics for a hot API path. - Medium risk: cross-tab polling uses browser coordination primitives, so older or unusual browser environments need fallback behavior to stay correct. - Low migration risk: the new index migration is ordered after current `master` and uses `CREATE INDEX IF NOT EXISTS`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent, GPT-5-based model, tool-enabled with shell/git execution. Exact hosted deployment identifier and context-window size were not surfaced in the agent runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.709.0-canary.14 |
||
|
|
9acce52aa5 |
[codex] Polish operator UI work state details (#9320)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Operators use the issue list, blocked-work notices, pipeline body documents, and issue documents to understand what agents are doing > - Some of those surfaces had low-information details: live-work blockers were visually easy to miss, revision history showed generic actor labels, and selected active subtasks carried an extra accent border > - These details matter because Paperclip's default UI should make agent state and work history legible without requiring users to inspect raw logs > - This pull request groups a small set of related operator UI polish changes for those work-state and document-history surfaces > - The benefit is clearer review context for blocked work, better attribution in document revision history, and cleaner active-subtask row styling ## Linked Issues or Issue Description No public GitHub issue found during dedup search. ### Bug Report **Pre-submission checklist** - Searched existing open PRs/issues for `document revision authors` and `blocked notice live work`; no duplicates found. - Reproduced against the current `master` base for this PR. - Confirmed this is core board UI behavior, not adapter/provider/local configuration. **What happened?** Blocked live-work notices, document revision history, pipeline body document revision history, and active subtask rows exposed technically correct but low-signal UI details. Revision menus could show generic `Board`/`Agent` labels instead of the specific author, and selected active subtasks had an extra accent border. **Expected behavior** Operators should see clear blocked-live-work copy, recognizable revision authors using available agent/user profile data, and visually consistent selected task rows. **Steps to reproduce** 1. Open an issue with a live-work blocker and inspect the blocked notice. 2. Open an issue document revision menu with agent/user-authored revisions. 3. Open a pipeline item body document revision menu with agent/user-authored revisions. 4. Inspect a selected active subtask row in the issue list. **Paperclip version or commit** `8b6a06ee2` (`origin/master` at branch creation). **Deployment mode** Local dev / self-hosted board UI. **Installation method** Built from source. **Agent adapter(s) involved** Not adapter-specific (core board UI). **Database mode** Not database-related. **Access context** Board operator UI. **Relevant logs or output** Not applicable. **Relevant config** Not applicable. **Additional context** This PR is a small polish/fix contribution. `ROADMAP.md` allows tightly scoped bugs and polish without roadmap-level coordination. **Privacy checklist** Reviewed the PR body for private instance references and omitted internal ticket links. ## What Changed - Polished the live-work blocked notice copy and tests so operators get a clearer explanation of what is blocking progress. - Show issue document revision authors with agent/user names and avatar treatment instead of generic actor labels. - Show pipeline body document revision authors with the same available agent/user profile data. - Removed the extra left accent border from selected active subtask rows. - Stabilized the document revision test helper for React runtimes where `act` is not exported as a function. ## Verification - `git diff --check origin/master..HEAD` - `git diff --name-only origin/master..HEAD -- .github/workflows pnpm-lock.yaml` produced no output. - `pnpm check:token-gates` passed: all color literal, arbitrary bracket value, and raw font-size gates clean. - `pnpm exec vitest run ui/src/components/IssueBlockedNotice.test.tsx ui/src/components/IssueDocumentsSection.test.tsx ui/src/components/IssueRow.test.tsx` passed: 3 test files, 32 tests. - `pnpm --filter @paperclipai/ui typecheck` passed. - Remote PR checks passed on head `3e18fe3f`: 22 complete, 0 pending, 0 failing. - Greptile completed at 5/5 with no blocking issues after addressing the pipeline revision author feedback. ## Risks Low risk. The change is limited to board UI rendering and component tests. The main risk is a subtle visual regression in issue/document surfaces that are not covered by screenshots; the affected component tests cover the intended copy, attribution, and row-style behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent, GPT-5 model family, tool-enabled CLI session. Exact runtime context window is not exposed by this execution environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4856558fd9 |
Fix skills routes to skip user secret resolution
Skip user_secret_ref bindings when resolving adapter config for agent skill listing and sync paths, while keeping normal runtime resolution strict. Add route and service regression tests for required user-secret refs in adapter config.canary/v2026.709.0-canary.13 |
||
|
|
8b6a06ee25 |
[codex] Add built-in agents and Reflection Coach bundle (#9206)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators need first-party agent capabilities for repeatable company work, not just manually created one-off agents. > - Built-in agents need to behave like normal company-scoped agents while preserving approval gates, permissions, budgets, and audit trails. > - Reflection and coaching work also needs bundled instructions, skill content, and a routine so the feature can be installed and reset predictably. > - The API, database, UI, portability, and tests all need to agree on the built-in lifecycle from not provisioned through setup, approval, ready, paused, and reset. > - This pull request adds built-in agent provisioning and the Reflection Coach bundle end-to-end. > - The benefit is a safer first-party path for Paperclip-managed agents without bypassing the same governance model used for operator-created agents. ## Linked Issues or Issue Description No public GitHub issue was found for this exact built-in agent and Reflection Coach bundle work. Problem/motivation: - Paperclip did not have a first-party built-in agent lifecycle for product-owned agents. - Bundled agent resources such as default instructions, skills, and routines needed managed ownership and reset semantics. - Approval-gated companies needed built-in setup to preserve requested adapter, budget, manager, and permission state through board approval. - The board UI needed clear built-in badges, setup affordances, readiness state, and bundle status without exposing secrets. Proposed solution: - Add a company-scoped built-in agent registry, provisioning/reset/reconcile/status APIs, and Reflection Coach bundled resources. - Track bundled managed resources in the database with idempotent migration behavior. - Reuse existing agent approval, authorization, budget, and activity-log paths instead of creating a bypass. - Add UI setup, badges, gates, bundle panels, and route coverage for built-in agents. Duplicate search: - Searched GitHub PRs for `built-in agents Reflection Coach repo:paperclipai/paperclip`; only this PR was returned. - Searched GitHub issues for the same query; no public issues were returned. ## What Changed - Added built-in agent definitions, lifecycle state derivation, provisioning, reset, reconcile, status, and routine-control routes. - Added the `built_in_managed_resources` migration and schema exports for bundled instructions, skill, and routine ownership. - Added the Reflection Coach built-in bundle with default instructions, skill catalog content, routine template, default permissions, and managed-resource drift handling. - Added approval-aware provisioning behavior that preserves requested adapter config, budgets, manager assignment, and built-in permissions through hire approval. - Added authorization and mutation gates for built-in agent and skill changes, including consented Reflection Coach change paths. - Added UI surfaces for built-in agent setup, roster/detail badges, readiness gates, bundle status, routine controls, and route filtering. - Added company import/export and validator coverage for built-in managed resources and low-trust/red-team presets. - Addressed Greptile follow-ups for pending approval reconciliation, consent-gate error propagation, config-read authorization fallback, approval-path manager preservation, and non-model adapter provisioning. ## Verification Local verification: - `git diff --check public/master..HEAD` passed. - `pnpm check:token-gates` passed with all gates clean. - `pnpm exec vitest run ui/src/components/ConfigureBuiltInAgentModal.test.tsx` passed: 1 file, 4 tests. - `pnpm exec vitest run ui/src/components/EntityRow.test.tsx ui/src/pages/Agents.test.tsx ui/src/components/BuiltInAgentGate.test.tsx ui/src/components/ConfigureBuiltInAgentModal.test.tsx ui/src/components/BuiltInBundlePanel.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx ui/src/pages/Routines.test.tsx` passed: 7 files, 64 tests. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/built-in-agents.test.ts src/__tests__/authorization-service.test.ts src/__tests__/company-skills-routes.test.ts` passed: 3 files, 91 tests. - `pnpm --filter @paperclipai/db check:migrations` passed. - `pnpm -r typecheck` passed after the rebase; `pnpm --filter ui typecheck` passed after the final UI review fix. Remote verification on latest head `1c61f693a4ec881d739022b0e75a8ca8bf8c2cd8`: - Merge state: `CLEAN`. - Greptile: `5/5`, zero unresolved Greptile threads. - PR check rollup: all checks successful, neutral, or skipped as expected. - Passing gates include Build, Typecheck + Release Registry, all server shards, all workspace shards, all serialized server suites, e2e, Canary Dry Run, policy, review, verify, Socket, Superagent, and Snyk. ## Risks - This adds a new managed-resource table and migration; the migration uses idempotent create/add/index guards and passed migration safety checks. - Built-in agent provisioning touches approval and authorization paths; tests cover pending approval preservation, stale retry rejection, consent gates, and config-read fallback behavior. - Reflection Coach creates managed instructions, skill, and routine resources; drift/reset behavior is covered by service tests and redacted API responses. - Non-model adapter setup now provisions a `needs_setup` built-in row before command/endpoint fields are complete; this matches the server lifecycle and is covered by the setup modal regression test. ## Model Used OpenAI Codex coding agent based on GPT-5. Exact hosted model ID, context-window size, and reasoning-mode labels are not exposed in this runtime; tool use, shell execution, GitHub CLI/API access, and local code editing were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.709.0-canary.12 |
||
|
|
53d09d4c34 |
fix(ui): inbox/task list parity, nesting alignment, hover perf, and routine detail polish (#9317)
## Thinking Path > - Paperclip's UI is governed by the design system merged in #9134 and the component convergence in #9240 — one Card, one Badge, one nav row, one `IssueRow`, a single multiplicative radius ladder. > - With those primitives in place, the remaining rough edges were interaction and alignment details on the surfaces people use every day: the inbox, the task list, the sidebar, and the routine/task detail pages. > - Each item here was reported from live use and fixed against a running instance, then verified by measurement (pixel alignment, frame timing) rather than by eye alone. > - The result is that the inbox and task lists now behave as one system (same hover, keyboard nav, tree-guide, and archive language), list nesting reads correctly, list hover is smooth, and the routine detail page scrolls and aligns like the rest of the app. ## Linked Issues or Issue Description No public GitHub issue exists for this work; describing per the feature template. - **Problem**: after the design-system foundation (#9134) and component convergence (#9240), the inbox/task lists still had interaction and alignment gaps — hover lag on long lists, keyboard navigation that only partly matched between the two lists, workspace/parent nesting whose guides and chevrons didn't line up, an inbox that sat offset from the task list, and a routine detail page with an odd double-scroll and a bespoke sub-nav. - **Proposed behavior**: the inbox and task lists share one interaction contract (hover, keyboard nav, collapse, archive), list nesting aligns to the status column with clean chevrons, list hover is CSS-only (no per-hover re-render), and the routine detail page uses a fixed header/sub-nav with a single scrolling content region and a sub-nav that matches the primary nav. ## What Changed - **List hover performance**: hover is now painted purely by CSS `:hover` and records the hovered row in a ref, instead of writing list-selection state on every `mouseenter` (which re-rendered 100–300 non-memoized rows per hover). Keyboard nav reads the ref so it still continues from the hovered row; the keyboard band clears on the first real mouse move so hover and keyboard selection never show two bands at once. The row's `transition-colors` fade was removed so the highlight snaps (no comet-tail). Measured on a 100-row scrub: worst frame **333 ms → 33 ms**, long frames (>50 ms) **12 → 0**, avg **41 → 60 fps**. - **Inbox ↔ task-list parity**: keyboard navigation works on every inbox tab (archive/read keys stay scoped to the archivable tab); group headers and parent tasks collapse/expand with the arrow keys in both lists; the task list gains the same j/k / arrows / Enter selection model as the inbox; hover selection bands match; the `g` then `i` go-to-inbox chord works app-wide. - **List nesting & alignment**: the workspace group-header chevron lines up exactly with the task chevrons below it (both lists); the parent→child connector line drops from under the parent's **status icon** rather than its chevron and breaks with a 14px gap around a nested row's own chevron; and inbox rows line up with the task list (read rows no longer reserve a mark-read column). - **Inbox archive affordance**: moved from a bare `x` left of the status icon to an `Archive` icon + label button on the right (before the timestamp), revealed on row hover; the left slot now carries only the unread dot; swipe-to-archive is unchanged. - **Sidebar**: when no agent has a live run, the AGENTS section shows 3 recent agents (was 5) plus "See all agents"; the working-agents view is unchanged. - **Routine detail page**: the layout is bounded to the main scroll area so the header (Run now + automation toggle) and the sub-nav stay fixed and only the section content scrolls (was a page-level scroll competing with a `sticky` sub-nav). The sub-nav items adopt the primary nav's rhythm — row padding/height, inset rounded pill, type scale, and 16px icons — and its background matches the main nav. - **Task detail**: dropped the redundant `🔵` prefix the breadcrumb added for live/in-progress tasks; the status glyph already conveys that state. ## Verification - `pnpm check:token-gates` → 3/3 CLEAN - `pnpm typecheck` → green (all packages) - `cd ui && npx vitest run` → 2255/2255 (assertions updated in lockstep where behavior changed) - `pnpm --filter @paperclipai/ui build` → exit 0 - Storybook visual regression: run locally throughout (the baseline-manifest archive is still unpublished, so CI cannot run this suite — pre-existing condition from #9134). Every visible delta was reviewed against a live instance and, where it was intentional (nesting alignment, routine sub-nav restyle, unread-row shift), the affected snapshots were re-baselined locally. - Manual/measured: inbox + task lists (grouped and nested, light + dark), hover-scrub frame timing, keyboard navigation, and the routine detail scroll/alignment were exercised on a running instance. ## Risks - Behavior-and-alignment changes concentrated in `IssueRow` / `IssuesList` / `Inbox` (the shared task-row surfaces). The riskiest area — the hover/keyboard-selection model — is covered by unit tests (updated in lockstep) and was measured and driven live. - The routine detail scroll change restructures that page's layout container; verified the page's own scroll stays fixed while only the section content scrolls. - The visual suite cannot yet run in CI (unpublished baseline archive — pre-existing); snapshot coverage is local-only until that lands. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic coding with tool use (file editing, test execution, Playwright measurement/screenshot verification); extended thinking enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bc85b456a1 |
fix(ui): keep agent names visible on mobile agents index (#9236)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Agents index (`ui/src/pages/Agents.tsx`) is the main roster view, with both a list layout and an org-tree layout > - On mobile viewports the roster could render rows with no visible agent names because shrink-resistant trailing controls competed with the fixed-width title cell > - A roster where names are unreadable is unusable on phones, and the org-tree view had a related narrow-width overflow risk for deeply indented names > - This pull request lets the list title flex below `xl`, hides nonessential row controls on mobile, keeps Join/Leave reachable, and truncates org-tree names safely > - The result is a readable agents index on mobile while preserving desktop meta-column alignment and truncation ## Linked Issues or Issue Description No existing public GitHub issue; describing the bug per the bug report template: **What happened?** On a mobile-width viewport, the agents index showed rows where agent names could become unreadable or disappear because trailing controls consumed the available row width. In the org view, deeply indented names could overflow the row. **Expected behavior** Agent names remain visible on every viewport, Join/Leave stays reachable, and desktop rows continue to align and truncate as before. **Steps to reproduce** Open Paperclip in a browser at a narrow viewport, navigate to the Agents index, and observe rows with long names or left-membership controls. **Paperclip version or commit** Reproducible on `master` prior to this fix. **Deployment mode** Local dev instance. The bug is viewport-width dependent, not deployment-mode dependent. ## What Changed - `EntityRow` now lets callers control title text, subtitle text, and the meta spacer classes while keeping the default truncation behavior unchanged. - Agents list rows now use `flex-1 xl:flex-none xl:w-56`, so names get mobile width while desktop meta columns keep their aligned fixed title column. - Long list-view names/subtitles wrap below `xl` but return to truncation at `xl` and above. - Join/Leave stays visible on mobile; run/status/star row controls stay hidden on mobile to avoid squeezing names. - Org-tree rows use `min-w-0 truncate`, so deep indentation shortens names with an ellipsis instead of overflowing. - Regression tests cover mobile name visibility, left-membership dimming, responsive title/meta behavior, and mobile Join/Leave reachability. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/EntityRow.test.tsx src/pages/Agents.test.tsx` — 2 files, 18 tests passed. - `pnpm check:token-gates` — all gates clean. - `git diff --cached | rg -n "(AKIA|ASIA|SECRET|TOKEN|PASSWORD|PRIVATE KEY|BEGIN RSA|BEGIN OPENSSH|api[_-]?key|bearer|paperclip_api_key|DATABASE_URL|postgres://|sk-[A-Za-z0-9]|xox[baprs]-)"` — no matches before push. ## Risks - Low risk: UI-only change scoped to the agents index. The main behavior change is that nonessential row controls remain hidden on mobile while Join/Leave remains available there and all controls remain available on larger screens/detail pages. ## Model Used - OpenAI GPT-5 via Codex (`codex_local` adapter), agentic code editing and local/GitHub verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.709.0-canary.11 |
||
|
|
4d50fa9f29 |
chore(lockfile): refresh pnpm-lock.yaml (#9304)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
b13eb5b2b5 |
Skill Studio: three-pane skill IDE with sandboxed test runs (#9241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Skills Manager gives operators a reusable skill layer, but iteration still required manual edits, ad hoc prompts, and indirect run inspection. > - Skill authors need a focused workflow for editing skill files, saving representative test inputs, and running those inputs through an agent without exposing harness tasks as normal company work. > - The backend therefore needs durable test inputs, reusable run templates, hidden harness issues, scoped run execution, retention metadata, and read-containment rules around hidden work. > - The frontend needs a three-pane Studio that keeps skill files, saved inputs/templates, and run output/history visible together while preserving the existing design system and token rules. > - This pull request ships that Skill Studio surface end to end: database migrations, shared contracts, server APIs/services, hidden harness execution behavior, UI routes/components, and focused tests. > - The benefit is faster and safer skill iteration, with inspectable outputs and fewer ways for internal harness work to leak into normal task lists, costs, or adjacent read APIs. ## Linked Issues or Issue Description No public GitHub issue exists for this feature. Feature request summary: - Problem: Skill authors need to edit and test company skills in one place instead of switching between the skill detail page, task creation, run output, and manual prompt history. - Proposed solution: Add a Skill Studio workbench with saved inputs, reusable templates, hidden sandboxed test runs, live run status, output inspection, run history, rerun/delete controls, and frontmatter-aware editing. - Expected users: Paperclip operators and agent-company maintainers who create, fork, import, and tune skills. - Related public PRs: Supersedes #9205, which was replaced so the public PR branch name follows contributor policy. - Duplicate search: searched public GitHub issues and PRs for "Skill Studio"; no other active public issue or PR directly covers this feature. ## What Changed - Added database migrations for Skill Studio test inputs, test runs, test run retention, and reusable run templates. - Added shared Skill Studio types, validators, route helpers, frontmatter utilities, and status handling. - Added server services and routes for saved inputs, test runs, templates, reruns, terminal-run deletion, hidden harness issue execution, and run-detail hydration. - Strengthened hidden-issue read containment across issue-adjacent routes and cost rollups used by skill test harness work. - Added the Skill Studio UI with skill file editing, frontmatter editing, saved inputs, templates, run creation/cancel/rerun/delete flows, output rendering, history, route support, and responsive pane behavior. - Added focused backend, shared, and UI tests for the new APIs, routing logic, editor/run behavior, hidden-issue containment, and migration safety. - Rebased onto current `master`, removed the generated lockfile diff from the PR, and verified no workflow files are changed. ## Verification - [x] `pnpm --filter @paperclipai/db check:migrations` - [x] `pnpm check:token-gates` - [x] `pnpm exec vitest run server/src/__tests__/company-skills-service.test.ts server/src/__tests__/company-skills-routes.test.ts server/src/__tests__/company-skill-test-runs-service.test.ts ui/src/lib/skill-studio.test.ts ui/src/pages/SkillStudio.test.tsx` — 5 files, 132 tests passed - [x] Greptile review on the latest PR head - [x] GitHub PR checks on the latest PR head ## Risks - Medium risk because this is a broad feature touching database schema, server orchestration, issue visibility, and a large UI surface. - Hidden harness issue containment is security-sensitive; this PR includes regression coverage for adjacent read paths and cost rollups. - The new migrations are additive and use idempotent guards where applicable, but deployed databases that previously tested draft migration numbers should still be checked carefully. - The UI depends on a new resizable panels package in `ui/package.json`; the lockfile is intentionally left to repository automation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent with shell, git, and GitHub CLI tool use. Earlier feature commits include assistance from other Paperclip coding agents; this PR preparation, rebase, cleanup commit, and PR body were completed by OpenAI Codex in a Paperclip worktree. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.709.0-canary.10 |
||
|
|
3571b6c38b |
feat(models): add gpt-5.4-mini to Codex and OpenCode selection (and openai/gpt-5.5 to OpenCode) (#4357)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runtimes are selected through adapters, and each adapter exposes the model IDs an operator can pick in the UI > - The `codex_local` and `opencode_local` adapters hardcode those lists, so a newly released model stays unreachable until it is added > - `gpt-5.4-mini` is available via both the Codex CLI and the OpenCode CLI, but neither adapter lists it; `openai/gpt-5.5` is likewise missing from `opencode_local` > - This pull request adds those entries, and deliberately keeps `gpt-5.4-mini` out of the Codex Fast mode allowlist because the model does not support Fast mode > - The benefit is that operators can select these models from the UI instead of falling back to a manual model ID, and Fast mode fails closed rather than sending unsupported overrides to the CLI ## Linked Issues or Issue Description No existing issue covers this. Describing it inline, following the feature request template: ### Problem or motivation `gpt-5.4-mini` is absent from the `models` list of both the `codex_local` and `opencode_local` adapters, and `openai/gpt-5.5` is absent from `opencode_local`. (`codex_local` already ships `gpt-5.5` — it is the default model on master.) Operators who want these models must type a manual model ID. For `codex_local` that has a real side effect. `isCodexLocalFastModeSupported` treats any *unknown* model as Fast-mode-capable and passes `service_tier="fast"` and `features.fast_mode=true` through to the CLI. Because `gpt-5.4-mini` does not support Codex Fast mode, an agent configured with a manual `gpt-5.4-mini` model ID and `fastMode` enabled silently sends overrides the CLI cannot honor. ### Proposed solution Add the three missing entries to the two `models` lists, and leave `CODEX_LOCAL_FAST_MODE_SUPPORTED_MODELS` untouched. Listing `gpt-5.4-mini` in `models` is precisely what makes it a *known* model, so `isCodexLocalFastModeSupported` returns `false`, `buildCodexExecArgs` omits the Fast mode overrides, and `fastModeIgnoredReason` is surfaced to the operator. ### Alternatives considered Adding `gpt-5.4-mini` to `CODEX_LOCAL_FAST_MODE_SUPPORTED_MODELS` as well — rejected, because the model does not support Fast mode and the overrides would be rejected at run time. Leaving the models unlisted so operators keep using manual IDs — rejected, because that is the path that silently enables Fast mode for a model that cannot use it. ### Roadmap alignment Not core roadmap work. `ROADMAP.md` does not plan adapter model-list maintenance; this is routine upkeep as upstream CLIs ship new models. ## What Changed - Add `gpt-5.4-mini` to the `codex_local` adapter's `models` list, positioned after `gpt-5.4` (newest-first ordering). - Add `openai/gpt-5.5` and `openai/gpt-5.4-mini` to the `opencode_local` adapter's `models` list. - Add a `buildCodexExecArgs` test asserting Fast mode is ignored for `gpt-5.4-mini`. `CODEX_LOCAL_FAST_MODE_SUPPORTED_MODELS` is intentionally unchanged. No behavior changes to existing models or adapter logic. ## Verification ``` pnpm --filter @paperclipai/adapter-codex-local --filter @paperclipai/adapter-opencode-local typecheck npx vitest run packages/adapters/codex-local packages/adapters/opencode-local ``` Both pass: typecheck clean on both packages, and 22 test files / 133 tests green, including the new `ignores fast mode for gpt-5.4-mini` case. ## Risks Low risk. The change is additive: three entries appended to two model-selection lists, plus one test. No default model changes, no adapter logic changes, no migrations. One behavioral shift is intended. An operator who had `gpt-5.4-mini` configured as a *manual* model ID with `fastMode` enabled was getting Fast mode overrides passed through to the Codex CLI. After this change `gpt-5.4-mini` is a known model, so those overrides are dropped and `fastModeIgnoredReason` explains why. ## Model Used - OpenAI Codex CLI with GPT-5 / GPT-5.5-assisted code editing (the original commits on this branch). - Anthropic Claude Opus 4.8 (`claude-opus-4-8`, 1M context, extended thinking, tool use) for the master merge, conflict resolution, and the scope reduction in the latest commit. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched the GitHub PR list for similar or duplicate PRs and confirmed this one is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (N/A; the adapter docs describe Fast mode support, which is unchanged) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>canary/v2026.709.0-canary.9 |
||
|
|
1c75a46c10 |
feat(ui): add waiting-on-live-work blocked notice (#9298)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators rely on the issue detail thread to understand whether a task is blocked, live, or waiting on another task. > - A blocked issue can have a healthy blocker chain where downstream work is actively running and the parent will resume automatically. > - Showing that case with the same amber blocked notice as a stalled or attention-needed blocker makes the state look more severe than it is. > - The UI already receives blocker-attention state, blocker summaries, and company live-run ids, so this can be clarified without a new API shape. > - This pull request adds a blue "Waiting on live work" notice for covered blocker chains while preserving the existing amber notice for the other blocked states. > - The benefit is that operators can distinguish healthy queued work from blocked work that needs intervention. ## Linked Issues or Issue Description Refs #3820 Refs #8271 Related PR: #3877 Supersedes #9295 ## What Changed - Added a blue `IssueBlockedNotice` variant when `blockerAttention.state` is `covered` and the blocker chain has live work. - Rendered blocker-chain progress as done, running, and queued steps, including a "Now running" row for live terminal blockers. - Preserved the existing amber blocked notice for stalled, attention-needed, ordinary blocked, and successful-run handoff states. - Plumbed the existing `liveIssueIds`, `blockedBy`, and `blockerAttention` data from issue detail into the chat-thread blocked notice. - Added regression coverage around the covered live-work state, the no-confirmed-live fallback, numeric step ordering, and amber fallback states. - Hardened a low-trust server route test cleanup helper so CI deletes heartbeat run events before deleting heartbeat runs. ## Verification - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm exec vitest run ui/src/components/IssueBlockedNotice.test.tsx` - GitHub PR workflow is green on head `52ab6d9076ce233c183bf7133fa666e8597b6765`. - Greptile check is green on head `52ab6d9076ce233c183bf7133fa666e8597b6765` with zero unresolved review threads. ## Risks Low runtime risk: the product change is frontend-only and uses data already returned to the issue detail page. The main risk is visual regression in the blocked notice; the change keeps non-covered states on the existing amber path and adds focused regression coverage. The server-side change is test-only cleanup for an existing CI shard failure. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex using GPT-5, tool-use enabled in a repository workspace. The runtime did not expose a more specific model build id or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.709.0-canary.8 |
||
|
|
97e0752158 |
Fix work timeline visible stats and kickoff attribution (#9271)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The work timeline helps operators understand what agents did, when they did it, and who kicked off each run > - Timeline statistics should describe the window the operator is viewing, not hidden work outside the viewport > - Kickoff chips should point to the run that was actually closest to the triggering edge, not stale later runs on the same issue > - This pull request scopes timeline summary stats to the visible window and deduplicates kickoff attribution to the nearest matching run > - The benefit is that operators get a more accurate, less misleading timeline while scanning agent work ## Linked Issues or Issue Description No matching public GitHub issue was found. Public duplicate search found no open PR for "work timeline visible window kickoff". ### What happened? Work timeline summary stats could count hidden spans outside the selected visible window, and kickoff attribution could appear on a stale matching run instead of the closest run associated with the edge. ### Expected behavior Visible timeline stats should reflect only the current visible range, and each kickoff edge should attach to the nearest matching run. ### Steps to reproduce 1. View a work timeline with agent spans that begin before or end after the selected visible window. 2. Compare the summary stats against only the spans visible in the viewport. 3. View repeated runs for the same actor and issue that share a kickoff edge. 4. Check which run gets the kickoff chip. ### Paperclip version or commit Current `origin/master` before this PR. ### Deployment mode Board UI. ## What Changed - Filters timeline summary stats to the visible time window. - Passes visible-window metadata into the work timeline chart and page state. - Assigns each kickoff edge to only the closest matching run, with deterministic tie-breaking. - Adds focused regression coverage for visible stats and stale kickoff attribution. ## Verification - `pnpm exec vitest run ui/src/components/timeline/WorkTimelineChart.test.tsx ui/src/lib/timeline/layout.test.ts ui/src/pages/Timeline.test.tsx` -- 38 tests passed. - `pnpm check:token-gates` -- all gates clean. - `git diff --check origin/master...HEAD` -- passed. ## Visual Evidence Public review artifacts for the user-visible timeline behaviours. The SVGs use sanitized run labels and no internal Paperclip issue identifiers.   Artifact gist: https://gist.github.com/cryppadotta/dc2645d54c4bc56ffc0323224b3ea4bf ## Risks - Low risk. The change affects timeline rendering and summary calculations only; the main risk is a subtle difference in which run gets a kickoff chip when multiple runs are very close together. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent based on GPT-5, tool-enabled shell workflow. Exact hosted model variant and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.709.0-canary.7 |
||
|
|
7cf0d3ebb0 |
Require health readiness for Paperclip dev services (#9269)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents often run against managed workspace runtime services, including reusable Paperclip dev servers > - A running process and an open root URL are not enough to prove the Paperclip API is actually ready > - If the API health endpoint is still failing, agents can reuse a service that looks alive but cannot safely serve the board or API clients > - This pull request makes Paperclip dev runtime readiness probe the resolved `/api/health` endpoint > - The benefit is that runtime service reuse waits for the same health signal operators and agents depend on ## Linked Issues or Issue Description No matching public GitHub issue was found. Public duplicate search found no open PR for "workspace runtime health readiness". Bug report: ### What happened? A managed Paperclip dev runtime service could satisfy HTTP readiness at the exposed base URL even when the Paperclip health endpoint was returning an unhealthy status. ### Expected behavior Paperclip dev runtime services should not be considered ready until their health endpoint succeeds. ### Steps to reproduce 1. Start a workspace runtime service named `paperclip-dev` whose base URL responds successfully. 2. Make that same service return HTTP 503 from `/api/health`. 3. Ask Paperclip to ensure the runtime service for a run. 4. Observe that the service can be reused even though the API health endpoint is not ready. ### Paperclip version or commit Current `origin/master` before this PR. ### Deployment mode Local workspace runtime service management. ## What Changed - Resolve Paperclip dev runtime readiness checks to the service health URL before polling. - Surface readiness errors with the actual health URL that failed. - Add a regression test that fails when `/api/health` returns HTTP 503 even if the service process is running. ## Verification - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts` — 81 tests passed. - `git diff --check origin/master...HEAD` — passed. ## Risks - Low to medium risk. This tightens readiness for Paperclip dev runtime services, so a service that previously looked ready while unhealthy will now fail fast instead of being reused. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent based on GPT-5, tool-enabled shell workflow. Exact hosted model variant and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.709.0-canary.6 |
||
|
|
3e63a7e3e5 |
Fail fast stalled Daytona Git network commands
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Daytona sandbox provider executes agent commands inside ephemeral Daytona workspaces via `executeCommand` > - Git operations inside these workspaces can block indefinitely when remotes are unreachable or when git prompts for credentials interactively (e.g. via askpass or terminal prompts) > - When git blocks, it consumes the full 900s adapter RPC ceiling, which surfaces as a hard timeout crash on the Paperclip side rather than an actionable error > - This PR adds noninteractive credential defaults and a 120s cap on detected Git network subcommands so stalled operations fail fast with a useful message > - The benefit is that engineers and agents see an actionable error pointing at missing credentials or unreachable remotes instead of an opaque 900s RPC crash ## Linked Issues or Issue Description <!-- No public GitHub issue exists for this internal infrastructure fix. Describing the issue inline. --> **Bug: Daytona sandbox Git network commands stall for up to 900 seconds** **What happened?** `executeCommand` hangs for the full 900 s adapter RPC ceiling when a Daytona workspace git network command (push/fetch/pull/clone) prompts for credentials interactively or the remote is unreachable. The command blocks silently for up to 900 s then crashes with a generic timeout error that names no actionable root cause. **Steps to reproduce** Run any agent handoff that includes a `git push`, `git fetch`, or `git pull` to a remote inside a Daytona workspace where the remote is unreachable or credentials are missing. **Expected behavior** Command fails fast (within ~120 s) with an actionable error naming the unreachable remote or the missing noninteractive credential. **Deployment mode** Daytona sandbox provider (`packages/plugins/sandbox-providers/daytona`). Root cause: Daytona one-shot execution wrappers did not set `GIT_TERMINAL_PROMPT=0`, `GCM_INTERACTIVE=Never`, or disabled askpass helpers, so git blocked waiting for interactive terminal input; no per-operation timeout existed for network-bound git subcommands. ## What Changed - Added `GIT_TERMINAL_PROMPT=0`, `GCM_INTERACTIVE=Never`, `GIT_ASKPASS=echo`, `SSH_ASKPASS=echo`, `SSH_ASKPASS_REQUIRE=force` to all Daytona one-shot execution wrapper invocations so git never blocks waiting for a credential prompt; callers may override via the `env` parameter - Detects Git network subcommands (`push`, `fetch`, `pull`, `ls-remote`, `clone`, `remote update`, `submodule update`) and caps their timeout at 120 s instead of the full 900 s adapter RPC ceiling - Returns an actionable timeout message that names the unreachable remote or the missing noninteractive credential rather than propagating the raw SDK error - Adds two new Vitest tests: one verifying noninteractive credential defaults are injected, one verifying the 120 s network cap and the improved timeout message ## Verification - `corepack pnpm exec vitest run packages/plugins/sandbox-providers/daytona/src/plugin.test.ts --config packages/plugins/sandbox-providers/daytona/vitest.config.ts` — 40 tests pass - `corepack pnpm exec tsc -p tsconfig.json --noEmit` from `packages/plugins/sandbox-providers/daytona` — clean - `corepack pnpm check:no-git-push` — clean > **Known CI note:** The standalone provider package is intentionally excluded from the root workspace. Direct `corepack pnpm test` from the provider directory fails before running tests due to Vitest tsconfig-root resolution in a grafted checkout. The root-config Vitest invocation above is the passing test signal. ## Risks - **Low overall risk.** The new `GIT_TERMINAL_PROMPT=0` / askpass defaults only affect Daytona one-shot execution; they do not touch any shared git config or host environment. - Callers that previously relied on interactive credential prompts inside Daytona (an unlikely pattern for agent workspaces) will now fail fast instead of prompting — this is the intended behavior. - The 120 s network timeout applies only when the command string starts with a recognized git network subcommand, so non-network git operations and all non-git commands are unaffected. - No PII, telemetry schema, crypto, auth flow, or new external endpoint changes. ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`) — extended thinking mode, tool use, code execution. Anthropic Claude running via Paperclip Claude Code adapter. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold.kim@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.709.0-canary.5 |
||
|
|
cf8b6e1bdd |
chore: trigger Docker build-and-push for lockfile refresh (post #9252)
Empty commit to trigger Docker workflow on the fixed lockfile lineage (
canary/v2026.709.0-canary.4
|
||
|
|
ea301442e8 |
chore(lockfile): refresh pnpm-lock.yaml (#9252)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
719da5f9b5 |
Harden environment deletion and expose delete blast radius (#9250)
## Thinking Path > - Paperclip manages AI agents that each have an associated execution environment (local, Kubernetes, etc.) > - Instance administrators can create and delete environments; currently the DELETE endpoint has no protection against deleting managed or in-use environments > - Deleting the managed local environment or the instance-default environment would break all agents using those environments with no path to recovery > - The endpoint also suffered a TOCTOU race: a check-then-delete pattern allowed the managed-local or default guard to pass if the environment's role changed between the read and the delete > - This pull request adds a blast-radius read endpoint so admins can preview impact, hard-blocks the dangerous deletes atomically, cleans up all dependent references after a valid delete, and fixes a concurrent creation race in ensureLocalEnvironment ## Linked Issues or Issue Description Fixes #9251 ## What Changed - **New endpoint** `GET /api/environments/:id/delete-blast-radius` (instance-admin gated): returns reference counts (agent defaults, workspace selections, issue selections, project selections, secret bindings, active leases, active setup sessions) and blocking reasons — no config, env-var values, or secret data returned. - **Atomic delete guard** `environmentService.removeIfDeletable(id)`: performs the DELETE with an inline `WHERE driver != 'local' AND NOT EXISTS (instanceSettings where defaultEnvironmentId = id)` predicate, eliminating the TOCTOU race between the app-level check and the DB write. - **Route hardening**: `DELETE /environments/:id` now calls `getDeleteBlastRadius` first (app-level check + logging), then calls `removeIfDeletable` (atomic guard). If the atomic guard returns null the route fetches a fresh blast-radius snapshot and rejects with a 409 Conflict carrying `deleteBlockedReasons`. - **Reference cleanup on valid delete**: after a successful delete, the route clears environment selections on all company execution workspaces, issues, and projects; syncs env-var secret bindings to `{}` (removing bindings for the deleted environment); syncs config secret refs to `[]` for the environment target; and removes the SSH private-key secret if one was stored. - **Race fix in `ensureLocalEnvironment`**: the insert-or-nothing path now catches a `environments_name_idx` unique-constraint violation and falls through to the existing SELECT, treating the name conflict as idempotent. - **Shared types**: `EnvironmentDeleteBlastRadius` and `EnvironmentDeleteBlockedReason` exported from `@paperclipai/shared`. - **OpenAPI**: registers the new blast-radius endpoint; updates the delete-environment response schema to document 403/404/409. - **Tests**: 56 existing environment-route and service tests continue to pass; new service-level regression tests assert the atomic guard rejects `local`-driver environments and instance-default environments and succeeds for deletable ones. ## Verification ``` corepack pnpm exec vitest run \ server/src/__tests__/environment-routes.test.ts \ server/src/__tests__/environment-service.test.ts # 56 tests, all passing corepack pnpm --filter @paperclipai/shared typecheck node scripts/ensure-plugin-build-deps.mjs cd server && ../node_modules/.bin/tsc --noEmit ``` ## Risks - **Blast-radius endpoint auth**: guarded by `assertCanAccessInstanceEnvironments`, the same gate as the existing environment-list and delete routes. Non-admin callers receive 401/403 before any data is returned. - **Atomic guard may reject a delete that the app-level check passed**: this is intentional — it means the environment became protected between the read and the write. The caller receives a fresh blast-radius snapshot explaining why. - **Secret cleanup ordering**: cleanup runs after the atomic DELETE succeeds, in parallel across companies. If cleanup partially fails the environment row is already gone; partial-cleanup state is recoverable by re-running the sync operations. Risk: low — these are idempotent upsert/sync operations. - **ensureLocalEnvironment race fix**: swapping a unique-constraint error for an idempotent SELECT adds one extra query on the conflict path. This path is rare (only fires during concurrent boot) and is significantly safer than the previous behavior. - **No migration**: all changes are application-level; no schema changes required. ## Model Used - Provider: Anthropic - Model: claude-sonnet-4-6 (Claude Sonnet 4.6) - Context window: 200k tokens - Mode: agentic tool use via Paperclip agent system (Claude Code) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Harold Kim <harold.kim@paperclip.ing>canary/v2026.709.0-canary.3 |
||
|
|
1b16f9611c |
feat(telemetry): document credential health retention (#9248)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The shared telemetry package (`packages/shared/src/telemetry`) defines the public contract for first-party telemetry events — event names, dimension shapes, and now retention windows > - A new telemetry event, `codex.credential_health`, carries credential-observability fields (enums, booleans, counts, coarse buckets) — no token material and no PII > - The retention window for this event and its class was undocumented at the contract level, leaving data-infra and reviewers without a discoverable source of truth > - This pull request adds `retention.ts` as the canonical retention-contract surface, assigns `codex.credential_health` to the `operational_enum_count` class (90-day window), exports the contract from the shared package, and updates the README > - The benefit is that the retention window is discoverable from the event definition rather than being implicit pipeline knowledge, and the no-token/no-PII note is locked in a contract test rather than relying on prose ## Issue Description No public GitHub issue exists for this change. Inline description follows the [feature request template](.github/ISSUE_TEMPLATE/feature_request.yml): ### Problem or motivation `packages/shared/src/telemetry/` applies a 90-day retention window to `codex.credential_health` events at the pipeline level, but this policy is not stated anywhere in the shared telemetry contract. Without a discoverable retention declaration, reviewers and data-infra must read pipeline configuration to understand retention behaviour — there is no contract-level source of truth. The `codex.credential_health` event carries only enums, booleans, counts, and coarse buckets — no token material and no PII. ### Proposed solution Add a `retention.ts` module to the shared telemetry package that defines the `operational_enum_count` retention class (90-day window) and maps `codex.credential_health` to it. Export the contract from the package index and add a focused contract test. Update the README with a Retention section and a pointer in the Public Sources table. This is additive documentation only — no runtime paths change. ### Alternatives considered Keep retention implicit in pipeline configuration only. Rejected: this leaves no discoverable, versioned contract for reviewers or data-infra, and means every consumer must read pipeline config to understand retention semantics. A schema-level declaration is the correct long-term home. ### Roadmap alignment Aligns with the telemetry contract hardening track — making implicit operational knowledge explicit and testable at the shared-package level. ## What Changed - **`packages/shared/src/telemetry/retention.ts`** (new): defines `RETENTION_DAYS` (class → days) and `EVENT_RETENTION_CLASS` (event name → class). `operational_enum_count` is the only class: 90-day window for enum/count/bucket events with no token material or PII. `codex.credential_health` is the first entry. The `string` key type is intentional to accommodate cross-system events (e.g. the Codex CLI) not yet promoted to the first-party `PaperclipEventName` schema. - **`packages/shared/src/telemetry/retention.test.ts`** (new): three focused assertions — `operational_enum_count` is 90 days, `codex.credential_health` is assigned that class, and the resolved window is 90 days. - **`packages/shared/src/telemetry/index.ts`**: exports `RETENTION_DAYS`, `EVENT_RETENTION_CLASS`, and `RetentionClass` from the shared package. - **`packages/shared/src/telemetry/README.md`**: adds `retention.ts` to the Public Sources table and a new Retention section with a class reference table and guidance for future assignments. ## Verification - `git diff --check` passes (no whitespace errors) - `pnpm --filter @paperclipai/shared exec vitest run src/telemetry/retention.test.ts src/telemetry/readme-contract.test.ts` — runs the new contract tests and the existing readme-contract test - `pnpm --filter @paperclipai/shared typecheck` — confirms the new exports compile cleanly - CI gates green ## Risks Low risk. This is a documentation-only addition: - No new telemetry event, dimension, emitter, or schema change - No runtime code paths changed - The new file is tree-shaken away in any consumer that doesn't import from it - `satisfies Record<string, number>` on `RETENTION_DAYS` ensures the type stays correct as new classes are added ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`) — Anthropic, 200k context window, tool use enabled, no extended thinking mode. Used for code authoring and PR composition. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold.kim@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Harold Kim <harold-kim@paperclip.ing>canary/v2026.709.0-canary.2 |
||
|
|
eedc7ddef2 |
Make ACP the default engine for local adapters (#9238)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter packages are the bridge between the control plane and local agent harnesses such as Claude Code, Codex, and Gemini CLI. > - ACP support was concentrated in a separate `acpx_local` adapter, which made ACP feel like a separate agent choice instead of an execution capability of the harness adapters. > - Claude, Codex, and Gemini now have ACP-capable harnesses, so the native adapter should own ACP selection, fallback, config, transcript parsing, and environment diagnostics. > - The standalone ACPX adapter still needs a compatibility path for existing rows, but it should not be offered as an active adapter for new agents. > - This pull request moves the shared ACP runtime into `@paperclipai/acpx-engine`, wires Claude/Codex/Gemini local adapters to prefer ACP when prerequisites are available, and retires `acpx_local` to a tombstone. > - The benefit is one adapter per harness, richer ACP transcripts by default where possible, and a migration path for existing Claude/Codex ACPX agents. ## Linked Issues or Issue Description Closes #5932 — the broken default `acpx_local` Claude path is replaced by native `claude_local` ACP support, existing Claude/Codex ACPX rows migrate to native adapters, and new agents no longer choose the standalone ACPX adapter. Refs #4893 — original merged ACPX local adapter runtime that this PR replaces with native per-harness ACP engines. Refs #6590 — prior ACPX-Claude seamlessness work folded into the new native Claude ACP path. Refs #197 — related open generic ACP/Kiro adapter work; this PR does not close it because Kiro/custom generic ACP remains a separate adapter decision. Refs #7018 — related Kimi-specific `acpx_local` shell failure; this PR retires the built-in standalone adapter but does not add a native Kimi adapter. Refs #8864 — related ACPX prompt/API guidance PR; this PR moves runtime guidance into the shared/native ACP engine path instead of the old standalone adapter. Refs #8881 — related `acpx_local` POSIX shell failure from the old `acpx` pin; this PR updates ACP dependencies but does not claim custom/OMP ACP support as a first-class native adapter. Refs #8964 — related open `acpx_local` stderr cleanup PR; this PR makes the old runtime path obsolete for new agents but keeps it as a non-closing reference. Problem description: - The standalone `acpx_local` adapter duplicates Claude/Codex agent choices that already have first-class local adapters. - ACP should be an execution engine capability of each harness adapter when the underlying harness supports ACP. - Existing `acpx_local` agents should either migrate to native harness adapters or fail with an explicit retirement message instead of silently falling back to the process adapter. ## What Changed - Added `@paperclipai/acpx-engine` as the shared ACP execution, session-codec, CLI formatter, and UI parser package. - Wired `claude_local`, `codex_local`, and `gemini_local` to auto-select ACP by default when prerequisites pass, with `engine=cli` opt-out and `engine=acp` strict mode. - Added ACP config schema/UI fields, environment checks, session-codec preservation, transcript parsing, and adapter capability metadata for the native adapters. - Retired `acpx_local` to a server tombstone, removed its UI/package/runtime image surface, and added a migration for existing Claude/Codex ACPX agents. - Updated package manifests, lockfile, release tooling, docs, Kubernetes sandbox defaults, and tests. ## Verification - `corepack pnpm --filter @paperclipai/acpx-engine typecheck` - `corepack pnpm --filter @paperclipai/adapter-claude-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-codex-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `corepack pnpm --filter @paperclipai/acpx-engine exec vitest run` - `corepack pnpm --filter @paperclipai/adapter-claude-local exec vitest run src/server/acp.test.ts src/server/execute.acp-fallback.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-codex-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-gemini-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts src/ui/parse-stdout.test.ts` - `corepack pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && corepack pnpm --filter @paperclipai/server exec tsc --noEmit` - `corepack pnpm --filter @paperclipai/server exec vitest run src/__tests__/adapter-routes.test.ts src/__tests__/adapter-session-codecs.test.ts src/__tests__/adapter-models.test.ts` - `corepack pnpm --filter @paperclipai/ui typecheck` - `corepack pnpm --filter @paperclipai/ui exec vitest run src/adapters/metadata.test.ts src/adapters/adapter-display-registry.test.ts src/components/AgentConfigForm.test.ts src/components/AgentConfigForm.render.test.tsx src/components/transcript/RunTranscriptView.test.tsx` - `node --test scripts/bootstrap-npm-package.test.mjs scripts/release-package-map.test.mjs scripts/verify-release-registry-state.test.mjs` Note: the server typecheck script calls `pnpm` internally; this dev shell exposes pnpm through Corepack only, so I ran the two script steps manually with `corepack pnpm`. ## Risks - Migration changes existing `acpx_local` Claude/Codex agents to native adapter types and clears old ACPX task sessions/runtime state. - Custom ACP commands remain on the retired tombstone and will need a separate future adapter/plugin path. - ACP auto-selection depends on local Node and ACP server command prerequisites; remote and unsupported environments fall back to CLI unless `engine=acp` is explicit. - `@paperclipai/acpx-engine` is a new public package and needs npm trusted-publishing bootstrap before release automation can publish it. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent. Exact hosted model build and context-window size are not exposed in this runtime. Tool use included shell execution, repository editing, GitHub CLI operations, and local test/typecheck execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.709.0-canary.1 |
||
|
|
187a929d97 |
docs: explain Codex and Claude credential ownership
Squash-merges the adapter credential ownership doc PR. - Documents codex_local host-owns-auth topology and CODEX_HOME auth precedence - Documents claude_local snapshot-owns-auth topology (sandbox targets only) - Adds API-key vs ChatGPT-subscription guidance for high-concurrency fleets - Adds deferred config-validation warning spec for sandbox credential mismatches - Fixes: SSH scope mismatch (Greptile P1) and codex auth.json materialization note (Greptile P2) Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3288ec4f47 |
chore: ignore materialized Paperclip runtime directory
Adds .paperclip-runtime/ to .gitignore so the materialized adapter credentials directory is never accidentally committed. |
||
|
|
c90e66bdd6 |
feat(ui): design-system component convergence — Card/Badge adoption, multiplicative radius ladder, unified list surfaces (#9240)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is governed by a design system (`DESIGN.md` + the token layer in `ui/src/index.css`, merged in #9134) whose first principle is "one way to say each thing" — one Card, one Badge, one nav row > - After the token extraction landed, ~35 files still hand-rolled card containers, ~45 files hand-rolled pill spans, the sidebar agents section duplicated the nav-row chrome, and the inbox and tasks lists rendered the same task rows two subtly different ways > - Each divergence is a place where a future design change (radius, hover language, status vocabulary) silently misses surfaces, defeating the "edit tokens + run checks" model the design system exists for > - This pull request converges those surfaces onto the shared primitives, codifies the radius scale as the modern multiplicative shadcn ladder, and unifies the row/hover/tree-guide language across the inbox and tasks lists — every visible delta was human-reviewed screen-by-screen against a live instance across nine feedback rounds > - The benefit is that the app's look is now steerable from single knobs (one `--radius` anchor, one Card, one Badge, one row component), and the visual regression suite covers the result (514 snapshots including a new AgentDetail page story) ## Linked Issues or Issue Description No public GitHub issue exists for this work; describing per the feature template: - **Problem**: after the design-token foundation (#9134), component-level drift remained — hand-rolled cards/pills, duplicated sidebar row chrome, and two different renderings of task rows (inbox vs tasks list) meant design changes had to be applied per-surface and frequently missed spots (e.g. status glyphs rendered 16px in the inbox but 20px in the tasks list because a slot override silently beat the component default). - **Proposed behavior**: all card-shaped containers render via `Card`, all label pills via `Badge`, sidebar rows via `SidebarNavItem`, and both task-list surfaces via one `IssueRow` configuration; the radius scale is a single multiplicative ladder anchored at `--radius: 0.5rem`. - **Alternatives considered**: converting interactive `<button>`/`<Link>` cards to `Card` divs (rejected — breaks semantics; documented inline with `design-allow` comments instead); keeping the legacy additive radius ladder (rejected in favor of the standard shadcn multiplicative mapping). ## What Changed - `Card` adoption across ~35 files (settings pages, auth/board flows, dashboards, list containers, KPI tiles); non-adoptable sites (interactive cards, `<li>` rows, class-string props, chart tooltip) carry documented `design-allow(card-pattern)` comments - `Card` gains an `interactive` prop — one quiet hover affordance for clickable cards (cursor, border darken, shadow lift, focus ring), applied to skills tiles, artifact cards, and the company selector; cards carry no resting shadow - `Badge` adoption for 113 hand-rolled pill spans across ~45 files; `PropertyChip` wraps `Badge` internally; `StatusBadge`, external-object chips, and match chips stay bespoke by documented decision (WCAG-tuned status mechanics) - Radius ladder becomes the multiplicative shadcn mapping (`sm/md/lg/xl/2xl/3xl/4xl = 0.6/0.8/1.0/1.4/1.8/2.2/2.6 × --radius`, anchor `0.5rem`); every card surface unifies on `rounded-lg`; the orphaned 8px literal token is deleted - Sidebar: agent rows render via `SidebarNavItem` (new additive props: `iconNode`, `active`, `trailing`, `liveAccessory`); live dots use `--status-agent-running`; one row rhythm and inset rounded pill highlight; right-aligned trailing badges; every labeled section is collapsible - Inbox + tasks lists unified: md status glyphs, `accent/50` rounded row hovers, vertical tree guides under parent rows (opaque underlay so dark-mode translucent borders don't stack), no horizontal dividers under expanded parents; the swipe-to-archive reveal layer shows only mid-swipe; board toggle uses the `SquareKanban` glyph - Kanban: every column carries a status-hued tint; lanes default expanded (including empty); compact mode collapses empty lanes to labeled rails (fixes a clipped, label-less empty-column state) - Storybook: new AgentDetail page story (realistic fixtures, light+dark) joins the visual suite; suite captures with `reducedMotion: 'reduce'` and the ux-lab reasoning ticker honors `prefers-reduced-motion`; a stale lexical alias in `storybook/main.ts` is fixed (build was broken since the lexical 0.46 bump) - Keyboard navigation, from live review of the unified lists: inbox navigation keys work on every tab (archive keys stay scoped to the archivable tab); keyboard-driven scrolling no longer hands the selection to whatever row lands under the stationary cursor (hover selects only after real pointer movement); the tasks list view gains the same j/k / arrows / Enter selection model as the inbox; and the `g` then `i` go-to-inbox chord works app-wide instead of only on the issue detail page - Token gates restored to 3/3 CLEAN (tokenized a post-#9134 regression in the recovery card); decisions recorded in `doc/design/DECISION-SHEET.md` and `doc/design/COMPONENT-INVENTORY.md` (investigation verdicts: FileTree vs WorkspaceFileBrowser and the four entity pickers stay separate — evidence included) ## Verification - `pnpm check:token-gates` → 3/3 CLEAN - `pnpm typecheck` → green (all packages) - `cd ui && npx vitest run` → 2106/2106 (assertions updated in lockstep where they documented superseded decisions; new tests for the global go-to-inbox chord) - `pnpm --filter @paperclipai/ui build` → exit 0 - `pnpm build-storybook` → succeeds (also fixes the lexical-alias break on master) - Visual regression: 514-snapshot Playwright suite green against the updated baseline (zero diffs from the keyboard-navigation round — those changes are purely behavioral). Note: baselines live outside git per the suite design and the baseline-manifest archive is not yet published, so CI cannot run this suite — it was run locally throughout; every visible delta was reviewed screen-by-screen in a live instance across nine review rounds. Review evidence (before/after triplets) intentionally kept out of the repo for size; available on request. - Manual: exercised dashboard, tasks (list + board), inbox, agents, skills, costs, settings, and artifact surfaces in light and dark themes ## Risks - Wide but shallow visual surface: most changes are class-string substitutions with behavior preserved (props, handlers, roles, test ids). The riskiest areas — dnd-kit card refs (React 19 ref-as-prop), inbox swipe-to-archive, and sidebar overlays — are covered by existing unit tests (all green) and were manually exercised. - Intentional visual deltas (rounded cards, tinted kanban columns, md status glyphs, unified hovers) are design decisions recorded in `doc/design/DECISION-SHEET.md`; each maps to a re-baselined snapshot set locally. - The visual suite cannot yet run in CI (unpublished baseline archive — pre-existing condition from #9134); until that lands, snapshot coverage is local-only. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic coding with tool use (file editing, test execution, Playwright screenshot verification); extended thinking enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (will confirm once CI runs) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending first review) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>canary/v2026.709.0-canary.0 |
||
|
|
cc17f29e7e |
fix(timeouts): raise sandbox wall-clock backstop to 4h and make acpx_local timeouts self-describing (#9232)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runs execute through adapters (e.g. `acpx_local`), which can run locally, over SSH, or inside sandbox execution targets, each with a wall-clock execution timeout > - Sandbox-backed runs defaulted to a 30-minute wall-clock backstop (`DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC = 1800`), which kills healthy long agent runs that are still making progress — long before the recovery watchdog's 4h critical threshold would even consider them stuck > - On top of that, `acpx_local` resolved its timeout directly from `adapterConfig.timeoutSec` instead of the shared execution-target resolver, and its timeout failures surfaced as a bare `Timed out after Ns` — giving operators no clue which timer fired or which knob raises it > - This pull request raises the sandbox backstop to 4h (aligned with the recovery watchdog), routes `acpx_local` through the shared timeout resolver, logs the effective timeout and its source at run start, and makes every timeout error message self-describing > - The benefit is that long-running sandbox agent runs no longer die at 30 minutes, and when a wall-clock timeout does fire, the run log states exactly which timer fired and how to configure it ## Linked Issues or Issue Description Refs #4535 (related: wall-clock execution timeouts killing agent runs that are still making progress — that issue covers a different hardcoded 600s timer, but the operator pain is the same). No exact public issue exists for this one, so describing it in-PR: **Bug:** A long sandbox-backed `acpx_local` agent run was killed with a bare `Timed out after 1800s` even though the agent was actively working. - **What happened:** The run hit the 30-minute `DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC` backstop. `acpx_local` never consulted the shared execution-target timeout resolution (it read `adapterConfig.timeoutSec` directly, default 0), so on sandbox targets the sandbox-provider default applied with no adapter-level say. The resulting error named neither the timer that fired nor the knob that controls it. - **Expected:** Healthy long runs should not be killed by a 30-minute wall-clock backstop when the recovery watchdog only treats runs as critically stuck after 4h of output silence; and any timeout error should say which timeout fired and how to raise it. - **Impact:** Long, legitimate agent runs in sandboxes fail mid-work; operators waste time reverse-engineering which of several timers produced "Timed out after Ns". ## What Changed - `packages/adapter-utils/src/execution-target.ts` - `DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC` raised from `1_800` to `14_400` (4h), with a comment explaining it intentionally matches the recovery watchdog's `ACTIVE_RUN_OUTPUT_CRITICAL_THRESHOLD_MS` (4h) so the adapter backstop never fires before the watchdog path. Output-inactivity monitors remain the primary hang detectors. - New `resolveAdapterExecutionTargetTimeout(target, configuredTimeoutSec)` returns `{ timeoutSec, source }` where `source` is `configured` / `sandbox_default` / `unlimited`. The existing `resolveAdapterExecutionTargetTimeoutSec` is preserved as a thin wrapper, so current callers are unaffected. - New `formatAdapterExecutionTimeoutErrorMessage(resolution)` and `formatAdapterExecutionTimeoutStartLogLine(resolution)` produce self-describing messages that name the timer that fired and the `adapterConfig.timeoutSec` knob that controls it. - `packages/adapters/acpx-local/src/server/execute.ts` - `buildRuntime` now resolves the wall-clock timeout through the shared resolver: sandbox targets default to the 4h backstop, local/SSH keep the historical "0 = no adapter timeout", and a configured `adapterConfig.timeoutSec` always wins. - The executor logs the effective timeout and its source at run start (`[paperclip] Adapter execution timeout: …`), so a later timeout is diagnosable from the run log alone. - All three bare timeout messages (timer cancel reason, turn result `errorMessage`, catch-path `messageOverride`) now use the self-describing format. - `packages/adapters/acpx-local/src/index.ts` — the adapter configuration doc for `timeoutSec` states the sandbox default and that the output-inactivity monitor remains the primary hang detector. - Tests: `packages/adapter-utils/src/execution-target-sandbox.test.ts` and `packages/adapters/acpx-local/src/server/execute.test.ts` (see Verification). ## Verification - `pnpm --filter @paperclipai/adapter-utils typecheck` — passes - `pnpm --filter @paperclipai/adapter-acpx-local typecheck` — passes - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapters/acpx-local/src/server/execute.test.ts` — 2 files, 40 tests, all pass - New/updated test coverage: - sandbox default resolves to 4h (and the constant is asserted to be `4 * 60 * 60`) - `resolveAdapterExecutionTargetTimeout` reports `configured` / `sandbox_default` / `unlimited` sources with the correct precedence (configured > sandbox default; local/SSH stay unlimited) - exact wording of the self-describing error message and the start-of-run log line - `acpx_local` runtime picks up the sandbox default into `timeoutMs`, keeps the unlimited local default, honors configured-over-default precedence, emits the start-of-run log line, and surfaces the self-describing `errorMessage`/cancel reason when the wall-clock timer kills a turn ## Risks - **Behavioral shift:** sandbox-backed adapter runs that previously hit the 30-minute backstop now run up to 4h before the adapter kills them. Genuinely hung runs are still caught much earlier by the adapters' output-inactivity monitors and by the recovery watchdog; the wall-clock timer is a last-resort kill switch. Operators who relied on the 30-minute default can restore it explicitly via `adapterConfig.timeoutSec`. - **Error-message consumers:** any tooling that pattern-matched the exact `Timed out after Ns` string from `acpx_local` will see the new self-describing message instead. - No API or schema changes; `resolveAdapterExecutionTargetTimeoutSec` keeps its exact signature and behavior (modulo the raised sandbox default). > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic) via Claude Code CLI — model ID `claude-fable-5`, extended thinking enabled, agentic tool use (file edits, shell, test execution) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.708.0-canary.10 |
||
|
|
4a7a732476 |
feat(skills): add company skill fork prechecks (#9235)
Adds company skill fork precheck metadata, fork result/reassignment contracts, selected-agent reassignment during fork creation, and targeted server/shared test coverage.canary/v2026.708.0-canary.9 |
||
|
|
f616b6746c |
fix(server): clean heartbeat run scratch directories (#9234)
Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.708.0-canary.8 |
||
|
|
38cca22b09 |
fix(server): avoid accepted-plan workspace branch freeze before child realization (#9233)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents tackle complex tasks via *plan* flows: a planner decomposes
work into child issues, which are accepted by the board and then
executed
> - When a plan is accepted, `createChild` in `issues.ts` inserts child
issues pre-bound to the parent's already-realized execution workspace —
carrying over its concrete branch ref
> - If the repository's base ref advances between plan acceptance and a
child's first heartbeat, the child inherits a stale branch that no
longer matches the current base
> - At first heartbeat the workspace validator detects the mismatch and
freezes the child ("branch freeze"), blocking it from starting any work
> - The real fix is to strip the concrete workspace binding when
creating accepted-plan children: they should receive only the unresolved
*intent* (mode, baseRef, branchTemplate) and realize a fresh workspace
from the current base on their own first heartbeat
> - This PR implements that strip, adds a regression test that proves a
post-base-advance child realizes cleanly, and also fixes
`parseIssueExecutionWorkspaceSettings` so `environmentId` is not
silently dropped on update round-trips (a latent bug that was masking
the original fix)
## Linked Issues or Issue Description
No public GitHub issue exists for this bug. Bug description follows the
bug-report template:
**What happened:** Accepted-plan decomposition pre-binds child issues to
the parent's realized execution workspace branch (`executionWorkspaceId`
+ `executionWorkspaceBranch`). When `origin/master` advances between
plan acceptance and the child's first heartbeat, the workspace branch
interlock fires and the child is permanently frozen before it can start.
**Expected behavior:** Accepted-plan children should receive only
unresolved workspace intent (mode, git strategy fields) and realize a
fresh isolated worktree from the current base on first heartbeat. A
base-ref advance between acceptance and first-run should be transparent.
**Steps to reproduce:**
1. Accept a plan that decomposes into one or more child issues
(isolated_workspace + git_worktree mode).
2. Allow `origin/master` to advance (new merge).
3. Observe the first child heartbeat: workspace validation fails with a
branch-freeze error.
**Paperclip version:** current `master` (pre-fix).
**Deployment mode:** any (affects all modes that use isolated workspace
+ git worktree strategy).
Supersedes #9227 (earlier attempt, now closed — the fix was incomplete
because `environmentId` was silently dropped during
`parseIssueExecutionWorkspaceSettings` update round-trips, causing the
child workspace to lose its environment binding; this PR includes that
fix).
## What Changed
- **`server/src/issues.ts` — `createChild` / accepted-plan decomposition
path:** strip resolved workspace fields (`executionWorkspaceId`,
concrete branch) when creating accepted-plan children; preserve only
unresolved intent fields (`mode`, `baseRef`, `branchTemplate`,
`environmentId`, runtime/provisioning settings).
- **`server/src/execution-workspace-policy.ts` —
`parseIssueExecutionWorkspaceSettings`:** preserve `environmentId`
through update round-trips (was silently dropped, causing environment to
detach on any workspace settings update).
-
**`server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts`:**
new regression test — accepted-plan child created after `origin/master`
moves realizes a fresh isolated worktree from the moved base and passes
workspace execution.
- **`server/src/__tests__/issues-service.test.ts`:** extended
workspace-linkage and `createChild` tests covering the accepted-plan
strip and the unchanged direct-child path.
## Verification
```sh
pnpm exec vitest run server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts
pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t "workspace linkage|accepted plan decomposition|createChild applies"
pnpm --filter @paperclipai/server typecheck
```
All three pass on this branch.
## Risks
**Low.** The change is scoped to the accepted-plan `createChild` code
path. The direct child / follow-up issue creation path (normal non-plan
decomposition) is unchanged and covered by existing tests. The
`parseIssueExecutionWorkspaceSettings` fix is additive — it now
preserves a field that was previously silently dropped, so no consumer
loses data.
## Model Used
Claude Sonnet 4.6 (`claude-sonnet-4-6`), Anthropic, 200K context window,
extended tool use + code generation.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (supersedes #9227)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.708.0-canary.7
|
||
|
|
88bf71b84e |
Reject projectless isolated git worktree tasks (#9231)
## Thinking Path
> - Paperclip is an open-source platform for orchestrating AI agents;
agents run inside execution workspaces that range from a shared
container to full git worktrees cloned from a project repository.
> - Isolated git-worktree workspaces require a project to determine
which repository to clone — without a project the worktree base path
cannot be computed.
> - A task pinned to `isolated_workspace` + `git_worktree` with no
project was previously accepted at creation time but failed late at
dispatch with the opaque `workspace_validation_failed` /
`git_worktree_base_agent_home` error — only after the heartbeat
attempted to provision the workspace.
> - Fail-closed validation should happen in two places: (1) explicit
create/update pins that contradict the requirement are rejected at the
HTTP layer with a structured 422; (2) rows that reach the heartbeat
dispatcher with this invalid combination (e.g. through inheritance or a
retroactively-removed project) are blocked before any heartbeat run or
adapter spawn.
> - This PR adds the shared detection policy, the create/update guard in
the issues service, and the heartbeat pre-dispatch guard, together with
focused unit tests for all three layers.
> - The benefit is deterministic early failure with a clear remediation
message instead of a late, cryptic runtime error.
## Linked Issues or Issue Description
No upstream public GitHub issue — describing the problem inline
(bug-report format).
**What happened?**
Creating an issue with `executionWorkspaceSettings: { mode:
"isolated_workspace", type: "git_worktree" }` and no `projectId` was
accepted without error. The issue then became blocked at dispatch time
with the opaque message `git_worktree_base_agent_home` /
`workspace_validation_failed` — surfaced only after the heartbeat
attempted to provision the workspace.
**Expected behavior**
The platform should reject the invalid combination at create/update time
with a structured 422 that includes a clear remediation message, before
any heartbeat resource is consumed.
**Steps to reproduce**
1. Call `POST /api/issues` (or `PATCH /api/issues/:id`) with
`executionWorkspaceSettings: { mode: "isolated_workspace", type:
"git_worktree" }` and omit `projectId` (or set it to `null`).
2. Observe: request succeeds (200/201).
3. Assign the issue to an agent and watch it enter `blocked` with a
cryptic `workspace_validation_failed` error at dispatch.
**Related prior fix** — Refs #4844 (`fix(validator): reject static cwd
combined with git_worktree strategy`) — same validation area, different
dimension (static cwd vs. missing project).
**Paperclip version / commit**
Latest `master` (pre-this-PR).
**Deployment mode**
Standard (app-global server).
## What Changed
- **`execution-workspace-policy.ts`** — new shared
`detectWorkspaceWorktreeRequiresProject` function returning a stable
`workspace_worktree_requires_project` policy violation when an isolated
git-worktree task has no project, project workspace, or reusable
execution workspace; exports canonical remediation text used by both the
HTTP guard and the heartbeat guard.
- **`issues.ts`** — create and update paths check the new policy before
persisting; explicit pins to `isolated_workspace` / `operator_branch` +
`git_worktree` with no project are rejected with a 422 including the
policy code and remediation text.
- **`heartbeat.ts`** — pre-dispatch preflight checks the same policy for
rows that reach the heartbeat with the invalid combination (e.g. through
inheritance); such rows are marked `blocked` with a skipped wakeup
request, durable issue comment, and activity log before any heartbeat
run or adapter spawn.
- **`execution-workspace-policy.test.ts`** — focused policy-layer unit
tests for detection logic and remediation text.
- **`issues-service.test.ts`** — create/update 422 guard tests for the
new policy.
- **`heartbeat-workspace-branch-containment.test.ts`** — pre-dispatch
blocking test for inherited/ambiguous invalid rows; also fixes a cleanup
race in the existing test suite.
## Verification
```sh
pnpm exec vitest run \
server/src/__tests__/execution-workspace-policy.test.ts \
server/src/__tests__/issues-service.test.ts \
server/src/__tests__/heartbeat-workspace-branch-containment.test.ts
pnpm --filter @paperclipai/server typecheck
# Targeted regression
pnpm exec vitest run \
server/src/__tests__/heartbeat-workspace-branch-containment.test.ts \
-t "blocks projectless isolated git-worktree issues before dispatch"
```
All three test files and typecheck passed locally before this PR was
opened.
## Risks
**Low risk.** The policy detection function is pure with no side
effects. The create/update guard only triggers on explicit
`isolated_workspace` or `operator_branch` + `git_worktree` pins combined
with a missing project — it does not fire on inherited settings (handled
by the heartbeat preflight), so there is no false-positive rejection
risk for valid tasks. The heartbeat guard fires before any resource is
provisioned; the only behavioral change for already-invalid rows is that
they receive a clear `blocked` status and durable comment instead of a
late cryptic error.
## Model Used
Claude Sonnet 4.6 (`claude-sonnet-4-6`), 200k context window, tool use
enabled (agentic coding). Used to implement all server-side changes and
tests in this PR.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.708.0-canary.6
|